Image-Text-to-Text
Transformers
GGUF
text-generation-inference
unsloth
qwen3_vl
trl
sft
chemistry
code
climate
art
biology
finance
legal
music
medical
agent
llama-cpp
gguf-my-repo
Instructions to use Lamapi/next-ocr-Q5_0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Lamapi/next-ocr-Q5_0-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Lamapi/next-ocr-Q5_0-GGUF")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Lamapi/next-ocr-Q5_0-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Lamapi/next-ocr-Q5_0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0 # Run inference directly in the terminal: llama cli -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0 # Run inference directly in the terminal: llama cli -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0 # Run inference directly in the terminal: ./llama-cli -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Lamapi/next-ocr-Q5_0-GGUF:Q5_0
Use Docker
docker model run hf.co/Lamapi/next-ocr-Q5_0-GGUF:Q5_0
- LM Studio
- Jan
- vLLM
How to use Lamapi/next-ocr-Q5_0-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Lamapi/next-ocr-Q5_0-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lamapi/next-ocr-Q5_0-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Lamapi/next-ocr-Q5_0-GGUF:Q5_0
- SGLang
How to use Lamapi/next-ocr-Q5_0-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Lamapi/next-ocr-Q5_0-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lamapi/next-ocr-Q5_0-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Lamapi/next-ocr-Q5_0-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lamapi/next-ocr-Q5_0-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use Lamapi/next-ocr-Q5_0-GGUF with Ollama:
ollama run hf.co/Lamapi/next-ocr-Q5_0-GGUF:Q5_0
- Unsloth Desktop
- Docker Model Runner
How to use Lamapi/next-ocr-Q5_0-GGUF with Docker Model Runner:
docker model run hf.co/Lamapi/next-ocr-Q5_0-GGUF:Q5_0
- Lemonade
How to use Lamapi/next-ocr-Q5_0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Lamapi/next-ocr-Q5_0-GGUF:Q5_0
Run and chat with the model
lemonade run user.next-ocr-Q5_0-GGUF-Q5_0
List all available models
lemonade list
- Atomic Chat
| tags: | |
| - text-generation-inference | |
| - transformers | |
| - unsloth | |
| - qwen3_vl | |
| - trl | |
| - sft | |
| - chemistry | |
| - code | |
| - climate | |
| - art | |
| - biology | |
| - finance | |
| - legal | |
| - music | |
| - medical | |
| - agent | |
| - llama-cpp | |
| - gguf-my-repo | |
| license: apache-2.0 | |
| language: | |
| - en | |
| - ab | |
| - aa | |
| - ae | |
| - af | |
| - ak | |
| - am | |
| - an | |
| - ar | |
| - as | |
| - av | |
| - ay | |
| - az | |
| - ba | |
| - be | |
| - bg | |
| - bh | |
| - bi | |
| - bm | |
| - bn | |
| - bo | |
| - br | |
| - bs | |
| - ca | |
| - ce | |
| - ch | |
| - co | |
| - cr | |
| - cs | |
| - cu | |
| - cv | |
| - cy | |
| - da | |
| - de | |
| - dv | |
| - dz | |
| - ee | |
| - el | |
| - eo | |
| - es | |
| - et | |
| - eu | |
| - fa | |
| - ff | |
| - fi | |
| - fj | |
| - fo | |
| - fr | |
| - fy | |
| - ga | |
| - gd | |
| - gl | |
| - gn | |
| - gv | |
| - ha | |
| - he | |
| - hi | |
| - ho | |
| - gu | |
| - hr | |
| - ht | |
| - hu | |
| - hz | |
| - hy | |
| - id | |
| - ia | |
| - ig | |
| - ie | |
| - ik | |
| - ii | |
| - is | |
| - io | |
| - iu | |
| - it | |
| - jv | |
| - ja | |
| - kg | |
| - ka | |
| - kj | |
| - ki | |
| - kl | |
| - kk | |
| - kn | |
| - km | |
| - kr | |
| - ko | |
| - ku | |
| - ks | |
| - kw | |
| - kv | |
| - la | |
| - ky | |
| - lg | |
| - lb | |
| - ln | |
| - li | |
| - lt | |
| - lo | |
| - lv | |
| - lu | |
| - mg | |
| - mi | |
| - mh | |
| - ml | |
| - mk | |
| - mr | |
| - mn | |
| - mt | |
| - ms | |
| - na | |
| - my | |
| - nd | |
| - nb | |
| - ng | |
| - nl | |
| - ne | |
| - 'no' | |
| - nn | |
| - nv | |
| - nr | |
| - oc | |
| - oj | |
| - om | |
| - ny | |
| - os | |
| - or | |
| - pa | |
| - pi | |
| - pl | |
| - ps | |
| - pt | |
| - rm | |
| - rn | |
| - qu | |
| - ro | |
| - ru | |
| - sn | |
| - rw | |
| - so | |
| - sa | |
| - sc | |
| - sd | |
| pipeline_tag: image-text-to-text | |
| library_name: transformers | |
| base_model: Lamapi/next-ocr | |
| # Lamapi/next-ocr-Q5_0-GGUF | |
| This model was converted to GGUF format from [`Lamapi/next-ocr`](https://huggingface.co/Lamapi/next-ocr) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space. | |
| Refer to the [original model card](https://huggingface.co/Lamapi/next-ocr) for more details on the model. | |
| ## Use with llama.cpp | |
| Install llama.cpp through brew (works on Mac and Linux) | |
| ```bash | |
| brew install llama.cpp | |
| ``` | |
| Invoke the llama.cpp server or the CLI. | |
| ### CLI: | |
| ```bash | |
| llama-cli --hf-repo Lamapi/next-ocr-Q5_0-GGUF --hf-file next-ocr-q5_0.gguf -p "The meaning to life and the universe is" | |
| ``` | |
| ### Server: | |
| ```bash | |
| llama-server --hf-repo Lamapi/next-ocr-Q5_0-GGUF --hf-file next-ocr-q5_0.gguf -c 2048 | |
| ``` | |
| Note: You can also use this checkpoint directly through the [usage steps](https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#usage) listed in the Llama.cpp repo as well. | |
| Step 1: Clone llama.cpp from GitHub. | |
| ``` | |
| git clone https://github.com/ggerganov/llama.cpp | |
| ``` | |
| Step 2: Move into the llama.cpp folder and build it with `LLAMA_CURL=1` flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux). | |
| ``` | |
| cd llama.cpp && LLAMA_CURL=1 make | |
| ``` | |
| Step 3: Run inference through the main binary. | |
| ``` | |
| ./llama-cli --hf-repo Lamapi/next-ocr-Q5_0-GGUF --hf-file next-ocr-q5_0.gguf -p "The meaning to life and the universe is" | |
| ``` | |
| or | |
| ``` | |
| ./llama-server --hf-repo Lamapi/next-ocr-Q5_0-GGUF --hf-file next-ocr-q5_0.gguf -c 2048 | |
| ``` | |