Instructions to use bnjlebron/chipotle-support-qwen3.5-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bnjlebron/chipotle-support-qwen3.5-0.8b # Run inference directly in the terminal: llama cli -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bnjlebron/chipotle-support-qwen3.5-0.8b # Run inference directly in the terminal: llama cli -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bnjlebron/chipotle-support-qwen3.5-0.8b # Run inference directly in the terminal: ./llama-cli -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bnjlebron/chipotle-support-qwen3.5-0.8b # Run inference directly in the terminal: ./build/bin/llama-cli -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Use Docker
docker model run hf.co/bnjlebron/chipotle-support-qwen3.5-0.8b
- LM Studio
- Jan
- vLLM
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bnjlebron/chipotle-support-qwen3.5-0.8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bnjlebron/chipotle-support-qwen3.5-0.8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bnjlebron/chipotle-support-qwen3.5-0.8b
- Ollama
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with Ollama:
ollama run hf.co/bnjlebron/chipotle-support-qwen3.5-0.8b
- Unsloth Desktop
- Pi
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bnjlebron/chipotle-support-qwen3.5-0.8b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with Docker Model Runner:
docker model run hf.co/bnjlebron/chipotle-support-qwen3.5-0.8b
- Lemonade
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bnjlebron/chipotle-support-qwen3.5-0.8b
Run and chat with the model
lemonade run user.chipotle-support-qwen3.5-0.8b-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bnjlebron/chipotle-support-qwen3.5-0.8b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bnjlebron/chipotle-support-qwen3.5-0.8b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bnjlebron/chipotle-support-qwen3.5-0.8b
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bnjlebron/chipotle-support-qwen3.5-0.8b" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Chipotle Support Qwen3.5 0.8B (Q8_0)
A Q8_0 quantized GGUF of Qwen3.5-0.8B fine-tuned on 685 distilled question/answer pairs harvested from the Chipotle support chatbot (Pepper).
This is a small persona model. It was trained as a personal project to see how far a 0.8B model could go with ~1M tokens of support-bot Q&A, adversarial red-teaming, and instruction-format training.
Model details
| Property | Value |
|---|---|
| Base model | Qwen3.5-0.8B |
| Architecture | qwen35 |
| File type | Q8_0 (8-bit quantization, ~1.4 GB) |
| Tensor count | 320 |
| Context length | 2048 (though it can be used with up to 256k because neither i nor my deepseek agent figured out how to pull down the context length) |
| Chat template | Qwen chat template (system / user / assistant) |
Training data
The training set was built from:
- ~685 cleaned Q&A pairs distilled from the Chipotle support chatbot
- Adversarial red-team rephrasings (edge cases, out-of-scope deflection)
- Identity / persona pairs
Known junk patterns (self-referential Q&A, accordion artifacts) were programmatically filtered out of the raw harvest.
Behavior
The model answers in the persona of a fast-food support chatbot. It:
- Handles common questions (menu, hours, orders, rewards) with ground-truth answers from the harvest
- Deflects out-of-scope questions with canned boundary replies
- Has been red-teamed against over-claiming and adversarial rephrasings
- May still hallucinate or get details wrong — it's a 0.8B, treat it accordingly
Usage
llama.cpp
llama-cli -m chipotle-support-qwen3.5-0.8b.gguf \
-p "system: You are Chipotle support. Respond helpfully and briefly.\nuser: Do you have vegan options?\nassistant:" \
-n 256000
LM Studio
Drop the GGUF in your models folder and load it — the Qwen chat template is baked into the file. it thinks the quant is nonexistent tho lmao
Disclaimer
This is an unofficial, fan-made project. It is not affiliated with, endorsed by, or connected to Chipotle Mexican Grill or its support systems. The model's answers are generated from a small distilled dataset and should not be treated as accurate, official, or current information about Chipotle.
License
The base model (Qwen3.5-0.8B) is licensed under Apache 2.0 (Qwen license). This fine-tune is provided for research / fun purposes. Training data was harvested from a public chatbot and cleaned; check your local terms before redistributing commercially.
- Downloads last month
- 52
We're not able to determine the quantization variants.