Instructions to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF # Run inference directly in the terminal: llama cli -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF # Run inference directly in the terminal: llama cli -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF # Run inference directly in the terminal: ./llama-cli -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Use Docker
docker model run hf.co/TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
- LM Studio
- Jan
- vLLM
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
- Ollama
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with Ollama:
ollama run hf.co/TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
- Unsloth Desktop
- Pi
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with Docker Model Runner:
docker model run hf.co/TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
- Lemonade
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Run and chat with the model
lemonade run user.TPAI-MiMo-v2.6-Qwen-9B-2-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
TPAI-MiMo-V2.6-Qwen-9B-2-GGUF
TPAI-MiMo-V2.6-Qwen-9B-2 is a coding-focused fine-tune of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B.
This release is intended primarily for local coding assistance, implementation, debugging, repair, and evidence-based code review.
The fine-tune was designed to improve:
- implementation correctness
- debugging and repair behavior
- conservative code review
- false-positive resistance
- severity calibration
- instruction following
- structured-output compliance
- concise technical responses
- distinguishing actual defects from style preferences
- recognizing when visible evidence does not support a claimed defect
This repository contains the GGUF release of the model.
Model Details
| Property | Value |
|---|---|
| Model | TPAI-MiMo-V2.6-Qwen-9B-2 |
| Base model | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Base architecture | Qwen3.5 / MiMo V2.6 |
| Base parameter count | approximately 9.4B |
| Fine-tuning method | LoRA supervised fine-tuning |
| Training framework | Unsloth |
| Primary specialization | Coding |
| Release format | GGUF |
| Quantization | Q8 |
| Natural-language training language | Primarily English |
| Upstream license | MIT |
The upstream model is multimodal. This fine-tune was trained and evaluated primarily for text-based coding behavior.
Multimodal capability has not been independently benchmarked for this fine-tune.
This model is not affiliated with or endorsed by XiaomiMiMo or Qwen.
Files
The GGUF release contains:
TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf
mmproj-TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf
The main GGUF contains the language model.
The mmproj file provides the corresponding multimodal projection component
for compatible llama.cpp-based runtimes.
For text-only coding use, the main GGUF is the important file.
For multimodal use, both files may be required depending on the inference backend.
Training
Training Data
The fine-tune used two curated JSONL training artifacts.
| Dataset artifact | Rows |
|---|---|
| TPAI merged coding corpus | 3,588 |
| TPAI additional training corpus | 1,647 |
| Total training rows | 5,235 |
| Unique instruction/input/output examples | 5,233 |
The 3,588-row artifact contains two exact duplicate instruction/input/output examples.
No exact instruction/input/output overlap was found between the two uploaded training artifacts.
The training schema is based on:
{
"instruction": "...",
"input": "...",
"output": "..."
}
A substantial subset additionally contains machine-readable provenance, validation, review, and source metadata.
Training focus
The combined corpus contains coding material emphasizing areas including:
- Python programming
- implementation
- debugging
- bug repair
- code review
- optimization
- refactoring
- security-oriented review
- counterexample reasoning
- instruction following
- API and library usage
- error handling
- data structures
- concurrency concepts
- algorithmic correctness
- defensive programming
- multi-language coding tasks
The corpus intentionally contains both direct implementation material and review/correction-oriented examples.
Curated Metadata Subset
Within the 3,588-row training artifact, 2,381 rows contain detailed machine-readable review metadata.
All 2,381 of those metadata-bearing rows are marked:
verified: true
keep_for_training: true
Their recorded review-quality distribution is:
| Review quality | Records |
|---|---|
| High | 2,071 |
| Medium | 310 |
Their recorded review-type distribution is:
| Review type | Records |
|---|---|
| Bug fix | 1,291 |
| Optimization | 851 |
| Security | 146 |
| Style | 47 |
| Refactor | 46 |
Their review-scope metadata includes:
- core reasoning
- minor or ambiguous issues
- style or preference
- maintainability/refactoring
The remaining training rows do not all contain the same machine-readable metadata structure. Their absence of this metadata should not be interpreted as evidence that they were rejected or unreviewed.
Training Data Licensing / Provenance Note
Some source-derived records in the training corpus retain explicit machine-readable content-license metadata.
Within the 2,381 metadata-bearing records:
| Recorded source-content license | Records |
|---|---|
| CC BY-SA 3.0 | 1,560 |
| CC BY-SA 4.0 | 821 |
These license fields describe the recorded source content associated with those training examples.
They should not be interpreted to mean that the entire training corpus is MIT licensed.
Not every training record contains equivalent source-license metadata, so this model card does not claim complete per-record licensing provenance for all 5,235 training rows.
The upstream XiaomiMiMo model itself is distributed under the MIT license.
Users redistributing training data or source-derived material should review the applicable source licenses separately from the model-weight license.
Training Configuration
Training was performed using Unsloth with the following configuration.
Core training parameters
| Parameter | Value |
|---|---|
| Maximum sequence length | 1,024 |
| Epochs | 3 |
| Learning rate | 0.0001 |
| Per-device batch size | 2 |
| Gradient accumulation | 8 |
| Nominal effective batch size | 16 |
| Warmup steps | 5 |
| Maximum steps | 0 |
| Save interval | 30 steps |
| Evaluation interval/config value | 0.01 |
| Weight decay | 0.01 |
| Random seed | 6977420 |
| Packing | Disabled |
| Train on completions only | Enabled |
| Gradient checkpointing | Unsloth |
| Optimizer | paged_adamw_8bit |
| LR scheduler | cosine |
max_steps: 0 leaves training governed by the configured epoch count rather
than imposing a separate fixed maximum-step limit.
The nominal effective batch size shown above is:
batch_size ร gradient_accumulation_steps
2 ร 8 = 16
for a single-worker training run.
LoRA Configuration
| Parameter | Value |
|---|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0 |
| rsLoRA | Disabled |
| LoftQ | Disabled |
| DoRA | Disabled |
LoRA was applied to:
q_proj
k_proj
v_proj
o_proj
gate_proj
up_proj
down_proj
This covers the principal attention projections and MLP projections.
Original Training Configuration
training:
max_seq_length: 1024
num_epochs: 3
learning_rate: 0.0001
batch_size: 2
gradient_accumulation_steps: 8
warmup_steps: 5
max_steps: 0
save_steps: 30
eval_steps: 0.01
weight_decay: 0.01
random_seed: 6977420
packing: false
train_on_completions: true
gradient_checkpointing: unsloth
optim: paged_adamw_8bit
lr_scheduler_type: cosine
lora:
lora_r: 16
lora_alpha: 32
lora_dropout: 0
target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
- gate_proj
- up_proj
- down_proj
use_rslora: false
use_loftq: false
use_dora: false
Evaluation
The model was evaluated against an independently converted Q8 version of its own upstream base model under the same local evaluation conditions.
Benchmark
tpai-qwen38-frozen-v1
Evaluation harness:
EvalPup 0.1.0
The evaluated slice contained:
- 12 frozen code-review cases
- 1 frozen implementation case
The same frozen prompts and evaluation harness were used for both models.
Current Frozen Evaluation Slice
| Metric | Base Q8 | TPAI fine-tune Q8 |
|---|---|---|
| Code-review semantic conclusions | 12 / 12 | 11 / 12 |
| Implementation task | Major correctness failure | Passed functional checks |
| Strict output-format compliance | Inconsistent | Strong |
| Severity calibration | Frequently inflated | Substantially improved |
| Average review latency | ~2.70 s | ~1.33 s |
| Review completion tokens | ~2,165 | ~985 |
Latency measurements are specific to the local evaluation environment and should not be interpreted as universal inference-performance numbers.
Evaluation Interpretation
The fine-tune showed clear improvements in the evaluated slice in:
- implementation correctness
- instruction adherence
- output-format compliance
- severity calibration
- concision
- local evaluation latency
However, the base model achieved 12/12 correct semantic review conclusions while the fine-tune achieved 11/12.
That regression is reported intentionally.
The fine-tune's observed code-review miss occurred on a clean Kotlin case where the visible implementation already represented failure explicitly, but the model nevertheless produced a defect finding.
This means the current evidence supports a nuanced conclusion:
Under identical local Q8 evaluation conditions, the fine-tune substantially improved implementation correctness, formatting, severity calibration, concision, and speed on this frozen slice, while showing one observed regression in conservative code-review judgment.
This benchmark is deliberately small and should not be interpreted as proof that the fine-tune universally outperforms the upstream model.
Broader evaluation is ongoing.
Implementation Evaluation
The frozen implementation task required construction of an expiring LRU cache with TTL behavior.
The base Q8 response had multiple fundamental correctness defects, including:
- inconsistent internal data representation
- incorrect LRU eviction behavior
- destructive retrieval behavior
- invalid assumptions about stored entry objects
- failure to obey the code-only output requirement
The fine-tuned model produced an implementation that passed functional checks covering:
- TTL expiration
- LRU eviction
- updating existing entries
NonekeysNonevalues- zero-second TTL behavior
- code-only output compliance
These checks represent a bounded functional test of one implementation task, not a comprehensive coding benchmark.
Evaluation Philosophy
TaskPuppyAI evaluation separates training material from frozen evaluation cases.
Frozen cases are intended to remain unchanged after evaluation begins so that model behavior cannot influence the benchmark definition retroactively.
For code review, the desired behavior is conservative and evidence-based:
- report concrete defects
- do not invent missing requirements
- do not assume unsupported deployment conditions
- do not treat style preferences as correctness bugs
- do not inflate severity
- recognize intentionally correct behavior
- recognize already-handled conditions
- request additional context when necessary
- allow clean code to remain clean
The objective is not simply to maximize the number of findings.
Intended Use
TPAI-MiMo-V2.6-Qwen-9B-2 is intended primarily for local or offline coding workflows including:
- code generation
- implementation from specifications
- debugging
- code repair
- code review
- refactoring assistance
- technical reasoning
- structured coding tasks
- explanation of code behavior
- developer-assistance experiments
- local AI coding tools
It was developed as a practical local coding model, not as a claim of frontier-model equivalence.
Multimodal Use
The upstream MiMo model uses a multimodal architecture and this repository
includes the corresponding mmproj GGUF.
However:
This fine-tune was trained and evaluated primarily on text-based coding tasks.
Image understanding and other multimodal capabilities have not yet been independently benchmarked for this fine-tuned model.
The presence of an mmproj file therefore indicates architectural support,
not a claim of validated multimodal performance.
Local Inference
The GGUF is intended for llama.cpp-compatible runtimes such as:
- llama.cpp
- LM Studio
- compatible GGUF frontends
For normal text coding use, load:
TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf
For compatible multimodal inference, also provide:
mmproj-TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf
Exact loading behavior depends on the inference frontend and llama.cpp version.
Quantization
This repository currently provides a high-quality Q8 GGUF build.
Quantization is treated as a deployment format rather than the source of the model's behavioral changes.
Where possible, evaluation comparisons were therefore performed against a Q8 conversion of the corresponding upstream base model rather than comparing different quantization levels.
Limitations
This model remains experimental.
Known limitations include:
- the currently reported frozen benchmark is small
- one observed false-positive code-review regression
- the training corpus is heavily coding-focused
- Python is strongly represented in the training material
- not every training row contains complete machine-readable provenance
- some source-derived records retain CC BY-SA source-license metadata
- multimodal behavior has not yet been independently benchmarked
- behavior can vary substantially by prompt format and inference backend
- benchmark success does not guarantee correct code
- generated fixes may introduce new bugs
- code-review findings may still contain false positives or missed defects
- 1024-token training sequences do not by themselves demonstrate equivalent training exposure across the model's full supported inference context
Generated code should be reviewed and tested before production or high-consequence use.
Reproducibility
Training artifacts used for this fine-tune:
TPAI-merged-3588-python-heavy-perplexity-qwen-deepseek-glm-debugging-v4.jsonl
TPAI-training-1647.jsonl
SHA-256 fingerprints of the uploaded training snapshots used to prepare this model card:
TPAI-merged-3588...
512a20a79e4490005f2ac222da3cea9a696ad1d1669cbba27ebb66d9e5d7427d
TPAI-training-1647...
c06c7c49a03c2d867e5a44ff7d40502db3550cf5a6355032a68fa5366abf20a1
Training configuration snapshot:
Lower-learn-rate-2-batch-3-epoch.yaml
SHA-256:
4a7b3aa072790a17713550c12d5e6ac6f0cc4fd6284a79d41ee180f12ae499a8
These hashes identify the exact files inspected while preparing this model card.
Project
Developed by TaskPuppyAI.
The broader project focuses on:
- local coding models
- high-quality coding datasets
- code-review calibration
- false-positive reduction
- reproducible fine-tuning
- frozen model evaluation
- local-first AI development tooling
Evaluation work is performed using TaskPuppyAI's evolving model-evaluation tooling, including EvalPup and related projects.
License
The upstream XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B model is released under the MIT License.
This model release uses the MIT license for the fine-tuned model weights.
Training-data licensing and model-weight licensing are distinct.
A subset of source-derived training records retains machine-readable CC BY-SA 3.0 or CC BY-SA 4.0 source-content metadata. This repository does not claim that those underlying source materials are MIT licensed.
Users should preserve and respect applicable source attribution and licensing requirements when redistributing underlying training data or source-derived content.
Acknowledgements
This work builds on:
- XiaomiMiMo's MiMo-V2.6-Distill-Qwen-9B
- Qwen's underlying architecture
- Unsloth for fine-tuning tooling
- llama.cpp and the GGUF ecosystem for local inference
Thanks to the open-source model, tooling, and coding communities that make experiments like this possible.
- Downloads last month
- 541
We're not able to determine the quantization variants.
Model tree for TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF
Base model
Qwen/Qwen3.5-9B-Base