TPAI-MiMo-V2.6-Qwen-9B-2-GGUF

TPAI-MiMo-V2.6-Qwen-9B-2 is a coding-focused fine-tune of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B.

This release is intended primarily for local coding assistance, implementation, debugging, repair, and evidence-based code review.

The fine-tune was designed to improve:

  • implementation correctness
  • debugging and repair behavior
  • conservative code review
  • false-positive resistance
  • severity calibration
  • instruction following
  • structured-output compliance
  • concise technical responses
  • distinguishing actual defects from style preferences
  • recognizing when visible evidence does not support a claimed defect

This repository contains the GGUF release of the model.


Model Details

Property Value
Model TPAI-MiMo-V2.6-Qwen-9B-2
Base model XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Base architecture Qwen3.5 / MiMo V2.6
Base parameter count approximately 9.4B
Fine-tuning method LoRA supervised fine-tuning
Training framework Unsloth
Primary specialization Coding
Release format GGUF
Quantization Q8
Natural-language training language Primarily English
Upstream license MIT

The upstream model is multimodal. This fine-tune was trained and evaluated primarily for text-based coding behavior.

Multimodal capability has not been independently benchmarked for this fine-tune.

This model is not affiliated with or endorsed by XiaomiMiMo or Qwen.


Files

The GGUF release contains:

TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf
mmproj-TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf

The main GGUF contains the language model.

The mmproj file provides the corresponding multimodal projection component for compatible llama.cpp-based runtimes.

For text-only coding use, the main GGUF is the important file.

For multimodal use, both files may be required depending on the inference backend.


Training

Training Data

The fine-tune used two curated JSONL training artifacts.

Dataset artifact Rows
TPAI merged coding corpus 3,588
TPAI additional training corpus 1,647
Total training rows 5,235
Unique instruction/input/output examples 5,233

The 3,588-row artifact contains two exact duplicate instruction/input/output examples.

No exact instruction/input/output overlap was found between the two uploaded training artifacts.

The training schema is based on:

{
  "instruction": "...",
  "input": "...",
  "output": "..."
}

A substantial subset additionally contains machine-readable provenance, validation, review, and source metadata.

Training focus

The combined corpus contains coding material emphasizing areas including:

  • Python programming
  • implementation
  • debugging
  • bug repair
  • code review
  • optimization
  • refactoring
  • security-oriented review
  • counterexample reasoning
  • instruction following
  • API and library usage
  • error handling
  • data structures
  • concurrency concepts
  • algorithmic correctness
  • defensive programming
  • multi-language coding tasks

The corpus intentionally contains both direct implementation material and review/correction-oriented examples.


Curated Metadata Subset

Within the 3,588-row training artifact, 2,381 rows contain detailed machine-readable review metadata.

All 2,381 of those metadata-bearing rows are marked:

verified: true
keep_for_training: true

Their recorded review-quality distribution is:

Review quality Records
High 2,071
Medium 310

Their recorded review-type distribution is:

Review type Records
Bug fix 1,291
Optimization 851
Security 146
Style 47
Refactor 46

Their review-scope metadata includes:

  • core reasoning
  • minor or ambiguous issues
  • style or preference
  • maintainability/refactoring

The remaining training rows do not all contain the same machine-readable metadata structure. Their absence of this metadata should not be interpreted as evidence that they were rejected or unreviewed.


Training Data Licensing / Provenance Note

Some source-derived records in the training corpus retain explicit machine-readable content-license metadata.

Within the 2,381 metadata-bearing records:

Recorded source-content license Records
CC BY-SA 3.0 1,560
CC BY-SA 4.0 821

These license fields describe the recorded source content associated with those training examples.

They should not be interpreted to mean that the entire training corpus is MIT licensed.

Not every training record contains equivalent source-license metadata, so this model card does not claim complete per-record licensing provenance for all 5,235 training rows.

The upstream XiaomiMiMo model itself is distributed under the MIT license.

Users redistributing training data or source-derived material should review the applicable source licenses separately from the model-weight license.


Training Configuration

Training was performed using Unsloth with the following configuration.

Core training parameters

Parameter Value
Maximum sequence length 1,024
Epochs 3
Learning rate 0.0001
Per-device batch size 2
Gradient accumulation 8
Nominal effective batch size 16
Warmup steps 5
Maximum steps 0
Save interval 30 steps
Evaluation interval/config value 0.01
Weight decay 0.01
Random seed 6977420
Packing Disabled
Train on completions only Enabled
Gradient checkpointing Unsloth
Optimizer paged_adamw_8bit
LR scheduler cosine

max_steps: 0 leaves training governed by the configured epoch count rather than imposing a separate fixed maximum-step limit.

The nominal effective batch size shown above is:

batch_size ร— gradient_accumulation_steps
2 ร— 8 = 16

for a single-worker training run.


LoRA Configuration

Parameter Value
LoRA rank 16
LoRA alpha 32
LoRA dropout 0
rsLoRA Disabled
LoftQ Disabled
DoRA Disabled

LoRA was applied to:

q_proj
k_proj
v_proj
o_proj
gate_proj
up_proj
down_proj

This covers the principal attention projections and MLP projections.


Original Training Configuration

training:
  max_seq_length: 1024
  num_epochs: 3
  learning_rate: 0.0001
  batch_size: 2
  gradient_accumulation_steps: 8
  warmup_steps: 5
  max_steps: 0
  save_steps: 30
  eval_steps: 0.01
  weight_decay: 0.01
  random_seed: 6977420
  packing: false
  train_on_completions: true
  gradient_checkpointing: unsloth
  optim: paged_adamw_8bit
  lr_scheduler_type: cosine

lora:
  lora_r: 16
  lora_alpha: 32
  lora_dropout: 0
  target_modules:
    - q_proj
    - k_proj
    - v_proj
    - o_proj
    - gate_proj
    - up_proj
    - down_proj
  use_rslora: false
  use_loftq: false
  use_dora: false

Evaluation

The model was evaluated against an independently converted Q8 version of its own upstream base model under the same local evaluation conditions.

Benchmark

tpai-qwen38-frozen-v1

Evaluation harness:

EvalPup 0.1.0

The evaluated slice contained:

  • 12 frozen code-review cases
  • 1 frozen implementation case

The same frozen prompts and evaluation harness were used for both models.


Current Frozen Evaluation Slice

Metric Base Q8 TPAI fine-tune Q8
Code-review semantic conclusions 12 / 12 11 / 12
Implementation task Major correctness failure Passed functional checks
Strict output-format compliance Inconsistent Strong
Severity calibration Frequently inflated Substantially improved
Average review latency ~2.70 s ~1.33 s
Review completion tokens ~2,165 ~985

Latency measurements are specific to the local evaluation environment and should not be interpreted as universal inference-performance numbers.


Evaluation Interpretation

The fine-tune showed clear improvements in the evaluated slice in:

  • implementation correctness
  • instruction adherence
  • output-format compliance
  • severity calibration
  • concision
  • local evaluation latency

However, the base model achieved 12/12 correct semantic review conclusions while the fine-tune achieved 11/12.

That regression is reported intentionally.

The fine-tune's observed code-review miss occurred on a clean Kotlin case where the visible implementation already represented failure explicitly, but the model nevertheless produced a defect finding.

This means the current evidence supports a nuanced conclusion:

Under identical local Q8 evaluation conditions, the fine-tune substantially improved implementation correctness, formatting, severity calibration, concision, and speed on this frozen slice, while showing one observed regression in conservative code-review judgment.

This benchmark is deliberately small and should not be interpreted as proof that the fine-tune universally outperforms the upstream model.

Broader evaluation is ongoing.


Implementation Evaluation

The frozen implementation task required construction of an expiring LRU cache with TTL behavior.

The base Q8 response had multiple fundamental correctness defects, including:

  • inconsistent internal data representation
  • incorrect LRU eviction behavior
  • destructive retrieval behavior
  • invalid assumptions about stored entry objects
  • failure to obey the code-only output requirement

The fine-tuned model produced an implementation that passed functional checks covering:

  • TTL expiration
  • LRU eviction
  • updating existing entries
  • None keys
  • None values
  • zero-second TTL behavior
  • code-only output compliance

These checks represent a bounded functional test of one implementation task, not a comprehensive coding benchmark.


Evaluation Philosophy

TaskPuppyAI evaluation separates training material from frozen evaluation cases.

Frozen cases are intended to remain unchanged after evaluation begins so that model behavior cannot influence the benchmark definition retroactively.

For code review, the desired behavior is conservative and evidence-based:

  • report concrete defects
  • do not invent missing requirements
  • do not assume unsupported deployment conditions
  • do not treat style preferences as correctness bugs
  • do not inflate severity
  • recognize intentionally correct behavior
  • recognize already-handled conditions
  • request additional context when necessary
  • allow clean code to remain clean

The objective is not simply to maximize the number of findings.


Intended Use

TPAI-MiMo-V2.6-Qwen-9B-2 is intended primarily for local or offline coding workflows including:

  • code generation
  • implementation from specifications
  • debugging
  • code repair
  • code review
  • refactoring assistance
  • technical reasoning
  • structured coding tasks
  • explanation of code behavior
  • developer-assistance experiments
  • local AI coding tools

It was developed as a practical local coding model, not as a claim of frontier-model equivalence.


Multimodal Use

The upstream MiMo model uses a multimodal architecture and this repository includes the corresponding mmproj GGUF.

However:

This fine-tune was trained and evaluated primarily on text-based coding tasks.

Image understanding and other multimodal capabilities have not yet been independently benchmarked for this fine-tuned model.

The presence of an mmproj file therefore indicates architectural support, not a claim of validated multimodal performance.


Local Inference

The GGUF is intended for llama.cpp-compatible runtimes such as:

  • llama.cpp
  • LM Studio
  • compatible GGUF frontends

For normal text coding use, load:

TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf

For compatible multimodal inference, also provide:

mmproj-TPAI-MiMo-V2.6-Distill-Qwen-9B-2.gguf

Exact loading behavior depends on the inference frontend and llama.cpp version.


Quantization

This repository currently provides a high-quality Q8 GGUF build.

Quantization is treated as a deployment format rather than the source of the model's behavioral changes.

Where possible, evaluation comparisons were therefore performed against a Q8 conversion of the corresponding upstream base model rather than comparing different quantization levels.


Limitations

This model remains experimental.

Known limitations include:

  • the currently reported frozen benchmark is small
  • one observed false-positive code-review regression
  • the training corpus is heavily coding-focused
  • Python is strongly represented in the training material
  • not every training row contains complete machine-readable provenance
  • some source-derived records retain CC BY-SA source-license metadata
  • multimodal behavior has not yet been independently benchmarked
  • behavior can vary substantially by prompt format and inference backend
  • benchmark success does not guarantee correct code
  • generated fixes may introduce new bugs
  • code-review findings may still contain false positives or missed defects
  • 1024-token training sequences do not by themselves demonstrate equivalent training exposure across the model's full supported inference context

Generated code should be reviewed and tested before production or high-consequence use.


Reproducibility

Training artifacts used for this fine-tune:

TPAI-merged-3588-python-heavy-perplexity-qwen-deepseek-glm-debugging-v4.jsonl
TPAI-training-1647.jsonl

SHA-256 fingerprints of the uploaded training snapshots used to prepare this model card:

TPAI-merged-3588...
512a20a79e4490005f2ac222da3cea9a696ad1d1669cbba27ebb66d9e5d7427d

TPAI-training-1647...
c06c7c49a03c2d867e5a44ff7d40502db3550cf5a6355032a68fa5366abf20a1

Training configuration snapshot:

Lower-learn-rate-2-batch-3-epoch.yaml
SHA-256:
4a7b3aa072790a17713550c12d5e6ac6f0cc4fd6284a79d41ee180f12ae499a8

These hashes identify the exact files inspected while preparing this model card.


Project

Developed by TaskPuppyAI.

The broader project focuses on:

  • local coding models
  • high-quality coding datasets
  • code-review calibration
  • false-positive reduction
  • reproducible fine-tuning
  • frozen model evaluation
  • local-first AI development tooling

Evaluation work is performed using TaskPuppyAI's evolving model-evaluation tooling, including EvalPup and related projects.


License

The upstream XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B model is released under the MIT License.

This model release uses the MIT license for the fine-tuned model weights.

Training-data licensing and model-weight licensing are distinct.

A subset of source-derived training records retains machine-readable CC BY-SA 3.0 or CC BY-SA 4.0 source-content metadata. This repository does not claim that those underlying source materials are MIT licensed.

Users should preserve and respect applicable source attribution and licensing requirements when redistributing underlying training data or source-derived content.


Acknowledgements

This work builds on:

  • XiaomiMiMo's MiMo-V2.6-Distill-Qwen-9B
  • Qwen's underlying architecture
  • Unsloth for fine-tuning tooling
  • llama.cpp and the GGUF ecosystem for local inference

Thanks to the open-source model, tooling, and coding communities that make experiments like this possible.

Downloads last month
541
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for TaskPuppyAI/TPAI-MiMo-v2.6-Qwen-9B-2-GGUF

Finetuned
Qwen/Qwen3.5-9B
Adapter
(1)
this model