Instructions to use AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4
- SGLang
How to use AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 with Docker Model Runner:
docker model run hf.co/AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4
AxionML MiMo-V2.6-Distill-Qwen-9B-NVFP4
Developed by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
This is an NVFP4-quantized version of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (9B parameters; a Xiaomi MiMo SFT of Qwen/Qwen3.5-9B for coding, general agents, visual coding and cybersecurity), quantized using NVIDIA Model Optimizer. Weights and activations of the language model's MLP layers are quantized to FP4; the checkpoint is 12 GB vs 19 GB in BF16.
About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection strategies such as dual-pass evaluation comparing "map max to 6" versus "map max to 4 with clipping." On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under the MIT License (inherited from the base model).
Quantization Details
NVFP4 (W4A4) is applied to the MLP gate/up/down_proj of all 32 language-model layers. Weight block scales come from an MSE sweep over FP8 scale candidates instead of max-of-abs. Activation block scales are computed at runtime under a calibrated per-tensor global scale. The KV-cache is not quantized.
Kept in BF16: full attention (q/k/v/o_proj), all Gated DeltaNet projections and conv1d, the vision encoder (model.visual*), embed_tokens and lm_head.
Why MLP-only: we also built an all-linear W4A4 variant (attention and Gated DeltaNet projections in NVFP4 too, 8.4 GB). It lost 5.3 points on GPQA-Diamond (46.3 vs 51.6 BF16, avg of 4 runs), while this MLP-only checkpoint stays within noise of BF16 on every benchmark below.
| Quantization format | NVFP4 W4A4, language-model MLP only, MSE FP8-sweep weight calibration |
| Weight micro-block / group size | 16 (FP8 E4M3 block scales + per-tensor FP32 global scale) |
| KV-cache | BF16 (not quantized) |
| Calibration dataset | nvidia/Nemotron-Post-Training-Dataset-v2 (stem/chat/math/code), 512 samples × 2048 tokens |
| Quantized checkpoint size | 12 GB (vs 19 GB BF16) |
| Tool | NVIDIA Model Optimizer main @ 23355eda9 |
| Target hardware | Blackwell (sm_100 / sm_103), native FP4 Tensor Cores |
The upstream checkpoint ships no MTP weights (mtp_num_hidden_layers: 1 in the config, but no mtp.* tensors), so none are included here.
Usage
Deploy with SGLang
Works on the stock SGLang release image, no branch or extra installs needed (verified on lmsysorg/sglang:v0.5.20-cu130, 1× B300):
docker run --gpus all --shm-size=32g --network=host \
-v ~/.cache/huggingface:/root/.cache/huggingface \
lmsysorg/sglang:v0.5.20-cu130 \
sglang serve \
--model-path AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 \
--quantization modelopt_fp4 \
--reasoning-parser mimo \
--tool-call-parser mimo \
--host 0.0.0.0 --port 30000
The server listens on http://0.0.0.0:30000. Sampling defaults come from generation_config.json (temperature=0.6, top_p=0.95, top_k=20). Thinking is on by default; pass chat_template_kwargs={"enable_thinking": false} to turn it off.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4",
messages=[{"role": "user", "content": "What is 15% of 240?"}],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")
Tool calls (--tool-call-parser mimo) and image inputs work through the standard OpenAI tools / image_url fields.
Reproduce with ModelOpt
From a Model-Optimizer checkout (main @ 23355eda9, transformers==5.12.1), on one Blackwell GPU. The recipe is ModelOpt's general/ptq/nvfp4_mlp_only_mse-kv_fp8_cast without the FP8 KV-cache, plus explicit vision/MTP exclusions (*mlp* would otherwise also match the vision tower's MLPs). Save it as qwen3_5_nvfp4_mlp_only_mse-kv_none.yaml:
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
imports:
base_disable_all: configs/ptq/units/base_disable_all
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
nvfp4: configs/numerics/nvfp4
nvfp4_static: configs/numerics/nvfp4_static
metadata:
description: NVFP4 W4A4 on Qwen3.5 dense MLP only; everything else BF16, KV cache unquantized.
quantize:
algorithm:
method: mse
fp8_scale_sweep: true
layerwise:
enable: false
quant_cfg:
- $import: base_disable_all
- quantizer_name: '*mlp*weight_quantizer'
cfg:
$import: nvfp4_static
- quantizer_name: '*mlp*input_quantizer'
cfg:
$import: nvfp4
- $import: default_disabled_quantizers
- quantizer_name: '*visual*'
enable: false
- quantizer_name: '*vision_tower*'
enable: false
- quantizer_name: '*mtp*'
enable: false
cd Model-Optimizer/examples/hf_ptq
python hf_ptq.py \
--pyt_ckpt_path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B \
--recipe /path/to/qwen3_5_nvfp4_mlp_only_mse-kv_none.yaml \
--dataset nemotron-post-training-dataset-v2 \
--calib_size 512 --calib_seq 2048 \
--export_path ./MiMo-V2.6-Distill-Qwen-9B-NVFP4
Load, calibration and export take about 5 minutes on one B300.
Accuracy
All numbers measured by us on the same stack for both columns: lmsysorg/sglang:v0.5.20-cu130 on 1× B300, sgl-eval 0.1.2, thinking on, the model's default sampling (temperature=0.6, top_p=0.95, top_k=20).
| Benchmark | Budget | BF16 | NVFP4 (this repo) |
|---|---|---|---|
| GSM8K (1319, pass@1) | 16K | 92.27 | 92.42 |
| GPQA-Diamond (avg of 4) | 32K | 51.64 | 51.39 |
| AIME 2025 (avg of 8) | 32K | 37.92 | 38.75 |
| MMMU-Pro (standard, 10 options) | 16K | 49.36 | 48.09 |
| Tool calls: right function + non-empty args (80 trials) | 8K | 67/80 | 67/80 |
This model reasons at length. At these budgets a large share of generations hit the token limit in both columns (GPQA 37% / 36%, AIME 55% / 56%, MMMU-Pro 24% / 23%, BF16 / NVFP4), so absolute scores are budget-limited and lower than Xiaomi's reported numbers; the BF16-vs-NVFP4 comparison is like for like. The 13/80 tool-call misses are the same failure in both columns (the model emits an empty argument object on one of the four prompts).
Reproduce (server from Deploy with SGLang on port 30000):
pip install sgl-eval==0.1.2
COMMON="--base-url http://localhost:30000/v1 --temperature 0.6 --top-p 0.95 --chat-template-kwarg enable_thinking=true --num-threads 256"
sgl-eval run gsm8k $COMMON --max-tokens 16384
sgl-eval run gpqa $COMMON --max-tokens 32768 --n-repeats 4
sgl-eval run aime25 $COMMON --max-tokens 32768 --n-repeats 8
sgl-eval run mmmu_pro $COMMON --max-tokens 16384
Model Overview
Architecture is unchanged from Qwen3.5-9B:
- Type: Causal Language Model with Vision Encoder
- Language Model
- Number of Parameters: 9B
- Hidden Dimension: 4096
- Token Embedding: 248320 (Padded)
- Number of Layers: 32
- Hidden Layout: 8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
- Gated DeltaNet:
- Number of Linear Attention Heads: 32 for V and 16 for QK
- Head Dimension: 128
- Gated Attention:
- Number of Attention Heads: 16 for Q and 4 for KV
- Head Dimension: 256
- Rotary Position Embedding Dimension: 64
- Feed Forward Network:
- Intermediate Dimension: 12288
- LM Output: 248320 (Padded)
- Vision Encoder: 27 layers, hidden 1152, patch 16
- Context Length: 262,144 tokens natively
Base model
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data (77.4B total tokens, 27.2B loss-bearing). It covers coding, general-purpose agent tasks, visual coding, and cybersecurity, and is released as a starting point for open research in agentic reinforcement learning. Results reported by Xiaomi for the BF16 checkpoint:
| Domain | Benchmark | Metric | Qwen3.5-9B | MiMo-V2.6-Distill-Qwen-9B (SFT) |
|---|---|---|---|---|
| Code | SWE Verified | avg@3 | 60.0 | 61.1 |
| Code | SWE Pro | avg@3 | 32.0 | 44.6 |
| Code | MiMo Code (mini)† | avg@3 | 19.5 | 51.6 |
| Cyber | MiMo Cyber (mini)† | avg@3 | 5.7 | 31.3 |
| General | AutomationBench v1.0.6 | avg@1 | 5.0 | 30.3 |
| General | Terminal Bench 2.1 | avg@1 | 27.0 | 37.1 |
| General | Toolathlon-Verified | avg@1 | 25.9 | 35.2 |
| General | OfficeQA | avg@1 | 9.0 | 19.5 |
| General | JobBench | avg@1 | 2.6 | 18.3 |
| General | MiMo General (mini)† | avg@1 | 28.5 | 62.2 |
| Visual | MiMo Visual Coding (mini)† | avg@1 | 61.7 | 64.0 |
† Xiaomi internal evaluation sets. These are the base model's numbers and were not re-measured on this checkpoint.
This repository changes only the numeric precision of the weights. The chat template, tokenizer and processor configs are inherited unchanged.
Limitations
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations and may generate inaccurate, biased, or offensive content. Quantization can introduce additional deviations from the base model's behavior. Please refer to the original model card for full details.
Citation
@misc{mimo2026v26,
title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}
- Downloads last month
- 143
Model tree for AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4
Base model
Qwen/Qwen3.5-9B-Base