AxionML MiMo-V2.6-Distill-Qwen-9B-NVFP4

Developed by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

This is an NVFP4-quantized version of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (9B parameters; a Xiaomi MiMo SFT of Qwen/Qwen3.5-9B for coding, general agents, visual coding and cybersecurity), quantized using NVIDIA Model Optimizer. Weights and activations of the language model's MLP layers are quantized to FP4; the checkpoint is 12 GB vs 19 GB in BF16.

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection strategies such as dual-pass evaluation comparing "map max to 6" versus "map max to 4 with clipping." On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under the MIT License (inherited from the base model).

Quantization Details

NVFP4 (W4A4) is applied to the MLP gate/up/down_proj of all 32 language-model layers. Weight block scales come from an MSE sweep over FP8 scale candidates instead of max-of-abs. Activation block scales are computed at runtime under a calibrated per-tensor global scale. The KV-cache is not quantized.

Kept in BF16: full attention (q/k/v/o_proj), all Gated DeltaNet projections and conv1d, the vision encoder (model.visual*), embed_tokens and lm_head.

Why MLP-only: we also built an all-linear W4A4 variant (attention and Gated DeltaNet projections in NVFP4 too, 8.4 GB). It lost 5.3 points on GPQA-Diamond (46.3 vs 51.6 BF16, avg of 4 runs), while this MLP-only checkpoint stays within noise of BF16 on every benchmark below.

Quantization format NVFP4 W4A4, language-model MLP only, MSE FP8-sweep weight calibration
Weight micro-block / group size 16 (FP8 E4M3 block scales + per-tensor FP32 global scale)
KV-cache BF16 (not quantized)
Calibration dataset nvidia/Nemotron-Post-Training-Dataset-v2 (stem/chat/math/code), 512 samples × 2048 tokens
Quantized checkpoint size 12 GB (vs 19 GB BF16)
Tool NVIDIA Model Optimizer main @ 23355eda9
Target hardware Blackwell (sm_100 / sm_103), native FP4 Tensor Cores

The upstream checkpoint ships no MTP weights (mtp_num_hidden_layers: 1 in the config, but no mtp.* tensors), so none are included here.

Usage

Deploy with SGLang

Works on the stock SGLang release image, no branch or extra installs needed (verified on lmsysorg/sglang:v0.5.20-cu130, 1× B300):

docker run --gpus all --shm-size=32g --network=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:v0.5.20-cu130 \
  sglang serve \
    --model-path AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4 \
    --quantization modelopt_fp4 \
    --reasoning-parser mimo \
    --tool-call-parser mimo \
    --host 0.0.0.0 --port 30000

The server listens on http://0.0.0.0:30000. Sampling defaults come from generation_config.json (temperature=0.6, top_p=0.95, top_k=20). Thinking is on by default; pass chat_template_kwargs={"enable_thinking": false} to turn it off.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4",
    messages=[{"role": "user", "content": "What is 15% of 240?"}],
    max_tokens=2048,
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)

message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")

Tool calls (--tool-call-parser mimo) and image inputs work through the standard OpenAI tools / image_url fields.

Reproduce with ModelOpt

From a Model-Optimizer checkout (main @ 23355eda9, transformers==5.12.1), on one Blackwell GPU. The recipe is ModelOpt's general/ptq/nvfp4_mlp_only_mse-kv_fp8_cast without the FP8 KV-cache, plus explicit vision/MTP exclusions (*mlp* would otherwise also match the vision tower's MLPs). Save it as qwen3_5_nvfp4_mlp_only_mse-kv_none.yaml:

# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
imports:
  base_disable_all: configs/ptq/units/base_disable_all
  default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
  nvfp4: configs/numerics/nvfp4
  nvfp4_static: configs/numerics/nvfp4_static

metadata:
  description: NVFP4 W4A4 on Qwen3.5 dense MLP only; everything else BF16, KV cache unquantized.

quantize:
  algorithm:
    method: mse
    fp8_scale_sweep: true
    layerwise:
      enable: false
  quant_cfg:
    - $import: base_disable_all
    - quantizer_name: '*mlp*weight_quantizer'
      cfg:
        $import: nvfp4_static
    - quantizer_name: '*mlp*input_quantizer'
      cfg:
        $import: nvfp4
    - $import: default_disabled_quantizers
    - quantizer_name: '*visual*'
      enable: false
    - quantizer_name: '*vision_tower*'
      enable: false
    - quantizer_name: '*mtp*'
      enable: false
cd Model-Optimizer/examples/hf_ptq
python hf_ptq.py \
  --pyt_ckpt_path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B \
  --recipe /path/to/qwen3_5_nvfp4_mlp_only_mse-kv_none.yaml \
  --dataset nemotron-post-training-dataset-v2 \
  --calib_size 512 --calib_seq 2048 \
  --export_path ./MiMo-V2.6-Distill-Qwen-9B-NVFP4

Load, calibration and export take about 5 minutes on one B300.

Accuracy

All numbers measured by us on the same stack for both columns: lmsysorg/sglang:v0.5.20-cu130 on 1× B300, sgl-eval 0.1.2, thinking on, the model's default sampling (temperature=0.6, top_p=0.95, top_k=20).

Benchmark Budget BF16 NVFP4 (this repo)
GSM8K (1319, pass@1) 16K 92.27 92.42
GPQA-Diamond (avg of 4) 32K 51.64 51.39
AIME 2025 (avg of 8) 32K 37.92 38.75
MMMU-Pro (standard, 10 options) 16K 49.36 48.09
Tool calls: right function + non-empty args (80 trials) 8K 67/80 67/80

This model reasons at length. At these budgets a large share of generations hit the token limit in both columns (GPQA 37% / 36%, AIME 55% / 56%, MMMU-Pro 24% / 23%, BF16 / NVFP4), so absolute scores are budget-limited and lower than Xiaomi's reported numbers; the BF16-vs-NVFP4 comparison is like for like. The 13/80 tool-call misses are the same failure in both columns (the model emits an empty argument object on one of the four prompts).

Reproduce (server from Deploy with SGLang on port 30000):

pip install sgl-eval==0.1.2
COMMON="--base-url http://localhost:30000/v1 --temperature 0.6 --top-p 0.95 --chat-template-kwarg enable_thinking=true --num-threads 256"
sgl-eval run gsm8k    $COMMON --max-tokens 16384
sgl-eval run gpqa     $COMMON --max-tokens 32768 --n-repeats 4
sgl-eval run aime25   $COMMON --max-tokens 32768 --n-repeats 8
sgl-eval run mmmu_pro $COMMON --max-tokens 16384

Model Overview

Architecture is unchanged from Qwen3.5-9B:

  • Type: Causal Language Model with Vision Encoder
  • Language Model
    • Number of Parameters: 9B
    • Hidden Dimension: 4096
    • Token Embedding: 248320 (Padded)
    • Number of Layers: 32
    • Hidden Layout: 8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 32 for V and 16 for QK
      • Head Dimension: 128
    • Gated Attention:
      • Number of Attention Heads: 16 for Q and 4 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
    • Feed Forward Network:
      • Intermediate Dimension: 12288
    • LM Output: 248320 (Padded)
  • Vision Encoder: 27 layers, hidden 1152, patch 16
  • Context Length: 262,144 tokens natively

Base model

XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data (77.4B total tokens, 27.2B loss-bearing). It covers coding, general-purpose agent tasks, visual coding, and cybersecurity, and is released as a starting point for open research in agentic reinforcement learning. Results reported by Xiaomi for the BF16 checkpoint:

Domain Benchmark Metric Qwen3.5-9B MiMo-V2.6-Distill-Qwen-9B (SFT)
Code SWE Verified avg@3 60.0 61.1
Code SWE Pro avg@3 32.0 44.6
Code MiMo Code (mini)† avg@3 19.5 51.6
Cyber MiMo Cyber (mini)† avg@3 5.7 31.3
General AutomationBench v1.0.6 avg@1 5.0 30.3
General Terminal Bench 2.1 avg@1 27.0 37.1
General Toolathlon-Verified avg@1 25.9 35.2
General OfficeQA avg@1 9.0 19.5
General JobBench avg@1 2.6 18.3
General MiMo General (mini)† avg@1 28.5 62.2
Visual MiMo Visual Coding (mini)† avg@1 61.7 64.0

† Xiaomi internal evaluation sets. These are the base model's numbers and were not re-measured on this checkpoint.

This repository changes only the numeric precision of the weights. The chat template, tokenizer and processor configs are inherited unchanged.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations and may generate inaccurate, biased, or offensive content. Quantization can introduce additional deviations from the base model's behavior. Please refer to the original model card for full details.

Citation

@misc{mimo2026v26,
  title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}
Downloads last month
143
Safetensors
Model size
7B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AxionML/MiMo-V2.6-Distill-Qwen-9B-NVFP4

Finetuned
Qwen/Qwen3.5-9B
Quantized
(57)
this model