AX-gemma-4-12b-MLX-AXQ-6bit-MTP

Checkpoint Tier 1 certified on df-macbookpro-m5 (2026-08-09). Rebuilt from google/gemma-4-12b-it (not the non-IT google/gemma-4-12b base). Size + quality retention ≥0.98 on authorizing host. MTP acceleration is not certified. Certificate: gemma4-12b-axq6-tier1.md.

Property Value
Product class AXQ 6bit
Source google/gemma-4-12b-it
Measured BPW 6.0001
Size ratio vs uniform 0.7568× (max 1.1)
Quality agent-coding retention 0.9921
Quality general retention 1.0000
Manifest SHA-256 dd328239bd1eebf5643642849b251d7ff8f4a33aa25815e9a29313d25c04a85d
Engine pair gemma-4-12b-itgemma-4-12b-it-assistant
MTP layout assistant/ + ax_gemma4_assistant_mtp.json

Certification

Claim Status
Checkpoint Tier 1 Certified on df-macbookpro-m5
MTP acceleration Tier 2 Not certified
Vision / multimodal quality Not claimed

Layout

./                          # AXQ target weights (+ vision sidecar)
assistant/                  # gemma4_assistant drafter
ax_gemma4_assistant_mtp.json
ax_composite_pack_manifest.json
axquant_*.json

Runtime (AX Engine 7.1.5)

Download the complete snapshot and serve its local root:

hf download AutomatosX/AX-gemma-4-12b-MLX-AXQ-6bit-MTP --local-dir ./AX-gemma-4-12b-MLX-AXQ-6bit-MTP
ax-engine serve ./AX-gemma-4-12b-MLX-AXQ-6bit-MTP --port 31418

AX Engine validates ax_gemma4_assistant_mtp.json and the exact-paired assistant/ bundle before attaching the drafter. Assistant-MTP is enabled by default in AX Engine 7.1.5 with a maximum draft depth of two. Set AX_MLX_GEMMA4_ASSISTANT_MTP=0 to force direct decode, or set AX_MLX_GEMMA4_ASSISTANT_MTP_MAX_DEPTH=1 to cap drafting at one token.

Stock MLX-LM does not consume the assistant/ bundle. This is not a Qwen sidecar, so the oMLX/MTPLX Qwen import workflow does not apply. The checkpoint card still makes no Tier 2 MTP acceleration claim.

Published / card updated: 2026-08-21.

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

Modality Claim Supported Reason
Vision present-not-certified true vision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-gemma-4-12b-MLX-AXQ-6bit-MTP.json
Audio not-applicable false audio not supported (no tower config and no sidecar weights)
Downloads last month
567
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collections including AutomatosX/AX-gemma-4-12b-MLX-AXQ-6bit-MTP