This is an experimental REAP.

Ling-3.0-flash REAP288 (73B total / 5.1B active) - GGUF

[288 of 512 routed experts kept per layer - 44% of experts pruned] from inclusionAI/Ling-3.0-flash (124B total / 5.1B active).

Method: one-shot REAP (Router-weighted Expert Activation Pruning) - experts scored by router-gate-value × output-L2-norm over calibration data, lowest-scoring deleted. No fine-tuning, no recovery training.

Calibration: 1M tokens, 50/25/25 ultrachat / wikitext / code

🎉 bailingmoe3 is supported in stock llama.cpp since PR #26608 (merged 2026-08-17, commit 3733366720). Any build from that commit onward loads these files directly.

🔔 2026-08-21: added reasoning_effort support (low = thinking off, high = on, default same). If you want reasoning_effort," re-download or override with chat_template.jinja.

Serving with experts in CPU RAM (attention on GPU, experts streamed from RAM):

llama-server -m Ling-3.0-flash-REAP288-71B-A5B-MXFP4.gguf   -ngl 99 -ot "ffn_.*_exps\.weight=CPU" --no-mmap -c 65536 --flash-attn on --jinja

Quants in this repo are all cut from a full-precision master: MXFP4, Q4_K_M, Q3_K_M, Q2_K

  • MXFP4 (experts MXFP4 / rest Q8_0) is the pick for CPU-offload serving.
Downloads last month
1,023
GGUF
Model size
73B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bloomer010/Ling-3.0-flash-REAP288-73B-A5B-GGUF

Quantized
(46)
this model

Paper for bloomer010/Ling-3.0-flash-REAP288-73B-A5B-GGUF