Ling-3.0-flash GGUF

GGUF conversions of inclusionAI/Ling-3.0-flash (124B total / 5.1B active, hybrid KDA + gated MLA, 512-expert MoE), converted directly from the released BF16 safetensors.

These are the reference conversions for the bailingmoe3 architecture, merged into llama.cpp in PR #26608 (2026-08-17). Every file bundles the MTP (NextN) block and Ling 3.0's trained per-layer SwiGLU clamp metadata, and no separate drafter file, nor fork required.

🦙🚨 llama.cpp 🦙🚨

Consistent agentic use (tool calling, reasoning split) currently requires two llama.cpp PRs:

  • Dedicated Ling parser: #28682 (✅ Merged as of 9/19)
  • Invalid UTF-8 handling in the PEG parser: #29161 (✅ Merged as of 9/20)

Without both, tool calls inside an unclosed think block are dropped and some turns fail with a 500.

To run with llama-server:

llama-server -hf bloomer010/Ling-3.0-flash-GGUF:Q4_K_S

Quant Sizing

Generally...
Larger files = More precision.
Smaller files = More compression = More slop and misbehavin'.

Weights and context share your memory, so be sure leave headroom.

your memory file size
192 GB+ UD-Q8_K_XL 177 GB
128 GB Q8_0 136 GB
96 GB UD-Q6_K_XL 116 GB
80 GB (A100/H100) Q5_K_M 92 GB
64 GB Q4_K_M 78 GB
56 GB Q4_K_S / MXFP4_MOE¹ 74 / 70 GB
48 GB Q3_K_M 63 GB
32 GB UD-Q2_K_XL / IQ2_M 43 / 42 GB
24 GB IQ1_M (with expert offload, see below) 30 GB

¹ MXFP4_MOE runs its native path on MXFP4-capable GPUs (Blackwell RTX 50-series, GB10/DGX Spark). Elsewhere it falls back to a slower dequant path — prefer Q4_K_S on older hardware.

With less VRAM than the file size, keep the experts on CPU and the rest on GPU, e.g.:

llama-server -hf bloomer010/Ling-3.0-flash-GGUF:IQ1_M \
  -ngl 99 -ot "ffn_.*_exps\.weight=CPU" -c 32768

Usage

Recommended sampling from the source model card: temperature 0.6, top_p 0.95, top_k 20. Thinking mode is on by default; disable per request with "chat_template_kwargs": {"enable_thinking": false}.

./build/bin/llama-server \
  -m Ling-3.0-flash-Q4_K_S.gguf \
  -c 262144 \
  -ngl auto \
  --flash-attn auto \
  --temp 0.6 --top-p 0.95 --top-k 20 \
  --jinja

MTP Drafting

🚨 Note that dspark was measured to be faster in at least one instance, and is likely faster on most setups. See the next section below.

Every quant bundles the MTP/NextN block. Enable it with --spec-type draft-mtp:

./build/bin/llama-server \
  -m Ling-3.0-flash-Q8_0.gguf \
  -c 262144 \
  -ngl auto \
  --flash-attn auto \
  --temp 0.6 --top-p 0.95 --top-k 20 \
  --jinja \
  --spec-type draft-mtp

During ordinary inference, llama.cpp skips the MTP tensors and may report them as unused. With --spec-type draft-mtp, the same GGUF is opened as an MTP draft model and block 42 is loaded and executed. No separate drafter file is required.

Speculative Decoding (DSpark)

Ling-3.0-flash-dspark-BF16.gguf (1.9 GB) and Ling-3.0-flash-dspark-Q4_K_M.gguf (665 MB) are speculative decoding drafts converted from inclusionAI/Ling-3.0-flash-dspark. Tested with llama.cpp as:

llama-server -m Ling-3.0-flash-Q5_K_M.gguf -md Ling-3.0-flash-dspark-Q4_K_M.gguf --spec-type draft-dspark --spec-draft-n-max 8

  • Measured against the Q5_K_M main model: draft acceptance 0.29 (BF16) and 0.26 (Q4_K_M), mean accepted length 3.28 and 3.09 tokens per step. The Q4_K_M draft is within about 6 percent of BF16 at a third of the size. Recent llama.cpp builds can fetch the draft automatically; with a quantized main model the auto-pick is the Q4_K_M draft.
  • Also measured against the bundled NextN head on the same target (acceptance 0.18, mean 2.43 tokens per step), the DSpark drafts accept substantially more: 0.29/3.28 at BF16 and 0.26/3.09 at Q4_K_M.

Additional MoE Information

MoE placement can be adjusted for available VRAM with -ncmoe N. Draft-model placement can be controlled separately with -ncmoed N and -ngld N.

Supports up to 256K context.

Conversion and Quantization

Taken directly from the released inclusionAI/Ling-3.0-flash BF16 safetensors.

Conversion-specific tensor transformations include:

  • A_log stored as exp(A_log)
  • MLA kv_b_proj split into separate K and V tensors, with the K tensor transposed
  • KDA convolution weights reshaped for llama.cpp
  • Per-expert tensors stacked into GGUF expert tensors
  • KDA and MLA g_proj tensors mapped separately

Norms, routing tensors, expert routing bias, KDA state scalars, dt_bias, and convolution weights remain F32.

Importance Matrix

Importance matrix generated from the Q8_0 model:

  • wiki.train.raw
  • 100 chunks
  • 512 tokens per chunk
  • 51,200 calibration tokens total
  • 573 matrix entries

Quants

MXFP4_MOE:

  • Quantized using llama.cpp's MXFP4_MOE quantization type (4.25 bpw)

Q8_0:

  • 8.51 BPW
  • 126.3 GiB
  • Includes MTP block

UD-Q2_K_XL:

  • Model-specific Unsloth-style mixed tensor recipe
  • Main expert gate/up tensors: IQ2_XS
  • Main expert down tensors: IQ3_XXS
  • Final target layer experts: IQ3_XXS and IQ4_XS
  • Attention, shared experts, and KDA projections retained at higher precision
  • MTP experts: Q3_K and Q4_K

IQ1_S:

  • Expected size: approximately 24.9 GiB
  • Preserves MTP functionality

Notes

The GGUF contains 43 blocks:

  • 42 target-model layers
  • 35 KDA layers
  • 7 gated MLA layers at zero-based indices 5, 11, 17, 23, 29, 35, and 41
  • One MTP/NextN block at index 42

The first two target layers use dense FFNs. The remaining target layers use 512 routed experts with top-8 selection plus one shared expert. Routing uses sigmoid scoring, expert bias, eight expert groups, and four selected groups.

The KDA safe gate is implemented as:

lower_bound * sigmoid(exp(A_log) * (f_proj(x) + dt_bias))

The lower bound is -5.0. The GGUF stores the positive exp(A_log) value, while the sign is supplied by the negative lower bound.

Validation Completed

  • BF16 architecture load and tensor round-trip
  • CPU and CUDA execution on a reduced-size BailingMoE3 fixture
  • Target next-token parity against the released Hugging Face implementation before the missing trained clamps were identified
  • Nonzero SwiGLU clamp execution and GGUF round-trip on the reduced-size BailingMoE3 fixture
  • First three recursive MTP proposals matched the Hugging Face implementation
  • Full MXFP4_MOE target and MTP graph smoke test
  • Q8_0 conversion completed successfully with all 938 tensors

Build

git clone https://github.com/ggml-org/llama.cpp.git   # bailingmoe3 merged 2026-08-17
# pre-merge builds:
# git clone --branch bailingmoe3-support https://github.com/aetherbird/llama.cpp.git

cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j --target llama-cli llama-server

Downloads last month
17,093
GGUF
Model size
127B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bloomer010/Ling-3.0-flash-GGUF

Quantized
(46)
this model