Instructions to use bloomer010/Ling-3.0-flash-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bloomer010/Ling-3.0-flash-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use bloomer010/Ling-3.0-flash-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bloomer010/Ling-3.0-flash-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bloomer010/Ling-3.0-flash-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
- Ollama
How to use bloomer010/Ling-3.0-flash-GGUF with Ollama:
ollama run hf.co/bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use bloomer010/Ling-3.0-flash-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bloomer010/Ling-3.0-flash-GGUF with Docker Model Runner:
docker model run hf.co/bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
- Lemonade
How to use bloomer010/Ling-3.0-flash-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.Ling-3.0-flash-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use bloomer010/Ling-3.0-flash-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bloomer010/Ling-3.0-flash-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bloomer010/Ling-3.0-flash-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ling-3.0-flash GGUF
GGUF conversions of inclusionAI/Ling-3.0-flash (124B total / 5.1B active, hybrid KDA + gated MLA, 512-expert MoE), converted directly from the released BF16 safetensors.
These are the reference conversions for the bailingmoe3 architecture, merged into llama.cpp in
PR #26608 (2026-08-17). Every file bundles the MTP (NextN)
block and Ling 3.0's trained per-layer SwiGLU clamp metadata, and no separate drafter file, nor fork
required.
🦙🚨 llama.cpp 🦙🚨
Consistent agentic use (tool calling, reasoning split) currently requires two llama.cpp PRs:
- Dedicated Ling parser: #28682 (✅ Merged as of 9/19)
- Invalid UTF-8 handling in the PEG parser: #29161 (✅ Merged as of 9/20)
Without both, tool calls inside an unclosed think block are dropped and some turns fail with a 500.
To run with llama-server:
llama-server -hf bloomer010/Ling-3.0-flash-GGUF:Q4_K_S
Quant Sizing
Generally...
Larger files = More precision.
Smaller files = More compression = More slop and misbehavin'.
Weights and context share your memory, so be sure leave headroom.
| your memory | file | size |
|---|---|---|
| 192 GB+ | UD-Q8_K_XL |
177 GB |
| 128 GB | Q8_0 |
136 GB |
| 96 GB | UD-Q6_K_XL |
116 GB |
| 80 GB (A100/H100) | Q5_K_M |
92 GB |
| 64 GB | Q4_K_M |
78 GB |
| 56 GB | Q4_K_S / MXFP4_MOE¹ |
74 / 70 GB |
| 48 GB | Q3_K_M |
63 GB |
| 32 GB | UD-Q2_K_XL / IQ2_M |
43 / 42 GB |
| 24 GB | IQ1_M (with expert offload, see below) |
30 GB |
¹ MXFP4_MOE runs its native path on MXFP4-capable GPUs (Blackwell RTX 50-series, GB10/DGX
Spark). Elsewhere it falls back to a slower dequant path — prefer Q4_K_S on older hardware.
With less VRAM than the file size, keep the experts on CPU and the rest on GPU, e.g.:
llama-server -hf bloomer010/Ling-3.0-flash-GGUF:IQ1_M \
-ngl 99 -ot "ffn_.*_exps\.weight=CPU" -c 32768
Usage
Recommended sampling from the source model card: temperature 0.6, top_p 0.95, top_k 20.
Thinking mode is on by default; disable per request with
"chat_template_kwargs": {"enable_thinking": false}.
./build/bin/llama-server \
-m Ling-3.0-flash-Q4_K_S.gguf \
-c 262144 \
-ngl auto \
--flash-attn auto \
--temp 0.6 --top-p 0.95 --top-k 20 \
--jinja
MTP Drafting
🚨 Note that dspark was measured to be faster in at least one instance, and is likely faster on most setups. See the next section below.
Every quant bundles the MTP/NextN block. Enable it with --spec-type draft-mtp:
./build/bin/llama-server \
-m Ling-3.0-flash-Q8_0.gguf \
-c 262144 \
-ngl auto \
--flash-attn auto \
--temp 0.6 --top-p 0.95 --top-k 20 \
--jinja \
--spec-type draft-mtp
During ordinary inference, llama.cpp skips the MTP tensors and may report them as unused. With
--spec-type draft-mtp, the same GGUF is opened as an MTP draft model and block 42 is loaded and
executed. No separate drafter file is required.
Speculative Decoding (DSpark)
Ling-3.0-flash-dspark-BF16.gguf (1.9 GB) and Ling-3.0-flash-dspark-Q4_K_M.gguf (665 MB) are speculative decoding drafts converted from inclusionAI/Ling-3.0-flash-dspark. Tested with llama.cpp as:
llama-server -m Ling-3.0-flash-Q5_K_M.gguf -md Ling-3.0-flash-dspark-Q4_K_M.gguf --spec-type draft-dspark --spec-draft-n-max 8
- Measured against the Q5_K_M main model: draft acceptance 0.29 (BF16) and 0.26 (Q4_K_M), mean accepted length 3.28 and 3.09 tokens per step. The Q4_K_M draft is within about 6 percent of BF16 at a third of the size. Recent llama.cpp builds can fetch the draft automatically; with a quantized main model the auto-pick is the Q4_K_M draft.
- Also measured against the bundled NextN head on the same target (acceptance 0.18, mean 2.43 tokens per step), the DSpark drafts accept substantially more: 0.29/3.28 at BF16 and 0.26/3.09 at Q4_K_M.
Additional MoE Information
MoE placement can be adjusted for available VRAM with -ncmoe N. Draft-model placement can be controlled separately with -ncmoed N and -ngld N.
Supports up to 256K context.
Conversion and Quantization
Taken directly from the released inclusionAI/Ling-3.0-flash BF16 safetensors.
Conversion-specific tensor transformations include:
A_logstored asexp(A_log)- MLA
kv_b_projsplit into separate K and V tensors, with the K tensor transposed - KDA convolution weights reshaped for llama.cpp
- Per-expert tensors stacked into GGUF expert tensors
- KDA and MLA
g_projtensors mapped separately
Norms, routing tensors, expert routing bias, KDA state scalars, dt_bias, and convolution weights
remain F32.
Importance Matrix
Importance matrix generated from the Q8_0 model:
wiki.train.raw- 100 chunks
- 512 tokens per chunk
- 51,200 calibration tokens total
- 573 matrix entries
Quants
MXFP4_MOE:
- Quantized using llama.cpp's MXFP4_MOE quantization type (4.25 bpw)
Q8_0:
- 8.51 BPW
- 126.3 GiB
- Includes MTP block
UD-Q2_K_XL:
- Model-specific Unsloth-style mixed tensor recipe
- Main expert gate/up tensors: IQ2_XS
- Main expert down tensors: IQ3_XXS
- Final target layer experts: IQ3_XXS and IQ4_XS
- Attention, shared experts, and KDA projections retained at higher precision
- MTP experts: Q3_K and Q4_K
IQ1_S:
- Expected size: approximately 24.9 GiB
- Preserves MTP functionality
Notes
The GGUF contains 43 blocks:
- 42 target-model layers
- 35 KDA layers
- 7 gated MLA layers at zero-based indices 5, 11, 17, 23, 29, 35, and 41
- One MTP/NextN block at index 42
The first two target layers use dense FFNs. The remaining target layers use 512 routed experts with top-8 selection plus one shared expert. Routing uses sigmoid scoring, expert bias, eight expert groups, and four selected groups.
The KDA safe gate is implemented as:
lower_bound * sigmoid(exp(A_log) * (f_proj(x) + dt_bias))
The lower bound is -5.0. The GGUF stores the positive exp(A_log) value, while the sign is
supplied by the negative lower bound.
Validation Completed
- BF16 architecture load and tensor round-trip
- CPU and CUDA execution on a reduced-size BailingMoE3 fixture
- Target next-token parity against the released Hugging Face implementation before the missing trained clamps were identified
- Nonzero SwiGLU clamp execution and GGUF round-trip on the reduced-size BailingMoE3 fixture
- First three recursive MTP proposals matched the Hugging Face implementation
- Full MXFP4_MOE target and MTP graph smoke test
- Q8_0 conversion completed successfully with all 938 tensors
Build
git clone https://github.com/ggml-org/llama.cpp.git # bailingmoe3 merged 2026-08-17
# pre-merge builds:
# git clone --branch bailingmoe3-support https://github.com/aetherbird/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j --target llama-cli llama-server
- Downloads last month
- 17,093
1-bit
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for bloomer010/Ling-3.0-flash-GGUF
Base model
inclusionAI/Ling-3.0-flash