R2_TQ3_4S — DeepSeek-V4-Flash-0731

Selective-imatrix TQ3_4S quant of DeepSeek-V4-Flash-0731 (284B MoE / 13B active), tuned per-tensor for coding quality at minimal size. 17% smaller than the previous TQ3_4S release and serves full 1M context on a single box.

Required Runtime

This model uses the custom TQ3_4S tensor type. Stock llama.cpp builds cannot load it — use the TurboQuant fork:

  • Repo: github.com/turbo-tan/llama.cpp-tq3 (branch main, deepseek4 support; tested build v10413)
  • This is a standard model — it does not contain an MTP draft block, so no --spec-type draft-mtp flags apply.

Versions

Variant Size bpw Context served Download
R2_TQ3_4S (this) 91 GB 2.73 1,048,576 Files
TQ3_4S (previous) 110 GB ~3.1 524,288 YTan2000/DeepSeek-V4-Flash-0731-TQ3_4S

Multi-file GGUF: download all 9 shards R2.gguf-00001-of-00009.gguf … 00009-of-00009 (11.8/10.9/10.9/9.5/10.2/10.8/11.1/12.0/9.6 GB) into one folder, then load -00001-of-00009:

from huggingface_hub import snapshot_download
snapshot_download(
  repo_id = "YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S",
  local_dir = "R2_TQ3_4S",
  allow_patterns = ["R2.gguf-*", "README.md"],
)

Quick start

Single consumer GPU + system RAM (RTX 3090/3090 Ti + 128 GB DDR4/5)

Full 1M context at 14.3 tok/s — experts run on CPU, attention on GPU:

sudo bash -c 'ulimit -l unlimited; exec ./llama-server \
  -m R2_TQ3_4S/R2.gguf-00001-of-00009.gguf \
  --host 127.0.0.1 --port 8080 \
  -c 1048576 -np 1 \
  -ngl 44 --n-cpu-moe 39 --load-mode mmap+mlock \
  -fa on -ctk q4_0 -ctv tq3_0 \
  --reasoning on --reasoning-budget 256 --reasoning-format deepseek --jinja \
  -t 16 -tb 16 -b 8192 --fit on'
  • ulimit -l unlimited is required — mlock must pin the full 91 GB or performance silently degrades
  • --n-cpu-moe 39 routes all MoE experts to CPU; -ngl 44 keeps attention/dense on GPU
  • compressed KV (-ctk q4_0 -ctv tq3_0) makes 1M context cost only ~14 GB

Single DGX Spark (GB10, 128 GB unified)

llama-server -m R2_TQ3_4S/R2.gguf-00001-of-00009.gguf \
  -ngl 99 -c 1048576 -np 1 --port 8085 \
  -ctk q4_0 -ctv tq3_0 --reasoning-format deepseek --reasoning-budget 256

With the DSpark drafter (--spec-type draft-dspark -md dspark-DeepSeek-V4-Flash-0731-Q8_0.gguf), decode rises to 22.9 tok/s @1M — faster than the previous TQ3_4S at 512K.

Benchmarks

Thinking ON, temp 0, official evalplus scorer.

benchmark

Benchmark R2_TQ3_4S (3090 @1M) R2_TQ3_4S (Spark) TQ3_4S (Spark, 512K)
HumanEval pass@1 93.3 90.9 94.5
HumanEval+ pass@1 89.0 86.6 90.9
MBPP pass@1 92.6 92.6 91.8
MBPP+ pass@1 77.5 75.9 77.2
Hard86 77/86 76/86 70/86
Decode tok/s @1M 14.3 18.4 (22.9 + drafter) 21.4 @512K

Task-level suite breakdown (raw openai_compat):

Task R2_TQ3_4S (3090 @1M)
coding 91.7
toolcall 90.0
dataextract 87.1
reasonmath 80.0
instructfollow 77.8
speed 49.4

Caveat: the 3090 suite run used reasoning budget 81,920 (garden server) rather than the 256 used for the Spark baselines.

Tested hardware

  • 1× RTX 3090 24 GB + Ryzen 5950X + 125 GB DDR4-3200 — full 1M, arithmetic gate 6/6
  • 1× DGX Spark (GB10) — full 1M, arithmetic gate 6/6, with and without drafter

Recipe (what "TQ3_4S" means)

Base: deepseek-v4-flash-0731 UD-Q8_K_XL, imatrix-guided requantization:

  • ffn_gate_exps / ffn_up_exps → IQ2_S (routed experts carry the coding quality)
  • ffn_down_exps → q2_K/q3_K
  • attention + dense layers → q4_K/q6_K
  • attention weights and embeddings → q6_K

Full recipe, tensor-type file and validation logs: github.com/turbo-tan/recipes

License

Base model license applies: deepseek-ai/DeepSeek-V4-Flash. Runtime: llama.cpp-tq3 (MIT).

Downloads last month
101
GGUF
Model size
284B params
Architecture
deepseek4
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S

Quantized
(127)
this model

Collection including YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S