Instructions to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S # Run inference directly in the terminal: llama cli -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S # Run inference directly in the terminal: llama cli -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S # Run inference directly in the terminal: ./llama-cli -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S # Run inference directly in the terminal: ./build/bin/llama-cli -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Use Docker
docker model run hf.co/YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
- LM Studio
- Jan
- Ollama
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with Ollama:
ollama run hf.co/YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
- Unsloth Desktop
- Pi
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with Docker Model Runner:
docker model run hf.co/YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
- Lemonade
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Run and chat with the model
lemonade run user.R2-DeepSeek-V4-Flash-0731-TQ3_4S-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
R2_TQ3_4S — DeepSeek-V4-Flash-0731
Selective-imatrix TQ3_4S quant of DeepSeek-V4-Flash-0731 (284B MoE / 13B active), tuned per-tensor for coding quality at minimal size. 17% smaller than the previous TQ3_4S release and serves full 1M context on a single box.
Required Runtime
This model uses the custom TQ3_4S tensor type. Stock llama.cpp builds cannot load it — use the TurboQuant fork:
- Repo: github.com/turbo-tan/llama.cpp-tq3 (branch
main, deepseek4 support; tested build v10413) - This is a standard model — it does not contain an MTP draft block, so no
--spec-type draft-mtpflags apply.
Versions
| Variant | Size | bpw | Context served | Download |
|---|---|---|---|---|
| R2_TQ3_4S (this) | 91 GB | 2.73 | 1,048,576 | Files |
| TQ3_4S (previous) | 110 GB | ~3.1 | 524,288 | YTan2000/DeepSeek-V4-Flash-0731-TQ3_4S |
Multi-file GGUF: download all 9 shards R2.gguf-00001-of-00009.gguf … 00009-of-00009 (11.8/10.9/10.9/9.5/10.2/10.8/11.1/12.0/9.6 GB) into one folder, then load -00001-of-00009:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id = "YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S",
local_dir = "R2_TQ3_4S",
allow_patterns = ["R2.gguf-*", "README.md"],
)
Quick start
Single consumer GPU + system RAM (RTX 3090/3090 Ti + 128 GB DDR4/5)
Full 1M context at 14.3 tok/s — experts run on CPU, attention on GPU:
sudo bash -c 'ulimit -l unlimited; exec ./llama-server \
-m R2_TQ3_4S/R2.gguf-00001-of-00009.gguf \
--host 127.0.0.1 --port 8080 \
-c 1048576 -np 1 \
-ngl 44 --n-cpu-moe 39 --load-mode mmap+mlock \
-fa on -ctk q4_0 -ctv tq3_0 \
--reasoning on --reasoning-budget 256 --reasoning-format deepseek --jinja \
-t 16 -tb 16 -b 8192 --fit on'
ulimit -l unlimitedis required — mlock must pin the full 91 GB or performance silently degrades--n-cpu-moe 39routes all MoE experts to CPU;-ngl 44keeps attention/dense on GPU- compressed KV (
-ctk q4_0 -ctv tq3_0) makes 1M context cost only ~14 GB
Single DGX Spark (GB10, 128 GB unified)
llama-server -m R2_TQ3_4S/R2.gguf-00001-of-00009.gguf \
-ngl 99 -c 1048576 -np 1 --port 8085 \
-ctk q4_0 -ctv tq3_0 --reasoning-format deepseek --reasoning-budget 256
With the DSpark drafter (--spec-type draft-dspark -md dspark-DeepSeek-V4-Flash-0731-Q8_0.gguf), decode rises to 22.9 tok/s @1M — faster than the previous TQ3_4S at 512K.
Benchmarks
Thinking ON, temp 0, official evalplus scorer.
| Benchmark | R2_TQ3_4S (3090 @1M) | R2_TQ3_4S (Spark) | TQ3_4S (Spark, 512K) |
|---|---|---|---|
| HumanEval pass@1 | 93.3 | 90.9 | 94.5 |
| HumanEval+ pass@1 | 89.0 | 86.6 | 90.9 |
| MBPP pass@1 | 92.6 | 92.6 | 91.8 |
| MBPP+ pass@1 | 77.5 | 75.9 | 77.2 |
| Hard86 | 77/86 | 76/86 | 70/86 |
| Decode tok/s @1M | 14.3 | 18.4 (22.9 + drafter) | 21.4 @512K |
Task-level suite breakdown (raw openai_compat):
| Task | R2_TQ3_4S (3090 @1M) |
|---|---|
| coding | 91.7 |
| toolcall | 90.0 |
| dataextract | 87.1 |
| reasonmath | 80.0 |
| instructfollow | 77.8 |
| speed | 49.4 |
Caveat: the 3090 suite run used reasoning budget 81,920 (garden server) rather than the 256 used for the Spark baselines.
Tested hardware
- 1× RTX 3090 24 GB + Ryzen 5950X + 125 GB DDR4-3200 — full 1M, arithmetic gate 6/6
- 1× DGX Spark (GB10) — full 1M, arithmetic gate 6/6, with and without drafter
Recipe (what "TQ3_4S" means)
Base: deepseek-v4-flash-0731 UD-Q8_K_XL, imatrix-guided requantization:
ffn_gate_exps/ffn_up_exps→ IQ2_S (routed experts carry the coding quality)ffn_down_exps→ q2_K/q3_K- attention + dense layers → q4_K/q6_K
- attention weights and embeddings → q6_K
Full recipe, tensor-type file and validation logs: github.com/turbo-tan/recipes
License
Base model license applies: deepseek-ai/DeepSeek-V4-Flash. Runtime: llama.cpp-tq3 (MIT).
- Downloads last month
- 101
We're not able to determine the quantization variants.
Model tree for YTan2000/R2-DeepSeek-V4-Flash-0731-TQ3_4S
Base model
deepseek-ai/DeepSeek-V4-Flash