Instructions to use tirex2001/Qwen3.8-Flash-Next-DACAN with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0 # Run inference directly in the terminal: llama cli -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0 # Run inference directly in the terminal: llama cli -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Use Docker
docker model run hf.co/tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
- LM Studio
- Jan
- Ollama
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with Ollama:
ollama run hf.co/tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
- Unsloth Desktop
- Pi
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with Docker Model Runner:
docker model run hf.co/tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
- Lemonade
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-DACAN-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tirex2001/Qwen3.8-Flash-Next-DACAN with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tirex2001/Qwen3.8-Flash-Next-DACAN:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next for DACAN
Weights of Qwen3.8-Flash-Next packed for DACAN — our engine, a fork of Strata rebuilt for two RTX 2080 Ti (22 GB, NVLink) and two AVX-512 Xeon sockets. DACAN runs this one model family only.
Status (03.10.2026): complete, 27 files (260.6 GB). The dense GGUF was checked against the file the DACAN service runs: three greedy answers (1 984, 1 931 and 596 tokens) came out identical character for character.
The model and its architecture
Qwen4ExpForConditionalGeneration, model_type: qwen4_exp, GGUF general.architecture: qwen4exp:
- 125B parameters with 6B activated per token, plus a 51B n-gram embedding table and a 4B MTP layer
- 48 layers: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE)), hidden size 2560
- Gated DeltaNet: 48 V heads, 16 QK heads; Qwen Sparse Attention: 24 Q heads, 2 KV heads
- MoE: 512 experts, 10 routed + 1 shared per token, expert width 640
- Gated residual (widened residual streams), n-gram embedding at layer 2, 1 MTP layer
- Context 262,144 tokens natively
DACAN checks this geometry on load (48 layers, 2560, 512 experts, 10 active, 24 / 2 heads) and refuses anything else, so fine-tunes of Qwen3.8-Flash-Next with the same shape work, other models and pruned variants (REAP etc.) do not.
Two layouts: NVFP4 + Q8 (DACAN) and Q8_0
NVFP4 + Q8 (DACAN) is our name for the layout the DACAN service runs: NVIDIA's NVFP4 routed experts, everything else that is large in Q8_0, the small sensitive matrices left in BF16. The experts are about 95 % of the weights and are read from RAM on every token, so their size sets the speed; the parts every token passes through stay at 8 bits.
| part | NVFP4 + Q8 (DACAN) | size |
|---|---|---|
| routed experts — gate, up, down × 512 × 48 layers | NVFP4 from NVIDIA's checkpoint: E2M1 values, an E4M3 scale per 16 values, an FP32 scale per tensor | 68 GB |
| attention and DeltaNet projections, shared expert, token embeddings, output head | Q8_0 | 4.1 GiB |
| n-gram embedding table | Q8_0 from FP8 | 54 GB |
routers, hyper-connections, sparse-attention indexer, DeltaNet ssm_alpha/beta, ple_key/value |
BF16 | 1.4 GiB |
ple_conv1d / norms |
F16 / F32 | 0.01 GiB |
The Q8_0 layout swaps the experts for ours in Q8_0 (128 GB) and keeps the rest; it needs 120 GiB of RAM for the experts instead of 63 GiB and answers slower (on Swift 1.5: 49 tok/s against 58, median of 9 runs). The same NVFP4 + Q8 (DACAN) mix of the fine-tune Swift 1.5, with experts we quantized from its BF16 weights, is in tirex2001/Swift-1.5-Qwen3.8-Flash-Next-DACAN.
Files
| file | what | size |
|---|---|---|
Qwen3.8-Flash-Next-Q8_0-experts.bin |
routed experts in Q8_0, requantized by us from Qwen's FP8 checkpoint (tools/q8_experts.py: 0.55–0.59 % from the FP8 weights) |
128 GB |
Qwen3.8-Flash-Next-NVFP4-experts.bin |
routed experts from NVIDIA's NVFP4 checkpoint, repacked for DACAN (tools/nvfp4_experts.py; unpacks to NVIDIA's values exactly, checked on layers 0, 21, 47) |
68 GB |
Qwen3.8-Flash-Next-dense-Q8_0.gguf |
everything but the routed experts: attention, DeltaNet, shared expert, router, head — Q8_0 (BF16/F32 where the model keeps them); cut from our 4-bit quant's first shard by dense_gguf.py, 1 079 of 1 223 tensors, sha256 3cceeeca…d470d525 |
5.99 GB |
Qwen3.8-Flash-Next-PLE-FP8-Q8_0.gguf |
the n-gram embedding table, Q8_0 from FP8 | 54 GB |
pack-nvfp4/, pack-q8_0/ |
DACAN packs (expert index, dense blob, tokenizer): pack-nvfp4 for NVFP4 + Q8 (DACAN), pack-q8_0 for Q8_0 |
1.5 GB |
mtp/ |
the MTP draft layer | 0.8 GB |
data/profile_other_2609.bin, data/usage_other_2609.bin |
routing tables for the run command | 0.3 MB |
Layout. Put the dense GGUF and the *-experts.bin you use in one folder: pack-*/native_experts.txt names the
experts file without a path and DACAN looks for it next to the --native GGUF. The service config we run (NVFP4 + Q8 (DACAN)) is
--pack pack-nvfp4 --native Qwen3.8-Flash-Next-dense-Q8_0.gguf --ple-gguf Qwen3.8-Flash-Next-PLE-FP8-Q8_0.gguf --mtp mtp --expert-profile data/profile_other_2609.bin --second-card 1 --second-card-usage data/usage_other_2609.bin --max-context 262144 --kv int8 --spec 4 (full list in the DACAN README).
How to run: see the DACAN README. Measured 03.10.2026 on 2× RTX 2080 Ti + 2× Xeon 8368 (NVFP4 + Q8 (DACAN), free node, greedy):
| tok/s | |
|---|---|
| answer: code / a list after a fresh 19K-token prompt | 76.9 / 80.4 |
| answer with reasoning low | 71.9 |
| answer at 98K tokens of context | 69.9 |
| answer: English / Russian prose | 60.4 / 47.0 |
| prompt read from scratch: 19K / 98K tokens | 786 / 715 |
The answer speed follows the MTP drafts: 85–89 % accepted on code and lists, 57 % on English prose, 32 % on Russian. Context 256K (a needle found up to 195K). Details: README-2x2080Ti.
Licenses
The model: Qwen Community License 1.0, copyright (c) 2026 Qwen. The NVFP4 experts are derived from NVIDIA's checkpoint and licensed by NVIDIA Corporation under the NVIDIA Open Model License.
Support the project · Поддержать проект
Everything here — the quants, the engine, the measurements — is made and published for free. If it helped you run a model on your own hardware, you can say thanks with a donation. Если это помогло вам запустить модель на своём железе, можно поблагодарить донатом.
| address | QR | |
|---|---|---|
| YooMoney / ЮMoney (roubles: a YooMoney wallet or any bank card) | 4100119356331418 · send / перевести |
|
| USDT (TRC-20, Tron network) | TBvoJHi7uyonSpvdH9Y6RAAXGWYVR2jeqw |
|
| Ethereum (ETH and Ethereum-network tokens, ERC-20) | 0x66CA7c683fbaF030b2300c918A3751209eA30dEa |
⚠️ Send only Tron-network assets (USDT TRC-20, TRX) to the Tron address and only Ethereum-network assets to the Ethereum address; anything sent over the wrong network is lost. Сеть важна: отправленное не в той сети пропадёт.
Thank you! Спасибо! 🙏
- Downloads last month
- 192
8-bit
Model tree for tirex2001/Qwen3.8-Flash-Next-DACAN
Base model
Qwen/Qwen3.8-Flash-Next