- Qwen3.8-Flash-Next on one or two RTX 3090s with 64 GB or 128 GB RAM: W4A16 + FP8 PLE + MTP3
Qwen3.8-Flash-Next on one or two RTX 3090s with 64 GB or 128 GB RAM: W4A16 + FP8 PLE + MTP3
Run Qwen 3.8 Flash Next locally on one or two 24 GB GPUs. The weights on this page need one of two custom vLLM runtimes; pick the row for your hardware:
| Hardware | Runtime (GitHub) | Prefill, 131K prompt | Decode | Context |
|---|---|---|---|---|
| 1ร RTX 3090 24 GB + 64 GB RAM | qwen38-flash-next-3090 | up to 2,106 tok/s | 43โ51 tok/s | 135,168 |
| 2ร RTX 3090 24 GB + 64 GB RAM | qwen38-flash-next-2x3090, 64 GB profile | 3,401โ3,413 tok/s | 84โ89 tok/s | 135,168 |
| 2ร RTX 3090 24 GB + 128 GB RAM | qwen38-flash-next-2x3090 | up to 4,191 tok/s | up to 111.5 tok/s | up to 262,144 |
New: opt-in uncensored mode
The two-GPU runtime can serve these weights with their refusals removed: add QWEN38_ABLITERATION=orcarouter
to .env. It applies the edit of
orcarouter/Qwen3.8-Flash-Next-Uncensored
at runtime: one refusal direction is projected out of every residual-stream write. The weights on this page stay
unchanged and nothing new is downloaded. On two RTX 3090s with the agent 128K profile, the model refused 7 of 8
mild borderline requests with the switch off and 0 of 8 with it on. Smoke tests passed and speed was unchanged
(131K prefill 2,909 vs 2,914 tok/s).
This removes the model's safety refusals: put your own safeguards in front of it before serving anyone else. It needs an image built from the GitHub repository's main branch until the next release image. See how it works, how it was verified and its limits.
New: two RTX 3090s with 64 GB of RAM
The 64 GB profile of the two-GPU runtime (v0.5.0) runs both cards with 64 GB of system RAM. The 128 GB profiles keep a pinned copy of every expert and the PLE table in RAM. Here each GPU owns its 88 most-used experts per layer outright, and the other experts of its half live once in RAM (38 GiB for both GPUs). During decode a CPU thread pool per GPU computes them while the GPU takes a share over PCIe. Prefill streams them to the GPUs in 8,192-token chunks. The larger chunks, not the RAM layout, are why it prefills faster than the 128 GB profile, which uses 4,096 (details). The FP8 PLE table is read in place from these files on NVMe.
| Two RTX 3090s, machine limited to 64 GB RAM, 3 runs | First token | Prefill | Decode |
|---|---|---|---|
| 131,099 + 512 tokens | 38.4โ38.5 s | 3,401โ3,413 tok/s | 84.2โ88.8 tok/s |
| 32,799 + 512 | 10.2โ10.3 s | 3,200โ3,211 tok/s | 77.0โ83.9 tok/s |
| 8,218 + 1,024 | 2.6 s | 3,114โ3,146 tok/s | 78.2โ79.7 tok/s |
With the server restricted to 12 CPU cores on 2 CCDs, decode ran at 71โ81 tok/s. Setup: clone the GitHub
repository, then cat configs/2x3090-64gb.env >> .env && make serve. No CUDA P2P and no swap are needed. See
the 64 GB benchmark report.
New: a single RTX 3090 with 64 GB of RAM
github.com/DominikBucko/qwen38-flash-next-3090 serves this checkpoint on one RTX 3090 (24 GB) and 64 GB of system RAM, with a 128K context. The GPU keeps the attention and dense weights, the 32 most-used experts of every layer and an INT8 KV cache (141,504 tokens in 3.05 GB); the CPU computes the other experts straight from RAM during decode, while prefill streams them through the GPU. The FP8 PLE table is read in place from these files on NVMe. The weights are unchanged.
| One RTX 3090, machine limited to 64 GB RAM (2 runs) | First token | Prefill | Decode |
|---|---|---|---|
| 131,099 + 512 tokens | 62.2 / 90.3 s | 2,106 / 1,451 tok/s | 47.4 / 43.3 tok/s |
| 32,799 + 512 | 15.3 / 14.8 s | 2,141 / 2,219 tok/s | 50.3 / 49.5 tok/s |
| 8,218 + 1,024 | 5.2 / 5.7 s | 1,570 / 1,452 tok/s | 49.1 / 48.5 tok/s |
Some prefill requests take a slower path (~1,450 instead of ~2,100 tok/s); see the known issue in the benchmark report. Decode depends mainly on RAM bandwidth: with the server restricted to 16 cores of the benchmark CPU it ran at 39โ47 tok/s, with 8 cores at 34โ36 tok/s. Quick start:
hf download albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE \
--revision ef554143369a706525336f6b42a09094835dc077 --local-dir /models/qwen38-flash-next
git clone https://github.com/DominikBucko/qwen38-flash-next-3090.git && cd qwen38-flash-next-3090
cp .env.example .env # set MODEL_DIR
make build-image && make serve
Or skip the build with the published image:
IMAGE=ghcr.io/dominikbucko/qwen38-flash-next-3090@sha256:7f176605b59c462af21b1b62fd854f9f3cca96e3d11a295581297483d45490a3 make serve.
See the benchmark report, how it works and the FAQ.
Two RTX 3090s: up to 4,191 tok/s prefill ยท up to 111.5 tok/s decode
2ร RTX 3090 + 128 GB RAM ยท v0.5.0 runtime (September 30)
One request with 131,072 input tokens and 2,048 output tokens, best of 2: the new prefill profile reaches the first token in 31.3 seconds (4,191 input tok/s) and decodes at 101.8 tok/s. The agent 128K profile takes 43.1 s (3,045 input tok/s) and decodes at 111.5 tok/s. The fast 256K profile reads the full 256K window (260,096 input tokens) in 90.8 seconds (2,865 input tok/s).
| Profile | Input + output tokens | First token | Input tok/s | Decode tok/s |
|---|---|---|---|---|
| Prefill (8K chunks, hot88), best of 2 | 131,072 + 2,048 | 31.3 s | 4,191 | 101.8 |
| Agent 128K (4K chunks, hot100), best of 2 | 131,072 + 2,048 | 43.1 s | 3,045 | 111.5 |
| Fast 256K, 3 runs (September 29) | 131,072 + 2,048 | 45.0โ46.0 s | 2,852โ2,916 | 98.9โ101.9 |
| Fast 256K, 3 runs (September 29) | 260,096 + 2,048 | 90.8 s | 2,864โ2,865 | 94.0โ98.3 |
The prefill profile streams experts in 8,192-token chunks and keeps 88 instead of 100 experts per layer on each GPU, so each decode step pulls more experts from system memory. Choose it when long prompts dominate the wait.
Clone the GitHub repository and use configs/fast-256k.env, configs/agent-128k.env or
configs/agent-128k-prefill.env, with a local build or the prebuilt
ghcr.io/dominikbucko/qwen38-flash-next-2x3090:v0.5.0 image (pin the digest
from the release notes). The model weights on this page are unchanged.
v0.4.0 fixes a KV-cache leak that kept one recurrent-state block per prefill chunk alive in each Mamba layer group. Long requests used to fill the cache and restart several times, and now they run without preemptions. The agent profile limits requests to 135,168 tokens and uses the freed cache memory to keep 16 more experts per GPU on the cards (hot100). Decode then fetches fewer experts from system memory. Target weights, BF16 KV, FP8 PLE, ten-expert routing and the approximate QSA budget are unchanged.
Decode depends on free RAM: the runtime keeps about 60 GB of expert weights and the 51.2 GB PLE table in system memory. Input tok/s means input tokens / time to first token. See the fix report, the agent profile report, the prefill profile results and the September 25 results.
Setup on two GPUs
Hardware requirements, 4090 guidance and community 5090 report ยท Performance tuning and public benchmark client ยท Share a hardware result
This hybrid serves one native 262,144-token context across both cards, with
NVMe-backed swap for loading headroom. It keeps Intel's AutoRound target
tensors exactly as published, replaces only the 102.4 GB BF16 n-gram/PLE table
with RadixArk's FP8 table, and adds a compact INT4 group-32 MTP draft under
runtime/mtp-int4-g32.
No target tensor was requantized or repacked during assembly.
Composition
| Component | Format | Pinned source |
|---|---|---|
| Target routed experts and eligible linear weights | AutoRound W4A16, INT4 symmetric group-128 | Intel/Qwen3.8-Flash-Next-W4A16-AutoRound@861536dda5bcb208376fc4cd879b2bf76bece9fe |
| Sensitive target layers | BF16, unchanged | Intel checkpoint above |
| 51.2B-parameter n-gram/PLE table | FP8 E4M3FN plus published scale | RadixArk/Qwen3.8-Flash-Next-NVFP4@7b719225242aacd3dbd3f9407468c2ee9a9d2594 |
| Optional MTP draft | Routed experts INT4 symmetric group-32; other tensors unchanged | runtime/mtp-int4-g32 |
The target contains 222,716 indexed tensors in 25 safetensors files with
124,750,778,874 bytes (116.183 GiB) of tensor payload. The compact MTP draft
contains 4,639 tensors in two files with 4,139,535,872 bytes (3.855 GiB) of
payload. hybrid_sources.json, runtime/mtp-int4-g32/compact_sources.json, and
runtime/repro.lock.json are machine-readable provenance records.
Runtime
This is not a stock Transformers checkpoint. Use the matching GitHub runtime
release and the digest-pinned vLLM image recorded in runtime/repro.lock.json.
For one GPU, use qwen38-flash-next-3090
(INT8 KV cache, CPU cold experts, 135,168-token context); the rest of this section
describes the two-GPU runtime.
The current GitHub default uses BF16 KV, TP2+EP2, UVA expert offload, an
84-expert GPU hot cache, prefix caching, and MTP3. The original bundled
runtime/README.md describes the older hot88 release; use the current GitHub
quickstart for the prefill-memory fixes and hot84 default.
The original measurements used runtime release
v0.1.0.
The current setup and measurement guide
adds reproducible probes and P2P/allocator diagnostics while retaining the
checkpoint tensor revision ef554143369a706525336f6b42a09094835dc077.
Configure at least 32 GiB of fast NVMe swap before loading the checkpoint;
48โ64 GiB is safer. If the first prompt raises a CUDA OOM, first check that an
old .env is not still selecting hot88. Start with
VLLM_WNA16_STATIC_HOT_CACHE_SIZE=84, then try 80 if needed. Each removed slot
saves roughly 116 MiB per GPU, with a decode-speed tradeoff. The
memory guide
documents host OOMs, KV-cache tuning, prefill transients, and two-client
capacity.
Measured performance
On 2ร RTX 3090 with 128 GB of system memory:
September 5 verified candidate
- 258,048 input + 4,096 output, three measured
repo-chatruns with no explicit warmup: 75.636 API-observed output tok/s by reciprocal mean TPOT (74.031โ76.707), with TTFT from 211.059 to 215.128 seconds; - 128 input + 4,096 output, one warmup and three measured
repo-chatruns: 77.2845 API-observed output tok/s (74.746โ79.707).
This native candidate used an 84-expert hot cache, CUDA P2P in both directions,
custom all-reduce enabled, and
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False. It ran the pinned vendor
vLLM plus the public overlay in a clean native environment with existing
dependencies, not a fresh Docker build. Current GitHub defaults use an
84-expert cache, custom all-reduce disabled, and expandable segments enabled.
All 27 model weight files matched the published SHA-256 manifest for canonical
tensor revision ef554143369a706525336f6b42a09094835dc077.
Four recoverable allocator warnings appeared during the first long prefill; all three measured streams completed with exact usage counts. Generated-token counts include reasoning and control tokens, and the forced 4,096-token capture can end during reasoning. These probes measure serving performance, not answer quality. See the benchmark bundle, machine-readable summary, and long-context chart.
Historical release measurements
- 262,016-token prompt: 1,275.6 prompt token/s;
- 128-output boundary probe after that prompt: 54.5 token/s;
- warmed 128-input/4,096-output greedy probes: 127.1โ134.0 output token/s;
- MTP acceptance on the warm probes: 86.3โ90.8%.
The figures above are historical single-request measurements. The 128-output boundary probe is too short to characterize sustained long-context generation. The public benchmark protocol uses 258,048 input + 4,096 output for that question and keeps new workload results separate. Agent quality evidence is single-run and provisional; private benchmark fixtures and traces are not included.
Limitations and license
- Maintainer validation uses SM86/RTX 3090. The hardware guide separately records a community dual-5090 report; it is not a maintainer benchmark.
- Dual RTX 4090 is not yet validated; no 4090 throughput claim is made.
- Optimized for one full-context request rather than high concurrency.
- PLE lives in host memory but can be paged to swap. Sustained paging can hurt performance; check residency and swap activity during serving.
- MTP is speculative: target verification preserves target token decisions, while the draft affects acceptance and speed.
- Review the Qwen Community License included in this repository and all upstream model cards before redistribution or commercial use.
- Downloads last month
- 6,188
Model tree for albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE
Base model
Qwen/Qwen3.8-Flash-Next