Qwen3.8-27B-Uncensored-NVFP4-FastLLM

An NVFP4 build of orcarouter/Qwen3.8-27B-Uncensored with every linear layer in NVFP4. It is 20.6 GB, down from 55.6 GB in BF16 and 29 GB in FP8.

I made it to serve the model on two RTX 2080 Ti 22 GB cards with FastLLM. Turing has no FP4 hardware. FastLLM stores the NVFP4 weights as-is and unpacks them to FP16 inside the kernel. Decoding is limited by weight reads, so the smaller file is faster on these cards even with the extra unpacking.

Every token reads the weights once: FP8 reads 29 GB and decodes 32.9 tok/s without speculation; NVFP4 reads 20 GB and decodes 45.1 tok/s

Write-up with the full numbers, the launch command and the debugging notes: NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s (中文: 2080 Ti 沒有 FP4 也能跑 NVFP4).

⚠️ Inherited disclaimer

The base model has had its safety alignment largely removed by abliteration. Quantizing it does not change that. Everything in the base model's disclaimer applies here. The model will comply with harmful requests. It is meant for research and controlled use, you are responsible for how you use it, and you should add your own moderation before exposing it to anyone else.

What's in it

Weights Format
All linear layers: attention q/k/v/o, GDN projections (in_proj_qkv/z/a/b, out_proj), MLP NVFP4 (E2M1 values, one FP8-E4M3 scale per 16 values, one FP32 global scale per tensor)
lm_head, embed_tokens, vision tower, conv1d, MTP head BF16, unchanged

Storage format per weight group for orcarouter's official NVFP4, lyf's NVFP4 and this build, with code tok/s 65.3, 136.0 and 149.9

  • The layout matches lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL exactly: tensor names, dtypes, shapes, shard map, quantization_config and hf_quant_config.json (2,687 tensors, total_size 20,558,935,392). Only the base model differs.
  • As in ModelOpt's recipe, q/k/v share one global scale per layer, the four GDN input projections share one, and MLP gate/up share one.
  • No calibration: input_global_scale is set to 1.0. FastLLM on sm_75 computes with FP16 activations and does not read it. If you run this on an engine that does W4A4 with real activation scales, expect worse results than a calibrated checkpoint.

How it was made

I converted it with a numpy-only script that runs on the CPU, no GPU needed: qwen38-nvfp4-convert.py.

hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir ./bf16
hf download lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL --include "*.json" "recipe.yaml" --local-dir ./nvfp4-ref
python3 qwen38-nvfp4-convert.py ./bf16 ./out --reference ./nvfp4-ref

To check the converter, I dequantized one attention, one GDN and one MLP layer of lyf's release to BF16 and quantized them again. The packed weights, block scales and global scales came out byte-identical to lyf's.

Results on two RTX 2080 Ti 22 GB (TP=2, FastLLM + DFlash2)

Same machine, same FastLLM build, temperature 0, 512 tokens, median of 3 runs, tok/s:

code math prose no speculation (code)
orcarouter FP8 116.1 135.5 74.2 32.9
this NVFP4 151.7 178.8 101.9 45.1

With fp8 KV cache and --max_batch 4: 2 concurrent streams run at 107.8 tok/s each, 4 streams at 74.3 each (about 297 combined).

Per-stream and combined tok/s for 1 to 4 concurrent code requests with fp8 KV cache

140-question eval (GSM8K 50, MATH-500 30, MuSR 20, HumanEval+ 40):

no thinking thinking (effort low)
orcarouter FP8 132 127
this NVFP4 131 128

Trade-off: prefilling a 54K-token prompt takes 50.6 s instead of 40.5 s. Prefill is compute-bound, so the unpacking is pure overhead there.

I have only tested this with FastLLM on sm_75. I have not tested vLLM, SGLang or Blackwell GPUs.

Serving with FastLLM

Tested on FastLLM at upstream a2bf07fd with d4b04876 reverted and PRs #749 and #756 merged (build steps in the previous post). The draft model is z-lab/Qwen3.8-27B-DFlash2.

export CUDA_VISIBLE_DEVICES=0,1
export FASTLLM_CUDA_GRAPH=0
export FASTLLM_DRAFT_QUANT=nvfp4
export FASTLLM_COOPERATIVE_LONG_PREFILL=1   # from PR #756; leave out on stock FastLLM

ftllm server -p ./Qwen3.8-27B-Uncensored-NVFP4-FastLLM \
  --tp 2 --max_batch 4 --chunked_prefill_size 4096 \
  --gpu_mem_ratio 0.95 --kv_cache_dtype fp8_e4m3 \
  --speculative_algorithm dflash \
  --speculative_draft_model_path ./Qwen3.8-27B-DFlash2 \
  --speculative_num_draft_tokens 8 \
  --prefix_cache true --port 8080

Don't pass --tokens: with it set, FastLLM clamps --max_batch back to 1 for this model.

Credits

Downloads last month
50
Safetensors
Model size
28B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coolthor/Qwen3.8-27B-Uncensored-NVFP4-FastLLM

Base model

Qwen/Qwen3.8-27B
Quantized
(61)
this model