GPTQ-Gauss4-NVFP4-W4A4 — NVFP4 W4A4 with GPTQ (Gaussian calibration)

GPTQ-Gauss4 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash at pinned revision 1a11b170eb65c8a62c80ecd01dfe22a5907298e6. This repository contains the drafter component; it is not a standalone chat model.

Variant

  • Quantization: NVFP4 W4A4 via GPTQModifier (nvfp4_expanded_mse).
  • Calibration: Synthetic Gaussian batches (2,027 batches, sequence length 2,048, seed 0).
  • Hessian Damping: 0.01 (36/36 modules quantized, 0 RTN fallbacks).
  • Source snapshot: Pinned newer BF16 drafter snapshot 1a11b170eb65c8a62c80ecd01dfe22a5907298e6 (block size 16, draft vocabulary 151,936, sliding-window attention).

Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-GPTQ-Gauss-NVFP4-W4A4 \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. Serving command and runtime patch are preserved in provenance/evaluation/.

Evaluation Performance

Evaluated with 7 speculative tokens per draft event against the Qwen/Qwen3-8B target:

Benchmark Mean Acceptance Length ($L = 1 + A/D$) Mean Accepted Ratio ($A/P$)
SPEED-Bench Qualitative (11 categories, 880 requests) 2.0459 tokens 14.94%
RedHatAI / speculator_benchmarks (9 subsets, 924 requests) 2.0926 tokens 15.61%

Reproducibility & Provenance

Full provenance and reproducibility artifacts are preserved:

  • Root directory contains model weights, configuration, tokenizer, quantization manifests, and recipe.
  • provenance/ contains:
    • train_command.txt: Source model training command.
    • train_command_scope.txt: Clarification of source drafter snapshot lineage.
    • quantization_command.txt: Exact command and environment used to run quantization.
    • quantization_source/: Source code snapshot of quantizer utilities.
    • evaluation/: Contains vLLM serving commands (vllm_command.txt), vLLM runtime patch (vllm.patch), speculators patch (speculators.patch), target/drafter SHA-256 hashes, and per-subset evaluation commands.
    • publication_sha256.txt: SHA-256 hashes of all published files.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card.

Downloads last month
33
Safetensors
Model size
1B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/Qwen3-8B-DFlash-GPTQ-Gauss-NVFP4-W4A4

Finetuned
Qwen/Qwen3-8B
Quantized
(454)
this model

Collection including inference-optimization/Qwen3-8B-DFlash-GPTQ-Gauss-NVFP4-W4A4