Quantized-Drafters
Collection
20 items • Updated
GPTQ-Gauss4 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash at pinned revision 1a11b170eb65c8a62c80ecd01dfe22a5907298e6. This repository contains the drafter component; it is not a standalone chat model.
GPTQModifier (nvfp4_expanded_mse).1a11b170eb65c8a62c80ecd01dfe22a5907298e6 (block size 16, draft vocabulary 151,936, sliding-window attention).Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:
vllm serve Qwen/Qwen3-8B \
--spec-model inference-optimization/Qwen3-8B-DFlash-GPTQ-Gauss-NVFP4-W4A4 \
--spec-tokens 7 \
--spec-method dflash
config.py provides the custom drafter configuration. Serving command and runtime patch are preserved in provenance/evaluation/.
Evaluated with 7 speculative tokens per draft event against the Qwen/Qwen3-8B target:
| Benchmark | Mean Acceptance Length ($L = 1 + A/D$) | Mean Accepted Ratio ($A/P$) |
|---|---|---|
| SPEED-Bench Qualitative (11 categories, 880 requests) | 2.0459 tokens | 14.94% |
| RedHatAI / speculator_benchmarks (9 subsets, 924 requests) | 2.0926 tokens | 15.61% |
Full provenance and reproducibility artifacts are preserved:
provenance/ contains:train_command.txt: Source model training command.train_command_scope.txt: Clarification of source drafter snapshot lineage.quantization_command.txt: Exact command and environment used to run quantization.quantization_source/: Source code snapshot of quantizer utilities.evaluation/: Contains vLLM serving commands (vllm_command.txt), vLLM runtime patch (vllm.patch), speculators patch (speculators.patch), target/drafter SHA-256 hashes, and per-subset evaluation commands.publication_sha256.txt: SHA-256 hashes of all published files.The source drafter lists Apache-2.0 licensing on its Hugging Face model card.