DiffusionGemma 26B A4B IT — ModelDeck self-contained GPTQ Q4 g32
This is the self-contained ModelDeck Q4/BF16 hybrid for google/diffusiongemma-26B-A4B-it. It
quantizes all 30 Mixture-of-Experts layers to symmetric GPTQ 4-bit weights with group
size 32 and packages the remaining model weights in BF16.
The pinned upstream model and revision remain recorded for provenance, but this release does not download or load the upstream checkpoint at runtime. All model, processor, tokenizer, and generation files needed by the ModelDeck loader are included here.
This package is not a standard Transformers or GPTQ checkpoint. Use the
custom ModelDeck direct Q4 loader; Ollama, llama.cpp, generic from_pretrained(), and
vLLM do not directly understand this hybrid checkpoint layout.
Quantization
- Method: GPTQ, 4 bits, symmetric
- Group size: 32
- Activation ordering: disabled
- Runtime: GPTQModel Triton V2 on ROCm
- Quantized tensors: expert
gate_up_projanddown_proj - Packed expert tensors: 12.347 GiB
- Packaged non-expert BF16 tensors: 5.561 GiB
- ModelDeck source commit:
c7ebfc89af267631449e99d8488a01d41e33e3b3
Validated configuration
- Base model:
google/diffusiongemma-26B-A4B-it - Base revision:
52de6b914ee1749a7d4933202505ddf5b414ec43 - Device: AMD Radeon 8060S Graphics
- Torch:
2.9.1+rocm7.2.1.gitff65f5bc - HIP:
7.2.53211-e1a6bc5663 - Transformers:
5.13.0 - Maximum output length: 256 tokens
- Maximum denoising steps: 48
- Temperature: 0.8
Release evaluation
| Measure | Q4 | BF16 |
|---|---|---|
| Contract passes | 9/9 | 9/9 |
| Constraint passes | 9/9 | 9/9 |
| Median wall time | 9.589 s | 5.202 s |
| Steady allocated memory | 18.061 GiB | 48.251 GiB |
- Q4/BF16 median latency ratio: 1.8434×
- Peak Q4 allocated memory: 19.952 GiB
- Steady allocation reduction versus BF16: 62.57%
- Mean token edit similarity versus BF16: 0.6341
- Exact same-seed deterministic replay: passed
- Additional Q4 stability contracts: 4/4
The complete prompt-level results and release gates are included in
q4-quality-evaluation.json.
ModelDeck usage
Download the immutable release into ModelDeck's expected checkpoint directory:
hf download ozyjay/diffusiongemma-26b-a4b-it-modeldeck-gptq-q4-g32 `
--revision v1.1.0 `
--local-dir var/diffusiongemma-26b-a4b-it-gptq-q4-g32
Then start the isolated worker from the ModelDeck repository:
./scripts/start_diffusiongemma_q4.ps1 -Smoke
Verify every packaged file before use:
./scripts/package_diffusiongemma_q4_release.ps1 -VerifyOnly
Scope and limitations
- Validated for text-diffusion generation on one AMD Radeon 8060S (
gfx1151) using the pinned ROCm stack above. - The Q4 experts reduce memory substantially but are slower than BF16 on the tested hardware.
- The package has not been validated for multimodal generation, other GPUs, other base revisions, or other GPTQ runtimes.
- Compatibility is tied to the ModelDeck source commit above; treat a different loader revision as unvalidated until the release gate passes again.
- Generated output can be inaccurate, biased, or unsafe. Apply task-appropriate safety and factuality checks.
Licence and provenance
The base DiffusionGemma model is published by Google DeepMind under Apache License 2.0.
This package contains transformed expert weights derived from the pinned base revision.
See LICENSE and THIRD_PARTY_NOTICES.md for redistribution information.
- Downloads last month
- 14
Model tree for ozyjay/diffusiongemma-26b-a4b-it-modeldeck-gptq-q4-g32
Base model
google/diffusiongemma-26B-A4B-it