DiffusionGemma 26B A4B IT — ModelDeck self-contained GPTQ Q4 g32

ModelDeck source and runtime

This is the self-contained ModelDeck Q4/BF16 hybrid for google/diffusiongemma-26B-A4B-it. It quantizes all 30 Mixture-of-Experts layers to symmetric GPTQ 4-bit weights with group size 32 and packages the remaining model weights in BF16.

The pinned upstream model and revision remain recorded for provenance, but this release does not download or load the upstream checkpoint at runtime. All model, processor, tokenizer, and generation files needed by the ModelDeck loader are included here.

This package is not a standard Transformers or GPTQ checkpoint. Use the custom ModelDeck direct Q4 loader; Ollama, llama.cpp, generic from_pretrained(), and vLLM do not directly understand this hybrid checkpoint layout.

Quantization

  • Method: GPTQ, 4 bits, symmetric
  • Group size: 32
  • Activation ordering: disabled
  • Runtime: GPTQModel Triton V2 on ROCm
  • Quantized tensors: expert gate_up_proj and down_proj
  • Packed expert tensors: 12.347 GiB
  • Packaged non-expert BF16 tensors: 5.561 GiB
  • ModelDeck source commit: c7ebfc89af267631449e99d8488a01d41e33e3b3

Validated configuration

  • Base model: google/diffusiongemma-26B-A4B-it
  • Base revision: 52de6b914ee1749a7d4933202505ddf5b414ec43
  • Device: AMD Radeon 8060S Graphics
  • Torch: 2.9.1+rocm7.2.1.gitff65f5bc
  • HIP: 7.2.53211-e1a6bc5663
  • Transformers: 5.13.0
  • Maximum output length: 256 tokens
  • Maximum denoising steps: 48
  • Temperature: 0.8

Release evaluation

Measure Q4 BF16
Contract passes 9/9 9/9
Constraint passes 9/9 9/9
Median wall time 9.589 s 5.202 s
Steady allocated memory 18.061 GiB 48.251 GiB
  • Q4/BF16 median latency ratio: 1.8434×
  • Peak Q4 allocated memory: 19.952 GiB
  • Steady allocation reduction versus BF16: 62.57%
  • Mean token edit similarity versus BF16: 0.6341
  • Exact same-seed deterministic replay: passed
  • Additional Q4 stability contracts: 4/4

The complete prompt-level results and release gates are included in q4-quality-evaluation.json.

ModelDeck usage

Download the immutable release into ModelDeck's expected checkpoint directory:

hf download ozyjay/diffusiongemma-26b-a4b-it-modeldeck-gptq-q4-g32 `
    --revision v1.1.0 `
    --local-dir var/diffusiongemma-26b-a4b-it-gptq-q4-g32

Then start the isolated worker from the ModelDeck repository:

./scripts/start_diffusiongemma_q4.ps1 -Smoke

Verify every packaged file before use:

./scripts/package_diffusiongemma_q4_release.ps1 -VerifyOnly

Scope and limitations

  • Validated for text-diffusion generation on one AMD Radeon 8060S (gfx1151) using the pinned ROCm stack above.
  • The Q4 experts reduce memory substantially but are slower than BF16 on the tested hardware.
  • The package has not been validated for multimodal generation, other GPUs, other base revisions, or other GPTQ runtimes.
  • Compatibility is tied to the ModelDeck source commit above; treat a different loader revision as unvalidated until the release gate passes again.
  • Generated output can be inaccurate, biased, or unsafe. Apply task-appropriate safety and factuality checks.

Licence and provenance

The base DiffusionGemma model is published by Google DeepMind under Apache License 2.0. This package contains transformed expert weights derived from the pinned base revision. See LICENSE and THIRD_PARTY_NOTICES.md for redistribution information.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ozyjay/diffusiongemma-26b-a4b-it-modeldeck-gptq-q4-g32

Quantized
(37)
this model