GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD

This model is based on GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD. GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD is licensed under the Local Inference Lab License, Version 1.0.

Multimodal checkpoint with QAD decoder weights, losslessly compressed routed-expert block scales, MXFP8 attention, NVFP4 MTP routed experts, and mixed-precision vision weights.

Component Stored precision
Decoder routed experts QAD NVFP4; E2M1 values, E4M3 block scales, FP32 global/input scales
Decoder attention MXFP8 projection selection from the pinned source
MTP routed experts NVFP4, BF16 activations; block scales stored directly
Vision attention linears MXFP8 E4M3 values; UE8M0 scales per 32 weights; 48 matrices
Vision MLP and merger linears NVFP4 E2M1 values; E4M3 scales per 16 weights and FP32 global scale; 76 matrices; BF16 activations
Other vision tensors BF16; 223 tensors, including normalization, biases, position embeddings, and convolutions
Normalization parameters 277 BF16 tensors; original bytes preserved

Vision uses weight-only post-training conversion for NVFP4 linears and dynamic MXFP8 attention activations. The vision weights occupy 390,629,680 bytes (372.53 MiB), saving 702.50 MiB of tensor storage versus their BF16 source. These numbers exclude runtime workspaces, activations, and KV cache.

The pinned input is local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD at revision fd660d51d1fc3caae26a4bf31b7451475bbb9bdc. All tensors outside the converted vision linear weights retain their source values and dtypes. CSF reconstructs the original scale bytes exactly.

Storage contract

The checkpoint contains 36 safetensors shards targeting 5 GB each, under tensors/. model.safetensors.index.json maps stored tensors to these files.

Scale encoding: weight_scale_encoding: "csf". Codec: byte-window4-fixed-stream-u24-exceptions/1. The 36,288 compressed matrices are exclusively main-decoder routed-expert E4M3 block scales. Each logical scale tensor is represented by a uint8 .nvfp4_csf_fixed stream and uint32 .nvfp4_csf_exceptions. MTP, vision, PLE, MXFP8, global, and activation scales remain uncompressed.

Runtime requirements

config.json uses the standard quant_method: "modelopt" and quant_algo: "MIXED_PRECISION". Decoder expert recipes declare quant_algo: "NVFP4", group_size: 16, and weight_scale_encoding: "csf". MTP and vision recipes retain their own precision settings without CSF encoding.

vLLM uses its standard safetensors loader and modelopt_mixed quantization. The quantization layer receives compressed scale tensors through normal weight-loading hooks and prepares TP-local scales for b12x. No checkpoint-root setting or CSF-specific load format is needed.

GLM runtimes that force the vision tower to BF16 must enable its explicit mixed-precision recipes, including the attention QKV projections.

Validation

Every stored tensor and shard passed SHA-256 verification. Retained tensors match the pinned input byte-for-byte. Every compressed scale passed independent lossless reconstruction, and reconstructed shard hashes match the reference serialization. Vision NVFP4 packing and rounding were checked against an independent E2M1 reference. GPU inference and multimodal quality evaluation have not been run for this checkpoint.

License

LICENSE, NOTICE, and LICENSES/ retain the source distribution's license and attribution materials, including the upstream model terms.

Downloads last month
336
Safetensors
Model size
175B params
Tensor type
F32
·
U32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD

Unable to build the model tree, the base model loops to the model itself. Learn more.