dmx-qwen2.5-1.5b-m7
DMX M=7 compressed version of Qwen/Qwen2.5-1.5B-Instruct.
Stats
- Source: Qwen/Qwen2.5-1.5B-Instruct (FP16)
- Format: DMX BFP M=7 (7 mantissa bits, block floating point)
- File size: 1.44 GB (53% smaller than FP16)
- Quality: Within GPU variance of FP16 (BF16-equivalent precision)
Usage
pip install dmx-compress dmx-runtime
from dmx_runtime import from_dmx_compressed
model = from_dmx_compressed(
"model.dmx",
model_id="Qwen/Qwen2.5-1.5B-Instruct"
)
Compressed with dmx-compress.
- Downloads last month
- 3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support