My findings for this setup.

#1
by s1arsky - opened

Interesting release. I will use it now instead of Unsloth Q4 UD equivalent, for now. I use 3090 single GPU.

My findings during testing this model Q3 UD XL with the recommended settings from its description page:

  1. --cache-reuse 256 \ is a no-op on qwen35 (Qwen says). Fact is that I get: 'W srv load_model: cache_reuse is not supported by this context, it will be disabled'. On top of that - cache reuse is not supported when multimodal.
  2. Dflash2 doesnt work with with vision. the spiritbun repo linked here - is missing the draft model. z-lab Q8 draft of DFlash2 works fine as replacement. Q8 draft is significantly faster than when I used Q2 draft. I do not look at acceptance rate because in my experience that doesn't translate to T/s speed llama logs.
  3. MTP is build-in and works with vision. Despite showing slightly less acceptance, I find it slightly faster than DFlash2 (for me) in terms of T/s (--spec-type draft-mtp --spec-draft-n-max 3 -ngld 99 )
  4. Without setting --parallel 1 \ there is OOM at max context
  5. I achieve speed of 30-60 T/s with MTP or DFlash2 with Q3_UD_XL Unleashed.
  6. Uncensor is flawless for me (with reasoning on).
  7. Vision with this model is MUCH faster than unsloth Q4. In fact it is first local vision (dense model) that has vision decoding speed I deem usable.

πŸ”₯

Sign up or log in to comment