IMO after trying it in VLLM with custom kernels

#1
by coughmedicine - opened

Not as good as the Q3.6 or Gemma4 models for RP. The best I could get was 80tg/s and pretty inconsistent results, but its a base/preview so maybe it gets better.
Or maybe my fp8 quant and claudes kernels ruined it and I should give the official kernels a try, but 30tg/s is pretty meh.

Gemma4 in realworld test is incredible but this is an interesting design by a US company worth keeping an eye on. Q3.5 beats every test I throw at it compared to 3.6. not sure what people are seeing. MTP enabled has a marked improvement in quality not just performance with that family of Qwen models. Very interesting for sure.

trohrbaugh changed discussion status to closed

Sign up or log in to comment