Avi Fenesh PRO
Avifenesh
AI & ML interests
https://inference.tiyuvta.ai/
Recent Activity
updated a model about 11 hours ago
Avifenesh/Qwen3.8-27B-DFlash2-memra liked a model about 13 hours ago
Avifenesh/Qwen3.8-27B-DFlash2-memra repliedto their post about 14 hours ago
Wrote this up because it still feels backwards. Acceptance dropped, tok/s went up.
Full draft head: 66.7% / 117.1 tok/s. Trimmed: 63.6% / 121.7. Later remeasure +5.1% on a Pro 6000, +6.4% on a 5090 laptop. Full head still accepts more. Loses anyway.
Not my idea. FR-Spec. I just wired it for the native MTP head.
https://huggingface.co/blog/Avifenesh/masked-mtp-drafts
https://huggingface.co/Avifenesh/Qwen3.8-27B-NVFP4-MTP-GGUFOrganizations
None yet