Community evaluation (2026-10-03): GSM8K 97.35%, GPQA 88.69% (NV-protocol 5-run mean), MMLU-Pro 83.70% — 4× RTX PRO 6000 Blackwell, Karmic Kraken serving stack

#5
by malaiwah - opened

Adds a community-evaluation section to the model card with independent accuracy measurements of this checkpoint (commit 175ae8ce), served end-to-end on 4× RTX PRO 6000 Blackwell (96 GB each) with the Karmic Kraken vLLM stack (ghcr.io/local-inference-lab/vllm:karmic-kraken-beta, digest-pinned, vLLM 0.1.dev22021+g93dabce32, MTP draft 3, LMCache engine-driven L1 64 GiB / L2 512 GiB).

Highlights (llm-inference-bench v0.7.6, pinned sha256-verified datasets, per-item scoring, Wilson 95% CIs):

  • GSM8K: 97.35% (1284/1319) greedy — at/above the BF16-band community anchors (96.9–97.2%); concurrency-invariant (C24 = C30)
  • GPQA Diamond: 88.69% 5-run mean at the NVIDIA sibling-table protocol (temp 1.0 / top_p 0.95 / 327,680 tokens); temp-0 greedy-loop truncations shown to be a sampling artifact
  • MMLU-Pro: 83.70% (837/1000) greedy — first published MMLU-Pro datapoint for this checkpoint

Full methodology, per-run numbers, and confidence intervals in the README diff. Raw JSON outputs available on request.

Closing: this discussion was created without a git reference (empty PR). The complete evaluation has been submitted as PR #6 with the README diff: https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4/discussions/6

malaiwah changed discussion status to closed

Sign up or log in to comment