Clef-Flash mixed-precision GGUF (measured-KL, text only)

Mixed precision, not ternary. Tensors use IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K and Q8_0; no ternary format (TQ1_0/TQ2_0) is used. This repo was renamed from Jakevin/clef-flash-ternary-GGUF on 2026-10-09; the old URL redirects here and the v1.0 tag is unchanged.

Version v1.0 (2026-10-09). Unofficial post-training quantization of Cloudflare/clef-flash revision 17f0b0ad64efb65d273590632833508766b2aae6. This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team. Licensed Apache-2.0 like the original (LICENSE). NOTICE.md lists the changes.

The file is the text backbone only. There is no vision tower. llama.cpp by itself does not produce Clef decisions: the joint schema head runs in Python on the backbone's final hidden states, and the option rows come from the original bf16 lm_head. The GGUF output.weight is Q2_K and the head does not read it.

general.architecture is qwen35. The llama.cpp clef graph does not export t_h_nextn, which is the hidden state this recipe reads, so the file is the Qwen3.5 text model rather than architecture clef. general.name is Snap Text because the conversion directory had that name. The weights are Clef-Flash text.

Size

bytes
clef-flash-mkl-3.30GB.gguf 3.300 GB (3,299,992,800)
original bf16 release 19.06 GB (18.82 GB shards + 0.24 GB head)

Scores

Author's private frozen development splits (kev). One seed. No confidence interval was computed.

suite n (clean) bf16 acc this model retained
decision-v7 1264 0.8861 0.8813 99.5%
transfer-v9 1046 0.8011 0.7859 98.1%
suite Brier NLL ECE
decision-v7 0.1755 0.3382 0.0278
transfer-v9 0.3061 0.6042 0.0421

Retention is this model's clean accuracy divided by the bf16 clean accuracy on the same suite.

decision-v7 can be optimistic. The imatrix and the KL allocation both used decision-v7 records: 128 calibration records, seed 1234, for the imatrix, and the first 32 of that draw (10,383 tokens) for the per-tensor KL. transfer-v9 was not used for calibration.

These are not the public Decision Index or Typesafe numbers.

Method

Per-tensor measured-KL sensitivity, then a knapsack that assigns each measured tensor a ggml type from IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K, and Q8_0. The shipped file is llama-quantize --tensor-type with that imatrix. alloc/selection.json is the knapsack result. alloc/tensor-types.txt is the pattern file passed to the quantizer (it pins ssm_alpha and ssm_beta to bf16; the other assignments are exact tensor names).

Counts in the GGUF (427 tensors):

type tensors what
IQ2_XXS 68 body linears
IQ2_S 23 body linears
IQ3_XXS 35 body linears and token_embd
Q2_K 16 15 body linears plus output.weight
Q3_K 13 body linears
Q4_K 36 body linears
Q8_0 11 body linears
BF16 48 ssm_alpha, ssm_beta (24 each)
F32 177 norms, ssm_conv1d, ssm_a, ssm_dt

Norms, conv, ssm_alpha, and ssm_beta are not quantized. In the file, norms and conv are F32; ssm_alpha and ssm_beta are BF16. output.weight is fixed at Q2_K and is not read by the classification head. token_embd is IQ3_XXS (IQ2_XXS asserts an imatrix, and llama-quantize does not pass one for the embedding).

Sum of the per-tensor KL values used by the knapsack (not a model KL): 0.009015. Dry-run file estimate was 3.300 GB.

Other files of the same source model

Listed side by side. Different formats and different calibration. This card does not say which allocation is better. Only numbers from the author's result files are shown; anything not measured says so.

release format size decision-v7 acc (retained) transfer-v9 acc (retained) KL(bf16‖·) d7 / t9 source
Cloudflare/clef-flash bf16 (reference) safetensors bf16 17.90 GB text 0.8861 (100%) 0.8011 (100%) 0 / 0 [a]
this repo, v1.0 GGUF, measured-KL mixed types 3.300 GB (3,299,992,800 B) 0.8813 (99.5%) 0.7859 (98.1%) not measured [b]
Jakevin/clef-flash-ternary-mlx v2.0 (b31) MLX packed mixed-bit GPTQ 3.13 GB 0.8766 (98.9%) 0.7361 (91.9%) 0.0562 / 0.1881 [c]
Jakevin/clef-flash-ternary-mlx v1.0 (T-cloq16) MLX ternary + CLoQ 3.13 GB 0.8497 (95.9%) 0.6788 (84.7%) 0.0896 / 0.2923 [a], [c]
bartowski/Cloudflare_clef-flash-GGUF Q2_K GGUF static Q2_K 4.19 GB (4,194,031,776 B) not measured not measured not measured [d]

Sizes exclude the 0.24 GB joint head, which every row needs separately. The MLX builds use a different calibration and packed format from this GGUF.

  • Each number is a single run on one seed. No confidence intervals are given for any row. (The MLX v2.0 card has a paired-bootstrap CI for its v2.0 − v1.0 difference only; no GGUF-vs-MLX difference was tested.)
  • This GGUF is scored with the original bf16 lm_head rows. The MLX rows use the 4-bit lm_head copy that ships in those packages (on the bf16 release the two give the same accuracy, 0.8861 / 0.8011 [a]).
  • KL(bf16‖model) over all question rows was measured for the MLX builds only. For this GGUF it was not measured; the 0.009015 in Method is a sum of per-tensor KLs against a Q8_0 reference, not a model KL. Per-row probability difference vs bf16 was recorded: mean |Δp| 0.0280 (decision-v7, 32 argmax flips in 1,468 rows) and 0.0622 (transfer-v9, 96 flips in 1,264 rows) (clef-gguf-mkl-20261008/out/mkl-*/summary.json).
  • Q8_0 GGUF (the per-tensor KL reference, not released) was not scored on the full suites. Only a path check on the first 100 decision-v7 dev records was run: accuracy 0.93, max |Δp| 0.0753, 1 argmax flip vs the bf16 rows [b]. (The mixed GGUF on the same 100 records: accuracy 0.93 [e].)
  • Static quants such as bartowski's were not scored. hard-v1 was not run on this GGUF.

Sources (author's run folders, kev evals, development splits): [a] ternary-clef-flash-20261006/RESULTS.md (bf16 baseline and MLX v1.0 T-cloq16 evals; sizes are packed files, head excluded). [b] clef-gguf-mkl-20261008/RESULTS.md (this GGUF, full decision-v7 and transfer-v9 dev evals). [c] clef-v2-release-20261008/RESULTS.md (MLX v2.0 b31 evals and KL). [d] clef-gguf-release-20261009/RESULTS.md (bartowski file size only; not scored). [e] clef-gguf-vision-20261009/RESULTS.md (100-record text consistency check and 6-case synthetic image check).

Use

The joint head needs torch, safetensors, numpy, and transformers. hsdump needs llama.cpp b11407 or later with text-only qwen35. Build it from hsdump.cpp in this folder. The two embedding calls are already in that llama.cpp (src/llama-ext.h); they are not in the installed llama.h, so the cpp file declares them. No llama.cpp patch is required at b11407 or later.

output.weight inside the GGUF is not the classification head. Pass --base as a checkout of Cloudflare/clef-flash (the same revision as above). Only lm_head.weight is read from it.

c++ -std=c++17 -O2 \
  -I llama.cpp/include -I llama.cpp/ggml/include \
  hsdump.cpp -o hsdump \
  -L /path/to/llama.cpp/build/bin -lllama \
  -Wl,-rpath,/path/to/llama.cpp/build/bin

python run_clef_gguf.py \
  --base /path/to/Cloudflare/clef-flash \
  --hsdump ./hsdump

The default record is the invoice example from the Cloudflare card (one choice question and one noul question). --record file.json scores another text record of the same shape. Hidden states are llama.cpp final-norm embeddings (llama_set_embeddings_nextn, masked false), the read used by Livesport/clef-flash-GGUF.

Limitations

  • llama.cpp does not run the classifier head by itself. Clef decisions need joint_head.safetensors + joint_schema_model.py in Python (run_clef_gguf.py), hidden states from hsdump (built from hsdump.cpp against llama.cpp b11407+), and the bf16 lm_head option rows from a Cloudflare/clef-flash checkout (--base). llama.cpp chat or completion on this file is not a Clef decision, and the GGUF output.weight (Q2_K) is not the head.
  • Calibration data. The imatrix used 128 records from the decision-v7 calibration split of the author's private kev suite (seed 1234, 44,619 tokens); the per-tensor KL used the first 32 of them (10,383 tokens) against a Q8_0 reference. These are Clef schema classification prompts from the same suite as decision-v7 (its dev split draws on AG News, DBpedia-14, Banking77, TREC, SST-5, IMDB, Yelp, Amazon, MNLI, BoolQ, plus compositional and policy items). decision-v7 can therefore be optimistic; transfer-v9 was not used for calibration. Other domains and languages were not checked.
  • Mixed precision, not ternary. Per-tensor IQ2_XXS / IQ2_S / Q2_K / IQ3_XXS / Q3_K / Q4_K / Q8_0, with BF16/F32 for norms, conv and ssm_alpha/beta. No TQ1_0/TQ2_0. "ternary" in the related MLX repo names does not describe this file.
  • Post-training quantization. No recovery training.
  • No confidence intervals. Single run, one calibration seed; no paired bootstrap against bf16 or against the MLX builds.
  • Evaluated sets only. Scores cover the decision-v7 (n=1264 clean) and transfer-v9 (n=1046 clean) development splits. Not run: hard-v1, the public Decision Index / Typesafe benchmarks, a model-level KL vs bf16, full-suite Q8_0 GGUF, and static quants (e.g. bartowski Q2_K).
  • Text only. No images, no video.

License and attribution

Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself post-trained from Qwen/Qwen3.5-9B (Apache-2.0). The hidden-state read follows Livesport/clef-flash-GGUF. Inference runs on llama.cpp.

繁中摘要

這是 Cloudflare/clef-flash 文字主幹的 measured-KL 混合精度 GGUF,3.300 GB(3,299,992,800 bytes)。decision-v7 正確率 0.8813(bf16 0.8861,保留 99.5%),transfer-v9 0.7859(bf16 0.8011,保留 98.1%)。單一 seed,沒有信賴區間。imatrix 與 KL 用了 decision-v7 的校正資料(imatrix 128 筆、seed 1234;KL 用其中前 32 筆),所以 decision-v7 可能偏樂觀;transfer-v9 沒有參與校正。沒有視覺。分類要跑 Python 的 JointSchemaHead,選項列取自原模型 bf16 lm_head;llama.cpp 單獨跑不出 clef 的判斷。output.weight 是 Q2_K,分類頭不讀它。norms、conv 在檔案裡是 F32,ssm_alpha / ssm_beta 是 BF16,都沒有再量化。

Downloads last month
44
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jakevin/clef-flash-mixed-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(48)
this model