Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Quazim0t0 
posted an update 5 days ago
Post
211
On-Fly-Jev is up: Quazim0t0/On-Fly-Jev

Second Jev-style model. First was Byrne-Jev (70M SpikeWhale). This one is a 96M spiking trunk - every unit is a copy of one of 100 real MaleCNS fly neurons - plus a typed-decision head. One forward pass. No generated text.

I built it for speed.

Typed-decisions test split, 400 cases, 2,000 decisions, 5 questions per case, local GPU:

• 7.3 ms p50 per decision (~137 / s)
• 36.6 ms p50 / 53.5 ms p95 per case of 5
• ~27 cases / s

Same protocol vs the others:

• On-Fly-Jev: 36.6 ms / case, ~137 decisions / s
• Byrne-Jev: 110.7 ms / case, ~45 / s
• ModernBERT-base: 349 ms / case, ~14 / s
• TypeSafe Jev 1.13 (hosted, so network is in it): 710 ms / case, ~7 / s

About 3x Byrne-Jev, 9.5x ModernBERT, 19x TypeSafe Jev per case.

Live ViZDoom, 1 question per tick including game I/O: 23.4 ms (~43 / s).

• Accuracy 0.666 (Byrne-Jev 0.630, Jev 1.13 0.727)
• ECE 0.045, same as Byrne-Jev, about a third of Jev 1.13

50/50 merge of two checkpoints from one run. Research artifact, not a chatbot. More videos are on the card.

Your choice:2 temperature never runs through your own agent.

decision_config.json ships it at 0.12. agent.py passes every temperature through clamp_temperature on load, which pins it to [0.5, 5.0] (TEMP_MIN, "Laya's runtime clamp"). eval_decisions.py scores through DecisionAgent, and serve.py does too. So every reported number uses 0.5 for 2-option choice questions, 4.2x softer than the fit.

It also never touches the headline. I bucketed the typed-decisions test split with your temp_bucket:

2,000 decisions    choice:3-5  600    noul:2  600    score:3-5  800

So the 0.045 ECE is set by three temperatures. choice:2, choice:6-10 and choice:11+ only act on the held-out sources, where ECE runs 0.08 to 0.40.

The 0.12 looks like a fit that ran to its floor. fit_one_temp clamps at 0.1. In the mix code, the 2-option choice items are the SST2 and IMDB sentiment ones, one-hot targets, and IMDB sits at 0.985 accuracy. On a set that is nearly all correct, NLL keeps falling as T goes to 0.

Which set was the choice:2 bucket fit on? And what does SST2's choice-form ECE look like at 0.5 against 1.0?

·

Good catch. Shipped decision_config.json was stale vs the weights.

DecisionAgent does not read that file. It reads temperatures from the checkpoint’s embedded decision_cfg. On-Fly-Jev.pt embeds [1.0, 1.0, 1.0] and has no per-bucket temps. So 0.12 never ran. The clamp never ran. Every number I reported is T = 1.0.

Accuracy does not care. Temp does not move argmax, and it does not move the noul 0.5 cut.

You were right about choice:2. That fit was junk. Hold-out was 12 items: 8 SST2, 2 IMDB, 2 other. It slammed into the 0.1 floor. SST2 got worse, not better:

• T = 1.0, choice-form ECE 0.254
• T = 0.5, 0.297
• T = 0.12, 0.310

One correction: about half of SST2 and IMDB are noul:2, not choice:2.

What I changed:

• Config now matches the weights.
• A bucket needs at least 100 items or it falls back to the question-type temperature.
• Fits clamp to [0.5, 5.0], same as the agent.

I refit on 4,093 held-out benchmark items, kept off the eval set. Ceiling is real. At T = 1, typed is fine (ECE 0.045) and the benchmarks are overconfident (0.197). Fitted temps pull benchmarks to 0.085 and wreck typed (0.245). One T per bucket cannot do both jobs.

Published model stays T = 1. Benchmark-calibrated copy is in calibrated/. README says both now.

The fix you describe isn't on main yet. Only calibrated/ changed.

At your 22:37Z commit (79367672), decision_config.json is the same blob as the commit before it (9d849b2248). choice:2 is still 0.1198, the type temperatures still [1.60, 0.98, 1.50].

README.md is the same blob too (944ee2ece5). It still says calibration is per type and per bucket from decision_config.json, and nothing points to calibrated/ or says T = 1.

The calibrated copy has the mirror image of the old bug. noul:2 fits to exactly 5.0, and so does the noul type temperature. That is your new clamp ceiling. The 0.12 sat on the old floor.

A 2-option fit that wants T above 5 is close to saying the logit gap carries no signal on that pool.

The eval file says otherwise for at least one source. IMDB is 0.985 accurate and has the worst held-out ECE of the ten, 0.154. Confidence can't go above 1, and it gets only 3 of 200 wrong, so at most 0.015 of that ECE can be overconfidence. The rest is the copy being unsure of answers it gets right.

You said about half of IMDB and SST2 are noul:2. So the pooled bucket may hold one source that wants T well under 1 and another that wants it far above 5.

What does the noul:2 NLL look like per source from 0.5 to 5? If one of them is still falling at 5, that bucket has your "one T can't do both jobs" problem inside it.