gemma-4-12B-it-plumb
google/gemma-4-12B-it with a plumb: a decision head built into the model. It answers
typed decisions (pick one of these options, yes or no, a score on a scale) in one forward pass, with a
calibrated probability for every option, reading the KV cache the conversation already filled. It is a
Jev-style decision model inside the LLM instead of next to it. The base weights are unchanged, so the model
chats, reasons and calls tools exactly like google/gemma-4-12B-it.
Made with Plumbify and served by vLLM through the Plumbify plugin.
Results
447 held-out decisions (task families and templates the plumb never trained on), on one vLLM server. Each cell is accuracy (latency p50 for one request). With thinking on, the model alone is scored in its better prompt setup, so the plumb is compared with the model at its best.
| google/gemma-4-12B-it alone | gemma-4-12B-it-plumb | |
|---|---|---|
| Thinking off | 0.736 (98 ms) | 0.826 (2.2 s) |
| Thinking on | 0.770 (24.1 s) | 0.872 (5.0 s) |
| Eval set | Off: alone | Off: plumbed | On: alone | On: plumbed |
|---|---|---|---|---|
| DecisionBench, held-out task families (150) | 0.833 | 0.833 | 0.793 | 0.840 |
| Hard-skill templates, held out (105) | 0.600 | 0.810 | 0.771 | 0.867 |
| General decisions, unseen templates (192) | 0.734 | 0.828 | 0.750 | 0.901 |
Use
pip install "plumbify[vllm]"
vllm serve totum-labs/gemma-4-12B-it-plumb
In a second terminal, chat with it and watch each decision as the model hands it to its plumb:
plumbify chat
Or ask it through the OpenAI API:
curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{
"model": "totum-labs/gemma-4-12B-it-plumb",
"messages": [{"role": "user", "content": "Ticket: I was charged twice.\n\nWhich team should handle this?\n- billing\n- technical\n- sales"}]
}' | jq .plumb.answer
Decisions stated in a message go to the plumb before the model replies; the model can also hand decisions to
it mid-conversation through a plumb_decide tool, answered from the KV cache. Each response carries a plumb
field with the probabilities, a conformal set and the final answer. "plumb": false in a request serves it
with the base model alone. Request options and response fields are in the guide.
What is in this repository
- The weights of google/gemma-4-12B-it, unchanged.
config.jsonwith the architecturePlumbGemma4UnifiedForConditionalGeneration, which tells vLLM and the plugin to attach the plumb.head.safetensors: the decision head (60.6M parameters). It reads the model's hidden states at layers 23, 29, 35, 41 and the final output.suffix_adapter.*: a LoRA (rank 32) on o_proj, k_proj, down_proj, up_proj, q_proj, v_proj, gate_proj, active only while the model reads a decision. It is never merged.plumb.json: taps, head shape and calibration (temperature 1.122, conformal threshold 0.794 at alpha 0.1).
Training
plumbify train on 10k decision rows drawn from about 200 decision families
(support, finance, coding, safety, medicine, law, engineering) plus solver-labelled hard-skill rows, one epoch,
one A100 80 GB, 76 minutes. The base model is frozen: only the decision head and the suffix adapter are trained, then the head's
temperature and a conformal threshold are fitted on a held-out development set.
Limitations
- Needs vLLM 0.30.0 with the Plumbify plugin, on one GPU (no tensor or pipeline parallelism).
- The results above are for decisions stated in the prompt. Decisions the model hands off mid-generation use the same head but see the whole conversation as context, which training did not include.
- The training decisions are in English. The plumb chooses among the options it is given; it is not a safety classifier and does not check that the options make sense.
License
Apache 2.0, inherited from google/gemma-4-12B-it (terms).
Citation
@misc{plumbify2026,
title = {Plumbify: a System 1 decision branch for open language models},
author = {Kamalakannan, Naveenraj},
year = {2026},
url = {https://github.com/therealnaveenkamal/plumbify}
}
- Downloads last month
- 267
Model tree for totum-labs/gemma-4-12B-it-plumb
Collection including totum-labs/gemma-4-12B-it-plumb
Evaluation results
- Accuracy (model alone on Plumbify held-out decisions (447)self-reported0.736
- Accuracy (plumbed on Plumbify held-out decisions (447)self-reported0.826
- Accuracy (model alone on Plumbify held-out decisions (447)self-reported0.770
- Accuracy (plumbed on Plumbify held-out decisions (447)self-reported0.872