gemma-4-12B-it-plumb

google/gemma-4-12B-it with a plumb: a decision head built into the model. It answers typed decisions (pick one of these options, yes or no, a score on a scale) in one forward pass, with a calibrated probability for every option, reading the KV cache the conversation already filled. It is a Jev-style decision model inside the LLM instead of next to it. The base weights are unchanged, so the model chats, reasons and calls tools exactly like google/gemma-4-12B-it.

Made with Plumbify and served by vLLM through the Plumbify plugin.

Results

447 held-out decisions (task families and templates the plumb never trained on), on one vLLM server. Each cell is accuracy (latency p50 for one request). With thinking on, the model alone is scored in its better prompt setup, so the plumb is compared with the model at its best.

google/gemma-4-12B-it alone gemma-4-12B-it-plumb
Thinking off 0.736 (98 ms) 0.826 (2.2 s)
Thinking on 0.770 (24.1 s) 0.872 (5.0 s)
Eval set Off: alone Off: plumbed On: alone On: plumbed
DecisionBench, held-out task families (150) 0.833 0.833 0.793 0.840
Hard-skill templates, held out (105) 0.600 0.810 0.771 0.867
General decisions, unseen templates (192) 0.734 0.828 0.750 0.901

Use

pip install "plumbify[vllm]"
vllm serve totum-labs/gemma-4-12B-it-plumb

In a second terminal, chat with it and watch each decision as the model hands it to its plumb:

plumbify chat

Or ask it through the OpenAI API:

curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{
  "model": "totum-labs/gemma-4-12B-it-plumb",
  "messages": [{"role": "user", "content": "Ticket: I was charged twice.\n\nWhich team should handle this?\n- billing\n- technical\n- sales"}]
}' | jq .plumb.answer

Decisions stated in a message go to the plumb before the model replies; the model can also hand decisions to it mid-conversation through a plumb_decide tool, answered from the KV cache. Each response carries a plumb field with the probabilities, a conformal set and the final answer. "plumb": false in a request serves it with the base model alone. Request options and response fields are in the guide.

What is in this repository

  • The weights of google/gemma-4-12B-it, unchanged.
  • config.json with the architecture PlumbGemma4UnifiedForConditionalGeneration, which tells vLLM and the plugin to attach the plumb.
  • head.safetensors: the decision head (60.6M parameters). It reads the model's hidden states at layers 23, 29, 35, 41 and the final output.
  • suffix_adapter.*: a LoRA (rank 32) on o_proj, k_proj, down_proj, up_proj, q_proj, v_proj, gate_proj, active only while the model reads a decision. It is never merged.
  • plumb.json: taps, head shape and calibration (temperature 1.122, conformal threshold 0.794 at alpha 0.1).

Training

plumbify train on 10k decision rows drawn from about 200 decision families (support, finance, coding, safety, medicine, law, engineering) plus solver-labelled hard-skill rows, one epoch, one A100 80 GB, 76 minutes. The base model is frozen: only the decision head and the suffix adapter are trained, then the head's temperature and a conformal threshold are fitted on a held-out development set.

Limitations

  • Needs vLLM 0.30.0 with the Plumbify plugin, on one GPU (no tensor or pipeline parallelism).
  • The results above are for decisions stated in the prompt. Decisions the model hands off mid-generation use the same head but see the whole conversation as context, which training did not include.
  • The training decisions are in English. The plumb chooses among the options it is given; it is not a safety classifier and does not check that the options make sense.

License

Apache 2.0, inherited from google/gemma-4-12B-it (terms).

Citation

@misc{plumbify2026,
  title  = {Plumbify: a System 1 decision branch for open language models},
  author = {Kamalakannan, Naveenraj},
  year   = {2026},
  url    = {https://github.com/therealnaveenkamal/plumbify}
}
Downloads last month
267
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for totum-labs/gemma-4-12B-it-plumb

Adapter
(101)
this model

Collection including totum-labs/gemma-4-12B-it-plumb

Evaluation results

  • Accuracy (model alone on Plumbify held-out decisions (447)
    self-reported
    0.736
  • Accuracy (plumbed on Plumbify held-out decisions (447)
    self-reported
    0.826
  • Accuracy (model alone on Plumbify held-out decisions (447)
    self-reported
    0.770
  • Accuracy (plumbed on Plumbify held-out decisions (447)
    self-reported
    0.872