Primus Decision 0.1 — Research Alpha
A compact, locally runnable model for typed probabilistic decisions.
Primus Decision 0.1 is AAME's first public Primus research component. Give it a structured JSON state and a supported typed question, and it returns a probability for each valid answer. Its two non-transformer encoders use 3,715,074 neural parameters, run on a CPU, and make no LLM call at inference.
The released model records 75.1% accuracy on 2,000 Typed Decisions test decisions. A separate invoice study records 26/32 correct structured reconciliation answers (81.25%) in its broader synthetic test, following 32/32 in the initial pilot. Each result has its own task, data and evaluation conditions; the public evidence documents what this version demonstrates and gives researchers a starting point for further work.
Built by Sai Tilak Pally at AAME, Hyderabad, India. Model artifacts, runtime and release documentation: Apache-2.0.
Website · GitHub · Invoice case study · Evidence download and reproduction · Benchmarks
What you can build and test
The public interface supports five question schemas in each of four workflows: invoice processing, customer service, security incidents and agent-trace observability.
| Question type | Output |
|---|---|
noul |
P(false) and P(true) |
choice |
One probability per supplied supported option |
score |
One probability per ordinal level and the expected level |
Workflow and question IDs identify the supported schemas. For a supported choice question, callers can supply a subset of its trained option keys in any order and provide question and option descriptions. The model normalizes over those supplied options. The complete schemas and option keys are in model/member0/config.json; the request and response shapes are in INTERFACE.md. The runtime validates this interface before inference.
Use the examples to integrate typed decisions into a local application, inspect confidence, and record outcomes in the published Decision History format. Evaluate the application's own records and decision criteria when adapting the component to a new use case.
Quickstart
CPU inference, Python 3.11+. Download from the public Hub; inference needs no hosted API or API key.
python3 -m venv .venv && . .venv/bin/activate
pip install -U huggingface_hub && hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1
cd primus-decision-0.1 && grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \
python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && \
python examples/predict_example.py
Every checked file should print OK. On Linux, use sha256sum -c - in place of shasum -a 256 -c -. The two LSA feature files use Python pickle; load them only from a trusted release and verify their hashes against its manifest. Hashes confirm file integrity, not the safety of an arbitrary source.
The example returns three distributions for an invoice whose amount exceeds its purchase order by 20%:
{"duplicate": {"probabilities": {"false": 0.986, "true": 0.014}},
"disposition": {"probabilities": {"approve": 0.002, "hold": 0.206, "manual_review": 0.708, "reject": 0.084}},
"discrepancy_severity": {"probabilities": {"0": 0.043, "1": 0.041, "2": 0.118, "3": 0.798}, "score": 2.672}}
Run Python from the downloaded model folder:
from primus_decision.ensemble import Ensemble
from primus_decision.predict import request_to_case, answers_from_predictions
model = Ensemble.load("model") # load once
case = request_to_case(state, questions, workflow="invoice_processing")
decisions, predictions = model.predict_cases([case])
answers = answers_from_predictions(decisions, predictions)
state and questions follow the runnable example and interface specification. To batch requests:
cases = [request_to_case(s, q, workflow=w) for s, q, w in requests]
decisions, predictions = model.predict_cases(cases)
answers, i = [], 0
for c in cases:
n = len(c.decisions)
answers.append(answers_from_predictions(decisions[i:i + n], predictions[i:i + n]))
i += n
python -X utf8 -m primus_decision.serve model starts the JSON-lines bridge: one request per line in and one response per line out.
Typed Decisions benchmark
Official test split of LocalLLaMA/typed-decisions, pinned to revision ea9306458d6e9563628369a3d1e72e362fb381d2: 400 cases, 2,000 decisions, one sealed run, raw probabilities. Primus and Laya use the task-trained specialist setting.
| Model | Accuracy | Brier | ECE (raw) | NLL | Score MAE | Neural parameters |
|---|---|---|---|---|---|---|
| Primus Decision 0.1 — measured by AAME | 0.751 | 0.059 | 0.127 | 0.637 | 0.275 | 3.7M |
| Laya — reproduced by AAME on the same harness | 0.766 | 0.066 | 0.213 | 0.707 | 0.242 | ~421M |
| Laya — published model card | 0.766 | 0.062 | 0.213 | — | 0.242 | ~421M |
Primus combines 75.1% accuracy with lower recorded NLL, raw ECE and Brier than the reproduced Laya checkpoint, using about 113× fewer neural parameters. Laya records 76.6% accuracy and the lower ordinal score error. The parameter ratio describes the neural networks; Primus's separate fitted feature tables and total runtime footprint are reported below.
Laya's published soft accuracy and Brier differ from the reproduced values under both its author's harness and AAME's: 0.509 and 0.066, respectively. The reproduced row is the comparison used here. Laya's public model card
Primus's 95% case-bootstrap accuracy interval is 72.9–77.4%, describing case-sampling uncertainty for this frozen run. It is not a test of equivalence between models. The protocol was fixed before the test split was opened: its accuracy target above 76.6% was not reached; its Brier/ECE target was reached. The original protocol and sealed record remain available in PROTOCOL.md and SEALED_RESULT.json.
The benchmark uses synthetic scenarios and teacher-generated reference labels. Its accuracy measures agreement with those references. These observed scores describe this evaluation; they do not establish an upper bound on future Primus versions or on performance in another application.
The dataset card also reports generalists such as TypeSafe Jev and meraGPT Decider. Their zero-shot setting and training histories differ from this specialist comparison; the full tables preserve those separately attributed results.
By workflow and question type
| Workflow | Primus accuracy | Reproduced Laya accuracy |
|---|---|---|
| Invoice processing | 0.808 | 0.804 |
| Security incidents | 0.736 | 0.766 |
| Agent-trace observability | 0.732 | 0.730 |
| Customer service | 0.728 | 0.764 |
| Question type | Primus accuracy | Reproduced Laya accuracy |
|---|---|---|
noul — yes/no probability |
0.837 | 0.857 |
choice — one of N |
0.738 | 0.733 |
score — ordinal level |
0.696 | 0.723 |
There are 500 decisions per workflow, grouped into cases. Use the recorded per-task results and case-level uncertainty when assessing a comparison.
Invoice case study
AAME evaluated the released weights without retraining on generated invoice records with arithmetic and membership labels. The broader study has 32 new invoices, presented as both structured records and narratives. Each task contains 16 true and 16 false cases.
| Presentation | Task | Primus 0.1 | TF-IDF classifier | Qwen token scoring | Qwen free answer | Explicit rules |
|---|---|---|---|---|---|---|
| Structured | Reconciliation | 26/32 | 27/32 | 16/32 | 16/32 | 32/32 |
| Structured | Duplicate ID | 16/32 | 16/32 | 18/32 | 18/32 | 32/32 |
| Narrative | Reconciliation | 17/32 | 24/32 | 18/32 | 18/32 | 32/32 |
| Narrative | Duplicate ID | 16/32 | 16/32 | 16/32 | 14/32 | 32/32 |
The initial, simpler pilot recorded 32/32 reconciliation answers for both the supported wording and a paraphrase. Structured reconciliation is the strongest Primus result in the broader test; input representation is a clear research direction from the paired structured/narrative results.
Primus and the classifier retain their original task supervision and were not fitted on these evaluation invoices. Qwen3-1.7B is a standalone competitor using its official non-thinking configuration; no Qwen output enters Primus. These task results describe the specified model configurations. The study includes all comparison conditions, a six-condition Qwen answer-format audit, matched-pair results and public-model evaluation scripts. The broader test caps generic derived-relation lines at 6 for both presentations, compared with the pilot default of 40, so changes between experiments reflect both cases and preprocessing.
Read the complete case study · Download and verify the evidence
Architecture and operating profile
Two members contribute equally to every prediction: a diagonal state-space S4D encoder (1,932,609 parameters) and a bidirectional GRU (1,782,465). Each combines flattened state text, deterministic derived-relation features and a 256-dimensional LSA case vector. Question-conditioned pooling and option scoring produce the typed answer distribution.
Both encoders were trained from scratch on the official training split: 1,005 fitting cases and 195 validation cases. Primus uses no pretrained language encoder or transformer. The original training labels are teacher-generated; inference runs entirely in the released local model.
| Item | Recorded value |
|---|---|
| Neural parameters | 3,715,074 |
| Fitted feature tables | Two LSA featurizers, 71.4 MB each |
| Installed model release | 157.9 MB |
| Warm CPU latency | 115–162 ms p50 per five-decision case across six recorded runs |
| Batched latency | 36–65 ms per case at batch size 32 |
| Resident memory after loading | 0.5–0.6 GB |
| Peak memory during recorded inference | 1.4–1.6 GB |
Timing uses a 4-vCPU Intel Xeon 2.10 GHz with two threads and the model loaded. Measure the published method on your own hardware. Compact structured states fit the tested input path; the encoder processes up to 640 tokens.
The robustness study recorded unchanged answers under choice-option shuffling and a 0.1-point accuracy change under instruction paraphrases. Renaming field keys to synonyms or camelCase changed accuracy by −4.1 and −5.4 points. Preserve or evaluate the application's field conventions as part of integration. Recorded experiments
Raw probabilities are the default. On the benchmark, the most-confident 50% of decisions achieved 90.9% accuracy. The optional validation-fitted temperature profile changes ECE from 0.127 to 0.040 and NLL from 0.637 to 0.568, alongside Brier from 0.059 to 0.114 and score MAE from 0.275 to 0.320. Choose confidence thresholds using outcomes from the intended application, with human review for consequential decisions.
Verify and extend
The model artifacts remain the frozen Decision 0.1 release. Use SHA256SUMS to verify downloaded files and PROVENANCE.json for the original model hashes, dataset revision and training configuration. The Hub's README is this model card; its .gitattributes and config.json serve Hub packaging. The frozen inference artifacts match GitHub tag primus-decision-0.1.
AAME checked bit-identical outputs between the public and internal frozen package on 600 decisions. A separately written metric implementation used by AAME reproduced the sealed numerical record within approximately 1e-16. These are implementation checks with published provenance.
The runtime, model artifacts, interface, architecture description, training configuration and recorded evidence are public. The invoice evidence includes its public-model evaluation scripts. For the original Typed Decisions benchmark, REPRODUCIBILITY.md specifies how to combine the pinned dataset with the public runtime and metrics; the original training pipeline and evaluation driver remain private.
Researchers can run new evaluations, modify the Apache-2.0 release and publish reproducible extensions. Give changed models or preprocessing a distinct version, retain the original result record, and report the data, conditions and relevant baselines. Share a reproduction or extension
The Road to Primus
Our longer-term goal is a broader architecture connecting persistent memory, recurrent graph reasoning, temporal state and history, salience and priority, imagination and simulation, consolidation and learning cycles, language and grounding, planning and decision layers, and agent and tool interfaces.
Decision 0.1 is the first public research alpha in that programme. The roadmap describes intended research directions; the released component demonstrates the typed decisions and measured results documented here.
Public vs Protected
We publish architecture descriptions, interfaces, benchmarks and reproducibility evidence for released components so their claims can be examined. Unreleased components, training methods, system integration details, internal datasets and implementation specifics remain private until AAME chooses to publish them. Public boundary
Frequently asked questions
What is Primus Decision 0.1? AAME's first public Primus component: a 3.7M-neural-parameter model that turns structured states and supported questions into typed probability distributions.
Does Primus need an LLM or cloud API to run? Inference runs locally on a CPU using the released S4D/GRU ensemble and feature tables. No LLM or hosted inference API is called. The original benchmark training labels were teacher-generated.
Can I try new inputs and question descriptions? Yes. Supply new records and descriptions within the supported workflow/question IDs and trained option keys, then evaluate the resulting answers. New schemas require an adapted model/interface and their own evaluation.
Is 75.1% the maximum Primus can achieve? It is the measured accuracy of the frozen 0.1 model on one named benchmark. The invoice study reports different task-specific scores; future experiments and versions must establish their own results.
Can I improve or extend the released model? Yes. Apache-2.0 permits use and modification of the published artifacts and runtime under its terms. Publish the changes and evaluation conditions with your result; AAME's private training pipeline is separate from that release.
Is Decision 0.1 the complete Primus architecture? It is the first public decision component. The broader architecture is the research programme described above.
Citation
@software{pally2026primusdecision,
author = {Pally, Sai Tilak and {AAME}},
title = {Primus Decision 0.1 (Research Alpha): an open 3.7M-parameter non-transformer model for typed probabilistic decisions},
year = {2026},
version = {0.1},
license = {Apache-2.0},
url = {https://github.com/pally-sai-tilak/primus-decision}
}
The Apache-2.0 license covers the released artifacts, runtime and documentation. Primus and AAME names and logos are trademarks and are not licensed.
- Downloads last month
- 18
Dataset used to train The-Aame/primus-decision-0.1
Evaluation results
- Accuracy (raw probabilities, one sealed run) on Typed Decisionstest set self-reported0.751
- Brier on Typed Decisionstest set self-reported0.059
- ECE (15 bins, raw) on Typed Decisionstest set self-reported0.127
- NLL on Typed Decisionstest set self-reported0.637
- Score MAE (ordinal questions) on Typed Decisionstest set self-reported0.275