privacy-gate-llm

1024 weights and a bias. A logistic head on top of BAAI/bge-m3 that answers one question about a piece of text:

Must this stay on this machine?

It catches health information, credentials and personal data written as ordinary English — the kind that pattern rules and secret scanners cannot see.

pip install privacy-gate && ollama pull bge-m3
privacy-gate check "the woman from Tuesday's clinic has a 7mm lesion on her shoulder"
# hold  score=6.7200 threshold=0.1209 margin=6.5991

On data other people made (run 12, the head unchanged): on an OpenShift AI router's evaluation corpus the head catches 0.70 of the prompts that must stay local at 0.14 friction, and 0.93 of the 60 prompts written with no marker to match, where Presidio catches 0.00; on a PII benchmark's formal sentences it over-holds (friction 0.45). The 99 % below is an in-domain number.

Current head: head-v1.json. head-v0.json is kept so earlier published numbers stay reproducible; use v1.

Why this exists

If you send prompts to a hosted model from a machine that also holds health data, you want something between the two. The usual answer is a regex ruleset. On this project's 206-example gold set, that ruleset catches 10 of 111 sensitive examples, 9.0%. The other 101 are the same information written as prose, and prose is what people actually type.

Results

Five-fold cross-validated: every score comes from a head that never saw that example.

AUC catch friction
regex ruleset alone — 9.0% 1.1%
hashed word unigrams, no model 0.7818 98% 83.2%
Qwen3.5-0.8B, prompted 0.4609 at chance
bge-m3 + logistic head, v1 0.9927 99.1% 5.3%
  • catch — of sensitive texts, the share held back. A miss is an incident.
  • friction — of ordinary texts, the share held back. A guard that fires on ordinary work gets switched off, and then it protects nothing.

Held out of training

v0 v1
21 short prompts: harmless held 8/15 0/15
21 short prompts: sensitive caught 6/6 6/6
20 sentences written after training: harmless held 2/9 1/9
20 sentences written after training: sensitive caught 11/11 11/11

v0 held ok, yes, continue and thanks!, because nothing that short was in its training data. v1 adds short prompts to training — different ones from the held-out probe — and short sensitive ones as well, so that brevity does not become a way to leak.

The negative results are the interesting part

Prompting a small chat model does not work, and looks like it does. Asked for a four-way label, Qwen3.5-0.8B answered HEALTH for 106 of 127 examples and never once SECRET or PII. Scored on logP(KEEP) − logP(SEND) it lands at AUC 0.4609, at chance and worse than a bag of words.

Always run a no-model baseline, and audit for length. The first attempt at adding short prompts let length alone reach AUC 0.70; a head fitted on that could have learned that short means safe. Balancing the short rows brought it to 0.63.

Intended use

A second gate, behind whatever deterministic checks you already have. It may only ever add a hold, never clear one, so a false negative leaves your existing protection exactly as strong as it was.

Out of scope — please read

A send verdict is not an assurance that text is safe. It is a statistical classifier with a measured miss rate: at its threshold, 1 of 111 sensitive examples leaks in cross-validation.

This is not production-validated. Every example — 206 for training, 41 held out — was written by one person. That is enough to choose an architecture, not enough to set a threshold that decides what leaves a machine holding real patient data. Measure it on your own traffic first.

It is not a medical device, and it says nothing about what a lesion is.

English only.

Usage

import json
import numpy as np
from huggingface_hub import hf_hub_download
from sentence_transformers import SentenceTransformer

head = json.load(open(hf_hub_download("YauhenBichel/privacy-gate-llm", "head-v1.json")))
encoder = SentenceTransformer("BAAI/bge-m3")
w = np.array(head["weights"]); mu = np.array(head["mean"]); sd = np.array(head["stdev"])

def decide(text: str) -> bool:
    # normalize_embeddings=True is required, not cosmetic. See below.
    v = encoder.encode([text], normalize_embeddings=True)[0]
    return float(((v - mu) / sd) @ w + head["bias"]) > head["threshold"]

decide("her biopsy is booked for the 20th")                    # True
decide("reformat this YAML and sort the keys alphabetically")  # False
decide("ok")                                                   # False

Two things that will silently give you wrong answers

The embeddings must be L2-normalised. The head was fitted on unit vectors; sentence-transformers does not normalise unless asked. An un-normalised vector scores confidently wrong rather than raising.

The embeddings must be bge-m3. Another encoder of the same width would also produce confident nonsense.

Choosing a threshold

threshold catch friction
−1.455 100% 12.6%
+0.121 99.1% 5.3% (shipped)
+0.881 90.1% 2.1%, with 11 leaks

Friction is flat at 5.3% from 95% to 99% catch, so there is nothing to gain below the shipped point.

Training data

206 hand-written examples, 95 clean and 111 sensitive, including short acknowledgements and short sensitive prompts; plus two held-out files never used for fitting (20 sentences written after training, and 21 short prompts). Every example is invented. Phone numbers come from Ofcom's 07700 900xxx drama range, and the AWS key is Amazon's own documentation placeholder.

Citation

@software{bichel2026privacygate,
  author  = {Bichel, Yauhen},
  title   = {privacy-gate-llm: a 1024-weight head that catches personal data written as prose},
  year    = {2026},
  url     = {https://github.com/MoleCare/privacy-gate-llm},
  license = {Apache-2.0}
}

Licence

Apache-2.0. The base encoder, BAAI/bge-m3, is MIT and is neither included nor redistributed here — only the head is.

Built for MoleCare. MoleCare is not a medical device and does not diagnose.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YauhenBichel/privacy-gate-llm

Base model

BAAI/bge-m3
Finetuned
(569)
this model

Dataset used to train YauhenBichel/privacy-gate-llm

Space using YauhenBichel/privacy-gate-llm 1

Evaluation results

  • ROC AUC (5-fold cross-validated) on privacy-gate-llm gold set (206 examples)
    self-reported
    0.993
  • ROC AUC on OpenShift AI router "privacy-plus" corpus, English (332 prompts, not ours)
    self-reported
    0.867
  • catch at the shipped threshold (friction 0.135) on OpenShift AI router "privacy-plus" corpus, English (332 prompts, not ours)
    self-reported
    0.702
  • ROC AUC on piimb/pii-masking-benchmark, balanced 2,000-sentence sample
    self-reported
    0.866
  • catch at the shipped threshold (friction 0.454) on piimb/pii-masking-benchmark, balanced 2,000-sentence sample
    self-reported
    0.929