Linda-RAID

English AI-generated-text detector tuned for the RAID benchmark. Part of the Linda family (Linda-RAID · Linda-Essay · Linda-Multi · Linda-Stylo).

Not a plain transformers model. Run it with linda_raid.py (below). The deberta/ folder is only one half of the detector: loaded on its own (e.g. pipeline(..., model=".../deberta")) it skips attack cleanup, the stylometric part, normalization and thresholds, and gives different, weaker scores — not Linda-RAID.

Scope, read first: this is a benchmark specialist. It is strong on the RAID generators, domains and attacks, and weak on recent frontier models and long formal essays (see Limitations). Do not use its verdict as evidence against a person.

How it works

Two models of different nature, averaged:

part what trained on
deberta/ DeBERTa-v3-large sequence classifier; the text is split into 256-token windows covering the whole document, score = mean logit margin (AI − human) ~590k texts of the RAID train split (8 domains, 11 generators, 4 decoding settings, 12 attacks; human texts of every attack labelled human), harder strata oversampled
stylometry/ ~230 hand-crafted style features (LightGBM) + hashed character / word / function-word n-grams (logistic regression); runs on CPU in milliseconds 200k texts of the RAID train split

Before scoring, the text is cleaned of character-level attacks (homoglyphs, zero-width characters, unusual whitespace) — canonical.py. Each part's score is z-normalized on held-out RAID human texts; the final score is the mean of the two z-scores. linda_config.json holds the normalization and two thresholds: 5% and 1% false positives on RAID human texts.

Usage

pip install torch transformers sentencepiece lightgbm scikit-learn scipy pyyaml numpy
python linda_raid.py --file essay.txt
from linda_raid import LindaRAID
det = LindaRAID()                       # GPU if available, else CPU
det.predict(["Some text to check ..."])
# [{'score': 0.09, 'verdict': 'human', 'deberta_z': 1.93, 'stylometry_z': -1.74}]

verdict: ai (above the 1%-FPR threshold), likely_ai (above 5%), human.

Results

RAID, local validation. 15% of RAID train source documents were held out: no text of those documents (human or generated) was used in training. 1,928 clean human texts + 40,000 AI texts from them, scored with the official raid.evaluate.run_evaluation (per-domain thresholds, FPR 5%):

model accuracy, all attacks no attack
DeBERTa-v3-large part 98.4 98.7
stylometry part 98.4 98.6
Linda-RAID 99.7 99.6

Remaining misses are mostly the paraphrase attack and base (non-chat) Mistral / Cohere outputs sampled without repetition penalty. Official leaderboard score: pending (submission Linda-RAID to liamdugan/raid).

Limitations

Out-of-distribution English sets (AUROC / TPR at 1% FPR, threshold set on each set's own human texts):

set Linda-RAID its DeBERTa part alone
essays (DAIGT / IvyPanda-LLM vs student essays) 96.9 / 62.3 99.4 / 92.6
frontier models (Claude 4.x, 2024–2025) 90.0 / 28.5 88.0 / 53.3
frontier models (Gemini 3, Claude 4.5, GLM) 86.9 / 3.5 89.3 / 3.5
Claude Opus 5.5 texts (2026) 82.7 / 13.3 90.4 / 50.0

On a small set of long (~3,000-word) competition essays with human and known-AI controls the model did not separate the known-AI control from human controls. For recent models and formal essays use Linda-Essay (in preparation). The stylometric part is the one that transfers worst; its RAID-specific n-grams do not describe modern model output.

Training data and license

License: CC BY-NC 4.0 — free for research and non-commercial use. Commercial license: lindapro.support@proton.me. Built on RAID (Dugan et al., ACL 2024; MIT) and microsoft/deberta-v3-large (MIT); their notices are in LICENSE.

@inproceedings{dugan-etal-2024-raid,
  title = {{RAID}: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors},
  author = {Dugan, Liam and Hwang, Alyssa and Trhl{\'i}k, Filip and Zhu, Andrew and Ludan, Josh Magnus and Xu, Hainiu and Ippolito, Daphne and Callison-Burch, Chris},
  booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics},
  year = {2024}
}

Contact: lindapro.support@proton.me · x.com/lindarcrusader

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lindarixon/Linda-RAID

Finetuned
(311)
this model

Dataset used to train Lindarixon/Linda-RAID