O5I
Purpose-embedded datasets for the AI era.
We publish refined versions of public datasets. Most public data ships with noisy, unverified, or undocumented labels — we
re-examine it and add the trust signals, quality audits, and routing metadata the raw data lacks, so teams can train on what's
reliable and route the rest to expert review.
Every dataset we publish documents its purpose, its limits, and what it teaches future AI.
Why we exist
- Public-interest counterweight. Data — medical, scientific, cultural, environmental — is being progressively enclosed:
gathered for free, processed behind paywalls, sold back at markup. We work the other direction: broadening access to data and
tools that should remain public, and documenting provenance so each dataset is verifiable public knowledge.
- Your labels become AI's reasoning. A label is a signal embedded in every model trained on it. We attach a traceable reason
to each label — what AI inherits is an instrument with an audit trail, not just a verdict.
- Open and auditable. We build on permissively-licensed public data and open models, so our work stays redistributable and
contestable — usable beyond the gated AI stack.
- 1% today, 10% tomorrow. 1% of every sale is donated to organizations advancing humanity and life, growing toward 10% as
revenue stabilizes.
Datasets
- cxr14-bpals-trial — an independent label-quality signal over NIH
ChestX-ray14: which existing labels to trust, and where to focus expert review. Free evaluation sample.
→ o5i.io