codybum commited on
Commit
20c67a3
·
verified ·
1 Parent(s): 0aa2095

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -5
README.md CHANGED
@@ -123,8 +123,7 @@ Token counts are char-estimates against the packed cache.
123
  | Hindawi OA journals | 2.93 B | | s2orc | 1.21 B |
124
  | MeDAL (PubMed abstracts) | 2.49 B | | | |
125
 
126
- **+ ~30 smaller sources:** clinical narratives (open_patients, mimic-iv-ed, mimic-cxr, ctrate, coral, tcga_reports,
127
- mts_dialog), knowledge / guidelines (stackexchange-science 0.27 B, dailymed, cpg, gene_ontology, medlineplus,
128
  orphapacket, medmentions, trialgpt), pharmacovigilance / relational rendered to NL (ctd, faers, aeolus, onsides,
129
  sider, cdc_places, cbioportal, civic, ade_corpus_v2), and deliberate register-diversity (locus legal-code).
130
 
@@ -134,9 +133,6 @@ license verification; they will be added once resolved.)*
134
  **Disclosed issue:** ~35% of tokens are duplicates (a FineWeb-Edu sharding build bug + PMC repetition); a deduped
135
  corpus (`v4_dedup`, 79.4 M docs, 0% dup) is ready but was **not** trained — this release is the original single-epoch corpus.
136
 
137
- **Governance:** includes PhysioNet-credentialed, de-identified, redistribution-restricted clinical sources (MIMIC-IV
138
- discharge / radiology, MIMIC-CXR) — any downstream use must confirm PhysioNet DUA compliance.
139
-
140
  ## Evaluation — 19-test suite, 75 metrics, vs a 16-model fleet
141
  KOS-V4-Base is benchmarked as a **from-scratch, single-epoch base against 16 external models trained on 1.7–200× more
142
  data** (0.3–36 T tokens). Pool of 17 models, 95 ranked metrics. BPB (bits-per-byte) is tokenizer-agnostic and included
 
123
  | Hindawi OA journals | 2.93 B | | s2orc | 1.21 B |
124
  | MeDAL (PubMed abstracts) | 2.49 B | | | |
125
 
126
+ **+ ~30 smaller sources:** open clinical narratives, knowledge / guidelines (stackexchange-science 0.27 B, dailymed, cpg, gene_ontology, medlineplus,
 
127
  orphapacket, medmentions, trialgpt), pharmacovigilance / relational rendered to NL (ctd, faers, aeolus, onsides,
128
  sider, cdc_places, cbioportal, civic, ade_corpus_v2), and deliberate register-diversity (locus legal-code).
129
 
 
133
  **Disclosed issue:** ~35% of tokens are duplicates (a FineWeb-Edu sharding build bug + PMC repetition); a deduped
134
  corpus (`v4_dedup`, 79.4 M docs, 0% dup) is ready but was **not** trained — this release is the original single-epoch corpus.
135
 
 
 
 
136
  ## Evaluation — 19-test suite, 75 metrics, vs a 16-model fleet
137
  KOS-V4-Base is benchmarked as a **from-scratch, single-epoch base against 16 external models trained on 1.7–200× more
138
  data** (0.3–36 T tokens). Pool of 17 models, 95 ranked metrics. BPB (bits-per-byte) is tokenizer-agnostic and included