Alberich-mini-v1 / DATA_SOURCES.md
wahnfried's picture
Release Alberich-mini-v1 adapter and full benchmark results
790a4b2 verified
|
Raw History Blame Contribute Delete
13.6 kB

Training sources, attributions and terms

Checked 2026-09-20 against the actual 100000-example training file and 6000-example continuation. Counts are presentations, not necessarily unique examples. Every training task is covered below. Source licenses are recorded separately from any license for the adapter. This is a provenance and terms inventory, not a declaration that every right has been cleared.

Source inventory

Dataset / task Main Repair Attribution Source terms
WinoGrande (winogrande) 20000 600 Sakaguchi et al.; Allen Institute for AI Apache-2.0
XNLI (German) (xnli_de) 13000 250 Conneau et al.; Facebook Research CC-BY-NC-4.0
ANLI (anli) 10000 500 Nie et al.; Facebook Research CC-BY-NC-4.0
MASSIVE (German) (massive_de) 4000 500 FitzGerald et al.; Amazon.com Inc. or its affiliates CC-BY-4.0
PAWS-X (German) (pawsx_de) 6000 200 Yang et al.; Google LLC Custom permissive dataset terms
HellaSwag (hellaswag) 8000 50 Zellers et al. MIT reported in pinned dataset card; upstream verification unavailable
PIQA (piqa) 5000 0 Bisk et al. AFL-3.0
Social IQa (social_iqa) 5000 50 Sap et al. CC-BY-4.0 (pinned source card)
BoolQ (boolq) 1006 0 Clark et al.; Google Research; Wikipedia contributors CC-BY-SA-3.0 (Google dataset card)
Stanford Sentiment Treebank / SST-2 via GLUE (sst2) 1000 0 Socher et al.; Stanford NLP; GLUE maintainers Unresolved for original dataset in reviewed sources
A-OKVQA via The Cauldron (aokvqa) 6000 100 Schwenk et al.; Allen Institute for AI Apache-2.0 repository; COCO images have separate rights
IconQA via The Cauldron (iconqa) 5000 100 Lu et al. CC-BY-NC-SA-4.0 (dataset)
ScienceQA via The Cauldron (scienceqa) 3000 100 Lu et al. CC-BY-NC-SA-4.0 (dataset); code separately MIT
AI2D via The Cauldron (ai2d) 1994 100 Kembhavi et al.; Allen Institute for AI CC-BY-SA-4.0 (AI2-managed AWS registry entry)
Visual Spatial Reasoning via The Cauldron (vsr) 1000 50 Liu, Emerson and Collier Apache-2.0 project; underlying COCO images separate
TextVQA via The Cauldron (textvqa_mc) 2000 150 Singh et al.; Facebook AI Research CC-BY-4.0 dataset; OpenImages/Flickr image-source rights
DocVQA via The Cauldron (docvqa_mc) 2000 150 Mathew, Karatzas and Jawahar; UCSF Industry Documents Library (documents) Original challenge download terms; not verified
GUM via CorefUD 1.4 / CRAC 2026 (corefud_en) 3368 400 Zeldes; GUM annotators and original text authors; CorefUD team CC-BY-NC-SA-4.0 collection; constituent text licenses vary
Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026 (corefud_de) 632 200 Bourgonje and Stede; Potsdam corpus contributors; newspaper authors; CorefUD team CC-BY-NC-SA-4.0 (license linked by original corpus page)
Locally generated coref_0 (coref_0) 88 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated coref_1 (coref_1) 176 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated coref_2 (coref_2) 192 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated coref_3 (coref_3) 160 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated coref_4 (coref_4) 176 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated coref_5 (coref_5) 208 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated synthetic_dialogue (synthetic_dialogue) 200 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated synthetic_numeric (synthetic_numeric) 100 1500 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated synthetic_ocr (synthetic_ocr) 500 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated synthetic_rules (synthetic_rules) 100 0 This project Project-authored templates; no separate third-party dataset grant asserted
Locally generated synthetic_temporal (synthetic_temporal) 100 1000 This project Project-authored templates; no separate third-party dataset grant asserted

Requirements, citations and changes

All external examples were selected, normalized into the project question/option format and mixed for supervised adapter training. Labels were mapped to answer options. CorefUD annotations were converted into contextual coreference questions. TextVQA and DocVQA answers were converted to multiple-choice questions with distractors. Images were processed with the model image processor; the benchmark additionally uses thumbnails. These modifications are ours, not endorsed by the dataset providers.

WinoGrande

Preserve license and applicable copyright/NOTICE notices; identify changes. Preserved notice. Citation / original work.

XNLI (German)

Attribution, license link, change notice; noncommercial restriction. Training translations derive from MultiNLI. Preserved notice. Citation / original work.

ANLI

Attribution, license link, change notice; noncommercial restriction. Preserved notice. Citation / original work.

MASSIVE (German)

Retain Amazon attribution, license link and change notice. Preserved notice. Citation / original work.

PAWS-X (German)

Google source acknowledgement requested; preserve supplied disclaimer. Do not mislabel as Apache. Preserved notice. Citation / original work.

HellaSwag

Original repository was unavailable during this audit. Original copyright notice and underlying source-text terms remain to be verified. Citation / original work.

PIQA

Original homepage explicitly names AFL-3.0, overriding unknown in the cached mirror card. Preserve attribution/change notices and applicable source/distribution obligations; see full terms. Preserved notice. Citation / original work.

Social IQa

Attribution, license link and change notice. Preserved notice. Citation / original work.

BoolQ

Attribution, license link, modification notice and share-alike for covered adaptations; retain underlying Wikipedia source credits. Preserved notice. Citation / original work.

Stanford Sentiment Treebank / SST-2 via GLUE

GLUE explicitly refers to original task licenses. Do not apply the GLUE code license to SST-2 data. Citation / original work.

A-OKVQA via The Cauldron

Preserve license/notices. This repository license does not establish permission for every underlying photograph. Preserved notice. Citation / original work.

IconQA via The Cauldron

Attribution, noncommercial use, change notice and share-alike for covered adaptations. Underlying icons/images may have additional terms. Preserved notice. Citation / original work.

ScienceQA via The Cauldron

Use dataset terms, not the code license. Attribution, noncommercial use, change notice and share-alike for covered adaptations. Preserved notice. Citation / original work.

AI2D via The Cauldron

Attribution, license link, change notice and share-alike for covered adaptations. Data acquired through The Cauldron; original AI2 registry/terms accessed 2026-09-20. Individual diagram provenance not reconstructed. Preserved notice. Citation / original work.

Visual Spatial Reasoning via The Cauldron

Preserve license/notices and changes; retain separate image-source obligations. Preserved notice. Citation / original work.

TextVQA via The Cauldron

Attribution, license link and modification notice. Original answers converted to choices with distractors; individual image-source attribution is not reconstructed here. Preserved notice. Citation / original work.

DocVQA via The Cauldron

Original site requires review of RRC download terms. Portal unavailable during audit; no blanket permissive license asserted. Individual document rights remain separate. Citation / original work.

GUM via CorefUD 1.4 / CRAC 2026

Credit original text sources and annotators. Selected document IDs, authors and source URLs are listed in coref_attributions.json. Document-level license verification remains incomplete. Preserved notice. Citation / original work.

Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026

Credit PCC and original newspaper sources; noncommercial/share-alike conditions. Preserve notices and identify transformations. Preserved notice. Citation / original work.

Collection and component credits

The Cauldron: HuggingFaceM4 / Laurençon et al., What matters when building vision-language models?. Collection terms: component datasets retain their own conditions; rights in collection-authored prompts are offered under CC BY 4.0. This does not relicense source photographs/documents.

CorefUD: Nedoluzhko, Novák, Popel, Žabokrtský, Zeldes and Zeman (2022), CorefUD 1.0: Coreference Meets Universal Dependencies; data version 1.4 / CRAC 2026 gold. Only GUM and PotsdamCC were used here, confirmed from the actual rows. GUM annotator credits, per-document attribution metadata. No Reddit documents occur in the selected GUM rows.

GLUE: Wang et al. (2019), GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. SST-2 uses the original Stanford dataset; the collection explicitly points to original task terms.

Image components: A-OKVQA and VSR use COCO; TextVQA uses Open Images. Their underlying images may retain photographer-specific rights. Individual-image rights and credits have not been fully reconstructed from The Cauldron shards. DocVQA documents originate in the UCSF Industry Documents Library. No source images or document text are redistributed here.

Evaluation-only sources

POPE and the SuperGLUE COPA, WiC and WSC subsets were used for evaluation, not adapter training. They remain in the broader source_provenance.json inventory so the benchmark is traceable; they are not additional training sources. Evaluation results are aggregates, not redistributions of these datasets.

Remaining gaps

DocVQA download terms, the original SST-2 grant, HellaSwag original notices and individual text/image rights are not fully resolved. Generic attribution does not settle those gaps. We do not claim commercial clearance or that an NC/SA dataset license automatically attaches to model weights. Full license texts and preserved notices accompany this inventory where available. See LICENSING.md.