# Training sources, attributions and terms Checked 2026-09-20 against the actual 100000-example training file and 6000-example continuation. Counts are presentations, not necessarily unique examples. Every training task is covered below. Source licenses are recorded separately from any license for the adapter. This is a provenance and terms inventory, not a declaration that every right has been cleared. ## Source inventory | Dataset / task | Main | Repair | Attribution | Source terms | |---|---:|---:|---|---| | [WinoGrande](https://github.com/allenai/winogrande/blob/master/LICENSE) (`winogrande`) | 20000 | 600 | Sakaguchi et al.; Allen Institute for AI | Apache-2.0 | | [XNLI (German)](https://github.com/facebookresearch/XNLI/blob/main/LICENSE) (`xnli_de`) | 13000 | 250 | Conneau et al.; Facebook Research | CC-BY-NC-4.0 | | [ANLI](https://github.com/facebookresearch/anli/blob/main/LICENSE) (`anli`) | 10000 | 500 | Nie et al.; Facebook Research | CC-BY-NC-4.0 | | [MASSIVE (German)](https://huggingface.co/datasets/AmazonScience/massive) (`massive_de`) | 4000 | 500 | FitzGerald et al.; Amazon.com Inc. or its affiliates | CC-BY-4.0 | | [PAWS-X (German)](https://github.com/google-research-datasets/paws/blob/master/LICENSE) (`pawsx_de`) | 6000 | 200 | Yang et al.; Google LLC | Custom permissive dataset terms | | [HellaSwag](https://huggingface.co/datasets/Rowan/hellaswag/blob/218ec52e09a7e7462a5400043bb9a69a41d06b76/README.md) (`hellaswag`) | 8000 | 50 | Zellers et al. | MIT reported in pinned dataset card; upstream verification unavailable | | [PIQA](https://yonatanbisk.com/piqa/) (`piqa`) | 5000 | 0 | Bisk et al. | AFL-3.0 | | [Social IQa](https://huggingface.co/datasets/allenai/social_i_qa/blob/537a2ec8ec565adc0b70b70752893e59e024df26/README.md) (`social_iqa`) | 5000 | 50 | Sap et al. | CC-BY-4.0 (pinned source card) | | [BoolQ](https://huggingface.co/datasets/google/boolq) (`boolq`) | 1006 | 0 | Clark et al.; Google Research; Wikipedia contributors | CC-BY-SA-3.0 (Google dataset card) | | [Stanford Sentiment Treebank / SST-2 via GLUE](https://nlp.stanford.edu/sentiment/) (`sst2`) | 1000 | 0 | Socher et al.; Stanford NLP; GLUE maintainers | Unresolved for original dataset in reviewed sources | | [A-OKVQA via The Cauldron](https://github.com/allenai/aokvqa/blob/main/LICENSE) (`aokvqa`) | 6000 | 100 | Schwenk et al.; Allen Institute for AI | Apache-2.0 repository; COCO images have separate rights | | [IconQA via The Cauldron](https://github.com/lupantech/IconQA#license) (`iconqa`) | 5000 | 100 | Lu et al. | CC-BY-NC-SA-4.0 (dataset) | | [ScienceQA via The Cauldron](https://github.com/lupantech/ScienceQA#-licenses) (`scienceqa`) | 3000 | 100 | Lu et al. | CC-BY-NC-SA-4.0 (dataset); code separately MIT | | [AI2D via The Cauldron](https://registry.opendata.aws/allenai-diagrams/) (`ai2d`) | 1994 | 100 | Kembhavi et al.; Allen Institute for AI | CC-BY-SA-4.0 (AI2-managed AWS registry entry) | | [Visual Spatial Reasoning via The Cauldron](https://github.com/cambridgeltl/visual-spatial-reasoning) (`vsr`) | 1000 | 50 | Liu, Emerson and Collier | Apache-2.0 project; underlying COCO images separate | | [TextVQA via The Cauldron](https://huggingface.co/datasets/facebook/textvqa) (`textvqa_mc`) | 2000 | 150 | Singh et al.; Facebook AI Research | CC-BY-4.0 dataset; OpenImages/Flickr image-source rights | | [DocVQA via The Cauldron](https://site.docvqa.org/datasets) (`docvqa_mc`) | 2000 | 150 | Mathew, Karatzas and Jawahar; UCSF Industry Documents Library (documents) | Original challenge download terms; not verified | | [GUM via CorefUD 1.4 / CRAC 2026](https://github.com/UniversalDependencies/UD_English-GUM/blob/master/LICENSE.txt) (`corefud_en`) | 3368 | 400 | Zeldes; GUM annotators and original text authors; CorefUD team | CC-BY-NC-SA-4.0 collection; constituent text licenses vary | | [Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026](https://angcl.ling.uni-potsdam.de/resources/pcc.html) (`corefud_de`) | 632 | 200 | Bourgonje and Stede; Potsdam corpus contributors; newspaper authors; CorefUD team | CC-BY-NC-SA-4.0 (license linked by original corpus page) | | Locally generated coref_0 (`coref_0`) | 88 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated coref_1 (`coref_1`) | 176 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated coref_2 (`coref_2`) | 192 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated coref_3 (`coref_3`) | 160 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated coref_4 (`coref_4`) | 176 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated coref_5 (`coref_5`) | 208 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated synthetic_dialogue (`synthetic_dialogue`) | 200 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated synthetic_numeric (`synthetic_numeric`) | 100 | 1500 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated synthetic_ocr (`synthetic_ocr`) | 500 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated synthetic_rules (`synthetic_rules`) | 100 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted | | Locally generated synthetic_temporal (`synthetic_temporal`) | 100 | 1000 | This project | Project-authored templates; no separate third-party dataset grant asserted | ## Requirements, citations and changes All external examples were selected, normalized into the project question/option format and mixed for supervised adapter training. Labels were mapped to answer options. CorefUD annotations were converted into contextual coreference questions. TextVQA and DocVQA answers were converted to multiple-choice questions with distractors. Images were processed with the model image processor; the benchmark additionally uses thumbnails. These modifications are ours, not endorsed by the dataset providers. ### WinoGrande Preserve license and applicable copyright/NOTICE notices; identify changes. [Preserved notice](NOTICE-WINOGRANDE.txt). [Citation / original work](https://arxiv.org/abs/1907.10641). ### XNLI (German) Attribution, license link, change notice; noncommercial restriction. Training translations derive from MultiNLI. [Preserved notice](NOTICE-XNLI.txt). [Citation / original work](https://aclanthology.org/D18-1269/). ### ANLI Attribution, license link, change notice; noncommercial restriction. [Preserved notice](NOTICE-ANLI.txt). [Citation / original work](https://arxiv.org/abs/1910.14599). ### MASSIVE (German) Retain Amazon attribution, license link and change notice. [Preserved notice](NOTICE-MASSIVE.txt). [Citation / original work](https://arxiv.org/abs/2204.08582). ### PAWS-X (German) Google source acknowledgement requested; preserve supplied disclaimer. Do not mislabel as Apache. [Preserved notice](NOTICE-PAWS.txt). [Citation / original work](https://aclanthology.org/D19-1382/). ### HellaSwag Original repository was unavailable during this audit. Original copyright notice and underlying source-text terms remain to be verified. [Citation / original work](https://aclanthology.org/P19-1472/). ### PIQA Original homepage explicitly names AFL-3.0, overriding unknown in the cached mirror card. Preserve attribution/change notices and applicable source/distribution obligations; see full terms. [Preserved notice](NOTICE-AFL-3.0.txt). [Citation / original work](https://arxiv.org/abs/1911.11641). ### Social IQa Attribution, license link and change notice. [Preserved notice](NOTICE-CC-BY-4.0.txt). [Citation / original work](https://aclanthology.org/D19-1454/). ### BoolQ Attribution, license link, modification notice and share-alike for covered adaptations; retain underlying Wikipedia source credits. [Preserved notice](NOTICE-CC-BY-SA-3.0.txt). [Citation / original work](https://aclanthology.org/N19-1300/). ### Stanford Sentiment Treebank / SST-2 via GLUE GLUE explicitly refers to original task licenses. Do not apply the GLUE code license to SST-2 data. [Citation / original work](https://aclanthology.org/D13-1170/). ### A-OKVQA via The Cauldron Preserve license/notices. This repository license does not establish permission for every underlying photograph. [Preserved notice](NOTICE-AOKVQA.txt). [Citation / original work](https://arxiv.org/abs/2206.01718). ### IconQA via The Cauldron Attribution, noncommercial use, change notice and share-alike for covered adaptations. Underlying icons/images may have additional terms. [Preserved notice](NOTICE-CC-BY-NC-SA-4.0.txt). [Citation / original work](https://arxiv.org/abs/2110.13214). ### ScienceQA via The Cauldron Use dataset terms, not the code license. Attribution, noncommercial use, change notice and share-alike for covered adaptations. [Preserved notice](NOTICE-CC-BY-NC-SA-4.0.txt). [Citation / original work](https://arxiv.org/abs/2209.09513). ### AI2D via The Cauldron Attribution, license link, change notice and share-alike for covered adaptations. Data acquired through The Cauldron; original AI2 registry/terms accessed 2026-09-20. Individual diagram provenance not reconstructed. [Preserved notice](NOTICE-CC-BY-SA-4.0.txt). [Citation / original work](https://arxiv.org/abs/1603.07396). ### Visual Spatial Reasoning via The Cauldron Preserve license/notices and changes; retain separate image-source obligations. [Preserved notice](NOTICE-VSR.txt). [Citation / original work](https://aclanthology.org/2023.tacl-1.37/). ### TextVQA via The Cauldron Attribution, license link and modification notice. Original answers converted to choices with distractors; individual image-source attribution is not reconstructed here. [Preserved notice](NOTICE-CC-BY-4.0.txt). [Citation / original work](https://arxiv.org/abs/1904.08920). ### DocVQA via The Cauldron Original site requires review of RRC download terms. Portal unavailable during audit; no blanket permissive license asserted. Individual document rights remain separate. [Citation / original work](https://arxiv.org/abs/2007.00398). ### GUM via CorefUD 1.4 / CRAC 2026 Credit original text sources and annotators. Selected document IDs, authors and source URLs are listed in coref_attributions.json. Document-level license verification remains incomplete. [Preserved notice](NOTICE-GUM.txt). [Citation / original work](https://doi.org/10.1007/s10579-016-9343-x). ### Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026 Credit PCC and original newspaper sources; noncommercial/share-alike conditions. Preserve notices and identify transformations. [Preserved notice](NOTICE-CC-BY-NC-SA-4.0.txt). [Citation / original work](https://aclanthology.org/2020.lrec-1.133/). ## Collection and component credits The Cauldron: HuggingFaceM4 / Laurençon et al., [What matters when building vision-language models?](https://arxiv.org/abs/2405.02246). [Collection terms](https://huggingface.co/datasets/HuggingFaceM4/the_cauldron): component datasets retain their own conditions; rights in collection-authored prompts are offered under CC BY 4.0. This does not relicense source photographs/documents. CorefUD: Nedoluzhko, Novák, Popel, Žabokrtský, Zeldes and Zeman (2022), [CorefUD 1.0: Coreference Meets Universal Dependencies](https://aclanthology.org/2022.lrec-1.520/); data version 1.4 / CRAC 2026 gold. Only GUM and PotsdamCC were used here, confirmed from the actual rows. [GUM annotator credits](https://gucorpling.org/gum/credits.html), [per-document attribution metadata](coref_attributions.json). No Reddit documents occur in the selected GUM rows. GLUE: Wang et al. (2019), [GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding](https://openreview.net/forum?id=rJ4km2R5t7). SST-2 uses the original Stanford dataset; the collection explicitly points to original task terms. Image components: A-OKVQA and VSR use [COCO](https://cocodataset.org/#termsofuse); TextVQA uses [Open Images](https://storage.googleapis.com/openimages/web/factsfigures.html). Their underlying images may retain photographer-specific rights. Individual-image rights and credits have not been fully reconstructed from The Cauldron shards. DocVQA documents originate in the [UCSF Industry Documents Library](https://www.industrydocuments.ucsf.edu/). No source images or document text are redistributed here. ## Evaluation-only sources POPE and the SuperGLUE COPA, WiC and WSC subsets were used for evaluation, not adapter training. They remain in the broader source_provenance.json inventory so the benchmark is traceable; they are not additional training sources. Evaluation results are aggregates, not redistributions of these datasets. ## Remaining gaps DocVQA download terms, the original SST-2 grant, HellaSwag original notices and individual text/image rights are not fully resolved. Generic attribution does not settle those gaps. We do not claim commercial clearance or that an NC/SA dataset license automatically attaches to model weights. Full license texts and preserved notices accompany this inventory where available. See LICENSING.md.