Instructions to use wahnfried/Alberich-mini-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use wahnfried/Alberich-mini-v1 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download DATA_SOURCES.md from wahnfried/Alberich-mini-v1: direct link, hf CLI and curl.
- Browser
- Download file 13.6 kB
-
https://huggingface.co/wahnfried/Alberich-mini-v1/resolve/main/DATA_SOURCES.md
- Command line
-
hf download hf://wahnfried/Alberich-mini-v1/DATA_SOURCES.md
-
curl -L -o DATA_SOURCES.md https://huggingface.co/wahnfried/Alberich-mini-v1/resolve/main/DATA_SOURCES.md
Training sources, attributions and terms
Checked 2026-09-20 against the actual 100000-example training file and 6000-example continuation. Counts are presentations, not necessarily unique examples. Every training task is covered below. Source licenses are recorded separately from any license for the adapter. This is a provenance and terms inventory, not a declaration that every right has been cleared.
Source inventory
| Dataset / task | Main | Repair | Attribution | Source terms |
|---|---|---|---|---|
WinoGrande (winogrande) |
20000 | 600 | Sakaguchi et al.; Allen Institute for AI | Apache-2.0 |
XNLI (German) (xnli_de) |
13000 | 250 | Conneau et al.; Facebook Research | CC-BY-NC-4.0 |
ANLI (anli) |
10000 | 500 | Nie et al.; Facebook Research | CC-BY-NC-4.0 |
MASSIVE (German) (massive_de) |
4000 | 500 | FitzGerald et al.; Amazon.com Inc. or its affiliates | CC-BY-4.0 |
PAWS-X (German) (pawsx_de) |
6000 | 200 | Yang et al.; Google LLC | Custom permissive dataset terms |
HellaSwag (hellaswag) |
8000 | 50 | Zellers et al. | MIT reported in pinned dataset card; upstream verification unavailable |
PIQA (piqa) |
5000 | 0 | Bisk et al. | AFL-3.0 |
Social IQa (social_iqa) |
5000 | 50 | Sap et al. | CC-BY-4.0 (pinned source card) |
BoolQ (boolq) |
1006 | 0 | Clark et al.; Google Research; Wikipedia contributors | CC-BY-SA-3.0 (Google dataset card) |
Stanford Sentiment Treebank / SST-2 via GLUE (sst2) |
1000 | 0 | Socher et al.; Stanford NLP; GLUE maintainers | Unresolved for original dataset in reviewed sources |
A-OKVQA via The Cauldron (aokvqa) |
6000 | 100 | Schwenk et al.; Allen Institute for AI | Apache-2.0 repository; COCO images have separate rights |
IconQA via The Cauldron (iconqa) |
5000 | 100 | Lu et al. | CC-BY-NC-SA-4.0 (dataset) |
ScienceQA via The Cauldron (scienceqa) |
3000 | 100 | Lu et al. | CC-BY-NC-SA-4.0 (dataset); code separately MIT |
AI2D via The Cauldron (ai2d) |
1994 | 100 | Kembhavi et al.; Allen Institute for AI | CC-BY-SA-4.0 (AI2-managed AWS registry entry) |
Visual Spatial Reasoning via The Cauldron (vsr) |
1000 | 50 | Liu, Emerson and Collier | Apache-2.0 project; underlying COCO images separate |
TextVQA via The Cauldron (textvqa_mc) |
2000 | 150 | Singh et al.; Facebook AI Research | CC-BY-4.0 dataset; OpenImages/Flickr image-source rights |
DocVQA via The Cauldron (docvqa_mc) |
2000 | 150 | Mathew, Karatzas and Jawahar; UCSF Industry Documents Library (documents) | Original challenge download terms; not verified |
GUM via CorefUD 1.4 / CRAC 2026 (corefud_en) |
3368 | 400 | Zeldes; GUM annotators and original text authors; CorefUD team | CC-BY-NC-SA-4.0 collection; constituent text licenses vary |
Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026 (corefud_de) |
632 | 200 | Bourgonje and Stede; Potsdam corpus contributors; newspaper authors; CorefUD team | CC-BY-NC-SA-4.0 (license linked by original corpus page) |
Locally generated coref_0 (coref_0) |
88 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated coref_1 (coref_1) |
176 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated coref_2 (coref_2) |
192 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated coref_3 (coref_3) |
160 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated coref_4 (coref_4) |
176 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated coref_5 (coref_5) |
208 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated synthetic_dialogue (synthetic_dialogue) |
200 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated synthetic_numeric (synthetic_numeric) |
100 | 1500 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated synthetic_ocr (synthetic_ocr) |
500 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated synthetic_rules (synthetic_rules) |
100 | 0 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Locally generated synthetic_temporal (synthetic_temporal) |
100 | 1000 | This project | Project-authored templates; no separate third-party dataset grant asserted |
Requirements, citations and changes
All external examples were selected, normalized into the project question/option format and mixed for supervised adapter training. Labels were mapped to answer options. CorefUD annotations were converted into contextual coreference questions. TextVQA and DocVQA answers were converted to multiple-choice questions with distractors. Images were processed with the model image processor; the benchmark additionally uses thumbnails. These modifications are ours, not endorsed by the dataset providers.
WinoGrande
Preserve license and applicable copyright/NOTICE notices; identify changes. Preserved notice. Citation / original work.
XNLI (German)
Attribution, license link, change notice; noncommercial restriction. Training translations derive from MultiNLI. Preserved notice. Citation / original work.
ANLI
Attribution, license link, change notice; noncommercial restriction. Preserved notice. Citation / original work.
MASSIVE (German)
Retain Amazon attribution, license link and change notice. Preserved notice. Citation / original work.
PAWS-X (German)
Google source acknowledgement requested; preserve supplied disclaimer. Do not mislabel as Apache. Preserved notice. Citation / original work.
HellaSwag
Original repository was unavailable during this audit. Original copyright notice and underlying source-text terms remain to be verified. Citation / original work.
PIQA
Original homepage explicitly names AFL-3.0, overriding unknown in the cached mirror card. Preserve attribution/change notices and applicable source/distribution obligations; see full terms. Preserved notice. Citation / original work.
Social IQa
Attribution, license link and change notice. Preserved notice. Citation / original work.
BoolQ
Attribution, license link, modification notice and share-alike for covered adaptations; retain underlying Wikipedia source credits. Preserved notice. Citation / original work.
Stanford Sentiment Treebank / SST-2 via GLUE
GLUE explicitly refers to original task licenses. Do not apply the GLUE code license to SST-2 data. Citation / original work.
A-OKVQA via The Cauldron
Preserve license/notices. This repository license does not establish permission for every underlying photograph. Preserved notice. Citation / original work.
IconQA via The Cauldron
Attribution, noncommercial use, change notice and share-alike for covered adaptations. Underlying icons/images may have additional terms. Preserved notice. Citation / original work.
ScienceQA via The Cauldron
Use dataset terms, not the code license. Attribution, noncommercial use, change notice and share-alike for covered adaptations. Preserved notice. Citation / original work.
AI2D via The Cauldron
Attribution, license link, change notice and share-alike for covered adaptations. Data acquired through The Cauldron; original AI2 registry/terms accessed 2026-09-20. Individual diagram provenance not reconstructed. Preserved notice. Citation / original work.
Visual Spatial Reasoning via The Cauldron
Preserve license/notices and changes; retain separate image-source obligations. Preserved notice. Citation / original work.
TextVQA via The Cauldron
Attribution, license link and modification notice. Original answers converted to choices with distractors; individual image-source attribution is not reconstructed here. Preserved notice. Citation / original work.
DocVQA via The Cauldron
Original site requires review of RRC download terms. Portal unavailable during audit; no blanket permissive license asserted. Individual document rights remain separate. Citation / original work.
GUM via CorefUD 1.4 / CRAC 2026
Credit original text sources and annotators. Selected document IDs, authors and source URLs are listed in coref_attributions.json. Document-level license verification remains incomplete. Preserved notice. Citation / original work.
Potsdam Commentary Corpus via CorefUD 1.4 / CRAC 2026
Credit PCC and original newspaper sources; noncommercial/share-alike conditions. Preserve notices and identify transformations. Preserved notice. Citation / original work.
Collection and component credits
The Cauldron: HuggingFaceM4 / Laurençon et al., What matters when building vision-language models?. Collection terms: component datasets retain their own conditions; rights in collection-authored prompts are offered under CC BY 4.0. This does not relicense source photographs/documents.
CorefUD: Nedoluzhko, Novák, Popel, Žabokrtský, Zeldes and Zeman (2022), CorefUD 1.0: Coreference Meets Universal Dependencies; data version 1.4 / CRAC 2026 gold. Only GUM and PotsdamCC were used here, confirmed from the actual rows. GUM annotator credits, per-document attribution metadata. No Reddit documents occur in the selected GUM rows.
GLUE: Wang et al. (2019), GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. SST-2 uses the original Stanford dataset; the collection explicitly points to original task terms.
Image components: A-OKVQA and VSR use COCO; TextVQA uses Open Images. Their underlying images may retain photographer-specific rights. Individual-image rights and credits have not been fully reconstructed from The Cauldron shards. DocVQA documents originate in the UCSF Industry Documents Library. No source images or document text are redistributed here.
Evaluation-only sources
POPE and the SuperGLUE COPA, WiC and WSC subsets were used for evaluation, not adapter training. They remain in the broader source_provenance.json inventory so the benchmark is traceable; they are not additional training sources. Evaluation results are aggregates, not redistributions of these datasets.
Remaining gaps
DocVQA download terms, the original SST-2 grant, HellaSwag original notices and individual text/image rights are not fully resolved. Generic attribution does not settle those gaps. We do not claim commercial clearance or that an NC/SA dataset license automatically attaches to model weights. Full license texts and preserved notices accompany this inventory where available. See LICENSING.md.