|
Download README.md from tiny-aya-translate/dualcodec-hindi-25hz: direct link, hf CLI and curl.
- Browser
- Download file 2.35 kB
-
https://huggingface.co/tiny-aya-translate/dualcodec-hindi-25hz/resolve/main/README.md
- Command line
-
hf download hf://tiny-aya-translate/dualcodec-hindi-25hz/README.md
-
curl -L -o README.md https://huggingface.co/tiny-aya-translate/dualcodec-hindi-25hz/resolve/main/README.md
2.35 kB
| language: | |
| - hi | |
| license: other | |
| license_name: see-upstream-dualcodec | |
| tags: | |
| - audio-codec | |
| - dualcodec | |
| - hindi | |
| - speech | |
| pipeline_tag: audio-to-audio | |
| pretty_name: DualCodec Hindi 25 Hz | |
| # DualCodec — Hindi, 25 Hz | |
| DualCodec neural speech codec fine-tuned on **Hindi**, operating at 25 Hz. | |
| Part of the TinyAya codec study, which asked whether fine-tuning a neural audio | |
| codec on a low-resource language improves reconstruction over the stock | |
| multilingual checkpoint — the same question that bounds the S2ST model, whose | |
| audio quality is capped by its frozen decoder. | |
| Weights: `model.safetensors`, `model_1.safetensors`. Produced by | |
| [`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning), which fine-tunes Mimi, DualCodec and | |
| Kanade on Turkish and Hindi across 8 optimizers with W&B Bayesian sweeps and | |
| bootstrap error bars. | |
| Upstream DualCodec licence terms apply. | |
| ## Code | |
| | repo | what it does | | |
| |---|---| | |
| | [`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning) | fine-tunes Mimi / DualCodec / Kanade on Turkish + Hindi | | |
| ## Project | |
| **TinyAya Stage 2** — Turkish⇄Hindi speech-to-speech translation with a text | |
| inner-monologue: a LoRA-adapted Cohere2 backbone driving a **frozen** Moshi depth | |
| decoder over Mimi codes. | |
| The v0.3 run covered **76,250 steps / 2.07 epochs** on a Cloud TPU v6e-16 | |
| (best val composite **2.8199** @ step 76,000). Read honestly: the text | |
| inner-monologue **learns to translate** (free-run chrF++ ~25.7 / 25.1), while | |
| **intelligible audio synthesis remains the frontier** (ASR-chrF++ 3.7 / 9.6 | |
| against a 92.1 / 86.6 ground-truth-audio ceiling) — bounded by the frozen depth | |
| decoder, not by translation understanding. | |
| - **Results:** [v0.3 evaluation report](https://github.com/tiny-aya-simultaneous-translation/model/blob/main/docs/v0.3-eval-report.md) | |
| - **Training run:** [W&B `xzcb60bl`](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/runs/xzcb60bl) · [emergence report](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/reports/TinyAya-v0.3-Emergence-and-Data-Efficiency--VmlldzoxNzU1OTU1NQ==) | |
| - **Blog:** [Adapting Moshi for Low-Resource Speech Translation](https://labscommunity.cohere.com/blog/2026/adapting-moshi-low-resource-speech-translation/) | |
| Compute for the v0.3 run was provided by **Google's TPU Research Cloud (TRC)**. | |