--- language: - hi license: other license_name: see-upstream-dualcodec tags: - audio-codec - dualcodec - hindi - speech pipeline_tag: audio-to-audio pretty_name: DualCodec Hindi 25 Hz --- # DualCodec — Hindi, 25 Hz DualCodec neural speech codec fine-tuned on **Hindi**, operating at 25 Hz. Part of the TinyAya codec study, which asked whether fine-tuning a neural audio codec on a low-resource language improves reconstruction over the stock multilingual checkpoint — the same question that bounds the S2ST model, whose audio quality is capped by its frozen decoder. Weights: `model.safetensors`, `model_1.safetensors`. Produced by [`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning), which fine-tunes Mimi, DualCodec and Kanade on Turkish and Hindi across 8 optimizers with W&B Bayesian sweeps and bootstrap error bars. Upstream DualCodec licence terms apply. ## Code | repo | what it does | |---|---| | [`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning) | fine-tunes Mimi / DualCodec / Kanade on Turkish + Hindi | ## Project **TinyAya Stage 2** — Turkish⇄Hindi speech-to-speech translation with a text inner-monologue: a LoRA-adapted Cohere2 backbone driving a **frozen** Moshi depth decoder over Mimi codes. The v0.3 run covered **76,250 steps / 2.07 epochs** on a Cloud TPU v6e-16 (best val composite **2.8199** @ step 76,000). Read honestly: the text inner-monologue **learns to translate** (free-run chrF++ ~25.7 / 25.1), while **intelligible audio synthesis remains the frontier** (ASR-chrF++ 3.7 / 9.6 against a 92.1 / 86.6 ground-truth-audio ceiling) — bounded by the frozen depth decoder, not by translation understanding. - **Results:** [v0.3 evaluation report](https://github.com/tiny-aya-simultaneous-translation/model/blob/main/docs/v0.3-eval-report.md) - **Training run:** [W&B `xzcb60bl`](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/runs/xzcb60bl) · [emergence report](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/reports/TinyAya-v0.3-Emergence-and-Data-Efficiency--VmlldzoxNzU1OTU1NQ==) - **Blog:** [Adapting Moshi for Low-Resource Speech Translation](https://labscommunity.cohere.com/blog/2026/adapting-moshi-low-resource-speech-translation/) Compute for the v0.3 run was provided by **Google's TPU Research Cloud (TRC)**.