cataluna84's picture
Card: full release metadata + code cross-links
8878bc0 verified
|
Raw History Blame Contribute Delete
2.35 kB
---
language:
- hi
license: other
license_name: see-upstream-dualcodec
tags:
- audio-codec
- dualcodec
- hindi
- speech
pipeline_tag: audio-to-audio
pretty_name: DualCodec Hindi 25 Hz
---
# DualCodec — Hindi, 25 Hz
DualCodec neural speech codec fine-tuned on **Hindi**, operating at 25 Hz.
Part of the TinyAya codec study, which asked whether fine-tuning a neural audio
codec on a low-resource language improves reconstruction over the stock
multilingual checkpoint — the same question that bounds the S2ST model, whose
audio quality is capped by its frozen decoder.
Weights: `model.safetensors`, `model_1.safetensors`. Produced by
[`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning), which fine-tunes Mimi, DualCodec and
Kanade on Turkish and Hindi across 8 optimizers with W&B Bayesian sweeps and
bootstrap error bars.
Upstream DualCodec licence terms apply.
## Code
| repo | what it does |
|---|---|
| [`codec-finetuning`](https://github.com/tiny-aya-simultaneous-translation/codec-finetuning) | fine-tunes Mimi / DualCodec / Kanade on Turkish + Hindi |
## Project
**TinyAya Stage 2** — Turkish⇄Hindi speech-to-speech translation with a text
inner-monologue: a LoRA-adapted Cohere2 backbone driving a **frozen** Moshi depth
decoder over Mimi codes.
The v0.3 run covered **76,250 steps / 2.07 epochs** on a Cloud TPU v6e-16
(best val composite **2.8199** @ step 76,000). Read honestly: the text
inner-monologue **learns to translate** (free-run chrF++ ~25.7 / 25.1), while
**intelligible audio synthesis remains the frontier** (ASR-chrF++ 3.7 / 9.6
against a 92.1 / 86.6 ground-truth-audio ceiling) — bounded by the frozen depth
decoder, not by translation understanding.
- **Results:** [v0.3 evaluation report](https://github.com/tiny-aya-simultaneous-translation/model/blob/main/docs/v0.3-eval-report.md)
- **Training run:** [W&B `xzcb60bl`](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/runs/xzcb60bl) · [emergence report](https://wandb.ai/cataluna84/tinyaya-stage2-tpu/reports/TinyAya-v0.3-Emergence-and-Data-Efficiency--VmlldzoxNzU1OTU1NQ==)
- **Blog:** [Adapting Moshi for Low-Resource Speech Translation](https://labscommunity.cohere.com/blog/2026/adapting-moshi-low-resource-speech-translation/)
Compute for the v0.3 run was provided by **Google's TPU Research Cloud (TRC)**.