cataluna84's picture
Card: full release metadata + code cross-links
8878bc0 verified
|
Raw History Blame Contribute Delete
2.35 kB
metadata
language:
  - hi
license: other
license_name: see-upstream-dualcodec
tags:
  - audio-codec
  - dualcodec
  - hindi
  - speech
pipeline_tag: audio-to-audio
pretty_name: DualCodec Hindi 25 Hz

DualCodec — Hindi, 25 Hz

DualCodec neural speech codec fine-tuned on Hindi, operating at 25 Hz.

Part of the TinyAya codec study, which asked whether fine-tuning a neural audio codec on a low-resource language improves reconstruction over the stock multilingual checkpoint — the same question that bounds the S2ST model, whose audio quality is capped by its frozen decoder.

Weights: model.safetensors, model_1.safetensors. Produced by codec-finetuning, which fine-tunes Mimi, DualCodec and Kanade on Turkish and Hindi across 8 optimizers with W&B Bayesian sweeps and bootstrap error bars.

Upstream DualCodec licence terms apply.

Code

repo what it does
codec-finetuning fine-tunes Mimi / DualCodec / Kanade on Turkish + Hindi

Project

TinyAya Stage 2 — Turkish⇄Hindi speech-to-speech translation with a text inner-monologue: a LoRA-adapted Cohere2 backbone driving a frozen Moshi depth decoder over Mimi codes.

The v0.3 run covered 76,250 steps / 2.07 epochs on a Cloud TPU v6e-16 (best val composite 2.8199 @ step 76,000). Read honestly: the text inner-monologue learns to translate (free-run chrF++ ~25.7 / 25.1), while intelligible audio synthesis remains the frontier (ASR-chrF++ 3.7 / 9.6 against a 92.1 / 86.6 ground-truth-audio ceiling) — bounded by the frozen depth decoder, not by translation understanding.

Compute for the v0.3 run was provided by Google's TPU Research Cloud (TRC).