File size: 15,019 Bytes
658c616
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
---
language:
  - en
  - de
  - fr
  - es
  - it
  - pt
  - nl
  - pl
  - ru
  - uk
  - cs
  - ro
  - hu
  - sv
  - da
  - fi
  - no
  - el
  - bg
  - sk
  - hr
  - sr
  - tr
license: mit
tags:
  - text-to-speech
  - tts
  - speech-synthesis
  - audio-generation
  - european-languages
  - diffusion
  - autoregressive
pipeline_tag: text-to-speech
inference: false
model-index:
  - name: kugelaudio-0-open
    results:
      - task:
          type: text-to-speech
        dataset:
          type: custom
          name: YODAS2
        metrics:
          - type: win-rate
            value: 78.0
            name: Human Preference vs ElevenLabs
---

# 🎙️ KugelAudio-0-Open

**Open-source text-to-speech for European languages**
7B parameter model powered by an AR + Diffusion architecture

<p align="center">
  <a href="https://github.com/Kugelaudio/kugelaudio-open"><img src="https://img.shields.io/badge/GitHub-Source_Code-black" alt="GitHub Source Code"></a>
  <a href="https://kugelaudio.com"><img src="https://img.shields.io/badge/🌐-Website-blue" alt="KugelAudio Website"></a>
</p>

<table align="center" style="border-collapse: collapse; border: none;">
  <tr style="border: none;">
    <td style="border: none; padding: 0 20px;">
      <a href="https://kugelaudio.com">
        <img src="https://www.kugelaudio.com/logos/Logo%20Short.svg" alt="KugelAudio" style="height: 60px; width: auto;">
      </a>
    </td>
    <td style="border: none; padding: 0 20px;">
      <a href="https://hpi.de/ki-servicezentrum/">
        <img src="https://docs.sc.hpi.de/attachments/aisc/aisc-logo.png" alt="KI-Servicezentrum Berlin-Brandenburg" style="height: 60px; width: auto;">
      </a>
    </td>
    <td style="border: none; padding: 0 20px;">
      <a href="https://www.bmftr.bund.de">
        <img src="https://hpi.de/fileadmin/_processed_/a/3/csm_BMFTR_de_Web_RGB_gef_durch_cd1f5345bd.jpg" alt="Gefördert durch BMFTR" style="height: 60px; width: auto;">
      </a>
    </td>
  </tr>
</table>

License: MIT Python 3.10+ Hosted API

KugelAudio KI-Servicezentrum Berlin-Brandenburg Gefördert durch BMFTR

---

## Motivation

**Open-source text-to-speech models for European languages are significantly lagging behind.** While English TTS has seen remarkable progress, speakers of German, French, Spanish, Polish, and dozens of other European languages have been underserved by the open-source community.

KugelAudio aims to change this. Building on the excellent foundation laid by the [VibeVoice team at Microsoft](https://github.com/microsoft/VibeVoice), we've trained a model specifically focused on European language coverage, using approximately **200,000 hours** of highly pre-processed and enhanced speech data from the [YODAS2 dataset](https://huggingface.co/datasets/espnet/yodas).

## 🏆 Benchmark Results: Outperforming ElevenLabs

**KugelAudio achieves state-of-the-art performance**, beating industry leaders including ElevenLabs in rigorous human preference testing. This breakthrough demonstrates that open-source models can now rival - and surpass - the best commercial TTS systems.

### Human Preference Benchmark (A/B Testing)

We conducted extensive A/B testing with **339 human evaluations** to compare KugelAudio against leading TTS models. Participants listened to a reference voice sample, then compared outputs from two models and selected which sounded more human and closer to the original voice.

### German Language Evaluation

The evaluation specifically focused on **German language samples** with diverse emotional expressions and speaking styles:

* **Neutral Speech**: Standard conversational tones
* **Shouting**: High-intensity, elevated volume speech
* **Singing**: Melodic and rhythmic speech patterns
* **Drunken Voice**: Slurred and irregular speech characteristics

These diverse test cases demonstrate the model's capability to handle a wide range of speaking styles beyond standard narration.

### OpenSkill Ranking Results

| Rank | Model | Score | Record | Win Rate |
|------|-------|-------|--------|----------|
| 🥇 1 | **KugelAudio** | **26** | 71W / 20L / 23T | **78.0%** |
| 🥈 2 | ElevenLabs Multi v2 | 25 | 56W / 34L / 22T | 62.2% |
| 🥉 3 | ElevenLabs v3 | 21 | 64W / 34L / 16T | 65.3% |
| 4 | Cartesia | 21 | 55W / 38L / 19T | 59.1% |
| 5 | VibeVoice | 10 | 30W / 74L / 8T | 28.8% |
| 6 | CosyVoice v3 | 9 | 15W / 91L / 8T | 14.2% |

_Based on 339 evaluations using Bayesian skill-rating system (OpenSkill)_

## Audio Samples

Listen to KugelAudio's diverse voice capabilities across different speaking styles and languages:

### German Voice Samples

| Sample | Description | Audio Player |
|--------|-------------|--------------|
| **Whispering** | Soft whispering voice | <audio controls><source src="https://huggingface.co/kugelaudio/kugelaudio-0-open/resolve/main/samples/258_Lukas_der_Flüsterer.wav" type="audio/wav"></audio> |
| **Female Narrator** | Professional female reader voice | <audio controls><source src="https://huggingface.co/kugelaudio/kugelaudio-0-open/resolve/main/samples/266_Petra_die_Vorleserin.wav" type="audio/wav"></audio> |
| **Angry Voice** | Irritated and frustrated speech | <audio controls><source src="https://huggingface.co/kugelaudio/kugelaudio-0-open/resolve/main/samples/261_Sauerer_Felix.wav" type="audio/wav"></audio> |
| **Radio Announcer** | Professional radio broadcast voice | <audio controls><source src="https://huggingface.co/kugelaudio/kugelaudio-0-open/resolve/main/samples/277_Radio_Lars.wav" type="audio/wav"></audio> |

*All samples are generated using pre-encoded voice embeddings.*

### Training Details

- **Base Model**: [Microsoft VibeVoice](https://github.com/microsoft/VibeVoice)
- **Training Data**: ~200,000 hours from [YODAS2](https://huggingface.co/datasets/espnet/yodas)
- **Hardware**: 8x NVIDIA H100 GPUs
- **Training Duration**: 5 days

### Supported Languages

This model supports the following European languages:

| Language | Code | Flag | Language | Code | Flag | Language | Code | Flag |
|----------|------|------|----------|------|------|----------|------|------|
| English | en | 🇺🇸 | German | de | 🇩🇪 | French | fr | 🇫🇷 |
| Spanish | es | 🇪🇸 | Italian | it | 🇮🇹 | Portuguese | pt | 🇵🇹 |
| Dutch | nl | 🇳🇱 | Polish | pl | 🇵🇱 | Russian | ru | 🇷🇺 |
| Ukrainian | uk | 🇺🇦 | Czech | cs | 🇨🇿 | Romanian | ro | 🇷🇴 |
| Hungarian | hu | 🇭🇺 | Swedish | sv | 🇸🇪 | Danish | da | 🇩🇰 |
| Finnish | fi | 🇫🇮 | Norwegian | no | 🇳🇴 | Greek | el | 🇬🇷 |
| Bulgarian | bg | 🇧🇬 | Slovak | sk | 🇸🇰 | Croatian | hr | 🇭🇷 |
| Serbian | sr | 🇷🇸 | Turkish | tr | 🇹🇷 | | | |

> **📊 Language Coverage Disclaimer**: Quality varies significantly by language. Spanish, French, English, and German have the strongest representation in our training data (~200,000 hours from YODAS2). Other languages may have reduced quality, prosody, or vocabulary coverage depending on their availability in the training dataset.

### Model Specifications

| Property              | Value                                                                       |
| --------------------- | --------------------------------------------------------------------------- |
| **Parameters**        | 7B                                                                          |
| **Architecture**      | AR + Diffusion (Qwen2.5-7B backbone)                                        |
| **Base Model**        | [Microsoft VibeVoice](https://github.com/microsoft/VibeVoice)               |
| **Audio Sample Rate** | 24kHz                                                                       |
| **Audio Format**       | Mono, float32                                                               |
| **VRAM Required**     | \~19GB                                                                      |
| **Training Hardware** | 8x NVIDIA H100                                                              |
| **Training Duration** | 5 days                                                                      |
| **Training Data**     | \~200,000 hours from [YODAS2](https://huggingface.co/datasets/espnet/yodas) |

## Quick Start

### Installation

```bash
# Install with pip
pip install kugelaudio-open

# Or with uv (recommended)
uv pip install kugelaudio-open
```

### Basic Usage

```python
from kugelaudio_open import (
    KugelAudioForConditionalGenerationInference,
    KugelAudioProcessor,
)
import torch

# Load model
device = "cuda" if torch.cuda.is_available() else "cpu"
model = KugelAudioForConditionalGenerationInference.from_pretrained(
    "kugelaudio/kugelaudio-0-open",
    torch_dtype=torch.bfloat16,
).to(device)
model.eval()

processor = KugelAudioProcessor.from_pretrained("kugelaudio/kugelaudio-0-open")

# Strip encoder weights to save VRAM (only decoders needed for inference)
model.model.strip_encoders()

# See available voices
print(processor.get_available_voices())  # ["default", "warm", "clear"]

# Generate speech with a specific voice
inputs = processor(text="Hallo Welt! Das ist KugelAudio.", voice="default", return_tensors="pt")
inputs = {k: v.to(device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}

with torch.no_grad():
    outputs = model.generate(**inputs, cfg_scale=3.0)

# Save audio
processor.save_audio(outputs.speech_outputs[0], "output.wav")
```

### Voices

KugelAudio provides pre-encoded voices that can be selected by name. The voices are stored as `.pt` files in the `voices/` folder and are automatically downloaded when needed.

```python
# List available voices
voices = processor.get_available_voices()
print(voices)  # ["default", "warm", "clear"]

# Generate with a specific voice
inputs = processor(text="Hallo, das ist eine warme Stimme!", voice="warm", return_tensors="pt")
inputs = {k: v.to(device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}

with torch.no_grad():
    outputs = model.generate(**inputs, cfg_scale=3.0)

processor.save_audio(outputs.speech_outputs[0], "warm_voice_output.wav")
```

> **Note:** Voice cloning from raw audio is not supported in this open-source release. Only the pre-encoded voices listed in `voices/voices.json` are available.

### Generation Parameters

| Parameter        | Default | Description                                                                |
| ---------------- | ------- | -------------------------------------------------------------------------- |
| cfg\_scale       | 3.0     | Classifier-free guidance scale (1.0-10.0). Higher = more adherence to text |
| max\_new\_tokens | 2048    | Maximum number of tokens to generate                                       |
| do\_sample       | False   | Whether to use sampling (vs greedy decoding)                               |
| temperature      | 1.0     | Sampling temperature (if do_sample=True)                                  |

## Architecture

KugelAudio uses a hybrid **Autoregressive + Diffusion** architecture based on Microsoft's VibeVoice:

```
Text Input → Qwen2.5-7B Backbone → Diffusion Head → Acoustic Decoder → Audio Output

                              Pre-encoded Voice Embedding
```

1. **Text Encoder**: Qwen2.5-7B language model encodes input text
2. **Diffusion Head**: Predicts speech latents using denoising diffusion (20 steps)
3. **Acoustic Decoder**: Hierarchical convolutional decoder converts latents to 24kHz audio

## Audio Watermarking

All audio generated by this model is automatically watermarked using Facebook's AudioSeal. The watermark is:

* **Imperceptible**: No audible difference in audio quality
* **Robust**: Survives compression, resampling, and editing
* **Detectable**: Can verify if audio was generated by KugelAudio

### Verify Watermark

```python
from kugelaudio_open.watermark import AudioWatermark

watermark = AudioWatermark()
result = watermark.detect(audio, sample_rate=24000)

print(f"Watermark detected: {result.detected}")
print(f"Confidence: {result.confidence:.1%}")
```

## Intended Use

### ✅ Appropriate Uses

* **Accessibility**: Text-to-speech for visually impaired users
* **Content Creation**: Podcasts, videos, audiobooks, e-learning
* **Voice Assistants**: Chatbots and virtual assistants
* **Language Learning**: Pronunciation practice and language education
* **Creative Projects**: With proper consent and attribution

### ❌ Prohibited Uses

* Creating deepfakes or misleading content
* Impersonating individuals without explicit consent
* Fraud, deception, or scams
* Harassment or abuse
* Any illegal activities

## Limitations

* **VRAM Requirements**: Requires \~19GB VRAM for inference (less with `strip_encoders()`)
* **Speed**: Approximately 1.0x real-time on modern GPUs
* **Language Quality Variation**: Quality may vary across languages based on training data distribution

## Hosted API

For production use without managing infrastructure, use our hosted API at kugelaudio.com:

***Ultra-low latency**: <100ms end-to-end
* 🌍 **Global edge deployment**
* 🔧 **Zero setup required**
* 📈 **Auto-scaling**

```python
from kugelaudio import KugelAudio

client = KugelAudio(api_key="your_api_key")
audio = client.tts.generate(text="Hello from KugelAudio!", model="kugel-1-turbo")
audio.save("output.wav")
```

## Acknowledgments

This model would not have been possible without the contributions of many individuals and organizations:

* **Microsoft VibeVoice Team**: For the excellent foundation architecture that this model builds upon
* **YODAS2 Dataset**: For providing the large-scale multilingual speech data
* **Qwen Team**: For the powerful language model backbone
* **Facebook AudioSeal**: For the audio watermarking technology

### Special Thanks

* **Carlos Menke**: For his invaluable efforts in gathering the first datasets and extensive work benchmarking the model
* **AI Service Center Berlin-Brandenburg (KI-Servicezentrum)**: For providing the GPU resources (8x H100) that made training this model possible

## Citation

```bibtex
@software{kugelaudio2026,
  title = {KugelAudio: Open-Source Text-to-Speech for European Languages},
  author = {Kratzenstein, Kajo and Menke, Carlos},
  year = {2026},
  institution = {Hasso-Plattner-Institut},
  url = {https://huggingface.co/kugelaudio/kugelaudio-0-open}
}
```

## License

This model is released under the MIT License.

## Author

**Kajo Kratzenstein**  
📧 [kajo@kugelaudio.com](mailto:kajo@kugelaudio.com)  
🌐 [kugelaudio.com](https://kugelaudio.com)

**Carlos Menke**

---

**Funding Notice**

Das zugrunde liegende Vorhaben wurde mit Mitteln des Bundesministeriums für Forschung, Technologie und Raumfahrt unter dem Förderkennzeichen »KI-Servicezentrum Berlin-Brandenburg« 16IS22092 gefördert.

_This project was funded by the German Federal Ministry of Research, Technology and Space under the funding code "AI Service Center Berlin-Brandenburg" 16IS22092._