Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
samsam55 's Collections
Video Models Capability Surveys
Long Horizon Agent Memory Harnesses & Techniques
Streaming Video Understanding
Cyber
VLM (image+text => text)
OCR
Image
Small but smart (?) models
Text to Music
Skills
Video Generation & Pipelines
Coding Agents (Games)
Reinforcement Learning Etc..
Datasets
Self Improving
Run on CPU Optimizations
Deep Search
World View Creation (out painting 3D)
Computer Use
Coding LLMs
Visual Multi Modal LLM
TTS & Speech to Text
Misc
Agents
3D Models & Modeling

TTS & Speech to Text

updated 7 days ago
Upvote
-

  • Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    Paper • 2510.03117 • Published Oct 3, 2025 • 12

  • ResembleAI/chatterbox

    Text-to-Speech • Updated Jun 10 • 1.65M • • 1.82k

  • Phonikud/phonikud

    0.3B • Updated Aug 24, 2025 • 90 • 1

  • UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

    Paper • 2510.13344 • Published Oct 15, 2025 • 65

  • k2-fsa/OmniVoice

    Text-to-Speech • 0.6B • Updated Jul 3 • 1.43M • 1.48k

  • Edge0/Audio8-ASR-Infinite

    Automatic Speech Recognition • 4B • Updated 11 days ago • 40.1k • 2.43k
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs