OpenAudio S1 is a leading text-to-speech (TTS) model trained on more than 2 million hours of audio data in multiple languages.
Supported languages:
Please refer to Fish Speech Github for more info. Demo available at Fish Audio Playground. Visit the OpenAudio website for blog & tech report.
OpenAudio S1 supports a variety of emotional, tone, and special markers to enhance speech synthesis:
1. Emotional markers: (angry) (sad) (disdainful) (excited) (surprised) (satisfied) (unhappy) (anxious) (hysterical) (delighted) (scared) (worried) (indifferent) (upset) (impatient) (nervous) (guilty) (scornful) (frustrated) (depressed) (panicked) (furious) (empathetic) (embarrassed) (reluctant) (disgusted) (keen) (moved) (proud) (relaxed) (grateful) (confident) (interested) (curious) (confused) (joyful) (disapproving) (negative) (denying) (astonished) (serious) (sarcastic) (conciliative) (comforting) (sincere) (sneering) (hesitating) (yielding) (painful) (awkward) (amused)
2. Tone markers: (in a hurry tone) (shouting) (screaming) (whispering) (soft tone)
3. Special markers: (laughing) (chuckling) (sobbing) (crying loudly) (sighing) (panting) (groaning) (crowd laughing) (background laughter) (audience laughing)
Special markers with corresponding onomatopoeia:
OpenAudio S1 includes the following models:
Both S1 and S1-mini incorporate online Reinforcement Learning from Human Feedback (RLHF).
Seed TTS Eval Metrics (English, auto eval, based on OpenAI gpt-4o-transcribe, speaker distance using Revai/pyannote-wespeaker-voxceleb-resnet34-LM):
This model is permissively licensed under the CC-BY-NC-SA-4.0 license.