Fish Audio releases preview of speech generator Drama 3
Fish Audio releases preview of speech generator Drama 3
Fish Audio has published a preview of the speech generator Drama 3, shifting its control scheme toward descriptive intonation cues. The principal change replaces short rigid audio tags with free text descriptions of tempo, tone, and character for more flexible control.
Key capabilities
The preview highlights several workflow and synthesis improvements that affect editing, multi-voice scenes and language support.
- Targeted regeneration allows replacing a single word or phrase inside an existing track without altering the remaining audio segment.
- The model can switch between different voices in mid-sentence, enabling expressive and dynamic speaker turns seamlessly inside one utterance.
- It generates multi-character dialogues in a single pass, reducing production time for scenes that require several distinct voices.
- The system supports Russian language generation, including phonetic and prosodic patterns common to contemporary Russian speech.
Availability and testing notes
Drama 3 is currently paywalled, whereas version 2.1 is publicly accessible and served as the basis for testing. During tests of version 2.1, an available voice closely resembled actor Sergey Burunov, indicating high sample fidelity for some presets.
However, certain tag directives such as 'crowd laughing' and 'excited' did not always take effect, and stress placement errors occurred frequently in trials. These limitations affected some expressive targets in the evaluated presets.
Access
The preview build is available through the API for developers and audio teams interested in early evaluation of the new voice synthesis workflows. Fish Audio positions the descriptive prompt approach as a way to simplify authoring while enabling more nuanced performance control.
Related posts

