ElevenLabs v3 is the most expressive text-to-speech model — multi-speaker dialogue, audio tags for emotion and SFX, and 70+ languages with studio-grade voice cloning.