Skip to main content
Clone a voice from a short reference audio sample and generate new speech in that voice. Useful for personalized TTS, localization, audiobook narration, and character voices.

Available models


Basic voice cloning

Provide a 10-30 second audio sample of the target voice:
reference_audio accepts a file path (auto base64-encoded), a URL, or raw base64 data.

All examples below reuse the same client.

Cross-lingual cloning


Real-time cloning (Zonos)


Batch narration


Tips

  • Reference quality matters. Clean recording, minimal noise, 10-30s of clear speech.
  • HiggsAudio for fidelity. When the clone must be indistinguishable from the original.
  • Chatterbox for languages. 30+ languages, cross-lingual cloning.