How does voice cloning work?

Last updated: November 3, 2025

A unique, high-quality voice for your AI agent builds brand identity, fosters customer trust, and creates a seamless, professional experience. We use state-of-the-art artificial intelligence to create lifelike, custom voices from a human speaker. This guide explains how it works.


The magic of creating a realistic digital voice begins with a high-quality audio recording. A comprehensive recording allows our AI to learn the unique characteristics, tone, and nuances of the speaker. We offer two convenient paths to provide this audio.

Instant Voice Clone

Set up a high-quality voice clone in minutes. This option provides a fast and effective way to validate a voice’s performance - ideal for both testing and production use. To ensure quality please adhere the following guidelines:

  • Recording Length: Only requires 2-5 minutes of speech sample

  • Solo Speaker: The recording must contain only one person speaking. Background voices or interruptions will make the audio unusable.

  • Avoid long pauses in the clip. Too many long pauses could result in the cloned voice drifting from the source clip.

  • Audio Quality: The sound must be clear and consistent. Record in a quiet, non-echoing space and ideally use a high-quality external microphone.

  • File Format: Export the final audio file in either .MP3, m4a or .WAV format.      

Submission: For instant clone recordings, audio files can be sent to [email protected]!

Professional Voice Clone

A Professional Voice Clone allows us to create an almost exact replica of a speaker’s voice, including their accent, speaking style, and audio quality. Unlike Instant Voice Cloning, which works from short samples, Pro Voice Cloning leverages hours of high-quality studio audio to capture subtle nuances and produce a more natural, expressive, and robust-sounding voice model.

Option A: Submitting Your Own Audio Recording

To create your Professional Voice Clone, please submit a high-quality audio recording that meets the following requirements:

  • Recording Length: Provide at least 60–90 minutes of continuous speech.

  • Recording Content: Use live, natural dialogue that reflects your intended use case for the voice.

  • Solo Speaker: The recording must feature only one speaker throughout. Background voices, interruptions, or overlapping speech will render the audio unusable.

  • Audio Quality: Ensure the sound is clear, consistent, and free of background noise or echo. Record in a quiet environment using a high-quality external microphone.

  • File Format: Export your audio file in .MP3 or .WAV format.

Submission: For Pro Voice Recordings, please upload your audio files to your dedicated Slack channel. If you don’t have one, contact us to get set up.

Option B: Guided In-Studio Recording Session

For guaranteed professional quality, we offer a guided recording session at our office. Our team handles all technical aspects to ensure we capture the perfect vocal performance.

A typical session lasts about 90 minutes and is structured to be efficient and effective:

  • Welcome & Setup (10 min): We ensure the voice actor is comfortable and perform a quick soundcheck with our professional-grade equipment.

  • Recording (75-90 min): This is divided into two phases:

    1. Natural Conversation: The speaker talks naturally, allowing us to capture the fundamental tonality and cadence of their voice.

    2. Scripted Reading: The speaker reads from a variety of provided scripts to capture specific emotions and phrases relevant to your use case.

  • Wrap-up (5 min): We conclude the session with a quick review.

After the voice is cloned, our team performs final quality checks. We then add the new custom voice to your account and will notify you as soon as it is ready to be used with your AI agents.