Voice Cloning in April 2026
Introduction
ElevenLabs’ probably the gold standard in voice cloning and TTS right now but it is a paid service and may not be accessible to everyone.
Luckily in early 2026, two open source alternatives emerged:
So let’s try these two models out…
Reference Audio
Arguably the most important part of voice cloning is the reference audio. The better the quality and the more representative of the target voice, the better the cloning results will be. For this demo, we’ll use Morgan Freeman’s voice from this video:
getting audio file from YouTube
As with most data science projects, the getting data part oftens takes the most effort.
To pull the audio, we’ll use yt-dlp1:
pipx upgrade yt-dlp
yt-dlp -x --audio-format mp3 \
--download-sections "*00:00:15.00-00:00:29.00" \
--audio-quality 0 \
-o "morgan_freeman.mp3" "https://www.youtube.com/watch?v=9cixE4DAmBw"- 1
-
tells
yt-dlpto extract the audio and convert it to mp3 format - 2
-
this option allows you to specify a time range for the audio you want to download but might not be precise enough for our needs. We’ll use
ffmpegto trim the audio a bit more precisely in the next step. - 3
- this option sets the audio quality to the highest possible (0 is the best quality for mp3)
- 4
-
make sure to upgrade
yt-dlpto the latest version to ensure that you have access to all the latest features and bug fixes.
This part is crucial because the reference audio should ideally be as clean and representative of the target voice as possible. Any background noise, music, or other voices can interfere with the cloning process and result in a less accurate clone.
While yt-dlp’s --download-sections option allows you to specify a time range for the audio you want to download, I personally found it to not be precise enough even after trying different time ranges.
Luckily we got ffmpeg to handle this:
ffmpeg -i morgan_freeman.mp3 -ss 00:00:05 -c copy "morgan_freeman_final.mp3"And that gives us this reference audio file:
transcription
reference text is not needed for Chatterbox Turbo
Reference text is required for the reference audio for the Qwen3TTS model. So whisper2 will do the work:
whisper morgan_freeman_final.mp3 --output_format txt -o .- 1
-
this will generate a
morgan_freeman_final.txtfile in the current directory with the transcribed text of the audio. The--output_format txtflag specifies that we want the output in plain text format, and the-o .flag specifies that we want to save the output in the current directory.
The resulting transcript looks like this:
morgan_freeman_final.txt
it tells me what's right and what's not when to leave and where to go it's not shakespeare
it does not speak in memorable lines my inner voice always gives it to me straightResults
for simplicity, I used both models’ HuggingFace Space demo to read the following passage from Mark Carney’s speech at Davos:
passage.txt
We know the old order is not coming back. We shouldn't mourn it. Nostalgia is not a strategy.
But we believe that from the fracture, we can build something better, stronger, more just.
This is the task of the middle powers. The countries that have the most to lose from a world of fortresses and the most to gain from genuine co-operation.
The powerful have their power. But we have something too — the capacity to stop pretending, to name reality, to build our strength at home and to act together.Then hosted the generated audio files on CropGif3 for embedding here.
Chatterbox Turbo
Chatterbox Turbo was easier to since it only requires reference audio and text to read. Reference text is not required, which saves the extra step of transcribing the reference audio. It also seems to be able to handle longer text input with the following was generated in one go.
Qwen3 TTS
Subjectively the Qwen3 TTS’s output sounds more natural and have more “emotion” but I did have to split the generation into two parts then stitch the output audios back together4
Takeaway
Chatterbox Turbo is more straightforward to use and can handle longer text input. While the output quality might not as good as Qwen3 TTS5, you only really need to send the reference audio and the text to the model and it can generate the output in one go.
This makes it a lot easier to turn it into an agentic tool for whatever voice application you are building.
Furthermore, it supports paralinguistic tags and the model’s suppose to be small and fast enough to run locally.
Footnotes
since it’s free, no account required, and the generated url is permanent↩︎
ran something like this:
↩︎ffmpeg -i part1.wav -i part2.wav \ -filter_complex "[0][1]concat=n=2:v=0:a=1" \ final.mp3subjectively speaking but you could play around with the
cfg_weightandexaggerationparams per the tips & tricks↩︎