← All Articles

AI Voice Cloning · 2026 Guide

Best AI voice cloning in 2026: how to clone your voice

Fine-tuned Dots.TTS was the best AI voice-cloning model in my testing, and it is Apache 2.0. MOSS-TTS v1.5 finished a close second. The model that sounds most like me, OmniVoice, is the one I cannot ship. You can play every ranked clip below, including the real recording.

Austin Starks Austin Starks ✦ Founder, NexusTrade ✦ August 24, 2026 ✦ 16 min read

I write about markets and AI for a living. Most of what I publish never becomes a video, because making one means setting up a camera and reading my own writing out loud, and I would rather spend those hours on the research. I wanted to feed in the text and get narration in my voice. A paid clone was the obvious first try and it was not good enough, so I trained my own.

Fine-tuned Dots.TTS won the final rounds. The model that sounds best is one I cannot use.

Press play on any row. I ranked one question: does it sound like me. Every verdict clip reads the identical 67-word script. The zero-shot arms use the same clean 20-second reference, the files were loudness matched and duration padded, and the model mapping stayed sealed until after I committed to a ranking.

The top two were separated over three sealed rounds, including one on a script neither had seen with three takes per model. In that round I flagged a "hiccup" artifact on two clips and asked whether they were the same model. They were, and it was MOSS. Dots.TTS took the round.

The useful answers
Best overall, and what I shipDots.TTS soar, fine-tuned
Runner-up, close enough to swap inMOSS-TTS v1.5, fine-tuned
Sounds the most like me, cannot shipOmniVoice + LoRA, checkpoint 800
Best with no trainingDots.TTS soar, zero-shot
Best paid optionElevenLabs Professional

This table merges blind rounds run over four days. The top three met each other directly. Rows further down never met, and are ordered from my notes rather than from a head-to-head, so read adjacent rows as ties. Every clip is one take, chosen the way anyone publishing a demo chooses one, which is exactly the problem the second half of this article is about.

Me, actually reading it
human"perfect", identified every time
The control. Recorded on a phone, reading the same 67-word script.
1Dots.TTS soar, fine-tuned
Apache 2.0, 2.2B"strong", "great, no crazy defects"
The keeper. Won the final round on a script it had never seen, and holds the tightest take-to-take spread of anything I trained: 1.3 seconds across five renders.
2MOSS-TTS v1.5, fine-tuned
Apache 2.0, 8.5B"great", "extremely good... it doesn't collapse"
Lost by a hair. Four times the parameters of the winner, and carries a faint hiccup I picked out blind across two separate takes.
3OmniVoice + LoRA, checkpoint 800
open code, CC-BY-NC weights"probably the best... but end of sentences still collapse hard"
The closest thing to my actual voice in this entire test, and I cannot use it. It garbles phrase endings on long text and silently drops words on short text. Listen to the end of this clip.
4CosyVoice 3 0.5B, fine-tuned
Apache 2.0"really good, no issues"
My first keeper, shipped for weeks, and it still holds up with no issues on either script. Two models eventually edged it out. It is 0.5B, the smallest thing here that I would genuinely ship.
5Dots.TTS soar, zero-shot
Apache 2.0"pretty good"
98.6% word coverage, zero added words, 23.84s.
6OmniVoice, zero-shot
open code, CC-BY-NC weights"really good", "decent"
100% word coverage after trimming 2.0s of hallucinated speech before the script.
7Audio8 0.6b, zero-shot
Apache 2.0"okay", "good-ish"
100% word coverage, 27.31s.
8Chatterbox, zero-shot
MIT"ok"
The fair rerun using the same clean session1 reference as every other zero-shot arm. This replaces the old scavenged-reference clip.
9Qwen3-TTS, zero-shot
Apache 2.0"okay"
The clean session1 reference, identical script.
10IndexTTS-2, zero-shot
custom model licence"okay, robotic, too fast"
The clean session1 reference, identical script.
11CosyVoice 3, zero-shot
Apache 2.0"okay"
The same architecture as the old fine-tuned keeper, with no adapter.
12Audio8 0.1b, zero-shot
Audio8 Community License v1.0"okay, robotic"
100% word coverage, 27.26s.
13MOSS-TTS v1.5, zero-shot (8.5B)
Apache 2.0"decent, worse than Audio8 0.6b"
21.76s. The largest model in the test, and near the bottom without training. Fine-tuned, the same model finished second overall. See below.
14ElevenLabs Professional Voice Clone
$22/moclose, not close enough
The better of the two paid clones I made, trained from cleaner source audio.
15Qwen3-TTS 1.7B, fine-tuned
Apache 2.0great audio, not my voice
Clean, natural synthesis. The generated speaker is somebody else.
16XTTS-v2, fine-tuned
non-commercial"hilariously bad"
Two separate takes landed at the bottom of a blind set.
The result that changed the article

Dots.TTS is the best zero-shot model I tested. It beat OmniVoice zero-shot in the round that ranked them together, and OmniVoice zero-shot in turn beat Chatterbox, Audio8 and MOSS-TTS in an earlier round. The previous Chatterbox claim came from an older clip rendered with a scavenged reference. When I reran Chatterbox against the same clean session1 reference as the field, I called it "ok" and ranked four zero-shot systems above it.

Fine-tuned Dots.TTS beat everything I trained. It won a head-to-head against fine-tuned MOSS on a 94-word article paragraph neither model had seen, three takes each, and it won the 124-word round before that. Both are Apache 2.0, so the licence does not decide it.

OmniVoice sounds the most like me and I still cannot ship it. Its code is Apache 2.0 but its pretrained weights are CC-BY-NC, and that is the smaller problem. The real one is that it garbles the ends of phrases on long text and silently drops words on short text, in every take I have ever rendered.

The fine-tunes had 45 minutes and the zero-shot models had 20 seconds. That measures my exact setup, not every possible zero-shot reference length. ElevenLabs also used cleaner source audio than my corpus.

Every one of these can sound like me. None of them can do it on demand.

Somewhere around the fortieth blind test I stopped learning anything from the question I had started with. Five different architectures now produce audio I call "great" on the first listen. Timbre is solved, and it is solved by more than one model.

What is not solved is doing it twice.

The same checkpoint, the same script, the same reference clip, run again, gives you a different result. Some of those results are unusable. Here is my own verbatim reaction to nine takes from the model I eventually shipped, all from one checkpoint, all reading the same 124 words:

ONE CHECKPOINT, NINE TAKES, SAME SCRIPT
23s   too short to hold the script
44s   "really good"
22s   too short
50s   "great, more stable"      <-- shipped
48s   "damn-near perfect"
46s   "perfect jk maybe slightly robotic"
30s   too short
47s   "cutoff bad"
 8s   too short

Four of the nine were too short to physically contain 124 words. Of the five that were long enough, four were good and one cut off sentence endings. A demo is one take, chosen by the person publishing it. That is true of every voice-cloning demo you have ever heard, including the ones in this article.

So the workflow that actually ships is not "generate narration". It is generate three to five takes, throw away anything materially shorter than the others, and listen to what survives. The machine discards the obvious failures. A human still has to approve the rest.

The screening rule is cruder than it sounds and it took me a while to get right. Sort the takes by duration and drop the short ones, but compare each model against its own takes rather than a fixed words-per-minute number. Dots reads at roughly 175 words per minute and MOSS at roughly 150, so a single threshold either throws away good Dots takes or waves through bad MOSS ones. On one 67-word script, five Dots takes came in at 22.56, 22.88, 22.88, 24.00 and 24.32 seconds. The 22.56 one had quietly dropped the closing sentence, "So I sold all of my Google calls." It sounded completely fine. A gap of one and a half seconds against its own siblings was the only evidence.

And duration only catches missing speech. It cannot catch a model mangling a word inside a full-length take, and it is useless for OmniVoice, which pads every render to the same length whether it succeeded or not: 25 renders, all between 45.85 and 46.18 seconds, some of them fine and some of them dropping words. There is no automatic check for that. There is only listening.

If you are choosing a model for real work, this is the axis to choose on. Not which one wins a single blind test, but which one loses the fewest takes.

The model I ranked near-last, zero-shot, finished second once I trained it

MOSS-TTS v1.5 was the largest model in the test at 8.5B. Two Reddit commenters recommended it in strong terms. Blind and untrained, I called it "decent, worse than Audio8 0.6b" and ranked it below every fine-tune and below a 0.6b model.

Then I fine-tuned it on the same 45 minutes. It finished second overall, and at one checkpoint I described it as "extremely good... honestly, it's amazing".

That is the most useful thing I learned about picking a model. A zero-shot ranking tells you almost nothing about how a model behaves once it has your voice. If I had trusted the zero-shot table, I would have thrown away the runner-up before training it. The winner, Dots.TTS, made the same jump: sixth zero-shot, first fine-tuned.

Parameter count bought nothing on its own. The 2.2B winner beat the 8.5B runner-up, and both beat models in between. What mattered was whether the architecture could absorb 45 minutes of one voice.

Try it yourself: one of these five is a real human

This is a reader exercise, not one of my ranking sessions. I assembled it from the same clips under the same conditions so you can try the thing I did: identical script, every file normalised to the same loudness, padded to the same length and the same byte size, and served from opaque filenames so the page source does not give it away. The session that actually produced the rankings is near the end of the article, and it used a different five files.

A
B
C
D
E

How to clone your voice, in one pass

  1. Record 30 to 45 minutes in one session. One room, one mic, one sitting. Mine was an iPhone with no interface and no treated room. Do not pool old clips from different days: that was the single biggest quality jump in my testing.
  2. Convert to 24 kHz mono and skip enhancement. ffmpeg -i src.mov -ar 24000 -ac 1 -c:a pcm_s16le audio.wav. I tested an enhanced version of the same recording and it scored worse.
  3. Cut into 4 to 12 second clips on sentence boundaries, then check the clip manifest against your recording length. My first chunker dropped over-long segments instead of splitting them and silently lost 27 of 52 minutes.
  4. Transcribe with openai-whisper, model large-v3, running on the same GPU. Each training row is a (text, audio) pair.
  5. Rent one 24 GB GPU and prove it is real before installing anything: python -c "import torch; assert torch.cuda.is_available()". A silent fall back to CPU will hide a dead machine for a whole run.
  6. Read the project's Dockerfile. It is the only authoritative record of the environment. For CosyVoice the load-bearing details are Python 3.10 and pynini from conda-forge.
  7. Train, saving a checkpoint per epoch. About 22 seconds per epoch for 45 minutes of audio. Do not assume the last checkpoint is the best one: mine peaked at epoch 4 of 31.
  8. Render three takes per checkpoint, then transcribe each one and diff it against your script. Generation is stochastic and most of these models take no seed, so takes differ in how much of the script they actually speak. I measured 25%, 54% and 90% word coverage across three renders of one input. Keep the take with the highest coverage and re-render anything below about 90%.
  9. Rank blind. Normalise loudness, pad to one duration, hide the labels until after you commit.
Copy-paste plan for an AI coding agent

Paste this into Claude Code, Cursor, or any agent with shell access. It encodes every failure that cost me time, including the ones that fail silently.

VOICE CLONE PLAN
Clone my voice with an open-source TTS model. Work through these steps in
order and verify each one before moving on. Tell me when something fails; do
not silently retry.

=== 0. SAFETY: THIS PROMPT CONTAINS NO SECRETS. DO NOT ADD ANY. ===

Do not paste an API key, token, or password into this prompt or into any chat
window. Everything below authenticates interactively or through a file the
agent never prints.

Agent: never echo, log, or repeat back the contents of a credential file. If
you need to confirm a token loaded, print its LENGTH, not its value.

=== 0b. TOOLING YOU NEED FIRST ===

RunPod (rents the GPU). Connect the MCP server so the agent can create and
destroy pods:

    claude mcp add --transport http runpod -s user https://mcp.getrunpod.io/

This uses a browser OAuth flow, so no API key is stored on disk and nothing
sensitive passes through this prompt. Then add an ACCOUNT-LEVEL SSH key at
runpod.io/console/user/settings. The SSH proxy authenticates against that key,
not the pod's PUBLIC_KEY, and without it every connection fails with
"Permission denied (publickey)".

Hugging Face (hosts the base models, your corpus, and your trained weights).
Create a token at huggingface.co/settings/tokens. Prefer a FINE-GRAINED token
scoped to only the repos you will create, with write access. A global write
token can delete every repo you own, and an agent will be handling it.

Log in INTERACTIVELY so the token never lands in shell history or in a process
argument list:

    huggingface-cli login

To place it on a pod, read it from stdin and write it to a file. Never pass a
token as a command-line flag:

    read -rs HF_TOKEN && printf '%s' "$HF_TOKEN" > /root/.hf_token && unset HF_TOKEN
    chmod 600 /root/.hf_token

=== 1. RECORD ===

30-45 minutes in ONE session. One room, one mic, one sitting. A phone is fine.
Do not pool clips from different days: mixed sessions teach the model the room
differences along with the voice.
Scraped audio does not scale: 14.5 min of clips pulled from old videos tied
2.5 min of the same. One clean session of 30-45 min is what actually worked.
No noise reduction, no enhancement. Enhanced audio scored worse.

=== 2. BUILD THE CORPUS ===

    ffmpeg -i SOURCE.mov -ar 24000 -ac 1 -c:a pcm_s16le audio_24k.wav

Chunk into 4-12s clips: split on . ! ?, merge forward until each clip is >= 4s,
force-split anything over 12s at a comma or a gap > 0.35s.
NEVER discard a segment for being too long.

Print a manifest and check it against the source length:

    clips=360 total_min=45.36 min=4.00 mean=7.56 max=12.00

STOP if total_min is not within ~10% of what I recorded. An early version of my
chunker dropped over-long segments instead of splitting them and silently lost
27 of 52 minutes with no error.

Cut restarts, coughs, off-mic lines and dead air. A bad clip hurts more than a
missing one.

=== 3. TRANSCRIBE ===

openai-whisper, model large-v3, on the GPU. Write train.jsonl with one row per
clip: {"text": ..., "audio": ...}
Correcting the transcripts afterwards is OPTIONAL. I tested it and it did not
improve output quality. Skip unless I ask.

=== 4. PIN THE CORPUS TO HUGGING FACE, BEFORE ANY TRAINING ===

    huggingface-cli repo create my-voice-corpus --type dataset --private
    # upload the clips + train.jsonl

This is not bookkeeping. The corpus must be ONE artifact that every training run
fetches, or you will end up comparing models trained on different data without
realising it. Every training script should fetch the pin and assert the clip
count before starting.

=== 5. START A POD ===

Use the RunPod MCP tools. Ask for a 24GB GPU (RTX 4090 is enough for a 0.5B
model). Community 4090 is about $0.34/hr, secure about $0.74/hr.

CUDA MUST BE >= 12.8. create-pod returns cudaVersion: if it is 12.2 or 12.4,
DELETE the pod and create another. Those hosts crash-loop the container while
the API still reports RUNNING. Each rejected pod costs under a cent.

Then prove the GPU is real before installing anything:

    python -c "import torch; \
      print('GPUCHECK', torch.cuda.is_available(), \
      (torch.randn(999,999,device='cuda') @ torch.randn(999,999,device='cuda')).sum().item())"

If that fails, destroy the pod and get another. Never let training fall back to
CPU: a silent CPU fallback hid a dead GPU from me for an entire run.

=== 6. BUILD THE ENVIRONMENT ===

FIND THE PROJECT'S docker/Dockerfile AND FOLLOW IT. It is the only authoritative
record of the environment. Guessing at it cost me many hours.

For the model I ended up shipping, dots.tts (note: the GitHub org is
studio-dots-ai, while the Hugging Face org is dots-studio, and guessing wrong
makes git prompt for a username instead of reporting a missing repo):

    curl -sL https://raw.githubusercontent.com/studio-dots-ai/dots.tts/main/constraints/recommended.txt \
      | grep -v '^gradio' > constraints.txt      # the package and its own
                                                 # constraints file disagree
    git clone --depth 1 https://github.com/studio-dots-ai/dots.tts.git
    pip install -e dots.tts -c constraints.txt

For CosyVoice 3 specifically:

    # the image may not ship conda at all; install miniconda if `conda` is missing
    conda create -y -n cosyvoice python=3.10 --override-channels -c conda-forge
    conda install -y -n cosyvoice --override-channels -c conda-forge pynini==2.1.5
    pip install "setuptools<81" wheel
    pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cu121
    pip install -r requirements.txt          # hold out openai-whisper
    pip install openai-whisper --no-build-isolation
    pip install huggingface_hub hf_transfer

Four traps, all of which fail SILENTLY or misleadingly:
  - `conda activate` does nothing in a non-interactive shell. Use absolute
    paths: /opt/conda/envs/NAME/bin/python. Otherwise your verification passes
    against the WRONG interpreter.
  - Anaconda default channels now need a ToS accept that fails unattended.
    Use --override-channels -c conda-forge.
  - openai-whisper fails to build inside isolation and aborts the ENTIRE
    requirements install, so nothing else lands.
  - Some images export HF_HUB_ENABLE_HF_TRANSFER=1 without shipping
    hf_transfer, so every Hugging Face download raises until you install it.

Assert every import before training:

    /opt/conda/envs/cosyvoice/bin/python -c \
      "import torch,torchaudio,hyperpyyaml,whisper,pkg_resources; print('DEPS_OK')"

=== 7. TRAIN ===

Save a checkpoint EVERY epoch. Expect ~22 seconds per epoch for a 45-minute
corpus on a 4090, so roughly 11 minutes total.

Checkpoints are ~2GB each, so prune to a sparse ladder as you go or a 150GB
disk fills at about epoch 36.

Do NOT assume the last checkpoint is the best. Audition early ones. Mine peaked at epoch 4 of 31.
Epoch count scales INVERSELY with corpus size: one epoch on 45 min is ~3x the
gradient steps of one epoch on 15 min, so a bigger corpus peaks at a LOWER
epoch number.

=== 8. PUSH EVERY CHECKPOINT TO HUGGING FACE ===

    huggingface-cli repo create my-voice-model --private
    # upload each kept checkpoint, winner first

NON-NEGOTIABLE, and do it BEFORE destroying the pod. Verify by reading the file
sizes back off Hugging Face.

A stopped RunPod pod is NOT a saved pod. /workspace is container overlay
storage, not a network volume. I have seen both failure modes in one hour: a
stopped pod that could not restart ("not enough free GPUs on the host"), and one
that restarted on a different host with /workspace completely empty.

=== 9. RENDER, 3+ TAKES PER CHECKPOINT ===

Generation is stochastic and most of these models take NO seed, so one take is
a lottery draw rather than a measurement.

Render 3 to 5. Then screen them in two passes.

PASS 1, MECHANICAL. Sort the takes by duration and drop any that are materially
shorter than the MEDIAN OF THAT MODEL'S OWN TAKES. Do not use an absolute
words-per-minute threshold: two models I trained read at ~175 and ~150 wpm, so
one cutoff either discards good takes from the fast model or passes bad ones
from the slow one. Five takes of a 67-word script came in at 22.56, 22.88,
22.88, 24.00 and 24.32 seconds; the 22.56 one had silently dropped the closing
sentence and sounded completely fine. A 1.5 second gap against its siblings was
the only evidence.

PASS 2, HUMAN. Listen to what survives and pick by ear.

Do NOT rank takes by Whisper word coverage. I did this for weeks and it is
worse than useless: it scores Whisper's transcription, not the model's speech,
so a robotic take Whisper parses cleanly outranks a natural one it stumbles on.
It put a take I described as sounding like Steve Urkel at the top of a ladder,
reported 24 missing words in a take where the full script was plainly present,
and scored 100% on clips where I could hear words being dropped.

If the inference call returns a GENERATOR, consume ALL of it and concatenate.
CosyVoice yields one segment per sentence group; taking only the first silently
truncated a 29.5s clip to 19.7s for me, with no error.

Duration only catches MISSING SPEECH. It cannot catch a mangled word inside a
full-length take, and it is useless for any model that pads its output to a
constant length: one model I tested returned 25 renders all between 45.85 and
46.18 seconds, some clean and some dropping words. There is no automatic check
for that. A human has to listen.

=== 10. BLIND TEST ===

Normalise every candidate to -16 LUFS, pad to one duration, and force an
identical byte size. Loudness, length and file size all leak which file is
which.

Shuffle the labels, seal the mapping, and do not reveal it until after I have
ranked. Keep sets to 5-6 files.

Balance the arms: do not put 4 checkpoints of one model against 1 of another.
One bad render sinks a lone entry.

=== 11. TEAR DOWN ===

Once checkpoints are pushed AND verified by read-back, TERMINATE the pod. Do
not stop it. Stopping still bills for storage and may lose the disk anyway.

=== REPORT BACK ===

Manifest numbers, per-epoch loss, the checkpoint list, take durations and word
coverage per render. Name what failed.

Two levers moved the result. The rest were a wash or worse.

Fine-tuning instead of zero-shotdecisive
Recording one clean sessionlarge effect
Scraping more clips from old videosa wash
More training stepsa wash
Correcting the transcriptsa wash
Enhancing the audio before trainingscored worse

One clean recording beat everything I scraped

I started by scraping audio out of videos I had already published, because it was free and it was there. My first clone came from 2 minutes 33 seconds of it and was good enough to be my keeper for weeks. Here it is:

2.55 minutes of training data
Built from 13 short takes, 2m33s total
Not bad. Recognisably in the right territory, and it was my keeper for weeks before anything beat it.
VibeVoice-1.5B · 22 clips · peaked at 200 steps

Scraping more of it stopped helping quickly. I built a second corpus of 14 minutes 28 seconds from the same videos and it tied with the 2.5 minute version in a blind ranking.

Then I sat down and recorded 45 minutes in one session, one room, one microphone, one sitting. That corpus produced the CosyVoice and VibeVoice keepers, then trained the OmniVoice adapter that beat them.

Record properly instead of scraping. Clips pulled from videos shot across different rooms and days carry all those rooms into the model. Forty five minutes recorded on purpose is a couple of hours of your time and it is the single biggest quality jump I found.

The peak arrives early

With one model I extended a run from 3200 to 6400 steps because the curve looked like it was still climbing. The first 3200 were the whole gain. My CosyVoice keeper peaked even earlier, at epoch 4 of 31, so plan on auditioning early checkpoints rather than training longer.

The ranking that picked the first keepers

This earlier five-file session contained both ElevenLabs voices, the CosyVoice and VibeVoice fine-tunes, and the real recording. It selected the two keepers I shipped before OmniVoice finished training. The newer sealed tests at the top of the article supersede its overall order.

FileWhat it wasWhat I said, before the reveal
AElevenLabs, scavenged source"bad"
BElevenLabs, cleaner source"good"
CVibeVoice, fine-tuned"good, maybe slightly better. honestly best"
DMe, the real recording"so good that I think it's probably legit me"
ECosyVoice, fine-tuned"good, maybe slightly better than B"

Read C and D together. I called the fine-tuned model "honestly best" and, separately, flagged the real recording as "probably legit me." I did identify the human. I also rated a synthetic clone at the top of the set while doing it. Both paid clones landed below both fine-tunes.

That established the first practical bar for narration, and it stood for weeks. Every model in this session has since been beaten by a fine-tune that did not exist when I ran it. The order at the top of the article is the current one.

What each option actually costs

OptionCostLimitYou own it?
ElevenLabs Creator$22/mo forever100k credits, ~100 minno
ElevenLabs Pro$99/mo forever500k creditsno
Fine-tune your ownunder $1 per run
~$5 to a first clone
plus GPU time to render
noneyour adapter,
subject to the base-model licence

The whole project, including every failed experiment and dozens of training runs, came to about $17 of GPU time. That is less than one month of the cheapest tier that supports professional voice cloning.

I did not provision or run any of this infrastructure by hand. I described what I wanted to an AI coding agent, it stood up the environment, rented the GPU, ran the training, and shut the machine down when it was done. The plan above is the instruction set, and everything it needs to get right is already in it.

The cost is a few dollars of rented GPU and the time it takes you to record.

One honest asterisk on that comparison: the $22 a month buys inference too, and self-hosting does not. Training is a one-off, but every minute of narration afterwards needs a GPU. At my volume that is pennies of pod time per batch. If you are generating a hundred minutes a month you should price it against the tier that covers it, not against the training run.

The 124-word script exposed silent drops

OmniVoice checkpoint 500 read the 67-word verdict script with 100% coverage and zero dropped words. On 124 unseen words, the same checkpoint fell to 86% to 91% coverage. Every take silently dropped 11 to 21 words.

A caveat I owe you about that number. "Coverage" here means the share of scripted words that Whisper heard in the render. I used it to pick takes for weeks, and it is worse than useless as a quality score. It measures Whisper's transcription, not the model's speech, so a robotic render that Whisper parses cleanly outscores a natural one it stumbles on. It ranked a take I described as sounding like Steve Urkel at the top of a ladder. It told me 24 words were missing from a take where the full script was plainly there. It scored 100% on OmniVoice clips where I could hear words being dropped. I keep the numbers below because the direction is real and the drop is large, but I no longer choose anything by them.

Model and scriptWord coverageWords dropped
OmniVoice ck500, 67 words100%0
OmniVoice ck500, 124 words86-91%11-21 every take
CosyVoice keeper, 124 words95-97%3-6
Dots.TTS fine-tuned, 124 wordsno drops heardjudged by ear
Unseen text · 124 words
Dots.TTS fine-tuned, the keeper
Chosen by ear from five takes. On the same 124 words I called it "actually good... because of the lack of collapsing, it's probably better".
45.01s · one of five takes · no coverage score used
Same unseen text · 124 words
OmniVoice + LoRA, checkpoint 800
The best OmniVoice I ever trained, on the same 124 words. "Probably the best... but end of sentences still collapse hard."
46.21s · sounds closest to me · collapses at phrase endings

Length is its own test, and it reorders the table. OmniVoice wins on speaker similarity and then loses continuity as the script grows. CosyVoice held long-form better and was my keeper for exactly that reason, until fine-tuned Dots.TTS beat it on both. A 23-second demo and a ten-minute narration are different tests, and most published demos are the former.

Licence, consent, and the things that decide this in practice

Check the licence before you pick a winner

"You own the weights" in the cost table is doing more work than it looks. You own the adapter you trained, on top of a base model whose terms you did not write. Checked on Hugging Face, 24 August 2026:

Base modelLicence
FunAudioLLM/Fun-CosyVoice3-0.5B-2512Apache 2.0
vibevoice/VibeVoice-1.5BMIT
IndexTeam/IndexTTS-2custom: "bilibili Model Use License Agreement"
Qwen/Qwen3-TTS-12Hz-1.7B-BaseApache 2.0
ResembleAI/chatterboxMIT
coqui/XTTS-v2Coqui Public Model License 1.0.0, non-commercial
Audio8/Audio8-TTS-Preview-0.6bApache 2.0
Audio8/Audio8-TTS-Preview-0.1bcustom: Audio8 Community License v1.0
OpenMOSS-Team/MOSS-TTS-v1.5Apache 2.0
k2-fsa/OmniVoiceCC-BY-NC on the pretrained weights
dots-studio/dots.tts-soarApache 2.0

Four of the eleven systems understate their restriction on the surface. IndexTTS-2 uses the bilibili Model Use License. XTTS-v2 is non-commercial, including its outputs. Audio8's 0.1b model uses a revenue-limited community licence even though its larger 0.6b sibling is Apache 2.0. OmniVoice is the easiest trap: the GitHub code is Apache 2.0 and the Hugging Face card has no licence tag, but the README licenses the pretrained weights under CC-BY-NC because the training data includes Emilia.

That code-versus-weights split also applies to F5-TTS: MIT code, CC-BY-NC weights, again because of Emilia. It remains an interesting research fine-tune, but its public weights are not the commercial next step a licence badge can make them look like.

One provenance note: I trained on vibevoice/VibeVoice-1.5B, a community copy, and Microsoft's first-party microsoft/VibeVoice-1.5B is also live and also MIT. Prefer the first-party repo, and read the licence on the specific checkpoint you pull, since a mirror can carry different terms than its source and terms change between releases.

Clone your own voice, and say that it is synthetic

Everything here is my voice, my recording, my consent. Cloning someone else's without permission is a legal problem in a growing number of places: Tennessee's ELVIS Act covers voice explicitly, right-of-publicity law applies in many US states, and the EU AI Act carries transparency obligations for synthetic media. None of that is legal advice, and all of it is worth ten minutes before you point this at a voice that is not yours. Label synthetic narration as synthetic.

Where this test stops

Eleven systems is what I ran, and the field is bigger. F5-TTS, Fish Speech, Higgs Audio, and Sesame CSM went untried. F5-TTS can be fine-tuned, but its public weights are CC-BY-NC. Higgs Audio v3 ships without training code and is research-only/non-commercial. Sesame CSM also ships without training code, and its public repository has been stale since 27 May 2025.

The long-form question now has a measured answer. OmniVoice was perfect on 67 words and dropped 11 to 21 words from every 124-word take. Fine-tuned Dots.TTS, the keeper, reads the same 124 words without dropping anything I can hear, and CosyVoice held 95% to 97% before it. Ten-minute narration still needs a stitching pass and per-chunk take selection, and I have not run one end to end yet.

Versions

Text to speech moves monthly, so these results are pinned to what I ran in August 2026: dots-studio/dots.tts-soar, OpenMOSS-Team/MOSS-TTS-v1.5, FunAudioLLM/Fun-CosyVoice3-0.5B-2512, vibevoice/VibeVoice-1.5B, IndexTeam/IndexTTS-2, Qwen/Qwen3-TTS-12Hz-1.7B-Base, coqui/XTTS-v2, Audio8/Audio8-TTS-Preview-0.6b, Audio8/Audio8-TTS-Preview-0.1b, OpenMOSS-Team/MOSS-TTS-v1.5, k2-fsa/OmniVoice, dots-studio/dots.tts-soar, and chatterbox-tts from PyPI. Assume a model released after this article behaves differently, and re-run the test rather than trusting the table.

What I would tell someone starting today

Fine-tune Dots.TTS soar. It is Apache 2.0, it is 2.2B so it trains and renders cheaply, and it won my final rounds. If it disappoints you, fine-tuned MOSS-TTS v1.5 is close enough to swap in without changing anything else.

Budget for three to five takes per clip, not one. This is the part I underestimated for weeks. Every model here is stochastic, and the same checkpoint gives you a usable take and an unusable one on the same script. Screen them by duration against each other, then listen to what survives.

For zero-shot, start with Dots.TTS. It is Apache 2.0, it needs no training, and it beat OmniVoice zero-shot head to head. It is also the model that won once fine-tuned, so it is a reasonable starting point either way.

Record one clean session before you touch a model. Forty five minutes from a phone in a quiet room beat everything I scavenged.

Then rank blind, and rank more than one take per model. Every time I evaluated with the labels visible, I confirmed what I already believed. Every time I ranked a single take, I got a result that did not reproduce.

Why I built this

So I can narrate the research instead of filming it

I run NexusTrade, an AI-powered trading research platform, and I publish what it finds. The writing is the work; recording it was the tax. Now the articles narrate themselves and I spend the time on the analysis.

The platform is built the same way this test was: run the experiment, rank it blind, publish the number even when it says you were wrong.

Explore NexusTrade

Build the strategy this article describes

Create a free account to backtest ideas against market history, inspect the risk, and deploy to paper or live markets when you're ready.

or

Free to browse. No credit card required.

Discussion

Sign in or create a free account to join the discussion.

No comments yet.