๐Ÿ“ The ๐Ÿ—ฃ๏ธ Open TTS Leaderboard benchmarks open-source text-to-speech models on intelligibility, speed and speaker similarity.

Every model is scored on the same texts across two multilingual eval sets (Seed-TTS eval and CV3-Eval) for the language it supports, both without voice cloning (the model's own default voice) and โ€” where the model supports it โ€” with voice cloning from a reference clip.

โš ๏ธ TO REQUEST A MODEL/DATASET/METRIC: open a pull request or issue on the GitHub repo! See here for adding a new model.

๐Ÿ“… Results version

Pick a past snapshot to see the leaderboard as it stood then.

๐Ÿ… Leaderboard

Averages across the selected languages and datasets. A model is ranked, with a headline average, only if it has results for every selected language on every selected dataset, so all ranked models are compared on the same set.

If batch inference is supported, the model is evaluated with batch size 32; otherwise, it is evaluated at batch size 1. See the ๐Ÿ—ฃ๏ธ Streaming tab for single-request, streaming-mode latency instead.

Evals are run with HF jobs, flavor h200: 1x H200 (141 GB), 23 vCPU, 256 GB RAM, 3000 GB storage

See the ๐Ÿค— About tab for metric definitions (WER, RTFx, SIM).

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

๐Ÿ“Š Pareto frontiers

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

๐Ÿ† Top models

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

Last updated on 9 October 2026