๐Ÿ“ The ๐Ÿ—ฃ๏ธ Open TTS Leaderboard benchmarks open-source text-to-speech models on intelligibility, speed and speaker similarity.

Every model is scored on the same texts across two multilingual eval sets (Seed-TTS eval and CV3-Eval) for the language it supports, both without voice cloning (the model's own default voice) and โ€” where the model supports it โ€” with voice cloning from a reference clip.

โš ๏ธ TO REQUEST A MODEL/DATASET/METRIC: please reach out via LinkedIn or X. Source code will be open-sourced shortly for requests via GitHub Issues/PRs!

๐Ÿ“… Results version

Pick a past snapshot to see the leaderboard as it stood then.

๐Ÿ… Leaderboard

Averages across the selected languages and datasets. A model is ranked, with a headline average, only if it has results for every selected language on every selected dataset, so all ranked models are compared on the same set.

If batch inference is supported, the model is evaluated with batch size 32; otherwise, it is evaluated at batch size 1. See the ๐Ÿ—ฃ๏ธ Streaming tab for single-request, streaming-mode latency instead.

Evals are run with HF jobs, flavor h200: 1x H200 (141 GB), 23 vCPU, 256 GB RAM, 3000 GB storage

See the ๐Ÿค— About tab for metric definitions (WER, RTFx, SIM).

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

๐Ÿ“Š Pareto frontiers

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

๐Ÿ† Top models

Datasets
Languages

Only models evaluated on every selected language and dataset are listed.

Show only models that can clone a reference voice, scored on the cloning runs (adds speaker similarity, SIM).

Last updated on 18 September 2026