Chatterbox TTS: voice cloning you can sell
Clone voices with 5s of audio — MIT license, commercial use allowed
- Video memory
- 8 GB
- Setup
- ~10 min
- Access
- port 7860
- Billing
- hourly, in BRL
Chatterbox, by Resemble AI, clones a voice from 5 to 15 seconds of sample audio, speaks 20+ languages and offers an expressiveness dial almost no open competitor has. The decisive part: an MIT license, so you can use the output commercially with no strings.
Voice cloning by Resemble AI under the MIT license: unrestricted commercial use. 20+ languages including Brazilian Portuguese, an expressiveness dial, and cloning from a 5–15s sample. Ready-to-use web UI, audio saved on the machine.
What it is for
- A consistent brand voice for video, podcast and phone menus
- Expressiveness control for livelier or more neutral reads
- MIT license: sell what you generate, no non-commercial clause
- Everything runs on your machine — the reference audio never leaves it
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the Chatterbox TTS template and a machine — the console hides the ones that do not meet the requirement.
- In about 10 minutes the setup finishes and the access address shows up in the panel, on port 7860.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
How is it different from F5-TTS?
Chatterbox is MIT licensed and has an expressiveness dial; F5-TTS tends to sound more neutral. Both fit on the same machine, so testing side by side is cheap.
Why not use XTTS-v2?
XTTS-v2 weights ship under a non-commercial license and the company behind it shut down in 2024. For professional work, Chatterbox is the safe pick.
How long until the first clip?
Setup takes about 10 minutes and the first click downloads the models (another 3 minutes or so). After that each clip takes seconds.