F5-TTS in the cloud, interface included
Clone any voice in 5 seconds of audio
- Video memory
- 8 GB
- Setup
- ~4 min
- Access
- port 7860
- Billing
- hourly, in BRL
F5-TTS clones a voice from a few seconds of audio and produces studio-grade narration. Spin up the machine, open the interface in your browser and start generating — nothing to install locally.
State-of-the-art Text-to-Speech model with instant voice cloning. Supports Brazilian Portuguese, English and 6 more languages. Ready-to-use web interface.
What it is for
- Video and audiobook narration in your own voice
- Ad and training voiceover without a studio
- Batch generation through the API for catalogs of thousands of lines
- Audio stays on your machine: nothing goes to a third-party service
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the F5-TTS template and a machine — the console hides the ones that do not meet the requirement.
- In about 4 minutes the setup finishes and the access address shows up in the panel, on port 7860.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
Which GPU is enough?
From 8 GB of video memory it runs well. 16 to 24 GB cards make batch generation noticeably faster.
How many languages does it support?
Pretrained weights cover several languages, including Portuguese. Clean reference samples with no background noise matter more than anything else.
How much does an hour of audio cost?
You pay for the machine by the hour, not per character. An entry-level GPU produces many minutes of audio per rented hour, which usually lands well below per-character pricing.