Qwen3 with your own API
Multilingual Qwen3-8B (great PT-BR) on a single GPU
- Video memory
- 20 GB
- Setup
- ~6 min
- Access
- port 8000
- Billing
- hourly, in BRL
Qwen3 is one of the strongest open families for multilingual work and ships under a permissive license. With vLLM in front, it serves your application through the same API you already use.
Alibaba's Qwen3-8B — strong multilingual model (great PT-BR) that runs on an RTX 4090. Served via vLLM with an OpenAI-compatible API. For more capability, switch to Qwen3-14B/32B (more VRAM) or see the Qwen3-235B template.
What it is for
- Support and summarisation without sending data outside
- Batch classification and extraction at a fixed hourly cost
- A base to fine-tune on your company's vocabulary
- Permissive license, no blocker for commercial use
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the Qwen3 template and a machine — the console hides the ones that do not meet the requirement.
- In about 6 minutes the setup finishes and the access address shows up in the panel, on port 8000.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
Which size should I pick?
The 8–14 billion versions fit on 24 GB and already perform well. Above that, look at 80 GB cards.
How is it in languages other than English?
It is one of the open models that slips least in grammar and vocabulary across languages, Portuguese included.
Can I use it commercially?
Yes, the license is permissive. Always check the exact release you download.