HomeTemplates › Qwen3
🧠 Self-hosted LLM

Qwen3 with your own API

Multilingual Qwen3-8B (great PT-BR) on a single GPU

From
loading…
cheapest machine that meets this template
Deploy in one click →
Video memory
20 GB
Setup
~6 min
Access
port 8000
Billing
hourly, in BRL

Qwen3 is one of the strongest open families for multilingual work and ships under a permissive license. With vLLM in front, it serves your application through the same API you already use.

Alibaba's Qwen3-8B — strong multilingual model (great PT-BR) that runs on an RTX 4090. Served via vLLM with an OpenAI-compatible API. For more capability, switch to Qwen3-14B/32B (more VRAM) or see the Qwen3-235B template.

What it is for

How to deploy

  1. Create your account and add balance (card or Pix, no subscription).
  2. In the console, pick the Qwen3 template and a machine — the console hides the ones that do not meet the requirement.
  3. In about 6 minutes the setup finishes and the access address shows up in the panel, on port 8000.

Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.

FAQ

Which size should I pick?

The 8–14 billion versions fit on 24 GB and already perform well. Above that, look at 80 GB cards.

How is it in languages other than English?

It is one of the open models that slips least in grammar and vocabulary across languages, Portuguese included.

Can I use it commercially?

Yes, the license is permissive. Always check the exact release you download.

Run Qwen3 today
You only pay for the hours the machine is running.
Deploy in one click →

Related templates

vLLM
OpenAI-compatible API for Llama, Qwen, Mistral and more
TGI (HuggingFace)
HuggingFace Text Generation Inference — vLLM alternative
LiteLLM Proxy
1 OpenAI endpoint routing to 100+ providers (cloud + local)
Ollama
Run DeepSeek, Qwen3, Llama and Mistral with one command