Kimi K2 for agents and code
Open-weight agentic model for code and multi-step tasks
- Video memory
- 1000 GB
- Setup
- ~15 min
- Access
- port 8000
- Billing
- hourly, in BRL
Kimi K2 stands out at tool calling and code — the two things that break most open models once you build a real agent. Here it comes served by vLLM.
Moonshot AI's Kimi K2 served via vLLM with an OpenAI-compatible API — focused on agents, tool use and multi-step coding. NOTE: large model — needs multiple H100/H200 GPUs or quantization. Set the model (e.g. K2.6) via the VLLM_MODEL env var.
What it is for
- An agent that uses tools without losing the thread
- Code generation and review over a large codebase
- An open alternative for automation that depends on a paid API today
- Context and tools inside your own infrastructure
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the Kimi K2 (Moonshot) template and a machine — the console hides the ones that do not meet the requirement.
- In about 15 minutes the setup finishes and the access address shows up in the panel, on port 8000.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
How much video memory?
Around 480 GB, on a block of large cards.
Is tool calling reliable?
It is one of the most consistent open models at this, though it still needs a well-specified prompt.
Is there a smaller version?
Not of K2 itself. To fit on one card, consider Qwen3 or DeepSeek-R1 Distill.