HomeTemplates › Kimi K2 (Moonshot)
🧠 Self-hosted LLM

Kimi K2 for agents and code

Open-weight agentic model for code and multi-step tasks

From
loading…
cheapest machine that meets this template
Deploy in one click →
Video memory
1000 GB
Setup
~15 min
Access
port 8000
Billing
hourly, in BRL

Kimi K2 stands out at tool calling and code — the two things that break most open models once you build a real agent. Here it comes served by vLLM.

Moonshot AI's Kimi K2 served via vLLM with an OpenAI-compatible API — focused on agents, tool use and multi-step coding. NOTE: large model — needs multiple H100/H200 GPUs or quantization. Set the model (e.g. K2.6) via the VLLM_MODEL env var.

What it is for

How to deploy

  1. Create your account and add balance (card or Pix, no subscription).
  2. In the console, pick the Kimi K2 (Moonshot) template and a machine — the console hides the ones that do not meet the requirement.
  3. In about 15 minutes the setup finishes and the access address shows up in the panel, on port 8000.

Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.

FAQ

How much video memory?

Around 480 GB, on a block of large cards.

Is tool calling reliable?

It is one of the most consistent open models at this, though it still needs a well-specified prompt.

Is there a smaller version?

Not of K2 itself. To fit on one card, consider Qwen3 or DeepSeek-R1 Distill.

Run Kimi K2 (Moonshot) today
You only pay for the hours the machine is running.
Deploy in one click →

Related templates

vLLM
OpenAI-compatible API for Llama, Qwen, Mistral and more
TGI (HuggingFace)
HuggingFace Text Generation Inference — vLLM alternative
LiteLLM Proxy
1 OpenAI endpoint routing to 100+ providers (cloud + local)
Ollama
Run DeepSeek, Qwen3, Llama and Mistral with one command