HomeTemplates › LLaMA-Factory
📓 Notebooks & Dev

LLM fine-tuning from a web interface

No-code LLM fine-tuning with LoRA/QLoRA

From
loading…
cheapest machine that meets this template
Deploy in one click →
Video memory
16 GB
Setup
~10 min
Access
port 7860
Billing
hourly, in BRL

LLaMA-Factory makes fine-tuning approachable: pick the base model, point at your dataset and watch the run on screen. LoRA and QLoRA let you adapt large models on a single card.

LLaMA-Factory is a web UI (LLaMA Board) for fine-tuning open LLMs (Llama, Qwen, Mistral, etc.) with LoRA/QLoRA, no code required. QLoRA of 7B-13B models fits a 24GB GPU (RTX 4090). Includes example datasets, evaluation and adapter export.

What it is for

How to deploy

  1. Create your account and add balance (card or Pix, no subscription).
  2. In the console, pick the LLaMA-Factory template and a machine — the console hides the ones that do not meet the requirement.
  3. In about 10 minutes the setup finishes and the access address shows up in the panel, on port 7860.

Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.

FAQ

How much video memory do I need?

With QLoRA, 16 GB already fine-tunes mid-sized models. Larger models or full training want 40 GB or more.

How many examples do I need?

A few thousand well-chosen examples usually beat tens of thousands of poorly filtered ones.

Where do the trained weights end up?

On your machine, ready to download or serve right there.

Run LLaMA-Factory today
You only pay for the hours the machine is running.
Deploy in one click →

Related templates

JupyterLab CUDA
GPU notebook with PyTorch, Transformers and Diffusers ready