AI & LLM Integration

Fine-Tuning LLMs with LoRA on Domain India VPS

By Domain India Team · DomainIndia EngineeringPublished 8 min read
Knowledge base article
Contents (12 sections)

Fine-tuning used to mean retraining a whole model on a cluster of GPUs. LoRA and QLoRA change that: you train a small adapter on a single rented GPU, then serve it wherever suits you.

Key takeaways

LoRA (Low-Rank Adaptation) makes it practical to fine-tune a 3B-13B open model on a single GPU. Instead of retraining billions of parameters, you train a small adapter: hours instead of days, and a file of tens or hundreds of megabytes instead of many gigabytes. Training needs a GPU, which you rent by the hour; a CPU VPS is useful for data preparation and for serving small models. This guide covers dataset prep, training, evaluation and deployment.

Why fine-tune instead of RAG

Before fine-tuning, ask: does RAG solve the problem? (See our RAG guide.)

NeedRAGFine-tuning
Inject recent factsBest fitWrong tool
Teach specific writing stylePossibleBest fit
Handle domain jargonOKBest fit
Follow strict output formatHit or missBest fit
Reduce hallucinationGoodMarginal
Lower per-request costNo effectYes (a smaller tuned model can do the job)

Fine-tune when you need the model to consistently behave a certain way (tone, format, refusals). RAG when you need it to know new things.

LoRA in one paragraph

LoRA freezes the full model and trains two small low-rank matrices (rank 8-64) alongside chosen weight matrices, usually the attention projections. The output is an adapter file, typically tens to a few hundred megabytes, that applies over the base model. You can keep many adapters on disk and switch between them per task without reloading the base. QLoRA does the same while the frozen base model is loaded in 4-bit, which cuts GPU memory sharply.

Hardware requirements

Rough GPU memory needs (they vary with sequence length and batch size):

Model sizeFull fine-tuneQLoRA (4-bit)Where to train
~3B24 GB+6-8 GBRented GPU
~7-8B60 GB+10-16 GBRented GPU
~13B100 GB+16-24 GBRented GPU

CPU fine-tuning is possible in theory but far too slow to be practical. Rent GPU time by the hour from a cloud GPU provider for training; prices and GPU models change often, so compare current offers.

Domain India VPS and GPUs

Domain India VPS plans are CPU-based: the catalogue lists vCPU, RAM and NVMe storage, with no GPU option. Use a rented GPU for training, then use a VPS for the parts that don't need a GPU: preparing data, storing adapters, and serving a small quantised model or the app that calls it.

Step 1 — Dataset preparation

Fine-tuning needs 100-10,000 high-quality (input, output) examples. Format as JSONL:

training.jsonl:

jsonl
{"instruction":"Write a support reply about a late delivery","input":"My order hasn't arrived yet","output":"Hi, sorry for the wait. Could you share your order number so I can check the courier status..."}
{"instruction":"Write a support reply about a refund","input":"I want to return a damaged item","output":"Hi, I'm sorry the item arrived damaged. Please reply with a photo of the damage and your order number..."}

Quality > quantity. 500 excellent examples beat 10,000 mediocre ones.

Good sources for data:

  • Past support tickets (anonymised: strip names, phone numbers, emails and any ID numbers)
  • KB articles (convert Q-style FAQ pairs)
  • Your brand voice guide examples
  • Customer-approved responses

Step 2 — Setup environment

On a rented GPU machine (a current Ubuntu LTS or AlmaLinux image with the NVIDIA driver installed):

bash
# Python + deps
python3 -m venv venv
source venv/bin/activate
# Install PyTorch with the command pytorch.org gives for your CUDA version
pip install torch
pip install transformers datasets peft accelerate bitsandbytes trl

# Hugging Face login (for gated models such as Llama)
hf auth login        # older versions: huggingface-cli login

Step 3 — Train (QLoRA)

train.py:

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, prepare_model_for_kbit_training
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig

BASE_MODEL = "meta-llama/Llama-3.2-3B-Instruct"

# 4-bit quantization config (a 3B model then fits in roughly 6-8 GB of VRAM)
# bf16 needs a recent GPU (Ampere or newer); on older GPUs use fp16=True, bf16=False
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    quantization_config=bnb_config,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL)
tokenizer.pad_token = tokenizer.eos_token

# LoRA config
lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)

model = prepare_model_for_kbit_training(model)
# SFTTrainer applies lora_config (passed as peft_config below)

# Load dataset
def format_prompt(example):
    return (f"### Instruction: {example['instruction']}\n"
            f"### Input: {example['input']}\n"
            f"### Response: {example['output']}{tokenizer.eos_token}")

dataset = load_dataset("json", data_files="training.jsonl", split="train")
dataset = dataset.map(lambda e: {"text": format_prompt(e)})

# Train
trainer = SFTTrainer(
    model=model,
    processing_class=tokenizer,
    train_dataset=dataset,
    args=SFTConfig(
        output_dir="./adapter",
        num_train_epochs=3,
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,
        learning_rate=2e-4,
        fp16=False,
        bf16=True,
        logging_steps=10,
        save_strategy="epoch",
        dataset_text_field="text",
        max_length=2048,        # called max_seq_length in older TRL versions
    ),
    peft_config=lora_config,
)

trainer.train()
trainer.save_model("./adapter-final")

Run:

bash
python train.py
# Watch loss decrease over epochs
# Output: ./adapter-final/adapter_model.safetensors plus adapter_config.json

Step 4 — Inference with the adapter

python
from peft import PeftModel

model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    quantization_config=bnb_config,
    device_map="auto",
)
model = PeftModel.from_pretrained(model, "./adapter-final")

def generate(prompt):
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
    outputs = model.generate(**inputs, max_new_tokens=300)
    return tokenizer.decode(outputs[0], skip_special_tokens=True)

print(generate("### Instruction: Write a support reply about a late delivery\n"
               "### Input: Where is my parcel?\n"
               "### Response:"))

Step 5 — Merge adapter for production

To simplify serving, merge the LoRA adapter into the base weights:

python
from peft import AutoPeftModelForCausalLM

# Load the base in 16-bit (not 4-bit) to merge; this needs enough RAM or VRAM for the full model
model = AutoPeftModelForCausalLM.from_pretrained("./adapter-final", torch_dtype=torch.bfloat16)
merged = model.merge_and_unload()
merged.save_pretrained("./merged-model")
tokenizer.save_pretrained("./merged-model")

Now ./merged-model is a standalone model tuned to your task. Check the base model's licence before you publish or sell it. Serve it with vLLM on a GPU, or convert and quantise it (for example to GGUF) to run small models on CPU with Ollama or llama.cpp. See the self-hosting LLMs guide.

Step 6 — Evaluate

Hold out 10% of your dataset as eval set. Measure:

  • Exact match — output exactly matches expected
  • BLEU / ROUGE — similarity scores (useful for format tasks, weak for open-ended answers)
  • Human eval — does a reviewer prefer fine-tuned output?

Simple script:

python
from rouge_score import rouge_scorer

scorer = rouge_scorer.RougeScorer(['rouge1', 'rougeL'], use_stemmer=True)
scores = []
for row in eval_set:
    predicted = generate(row['prompt'])
    s = scorer.score(row['expected'], predicted)
    scores.append(s['rougeL'].fmeasure)
print(f"Average ROUGE-L: {sum(scores)/len(scores):.3f}")

Common pitfalls

Overfitting
The model memorises training data and fails on new inputs. Use a small rank (r=8-16), add dropout, and stop after 2-3 epochs.
Catastrophic forgetting
The model loses general ability. Keep training data varied; include some general examples, not only domain-specific ones.
Low-quality data
Garbage in, garbage out. Spend most of your time on data quality, not training settings.
Wrong target modules
q_proj and v_proj alone work for many models; adding k_proj and o_proj (or all linear layers) often improves results at extra cost.
Training on CPU
Possible in theory, impractical in practice. Rent GPU time.
Loading an adapter without its base model
Adapters are not standalone. Always load the same base model first.

FAQ

How much does LoRA fine-tuning cost?

Mostly GPU rental time. A QLoRA run on a 7B model with about a thousand examples usually takes a few GPU hours, so check current hourly GPU prices and multiply. Open-weight models have no per-token fee, but read each model's licence.

Can I fine-tune hosted models from OpenAI or Anthropic?

Some providers offer fine-tuning for some of their models, through their own platforms or cloud partners, and the list changes. Check each provider's current documentation. If you need full control over the weights and cost, an open-weight model with LoRA is the usual route.

Should I fine-tune for RAG?

Generally no. RAG + a good general model outperforms fine-tuned-without-RAG for factual queries. Fine-tune for style/format; use RAG for knowledge.

How often to re-train?

When you move to a newer base model, or when your data or requirements change. Keep your dataset and training script in version control so re-training is routine.

What about DPO / RLHF?

They tune a model on preferences (a better and a worse answer) rather than on single examples. Start with supervised fine-tuning with LoRA; try DPO, which TRL also supports, only if you need preference-based tuning.

Running this on Domain India

  • Training: needs a GPU. Domain India VPS plans are CPU-only, so rent GPU time elsewhere for the training run.
  • Data preparation and storage: a self-managed VPS with full root access works well for cleaning datasets, running evaluation scripts against an API, and keeping adapters and datasets.
  • Serving: small quantised models can run on CPU with Ollama or llama.cpp on a VPS; size RAM to the model. Larger models need a GPU host.
  • Shared hosting: not suitable. Long-running processes are stopped there, and there is no GPU.
  • Backups: VPS plans include no backups or snapshots, so copy your datasets and adapters somewhere safe.

Ready to put a tuned model to work? Read the self-hosting LLMs guide, compare VPS plans, or open a ticket with hosting questions.

Serve your models on a VPS

A self-managed VPS with full root access for data pipelines, evaluation and CPU inference of small models.

View VPS plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app