AI & LLM Integration

Self-Hosting Llama 3 and Mistral LLMs on Domain India VPS (Ollama + vLLM)

By Domain India Team · DomainIndia EngineeringPublished 10 min read
Knowledge base article
Contents (14 sections)

Running an open-weight language model (Llama, Mistral, Qwen, Gemma) on your own server keeps prompts and documents on hardware you control and replaces per-token API bills with a fixed monthly cost. This guide covers sizing, Ollama for the quickest setup, llama.cpp for the best CPU performance, vLLM for GPU throughput, and how to put the model safely behind an API.

Key takeaways

Small open models (3B–8B parameters, 4-bit quantised) run usefully on a CPU-only Linux VPS with enough RAM; the RAM, not the CPU, decides which model fits. Start with Ollama, keep it bound to localhost, and put nginx with authentication or your own API gateway in front. Use vLLM only when you have a GPU and many concurrent users.

Why self-host an LLM

Hosted APIs (OpenAI, Anthropic Claude, Google Gemini) are fast and very capable, but a self-hosted model makes sense when:

  • Privacy — customer data, medical records or contracts must not leave your server
  • Predictable cost — heavy, steady usage costs a fixed server fee instead of a per-token bill
  • Customisation — you want to fine-tune on your own vocabulary or documents
  • Offline or sovereign use — no dependency on a foreign API being reachable

Trade-offs:

  • Quality gap — small open models are well behind the largest hosted models on hard reasoning and coding
  • Ops burden — you handle updates, downtime and scaling
  • Speed on CPU — expect a few tokens per second on a 7B–8B model, not the near-instant replies of a hosted API

What fits in how much RAM

Rule of thumb: a 4-bit quantised model needs roughly (parameters in billions × 0.6) GB of RAM, plus room for the context and the operating system. An 8B model needs about 5 GB; a 70B model needs 40 GB or more.

Server RAMComfortable 4-bit model (CPU)Rough speed on CPUTypical use
2 GB1B–3B (Llama 3.2 3B, Gemma 3 1B)ModerateClassification, short replies
4 GB3B–4B (Phi-4-mini, Gemma 3 4B, Qwen3 4B)ModerateChatbot, summarisation
8 GB7B–8B (Mistral 7B, Llama 3.1 8B, Qwen3 8B)Slow (a few tokens/sec)Q&A, RAG over your documents
16 GB12B–14B (Gemma 3 12B, Qwen3 14B)SlowBetter quality RAG, drafting
32 GBUp to ~27B–32B (Gemma 3 27B, Qwen3 32B)Very slow on CPUBatch jobs, overnight processing

Speeds depend on the CPU generation, the number of cores and the quantisation, so benchmark on your own server before you promise response times.

Insight

CPU inference is slower but workable for many jobs. A support bot that answers in a few seconds is still usable, and batch work such as nightly summarisation or tagging does not care about speed. For many users at once, or large models, you need a GPU.

Option A — Ollama (easiest)

Ollama bundles model download, serving and a REST API in one program.

  1. SSH into your server as root (or a sudo user).
  2. Install:
    bash
    curl -fsSL https://ollama.com/install.sh | sh
  3. Download a model (sizes are approximate for the default 4-bit builds):
    bash
    ollama pull llama3.2:3b     # ~2 GB, fast
    ollama pull mistral:7b      # ~4 GB, balanced
    ollama pull qwen3:8b        # ~5 GB, strong multilingual
  4. Test:
    bash
    ollama run mistral:7b "Explain how a TCP handshake works in 3 sentences."
  5. Call the REST API (Ollama listens on 127.0.0.1:11434 by default):
    bash
    curl http://127.0.0.1:11434/api/generate -d '{
      "model": "mistral:7b",
      "prompt": "Summarise: ...",
      "stream": false
    }'

Ollama also offers an OpenAI-compatible endpoint at /v1/chat/completions, so most OpenAI SDKs work by changing the base URL.

systemd + nginx front

On Linux the install script creates and starts an ollama systemd service. To keep a model loaded instead of unloading it after 5 idle minutes, set OLLAMA_KEEP_ALIVE:

bash
sudo systemctl edit ollama
# add:
# [Service]
# Environment="OLLAMA_KEEP_ALIVE=30m"
sudo systemctl restart ollama

To reach it from outside, leave Ollama on localhost and publish it through nginx with TLS and authentication:

nginx
server {
    listen 443 ssl;
    server_name llm.yourcompany.com;
    # ssl_certificate / ssl_certificate_key: issue with certbot (Let's Encrypt)

    location / {
        auth_basic "LLM";
        auth_basic_user_file /etc/nginx/.htpasswd-llm;
        proxy_pass http://127.0.0.1:11434;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_buffering off;          # stream tokens as they arrive
        proxy_read_timeout 300s;
    }
}
Warning

Never expose Ollama unauthenticated to the internet. Open Ollama ports are scanned for and abused: strangers run models on your CPU, pull or delete models and use your bandwidth. Do not set OLLAMA_HOST=0.0.0.0 on a public server; keep it on localhost behind nginx with authentication or your own gateway.

Option B — llama.cpp (best CPU performance)

Ollama is built on llama.cpp. Use llama.cpp directly when you want full control of threads, context and quantisation.

bash
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config Release -j "$(nproc)"

# Download a GGUF model file (Q4_K_M is a good default) from Hugging Face,
# then serve it on localhost only:
./build/bin/llama-server -m models/your-model-Q4_K_M.gguf \
    -c 4096 --host 127.0.0.1 --port 8080

llama-server also exposes an OpenAI-compatible API. Put it behind nginx with authentication, exactly as for Ollama.

Option C — vLLM (GPU throughput)

vLLM serves many concurrent requests efficiently with continuous batching. It is built for GPUs; its CPU backend exists but is rarely the right choice, so on a CPU-only server prefer Ollama or llama.cpp.

bash
python3 -m venv ~/vllm && source ~/vllm/bin/activate
pip install vllm

# OpenAI-compatible server
vllm serve meta-llama/Llama-3.2-3B-Instruct \
    --host 127.0.0.1 --port 8000 \
    --max-model-len 8192 \
    --dtype auto

Meta's Llama models on Hugging Face are gated: accept the licence on the model page and log in with huggingface-cli login first.

Then call it with any OpenAI SDK:

python
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
    model="meta-llama/Llama-3.2-3B-Instruct",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)

The rest of your app does not need to know the model is self-hosted.

Picking a model

Open models change quickly; check each model's licence and the current release before you commit.

Model familySizes to try on CPUStrengthsWatch out for
Llama 3.1 / 3.2 (Meta)3B, 8BGood general EnglishCommunity licence terms; weaker in Indian languages
Mistral 7B7BBalanced, Apache 2.0 licenceOlder; weaker multilingual
Qwen3 (Alibaba)4B, 8B, 14BStrong multilingual and codeLarger sizes are slow on CPU
Gemma 3 (Google)1B, 4B, 12BEfficient, multilingualGemma licence terms
Phi-4-mini (Microsoft)3.8BSmall, capable reasoningLimited world knowledge

For Indian languages, test multilingual families such as Qwen3 and Gemma 3, and models built in India for Indic languages, for example from Sarvam AI. Always evaluate on your own sample prompts in Hindi, Tamil or whichever language you need.

Production patterns

Pattern 1 — API key gateway

Don't expose the model server directly. Put a small API in front that validates keys, rate-limits per key, logs usage and can fall back to a paid API.

python
# FastAPI gateway example (pip install fastapi uvicorn httpx)
import os, time
from collections import defaultdict
from fastapi import FastAPI, HTTPException, Depends, Header
import httpx

app = FastAPI()
VALID_KEYS = set(os.environ["GATEWAY_KEYS"].split(","))   # comma-separated
hits = defaultdict(int)   # in-memory; use a database for several processes

async def validate_key(x_api_key: str = Header()):
    if x_api_key not in VALID_KEYS:
        raise HTTPException(401, "Invalid key")
    return x_api_key

@app.post("/v1/chat/completions")
async def chat(req: dict, api_key: str = Depends(validate_key)):
    minute = int(time.time() // 60)
    hits[(api_key, minute)] += 1
    if hits[(api_key, minute)] > 60:
        raise HTTPException(429, "Rate limit")

    async with httpx.AsyncClient(timeout=300) as client:
        resp = await client.post("http://127.0.0.1:11434/v1/chat/completions", json=req)
    return resp.json()

Pattern 2 — Keep the model warm

Loading a model into RAM takes seconds to tens of seconds. Rather than pinging it from cron, set OLLAMA_KEEP_ALIVE (above), or pass "keep_alive": "30m" in API requests. Leave enough free RAM for the model to stay resident.

Pattern 3 — Fallback chain

python
try:
    return await local_chat(prompt)     # self-hosted
except (TimeoutError, httpx.HTTPError):
    return await hosted_chat(prompt)    # paid fallback

You pay the API only during peaks or outages. Tell users if their data may go to a third-party API in that case.

Common pitfalls

Inference crawls at 1 token/sec
The server is swapping. Use a smaller model or quantisation, or more RAM.
"Out of memory" on model load
Stop other services, pick a smaller model, or move to a server with more RAM.
Terrible output quality
Over-aggressive quantisation (Q2 is often unusable) or the wrong model. Start with Q4_K_M.
Slow with several users
Ollama and llama.cpp on CPU handle few concurrent requests well. Queue requests, or move to a GPU with vLLM.
Disk fills up
Models are large (Ollama stores them under /usr/share/ollama/.ollama/models for the service). Remove unused ones with ollama rm.
Context silently truncated
Long prompts beyond the context window give wrong answers. Set num_ctx and check input length.

Running this on Domain India

LLM serving needs root access, several GB of RAM and a long-running process, so it belongs on a VPS, not shared hosting or the App Platform.

  • Domain India VPS plans are self-managed KVM servers with full root access and NVMe storage. The catalogue lists RAM from 2 GB (VPS Starter) up to 32 GB (VPS Enterprise); no GPU is listed, so plan for CPU inference with the model sizes in the table above.
  • You install and update the software yourself. No backups or snapshots are included with a VPS, so keep copies of your configuration and data elsewhere.
  • Shared hosting (cPanel, DirectAdmin, Webuzo) stops long-running processes and cannot run a model server. You can still call a model on your VPS, or a hosted API, from a PHP or Node.js site on shared hosting over HTTPS.

FAQ

Do I need a GPU to self-host an LLM?

No. Small quantised models (about 3B to 8B parameters) run on a CPU-only server with enough RAM, at a few tokens per second for the larger ones. That is fine for many chatbots and batch jobs. You need a GPU for many concurrent users, fast responses or large models.

Is self-hosting cheaper than a paid API?

It depends on volume. At low usage, a pay-per-token API is usually cheaper and better. With steady, heavy usage, a fixed-price server can cost less, especially if you already need one. Compare your real token usage against the provider's current prices and the server's monthly cost.

Can I fine-tune on my own data?

Yes. LoRA or QLoRA fine-tuning is practical on a GPU; on a CPU it is very slow. Try retrieval-augmented generation (RAG) over your documents first: it is cheaper and often enough.

Llama, Mistral or Qwen?

Llama for general English, Mistral 7B for a balanced, permissively licensed model, and Qwen3 or Gemma 3 for multilingual work including Indian languages. Test two or three on your own prompts before you choose.

Can Ollama serve more than one model?

Yes. Pull the models you need; Ollama loads them on demand and unloads idle ones after 5 minutes by default. Change that with the OLLAMA_KEEP_ALIVE setting, and make sure the server has RAM for every model you keep loaded.

Can I run Ollama on shared hosting?

No. Shared hosting stops long-running background processes and gives no root access. Run the model on a VPS and call it from your website over HTTPS.

Ready to run your own model? Compare VPS plans, or open a support ticket if you are unsure which size fits.

Self-host your own LLM

Self-managed KVM VPS with full root access and NVMe storage, from ₹553 a month excluding GST.

View VPS plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app
Self-Host Llama and Mistral on a VPS: Ollama, vLLM