Running an open-weight language model (Llama, Mistral, Qwen, Gemma) on your own server keeps prompts and documents on hardware you control and replaces per-token API bills with a fixed monthly cost. This guide covers sizing, Ollama for the quickest setup, llama.cpp for the best CPU performance, vLLM for GPU throughput, and how to put the model safely behind an API.
Small open models (3B–8B parameters, 4-bit quantised) run usefully on a CPU-only Linux VPS with enough RAM; the RAM, not the CPU, decides which model fits. Start with Ollama, keep it bound to localhost, and put nginx with authentication or your own API gateway in front. Use vLLM only when you have a GPU and many concurrent users.
Why self-host an LLM
Hosted APIs (OpenAI, Anthropic Claude, Google Gemini) are fast and very capable, but a self-hosted model makes sense when:
- Privacy — customer data, medical records or contracts must not leave your server
- Predictable cost — heavy, steady usage costs a fixed server fee instead of a per-token bill
- Customisation — you want to fine-tune on your own vocabulary or documents
- Offline or sovereign use — no dependency on a foreign API being reachable
Trade-offs:
- Quality gap — small open models are well behind the largest hosted models on hard reasoning and coding
- Ops burden — you handle updates, downtime and scaling
- Speed on CPU — expect a few tokens per second on a 7B–8B model, not the near-instant replies of a hosted API
What fits in how much RAM
Rule of thumb: a 4-bit quantised model needs roughly (parameters in billions × 0.6) GB of RAM, plus room for the context and the operating system. An 8B model needs about 5 GB; a 70B model needs 40 GB or more.
| Server RAM | Comfortable 4-bit model (CPU) | Rough speed on CPU | Typical use |
|---|---|---|---|
| 2 GB | 1B–3B (Llama 3.2 3B, Gemma 3 1B) | Moderate | Classification, short replies |
| 4 GB | 3B–4B (Phi-4-mini, Gemma 3 4B, Qwen3 4B) | Moderate | Chatbot, summarisation |
| 8 GB | 7B–8B (Mistral 7B, Llama 3.1 8B, Qwen3 8B) | Slow (a few tokens/sec) | Q&A, RAG over your documents |
| 16 GB | 12B–14B (Gemma 3 12B, Qwen3 14B) | Slow | Better quality RAG, drafting |
| 32 GB | Up to ~27B–32B (Gemma 3 27B, Qwen3 32B) | Very slow on CPU | Batch jobs, overnight processing |
Speeds depend on the CPU generation, the number of cores and the quantisation, so benchmark on your own server before you promise response times.
CPU inference is slower but workable for many jobs. A support bot that answers in a few seconds is still usable, and batch work such as nightly summarisation or tagging does not care about speed. For many users at once, or large models, you need a GPU.
Option A — Ollama (easiest)
Ollama bundles model download, serving and a REST API in one program.
- SSH into your server as root (or a sudo user).
- Install:
bash curl -fsSL https://ollama.com/install.sh | sh - Download a model (sizes are approximate for the default 4-bit builds):
bash ollama pull llama3.2:3b # ~2 GB, fast ollama pull mistral:7b # ~4 GB, balanced ollama pull qwen3:8b # ~5 GB, strong multilingual - Test:
bash ollama run mistral:7b "Explain how a TCP handshake works in 3 sentences." - Call the REST API (Ollama listens on
127.0.0.1:11434by default):bash curl http://127.0.0.1:11434/api/generate -d '{ "model": "mistral:7b", "prompt": "Summarise: ...", "stream": false }'
Ollama also offers an OpenAI-compatible endpoint at /v1/chat/completions, so most OpenAI SDKs work by changing the base URL.
systemd + nginx front
On Linux the install script creates and starts an ollama systemd service. To keep a model loaded instead of unloading it after 5 idle minutes, set OLLAMA_KEEP_ALIVE:
sudo systemctl edit ollama
# add:
# [Service]
# Environment="OLLAMA_KEEP_ALIVE=30m"
sudo systemctl restart ollamaTo reach it from outside, leave Ollama on localhost and publish it through nginx with TLS and authentication:
server {
listen 443 ssl;
server_name llm.yourcompany.com;
# ssl_certificate / ssl_certificate_key: issue with certbot (Let's Encrypt)
location / {
auth_basic "LLM";
auth_basic_user_file /etc/nginx/.htpasswd-llm;
proxy_pass http://127.0.0.1:11434;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_buffering off; # stream tokens as they arrive
proxy_read_timeout 300s;
}
}Never expose Ollama unauthenticated to the internet. Open Ollama ports are scanned for and abused: strangers run models on your CPU, pull or delete models and use your bandwidth. Do not set OLLAMA_HOST=0.0.0.0 on a public server; keep it on localhost behind nginx with authentication or your own gateway.
Option B — llama.cpp (best CPU performance)
Ollama is built on llama.cpp. Use llama.cpp directly when you want full control of threads, context and quantisation.
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config Release -j "$(nproc)"
# Download a GGUF model file (Q4_K_M is a good default) from Hugging Face,
# then serve it on localhost only:
./build/bin/llama-server -m models/your-model-Q4_K_M.gguf \
-c 4096 --host 127.0.0.1 --port 8080llama-server also exposes an OpenAI-compatible API. Put it behind nginx with authentication, exactly as for Ollama.
Option C — vLLM (GPU throughput)
vLLM serves many concurrent requests efficiently with continuous batching. It is built for GPUs; its CPU backend exists but is rarely the right choice, so on a CPU-only server prefer Ollama or llama.cpp.
python3 -m venv ~/vllm && source ~/vllm/bin/activate
pip install vllm
# OpenAI-compatible server
vllm serve meta-llama/Llama-3.2-3B-Instruct \
--host 127.0.0.1 --port 8000 \
--max-model-len 8192 \
--dtype autoMeta's Llama models on Hugging Face are gated: accept the licence on the model page and log in with huggingface-cli login first.
Then call it with any OpenAI SDK:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="meta-llama/Llama-3.2-3B-Instruct",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)The rest of your app does not need to know the model is self-hosted.
Picking a model
Open models change quickly; check each model's licence and the current release before you commit.
| Model family | Sizes to try on CPU | Strengths | Watch out for |
|---|---|---|---|
| Llama 3.1 / 3.2 (Meta) | 3B, 8B | Good general English | Community licence terms; weaker in Indian languages |
| Mistral 7B | 7B | Balanced, Apache 2.0 licence | Older; weaker multilingual |
| Qwen3 (Alibaba) | 4B, 8B, 14B | Strong multilingual and code | Larger sizes are slow on CPU |
| Gemma 3 (Google) | 1B, 4B, 12B | Efficient, multilingual | Gemma licence terms |
| Phi-4-mini (Microsoft) | 3.8B | Small, capable reasoning | Limited world knowledge |
For Indian languages, test multilingual families such as Qwen3 and Gemma 3, and models built in India for Indic languages, for example from Sarvam AI. Always evaluate on your own sample prompts in Hindi, Tamil or whichever language you need.
Production patterns
Pattern 1 — API key gateway
Don't expose the model server directly. Put a small API in front that validates keys, rate-limits per key, logs usage and can fall back to a paid API.
# FastAPI gateway example (pip install fastapi uvicorn httpx)
import os, time
from collections import defaultdict
from fastapi import FastAPI, HTTPException, Depends, Header
import httpx
app = FastAPI()
VALID_KEYS = set(os.environ["GATEWAY_KEYS"].split(",")) # comma-separated
hits = defaultdict(int) # in-memory; use a database for several processes
async def validate_key(x_api_key: str = Header()):
if x_api_key not in VALID_KEYS:
raise HTTPException(401, "Invalid key")
return x_api_key
@app.post("/v1/chat/completions")
async def chat(req: dict, api_key: str = Depends(validate_key)):
minute = int(time.time() // 60)
hits[(api_key, minute)] += 1
if hits[(api_key, minute)] > 60:
raise HTTPException(429, "Rate limit")
async with httpx.AsyncClient(timeout=300) as client:
resp = await client.post("http://127.0.0.1:11434/v1/chat/completions", json=req)
return resp.json()Pattern 2 — Keep the model warm
Loading a model into RAM takes seconds to tens of seconds. Rather than pinging it from cron, set OLLAMA_KEEP_ALIVE (above), or pass "keep_alive": "30m" in API requests. Leave enough free RAM for the model to stay resident.
Pattern 3 — Fallback chain
try:
return await local_chat(prompt) # self-hosted
except (TimeoutError, httpx.HTTPError):
return await hosted_chat(prompt) # paid fallbackYou pay the API only during peaks or outages. Tell users if their data may go to a third-party API in that case.
Common pitfalls
/usr/share/ollama/.ollama/models for the service). Remove unused ones with ollama rm.num_ctx and check input length.Running this on Domain India
LLM serving needs root access, several GB of RAM and a long-running process, so it belongs on a VPS, not shared hosting or the App Platform.
- Domain India VPS plans are self-managed KVM servers with full root access and NVMe storage. The catalogue lists RAM from 2 GB (VPS Starter) up to 32 GB (VPS Enterprise); no GPU is listed, so plan for CPU inference with the model sizes in the table above.
- You install and update the software yourself. No backups or snapshots are included with a VPS, so keep copies of your configuration and data elsewhere.
- Shared hosting (cPanel, DirectAdmin, Webuzo) stops long-running processes and cannot run a model server. You can still call a model on your VPS, or a hosted API, from a PHP or Node.js site on shared hosting over HTTPS.
FAQ
Do I need a GPU to self-host an LLM?
No. Small quantised models (about 3B to 8B parameters) run on a CPU-only server with enough RAM, at a few tokens per second for the larger ones. That is fine for many chatbots and batch jobs. You need a GPU for many concurrent users, fast responses or large models.
Is self-hosting cheaper than a paid API?
It depends on volume. At low usage, a pay-per-token API is usually cheaper and better. With steady, heavy usage, a fixed-price server can cost less, especially if you already need one. Compare your real token usage against the provider's current prices and the server's monthly cost.
Can I fine-tune on my own data?
Yes. LoRA or QLoRA fine-tuning is practical on a GPU; on a CPU it is very slow. Try retrieval-augmented generation (RAG) over your documents first: it is cheaper and often enough.
Llama, Mistral or Qwen?
Llama for general English, Mistral 7B for a balanced, permissively licensed model, and Qwen3 or Gemma 3 for multilingual work including Indian languages. Test two or three on your own prompts before you choose.
Can Ollama serve more than one model?
Yes. Pull the models you need; Ollama loads them on demand and unloads idle ones after 5 minutes by default. Change that with the OLLAMA_KEEP_ALIVE setting, and make sure the server has RAM for every model you keep loaded.
Can I run Ollama on shared hosting?
No. Shared hosting stops long-running background processes and gives no root access. Run the model on a VPS and call it from your website over HTTPS.
Ready to run your own model? Compare VPS plans, or open a support ticket if you are unsure which size fits.
Self-managed KVM VPS with full root access and NVMe storage, from ₹553 a month excluding GST.
View VPS plans