When a site slows down or a server fills its disk at 3 AM, you want a graph and a log line that tell you why, and an alert that reached you first. This guide builds that stack on your own Linux VPS with Prometheus, Grafana, Loki, Grafana Alloy and Alertmanager.
You can't fix what you can't see. Run Prometheus for metrics, Loki for logs and Grafana for dashboards on one monitoring VPS, ship data from every server with Node Exporter and Grafana Alloy, and let Alertmanager page you before customers complain. Keep every component on localhost or a firewalled private address, and put Grafana behind HTTPS.
The three pillars of observability
| Pillar | Tool | Answers |
|---|---|---|
| Metrics | Prometheus | "How fast? How many? How often?" |
| Logs | Loki | "What happened? What did it say?" |
| Traces | Jaeger / Tempo | "Where did this slow request spend its time?" |
Start with metrics and logs. Add traces when requests pass through several services and you need to see where the time goes.
What to monitor
The four golden signals (from Google's SRE book):
- Latency — how long requests take (p50, p95, p99)
- Traffic — requests per second
- Errors — failure rate
- Saturation — resource use (CPU, RAM, disk, queue depth)
Plus infrastructure:
- CPU / RAM / disk / network per server
- Database connections, slow query count
- Cache hit rate
- Queue size (Sidekiq, BullMQ)
- External API latency
Option A — Self-hosted stack on a VPS
Install Prometheus, Grafana and Loki either on the same VPS as a small app or on a separate monitoring VPS. A separate one is better: if the app server dies, your monitoring still sees it.
Version numbers below are placeholders. Check each project's releases page and use the current stable version.
Step 1 — Install Prometheus
PROM_VERSION=3.x.y # current release from github.com/prometheus/prometheus/releases
wget https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
tar xzf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
sudo mv prometheus-${PROM_VERSION}.linux-amd64 /opt/prometheus
sudo useradd -r -s /sbin/nologin prometheus
sudo chown -R prometheus:prometheus /opt/prometheus/opt/prometheus/prometheus.yml:
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'node'
static_configs:
- targets: ['localhost:9100', 'vps2.internal:9100', 'vps3.internal:9100']
- job_name: 'app'
metrics_path: '/metrics'
static_configs:
- targets: ['app-vps:8080']
- job_name: 'postgres'
static_configs:
- targets: ['db-vps:9187']
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
rule_files:
- 'alerts.yml'systemd service /etc/systemd/system/prometheus.service:
[Unit]
Description=Prometheus
After=network-online.target
[Service]
User=prometheus
ExecStart=/opt/prometheus/prometheus \
--config.file=/opt/prometheus/prometheus.yml \
--storage.tsdb.path=/opt/prometheus/data \
--web.listen-address=127.0.0.1:9090 \
--storage.tsdb.retention.time=30d
Restart=on-failure
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable --now prometheusStep 2 — Node Exporter (on every server)
On every VPS you want to monitor:
NE_VERSION=1.x.y # current release from github.com/prometheus/node_exporter/releases
wget https://github.com/prometheus/node_exporter/releases/download/v${NE_VERSION}/node_exporter-${NE_VERSION}.linux-amd64.tar.gz
tar xzf node_exporter-${NE_VERSION}.linux-amd64.tar.gz
sudo mv node_exporter-${NE_VERSION}.linux-amd64/node_exporter /usr/local/bin/
sudo useradd -r -s /sbin/nologin node_exporter/etc/systemd/system/node_exporter.service:
[Unit]
Description=Node Exporter
After=network-online.target
[Service]
User=node_exporter
ExecStart=/usr/local/bin/node_exporter --web.listen-address=:9100
Restart=on-failure
[Install]
WantedBy=multi-user.targetIt exposes CPU, RAM, disk, network, filesystem and more out of the box.
Node Exporter listens on port 9100 on every interface. Allow that port only from your monitoring server's IP (firewalld, ufw or nftables), or bind it to a private address. The same applies to Loki (3100), Alertmanager (9093) and any app /metrics route: none of them should be open to the whole internet.
Step 3 — Install Grafana, Loki and Alloy from Grafana's repository
Grafana Labs publishes packages for Grafana, Loki and Alloy in one repository.
# Debian / Ubuntu
sudo mkdir -p /etc/apt/keyrings
wget -qO - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install -y grafana loki # on the monitoring server
sudo apt install -y alloy # on every server that ships logs
# AlmaLinux / Rocky: add the repository described at https://rpm.grafana.com,
# then: sudo dnf install -y grafana loki alloy
sudo systemctl enable --now grafana-serverGrafana listens on port 3000. Log in with the default admin/admin and change the password at once. Rather than opening port 3000 to the internet, put Grafana behind nginx with a Let's Encrypt certificate, or reach it over an SSH tunnel.
Add Prometheus as a data source (http://localhost:9090), then import community dashboards from grafana.com, for example #1860 (Node Exporter Full) and #9628 (PostgreSQL).
Step 4 — Loki (log store)
A minimal single-server Loki 3 configuration at /etc/loki/config.yml:
auth_enabled: false
server:
http_listen_address: 127.0.0.1 # use a private IP if other servers push logs
http_listen_port: 3100
common:
path_prefix: /var/lib/loki
replication_factor: 1
ring:
kvstore:
store: inmemory
storage:
filesystem:
chunks_directory: /var/lib/loki/chunks
rules_directory: /var/lib/loki/rules
schema_config:
configs:
- from: 2024-01-01
store: tsdb
object_store: filesystem
schema: v13
index:
prefix: index_
period: 24h
limits_config:
retention_period: 336h # 14 days
compactor:
working_directory: /var/lib/loki/compactor
retention_enabled: true
delete_request_store: filesystemsudo systemctl enable --now lokiStep 5 — Ship logs with Grafana Alloy
Promtail, the old log shipper, is deprecated and has reached end of life; Grafana Alloy replaces it (and the old Grafana Agent). Configure it at /etc/alloy/config.alloy on each server:
local.file_match "logs" {
path_targets = [
{"__path__" = "/var/log/*.log", "job" = "varlogs", "host" = "vps1"},
{"__path__" = "/home/app/logs/*.log", "job" = "app", "host" = "vps1"},
]
}
loki.source.file "files" {
targets = local.file_match.logs.targets
forward_to = [loki.write.default.receiver]
}
loki.write "default" {
endpoint {
url = "http://10.0.0.5:3100/loki/api/v1/push" // Loki's private address
}
}sudo systemctl enable --now alloyIf logs must cross the public internet, send them to an HTTPS endpoint (nginx with TLS and authentication in front of Loki) instead of plain HTTP.
Add Loki as a data source in Grafana (http://localhost:3100) and query logs with LogQL:
{job="app"} |= "error"Option B — Grafana Cloud free tier
If your VPS resources are tight, Grafana Cloud has a free plan with limits on metric series, log volume and retention (check their pricing page for the current numbers). Install Grafana Alloy on your VPS and point it at your Grafana Cloud stack: no self-hosting burden, but your metrics and logs are stored by a third party.
Step 6 — Instrument your app
Node.js:
import express from 'express';
import client from 'prom-client';
const register = new client.Registry();
client.collectDefaultMetrics({ register });
const httpRequestsTotal = new client.Counter({
name: 'http_requests_total',
help: 'Total HTTP requests',
labelNames: ['method', 'route', 'status'],
registers: [register],
});
const httpDuration = new client.Histogram({
name: 'http_request_duration_seconds',
help: 'HTTP request duration',
labelNames: ['method', 'route', 'status'],
buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10],
registers: [register],
});
const app = express();
app.use((req, res, next) => {
const start = process.hrtime.bigint();
res.on('finish', () => {
const duration = Number(process.hrtime.bigint() - start) / 1e9;
// Use the route pattern, never the raw path, to keep label values bounded
const labels = { method: req.method, route: req.route?.path || 'unmatched', status: res.statusCode };
httpRequestsTotal.inc(labels);
httpDuration.observe(labels, duration);
});
next();
});
app.get('/metrics', async (req, res) => {
res.set('Content-Type', register.contentType);
res.send(await register.metrics());
});Python (FastAPI):
from prometheus_fastapi_instrumentator import Instrumentator
Instrumentator().instrument(app).expose(app)PHP (on a VPS; this example stores counters in Redis, which you run yourself):
composer require promphp/prometheus_client_phpuse Prometheus\CollectorRegistry;
use Prometheus\RenderTextFormat;
use Prometheus\Storage\Redis;
$registry = new CollectorRegistry(new Redis(['host' => '127.0.0.1']));
$counter = $registry->getOrRegisterCounter('app', 'requests_total', 'Total requests', ['route']);
$counter->inc(['/api/users']);
// In your /metrics endpoint:
$renderer = new RenderTextFormat();
header('Content-Type: ' . RenderTextFormat::MIME_TYPE);
echo $renderer->render($registry->getMetricFamilySamples());Step 7 — Alerting with Alertmanager
Install Alertmanager from its GitHub releases the same way as Prometheus (listen on 127.0.0.1:9093). Then define rules in /opt/prometheus/alerts.yml:
groups:
- name: infrastructure
rules:
- alert: HighCPU
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 10m
labels: { severity: warning }
annotations:
summary: "CPU > 80% on {{ $labels.instance }}"
- alert: DiskFull
expr: node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} < 0.10
for: 5m
labels: { severity: critical }
annotations:
summary: "Disk < 10% free on {{ $labels.instance }}"
- alert: HighErrorRate
expr: |
sum by(route) (rate(http_requests_total{status=~"5.."}[5m]))
/ sum by(route) (rate(http_requests_total[5m])) > 0.05
for: 5m
labels: { severity: critical }
annotations:
summary: "5xx rate > 5% on {{ $labels.route }}"
- alert: TargetDown
expr: up == 0
for: 2m
labels: { severity: critical }
annotations:
summary: "{{ $labels.job }} on {{ $labels.instance }} is DOWN"Alertmanager config /etc/alertmanager/alertmanager.yml:
route:
receiver: default
group_by: ['alertname', 'instance']
group_wait: 30s
receivers:
- name: default
email_configs:
- to: '[email protected]'
from: '[email protected]'
smarthost: 'smtp.gmail.com:587'
auth_username: '[email protected]'
auth_password: 'app-password-here' # a Google app password, not your login
slack_configs:
- api_url: 'https://hooks.slack.com/services/YOUR/WEBHOOK'
channel: '#alerts'Step 8 — External uptime checks
Internal metrics can't tell you the site is down if the monitoring server goes down with it. Run an uptime checker from a different server or a hosted service. A popular self-hosted option is Uptime Kuma:
# use the image tag recommended in the Uptime Kuma README
docker run -d --restart=always -p 127.0.0.1:3001:3001 \
-v uptime-kuma:/app/data --name uptime-kuma louislam/uptime-kuma:1Put it behind nginx with HTTPS, then configure HTTP checks, keyword checks, SSL certificate expiry and DNS checks, with alerts by email, Slack or Telegram.
Common pitfalls
user_id label means millions of series. Keep label values bounded: status code and route pattern, never per-user or raw URL.retention_period and enable the compactor's retention, as in the config above.Running this on Domain India
This stack needs root access and long-running daemons, so it runs on a VPS, not on shared hosting.
- A Domain India VPS is a self-managed KVM server with full root access; you install, secure and update the monitoring stack yourself. Monitoring, backups and snapshots are not included with a VPS, so back up your Grafana dashboards and configuration elsewhere.
- Shared hosting (cPanel, DirectAdmin, Webuzo) can't run exporters or log shippers. For a site on shared hosting, use an external uptime checker, and use the control panel's own resource usage and log tools.
- If you monitor several servers, a small separate monitoring VPS keeps the monitoring alive when an app server fails.
FAQ
How much RAM does this stack need?
For a small setup (one to three servers monitored, about 30 days of metrics and two weeks of logs), a VPS with around 2 GB of RAM is usually enough. With ten or more servers or long retention, give the monitoring server 4 GB or more and watch its own disk and memory graphs.
Should I use Datadog or New Relic, or self-host?
Self-host if you have the time to run it and want a predictable cost. A managed service (Datadog, New Relic, Grafana Cloud) suits you if you want no operations work and accept paying per host, metric or gigabyte.
What about application performance monitoring (APM) and tracing?
OpenTelemetry is the standard way to instrument an app. Send traces to Grafana Tempo, logs to Loki and metrics to Prometheus, and Grafana links all three so you can jump from a slow request to its logs.
Is Promtail still the right log shipper?
No. Promtail is deprecated and past its end of life, and the Grafana Agent has also been replaced. Use Grafana Alloy, which ships logs to Loki and can also collect metrics and traces.
How do I monitor Cloudflare or the edge?
The Cloudflare dashboard has its own analytics. For a combined view, Cloudflare's Logpush (available on some plans) can send request logs to a destination you then load into Loki.
Do I need all this for a small website?
No. For a small site, an external uptime checker plus Node Exporter and one basic Grafana dashboard on your VPS is plenty. Add Loki, alerting and tracing as the site and team grow.
Ready to build your monitoring server? Compare VPS plans, or open a support ticket if you need help choosing a size.
Self-managed KVM VPS with full root access and NVMe storage, from ₹553 a month excluding GST.
See VPS plans