DevOps & CI/CD Pipelines

Production Observability on Domain India VPS: Prometheus, Grafana, and Loki

By Domain India Team · DomainIndia EngineeringPublished 11 min read
Knowledge base article
Contents (15 sections)

When a site slows down or a server fills its disk at 3 AM, you want a graph and a log line that tell you why, and an alert that reached you first. This guide builds that stack on your own Linux VPS with Prometheus, Grafana, Loki, Grafana Alloy and Alertmanager.

Key takeaways

You can't fix what you can't see. Run Prometheus for metrics, Loki for logs and Grafana for dashboards on one monitoring VPS, ship data from every server with Node Exporter and Grafana Alloy, and let Alertmanager page you before customers complain. Keep every component on localhost or a firewalled private address, and put Grafana behind HTTPS.

The three pillars of observability

PillarToolAnswers
MetricsPrometheus"How fast? How many? How often?"
LogsLoki"What happened? What did it say?"
TracesJaeger / Tempo"Where did this slow request spend its time?"

Start with metrics and logs. Add traces when requests pass through several services and you need to see where the time goes.

What to monitor

The four golden signals (from Google's SRE book):

  1. Latency — how long requests take (p50, p95, p99)
  2. Traffic — requests per second
  3. Errors — failure rate
  4. Saturation — resource use (CPU, RAM, disk, queue depth)

Plus infrastructure:

  • CPU / RAM / disk / network per server
  • Database connections, slow query count
  • Cache hit rate
  • Queue size (Sidekiq, BullMQ)
  • External API latency

Option A — Self-hosted stack on a VPS

Install Prometheus, Grafana and Loki either on the same VPS as a small app or on a separate monitoring VPS. A separate one is better: if the app server dies, your monitoring still sees it.

Version numbers below are placeholders. Check each project's releases page and use the current stable version.

Step 1 — Install Prometheus

bash
PROM_VERSION=3.x.y   # current release from github.com/prometheus/prometheus/releases
wget https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
tar xzf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
sudo mv prometheus-${PROM_VERSION}.linux-amd64 /opt/prometheus
sudo useradd -r -s /sbin/nologin prometheus
sudo chown -R prometheus:prometheus /opt/prometheus

/opt/prometheus/prometheus.yml:

yaml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

scrape_configs:
  - job_name: 'prometheus'
    static_configs:
      - targets: ['localhost:9090']

  - job_name: 'node'
    static_configs:
      - targets: ['localhost:9100', 'vps2.internal:9100', 'vps3.internal:9100']

  - job_name: 'app'
    metrics_path: '/metrics'
    static_configs:
      - targets: ['app-vps:8080']

  - job_name: 'postgres'
    static_configs:
      - targets: ['db-vps:9187']

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['localhost:9093']

rule_files:
  - 'alerts.yml'

systemd service /etc/systemd/system/prometheus.service:

ini
[Unit]
Description=Prometheus
After=network-online.target

[Service]
User=prometheus
ExecStart=/opt/prometheus/prometheus \
    --config.file=/opt/prometheus/prometheus.yml \
    --storage.tsdb.path=/opt/prometheus/data \
    --web.listen-address=127.0.0.1:9090 \
    --storage.tsdb.retention.time=30d
Restart=on-failure

[Install]
WantedBy=multi-user.target
bash
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus

Step 2 — Node Exporter (on every server)

On every VPS you want to monitor:

bash
NE_VERSION=1.x.y     # current release from github.com/prometheus/node_exporter/releases
wget https://github.com/prometheus/node_exporter/releases/download/v${NE_VERSION}/node_exporter-${NE_VERSION}.linux-amd64.tar.gz
tar xzf node_exporter-${NE_VERSION}.linux-amd64.tar.gz
sudo mv node_exporter-${NE_VERSION}.linux-amd64/node_exporter /usr/local/bin/
sudo useradd -r -s /sbin/nologin node_exporter

/etc/systemd/system/node_exporter.service:

ini
[Unit]
Description=Node Exporter
After=network-online.target

[Service]
User=node_exporter
ExecStart=/usr/local/bin/node_exporter --web.listen-address=:9100
Restart=on-failure

[Install]
WantedBy=multi-user.target

It exposes CPU, RAM, disk, network, filesystem and more out of the box.

Firewall the exporters

Node Exporter listens on port 9100 on every interface. Allow that port only from your monitoring server's IP (firewalld, ufw or nftables), or bind it to a private address. The same applies to Loki (3100), Alertmanager (9093) and any app /metrics route: none of them should be open to the whole internet.

Step 3 — Install Grafana, Loki and Alloy from Grafana's repository

Grafana Labs publishes packages for Grafana, Loki and Alloy in one repository.

bash
# Debian / Ubuntu
sudo mkdir -p /etc/apt/keyrings
wget -qO - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install -y grafana loki      # on the monitoring server
sudo apt install -y alloy             # on every server that ships logs

# AlmaLinux / Rocky: add the repository described at https://rpm.grafana.com,
# then: sudo dnf install -y grafana loki alloy

sudo systemctl enable --now grafana-server

Grafana listens on port 3000. Log in with the default admin/admin and change the password at once. Rather than opening port 3000 to the internet, put Grafana behind nginx with a Let's Encrypt certificate, or reach it over an SSH tunnel.

Add Prometheus as a data source (http://localhost:9090), then import community dashboards from grafana.com, for example #1860 (Node Exporter Full) and #9628 (PostgreSQL).

Step 4 — Loki (log store)

A minimal single-server Loki 3 configuration at /etc/loki/config.yml:

yaml
auth_enabled: false

server:
  http_listen_address: 127.0.0.1   # use a private IP if other servers push logs
  http_listen_port: 3100

common:
  path_prefix: /var/lib/loki
  replication_factor: 1
  ring:
    kvstore:
      store: inmemory
  storage:
    filesystem:
      chunks_directory: /var/lib/loki/chunks
      rules_directory: /var/lib/loki/rules

schema_config:
  configs:
    - from: 2024-01-01
      store: tsdb
      object_store: filesystem
      schema: v13
      index:
        prefix: index_
        period: 24h

limits_config:
  retention_period: 336h           # 14 days

compactor:
  working_directory: /var/lib/loki/compactor
  retention_enabled: true
  delete_request_store: filesystem
bash
sudo systemctl enable --now loki

Step 5 — Ship logs with Grafana Alloy

Promtail, the old log shipper, is deprecated and has reached end of life; Grafana Alloy replaces it (and the old Grafana Agent). Configure it at /etc/alloy/config.alloy on each server:

code
local.file_match "logs" {
  path_targets = [
    {"__path__" = "/var/log/*.log",       "job" = "varlogs", "host" = "vps1"},
    {"__path__" = "/home/app/logs/*.log", "job" = "app",     "host" = "vps1"},
  ]
}

loki.source.file "files" {
  targets    = local.file_match.logs.targets
  forward_to = [loki.write.default.receiver]
}

loki.write "default" {
  endpoint {
    url = "http://10.0.0.5:3100/loki/api/v1/push"   // Loki's private address
  }
}
bash
sudo systemctl enable --now alloy

If logs must cross the public internet, send them to an HTTPS endpoint (nginx with TLS and authentication in front of Loki) instead of plain HTTP.

Add Loki as a data source in Grafana (http://localhost:3100) and query logs with LogQL:

logql
{job="app"} |= "error"

Option B — Grafana Cloud free tier

If your VPS resources are tight, Grafana Cloud has a free plan with limits on metric series, log volume and retention (check their pricing page for the current numbers). Install Grafana Alloy on your VPS and point it at your Grafana Cloud stack: no self-hosting burden, but your metrics and logs are stored by a third party.

Step 6 — Instrument your app

Node.js:

javascript
import express from 'express';
import client from 'prom-client';

const register = new client.Registry();
client.collectDefaultMetrics({ register });

const httpRequestsTotal = new client.Counter({
  name: 'http_requests_total',
  help: 'Total HTTP requests',
  labelNames: ['method', 'route', 'status'],
  registers: [register],
});

const httpDuration = new client.Histogram({
  name: 'http_request_duration_seconds',
  help: 'HTTP request duration',
  labelNames: ['method', 'route', 'status'],
  buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10],
  registers: [register],
});

const app = express();

app.use((req, res, next) => {
  const start = process.hrtime.bigint();
  res.on('finish', () => {
    const duration = Number(process.hrtime.bigint() - start) / 1e9;
    // Use the route pattern, never the raw path, to keep label values bounded
    const labels = { method: req.method, route: req.route?.path || 'unmatched', status: res.statusCode };
    httpRequestsTotal.inc(labels);
    httpDuration.observe(labels, duration);
  });
  next();
});

app.get('/metrics', async (req, res) => {
  res.set('Content-Type', register.contentType);
  res.send(await register.metrics());
});

Python (FastAPI):

python
from prometheus_fastapi_instrumentator import Instrumentator
Instrumentator().instrument(app).expose(app)

PHP (on a VPS; this example stores counters in Redis, which you run yourself):

bash
composer require promphp/prometheus_client_php
php
use Prometheus\CollectorRegistry;
use Prometheus\RenderTextFormat;
use Prometheus\Storage\Redis;

$registry = new CollectorRegistry(new Redis(['host' => '127.0.0.1']));

$counter = $registry->getOrRegisterCounter('app', 'requests_total', 'Total requests', ['route']);
$counter->inc(['/api/users']);

// In your /metrics endpoint:
$renderer = new RenderTextFormat();
header('Content-Type: ' . RenderTextFormat::MIME_TYPE);
echo $renderer->render($registry->getMetricFamilySamples());

Step 7 — Alerting with Alertmanager

Install Alertmanager from its GitHub releases the same way as Prometheus (listen on 127.0.0.1:9093). Then define rules in /opt/prometheus/alerts.yml:

yaml
groups:
  - name: infrastructure
    rules:
      - alert: HighCPU
        expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "CPU > 80% on {{ $labels.instance }}"

      - alert: DiskFull
        expr: node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} < 0.10
        for: 5m
        labels: { severity: critical }
        annotations:
          summary: "Disk < 10% free on {{ $labels.instance }}"

      - alert: HighErrorRate
        expr: |
          sum by(route) (rate(http_requests_total{status=~"5.."}[5m]))
            / sum by(route) (rate(http_requests_total[5m])) > 0.05
        for: 5m
        labels: { severity: critical }
        annotations:
          summary: "5xx rate > 5% on {{ $labels.route }}"

      - alert: TargetDown
        expr: up == 0
        for: 2m
        labels: { severity: critical }
        annotations:
          summary: "{{ $labels.job }} on {{ $labels.instance }} is DOWN"

Alertmanager config /etc/alertmanager/alertmanager.yml:

yaml
route:
  receiver: default
  group_by: ['alertname', 'instance']
  group_wait: 30s

receivers:
  - name: default
    email_configs:
      - to: '[email protected]'
        from: '[email protected]'
        smarthost: 'smtp.gmail.com:587'
        auth_username: '[email protected]'
        auth_password: 'app-password-here'   # a Google app password, not your login
    slack_configs:
      - api_url: 'https://hooks.slack.com/services/YOUR/WEBHOOK'
        channel: '#alerts'

Step 8 — External uptime checks

Internal metrics can't tell you the site is down if the monitoring server goes down with it. Run an uptime checker from a different server or a hosted service. A popular self-hosted option is Uptime Kuma:

bash
# use the image tag recommended in the Uptime Kuma README
docker run -d --restart=always -p 127.0.0.1:3001:3001 \
  -v uptime-kuma:/app/data --name uptime-kuma louislam/uptime-kuma:1

Put it behind nginx with HTTPS, then configure HTTP checks, keyword checks, SSL certificate expiry and DNS checks, with alerts by email, Slack or Telegram.

Common pitfalls

Metrics cardinality explosion
A user_id label means millions of series. Keep label values bounded: status code and route pattern, never per-user or raw URL.
Alert fatigue
Too many low-severity alerts and the team ignores them all. Start with 5–10 critical alerts and grow slowly.
No runbook
"CPU high" fires at 3 AM and nobody knows what to do. Link every alert to a short runbook.
Retention too long for the disk
Long retention of 15-second metrics needs a lot of disk. Keep weeks locally; use Thanos or Grafana Mimir for long-term storage.
Loki disk fills
Unbounded log ingestion. Set retention_period and enable the compactor's retention, as in the config above.
Grafana open to the internet
Always behind HTTPS and a strong login. Add nginx basic auth or an identity-aware proxy for extra protection.

Running this on Domain India

This stack needs root access and long-running daemons, so it runs on a VPS, not on shared hosting.

  • A Domain India VPS is a self-managed KVM server with full root access; you install, secure and update the monitoring stack yourself. Monitoring, backups and snapshots are not included with a VPS, so back up your Grafana dashboards and configuration elsewhere.
  • Shared hosting (cPanel, DirectAdmin, Webuzo) can't run exporters or log shippers. For a site on shared hosting, use an external uptime checker, and use the control panel's own resource usage and log tools.
  • If you monitor several servers, a small separate monitoring VPS keeps the monitoring alive when an app server fails.

FAQ

How much RAM does this stack need?

For a small setup (one to three servers monitored, about 30 days of metrics and two weeks of logs), a VPS with around 2 GB of RAM is usually enough. With ten or more servers or long retention, give the monitoring server 4 GB or more and watch its own disk and memory graphs.

Should I use Datadog or New Relic, or self-host?

Self-host if you have the time to run it and want a predictable cost. A managed service (Datadog, New Relic, Grafana Cloud) suits you if you want no operations work and accept paying per host, metric or gigabyte.

What about application performance monitoring (APM) and tracing?

OpenTelemetry is the standard way to instrument an app. Send traces to Grafana Tempo, logs to Loki and metrics to Prometheus, and Grafana links all three so you can jump from a slow request to its logs.

Is Promtail still the right log shipper?

No. Promtail is deprecated and past its end of life, and the Grafana Agent has also been replaced. Use Grafana Alloy, which ships logs to Loki and can also collect metrics and traces.

How do I monitor Cloudflare or the edge?

The Cloudflare dashboard has its own analytics. For a combined view, Cloudflare's Logpush (available on some plans) can send request logs to a destination you then load into Loki.

Do I need all this for a small website?

No. For a small site, an external uptime checker plus Node Exporter and one basic Grafana dashboard on your VPS is plenty. Add Loki, alerting and tracing as the site and team grow.

Ready to build your monitoring server? Compare VPS plans, or open a support ticket if you need help choosing a size.

Monitor your servers from one VPS

Self-managed KVM VPS with full root access and NVMe storage, from ₹553 a month excluding GST.

See VPS plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app