Server Management (Unmanaged)

A Comprehensive Guide to Setting Up a Multi-Server VPS Cluster

By the Domain India teamPublished 11 min read
Knowledge base article
Contents (13 sections)

When one server is no longer enough, or when a single failure would take your business offline, the next step is a small cluster: several VPSes that share the work. This guide shows how to plan and build a practical multi-server setup with a load balancer, two or more application servers and a replicated database. It applies to servers you run yourself, such as your own VPSes, not to shared hosting, where the server setup is managed for you and cannot be changed.

Key takeaways

Start with three roles: a load balancer (HAProxy or nginx), two or more identical application servers, and a database with a primary and a replica. Connect the servers over an encrypted private network such as WireGuard, keep the app servers stateless (sessions in Redis or the database, uploads in shared storage), and monitor everything. A cluster only adds reliability if you remove single points of failure one at a time and test failover before you need it.

1. Do you actually need a cluster?

A cluster adds moving parts, cost and work. Before building one, check whether a bigger single server, a cache or query tuning would solve the problem.

A cluster helps when
  • Traffic regularly exceeds what the largest sensible single server can handle
  • Downtime during updates or a server failure costs you real money
  • You need to deploy without taking the site offline
It is overkill when
  • A single VPS is under 50% busy most of the time
  • Slowness comes from unindexed queries or no caching
  • Nobody on the team has time to operate several servers

Upgrading one server (vertical scaling) is simpler and often the right first step. Adding servers (horizontal scaling) is what you do when that runs out, or when you need redundancy.

2. A practical reference layout

Load balancer
One VPS running HAProxy or nginx, holding the public IP and the TLS certificate
App servers
Two or more identical VPSes running your application, reachable only from the load balancer
Database
A primary that takes writes and a replica that stays in sync, for reads and failover

Add Redis for sessions and caching, and object storage or a shared file service for uploads. Use the same Linux distribution and version on every node. Ubuntu LTS, Debian, AlmaLinux and Rocky Linux are all good choices; CentOS Linux has reached end of life and should not be used for new servers.

3. Prepare every node

On each server:

  • Create a sudo user, log in with SSH keys only and disable password and root login.
  • Apply updates and turn on automatic security updates.
  • Set the hostname and time sync (chrony or systemd-timesyncd), because replication and logs depend on accurate clocks.
  • Enable a firewall that allows only what each role needs.
bash
# Ubuntu/Debian example for an app server
sudo apt update && sudo apt full-upgrade -y
sudo ufw default deny incoming
sudo ufw allow OpenSSH
sudo ufw allow in on wg0 to any port 8080 proto tcp   # app port, private network only
sudo ufw enable

On AlmaLinux or Rocky Linux use dnf and firewalld instead. Tools such as Ansible make this repeatable across nodes; see infrastructure as code with Terraform and Ansible.

4. Connect the servers privately

Traffic between your nodes (app to database, load balancer to app) carries passwords and customer data, so it must not travel in the clear over the internet. WireGuard creates an encrypted private network between servers using only their public IPs.

ini
# /etc/wireguard/wg0.conf on the app server (10.10.0.2)
[Interface]
Address = 10.10.0.2/24
PrivateKey = <this server's private key>
ListenPort = 51820

[Peer]   # database server
PublicKey = <db server's public key>
Endpoint = <db public IP>:51820
AllowedIPs = 10.10.0.3/32

Generate keys with wg genkey | tee private.key | wg pubkey > public.key, add a [Peer] block for each node, open UDP 51820 between the servers, and start it with sudo systemctl enable --now wg-quick@wg0. From then on, bind databases and internal services to the 10.10.0.x addresses only.

5. Load balancing

HAProxy is purpose-built for this and checks the health of each backend actively, taking a failed server out of rotation within seconds.

cfg
# /etc/haproxy/haproxy.cfg (excerpt)
frontend web
    bind :80
    bind :443 ssl crt /etc/haproxy/certs/example.com.pem
    http-request redirect scheme https unless { ssl_fc }
    default_backend app

backend app
    balance roundrobin
    option httpchk GET /healthz
    server app1 10.10.0.2:8080 check
    server app2 10.10.0.4:8080 check

Give your app a /healthz endpoint that returns 200 only when it can reach the database. nginx works too: define an upstream block with the app servers and proxy_pass to it. Open-source nginx removes a failing server only after real requests to it fail (max_fails), rather than probing it in advance. Round-robin DNS (several A records for one name) spreads load but has no health checks, so a dead server keeps receiving visitors. For a deeper HAProxy walkthrough, see mastering load balancing with HAProxy.

6. Keep the app servers stateless

Any request can land on any app server, so none of them may keep something the others need.

  • Sessions: store them in Redis or the database, not in local files or memory. Sticky sessions work as a stopgap but break when a server goes down.
  • Uploads: write them to S3-compatible object storage, or to one file server the app nodes share over the private network. For a few servers with mostly static files, deploying the same release to every node plus lsyncd or scheduled rsync for uploads is enough.
  • Deploys: ship the same build to every node, then take nodes out of the load balancer one at a time to update them. That is your zero-downtime deploy.

Distributed file systems such as GlusterFS or Ceph are powerful but add real operational complexity; use them only when simpler options cannot keep up.

7. Database replication

MySQL 8.4 source and replica

Enable GTIDs on both servers so a replica always knows where it is:

ini
# /etc/mysql/mysql.conf.d/replication.cnf  (server_id = 2 on the replica)
[mysqld]
server_id = 1
gtid_mode = ON
enforce_gtid_consistency = ON
bind-address = 10.10.0.3

On the source, create a replication user that may connect only from the private network:

sql
CREATE USER 'repl'@'10.10.0.%' IDENTIFIED BY 'a-long-random-password' REQUIRE SSL;
GRANT REPLICATION SLAVE ON *.* TO 'repl'@'10.10.0.%';

Load a consistent copy of the data onto the replica (for example with mysqldump --single-transaction --source-data, or MySQL's clone plugin), then point it at the source:

sql
CHANGE REPLICATION SOURCE TO SOURCE_HOST='10.10.0.3', SOURCE_USER='repl',
  SOURCE_PASSWORD='a-long-random-password', SOURCE_AUTO_POSITION=1, SOURCE_SSL=1;
START REPLICA;
SHOW REPLICA STATUS\G

Older tutorials use CHANGE MASTER TO and START SLAVE with a binary log file and position; those commands are removed in current MySQL. MariaDB still uses the older CHANGE MASTER TO syntax with its own GTID option, so follow the MariaDB documentation there.

PostgreSQL streaming replication

Create a role with REPLICATION, allow it from the replica's private IP in pg_hba.conf, then, with the replica stopped and its data directory empty, initialise it with pg_basebackup -h 10.10.0.3 -U replicator -D "$PGDATA" -R -X stream and start it. The -R flag writes the connection settings so the replica starts streaming straight away.

Multi-primary

Galera Cluster (MariaDB) lets every node accept writes, but needs at least three nodes, is sensitive to network latency and handles some write patterns badly. Most applications are better served by one primary and one or more replicas.

A replica is not a backup

Replication copies mistakes instantly: a dropped table or a bad migration disappears from the replica too. Keep scheduled backups, such as nightly database dumps plus binary logs or WAL archives, stored off the cluster, and test a restore regularly.

8. Removing the last single point of failure

With two app servers and a replica, the load balancer and the primary database are still single points of failure.

  • Load balancer: a second load balancer can take over a shared virtual IP with Keepalived (VRRP). This needs an IP address that can move between servers and a network that allows it, which not every VPS provider supports; ask your provider first. The alternative is DNS failover through a DNS provider with health checks.
  • Database: promoting a replica can be manual (a runbook you have practised) or automated with a tool such as Patroni for PostgreSQL or Orchestrator for MySQL. Automated failover done badly causes split-brain, so start manual.
  • Test it. Stop a node deliberately in a quiet hour and watch what happens. A failover you have never tested will not work when you need it.

9. Monitoring, logs and ongoing care

Watch CPU, memory, disk, replication lag and HTTP error rates on every node, with alerts that reach a person. Prometheus with Grafana is a common open-source stack, and Loki or a similar tool centralises logs so you are not logging into five servers during an incident; see production observability with Prometheus, Grafana and Loki. Patch nodes one at a time using the rolling method from section 6, and keep a written record of every configuration change. If you later need automatic scaling and self-healing containers, lightweight Kubernetes with k3s is the next step.

10. Building your cluster on Domain India VPS

Every Domain India VPS is a KVM server with full root access, so you can run HAProxy, WireGuard, MySQL or PostgreSQL and anything else in this guide. VPS hosting is self-managed: you install, secure, update and back up the servers yourself. Ask support whether a movable IP or private networking between your VPSes is available before you design around them; the WireGuard approach in section 4 works over ordinary public IPs either way. Harden each node with the VPS security checklist and the VPS firewall guide.

A small cluster might use a smaller plan for the load balancer and larger ones for the app and database servers:

VPS Starter
₹552.65/mo + GST
  • 1 vCPU
  • 2 GB DDR4 RAM
  • 64 GB NVMe SSD Storage
  • 2 TB Monthly Bandwidth
See plan details
VPS Standard
₹2,210.60/mo + GST
  • 4 vCPU
  • 8 GB DDR4 RAM
  • 256 GB NVMe SSD Storage
  • 5 TB Monthly Bandwidth
See plan details

Prices on the cards are Domain India list prices and exclude 18% GST.

How many servers do I need for a basic high-availability setup?

A common starting point is four: one load balancer, two application servers and one database server, plus a database replica for failover. The load balancer remains a single point of failure until you add a second one with a movable IP or DNS failover.

Should I use HAProxy or nginx as the load balancer?

Both work. HAProxy actively health-checks backends and removes a failed server within seconds, which makes it a natural choice for a dedicated load balancer. nginx is convenient if you already use it, but its open-source version detects failures only when real requests fail.

How do I keep user sessions working across several app servers?

Store sessions in a shared store such as Redis or the database instead of on each server's disk or memory. Then any server can handle any request, and losing one server does not log users out.

Is database replication the same as a backup?

No. Replication copies every change, including accidental deletions, to the replica within seconds. Keep separate scheduled backups stored away from the cluster and test restoring them.

How should servers in a cluster talk to each other securely?

Over an encrypted private network. WireGuard connects servers using their public IPs, and you then bind databases and internal services to the private addresses only, with a firewall blocking them from the internet.

Can I build a server cluster on shared hosting?

No. Shared hosting is one managed environment and you cannot change its server configuration. A cluster needs servers you control, such as several VPSes with root access.

Ready to plan your cluster? Compare VPS plans to size each role, or open a support ticket to ask about IP addresses and networking before you build.

Build your cluster on KVM VPS

Full root access on every node, so you choose the load balancer, database and failover design.

See VPS plans

Ready when you are

Get VPS from ₹552.65/mo + GST

See plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app