Rust Development

Rust for Web Scraping and Data Pipelines on Domain India VPS

By Domain India Team · DomainIndia EngineeringPublished 9 min read
Knowledge base article
Contents (12 sections)

When a scraping or ETL job outgrows a quick Python script, Rust is a strong next step: fast, low on memory and predictable. This guide builds a small but production-shaped pipeline and schedules it on a Linux VPS.

Key takeaways

Rust's speed + memory safety makes it ideal for data pipelines that scrape, transform, or ingest millions of records. This guide shows async scraping with reqwest + scraper, parsing JSON/CSV at scale, writing to PostgreSQL with sqlx, and scheduling it with systemd on your own VPS.

Why Rust for data pipelines

Typical scraping/ETL workloads in Python:

  • scraping thousands of URLs one at a time with requests and BeautifulSoup is slow
  • parsing multi-gigabyte JSON files is CPU-bound in pure Python
  • loading into a database is bottlenecked by per-row inserts

The Rust versions of the same jobs are usually much faster:

  • async reqwest fetches many URLs concurrently, so network waits overlap
  • serde_json parses quickly, and rayon spreads CPU work over all cores
  • PostgreSQL's COPY FROM loads rows in bulk instead of one INSERT at a time

(Python can close much of the gap with asyncio and bulk loading too; the exact speed-up depends on your data and the target sites.) Rust pays off when data work is CPU- or IO-heavy and runs often. For light scraping, say a hundred URLs a day, Python is fine and quicker to write.

Project scaffold

bash
cargo new --bin scraper && cd scraper

Cargo.toml:

toml
[dependencies]
tokio = { version = "1", features = ["full"] }
reqwest = { version = "0.12", features = ["json", "gzip", "rustls-tls"] }
scraper = "0.23"           # HTML parsing
serde = { version = "1", features = ["derive"] }
serde_json = "1"
anyhow = "1"               # error handling
tracing = "0.1"
tracing-subscriber = "0.3"
futures = "0.3"
sqlx = { version = "0.8", features = ["runtime-tokio", "tls-rustls", "postgres", "chrono"] }
csv = "1.3"
rayon = "1"                # parallel CPU work
governor = "0.8"           # rate limiting

Version numbers move quickly; run cargo add crate-name or check crates.io for the current release of each.

Pattern 1 — Async bulk scraping

Scrape 10,000 URLs concurrently, extract title + price, save to CSV.

rust
use futures::stream::{self, StreamExt};
use reqwest::Client;
use scraper::{Html, Selector};
use serde::Serialize;

#[derive(Serialize, Debug)]
struct Product {
    url: String,
    title: String,
    price: Option<String>,
}

async fn scrape_one(client: &Client, url: &str) -> anyhow::Result<Product> {
    let html = client.get(url).send().await?.error_for_status()?.text().await?;
    let doc = Html::parse_document(&html);
    let title_sel = Selector::parse("h1.product-title").unwrap();
    let price_sel = Selector::parse("span.price").unwrap();

    let title = doc.select(&title_sel)
        .next()
        .map(|n| n.text().collect::<String>().trim().to_string())
        .unwrap_or_default();
    let price = doc.select(&price_sel)
        .next()
        .map(|n| n.text().collect::<String>().trim().to_string());

    Ok(Product { url: url.to_string(), title, price })
}

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt::init();

    let urls: Vec<String> = std::fs::read_to_string("urls.txt")?
        .lines().map(String::from).collect();

    let client = Client::builder()
        .user_agent("MyScraper/1.0 (+https://yourcompany.com/bot)")
        .timeout(std::time::Duration::from_secs(10))
        .gzip(true)
        .build()?;

    let results: Vec<Product> = stream::iter(urls.iter())
        .map(|url| {
            let client = &client;
            async move {
                scrape_one(client, url).await
                    .unwrap_or_else(|e| {
                        tracing::warn!("failed {url}: {e}");
                        Product { url: url.to_string(), title: String::new(), price: None }
                    })
            }
        })
        .buffer_unordered(50)  // 50 concurrent requests
        .collect()
        .await;

    let mut wtr = csv::Writer::from_path("products.csv")?;
    for p in &results { wtr.serialize(p)?; }
    wtr.flush()?;

    tracing::info!("scraped {} products", results.len());
    Ok(())
}

buffer_unordered(50) caps concurrency — essential to avoid being rate-limited or overwhelming target servers.

Pattern 2 — Streaming JSON parsing

For a 5 GB JSON file you can't load into memory, stream line-by-line (if NDJSON) or use serde_json::Deserializer:

rust
use std::fs::File;
use std::io::{BufRead, BufReader};
use serde::Deserialize;

#[derive(Deserialize, Debug)]
struct Event {
    user_id: u64,
    action: String,
    timestamp: i64,
}

fn process_ndjson(path: &str) -> anyhow::Result<u64> {
    let file = File::open(path)?;
    let reader = BufReader::new(file);
    let mut count = 0;

    for line in reader.lines() {
        let line = line?;
        if line.is_empty() { continue; }
        let event: Event = serde_json::from_str(&line)?;
        // process event...
        count += 1;
        if count % 100_000 == 0 {
            tracing::info!("processed {}", count);
        }
    }
    Ok(count)
}

Memory usage: flat, regardless of file size.

Pattern 3 — Parallel CPU work with rayon

For transforming millions of rows, Tokio isn't the right tool (async is for IO). Use rayon for CPU:

rust
use rayon::prelude::*;

let rows: Vec<RawRow> = load_all();
let processed: Vec<ProcessedRow> = rows
    .par_iter()
    .map(|row| expensive_transform(row))
    .collect();

.par_iter() uses all CPU cores automatically.

Pattern 4 — Bulk DB inserts with sqlx + COPY

Per-row inserts are slow. For 1M rows, use PostgreSQL's COPY FROM:

rust
use sqlx::postgres::PgPool;
use std::io::Write;   // for writeln! into a Vec<u8>

async fn bulk_insert(pool: &PgPool, rows: &[Product]) -> anyhow::Result<()> {
    let mut conn = pool.acquire().await?;
    let mut copy = conn.copy_in_raw(
        "COPY products (url, title, price) FROM STDIN (FORMAT csv)"
    ).await?;

    let mut buf = Vec::with_capacity(1024 * 1024);
    for p in rows {
        writeln!(
            &mut buf,
            "{},{},{}",
            csv_escape(&p.url),
            csv_escape(&p.title),
            csv_escape(p.price.as_deref().unwrap_or(""))
        )?;
    }

    copy.send(buf.as_slice()).await?;
    copy.finish().await?;
    Ok(())
}

fn csv_escape(s: &str) -> String {
    if s.contains([',', '"', '\n', '\r']) {
        format!("\"{}\"", s.replace('"', "\"\""))
    } else {
        s.to_string()
    }
}

COPY is typically orders of magnitude faster than individual INSERTs for large batches. For very large loads, send the buffer in chunks rather than building one huge Vec.

Pattern 5 — Rate limiting yourself

Be a good citizen — don't hammer target servers.

rust
use governor::{Quota, RateLimiter};
use std::num::NonZeroU32;

let quota = Quota::per_second(NonZeroU32::new(10).unwrap()); // 10 req/sec
let limiter = RateLimiter::direct(quota);

for url in &urls {
    limiter.until_ready().await;
    scrape_one(&client, url).await?;
}

Scheduling on a VPS

Option 1 — cron (simple):

bash
# /etc/cron.d/scraper
0 */6 * * *  scraper  /opt/scraper/bin/scraper >> /var/log/scraper.log 2>&1

Option 2 — systemd timer (better — journald logs, retry semantics):

/etc/systemd/system/scraper.service:

ini
[Unit]
Description=Web Scraper
[Service]
Type=oneshot
User=scraper
ExecStart=/opt/scraper/bin/scraper
EnvironmentFile=/opt/scraper/.env

/etc/systemd/system/scraper.timer:

ini
[Unit]
Description=Run scraper every 6 hours

[Timer]
OnCalendar=00/6:00
Persistent=true
RandomizedDelaySec=300

[Install]
WantedBy=timers.target
bash
sudo systemctl enable --now scraper.timer
systemctl list-timers

Respecting robots.txt and ethics

rust
// Sketch: fetch robots.txt before scraping a site
let robots = client.get("https://target.example/robots.txt").send().await?.text().await?;
// Parse it with a robots.txt crate (for example `texting_robots`)
// and skip any URL your user agent is not allowed to fetch

Ethical scraping rules:

  • Respect robots.txt
  • Rate limit per domain (about one request a second is a polite default)
  • Identify yourself with a User-Agent + contact URL
  • Cache responses — don't re-scrape unchanged pages
  • Don't scrape private, paywalled or login-protected content
  • Don't collect personal data you don't need; in India the Digital Personal Data Protection Act 2023 applies to personal data you process

Ignoring these is not only bad manners: it can get your IP blocked or lead to legal action. Heavy or abusive scraping from a server can also breach your hosting provider's acceptable-use terms.

Common pitfalls

Running out of file descriptors
Many concurrent connections use many file descriptors. Raise the limit (LimitNOFILE= in the systemd unit) and reuse one Client, which pools connections.
Memory blows up
Holding every row in memory. Stream instead: process and write as you go.
No retry on transient errors
503s and timeouts should be retried with backoff. Use a crate such as backon or tokio-retry, or a small loop.
Blocking inside async code
A sync call (heavy file IO, CPU work) inside async code stalls the runtime thread. Use tokio::task::spawn_blocking or rayon.
Hitting rate limits
buffer_unordered(100) is too aggressive for one site. Drop to a handful per domain, or add a governor rate limiter.
Cookies leaking between sites
Cookie storage is off by default in reqwest. Turn it on (the cookies feature plus cookie_store(true)) only for the site that needs a session, and use a separate Client per site.

FAQ

Rust or Python for scraping?

Python for <10K URLs and fast development. Rust for >100K URLs, huge files, or when reliability matters (no GC pauses, predictable memory). Both are valid.

How much RAM does a Rust scraper need?

Usually very little: a plain HTTP scraper often runs in tens of megabytes, depending on concurrency and how much data you hold in memory. Headless browsers are the exception; each browser instance needs hundreds of megabytes.

Is scraping legal?

Depends on target's terms + local law. Public data with proper rate limiting + robots.txt compliance is generally OK. Always consult a lawyer for commercial scraping.

Headless browser needed?

For plain HTML, reqwest and scraper are enough. For pages rendered by JavaScript, use chromiumoxide (Chrome DevTools Protocol) or fantoccini (WebDriver). Both drive a real browser, which is much heavier. First check whether the site loads its data from a JSON API you can call directly.

Can I run a Rust scraper on Domain India shared hosting?

No. Shared hosting stops long-running processes, cron jobs run at most every 4 minutes, and you can't install a Rust toolchain or system services. Run scrapers and data pipelines on a VPS, where you have full root access.

Running this on Domain India

  • VPS: self-managed with full root access, so you can install the Rust toolchain (or copy a compiled binary), PostgreSQL and systemd timers. Plans differ in vCPU and RAM; rayon uses every core you give it.
  • Build once, copy the binary: compile with cargo build --release on your computer or in CI for the same Linux target, then copy the binary to the VPS, so the server doesn't need a compiler.
  • Backups: VPS plans include no backups or snapshots. Export important data regularly and store a copy elsewhere.
  • Shared hosting and App Platform: not suited to scheduled scrapers. Shared hosting stops long-running processes, and the App Platform is for web apps.

Ready to run your pipeline? Choose a Domain India VPS, read Rust web applications with Actix and Axum to add an API on top, or open a ticket with hosting questions.

Run your data pipelines on a VPS

A self-managed VPS with full root access for Rust binaries, PostgreSQL and scheduled jobs.

View VPS plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app
Rust for Web Scraping and Data Pipelines on DomainIndia VPS