When a scraping or ETL job outgrows a quick Python script, Rust is a strong next step: fast, low on memory and predictable. This guide builds a small but production-shaped pipeline and schedules it on a Linux VPS.
Rust's speed + memory safety makes it ideal for data pipelines that scrape, transform, or ingest millions of records. This guide shows async scraping with reqwest + scraper, parsing JSON/CSV at scale, writing to PostgreSQL with sqlx, and scheduling it with systemd on your own VPS.
Why Rust for data pipelines
Typical scraping/ETL workloads in Python:
- scraping thousands of URLs one at a time with requests and BeautifulSoup is slow
- parsing multi-gigabyte JSON files is CPU-bound in pure Python
- loading into a database is bottlenecked by per-row inserts
The Rust versions of the same jobs are usually much faster:
- async reqwest fetches many URLs concurrently, so network waits overlap
- serde_json parses quickly, and rayon spreads CPU work over all cores
- PostgreSQL's
COPY FROMloads rows in bulk instead of oneINSERTat a time
(Python can close much of the gap with asyncio and bulk loading too; the exact speed-up depends on your data and the target sites.) Rust pays off when data work is CPU- or IO-heavy and runs often. For light scraping, say a hundred URLs a day, Python is fine and quicker to write.
Project scaffold
cargo new --bin scraper && cd scraperCargo.toml:
[dependencies]
tokio = { version = "1", features = ["full"] }
reqwest = { version = "0.12", features = ["json", "gzip", "rustls-tls"] }
scraper = "0.23" # HTML parsing
serde = { version = "1", features = ["derive"] }
serde_json = "1"
anyhow = "1" # error handling
tracing = "0.1"
tracing-subscriber = "0.3"
futures = "0.3"
sqlx = { version = "0.8", features = ["runtime-tokio", "tls-rustls", "postgres", "chrono"] }
csv = "1.3"
rayon = "1" # parallel CPU work
governor = "0.8" # rate limitingVersion numbers move quickly; run cargo add crate-name or check crates.io for the current release of each.
Pattern 1 — Async bulk scraping
Scrape 10,000 URLs concurrently, extract title + price, save to CSV.
use futures::stream::{self, StreamExt};
use reqwest::Client;
use scraper::{Html, Selector};
use serde::Serialize;
#[derive(Serialize, Debug)]
struct Product {
url: String,
title: String,
price: Option<String>,
}
async fn scrape_one(client: &Client, url: &str) -> anyhow::Result<Product> {
let html = client.get(url).send().await?.error_for_status()?.text().await?;
let doc = Html::parse_document(&html);
let title_sel = Selector::parse("h1.product-title").unwrap();
let price_sel = Selector::parse("span.price").unwrap();
let title = doc.select(&title_sel)
.next()
.map(|n| n.text().collect::<String>().trim().to_string())
.unwrap_or_default();
let price = doc.select(&price_sel)
.next()
.map(|n| n.text().collect::<String>().trim().to_string());
Ok(Product { url: url.to_string(), title, price })
}
#[tokio::main]
async fn main() -> anyhow::Result<()> {
tracing_subscriber::fmt::init();
let urls: Vec<String> = std::fs::read_to_string("urls.txt")?
.lines().map(String::from).collect();
let client = Client::builder()
.user_agent("MyScraper/1.0 (+https://yourcompany.com/bot)")
.timeout(std::time::Duration::from_secs(10))
.gzip(true)
.build()?;
let results: Vec<Product> = stream::iter(urls.iter())
.map(|url| {
let client = &client;
async move {
scrape_one(client, url).await
.unwrap_or_else(|e| {
tracing::warn!("failed {url}: {e}");
Product { url: url.to_string(), title: String::new(), price: None }
})
}
})
.buffer_unordered(50) // 50 concurrent requests
.collect()
.await;
let mut wtr = csv::Writer::from_path("products.csv")?;
for p in &results { wtr.serialize(p)?; }
wtr.flush()?;
tracing::info!("scraped {} products", results.len());
Ok(())
}buffer_unordered(50) caps concurrency — essential to avoid being rate-limited or overwhelming target servers.
Pattern 2 — Streaming JSON parsing
For a 5 GB JSON file you can't load into memory, stream line-by-line (if NDJSON) or use serde_json::Deserializer:
use std::fs::File;
use std::io::{BufRead, BufReader};
use serde::Deserialize;
#[derive(Deserialize, Debug)]
struct Event {
user_id: u64,
action: String,
timestamp: i64,
}
fn process_ndjson(path: &str) -> anyhow::Result<u64> {
let file = File::open(path)?;
let reader = BufReader::new(file);
let mut count = 0;
for line in reader.lines() {
let line = line?;
if line.is_empty() { continue; }
let event: Event = serde_json::from_str(&line)?;
// process event...
count += 1;
if count % 100_000 == 0 {
tracing::info!("processed {}", count);
}
}
Ok(count)
}Memory usage: flat, regardless of file size.
Pattern 3 — Parallel CPU work with rayon
For transforming millions of rows, Tokio isn't the right tool (async is for IO). Use rayon for CPU:
use rayon::prelude::*;
let rows: Vec<RawRow> = load_all();
let processed: Vec<ProcessedRow> = rows
.par_iter()
.map(|row| expensive_transform(row))
.collect();.par_iter() uses all CPU cores automatically.
Pattern 4 — Bulk DB inserts with sqlx + COPY
Per-row inserts are slow. For 1M rows, use PostgreSQL's COPY FROM:
use sqlx::postgres::PgPool;
use std::io::Write; // for writeln! into a Vec<u8>
async fn bulk_insert(pool: &PgPool, rows: &[Product]) -> anyhow::Result<()> {
let mut conn = pool.acquire().await?;
let mut copy = conn.copy_in_raw(
"COPY products (url, title, price) FROM STDIN (FORMAT csv)"
).await?;
let mut buf = Vec::with_capacity(1024 * 1024);
for p in rows {
writeln!(
&mut buf,
"{},{},{}",
csv_escape(&p.url),
csv_escape(&p.title),
csv_escape(p.price.as_deref().unwrap_or(""))
)?;
}
copy.send(buf.as_slice()).await?;
copy.finish().await?;
Ok(())
}
fn csv_escape(s: &str) -> String {
if s.contains([',', '"', '\n', '\r']) {
format!("\"{}\"", s.replace('"', "\"\""))
} else {
s.to_string()
}
}COPY is typically orders of magnitude faster than individual INSERTs for large batches. For very large loads, send the buffer in chunks rather than building one huge Vec.
Pattern 5 — Rate limiting yourself
Be a good citizen — don't hammer target servers.
use governor::{Quota, RateLimiter};
use std::num::NonZeroU32;
let quota = Quota::per_second(NonZeroU32::new(10).unwrap()); // 10 req/sec
let limiter = RateLimiter::direct(quota);
for url in &urls {
limiter.until_ready().await;
scrape_one(&client, url).await?;
}Scheduling on a VPS
Option 1 — cron (simple):
# /etc/cron.d/scraper
0 */6 * * * scraper /opt/scraper/bin/scraper >> /var/log/scraper.log 2>&1Option 2 — systemd timer (better — journald logs, retry semantics):
/etc/systemd/system/scraper.service:
[Unit]
Description=Web Scraper
[Service]
Type=oneshot
User=scraper
ExecStart=/opt/scraper/bin/scraper
EnvironmentFile=/opt/scraper/.env/etc/systemd/system/scraper.timer:
[Unit]
Description=Run scraper every 6 hours
[Timer]
OnCalendar=00/6:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.targetsudo systemctl enable --now scraper.timer
systemctl list-timersRespecting robots.txt and ethics
// Sketch: fetch robots.txt before scraping a site
let robots = client.get("https://target.example/robots.txt").send().await?.text().await?;
// Parse it with a robots.txt crate (for example `texting_robots`)
// and skip any URL your user agent is not allowed to fetchEthical scraping rules:
- Respect robots.txt
- Rate limit per domain (about one request a second is a polite default)
- Identify yourself with a User-Agent + contact URL
- Cache responses — don't re-scrape unchanged pages
- Don't scrape private, paywalled or login-protected content
- Don't collect personal data you don't need; in India the Digital Personal Data Protection Act 2023 applies to personal data you process
Ignoring these is not only bad manners: it can get your IP blocked or lead to legal action. Heavy or abusive scraping from a server can also breach your hosting provider's acceptable-use terms.
Common pitfalls
LimitNOFILE= in the systemd unit) and reuse one Client, which pools connections.backon or tokio-retry, or a small loop.tokio::task::spawn_blocking or rayon.buffer_unordered(100) is too aggressive for one site. Drop to a handful per domain, or add a governor rate limiter.cookies feature plus cookie_store(true)) only for the site that needs a session, and use a separate Client per site.FAQ
Rust or Python for scraping?
Python for <10K URLs and fast development. Rust for >100K URLs, huge files, or when reliability matters (no GC pauses, predictable memory). Both are valid.
How much RAM does a Rust scraper need?
Usually very little: a plain HTTP scraper often runs in tens of megabytes, depending on concurrency and how much data you hold in memory. Headless browsers are the exception; each browser instance needs hundreds of megabytes.
Is scraping legal?
Depends on target's terms + local law. Public data with proper rate limiting + robots.txt compliance is generally OK. Always consult a lawyer for commercial scraping.
Headless browser needed?
For plain HTML, reqwest and scraper are enough. For pages rendered by JavaScript, use chromiumoxide (Chrome DevTools Protocol) or fantoccini (WebDriver). Both drive a real browser, which is much heavier. First check whether the site loads its data from a JSON API you can call directly.
Can I run a Rust scraper on Domain India shared hosting?
No. Shared hosting stops long-running processes, cron jobs run at most every 4 minutes, and you can't install a Rust toolchain or system services. Run scrapers and data pipelines on a VPS, where you have full root access.
Running this on Domain India
- VPS: self-managed with full root access, so you can install the Rust toolchain (or copy a compiled binary), PostgreSQL and systemd timers. Plans differ in vCPU and RAM; rayon uses every core you give it.
- Build once, copy the binary: compile with
cargo build --releaseon your computer or in CI for the same Linux target, then copy the binary to the VPS, so the server doesn't need a compiler. - Backups: VPS plans include no backups or snapshots. Export important data regularly and store a copy elsewhere.
- Shared hosting and App Platform: not suited to scheduled scrapers. Shared hosting stops long-running processes, and the App Platform is for web apps.
Ready to run your pipeline? Choose a Domain India VPS, read Rust web applications with Actix and Axum to add an API on top, or open a ticket with hosting questions.
A self-managed VPS with full root access for Rust binaries, PostgreSQL and scheduled jobs.
View VPS plans