← Back to Blog

A VPS for Web Scraping: Responsible Monitors with Python and Playwright

Published · by RS Computers

Web scraping Python Playwright

If you want to know when a supplier changes a price or a new tender appears, a small VPS for web scraping is enough: a Linux server that fetches pages on a timer, from its own IP address, and stores what it finds in a database. Plain HTTP requests do almost all of that work. A headless browser (Chromium driven by code, with no screen) is the exception. The rules come first, because a monitor that gets blocked in its second week is worthless, and after them three small monitors on Debian 13 that you can copy line by line.

Summary:

Rules of the road before the first request

Site owners are tired of bots. Imperva's 2026 Bad Bot Report counts automated traffic at more than 53% of all web traffic in 2025, and Read the Docs reported that one AI crawler downloaded 73 TB from it in a single month of 2024.

robots.txt, by the standard

robots.txt is the file at a site's root that tells crawlers which paths they may fetch, standardised in 2022 as RFC 9309. Its rules are not access authorization, so an "allow" never overrides a site's terms. If the request for robots.txt itself gets a 4xx answer, you may crawl; a 5xx answer means you treat the whole site as off limits. Cache the file for 24 hours at most and parse at least 500 KiB of it. Crawl-delay is not in the standard, but honour it anyway.

Identify yourself and go slowly

httpx2, the HTTP client this guide uses, announces itself as python-httpx2/2.13.1, which some sites block on sight. Send a User-Agent that names your bot and links to a page about it, such as AcmePriceBot/1.0 (+https://acme.example/bot), plus a From header with a mailbox someone reads. One static page behind Caddy with a free certificate is enough for that bot page. Never borrow a browser's User-Agent. Fetch one page at a time per site, with a pause, and stop when a site answers 429 Too Many Requests. On repeat visits, send back the page's ETag or Last-Modified value; an unchanged page then answers 304 Not Modified with no body.

Anyone can copy a User-Agent, so bot identity is moving to signatures. On 1 September 2026 the IETF's Web Bot Auth working group adopted HTTP Message Signatures for automated traffic as a standards-track draft: the bot signs each request with a private key and publishes the public key on its own HTTPS site. Cloudflare already accepts these signatures in its verified bots programme. A monitor that visits a few pages a day can skip it for now; an honest name and a mailbox that answers still do the job.

A CAPTCHA or a login wall means no. Proxies and CAPTCHA solvers are left out of this guide on purpose; ask the owner for an export or an allow-list entry instead.

Often, but "it was public" is not the test, and this is not legal advice. Terms of service can bind you even when the data has no legal protection. In Ryanair v PR Aviation (C-30/14, 15 January 2015), a flight price scraping case, the EU Court of Justice held that the owner of a database protected by neither copyright nor the database right may restrict its use by contract. Where that right does apply, CV-Online Latvia v Melons (C-762/19, 3 June 2021) held that a search engine copying and indexing a substantial part of a job-ad site extracts its database, which the owner can forbid when that harms the investment behind it.

Personal data is the bigger issue: names and contacts in job ads or tender notices count even when public. The GDPR requires a lawful basis, usually legitimate interest, and Article 14 obliges you to inform the people concerned within one month. The CNIL's guidance on scraping for AI training sets a bar worth meeting for any scraper: define in advance what you collect, delete anything else straight after collection, and leave out sites that object through robots.txt or CAPTCHAs. Kosovo's Law No. 06/L-082 follows the GDPR.

And hiQ Labs v. LinkedIn, the US case quoted as proof that public data is fair game, ended in December 2022 with a permanent injunction against hiQ for breaching LinkedIn's User Agreement.

Laptop, scraping API or a VPS for web scraping?

Close the laptop and the scraper on it stops. Scraping APIs keep running but bill every request, and on one popular API a JavaScript-rendered page costs five credits against one for plain HTML.

On your own server, the limits are your politeness and your RAM. The HTTP Archive Web Almanac 2025 puts the median mobile inner page, the kind a product page is, at 1.8 MB, of which only about 20 KB is HTML. So 500 product pages cost about 10 MB as HTML and roughly 0.9 GB as full browser loads (our arithmetic). No polite scraper fills a 1 Gb/s port; the 10 Gb/s port on VDS plans is for bulk research downloads.

httpx2, Scrapy or Playwright?

Check the browser's Network tab first. If the data arrives as JSON, request that endpoint (terms permitting) and skip HTML parsing. A JSON endpoint is also simple enough to poll without code, from a Schedule trigger and an HTTP Request node in self-hosted n8n.

ToolPick it whenCost and trap
httpx2 + BeautifulSoupThe data is in the HTML and you know the URLsTens of MB. Redirects are not followed by default, and the timeout is 5 seconds.
ScrapyYou follow links across a whole catalogueOne process. start_requests() was removed in 2.16 in favour of start(), so old tutorials crawl nothing.
PlaywrightContent appears only after JavaScript runsThe heaviest, with no official memory figure. Unclosed pages leak memory.

We default to httpx2, switch to Scrapy when a job has to follow links, and use Playwright only when nothing else gets the data. A new Scrapy project already obeys robots.txt with one request per domain and a one-second delay; just set your USER_AGENT and turn on AUTOTHROTTLE_ENABLED in settings.py.

Why httpx2 and not the better-known httpx? httpx has had no stable release since 0.28.1 in December 2024, and its 1.0 previews are a redesign still in early development. So in May 2026 Pydantic took the familiar code over as httpx2 to keep fixes, security ones included, coming. The API is the same under the new name (import httpx becomes import httpx2), and Scrapy 2.18 moved its httpx download handler to it in August. One difference: httpx2 checks HTTPS certificates against the server's own certificate store, which is why step 2 installs ca-certificates.

Where to run it: Amsterdam, Dublin or Prishtina

RS Computers runs KVM virtual machines with NVMe storage in Amsterdam (Netherlands), Dublin (Ireland) and Prishtina (Kosovo). Every server gets its own IPv4 and IPv6 address. No other customer's bot sends requests from it, so what a site sees from that address is your bot alone, and an owner who agrees to let you in can allow-list exactly that address. Traffic is unmetered on a 1 Gb/s port (VPS plans) or 10 Gb/s (VDS plans), so a big download costs nothing extra. If the browser job you measure further down creeps toward the plan's RAM, move up a plan from the client area; it takes a short reboot on the same server.

Amsterdam and Dublin are inside the EU, which keeps things simple when you collect personal data; Prishtina suits Kosovo and the region. The plans page shows which locations have servers ready.

Picking a plan: HTTP jobs, browser jobs and research crawls

Debian 13 runs in 1 GB without a desktop, and PostgreSQL's default shared_buffers is 128 MB; browsers are what fill a server.

PlanCrawls comfortably
VPS Nano (1 vCPU, 1 GB, 20 GB NVMe, 1 Gb/s)The price monitor and tender tracker. No browser.
VPS Micro (2 vCPU, 2 GB, 40 GB NVMe, 1 Gb/s)All three examples, rendering one page at a time.
VPS Mini (4 vCPU, 4 GB, 80 GB NVMe, 1 Gb/s)Several monitors, or a few browser pages in parallel.
VDS Small (4 vCPU, 8 GB, 240 GB NVMe, 10 Gb/s)Browser-heavy monitoring of many sites.
VDS Medium (8 vCPU, 16 GB, 480 GB NVMe, 10 Gb/s)Research datasets and Common Crawl extracts.
VDS Large (16 vCPU, 32 GB, 960 GB NVMe, 10 Gb/s)Large research pipelines or many parallel browser jobs.

For broad research, start from Common Crawl's September 2026 archive, published on 19 September with 2.17 billion pages from 33.2 million domains, and crawl only the gaps.

Setting up a web scraping server on Debian 13

Everything below is typed in one root SSH session. Your scrapers never run as root, though: wherever a line begins with runuser -u scraper --, root hands that one command to the unprivileged scraper user. If SSH is still new to you, the first week of our Linux practice lab guide covers it.

1. Update and allow only SSH

Scrapers only make outgoing connections, so the firewall needs a single open port: SSH.

# as root
apt update && apt full-upgrade -y
apt install -y ufw
ufw allow 22/tcp
ufw --force enable
ufw status verbose

Expect Status: active and 22/tcp ALLOW IN. Always allow 22/tcp before enabling ufw, or you lock yourself out.

2. Install Python and PostgreSQL

Debian 13's own repository has everything this step needs: Python 3.13 with its venv module, PostgreSQL 17 and the list of trusted certificate authorities that httpx2 checks HTTPS sites against.

# as root
apt install -y python3 python3-venv postgresql ca-certificates
python3 --version
pg_lsclusters

Expect Python 3.13.5 and a line starting 17 main 5432 online.

3. Create the user and install the tools

Next comes a scraper user that cannot log in, with pinned package versions in its own virtual environment, a private Python folder that apt leaves alone.

# as root (each runuser line runs as the scraper user)
useradd --system --create-home --home-dir /opt/scraper --shell /usr/sbin/nologin scraper
mkdir -p /opt/scraper/app /opt/scraper/data
cat > /opt/scraper/app/requirements.txt <<'EOF'
httpx2==2.13.1
beautifulsoup4==4.15.0
lxml==6.1.3
scrapy==2.19.0
playwright==1.63.0
psycopg[binary]==3.3.6
protego==0.7.0
EOF
chown -R scraper:scraper /opt/scraper
cd /opt/scraper
runuser -u scraper -- python3 -m venv venv
runuser -u scraper -- venv/bin/pip install -r app/requirements.txt
runuser -u scraper -- venv/bin/scrapy version

It prints Scrapy 2.19.0. Debian's python3-scrapy (2.12.0) lacks recent security fixes, hence the venv. externally-managed-environment means pip ran outside the venv; ensurepip is not available means python3-venv is missing: install it, delete /opt/scraper/venv, retry.

4. Install the browser (example three only)

Skip this step unless a page needs JavaScript. Root installs Chromium's system libraries, then scraper downloads only the light headless shell.

# as root (the runuser line runs as the scraper user)
cd /opt/scraper
venv/bin/playwright install-deps chromium
export PLAYWRIGHT_BROWSERS_PATH=/opt/scraper/ms-playwright
runuser -u scraper -- venv/bin/playwright install --only-shell chromium
ls /opt/scraper/ms-playwright

You should see a folder starting with chromium. If a job later reports Executable doesn't exist, the install and run paths differ.

5. Create the database

PostgreSQL gets a scraper role and a database of the same name, and the last line connects to check both.

# as root (runuser runs each line as the postgres or scraper user)
cd /tmp
runuser -u postgres -- createuser scraper
runuser -u postgres -- createdb -O scraper scraper
runuser -u scraper -- psql -c 'select version();'

A row starting PostgreSQL 17 appears without a password: Debian's peer authentication maps the Linux user to the same-named role, and PostgreSQL listens only on localhost. If peer authentication fails, the names differ; never open port 5432 to fix it. More in our database server guide.

6. Store the bot identity and the shared helper

Create a Telegram bot with @BotFather first; our Telegram bot guide shows how to find your chat ID. This stores the bot details in a root-only file and writes common.py, which every example imports.

# as root
cat > /etc/scraper.env <<'EOF'
BOT_NAME=YOUR_BOT_NAME
BOT_URL=https://YOUR_DOMAIN/bot
BOT_EMAIL=bot@YOUR_DOMAIN
TG_TOKEN=YOUR_BOT_TOKEN
TG_CHAT=YOUR_CHAT_ID
EOF
chmod 600 /etc/scraper.env
cat > /opt/scraper/app/common.py <<'EOF'
import os, time
from urllib.parse import urljoin
import httpx2, psycopg
from protego import Protego

env = os.environ
UA = f"{env['BOT_NAME']}/1.0 (+{env['BOT_URL']}; {env['BOT_EMAIL']})"
robots = {}

def client():
    return httpx2.Client(headers={"User-Agent": UA, "From": env["BOT_EMAIL"]},
                         timeout=20, follow_redirects=True)

def allowed(c, url):
    txt = urljoin(url, "/robots.txt")
    if txt not in robots:
        r = c.get(txt)
        robots[txt] = None if r.status_code >= 500 else Protego.parse(
            r.text if r.status_code < 400 else "")
    return robots[txt] is not None and robots[txt].can_fetch(url, env["BOT_NAME"])

def get(c, url, **kw):
    time.sleep(5)
    r = c.get(url, **kw)
    if r.status_code == 429:
        raise SystemExit(f"429 Too Many Requests from {url}")
    return r

def db():
    return psycopg.connect("dbname=scraper")

def alert(text):
    httpx2.post(f"https://api.telegram.org/bot{env['TG_TOKEN']}/sendMessage",
                data={"chat_id": env["TG_CHAT"], "text": text[:4000]}).raise_for_status()
EOF
/opt/scraper/venv/bin/python -m py_compile /opt/scraper/app/common.py

No output means no syntax errors. Use one word for the bot name.

7. Add the systemd units

One template unit, scrape@.service, runs any script in /opt/scraper/app by name. A second unit sends the alert when a job fails, and the last line fires a test alert.

# as root
cat > /etc/systemd/system/scrape@.service <<'EOF'
[Unit]
Wants=network-online.target
After=network-online.target postgresql.service
OnFailure=scrape-failed@%i.service

[Service]
Type=oneshot
User=scraper
WorkingDirectory=/opt/scraper/app
EnvironmentFile=/etc/scraper.env
Environment=PLAYWRIGHT_BROWSERS_PATH=/opt/scraper/ms-playwright
ExecStart=/opt/scraper/venv/bin/python %i.py
TimeoutStartSec=30min
MemoryHigh=1200M
MemoryMax=1500M
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/opt/scraper/data
EOF
cat > /etc/systemd/system/scrape-failed@.service <<'EOF'
[Service]
Type=oneshot
User=scraper
WorkingDirectory=/opt/scraper/app
EnvironmentFile=/etc/scraper.env
ExecStart=/opt/scraper/venv/bin/python -c "import common; common.alert('Scraper job %i failed on %H')"
EOF
systemctl daemon-reload
systemctl start scrape-failed@test.service

A Telegram message about job test should arrive; if not, journalctl -u scrape-failed@test says why. 401 Unauthorized there means a wrong token; 400 Bad Request or 403 Forbidden usually means a wrong chat ID or that you never sent your bot a message. TimeoutStartSec stops one hung run from blocking later ones, MemoryMax kills only this job, and ProtectSystem=strict leaves only /opt/scraper/data writable. Never add MemoryDenyWriteExecute; it breaks Chromium.

Example 1: a self-hosted price and stock monitor

Say a Prishtina shop resells two suppliers' products and needs to know the day either of them changes a price. In the European Commission's 2017 e-commerce sector inquiry, 53% of the retailers that answered tracked competitors' online prices, and two thirds of those used software for it. This monitor reads the schema.org product data (JSON-LD) many shops embed for search engines, which survives redesigns better than CSS selectors. It fails loudly when a page yields no price, because a selector that quietly matches nothing can go unnoticed for weeks.

Put real product pages in URLS first. The block writes the monitor, adds a timer for 06:00 and 18:00 (up to 10 minutes late at random) and runs it once.

# as root
cat > /opt/scraper/app/price_monitor.py <<'EOF'
import json, sys
from bs4 import BeautifulSoup
from common import client, allowed, get, db, alert

URLS = ["https://SUPPLIER_ONE/product/PRODUCT_A", "https://SUPPLIER_TWO/item/PRODUCT_B"]

def read_offer(html):
    for tag in BeautifulSoup(html, "lxml").find_all("script", type="application/ld+json"):
        data = json.loads(tag.string or "{}")
        for item in data.get("@graph", [data]) if isinstance(data, dict) else data:
            offer = item.get("offers") if isinstance(item, dict) else None
            offer = offer[0] if isinstance(offer, list) and offer else offer
            if isinstance(offer, dict) and "price" in offer:
                return float(offer["price"]), "InStock" in str(offer.get("availability"))

broken, changes = [], []
with db() as conn, client() as c:
    conn.execute("CREATE TABLE IF NOT EXISTS price_seen (url text, price numeric, "
                 "in_stock boolean, seen_at timestamptz DEFAULT now())")
    for url in URLS:
        if not allowed(c, url):
            continue
        offer = read_offer(get(c, url).text)
        if not offer:
            broken.append(url)
            continue
        last = conn.execute("SELECT price, in_stock FROM price_seen WHERE url = %s "
                            "ORDER BY seen_at DESC LIMIT 1", (url,)).fetchone()
        if last is None or (float(last[0]), last[1]) != offer:
            conn.execute("INSERT INTO price_seen (url, price, in_stock) "
                         "VALUES (%s, %s, %s)", (url, *offer))
            if last:
                changes.append(f"{url}: {last[0]} -> {offer[0]}, in stock: {offer[1]}")

if changes:
    alert("Changed:\n" + "\n".join(changes))
if broken:
    sys.exit("No price found on: " + " ".join(broken))
EOF
cat > /etc/systemd/system/prices.timer <<'EOF'
[Timer]
OnCalendar=*-*-* 06,18:00:00
RandomizedDelaySec=10min
Persistent=true
Unit=scrape@price_monitor.service

[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now prices.timer
systemctl start scrape@price_monitor.service
journalctl -u scrape@price_monitor -n 20

The journal shows the job finishing; later runs alert on changes. No JSON-LD? Use soup.select_one() selectors. Our e-commerce hosting guide covers the shop side.

Example 2: a public tender tracker that needs no scraping

Check for an API before you write a scraper. For EU tenders there is one: TED (Tenders Electronic Daily), the EU's journal of public tenders, has an API that searches published notices without an API key. For a hypothetical IT firm bidding in the Netherlands and Ireland, this tracker filters by country, date and CPV code, the EU's numbered list of what a tender buys (72000000 is IT services, 48000000 software). It runs daily at 07:30 and once right now:

# as root
cat > /opt/scraper/app/tenders.py <<'EOF'
from psycopg.types.json import Jsonb
from common import client, db, alert

QUERY = ("classification-cpv IN (72000000 48000000) AND buyer-country IN (NLD IRL) "
         "AND publication-date>=today(-3)")
r = client().post("https://api.ted.europa.eu/v3/notices/search", json={
    "query": QUERY, "limit": 100,
    "fields": ["publication-number", "notice-title", "deadline-receipt-tender-date-lot"]})
r.raise_for_status()

new = []
with db() as conn:
    conn.execute("CREATE TABLE IF NOT EXISTS tender (id text PRIMARY KEY, data jsonb)")
    for n in r.json()["notices"]:
        nid = n["publication-number"]
        if conn.execute("INSERT INTO tender VALUES (%s, %s) ON CONFLICT DO NOTHING",
                        (nid, Jsonb(n))).rowcount:
            title = (n.get("notice-title") or {}).get("eng", nid)
            new.append(f"{title}\nhttps://ted.europa.eu/en/notice/-/detail/{nid}")

if new:
    alert(f"{len(new)} new tender(s):\n\n" + "\n\n".join(new[:10]))
EOF
cat > /etc/systemd/system/tenders.timer <<'EOF'
[Timer]
OnCalendar=*-*-* 07:30:00
RandomizedDelaySec=10min
Persistent=true
Unit=scrape@tenders.service

[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now tenders.timer
systemctl start scrape@tenders.service

Telegram should show recent matching notices. ON CONFLICT DO NOTHING skips notices already stored, so later runs send only new ones. An HTTP 400 in journalctl -u scrape@tenders means a typo in the query.

Example 3: website change monitoring with alerts

Some pages matter only when they change, like a supplier's terms. This watcher sends back the stored ETag and Last-Modified values, so an unchanged page costs a 304, and hashes only the visible text of <main>, so a new tracking script is not a change. Pages marked True are rendered in Chromium with images, fonts and media blocked, then given three seconds for their scripts to fill in the text. Waiting for the network to go quiet instead timed out after 30 seconds on a real shop page in our test, because that page never stops making requests. The watcher runs every six hours. Put your own pages in PAGES.

# as root
cat > /opt/scraper/app/page_watch.py <<'EOF'
import hashlib
from bs4 import BeautifulSoup
from common import client, allowed, get, db, alert, UA

PAGES = {"https://SUPPLIER_ONE/terms": False, "https://REGULATOR_SITE/notices": True}

def render(url):
    from playwright.sync_api import sync_playwright
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_context(user_agent=UA).new_page()
        page.route("**/*", lambda r: r.abort() if r.request.resource_type
                   in ("image", "font", "media") else r.continue_())
        page.goto(url)
        page.wait_for_timeout(3000)
        html = page.content()
        page.context.close()
        browser.close()
        return html

with db() as conn, client() as c:
    conn.execute("CREATE TABLE IF NOT EXISTS watch (url text PRIMARY KEY, "
                 "etag text, modified text, hash text)")
    for url, needs_browser in PAGES.items():
        if not allowed(c, url):
            continue
        etag, modified, old = conn.execute("SELECT etag, modified, hash FROM watch "
                                           "WHERE url = %s", (url,)).fetchone() or (None,) * 3
        if needs_browser:
            html = render(url)
        else:
            ask = {"If-None-Match": etag, "If-Modified-Since": modified}
            r = get(c, url, headers={k: v for k, v in ask.items() if v})
            if r.status_code == 304:
                continue
            r.raise_for_status()
            html, etag, modified = r.text, r.headers.get("ETag"), r.headers.get("Last-Modified")
        soup = BeautifulSoup(html, "lxml")
        text = (soup.select_one("main") or soup).get_text(" ", strip=True)
        digest = hashlib.sha256(text.encode()).hexdigest()
        if old and digest != old:
            alert(f"Page changed: {url}")
        conn.execute("INSERT INTO watch VALUES (%s, %s, %s, %s) ON CONFLICT (url) DO UPDATE "
                     "SET etag = excluded.etag, modified = excluded.modified, "
                     "hash = excluded.hash", (url, etag, modified, digest))
        conn.commit()
EOF
cat > /etc/systemd/system/pagewatch.timer <<'EOF'
[Timer]
OnCalendar=*-*-* 00/6:15:00
RandomizedDelaySec=10min
Persistent=true
Unit=scrape@page_watch.service

[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now pagewatch.timer
systemctl start scrape@page_watch.service
systemctl list-timers

The list should show all three timers. A page that alerts every run has a clock or banner in its text; target a narrower element. For many browser pages, reuse one browser with a fresh context per page, and close what you open.

Keeping it running: memory, updates, backups and alerts

Headless Chromium has no published per-page memory figure, so measure: this runs the page watcher once and reports its peak.

# as root
systemd-run --wait -p User=scraper -p EnvironmentFile=/etc/scraper.env -p Environment=PLAYWRIGHT_BROWSERS_PATH=/opt/scraper/ms-playwright /opt/scraper/venv/bin/python /opt/scraper/app/page_watch.py

Set MemoryHigh above the Memory peak it prints and MemoryMax a little higher, below the plan's RAM. On our 8 vCPU / 16 GB test server, a watcher run that rendered one real page peaked at about 320 MB for a news front page and 480 MB for a shop product page with Playwright 1.63, while the price monitor and tender tracker peaked at about 40 and 28 MB. On Ubuntu 24.04 the setup is the same, with Python 3.12.3 and PostgreSQL 16, but its systemd 255 prints a meaningless Memory peak; read the peak: value in systemctl status scrape@page_watch while a run is still going.

Once a week or so, update the system and see which Python packages have newer releases:

# as root (the runuser line runs as the scraper user)
apt update && apt full-upgrade -y
cd /opt/scraper
runuser -u scraper -- venv/bin/pip list --outdated

Each line is a package with a newer release; pip itself nearly always appears and can be ignored. Read release notes before raising a pin in requirements.txt, then repeat the install lines from steps 3 and 4; Playwright upgrades need a new browser download.

The free weekly backup covers the whole server. Between those, this dumps the database, checks it and deletes dumps older than 30 days (old data is a GDPR risk):

# as root (runuser runs pg_dump as the postgres user)
cd /tmp
runuser -u postgres -- pg_dump -Fc scraper > /var/backups/scraper-$(date +%F).dump
pg_restore --list /var/backups/scraper-$(date +%F).dump | grep TABLE
find /var/backups -name 'scraper-*.dump' -mtime +30 -delete

Expect TABLE lines for price_seen, tender and watch. Copy dumps off the server, as our off-site backup guide does with restic. OnFailure catches crashes but not a job that never runs; an Uptime Kuma push monitor called at the end of each script covers that (see our monitoring guide).

Locking down a scraping server

Frequently asked questions

Do I have to follow robots.txt?

RFC 9309 says its rules are not access authorization, but follow them. Ignoring robots.txt gets you blocked fastest, the CNIL treats it as an objection when data is scraped to train AI, and in the EU a machine-readable opt-out switches off the general text and data mining exception (Article 4 of the 2019 Copyright Directive) that businesses rely on.

How much RAM do I need to run Playwright or headless Chrome on a server?

Measure it: on Debian 13, run your job with systemd-run --wait and read the Memory peak line. On our test server, a job rendering one real page in headless Chromium peaked at 320 to 480 MB. HTTP scrapers with PostgreSQL fit in 1 to 2 GB; a few parallel browser pages want 2 to 4 GB or more.

Why does my scraper work on my laptop but get blocked on the server?

Many sites distrust data centre IP ranges, and a generic User-Agent makes it worse. If a slow, honestly named bot is still blocked, the site has chosen not to serve bots: use its API or ask for access.

Should I use cron or systemd timers for scrapers?

Either works, but systemd timers log to the journal, catch up on missed runs, spread start times, alert through OnFailure and never overlap a running job.

Can RS Computers set up the monitors for me?

Yes. Message us on Telegram or email info@rscomputers-ks.com with the sites you want watched and how often, and we will suggest a plan and quote the setup.

Which monitor should you build first?

The tender tracker. It scrapes nothing, so it safely proves that the database and the alerts work. Add the price monitor next, and save Playwright for pages that leave you no choice. For the server, take the smallest plan that fits your heaviest job, install Debian 13 and work through the seven steps. All six plans, for each location, are on the VPS and VDS plans page.

← All articles

Chat on Telegram