← Back to Blog

Paperless-ngx on a VPS: Scan, Search and Archive Every Business Document in One Evening

Published · by RS Computers

Paperless-ngx Docker Document management

Paperless-ngx is a free, open-source document archive: you feed it scans, photos and PDFs, it reads the text with OCR (optical character recognition, the step that turns a picture of a page into real text), and from then on every invoice, contract and letter is one search away. On a VPS it runs as three Docker containers, and a small office can go from a shoebox of paper to a searchable archive in one evening. We timed every step below on a Debian 13 test server in Amsterdam with 2 vCPUs and 3 GB of RAM: a scanned one-page invoice in Albanian, German and English was readable and searchable 24 seconds after we dropped it into the inbox folder.

Key facts, checked on 9 October 2026:

The plan for the evening

Think of this as five blocks of work. The times are what it took us, plus a margin for reading.

BlockWhat you doRough time
1. ServerInstall Docker, create a normal userAbout 5 minutes (the package install itself took 17 seconds)
2. PaperlessDownload the official compose files, set passwords and OCR languages, startAbout 15 minutes, of which 60 seconds is the first start
3. First documentsUpload a scan, check the OCR text, search for it10 minutes
4. OrderCorrespondents, tags, document types and their matching rules, plus email import30 minutes or more, depending on your paperwork
5. BackupsNightly export, and one real restore test15 minutes

A few words you will meet. A correspondent is who sent or received the document, such as your electricity company. A document type is what it is: invoice, contract, payslip. A tag is any label you like, and one document can have many. The consumption folder is an inbox on the server: anything placed there is imported and then removed from the folder.

Block 1: a server with Docker

Start with a fresh Debian 13 VPS. Paperless needs about 1 GB of RAM for itself, so a 1 GB server is too tight once the operating system is counted. Our sizing advice is further down. Log in as root and install Docker from Debian's own packages, then create a user called paper that will own everything:

# as root
apt update
apt install -y docker.io docker-compose curl
useradd -m -s /bin/bash -G docker paper
docker compose version

The last command printed Docker Compose version 2.26.1-4 on Debian 13, and docker --version reported 26.1.5. Being in the docker group lets the paper user run containers without sudo. Switch to that user for everything that follows:

# as root
su - paper

Block 2: install Paperless-ngx on the VPS with Docker Compose

The project publishes ready-made compose files on GitHub. We use the PostgreSQL variant, which the setup guide recommends for new installations. It starts three containers: the Paperless web server and worker, PostgreSQL 18 as the database, and Valkey 9 (a Redis-compatible store) as the task queue.

# as the paper user
mkdir -p ~/paperless && cd ~/paperless
BASE=https://raw.githubusercontent.com/paperless-ngx/paperless-ngx/main/docker/compose
curl -fsSL -o docker-compose.yml $BASE/docker-compose.postgres.yml
curl -fsSL -o docker-compose.env $BASE/docker-compose.env
curl -fsSL -o .env $BASE/.env
mkdir -p consume export

Create the consume and export folders yourself, as above. If Docker creates them on first start, they belong to root and the paper user cannot drop files into them.

Three changes before the first start

The downloaded files work, but they publish port 8000 to the whole internet, use the database password paperless, and contain the placeholder secret key change-me. Fix all three:

# as the paper user, in ~/paperless
sed -i 's/- "8000:8000"/- "127.0.0.1:8000:8000"/' docker-compose.yml
DBPASS=$(openssl rand -hex 16)
sed -i "s/POSTGRES_PASSWORD: paperless/POSTGRES_PASSWORD: $DBPASS/" docker-compose.yml
sed -i "/PAPERLESS_DBENGINE: postgresql/a\      PAPERLESS_DBPASS: $DBPASS" docker-compose.yml
SECRET=$(openssl rand -hex 32)
sed -i "s/^PAPERLESS_SECRET_KEY=change-me/PAPERLESS_SECRET_KEY=$SECRET/" docker-compose.env

The first line makes Paperless listen only on the server itself (127.0.0.1). That matters because Docker writes its own firewall rules, and published ports bypass ufw, "effectively ignoring your firewall configuration" (Docker docs). Afterwards, ss -tlnp showed only 127.0.0.1:8000. The other lines set a random database password in both places and a random secret key.

OCR languages: put your main language first

Now add your time zone and languages to the end of docker-compose.env. For a business in Kosovo that also deals with German and English paperwork, that looks like this:

# as the paper user, in ~/paperless
cat >> docker-compose.env <<'EOF'
PAPERLESS_TIME_ZONE=Europe/Amsterdam
PAPERLESS_OCR_LANGUAGES=sqi
PAPERLESS_OCR_LANGUAGE=sqi+deu+eng
EOF

Our lab server sat in Amsterdam, so use your own time zone name. The time zone database has no separate entry for Kosovo, which uses Europe/Belgrade. The two language settings look alike but do different jobs. PAPERLESS_OCR_LANGUAGES (plural) installs extra language packs when the container starts, because Albanian is not in the image. PAPERLESS_OCR_LANGUAGE (singular) chooses which languages are used to read each page. The codes come from Tesseract's list (Tesseract is the OCR engine inside Paperless): sqi for Albanian, deu for German, eng for English.

The order is not cosmetic. Our first attempt used eng+deu+sqi, and the archived text read "IVSH" instead of "TVSH" (the Albanian VAT abbreviation) and "Ruckseite" without its umlaut. We then ran Tesseract inside the container on the same scan with six combinations:

LanguagesShitësi / PrishtinëRückseiteTime for one page
engWrong (é instead of ë)Wrong2.1 s
eng+deuWrongCorrect2.7 s
eng+deu+sqiWrongCorrect3.8 s
sqiCorrectWrong1.7 s
sqi+engWrongWrong2.3 s
sqi+deu+engCorrectCorrect2.8 s

Our lab, 9 October 2026: one A4 page at 300 dpi, Tesseract alone on the image file, timed with docker compose exec. Paperless runs extra steps around Tesseract, so its archived text can differ slightly from this table. Put the language most of your documents are written in first, and add only the languages you really receive, because each extra language costs CPU time.

Start it

# as the paper user, in ~/paperless
docker compose pull
docker compose up -d
docker compose ps

The pull downloads about 3 GB: 2.51 GB for the Paperless image, 457 MB for PostgreSQL and 45 MB for Valkey. Right after up -d, docker compose ps shows the web server as (health: starting). Wait about a minute and run it again until it says (healthy); ours took 60 seconds. You can watch the Albanian pack being added:

# as the paper user, in ~/paperless
docker compose logs webserver | grep tesseract

Expected lines include Installing package tesseract-ocr-sqi... and Installed tesseract-ocr-sqi. Now create your admin account. The command asks for a user name, an email address and a password twice:

# as the paper user, in ~/paperless
docker compose exec webserver createsuperuser

It ends with Superuser created successfully. A quick check that the web server answers:

# as the paper user
curl -sI http://127.0.0.1:8000/ | head -3

You should see HTTP/1.1 302 Found and a redirect to /accounts/login/. Because the port only listens on 127.0.0.1, the browser cannot reach it directly from outside. For a quick look, an SSH tunnel (SSH port forwarding from your own computer to port 8000 on the server) works. For daily use, put a reverse proxy with a real certificate in front, as in our Caddy, Nginx and Traefik guide, and set PAPERLESS_URL to your public address, as the configuration docs describe. We did not test either of these in this lab. A WireGuard VPN is the other good option if only your team needs access.

Block 3: the first scan, and proof that search works

For a fair test we made a scan the way a cheap office scanner would: an A4 grayscale image at 300 dpi, slightly crooked and noisy, holding an invoice in Albanian with German and English lines. The PDF contained no text at all. We copied it into the inbox:

# as the paper user
cp ~/scan-fatura.pdf ~/paperless/consume/

After 24.3 seconds the document was in the archive, the consume folder was empty again, and the log said Output file is a PDF/A-2b (as expected). PDF/A is the archive flavour of PDF, made to stay readable for decades; Paperless stores it next to your untouched original. It also picked up the date printed on the page, 2 October 2026.

Every search for words from the page found it: Faleminderit, pagesë, the invoice number RS-2026-0417, and even bashkepunimin typed without its accent. The search ignores accents, which forgives small OCR slips.

In the browser you simply drag files onto the dashboard. Scripts use the REST API; this is how we uploaded a JPG:

# as the paper user
curl -u admin:YOUR_PASSWORD -F "document=@scan-fatura.jpg" \
  -F "title=Fatura RS-2026-0417 (photo)" \
  http://127.0.0.1:8000/api/documents/post_document/

It answers with a task ID in quotes. That image went through OCR in 18.0 seconds. A five-page scanned PDF of 4.9 MB took 38.2 seconds, so the first page costs about 18 seconds and each extra page about 5 more on 2 vCPUs.

Block 4: let Paperless file things for you

Each correspondent, tag and document type can carry a matching rule. When a new document arrives, Paperless checks its text against every rule and fills in the fields by itself. The matching options are Any (one of the words is enough), All (every word, any order), Exact (the phrase as written), Regular expression, Fuzzy, and Auto, which learns from how you filed earlier documents and, per the docs, retrains about once an hour.

In the web interface you set the matching algorithm and match text on each correspondent, tag or document type. We created ours through the API with the same fields: correspondent "Kafeneja Dardania" matching Dardania, document type "Invoice" matching FATURË RECHNUNG INVOICE, tag "VAT invoice" matching TVSH MwSt VAT, all with Any.

The first scan got its correspondent and document type straight away. The tag did not, because matching runs on the OCR text, not on the paper: OCR had read "TVSH" as "IVSH". Open a document's details to see the text Paperless actually read, and build rules from words that come out reliably. After we added the misread form to the tag, this command applied the rules to documents already in the archive:

# as the paper user, in ~/paperless
docker compose exec -T webserver document_retagger --tags --no-progress-bar

Its summary said Tags added 1. After we switched the language order to sqi+deu+eng, the next upload read "TVSH" correctly and received the correspondent, the document type and the tag on its own.

Email import

Many invoices never exist on paper. In the Mail settings you add an IMAP account, then mail rules: which folder to watch, which senders or subjects to filter, and what happens to the email afterwards (mark as read, move, flag or delete). Paperless checks every 10 minutes by default, and Gmail and Outlook can also connect with OAuth2 (mail docs). A tidy setup is a folder called "To Paperless" and a rule that moves processed mail to "Archived", so filing an invoice means dragging one email. Fetching uses IMAP, usually on port 993. Outbound port 25 is blocked on new RS Computers servers (opened on request via Telegram), but that only affects sending mail, as in our mailcow setup. We did not connect a real mailbox in the lab, so this part follows the documentation.

Block 5: backups you have actually restored

An archive on one disk is one failure away from gone. Paperless includes a document exporter that writes every original, every archive PDF, the thumbnails and a manifest.json with all tags, correspondents, users and settings into a plain folder:

# as the paper user, in ~/paperless
docker compose exec -T webserver document_exporter ../export --no-progress-bar
ls -la export/

Ours took 5.8 seconds and produced 3.2 MB for two documents. The files have readable names such as 2026-10-02 Kafeneja Dardania scan-fatura.pdf, so even without Paperless you could open the folder and find an invoice. Run it again and it updates only what changed, which suits rsync and other incremental backup tools. To run it every night at 02:30:

# as the paper user (as root, run "apt install -y cron" first if crontab is missing)
(crontab -l 2>/dev/null; echo "30 2 * * * cd /home/paper/paperless && docker compose exec -T webserver document_exporter ../export --delete --no-progress-bar >> /home/paper/export.log 2>&1") | crontab -
crontab -l

-T is needed because cron has no terminal, and --delete removes files from the export folder that are no longer part of the archive. Then copy the export folder off the server, encrypted, with restic as in our offsite backup guide.

The restore test most guides skip

A backup is a theory until you restore it. The importer needs a completely empty Paperless, so we started a second copy next to the first, on port 8001, and fed it the export:

# as the paper user
mkdir -p ~/paperless-restore && cd ~/paperless-restore
cp ~/paperless/docker-compose.yml ~/paperless/docker-compose.env .
echo "COMPOSE_PROJECT_NAME=paperless-restore" > .env
sed -i 's/127.0.0.1:8000:8000/127.0.0.1:8001:8000/' docker-compose.yml
mkdir -p consume export && cp -a ~/paperless/export/. export/
docker compose up -d

Wait until docker compose ps shows (healthy) (63 seconds for us), then import and check. Skip createsuperuser here, because the users come with the import.

# as the paper user, in ~/paperless-restore
docker compose exec -T webserver document_importer ../export --no-progress-bar
docker compose exec -T webserver document_sanity_checker --no-progress-bar

The importer printed Checking the manifest, Copy files into paperless... and Updating search index..., and finished in 6.2 seconds. The sanity checker answered No issues detected. The original admin password worked on port 8001, and a search for "Faleminderit" found both documents with their tags. Afterwards, docker compose down -v in ~/paperless-restore removed the test copy. Import into the same Paperless version that made the export, and repeat the test after upgrades.

How much server you need: our measurements

MeasurementResult in our lab
RAM at restWeb server and worker 767 to 789 MiB, PostgreSQL 49 MiB, Valkey 14 MiB; 873 to 918 MB used in the whole 3 GB machine
RAM while reading a pageMain container peaked at 1.04 GiB, using 130 to 143% CPU (out of 200% for 2 vCPUs)
Disk for the software3.0 GB of images, 69 MB for an almost empty database
Disk per scanned pageAbout 1.6 MB for our noisy 300 dpi scan: 976 KB original, 603 KB archive PDF, 10 KB thumbnail
Time per document18 seconds for one page, 38 seconds for five pages

Measured on 9 October 2026 with docker stats, free -m and docker system df in a Debian 13 container limited to 2 vCPUs and 3 GB of RAM. Clean office scans compress better than our deliberately noisy page, so treat 1.6 MB as an upper estimate: 5,000 pages a year would need about 8 GB.

What that means for RS Computers plans, all with NVMe storage and shared vCPUs:

How long must you keep business records?

This is general information, not legal advice. Ask your accountant which rules apply to your company. Here is what the official sources say for three countries our customers often work in:

CountryRetention periodSource
Netherlands7 years for basic records (ledger, sales and purchase records, invoices, payroll); 10 years for real estate records and One Stop Shop VAT dataBelastingdienst
Germany10 years for books, annual accounts and inventories; 8 years for accounting vouchers such as invoices; 6 years for business letters§ 147 AO
KosovoAt least 10 years for the journal and general ledger, from the last day of the financial year; at least 5 years for the accounting documents behind them; payroll records indefinitely; at least 3 years for sales books and payment documentsLaw No. 06/L-032, Article 20 (Official Gazette, 19 April 2018)

A few points matter for a digital archive. The Belastingdienst says digital records must remain usable for the whole period, and a printout alone is not enough. German law accepts digital copies of letters and vouchers that match the original visually and can be made readable and analysed by machine without delay, but annual accounts and the opening balance sheet must be kept as originals. In Kosovo, Article 20 says an electronic general ledger must be locked against changes after year end and signed with an electronic signature, or printed and bound. The Government published a draft amendment to Law 06/L-032 in April 2026, so check whether it has passed. Paperless is not a certified archive system; ask your accountant whether it satisfies an audit.

Where the server sits matters too. These documents are full of personal data. Our Amsterdam and Dublin locations are inside the EU, and a VPS in Prishtina keeps a Kosovo company's records in Kosovo. Availability by city is on the plans page.

Frequently asked questions

Is Paperless-ngx free for business use?

Yes. Paperless-ngx is open source under the GPL-3.0 licence, with no paid edition or per-user fee. You pay only for the server it runs on and your own time.

How much RAM does Paperless-ngx need?

In our lab test of version 3.3.0 with PostgreSQL and Valkey, the three containers used about 830 MB at rest, and the main container peaked at 1.04 GB while reading a scanned page. A 2 GB server is the practical minimum, and 4 GB is comfortable for a small business.

Can Paperless-ngx read Albanian documents?

Yes. Add PAPERLESS_OCR_LANGUAGES=sqi to install the Albanian language pack, and set PAPERLESS_OCR_LANGUAGE=sqi+deu+eng or a similar combination with Albanian first. In our test, putting Albanian first was the difference between "Prishtinë" and "Prishtiné".

How do I back up Paperless-ngx?

Run the built-in document_exporter every night with cron. It writes all originals, archive PDFs and a manifest with tags and settings into a folder. Copy that folder to another location, and test it at least once with document_importer on an empty instance of the same version.

Should I expose Paperless-ngx directly to the internet?

No. Bind port 8000 to 127.0.0.1 and put a reverse proxy with HTTPS in front, or reach it over a VPN or SSH tunnel. Docker's published ports bypass ufw, so a firewall rule alone does not protect a port published on all addresses.

Can Paperless-ngx import invoices from email automatically?

Yes. You add an IMAP account and mail rules in the settings. Paperless checks every 10 minutes by default, imports matching attachments, and then marks the email as read, moves it, flags it or deletes it, depending on the rule.

Tomorrow morning, the shoebox goes

By the end of the evening you have a private archive that reads Albanian, German and English, files documents by itself, takes scans from a folder, an upload or your mailbox, and exports a backup every night that you have already restored once. If you want it on a server near your office, pick a plan with at least 2 GB of RAM on the VPS and VDS plans page. Prefer that someone else sets it up? Message us on Telegram or email us, and we will quote the installation for your setup.

← All articles

Chat on Telegram