Most small teams that self-host Open WebUI on a VPS can skip the GPU entirely. The server keeps the accounts, the chat history and the documents, while the heavy thinking happens at OpenAI, Anthropic, Mistral or OpenRouter on your own API keys. What a CPU-only server won't do is run a ChatGPT-class model locally at a speed anyone will sit through. It runs small models, embeddings and a vector database fine, and those are worth keeping in-house.
- Open WebUI needs about 2 GB of RAM before any local model is loaded.
- On shared vCPUs, a 3B model at 4-bit (a 2.0 GB download) is the practical local model. On our platform it wrote about 13 tokens per second on 4 vCPU and 21 on 8 vCPU.
- Ollama's default context on a CPU-only server is 4,096 tokens, which silently cuts long RAG prompts.
- With a cloud API, prompts, history and retrieved document chunks reach the provider. Accounts, files and vectors stay with you.
- Run Open WebUI 0.11.4 or later: it fixes 19 security advisories published at the end of September 2026.
Can you self-host Open WebUI on a VPS without a GPU?
Yes, because the interface and the model are separate programs. Open WebUI is a self-hosted, ChatGPT-style web interface with user accounts, document upload and chat history. It talks to any OpenAI-compatible API and to Ollama, which runs open-weight models on your own machine. Only Ollama cares about graphics cards. Open WebUI wants memory: its performance guide calls 1 GB "very low for anything doing RAG or embeddings" and says to aim for 2 GB or more.
In practice, a CPU server runs Open WebUI for a team, embeddings, a vector database and a 1B to 4B model for background jobs. It runs an 8B model for one patient person. It won't run Llama 3.1 70B, a 43 GB download, at any useful speed, and since Ollama serves one request at a time by default, ten colleagues on one local model mostly queue.
Why host it yourself? In Microsoft and LinkedIn's 2024 Work Trend Index, 78% of people using AI at work brought their own tools. A shared server gives them one login, history on a disk you control and API keys the company owns.
How much RAM does Ollama need? Model size against memory
Start with the download size. Ollama pulls 4-bit quantised models by default; quantisation stores each weight in fewer bits, and llama.cpp's documentation puts Llama 3.1 8B at 4.58 GiB in Q4_K_M against 14.96 GiB unquantised. Add the KV cache, the model's working memory for the conversation, and Open WebUI itself. Sizes are from Ollama's model library; the plan column is our sizing with Open WebUI alongside.
| Model (Ollama tag) | Download | Smallest plan we'd use | Sensible job on a CPU |
|---|---|---|---|
llama3.2:1b | 1.3 GB | VPS Mini | Chat titles, tags, sorting messages |
llama3.2:3b | 2.0 GB | VDS Small | Open WebUI task model, short summaries |
llama3.1:8b | 4.9 GB | VDS Medium | One person chatting fully locally |
qwen3:14b | 9.3 GB | VDS Large | Batch work nobody waits for |
qwen3:30b | 19 GB | VDS Large, tight | Mixture-of-experts with about 3B active parameters, so quicker than its size suggests |
nomic-embed-text | 274 MB | Any plan running Ollama | RAG embeddings, 768 dimensions |
LlamaIndex's local starter asks for about 32 GB to run Llama 3.1 8B, a 4.9 GB download. Ollama's own README used to say 8 GB for 7B models, 16 GB for 13B and 32 GB for 33B, still a fair rule for one user.
Context costs memory too. A token is a chunk of text, roughly a short word. Ollama's context-length docs say machines with under 24 GiB of VRAM, so every CPU-only server, default to a 4,096-token context. Raising OLLAMA_CONTEXT_LENGTH costs RAM, multiplied by OLLAMA_NUM_PARALLEL; Ollama's FAQ says OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves that cache.
Myth: twice the vCPUs means twice the speed
Every new token means reading the model's weights from RAM, so once a few vCPUs are busy, memory bandwidth limits CPU generation more than the vCPU count does. A 2023 Intel paper on CPU inference (arXiv 2311.00502) says LLMs need "large memory capacity and high memory bandwidth".
Jeff Geerling's open benchmarks show it nicely. On the CPU alone, Llama 3.2 3B ran at 9.06 tokens per second on an Intel N150 mini PC, 23.81 on a Ryzen AI 5 340 laptop and 23.52 on a 192-CPU AmpereOne server. A 14B model on the N150 managed 2.13. The big server tied with the laptop.
Those are physical machines, so we measured our own. The figures below come from an 8 vCPU / 16 GB server on our platform (the size of VDS Medium), then on the same server limited to 4 vCPU (VDS Small), with the stack from this guide running. Each model got the same one-line prompt three times through ollama run MODEL --verbose, as in step 7; the table shows the middle result, in tokens per second.
| Model | Writing, 8 vCPU (VDS Medium size) | Writing, limited to 4 vCPU (VDS Small) | Prompt reading, 8 vCPU | Prompt reading, 4 vCPU | RAM in use, model loaded |
|---|---|---|---|---|---|
llama3.2:1b | 28.7 | 22.0 | 108 | 95 | 2.5 GB |
llama3.2:3b | 21.1 | 13.3 | 64 | 38 | 3.5 GB |
llama3.1:8b | 10.5 | 6.5 | 26 | 15 | 6.2 GB |
Writing is the eval rate line, prompt reading the prompt eval rate. RAM is what free -m showed with Open WebUI, Qdrant and the model loaded, the same on 4 and 8 vCPU; the model process alone took 1.5, 2.4 and 5.2 GB. Halving the vCPUs cut the writing speed of the 3B and 8B models by 37 to 38%, not by half, and prompt reading by 40 to 44%.
Watch the prompt, too. We gave llama3.2:3b a 1,392-token prompt, roughly a RAG question with a few retrieved chunks: it read for 19 seconds on 8 vCPU and 35 seconds on 4 vCPU before the first word appeared. A 2023 issue on Ollama's GitHub describes an Azure VM with 8 vCPUs and 32 GiB running Mistral 7B where, by the fourth short question in a chat, replies took over 60 seconds to start. RAG wraps every question in pages of retrieved text, so it suffers most.
The setup worth running: cloud models for chat, local models for chores
Each Open WebUI message can trigger calls you never see: a chat title, tags, follow-up suggestions, a rewritten search query. By default they use the model the user chats with. Open WebUI's FAQ warns this "can significantly increase your costs", and with a cloud model the conversation leaves the server again each time.
So split the work. Keep conversations on a strong cloud model and give the chores to a local llama3.2:3b. Open WebUI's task model guide now also suggests qwen3.5:2b and gemma4:e2b; check the download first, because they are 2.7 GB and 4.6 GB against 2.0 GB, and the speeds above are for Llama only. Embed documents locally with nomic-embed-text, so files never go to an embeddings API. Both fit beside Open WebUI on VDS Small and keep working when a provider is down.
Once the stack below runs, go to Admin Panel, Settings, Interface: set the task model for external models to llama3.2:3b and switch off unused generators. Under Documents, pick Ollama and nomic-embed-text before the first upload, because switching embedders later means re-indexing everything.
The default embedder, all-MiniLM-L6-v2, was trained mostly on English. For Albanian, Serbian or Dutch files, try embeddinggemma (622 MB, over 100 languages), then test it on real questions.
Sizing a plan for API-only, hybrid or fully local AI
RS Computers runs KVM virtual servers in Amsterdam, Prishtina and Dublin. Current capacity per location is listed on the plans page. Each plan is a full virtual machine with NVMe storage (quicker model loads), unmetered traffic on a 1 Gb/s port on VPS plans and 10 Gb/s on VDS plans, and your own IPv4 and IPv6. Free weekly backups restore from the client area, where you can also upgrade. You choose a Linux distribution when ordering; Windows Server is offered on VPS Mini and all VDS plans. There is no GPU, and the vCPUs are shared.
| Plan | Realistic use |
|---|---|
| VPS Nano (1 vCPU, 1 GB RAM, 20 GB NVMe, 1 Gb/s) | A bot calling a cloud API, as in our guide to hosting a Telegram or Discord bot. Too small for Open WebUI. |
| VPS Micro (2 vCPU, 2 GB RAM, 40 GB NVMe, 1 Gb/s) | API-only Open WebUI for a few people, without the Ollama service. |
| VPS Mini (4 vCPU, 4 GB RAM, 80 GB NVMe, 1 Gb/s) | Open WebUI on cloud APIs for a small team, with built-in RAG and llama3.2:1b as task model. |
| VDS Small (4 vCPU, 8 GB RAM, 240 GB NVMe, 10 Gb/s) | The hybrid setup: cloud chat, local llama3.2:3b and embeddings, plus Qdrant. Our default pick. |
| VDS Medium (8 vCPU, 16 GB RAM, 480 GB NVMe, 10 Gb/s) | Add an 8B model for one user at a time, a longer context and PostgreSQL with pgvector. |
| VDS Large (16 vCPU, 32 GB RAM, 960 GB NVMe, 10 Gb/s) | 14B models or qwen3:30b for batch jobs, several models at once, or a RAG back end for other apps. |
Install Open WebUI and Ollama on Debian 13
Use Docker, as Open WebUI's docs recommend: v0.11.4 supports Python 3.11 and 3.12, Debian 13 ships 3.13, and pip install open-webui fails (Ubuntu 24.04 has 3.12, which works in a venv). We ran the blocks below on Debian 13 in October 2026 with Open WebUI 0.11.4, Ollama 0.35.1 and Qdrant 1.19.2. Commands run as root over SSH unless a block says otherwise; if that is new to you, our Linux lab guide starts from the first login.
Step 1: firewall and base packages
This updates the system, installs the tools used later and opens only SSH, HTTP and HTTPS.
# as root
apt update && apt upgrade -y
apt install -y ca-certificates curl gnupg nano openssl ufw python3-venv
ufw allow 22/tcp
ufw allow 80/tcp
ufw allow 443/tcp
ufw --force enable
ufw status verbose
You should see Status: active with ALLOW lines for 22, 80 and 443. Port 3000 is missing on purpose.
Step 2: Docker Engine
This adds Docker's own apt repository, following Docker's Debian install guide, and installs the engine with Compose.
# as root
install -m 0755 -d /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/debian/gpg -o /etc/apt/keyrings/docker.asc
chmod a+r /etc/apt/keyrings/docker.asc
tee /etc/apt/sources.list.d/docker.sources <<EOF
Types: deb
URIs: https://download.docker.com/linux/debian
Suites: $(. /etc/os-release && echo "$VERSION_CODENAME")
Components: stable
Architectures: $(dpkg --print-architecture)
Signed-By: /etc/apt/keyrings/docker.asc
EOF
apt update
apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
docker run hello-world
Look for "Hello from Docker!". On Ubuntu 24.04, write ubuntu instead of debian in both URLs and remove the distribution's docker.io package first. If volumes and Compose files are new to you, our guide to Docker and Compose on a VPS explains them before you go on.
Step 3: the Compose stack
First come the stack's folder and a random secret key; without a fixed key, updates log everybody out.
# as root
mkdir -p /opt/ai-stack
cd /opt/ai-stack
echo "WEBUI_SECRET_KEY=$(openssl rand -hex 32)" > .env
chmod 600 .env
cat .env
You'll see WEBUI_SECRET_KEY= and 64 characters. Now run nano /opt/, paste this, replace YOUR_DOMAIN with your chat's name (chat.example.com, say) and save with Ctrl+O, Enter, Ctrl+X:
# as root: contents of /opt/ai-stack/compose.yaml
services:
open-webui:
image: ghcr.io/open-webui/open-webui:v0.11.4
container_name: open-webui
restart: unless-stopped
ports:
- "127.0.0.1:3000:8080"
volumes:
- open-webui:/app/backend/data
environment:
- WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY}
- OLLAMA_BASE_URL=http://ollama:11434
- WEBUI_URL=https://YOUR_DOMAIN
- ENABLE_CODE_EXECUTION=false
- ENABLE_CODE_INTERPRETER=false
# step 6: delete the "# " in front of the next two lines
# - CORS_ALLOW_ORIGIN=https://YOUR_DOMAIN
# - WEBUI_SESSION_COOKIE_SECURE=true
depends_on:
- ollama
ollama:
image: ollama/ollama:0.35.1
container_name: ollama
restart: unless-stopped
ports:
- "127.0.0.1:11434:11434"
volumes:
- ollama:/root/.ollama
environment:
- OLLAMA_NO_CLOUD=1
volumes:
open-webui: {}
ollama: {}
The 127.0.0.1 prefixes matter: Docker's docs call published ports "insecure by default", and a plain 3000:8080 would bypass ufw. OLLAMA_NO_CLOUD=1 disables Ollama's cloud models. For API-only on VPS Micro, delete the ollama service and both depends_on lines, and add ENABLE_OLLAMA_API=false. Since 0.11.4 there is also a slim Open WebUI image of about 175 MB, but it drops the local embedding models, PDF and Word reading and the built-in vector search (knowledge search then needs PostgreSQL with pgvector), so this guide uses the standard one. Then download the images (about 16 GB on disk in our test, 9 GB of it Ollama) and start them:
# as root
cd /opt/ai-stack
docker compose up -d
docker compose ps
curl -s http://127.0.0.1:3000/health
Straight after the start, docker compose ps shows open-webui with (health: starting) and curl prints nothing, because Open WebUI is still setting itself up. Wait a minute or two and run the last two lines again: open-webui shows (healthy) and curl prints {"status":true}. Still nothing after five minutes? Check docker compose logs open-webui.
Step 4: create the admin through an SSH tunnel
The first account registered becomes the administrator, and Open WebUI's docs say never to expose an unconfigured instance publicly. Run on your own computer, this forwards your port 3000 to the server's private one:
# as the local user on your own computer (not on the server)
ssh -L 3000:127.0.0.1:3000 root@YOUR_SERVER_IP
Type yes if SSH asks about the server's fingerprint, then log in. Keep that window open, browse to http://localhost:3000 and register. Open WebUI then switches public sign-up off by itself: add colleagues under Admin Panel, Users, or turn New Sign Ups back on under Admin Panel, Settings, General, where new accounts wait as "pending" until you approve them.
Step 5: connect your API keys
Under Admin Panel, Settings, Connections, add an OpenAI-compatible connection per provider; "Verify Connection" catches a mistyped key.
- OpenAI:
https://api.openai.com/ v1 - Anthropic:
https://, a compatibility layer Anthropic describes as mainly for testing, without prompt caching.api.anthropic.com/ v1 - Mistral:
https://api.mistral.ai/ v1 - OpenRouter:
https://. Add a model allowlist, or users scroll through more than 400 models.openrouter.ai/ api/ v1
Step 6: HTTPS with Caddy
Open WebUI has no TLS of its own. Point an A and AAAA record for YOUR_DOMAIN at the server, then install Caddy from its official repository; it fetches the certificate itself. Our reverse proxy and SSL guide covers nginx and Traefik.
# as root
apt install -y debian-keyring debian-archive-keyring apt-transport-https curl
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/gpg.key' | gpg --dearmor -o /usr/share/keyrings/caddy-stable-archive-keyring.gpg
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/debian.deb.txt' | tee /etc/apt/sources.list.d/caddy-stable.list
chmod o+r /usr/share/keyrings/caddy-stable-archive-keyring.gpg
chmod o+r /etc/apt/sources.list.d/caddy-stable.list
apt update
apt install -y caddy
Check with systemctl is-active caddy, which prints active. Replace everything in /etc/ with:
# as root: contents of /etc/caddy/Caddyfile
YOUR_DOMAIN {
reverse_proxy 127.0.0.1:3000
}
Remove the # before the two step 6 lines in compose.yaml. Now reload Caddy, restart the stack and test HTTPS:
# as root
systemctl reload caddy
cd /opt/ai-stack
docker compose up -d
until curl -sf http://127.0.0.1:3000/health; do sleep 5; done; echo
curl -I https://YOUR_DOMAIN
The until line waits while Open WebUI restarts (about half a minute on our test server) and prints {"status":true} when it is back. Then expect HTTP/2 200 and a padlock in the browser. If journalctl -u caddy shows a failed challenge, DNS isn't pointing here yet.
Step 7: pull a small model and measure it
This checks the CPU flags (Ollama runs fastest with AVX2), pulls the task and embedding models, and opens a chat that reports its speed.
# as root
grep -o -w -E 'avx2|avx512f|avx512_vnni' /proc/cpuinfo | sort -u
docker exec ollama ollama pull llama3.2:3b
docker exec ollama ollama pull nomic-embed-text
docker exec -it ollama ollama run llama3.2:3b --verbose
The grep should print avx2, as it does on our servers; if it prints nothing, keep that server to API models. Each pull ends in success. Ask something short and read the eval rate under the answer, in tokens per second, and compare it with our table above; /bye leaves. Run docker exec ollama ollama ps straight after and the model is still loaded, showing 100% CPU and a context of 4096, the CPU-only default from earlier. If Open WebUI lists no Ollama models, check OLLAMA_BASE_URL reads http://ollama:11434, since 127.0.0.1 inside a container is the container itself.
Adding RAG: Qdrant or pgvector, and a minimal LangChain app
RAG (retrieval-augmented generation) answers from your documents. An embedding model turns each chunk of text into a vector, a list of numbers capturing its meaning; a vector database stores them; each question pulls the closest chunks into the prompt. Open WebUI's built-in store is enough until a bot, a helpdesk widget or a self-hosted n8n workflow needs the same index.
Choose pgvector, the PostgreSQL extension that stores and searches vectors, if PostgreSQL already holds your data: one thing to back up, SQL joins included. Debian 13's own pgvector package is affected by CVE-2026-3172 (fixed in 0.8.2), so use the PostgreSQL project's repository, as in our guide to running PostgreSQL on a VPS. Choose Qdrant, an open-source vector database that runs as its own service, for a standalone index. By Qdrant's capacity planning formula, a million 768-dimension vectors need about 2.86 GiB and 100,000 chunks about 0.29 GiB. Qdrant 1.19, released in August 2026, added a turbo4 datatype that stores half a byte per dimension instead of four, so the same million vectors fit in about 0.36 GiB, at some cost in precision. pgvector's HNSW index stops at 2,000 dimensions (4,000 with halfvec), below the 3,072 of OpenAI's text-embedding-3-large.
Most RAG disappointment isn't the database. Hacker News threads on production RAG keep returning to one point: chunking is the hard part.
Step 8: Qdrant with an API key
Self-hosted Qdrant has no authentication by default, so this creates a key and binds the ports to localhost.
# as root
mkdir -p /srv/qdrant/storage
openssl rand -hex 32 > /root/qdrant.key
chmod 600 /root/qdrant.key
docker run -d --name qdrant --restart unless-stopped \
-p 127.0.0.1:6333:6333 -p 127.0.0.1:6334:6334 \
-e QDRANT__SERVICE__API_KEY="$(cat /root/qdrant.key)" \
-e QDRANT__TELEMETRY_DISABLED=true \
-v /srv/qdrant/storage:/qdrant/storage \
qdrant/qdrant:v1.19.2
sleep 5
curl -s -H "api-key: $(cat /root/qdrant.key)" http://127.0.0.1:6333/
echo
cat /root/qdrant.key
Docker shows the download, then a long container ID, which is not the key. The curl returns JSON with "version":"1.19.2", and the last line prints the key for step 9.
Step 9: a minimal LangChain app
LangChain, a Python framework that wires models, documents and vector stores together, is often criticised for layers of abstraction, so this app uses only its smallest packages. Apps don't run as root, so create a user and switch to it:
# as root
adduser --disabled-password --gecos "" ai
su - ai
A usermod: no changes line is harmless. The prompt changes to ai@. Next comes a virtual environment with pinned package versions (Python 3.13 is fine for LangChain):
# as the ai user
mkdir -p ~/rag && cd ~/rag
python3 -m venv .venv
. .venv/bin/activate
pip install langchain-core==1.6.6 langchain-text-splitters==1.1.3 langchain-ollama==1.1.0 langchain-qdrant==1.1.0 qdrant-client==1.19.1
It ends with Successfully installed. If you see externally-managed-environment, run . .venv/bin/activate again. Put a text document at ~/rag/ and save this as ~/rag/:
# as the ai user: contents of ~/rag/app.py
import os
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_ollama import OllamaEmbeddings, ChatOllama
from langchain_qdrant import QdrantVectorStore
OLLAMA = "http://127.0.0.1:11434"
text = open("handbook.txt", encoding="utf-8").read()
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
chunks = splitter.split_documents([Document(page_content=text)])
store = QdrantVectorStore.from_documents(
chunks, OllamaEmbeddings(model="nomic-embed-text", base_url=OLLAMA),
url="http://127.0.0.1:6333", api_key=os.environ["QDRANT_API_KEY"], collection_name="docs",
force_recreate=True)
question = "How many vacation days do we get?"
context = "\n----\n".join(d.page_content for d in store.similarity_search(question, k=4))
llm = ChatOllama(model="llama3.2:3b", base_url=OLLAMA)
print(llm.invoke("Answer only from this context:\n" + context + "\n\nQuestion: " + question).content)
force_recreate=True rebuilds the collection on every run; without it, each run stores every chunk again. Run it, then ask Qdrant what was stored:
# as the ai user
cd ~/rag
. .venv/bin/activate
export QDRANT_API_KEY=YOUR_QDRANT_KEY
python app.py
curl -s -H "api-key: $QDRANT_API_KEY" http://127.0.0.1:6333/collections/docs
The answer can take a minute on a CPU. If you see a warning about an API key over an insecure connection, ignore it: the traffic stays on the server. The curl shows points_count above zero. To answer with a cloud model, swap ChatOllama for ChatOpenAI from langchain-openai.
Is Open WebUI private? What leaves the server
Open WebUI's FAQ is plain: nothing goes out by default, and once you connect a provider, prompts and responses go there. What goes out: the conversation so far with every request, the retrieved chunks and, unless a local model handles them, the title and tag calls. Cloud embeddings send whole documents. Accounts, history, files and vectors stay on your server. Admins can read every chat by default (ENABLE_ADMIN_CHAT_ACCESS): tell your team, or switch it off.
OpenAI says API data hasn't been used for training since 1 March 2023 unless you opt in, and keeps abuse-monitoring logs for up to 30 days. Anthropic says API inputs and outputs are deleted within 30 days, with exceptions such as policy enforcement.
Under the GDPR the provider is your processor, which needs an Article 28 agreement, and processing outside the EEA falls under Chapter V. Amsterdam and Dublin are inside the EU. Kosovo has no EU adequacy decision, so EU personal data on a Prishtina server needs safeguards such as standard contractual clauses; Kosovo's own Law No. 06/L-082 is aligned with the GDPR.
The EU AI Act (Regulation (EU) 2024/1689) has applied since 2 August 2026, although Regulation (EU) 2026/1744, in force since 27 July 2026, pushed most high-risk rules back to December 2027 and August 2028. For an internal assistant, two duties matter: Article 4, which that amendment softened into taking measures that support your staff's AI literacy, and making clear people are dealing with AI unless that is obvious (Article 50(1)). This is orientation, not legal advice.
API cost is counted in tokens, and the rate limit is shared
Cloud models bill input and output tokens separately, and every turn resends the conversation, so message 40 of a thread costs far more than message 2. New topic, new chat. Rate limits apply per organisation or project, not per person, and one Open WebUI usually means one key. OpenAI counts requests and tokens per minute and per day; Anthropic counts requests, input tokens and output tokens per minute for each model. Over the limit, both return HTTP 429. A second key in the same project or workspace shares the same limit, so to stop one busy department starving the rest, give it its own OpenAI project or Anthropic workspace with a lower limit and add that key as a separate connection.
How do you update Open WebUI without losing chats?
Back up before every update: Open WebUI's docs warn that rolling back the container doesn't undo a database migration. Everything lives in the volume ai-stack_open-webui. The backup stops the app so SQLite is consistent, archives the volume and starts the app again:
# as root
mkdir -p /root/backups
cd /opt/ai-stack
docker compose stop open-webui
docker run --rm -v ai-stack_open-webui:/data -v /root/backups:/backup alpine tar czf /backup/openwebui-$(date +%Y%m%d).tar.gz /data
docker compose start open-webui
ls -lh /root/backups
You'll see an archive with today's date, close to 1 GB on our test server: almost all of it is embedding and speech-to-text models Open WebUI keeps in its cache, while chats and settings took under 1 MB. The removing leading '/' line from tar is normal. Copy it off the server with our restic offsite backups guide; our free weekly backups, restorable from the client area, cover the whole server. Ollama models just download again. To update, read the release notes and change the image tags in compose.yaml; this pulls and recreates the containers:
# as root
cd /opt/ai-stack
docker compose pull
docker compose up -d
until curl -sf http://127.0.0.1:3000/health; do sleep 5; done; echo
curl -s http://127.0.0.1:3000/api/version; echo
The last line shows the version now running, such as {"version":"0.11.4",; it should match the tag you set. Move Qdrant one minor version at a time. For monitoring, point an Uptime Kuma check at https:// (see our Uptime Kuma and Grafana guide) and watch memory with docker stats --no-stream. An OOM-killer line in journalctl -k means a model is too big for the plan.
What to lock down before anyone else logs in
These tools ship open. Cisco's 2025 Shodan study found 1,139 exposed Ollama servers, and about a fifth were hosting models anyone could use. Ollama's local API has no authentication at all, which made CVE-2026-7482, "Bleeding Llama", so serious: disclosed by Cyera in May 2026 and rated 9.1, it let anyone who could reach an Ollama older than 0.17.1 read its process memory, prompts and environment variables included, with three API calls.
- Keep every published port on 127.0.0.1, and never set
OLLAMA_HOST=0.0.0.0on a public server. - If everyone who uses the chat can install a VPN app, you can skip the public domain and reach Open WebUI only through a tunnel; our WireGuard guide for Debian 13 sets up the server and the first clients.
- Leave sign-up off and add people yourself, approve "pending" users by hand if you turn it on, or connect SSO or LDAP.
- Tools and Functions run arbitrary Python inside the app. Install only ones you've read.
- Set
RAG_FILE_MAX_SIZE; uploads have no size limit by default. - Read the release notes and security advisories monthly. On 27 and 28 September 2026 Open WebUI published 19 advisories fixed in 0.11.4, four rated high; in one of them (rated 8.1), a malicious website could steal the session of a signed-in user who visited it, on instances with community sharing switched on.
- Use SSH keys and turn off root password login, which our Debian 13 image allows at first boot. On a fresh test server we booted, the first SSH password guess arrived less than six minutes after boot.
Frequently asked questions
Can you run Open WebUI without a GPU?
Yes. Open WebUI is a web app needing about 2 GB of RAM and no graphics card. A GPU only matters for running large models locally through Ollama.
How many tokens per second can a CPU-only server generate?
It depends on model size and memory bandwidth more than vCPU count. On an 8 vCPU / 16 GB server on our platform (the size of VDS Medium), Llama 3.2 3B wrote about 21 tokens per second and Llama 3.1 8B about 10; limited to 4 vCPU, about 13 and 6.5. Published CPU results for Llama 3.2 3B run from about 9 tokens per second on a mini PC to about 24 on a laptop. Measure yours with ollama run MODEL --verbose.
Can I use Open WebUI with my own OpenAI, Claude, Mistral or OpenRouter API key?
Yes. Add each as an OpenAI-compatible connection using the base URLs in step 5. Usage is billed per token to your own key.
Why does my RAG ignore most of my documents with Ollama?
Usually because Ollama's default context on a CPU-only server is 4,096 tokens, so part of the prompt is cut. Raise OLLAMA_CONTEXT_LENGTH as far as RAM allows or retrieve fewer chunks.
Is Open WebUI free for commercial use or for a team?
There's no licence fee, but since version 0.6.6 its licence requires keeping the Open WebUI branding unless you have 50 or fewer users in 30 days or an enterprise licence.
Can RS Computers set up Open WebUI for our team?
Yes. Message us on Telegram or email info@rscomputers-ks.com with your team size and API providers, and we'll suggest a plan and quote the setup.
Ten questions before you invite the team
Order VDS Small with Debian 13 in the location closest to your team. Run steps 1 to 7, connect one API key and set llama3.2:3b as task model. Upload one document your team really asks about and try ten questions you already know the answers to. If the answers hold up, invite the team; if they don't, fix chunking and context before blaming the model. Compare the plans and order on our VPS and VDS plans page.