← Back to Blog

Why Is My VPS Slow? Diagnose CPU Steal, RAM, Disk and Network Bottlenecks

Published · by RS Computers

Linux Troubleshooting VPS

A slow VPS can be short of CPU, memory or storage performance, but it can also be waiting on DNS, a database lock or another website's API. Capture what the server is doing during the slowdown before restarting it. The useful question is which resource or dependency the application is waiting for.

This guide is for diagnosing a Linux VPS while it is running. Connect using the SSH service already installed on your RS Computers Linux VPS. It complements our benchmark guide: a benchmark tests capacity, while the commands below help explain a real incident. We checked the commands on 11 October 2026 using a Debian 13 lab container and read-only samples from its KVM host.

Why is my VPS slow? Match the symptom first

SymptomFirst checkPossible next step
Everything, including SSH, feels slowCPU, memory pressure, disk waits and the network pathCapture samples before changing anything
One application is slowIts logs, workers, queries and upstream servicesTrace one slow request
It slows down at the same time dailyBackups, scheduled jobs and traffic peaksReschedule or limit overlapping work
It is fast on the server but slow from your officeDNS, routing, loss, proxy and TLS timingCompare another client network

Collect a short baseline

As a normal user, start with:

uptime
nproc
free -h
df -h /
df -i /
vmstat -y 1 3

df -h shows space; df -i shows inode usage. A filesystem can run out of inodes even with gigabytes free. The -y option in vmstat omits the initial average since boot, so the output represents fresh one-second samples. For a longer incident capture, increase the final count to 30.

Our small container reported 2 CPUs available and 2.0 GiB of RAM, with about 1.8 GiB available immediately after setup. The underlying host had roughly 15 GiB available and about 63 MiB of swap in use in an earlier sample. Its later one-second samples showed no current swap-in or swap-out. That is a useful example of why swap occupied and active swapping are different observations.

See the vmstat manual for column definitions and free's documentation for memory accounting.

CPU: distinguish your work from time waiting for a vCPU

In vmstat, us and sy describe time spent on user and kernel work; id is idle and st is steal time. Steal means time a virtual CPU could not run because the hypervisor was servicing something else. Sustained steal alongside slow requests is worth reporting to the provider, but a universal percentage threshold cannot diagnose every workload.

Find candidate processes with:

ps -eo pid,comm,%cpu,%mem --sort=-%cpu | head

The percentages from ps are not the same as a live one-second profiler sample. Use them to locate suspects, then inspect the application. A process with one busy thread can saturate its useful CPU capacity while the machine still shows spare cores. More vCPUs will not automatically make that thread run faster.

Load average is also not a CPU percentage. It includes runnable tasks and tasks in uninterruptible sleep. A high load with storage waits needs a different fix from a high load caused by CPU work. Our host samples recorded st=0; they do not establish how that server behaves all day.

Follow a suspect process over time

For current process samples, use pidstat from the sysstat package. If it is missing, install sysstat as root with apt update followed by apt install -y sysstat. Replace 1234 with the process ID you found above and run the next command as the process owner, or as root when inspecting another user's service:

pidstat -u -r -d -p 1234 1 3

This takes three samples with CPU, memory and I/O fields. In our isolated lab we deliberately ran a single Python calculation for six seconds. The three sampled CPU readings were 99%, 100% and 100%, averaging 99.67%. Resident memory was 8,368 KiB, and the sampled disk read/write rates were zero. That is an example of a process busy with computation, not a storage-throughput benchmark.

Without normalizing by processor count, about 100% CPU in this output means approximately one logical CPU's worth of work. It does not mean that every vCPU in a larger VPS is saturated. A parallel process can report more than 100%. The pidstat reference describes the reporting options.

Check service and container resource limits as well. A worker with a CPU quota or a restrictive CPU set can be constrained while other parts of the VPS remain idle. In cgroup v2, the relevant group's cpu.max and cpu.stat help distinguish a configured quota from host steal time. Read them in the correct service group; the root cgroup may not carry that service's limit. The kernel's cgroup documentation explains those controls.

Memory: look at available RAM and pressure

Linux uses otherwise idle memory for caches. A small free number is not enough to conclude that RAM is exhausted. Read the available estimate, current swap activity and application errors together. A database cache can make memory usage look high while improving response times.

On kernels exposing pressure stall information, run:

cat /proc/pressure/cpu
cat /proc/pressure/memory
cat /proc/pressure/io

The some line records time when at least some tasks were stalled on the resource. The averages cover 10, 60 and 300 seconds. Memory pressure rising at the same time as slow requests is stronger evidence than one memory percentage. Read the Linux PSI documentation before comparing hosts and containers: a container's view may not describe only your application, and cgroup-level counters may be more suitable.

Check the system journal for memory kills and application restarts if you have administrative access. Reduce avoidable worker counts or fix a leak before treating extra RAM as a permanent cure. Adding swap can provide a buffer in some configurations, but storage-backed paging is not equivalent to RAM.

Look for a killed process or failed service

As root, inspect recent kernel messages and failed services:

journalctl -k -b --since '-1 hour' \
  --grep='Out of memory|Killed process|oom-kill' --no-pager
systemctl --failed --no-pager

We checked both commands in the fresh lab: they reported no matching kernel messages and no failed units. On an affected server, look for the process name, time and any memory-limit context. No matching log is not proof that no kill occurred: a container may lack host logs, the event may predate this boot, or the journal may have rotated.

A database that keeps disappearing deserves a restart-history and memory investigation before an upgrade. If it is being killed inside a container's memory limit, adding RAM to the VPS will not necessarily increase that limit. Conversely, a continuously growing application process may have a leak that a larger plan only postpones.

Storage: inspect latency without adding a benchmark workload

Install sysstat as root if iostat is unavailable:

apt update
apt install -y sysstat

Then run as a normal user:

iostat -xz -y 1 2

Review read and write wait times, queue size and throughput while the problem occurs. The wait figures include queuing and service time. Compare with your own healthy baseline. A device's utilization percentage is not a universal saturation meter for modern parallel storage. The iostat manual explains the fields and their limits.

Our container printed CPU rows but no disk-device rows. That was a visibility limitation, not proof of an idle or infinitely fast disk. Run this check in the VPS operating system when container isolation hides block-device statistics. We did not run a destructive storage test or infer disk performance from the empty table.

If waits line up with a backup or large import, move that job away from peak traffic or limit its concurrency. Avoid running a heavy fio test on an already struggling production database: it changes the conditions you are trying to observe.

Do not confuse a full filesystem with a slow device

A full filesystem can stop sessions, logging, uploads or database writes even when the storage hardware is fast. Check the filesystem containing the application as well as / if your data lives on another mount. A growing log, an abandoned SQL export or a backup archive inside the site directory can explain the change.

Identify what owns the files before removing anything. Deleting a database file manually can destroy the service; deleting a log that a process still holds open may not immediately release its disk space. Use the application's retention or log-rotation procedure, then check the same filesystem again.

Network and application: time the request

From a normal shell, time an endpoint you control:

curl -4 -sS -o /dev/null \
  -w 'dns=%{time_namelookup} tcp=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' \
  https://rscomputers-ks.com/

The timestamps are cumulative from request start. A long gap after connection setup but before the first byte can include application work, an upstream wait and network delay. It is a clue, not a diagnosis. Compare the same request from the server and a client machine, check the HTTP result and consult application logs.

A CDN can serve a fast cached homepage while an uncached account page remains slow. Test the actual failing path with appropriate access, keeping credentials and customer data out of shared diagnostics. A timing result from a CDN edge does not establish how quickly your origin server generates uncached pages.

ObservationInvestigate next
DNS completion is slow on one client networkThe client's resolver and DNS path; compare the same name from another network
Connection setup is slow, but the server is idlePacket loss, routing, listener queues and firewall behavior
Connection setup is quick, then the first byte is lateApplication queues, database work, session locks and external API calls
The first byte is quick, but a large transfer is slowPayload size, transfer path, compression and bandwidth to that destination

Also inspect the response itself. A quickly returned error is not a healthy transaction, and a slow redirect chain may be an application configuration problem. Change one variable at a time: the client network, endpoint or application action. Otherwise you cannot tell which difference explains the result.

Decide whether to tune the app or change the plan

Evidence during the problemAction to test
A query or external API dominates request timeImprove that query, cache suitable results, or reduce sequential calls before buying more CPU
Many useful workers keep the available CPUs busyCheck parallel scaling and consider more vCPUs if the application can use them
Memory pressure and active paging coincide with slow requestsReview worker counts, caches and leaks; add RAM when the required working set exceeds the allocation
A backup overlaps with the busiest hourChange the schedule or concurrency, then compare the next busy period
Repeated host steal occurs during otherwise unexplained slowdownsSend timestamped samples to the provider for investigation

Make one change and measure again

Record the incident time, time zone, application symptom, affected endpoint and command output. Change one likely cause, then repeat the same request. If the problem disappears after a restart, keep the evidence: restarting can erase the queue or memory state without fixing the cause.

A useful support report can be short: "At 14:10 Europe/Warsaw, the product-search page took four seconds instead of its usual half second. SSH was responsive. The attached 30-second vmstat capture was taken during the delay. The daily import started at 14:00." These numbers are an illustrative report format, not a result from our lab. That context is much more actionable than "the VPS feels slow".

For RS Computers VPS and VDS customers in Kosovo, the Netherlands and Ireland, send a concise report through support. Include repeated samples instead of a single cropped screenshot. Redact tokens, passwords and personal data. When the evidence points to resource limits, the server sizing guide helps identify what an upgrade would actually change.

Frequently asked questions

Does high RAM usage mean my VPS needs an upgrade?

Not by itself. Check available memory, active swapping, pressure and application behavior. File caching can account for memory usage on a healthy Linux server.

Does zero steal time prove the host is healthy?

No. It means that the sampled CPU counters did not report steal at that time. Storage, networking and application dependencies can still be slow.

Should I run a benchmark during an incident?

Start with passive measurements. A benchmark adds load and can make the incident worse or obscure the original bottleneck. Schedule capacity testing separately.

Why is a website slow while SSH is fast?

The website can be waiting on its own workers, database or upstream API while the operating system has spare resources. Time the failing request and inspect that application's logs rather than assuming the whole VPS is overloaded.

Can a CPU limit look like a slow server?

Yes. A service or container can exhaust its assigned CPU quota while the VPS as a whole has idle capacity. Inspect the application's resource limits separately from hypervisor steal time.

← All articles

Chat on Telegram