Ceph block storage is RBD, the RADOS Block Device: a virtual disk that Ceph cuts into 4 MiB objects and spreads over many disks and servers, so a VM, a Kubernetes pod or a plain Linux host sees one ordinary drive while every write is copied to two or three places. That copying is what makes Ceph block storage performance different from a local NVMe disk. In our lab, an RBD disk delivered about one seventh of the random read IOPS of a plain virtual disk in the same VM, and a single small write took 2 ms instead of 0.26 ms. Below: the fio commands, the numbers, CephFS vs RBD on the same cluster, and two traps that crashed our test cluster.
Key facts, checked on 11 October 2026:
- RBD stores an image as objects of 4 MiB by default; the smallest allowed object size is 4K and the largest 32M (rbd man page).
- Ceph keeps three copies of each object by default, and its docs warn that "setting size to 2 or min_size to 1 in production risks data loss" (Ceph pools documentation).
- Ceph asks for at least 10 Gb/s between hosts and clients. Re-copying 1 TiB takes about three hours over 1 Gb/s and about twenty minutes over 10 Gb/s (hardware recommendations).
- The current release is Tentacle 20.2.4 from 19 August 2026, supported to about 1 June 2027 (Ceph releases). It added a new key type, aes256k, that the Linux kernel client reads only from kernel 7.0 onward (20.2.4 release notes).
- Ceph-CSI, the Kubernetes driver, is at v3.18.1 (1 October 2026) and serves both RBD and CephFS volumes (ceph-csi on GitHub).
For the architecture, see Ceph block and object storage; for a real three-node cluster that loses a node, see our three-VPS Ceph lab. This post stays on the block device and its speed.
What is Ceph RBD, and what is the "block size"?
An RBD image is a disk-shaped file that lives inside a Ceph pool. It is thin-provisioned, which means it only uses space for blocks that have been written. In our lab we created a 10 GiB image and rbd du reported 10 GiB provisioned and 0 B used. A 4 GiB image that held an ext4 filesystem with a 2 GiB test file showed 2.1 GiB used.
When people search for "ceph block size" they usually mean the object size. RBD stores each slice of the image as one RADOS object, the unit Ceph replicates and moves. rbd info prints it as an "order":
# as root inside the lab VM (the image is created in the lab section below)
./cephadm shell -- rbd info rbd/vm-disk
Our output said size 4 GiB in 1024 objects and order 22 (4 MiB objects), because 2 to the power of 22 bytes is 4 MiB. You can pick another size per image when you create it; rbd create --size 1G --object-size 1M rbd/small-objects gave us order 20 (1 MiB objects). It cannot be changed later.
The object size is not the sector size your filesystem sees. After mapping the image, the kernel reported a 4 MiB "optimal I/O size" (the object size) and a 64 KiB "minimum I/O size", which matches the krbd alloc_size default of 64K described in the man page. Your filesystem still writes in 4 KiB blocks. Leave it at 4 MiB unless you have measured a reason.
Three ways a machine uses an RBD image
| Client | How it works | Typical user | Watch out for |
|---|---|---|---|
| krbd (kernel module) | rbd map creates /dev/rbd0, which you format and mount like any disk | A plain Linux server, Kubernetes nodes through Ceph-CSI | Kernel version decides which image features and key types work |
| librbd (user space) | QEMU/KVM opens the image directly through the library; no device node on the host | OpenStack, Proxmox VE and libvirt VMs | Its client cache can make benchmarks look far better than the cluster is |
| Ceph-CSI | A Kubernetes driver that creates an RBD image per PersistentVolumeClaim and maps it on the node | Kubernetes, usually deployed by Rook | RBD volumes are ReadWriteOnce in file mode; shared folders need CephFS |
Ceph-CSI also offers block-mode RBD volumes as ReadWriteMany, meant for software that coordinates its own writes, such as VM live migration.
CephFS vs RBD vs RGW: start from who needs the data
All three sit on the same cluster. The question that decides between them is how many machines touch the data and how they talk to it.
| Your situation | Use | Why |
|---|---|---|
| One VM or one database needs a fast disk | RBD | Lowest latency of the three, works with any filesystem, snapshots and clones per image |
| Ten web servers must read and write the same folder | CephFS | A real shared POSIX filesystem; RBD with ext4 would corrupt if two hosts mounted it |
| An app uploads files through an API (backups, media, logs) | RGW (S3) | No filesystem to size or repair; see our Ceph object storage guide |
| Kubernetes pods, one per volume | RBD through Ceph-CSI | ReadWriteOnce volumes with the block device's speed |
| Kubernetes pods sharing a volume | CephFS through Ceph-CSI | ReadWriteMany volumes |
| Millions of tiny files, created and deleted all day | RBD, or rethink the design | CephFS asks its metadata server (MDS) for every create and delete; our test below shows the cost |
Our lab: one VM, three OSDs, Ceph 20.2.4
To compare RBD, CephFS and a plain disk on identical hardware, everything ran inside one Incus virtual machine on our Amsterdam test server: Debian 13 with kernel 6.12, 2 vCPU, 4 GiB of RAM, a 30 GiB system disk and four extra 8 GiB virtual disks. Three became OSDs (the daemons that store data, one per disk), and the fourth stayed a plain ext4 disk as the baseline.
# as root on the Incus host (lab setup only)
incus launch images:debian/13 wf-rbd --vm -c limits.memory=4GiB -c limits.cpu=2 -d root,size=30GiB
incus storage volume create default wf-rbd-osd1 --type=block size=8GiB
incus storage volume attach default wf-rbd-osd1 wf-rbd
We repeated the last two lines for wf-rbd-osd2, wf-rbd-osd3 and wf-rbd-base. Inside the VM they appear as /dev/sdb to /dev/sde. Then the packages, with Debian's own Ceph client tools (version 18.2.7 Reef, because Ceph publishes no Tentacle packages for Debian 13 yet) and cephadm downloaded from Ceph:
# as root inside the VM
apt-get install -y podman lvm2 chrony curl fio jq ceph-common
apt-get install -y openssh-server
curl -s --remote-name --location https://download.ceph.com/rpm-tentacle/el9/noarch/cephadm
chmod +x cephadm
./cephadm bootstrap --mon-ip 10.141.188.173 --single-host-defaults --skip-monitoring-stack --skip-dashboard
Use your server's own IP address after --mon-ip. Our first bootstrap failed after two minutes with Connect call failed ('10.141.188.173', 22): cephadm manages even its own host over SSH, and the minimal Debian image had no SSH server. With openssh-server installed, bootstrap finished in 33 seconds and printed Bootstrap complete. The flag --single-host-defaults lets Ceph place copies on different disks of one machine instead of different machines, and it sets the default pool size to 2. We checked that with ceph config dump.
The key-type trap with older clients
A fresh 20.2.4 cluster allows only the new aes256k keys. Debian's 18.2.7 client tools cannot even read them (error parsing file /etc/ceph/ceph.client.admin.keyring ... Malformed input), so we ran every admin command through ./cephadm shell --, which starts a Tentacle container. The release notes say the kernel client reads them only from Linux 7.0, and our VM runs 6.12. For the lab, we allowed the older key type for one client that may only touch the rbd pool:
# as root inside the VM
./cephadm shell -- ceph config set osd osd_memory_target 1073741824
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdb
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdc
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdd
./cephadm shell -- ceph osd pool create rbd
./cephadm shell -- rbd pool init rbd
./cephadm shell -- rbd create --size 4G rbd/vm-disk
./cephadm shell -- ceph mon set auth_allowed_ciphers aes,aes256k
./cephadm shell -- ceph auth get-or-create client.vmdisk mon 'profile rbd' osd 'profile rbd pool=rbd'
./cephadm shell -- ceph auth rotate --key-type=aes client.vmdisk > /etc/ceph/ceph.client.vmdisk.keyring
Each OSD line answered Created osd(s) N on host 'wf-rbd'. The last two commands lower security, and Ceph said so at once with three health warnings, AUTH_INSECURE_CLIENT_KEY_TYPE, AUTH_INSECURE_KEYS_ALLOWED and AUTH_INSECURE_KEYS_CREATABLE, because the old type is the one the August security release replaces. That is fine for a throwaway lab. In production, use clients with kernel 7.0 or a vendor kernel with the backport, as the CephX key rotation guide describes.
Now map the image with the kernel client, format it and mount it:
# as root inside the VM
rbd --id vmdisk map rbd/vm-disk
mkfs.ext4 -q /dev/rbd0
mkdir -p /mnt/rbd
mount /dev/rbd0 /mnt/rbd
rbd map printed /dev/rbd0, and dmesg showed rbd: rbd0: capacity 4294967296 features 0x3d. The baseline disk got the same mkfs.ext4 and was mounted at /mnt/base.
Adding a small CephFS on the same cluster
One command creates the filesystem, its two pools and the metadata servers (MDS). Our 4 GiB VM ran out of memory during this step, so we added a swap file and kept a single MDS.
# as root inside the VM
fallocate -l 4G /swapfile
chmod 600 /swapfile
mkswap -q /swapfile
swapon /swapfile
./cephadm shell -- ceph fs volume create labfs
./cephadm shell -- ceph orch apply mds labfs --placement=1
./cephadm shell -- ceph fs authorize labfs client.fsuser / rw
./cephadm shell -- ceph auth rotate --key-type=aes client.fsuser > /etc/ceph/ceph.client.fsuser.keyring
mkdir -p /mnt/cephfs
mount -t ceph fsuser@.labfs=/ /mnt/cephfs
ceph -s then listed volumes: 1/1 healthy and two new pools, cephfs.labfs.meta with 16 PGs and cephfs.labfs.data with 64, both with 2 copies like the RBD pool. df -h /mnt/cephfs showed 5.3G free, because every byte stored costs two bytes of raw disk. Idle after the tests, each OSD used about 395 MiB, the manager 170 MiB and the monitor 37 MiB (ceph orch ps).
Ceph block storage performance: real fio numbers
We ran five fio tests on each target, each 45 seconds after a 5-second warm-up, on a 2 GiB file with direct I/O so the page cache stays out of it. These are the exact command lines for the RBD mount; for the other targets only the path changes. We also added --output-format=json to read the results with jq.
# as root inside the VM
fio --name=randread4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randread --bs=4k --iodepth=32
fio --name=randwrite4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=32
fio --name=qd1write4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1
fio --name=seqread1m --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=read --bs=1M --iodepth=8
fio --name=seqwrite1m --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=write --bs=1M --iodepth=8
The third test, one 4 KiB write at a time, is closest to how a database commit feels.
Read this table with care: one VM with 2 vCPU on a shared host, all OSDs on one machine, and no real network between client and OSDs. It shows the relative overhead of the Ceph layers, not what a production cluster delivers. Both Ceph targets used pools with 2 copies.
| fio test | Plain virtual disk (ext4) | RBD, kernel client (ext4) | CephFS, kernel mount |
|---|---|---|---|
| 4K random read, QD32 | 49,672 IOPS | 6,867 IOPS | 7,066 IOPS |
| 4K random write, QD32 | 31,942 IOPS | 1,934 IOPS | 2,084 IOPS |
| 4K write, QD1 (avg latency) | 3,271 IOPS (0.26 ms) | 491 IOPS (1.99 ms) | 525 IOPS (1.86 ms) |
| 1M sequential read, QD8 | 3,894 MiB/s | 930 MiB/s | 983 MiB/s |
| 1M sequential write, QD8 | 3,220 MiB/s | 414 MiB/s | 377 MiB/s |
| 5,000 small files: create + sync | 0.31 s | 0.50 s | 3.85 s |
| Same files: list + delete + sync | 0.11 s | 0.13 s | 2.75 s |
What the numbers say:
- RBD kept 14% of the plain disk's random read IOPS and 6% of its random write IOPS. Two vCPU had to run the client, three OSDs, a monitor and a manager at once, so the OSDs ran short of CPU long before the virtual disks did.
- A single small write took 1.99 ms on RBD against 0.26 ms locally, about 7.6 times longer, because it has to reach the primary OSD and be confirmed by the second copy before fio hears back.
- For one big file written by one client, CephFS and RBD were within 10% of each other. CephFS file data goes straight to the OSDs, just like RBD.
- The gap opens with metadata. Creating 5,000 small files took 7.7 times longer on CephFS than on ext4 over RBD, and listing and deleting them took 21 times longer, because each step is a request to the MDS.
The small-file test is a plain shell function that we ran once per mount point (it needs the bc package):
# as root inside the VM
smallfiles() { d=$1/small; mkdir -p $d; s=$(date +%s.%N); for i in $(seq 1 5000); do echo hello > $d/f$i; done; sync; m=$(date +%s.%N); ls -l $d | wc -l >/dev/null; rm -rf $d; sync; e=$(date +%s.%N); echo "$1 create5000+sync=$(echo "$m - $s" | bc)s list+delete=$(echo "$e - $m" | bc)s"; }
smallfiles /mnt/base; smallfiles /mnt/rbd; smallfiles /mnt/cephfs
Repeat your runs: our first RBD run, minutes after the autoscaler had split the pool into 32 placement groups, was 24% slower on random reads and 41% slower on random writes than the run in the table.
krbd or librbd: the cache that flatters benchmarks
fio can also talk to Ceph through librbd, the same library QEMU uses, with --ioengine=rbd. We wrote a second 2 GiB image full, then tested it raw (no filesystem) three ways: through the kernel device, through librbd with default settings, and through librbd with its cache switched off.
# as root inside the VM
fio --name=qd1write4k --ioengine=rbd --clientname=vmdisk --pool=rbd --rbdname=fio-img --size=2G --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1
CEPH_ARGS="--rbd_cache=false" fio --name=qd1write4k --ioengine=rbd --clientname=vmdisk --pool=rbd --rbdname=fio-img --size=2G --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1
The other rows use the same --rw, --bs and --iodepth values as the first set of tests. For the krbd column we mapped the same image (it became /dev/rbd1) and used --filename=/dev/rbd1 --direct=1 --ioengine=libaio instead of the three rbd options.
| Raw image, 2 copies | krbd /dev/rbd1 | librbd, cache off | librbd, default cache |
|---|---|---|---|
| 4K random read, QD32 | 6,804 IOPS | 5,948 IOPS | 6,498 IOPS |
| 4K random write, QD32 | 1,848 IOPS | 1,607 IOPS | 3,266 IOPS |
| 4K write, QD1 (avg latency) | 509 IOPS (1.93 ms) | 403 IOPS (2.44 ms) | 3,528 IOPS (0.26 ms) |
| 1M sequential write, QD8 | 386 MiB/s | 359 MiB/s | 420 MiB/s |
With the cache off, the kernel client was slightly faster than librbd on this small VM. With librbd's default cache, single writes looked almost nine times faster, because librbd's client cache (32 MiB per image by default, see the RBD config reference) absorbed them in memory. A database that waits for fsync on every commit would not see that speed. Benchmark with the cache off, or with a workload that flushes, or you are measuring RAM.
The built-in rbd bench uses librbd with the cache on. On the same image it reported 3,404 random 4K writes per second with 16 threads, 569 MiB/s for 4 MiB writes and 1.4 GiB/s for 4 MiB reads:
# as root inside the VM
rbd --id vmdisk bench --io-type write --io-size 4K --io-threads 16 --io-total 256M --io-pattern rand rbd/fio-img
rbd --id vmdisk bench --io-type write --io-size 4M --io-threads 16 --io-total 2G rbd/fio-img
rbd --id vmdisk bench --io-type read --io-size 4M --io-threads 16 --io-total 2G rbd/fio-img
Fine as a smoke test for a new pool; plan with fio.
Replication size 3 vs 2: space and write latency
The pool "size" is how many copies Ceph keeps; min_size is how many must be available before Ceph accepts writes. We switched the RBD pool from 2 to 3 copies and ran the raw krbd tests again:
# as root inside the VM
./cephadm shell -- ceph osd pool set rbd size 3
./cephadm shell -- ceph osd pool set rbd min_size 2
| Raw krbd image | size 2 | size 3 | Change |
|---|---|---|---|
| 4K random read, QD32 | 6,804 IOPS | 7,131 IOPS | No real change (reads come from one copy) |
| 4K random write, QD32 | 1,848 IOPS | 1,494 IOPS | 19% fewer |
| 4K write, QD1 latency | 1.93 ms | 2.42 ms | 25% slower |
| 1M sequential write | 386 MiB/s | 263 MiB/s | 32% less |
| Raw space for 4.0 GiB stored | 8 GiB (calculated) | 12 GiB (from ceph df) | Usable space falls from 1/2 to 1/3 of raw |
Size 2 is faster and cheaper, and the Ceph docs still advise against it in production: with two copies, one failed disk during the rebuild of another loses data. Keep size 3 and min_size 2 for anything you care about, and buy the disks for it. If space hurts, erasure coding is the tool, which our architecture article covers.
Switching to size 3 also gave us our first crash. Ceph started copying 4 GiB to the new third copies, one OSD grew past 1.3 GB of memory, and the kernel's out-of-memory killer stopped it inside our 4 GiB VM, which had no swap.
What limits Ceph performance, in the order you will meet it
- Network latency. Every write travels client to primary OSD, then primary to the replicas, so single writes on Ceph are always slower than on a local NVMe disk.
- Bandwidth. A 1 Gb/s link carries about 120 MB/s, less than one SATA SSD, and with 3 copies each client write becomes two more writes between OSD hosts. Rebuilds use the same link, which is why Ceph names 10 Gb/s as the minimum and suggests 25 Gb/s for busier clusters.
- CPU per OSD. Our lab ran out of CPU long before its disks were busy. On flash, each OSD needs several CPU threads.
- Drives that confirm flushes quickly. Ceph waits for every write to be safe on disk; enterprise NVMe with power-loss protection does that fast, many consumer SSDs do not.
- Replication size: 19% of random write IOPS in our test.
- Placement group (PG) count. A pool is split into PGs, and each PG lives on a few OSDs. Our new pool started with 1 PG, so every request hit the same two OSDs until the autoscaler raised it to 32. The placement group docs set a default target of 100 PGs per OSD and recommend 200 for all but the smallest clusters.
Safe tuning for beginners
What we would do on a first cluster:
- Leave the PG autoscaler on, and let it finish before you benchmark.
ceph -sshould show every PG asactive+clean. - Keep
osd_memory_targetat its 4 GiB default on real servers. The hardware guide says not to go below 2 GB. Our 1 GiB lab setting worked until recovery started, then caused the crash described above. - Never map RBD images or mount CephFS with the kernel client on a host that also runs OSDs, except in a lab. We learned this the hard way: after the out-of-memory kill, the OSD restart scanned
/dev/rbd0, a device served by those same stopped OSDs, and hung until we reset the VM. The ceph-users list has the background on this deadlock risk. - Give VMs librbd with the default cache settings;
rbd_cache_writethrough_until_flushkeeps the cache safe for guests that never send flushes. Just do not benchmark through it. - Keep the cluster network separate from client traffic once you grow past a lab, and give it 10 Gb/s or more.
To measure a plain server disk first, see our VPS benchmark commands.
Frequently asked questions
What is Ceph RBD?
RBD, the RADOS Block Device, is Ceph's block storage. It presents a virtual disk that is thin-provisioned and split into objects, 4 MiB each by default, stored with two or three copies across the cluster. Linux maps it as /dev/rbdX, while QEMU/KVM uses it through the librbd library.
Is CephFS slower than RBD?
For large files read or written by one client, they are close: in our lab, CephFS and RBD were within 10% on 4K random I/O and 1M sequential tests. CephFS is much slower with many small files, because every create and delete goes through the metadata server. Creating 5,000 files took 3.85 s on CephFS and 0.50 s on ext4 over RBD.
What is the default block size in Ceph?
RBD images use a 4 MiB object size by default, shown as "order 22" in rbd info. It can be set per image from 4K to 32M with --object-size when the image is created. The filesystem on top still uses its own block size, usually 4 KiB.
Can two servers mount the same RBD image?
Not with ext4 or XFS; two hosts writing to it will corrupt the filesystem. Use CephFS when several machines need the same files.
How much space do I lose with replication 3?
Two thirds of the raw capacity. With size 3, every terabyte you store uses three terabytes of disk; in our lab, 4.0 GiB of data used 12 GiB. Size 2 uses half, but the Ceph docs warn that it risks data loss in production.
Why does rbd map fail on Ceph 20.2.4?
Ceph 20.2.4 creates keys of the new aes256k type, and kernel clients older than Linux 7.0 cannot use them unless the vendor has backported support. Use a newer kernel, or, in a test cluster only, allow the older aes type and rotate one client key to it, as shown in this article.
Try it on servers you control
One VM teaches the commands; three servers show real replication and recovery. RS Computers runs KVM VPS and VDS in Amsterdam, Dublin and Prishtina, all on NVMe with unmetered traffic and their own IPv4 and IPv6, and the VDS plans come with 10 Gb/s ports, which is the speed Ceph asks for between nodes. Three VDS Small or Medium servers in one city make a realistic test cluster; follow our three-node lab guide to build it. vCPUs on our plans are shared, so use them for labs and proofs of concept, and put production OSDs on bare metal. Availability by city is on the plans page, and if you want help sizing a lab, message us on Telegram.