← Back to Blog

Ceph Block Storage (RBD): How It Works, CephFS vs RBD, and Real fio Numbers

Published · by RS Computers

Ceph Storage Benchmarks

Ceph block storage is RBD, the RADOS Block Device: a virtual disk that Ceph cuts into 4 MiB objects and spreads over many disks and servers, so a VM, a Kubernetes pod or a plain Linux host sees one ordinary drive while every write is copied to two or three places. That copying is what makes Ceph block storage performance different from a local NVMe disk. In our lab, an RBD disk delivered about one seventh of the random read IOPS of a plain virtual disk in the same VM, and a single small write took 2 ms instead of 0.26 ms. Below: the fio commands, the numbers, CephFS vs RBD on the same cluster, and two traps that crashed our test cluster.

Key facts, checked on 11 October 2026:

For the architecture, see Ceph block and object storage; for a real three-node cluster that loses a node, see our three-VPS Ceph lab. This post stays on the block device and its speed.

What is Ceph RBD, and what is the "block size"?

An RBD image is a disk-shaped file that lives inside a Ceph pool. It is thin-provisioned, which means it only uses space for blocks that have been written. In our lab we created a 10 GiB image and rbd du reported 10 GiB provisioned and 0 B used. A 4 GiB image that held an ext4 filesystem with a 2 GiB test file showed 2.1 GiB used.

When people search for "ceph block size" they usually mean the object size. RBD stores each slice of the image as one RADOS object, the unit Ceph replicates and moves. rbd info prints it as an "order":

# as root inside the lab VM (the image is created in the lab section below)
./cephadm shell -- rbd info rbd/vm-disk

Our output said size 4 GiB in 1024 objects and order 22 (4 MiB objects), because 2 to the power of 22 bytes is 4 MiB. You can pick another size per image when you create it; rbd create --size 1G --object-size 1M rbd/small-objects gave us order 20 (1 MiB objects). It cannot be changed later.

The object size is not the sector size your filesystem sees. After mapping the image, the kernel reported a 4 MiB "optimal I/O size" (the object size) and a 64 KiB "minimum I/O size", which matches the krbd alloc_size default of 64K described in the man page. Your filesystem still writes in 4 KiB blocks. Leave it at 4 MiB unless you have measured a reason.

Three ways a machine uses an RBD image

ClientHow it worksTypical userWatch out for
krbd (kernel module)rbd map creates /dev/rbd0, which you format and mount like any diskA plain Linux server, Kubernetes nodes through Ceph-CSIKernel version decides which image features and key types work
librbd (user space)QEMU/KVM opens the image directly through the library; no device node on the hostOpenStack, Proxmox VE and libvirt VMsIts client cache can make benchmarks look far better than the cluster is
Ceph-CSIA Kubernetes driver that creates an RBD image per PersistentVolumeClaim and maps it on the nodeKubernetes, usually deployed by RookRBD volumes are ReadWriteOnce in file mode; shared folders need CephFS

Ceph-CSI also offers block-mode RBD volumes as ReadWriteMany, meant for software that coordinates its own writes, such as VM live migration.

CephFS vs RBD vs RGW: start from who needs the data

All three sit on the same cluster. The question that decides between them is how many machines touch the data and how they talk to it.

Your situationUseWhy
One VM or one database needs a fast diskRBDLowest latency of the three, works with any filesystem, snapshots and clones per image
Ten web servers must read and write the same folderCephFSA real shared POSIX filesystem; RBD with ext4 would corrupt if two hosts mounted it
An app uploads files through an API (backups, media, logs)RGW (S3)No filesystem to size or repair; see our Ceph object storage guide
Kubernetes pods, one per volumeRBD through Ceph-CSIReadWriteOnce volumes with the block device's speed
Kubernetes pods sharing a volumeCephFS through Ceph-CSIReadWriteMany volumes
Millions of tiny files, created and deleted all dayRBD, or rethink the designCephFS asks its metadata server (MDS) for every create and delete; our test below shows the cost

Our lab: one VM, three OSDs, Ceph 20.2.4

To compare RBD, CephFS and a plain disk on identical hardware, everything ran inside one Incus virtual machine on our Amsterdam test server: Debian 13 with kernel 6.12, 2 vCPU, 4 GiB of RAM, a 30 GiB system disk and four extra 8 GiB virtual disks. Three became OSDs (the daemons that store data, one per disk), and the fourth stayed a plain ext4 disk as the baseline.

# as root on the Incus host (lab setup only)
incus launch images:debian/13 wf-rbd --vm -c limits.memory=4GiB -c limits.cpu=2 -d root,size=30GiB
incus storage volume create default wf-rbd-osd1 --type=block size=8GiB
incus storage volume attach default wf-rbd-osd1 wf-rbd

We repeated the last two lines for wf-rbd-osd2, wf-rbd-osd3 and wf-rbd-base. Inside the VM they appear as /dev/sdb to /dev/sde. Then the packages, with Debian's own Ceph client tools (version 18.2.7 Reef, because Ceph publishes no Tentacle packages for Debian 13 yet) and cephadm downloaded from Ceph:

# as root inside the VM
apt-get install -y podman lvm2 chrony curl fio jq ceph-common
apt-get install -y openssh-server
curl -s --remote-name --location https://download.ceph.com/rpm-tentacle/el9/noarch/cephadm
chmod +x cephadm
./cephadm bootstrap --mon-ip 10.141.188.173 --single-host-defaults --skip-monitoring-stack --skip-dashboard

Use your server's own IP address after --mon-ip. Our first bootstrap failed after two minutes with Connect call failed ('10.141.188.173', 22): cephadm manages even its own host over SSH, and the minimal Debian image had no SSH server. With openssh-server installed, bootstrap finished in 33 seconds and printed Bootstrap complete. The flag --single-host-defaults lets Ceph place copies on different disks of one machine instead of different machines, and it sets the default pool size to 2. We checked that with ceph config dump.

The key-type trap with older clients

A fresh 20.2.4 cluster allows only the new aes256k keys. Debian's 18.2.7 client tools cannot even read them (error parsing file /etc/ceph/ceph.client.admin.keyring ... Malformed input), so we ran every admin command through ./cephadm shell --, which starts a Tentacle container. The release notes say the kernel client reads them only from Linux 7.0, and our VM runs 6.12. For the lab, we allowed the older key type for one client that may only touch the rbd pool:

# as root inside the VM
./cephadm shell -- ceph config set osd osd_memory_target 1073741824
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdb
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdc
./cephadm shell -- ceph orch daemon add osd wf-rbd:/dev/sdd
./cephadm shell -- ceph osd pool create rbd
./cephadm shell -- rbd pool init rbd
./cephadm shell -- rbd create --size 4G rbd/vm-disk
./cephadm shell -- ceph mon set auth_allowed_ciphers aes,aes256k
./cephadm shell -- ceph auth get-or-create client.vmdisk mon 'profile rbd' osd 'profile rbd pool=rbd'
./cephadm shell -- ceph auth rotate --key-type=aes client.vmdisk > /etc/ceph/ceph.client.vmdisk.keyring

Each OSD line answered Created osd(s) N on host 'wf-rbd'. The last two commands lower security, and Ceph said so at once with three health warnings, AUTH_INSECURE_CLIENT_KEY_TYPE, AUTH_INSECURE_KEYS_ALLOWED and AUTH_INSECURE_KEYS_CREATABLE, because the old type is the one the August security release replaces. That is fine for a throwaway lab. In production, use clients with kernel 7.0 or a vendor kernel with the backport, as the CephX key rotation guide describes.

Now map the image with the kernel client, format it and mount it:

# as root inside the VM
rbd --id vmdisk map rbd/vm-disk
mkfs.ext4 -q /dev/rbd0
mkdir -p /mnt/rbd
mount /dev/rbd0 /mnt/rbd

rbd map printed /dev/rbd0, and dmesg showed rbd: rbd0: capacity 4294967296 features 0x3d. The baseline disk got the same mkfs.ext4 and was mounted at /mnt/base.

Adding a small CephFS on the same cluster

One command creates the filesystem, its two pools and the metadata servers (MDS). Our 4 GiB VM ran out of memory during this step, so we added a swap file and kept a single MDS.

# as root inside the VM
fallocate -l 4G /swapfile
chmod 600 /swapfile
mkswap -q /swapfile
swapon /swapfile
./cephadm shell -- ceph fs volume create labfs
./cephadm shell -- ceph orch apply mds labfs --placement=1
./cephadm shell -- ceph fs authorize labfs client.fsuser / rw
./cephadm shell -- ceph auth rotate --key-type=aes client.fsuser > /etc/ceph/ceph.client.fsuser.keyring
mkdir -p /mnt/cephfs
mount -t ceph fsuser@.labfs=/ /mnt/cephfs

ceph -s then listed volumes: 1/1 healthy and two new pools, cephfs.labfs.meta with 16 PGs and cephfs.labfs.data with 64, both with 2 copies like the RBD pool. df -h /mnt/cephfs showed 5.3G free, because every byte stored costs two bytes of raw disk. Idle after the tests, each OSD used about 395 MiB, the manager 170 MiB and the monitor 37 MiB (ceph orch ps).

Ceph block storage performance: real fio numbers

We ran five fio tests on each target, each 45 seconds after a 5-second warm-up, on a 2 GiB file with direct I/O so the page cache stays out of it. These are the exact command lines for the RBD mount; for the other targets only the path changes. We also added --output-format=json to read the results with jq.

# as root inside the VM
fio --name=randread4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randread --bs=4k --iodepth=32
fio --name=randwrite4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=32
fio --name=qd1write4k --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1
fio --name=seqread1m --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=read --bs=1M --iodepth=8
fio --name=seqwrite1m --filename=/mnt/rbd/fio.test --size=2G --direct=1 --ioengine=libaio --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=write --bs=1M --iodepth=8

The third test, one 4 KiB write at a time, is closest to how a database commit feels.

Read this table with care: one VM with 2 vCPU on a shared host, all OSDs on one machine, and no real network between client and OSDs. It shows the relative overhead of the Ceph layers, not what a production cluster delivers. Both Ceph targets used pools with 2 copies.

fio testPlain virtual disk (ext4)RBD, kernel client (ext4)CephFS, kernel mount
4K random read, QD3249,672 IOPS6,867 IOPS7,066 IOPS
4K random write, QD3231,942 IOPS1,934 IOPS2,084 IOPS
4K write, QD1 (avg latency)3,271 IOPS (0.26 ms)491 IOPS (1.99 ms)525 IOPS (1.86 ms)
1M sequential read, QD83,894 MiB/s930 MiB/s983 MiB/s
1M sequential write, QD83,220 MiB/s414 MiB/s377 MiB/s
5,000 small files: create + sync0.31 s0.50 s3.85 s
Same files: list + delete + sync0.11 s0.13 s2.75 s

What the numbers say:

The small-file test is a plain shell function that we ran once per mount point (it needs the bc package):

# as root inside the VM
smallfiles() { d=$1/small; mkdir -p $d; s=$(date +%s.%N); for i in $(seq 1 5000); do echo hello > $d/f$i; done; sync; m=$(date +%s.%N); ls -l $d | wc -l >/dev/null; rm -rf $d; sync; e=$(date +%s.%N); echo "$1 create5000+sync=$(echo "$m - $s" | bc)s list+delete=$(echo "$e - $m" | bc)s"; }
smallfiles /mnt/base; smallfiles /mnt/rbd; smallfiles /mnt/cephfs

Repeat your runs: our first RBD run, minutes after the autoscaler had split the pool into 32 placement groups, was 24% slower on random reads and 41% slower on random writes than the run in the table.

krbd or librbd: the cache that flatters benchmarks

fio can also talk to Ceph through librbd, the same library QEMU uses, with --ioengine=rbd. We wrote a second 2 GiB image full, then tested it raw (no filesystem) three ways: through the kernel device, through librbd with default settings, and through librbd with its cache switched off.

# as root inside the VM
fio --name=qd1write4k --ioengine=rbd --clientname=vmdisk --pool=rbd --rbdname=fio-img --size=2G --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1
CEPH_ARGS="--rbd_cache=false" fio --name=qd1write4k --ioengine=rbd --clientname=vmdisk --pool=rbd --rbdname=fio-img --size=2G --runtime=45 --ramp_time=5 --time_based --group_reporting --rw=randwrite --bs=4k --iodepth=1

The other rows use the same --rw, --bs and --iodepth values as the first set of tests. For the krbd column we mapped the same image (it became /dev/rbd1) and used --filename=/dev/rbd1 --direct=1 --ioengine=libaio instead of the three rbd options.

Raw image, 2 copieskrbd /dev/rbd1librbd, cache offlibrbd, default cache
4K random read, QD326,804 IOPS5,948 IOPS6,498 IOPS
4K random write, QD321,848 IOPS1,607 IOPS3,266 IOPS
4K write, QD1 (avg latency)509 IOPS (1.93 ms)403 IOPS (2.44 ms)3,528 IOPS (0.26 ms)
1M sequential write, QD8386 MiB/s359 MiB/s420 MiB/s

With the cache off, the kernel client was slightly faster than librbd on this small VM. With librbd's default cache, single writes looked almost nine times faster, because librbd's client cache (32 MiB per image by default, see the RBD config reference) absorbed them in memory. A database that waits for fsync on every commit would not see that speed. Benchmark with the cache off, or with a workload that flushes, or you are measuring RAM.

The built-in rbd bench uses librbd with the cache on. On the same image it reported 3,404 random 4K writes per second with 16 threads, 569 MiB/s for 4 MiB writes and 1.4 GiB/s for 4 MiB reads:

# as root inside the VM
rbd --id vmdisk bench --io-type write --io-size 4K --io-threads 16 --io-total 256M --io-pattern rand rbd/fio-img
rbd --id vmdisk bench --io-type write --io-size 4M --io-threads 16 --io-total 2G rbd/fio-img
rbd --id vmdisk bench --io-type read --io-size 4M --io-threads 16 --io-total 2G rbd/fio-img

Fine as a smoke test for a new pool; plan with fio.

Replication size 3 vs 2: space and write latency

The pool "size" is how many copies Ceph keeps; min_size is how many must be available before Ceph accepts writes. We switched the RBD pool from 2 to 3 copies and ran the raw krbd tests again:

# as root inside the VM
./cephadm shell -- ceph osd pool set rbd size 3
./cephadm shell -- ceph osd pool set rbd min_size 2
Raw krbd imagesize 2size 3Change
4K random read, QD326,804 IOPS7,131 IOPSNo real change (reads come from one copy)
4K random write, QD321,848 IOPS1,494 IOPS19% fewer
4K write, QD1 latency1.93 ms2.42 ms25% slower
1M sequential write386 MiB/s263 MiB/s32% less
Raw space for 4.0 GiB stored8 GiB (calculated)12 GiB (from ceph df)Usable space falls from 1/2 to 1/3 of raw

Size 2 is faster and cheaper, and the Ceph docs still advise against it in production: with two copies, one failed disk during the rebuild of another loses data. Keep size 3 and min_size 2 for anything you care about, and buy the disks for it. If space hurts, erasure coding is the tool, which our architecture article covers.

Switching to size 3 also gave us our first crash. Ceph started copying 4 GiB to the new third copies, one OSD grew past 1.3 GB of memory, and the kernel's out-of-memory killer stopped it inside our 4 GiB VM, which had no swap.

What limits Ceph performance, in the order you will meet it

Safe tuning for beginners

What we would do on a first cluster:

  1. Leave the PG autoscaler on, and let it finish before you benchmark. ceph -s should show every PG as active+clean.
  2. Keep osd_memory_target at its 4 GiB default on real servers. The hardware guide says not to go below 2 GB. Our 1 GiB lab setting worked until recovery started, then caused the crash described above.
  3. Never map RBD images or mount CephFS with the kernel client on a host that also runs OSDs, except in a lab. We learned this the hard way: after the out-of-memory kill, the OSD restart scanned /dev/rbd0, a device served by those same stopped OSDs, and hung until we reset the VM. The ceph-users list has the background on this deadlock risk.
  4. Give VMs librbd with the default cache settings; rbd_cache_writethrough_until_flush keeps the cache safe for guests that never send flushes. Just do not benchmark through it.
  5. Keep the cluster network separate from client traffic once you grow past a lab, and give it 10 Gb/s or more.

To measure a plain server disk first, see our VPS benchmark commands.

Frequently asked questions

What is Ceph RBD?

RBD, the RADOS Block Device, is Ceph's block storage. It presents a virtual disk that is thin-provisioned and split into objects, 4 MiB each by default, stored with two or three copies across the cluster. Linux maps it as /dev/rbdX, while QEMU/KVM uses it through the librbd library.

Is CephFS slower than RBD?

For large files read or written by one client, they are close: in our lab, CephFS and RBD were within 10% on 4K random I/O and 1M sequential tests. CephFS is much slower with many small files, because every create and delete goes through the metadata server. Creating 5,000 files took 3.85 s on CephFS and 0.50 s on ext4 over RBD.

What is the default block size in Ceph?

RBD images use a 4 MiB object size by default, shown as "order 22" in rbd info. It can be set per image from 4K to 32M with --object-size when the image is created. The filesystem on top still uses its own block size, usually 4 KiB.

Can two servers mount the same RBD image?

Not with ext4 or XFS; two hosts writing to it will corrupt the filesystem. Use CephFS when several machines need the same files.

How much space do I lose with replication 3?

Two thirds of the raw capacity. With size 3, every terabyte you store uses three terabytes of disk; in our lab, 4.0 GiB of data used 12 GiB. Size 2 uses half, but the Ceph docs warn that it risks data loss in production.

Why does rbd map fail on Ceph 20.2.4?

Ceph 20.2.4 creates keys of the new aes256k type, and kernel clients older than Linux 7.0 cannot use them unless the vendor has backported support. Use a newer kernel, or, in a test cluster only, allow the older aes type and rotate one client key to it, as shown in this article.

Try it on servers you control

One VM teaches the commands; three servers show real replication and recovery. RS Computers runs KVM VPS and VDS in Amsterdam, Dublin and Prishtina, all on NVMe with unmetered traffic and their own IPv4 and IPv6, and the VDS plans come with 10 Gb/s ports, which is the speed Ceph asks for between nodes. Three VDS Small or Medium servers in one city make a realistic test cluster; follow our three-node lab guide to build it. vCPUs on our plans are shared, so use them for labs and proofs of concept, and put production OSDs on bare metal. Availability by city is on the plans page, and if you want help sizing a lab, message us on Telegram.

← All articles

Chat on Telegram