Ceph object storage is the Ceph Object Gateway (RGW), an HTTP service that speaks a large subset of the Amazon S3 and OpenStack Swift APIs and keeps every bucket and object inside a Ceph cluster. Ceph block storage is RBD, the RADOS Block Device, which cuts each virtual disk into objects, 4 MiB by default, and spreads them over the same cluster. Both sit on RADOS, the object store underneath everything Ceph does, so one cluster can hand out VM disks, S3 buckets and a shared file system (CephFS) at the same time.
Where Ceph stands in October 2026:
- Tentacle (20.2.x) is the release to deploy. Squid reaches its estimated end of life on 31 October 2026, and Reef ended in March.
- Patch to 20.2.4 or 19.2.6, which fixed four CVEs in August 2026.
- Plan for three monitors, at least three OSD hosts (five give the cluster room to heal), three copies of every object and at least 10 Gb/s networking.
- Erasure coding 4+2 uses 1.5 TB of raw disk per TB stored instead of 3 TB, and suits RGW buckets better than VM disks.
- Don't build new cache tiers or iSCSI gateways; both are on their way out.
What is the Ceph storage system?
Ceph is open-source, software-defined storage that runs on ordinary servers and disks. Sage Weil first presented it in 2006, Red Hat bought his company Inktank in 2014, and the code is licensed under the LGPL. In 2026 the Ceph trademark moved from IBM and Red Hat to the Linux Foundation, so no single vendor owns the name any more.
One cluster can be "unified" because everything ends up as objects in RADOS (Reliable Autonomic Distributed Object Store). RBD turns those objects into block devices, RGW into S3 and Swift buckets and CephFS into a POSIX file system, while librados lets an application talk to RADOS directly. When a drive or a whole server dies, Ceph keeps serving data from the surviving copies and rebuilds the missing ones in the background. That self-healing is why people put up with how complex Ceph is to run.
Ceph object storage architecture: monitors, OSDs, placement groups and CRUSH
A cluster is built from a few daemon types. The first three counts are Ceph's documented minimums; the last two are what we would start with.
| Daemon | What it does | Count to plan for |
|---|---|---|
Monitor (ceph-mon) | Keeps the cluster maps, handles authentication | 3, always an odd number |
Manager (ceph-mgr) | Metrics, dashboard, orchestration modules | 2, active and standby |
OSD (ceph-osd) | Stores data on one drive, replicates, recovers, rebalances | 3 at the very least, one per drive |
MDS (ceph-mds) | Metadata for CephFS only | 1 active plus a standby, if you use CephFS |
RGW (radosgw) | The S3 and Swift HTTP endpoint | 2 behind a load balancer, if you use object storage |
OSD means object storage daemon; "object storage device" is an older expansion of the same letters. There is one OSD per drive, formatted with BlueStore, Ceph's default back end, which writes to the raw device with no file system in between.
No central table says where an object lives, which is what lets Ceph scale. Each object belongs to a pool, its name is hashed into one of the pool's placement groups (PGs), and CRUSH, a placement algorithm fed with the cluster map, turns each PG into a list of OSDs on different hosts or racks. Clients run CRUSH themselves and go straight to the right OSD. A write lands on the primary OSD, which forwards it to the replicas and acknowledges only when they have confirmed it. The PG autoscaler is on by default and aims for about 100 PGs per OSD, although the placement group docs recommend 200 for all but the smallest clusters. The Umbrella RC1 build makes 200 the default, so expect the autoscaler to split PGs after that upgrade, and pin the old values beforehand if your OSDs run below the 4 GiB memory target.
What is Ceph RBD block storage?
RBD (RADOS Block Device) is Ceph's block storage: disk images that are thin-provisioned, resizable and striped over many OSDs. A 100 GB image only takes the space you have written to it. Snapshots and copy-on-write clones are part of the format, which is how OpenStack and Kubernetes create a new disk in seconds. A Linux host can map an image as /dev/rbd0 through the kernel module, while QEMU/KVM uses librbd, so the VM's disk talks to the cluster directly with its own client-side cache. OpenStack, CloudStack and Proxmox VE attach VM disks that way.
Ceph block size: the 4 MiB object
The "block size" people ask about is usually the object size. RBD splits images into 4 MiB objects by default. The rbd man page allows 4K to 32M, set per image with --object-size or cluster-wide with rbd_default_order (default 22, since 222 bytes is 4 MiB), and rbd info shows it as order 22 (4 MiB objects). The VM never sees this number. Leave it alone unless you have measured a reason: smaller objects mean more to track and recover, larger ones pile a busy image onto fewer OSDs.
Gateways, and what changed in Tentacle
VMware hosts cannot run librbd, so they need a gateway. The iSCSI gateway has been in maintenance since November 2022, and Ceph's docs point new deployments to the NVMe-oF gateway, which exports RBD images over NVMe/TCP with high availability on by default. Tentacle also added live migration that imports an image instantly from another Ceph cluster or over NBD.
Looking for open-source block storage in general? RBD is the option that scales across many hosts and gives you object and file storage from the same hardware. For two servers mirroring one disk, DRBD is simpler, and inside Kubernetes alone Longhorn is easier to live with.
How does Ceph object storage (RGW) work?
The Ceph Object Gateway (RGW, or RADOS Gateway) is a web server built on librados that accepts S3 and Swift requests and stores the data as RADOS objects. Ceph's docs promise compatibility with a large subset of each API, not all of it. An earlier version of this article claimed full Swift support; that was wrong, so test the exact calls your software makes. S3 and Swift share one namespace, so an object uploaded through Swift can be read with an S3 client.
RGW daemons keep no state, so you scale by adding more behind a load balancer. Performance is decided by the pools underneath. The bucket index is stored as OMAP data in RocksDB, split into 11 shards by default, and belongs on replicated SSD or NVMe (the index pool has to be replicated anyway, because erasure-coded pools do not support OMAP by default). The data pool is the one most clusters put on erasure coding, because objects are written once and read many times.
The capabilities buyers compare between object stores are mostly there: bucket policies, STS, lifecycle rules, versioning, Object Lock, server-side encryption, bucket notifications and multisite replication. Tentacle made accounts the main IAM model and deprecated the tenant-level IAM APIs, so start new deployments with accounts. Two of the August 2026 CVEs were RGW bugs (CVE-2026-39944 and CVE-2026-54330, both in signature checks), and multisite users must set rgw_sigv4_insecure to true before upgrading and back to false afterwards. That is the step people forget.
With MinIO's GitHub repository archived on 25 April 2026, RGW is one of the most complete S3 servers still developed in the open, but it comes with a cluster attached. For a single backup target, a small single-binary store is far less work; our guide to off-site backups after MinIO compares the options.
Block, object or CephFS: which one should you use?
| Question | Block (RBD) | Object (RGW) | File (CephFS) |
|---|---|---|---|
| How software reaches it | A disk device | HTTP calls to the S3 or Swift API | A mounted POSIX file system |
| Shared between machines? | One writer at a time, normally | Any number of clients | Many clients mount it at once |
| Typical workloads | VM disks, databases, ReadWriteOnce volumes | Backups, media, logs, data lakes | Shared folders, home directories, ReadWriteMany volumes |
| Extra daemons | None | RGW plus a load balancer | MDS, active and standby |
| Poor fit | Data many servers must change together | Apps that rewrite small parts of a file | Millions of tiny files with constant metadata churn |
Our rule of thumb: if the software asks for a disk, give it RBD. If it speaks S3, give it RGW. Use CephFS only when several machines really need the same directory tree. Given the choice, prefer object storage: a bucket has no size to provision and no file system to repair. The catch is that you cannot change four bytes in the middle of an object; you upload a new one.
How is Ceph used in cloud computing?
Most people meet Ceph as the storage layer of a cloud: OpenStack, Kubernetes or Proxmox VE. The platform asks Ceph for disks and buckets, and Ceph makes sure a lost server does not mean a lost VM.
What is Ceph storage in OpenStack?
In OpenStack, Ceph usually provides every kind of storage at once. The Ceph block devices and OpenStack guide sets up four RBD pools, and RGW adds object storage through Keystone integration:
| OpenStack service | What it stores | Ceph side |
|---|---|---|
| Glance | VM images | RBD, pool images |
| Cinder | Persistent volumes | RBD, pool volumes |
| Cinder backup | Volume backups | RBD, pool backups |
| Nova | Ephemeral and boot disks | RBD, pool vms |
| Swift API with Keystone | Object storage for tenants | RGW, authenticating against Keystone |
Two details decide whether it feels fast. Upload Glance images as raw, not QCOW2: only raw images let Cinder and Nova create a disk as a copy-on-write clone, which takes seconds whatever the image size. (Our old version said Glance stores images "as RBD snapshots"; to be exact, each is an RBD image with a protected snapshot that new volumes clone from.) And keep versions in step, because the Cinder RBD driver supports the active stable Ceph releases plus the two before them. OpenStack 2026.2 Hibiscus came out on 30 September 2026; pair it with Tentacle. Leaving VMware? Read our article on moving from VMware to OpenStack.
Kubernetes: Rook and Ceph-CSI
Rook, a CNCF graduated project, deploys and runs Ceph inside Kubernetes, and Ceph-CSI turns a PersistentVolumeClaim into an RBD image or a CephFS subvolume. Since Rook v1.20, Rook no longer deploys the CSI drivers; the ceph-csi-operator does, and old ConfigMap settings have to be migrated. Rook's upgrade guide also tells users to move to v1.20.6 or later and to Ceph 20.2.4 or 19.2.6 because of CVE-2025-30156. Rook supports Squid and Tentacle today. For OpenShift-style Kubernetes, see running OKD on OpenStack.
Proxmox VE
Proxmox VE installs and manages Ceph from its own web interface, which is how many three-node clusters start. Three nodes is the minimum, not a comfortable size. Our Proxmox vs OpenStack comparison helps if you are still choosing a platform.
Replication or erasure coding: how much disk do you need?
By default Ceph keeps three copies of every object (pool size 3), and its pool docs warn that size 2, or min_size 1, risks data loss in production. Erasure coding (EC) splits each object into K data chunks plus M coding chunks, and any M of them can be lost without losing data. The default profile is k=2, m=2 with the ISA-L plugin, which replaced the unmaintained Jerasure library as the default in Tentacle.
| Layout | Raw disk per 1 TB stored | Hosts it can lose without data loss | Where it fits |
|---|---|---|---|
| 3-way replication | 3 TB | 2 | VM disks, databases, every index and metadata pool |
| EC 2+2 (default) | 2 TB | 2 | Trying EC on a small cluster; needs 4 hosts |
| EC 4+2 | 1.5 TB | 2 | RGW data, backups, archives; needs 6 hosts |
| EC 8+3 | about 1.4 TB | 3 | Large archives; needs 11 hosts |
The host counts assume the default failure domain of host; add one more if you want the cluster to heal itself after a failure. EC has costs too: a write touches K+M drives, and rebuilding one chunk reads K others. RBD and CephFS can use EC data pools only with allow_ec_overwrites on and their metadata in a replicated pool. Tentacle's FastEC, switched on per pool with allow_ec_optimizations, speeds up small reads and writes, the first good reason in years to try EC for VM disks. We would still keep databases on replicated pools. Across two data centres, stretch mode keeps two copies in each (size 4), with a tiebreaker monitor in a third site.
Ceph hardware and network sizing
These figures come from Ceph's hardware recommendations for Tentacle, which tell you to size for peak use, not averages.
| Daemon | CPU | RAM | Storage |
|---|---|---|---|
| OSD on HDD | 1 thread minimum, 3 recommended | 4 GiB memory target per OSD | Drives of 1 TiB or more; DB/WAL on SSD, 4 to 5 HDDs per SATA SSD or up to 15 per NVMe |
| OSD on NVMe | 4 threads minimum, 6 recommended | 4 GiB memory target per OSD | Drives of 1 TiB or more, enterprise class |
| Monitor | 2 cores minimum | 5 GB or more | 100 GB, SSD strongly urged |
| MDS (CephFS) | 2 cores, clock speed matters most | 8 GiB or more | Metadata pool on SSD |
Two rules on that page catch people out. Give each server more RAM than its OSD count times the memory target times two, so a node with 12 NVMe OSDs wants about 96 GiB. And treat 10 Gb/s as the floor: the docs work out that re-replicating 10 TiB takes about 30 hours over 1 Gb/s and about 3 hours over 10 Gb/s, and a second failure in those 30 hours is what you bought Ceph to survive. Use 25 Gb/s for substantial workloads and 100 Gb/s for dense NVMe nodes, bonded across two switches. We would add three rules of our own:
- Plan five OSD hosts, not three. With three hosts and three copies, a dead host leaves nowhere to rebuild.
- Buy enterprise SSDs with power-loss protection. Ceph waits for every write to be safely on disk, and many consumer drives crawl under that load.
- Keep drive sizes similar. CRUSH weights by capacity, so a 16 TB drive among 4 TB ones gets four times the traffic.
Ceph block storage performance: what to expect
At scale, Ceph is very fast. A test on 63 nodes of a 68-node cluster published in January 2024 (630 NVMe OSDs, two 100 GbE links per node, Quincy 17.2.7, fio against RBD images through librbd) read in 4 MB blocks at 1,025 GiB/s and served 25.5 million 4K random reads per second with 3-way replication. With 6+2 erasure coding it wrote 4 MB blocks faster, 387 GiB/s against 270 GiB/s, because it writes 1.33 bytes per byte stored instead of 3. Small random writes went the other way: 936,000 per second against 4.9 million. That is the replication versus EC trade-off in two numbers.
On a small cluster, latency matters more. Every write crosses the network to the primary OSD and then to two replicas, so a single-threaded database commit on Ceph is always slower than on local NVMe. The "why is Ceph so slow" threads usually involve 1 Gb/s networks, consumer SSDs or a three-node lab on one link. What helps, roughly in order: a faster network and enterprise NVMe, more OSDs and parallel clients, enough CPU (at 4 to 6 threads per NVMe OSD, Ceph on flash is CPU hungry) and librbd caching for VMs. Cache tiering does not help: it has been deprecated since Reef, and the docs say not to deploy new cache tiers.
Which Ceph release should you run in October 2026?
The Ceph releases index lists two active lines. A new named release ships roughly once a year and gets about two years of fixes.
| Release | Status | Latest version | End of life |
|---|---|---|---|
| Tentacle (20.x) | Active, the one to deploy | 20.2.4, 19 August 2026 | About 1 June 2027 (estimate) |
| Squid (19.x) | Active until the end of October | 19.2.6, 19 August 2026 | 31 October 2026 (estimate) |
| Reef (18.x) | Archived | 18.2.8, 20 March 2026 | 20 March 2026 |
| Quincy (17.x) | Archived | 17.2.9, a final hotfix after end of life | 13 January 2025 |
| Umbrella (21.x) | Release candidates only | 21.1.1 RC1, 28 September 2026 | Not set |
Reef and Squid clusters can upgrade straight to Tentacle; Quincy and older go through Reef or Squid first. Whatever you run, take the 19 August 2026 point release. Besides the two RGW bugs, it fixes a CephX authentication bypass (CVE-2025-30156) and a monitor authorization flaw (CVE-2026-50152), and it adds a new CephX key type, aes256k, with a key rotation procedure to follow afterwards.
The Umbrella candidates bring faster erasure-coded reads and client-side encryption for CephFS, and a cephadm change merged on 6 October 2026 says Umbrella's official builds will need x86-64-v3 (AVX2-era) CPUs, so check your oldest servers. Vampire follows in spring 2027.
Ceph enterprise storage: subscription, self-run, or something simpler?
Ceph enterprise storage usually means one of three things: IBM Storage Ceph (version 9 is built on Tentacle, version 8 on Squid), Ceph inside Canonical's or Proxmox's platforms with their support, or upstream Ceph deployed with cephadm, where you are the support contract. Running it yourself makes sense when more than one person can be woken up for it and the data fills five or more servers. If your data fits on one server's NVMe and an hour of restore time is acceptable, one well-backed-up server is far less work. Ceph is the right answer when the data outgrows one machine, or one machine failing is not acceptable.
Where RS Computers fits, and where it doesn't
We don't sell Ceph clusters or S3 buckets. RS Computers rents KVM virtual servers, VPS and VDS, in Amsterdam (Netherlands), Prishtina (Kosovo) and Dublin (Ireland), with NVMe storage, your own IPv4 and IPv6 address, free weekly backups and a 1 Gb/s port on VPS plans or 10 Gb/s on VDS plans. The plans page shows which of the three cities has room for new servers today. Production OSDs belong on bare metal with whole drives and their own network, and our vCPUs are shared, so we would not build a production Ceph cluster from virtual servers, ours or anyone else's. Around Ceph, they are useful:
- As a lab, to learn the commands, the failure modes and the upgrade steps before you touch production hardware.
- As an RGW proof of concept, to see which S3 calls your tools and apps really make.
- For tooling outside the cluster, such as copying critical buckets to another city (replication protects against failed disks, not deleted buckets) or checking your RGW endpoint from outside.
- As the simpler option when a private cloud is overkill: a VDS with up to 960 GB of NVMe and free weekly backups, plus your own off-site copy on a second server in another city.
A one-server Ceph lab with MicroCeph
Canonical's MicroCeph packages Ceph as a snap and can run a whole cluster on one machine, with loop files standing in for drives. These are the steps from its README, for an Ubuntu server (if the snap command is missing, apt install snapd adds it). The third line creates three 4 GiB loop-file OSDs, so leave at least 12 GiB free on the root disk:
# as root
snap install microceph
microceph cluster bootstrap
microceph disk add loop,4G,3
microceph status
ceph status
microceph status should list your server with its services and three disks, and ceph status should report three OSDs up and in; if health shows a warning in the first minute, run it again. Three OSDs at the default 4 GiB memory target can grow to about 12 GiB, so give the lab a VDS Medium (8 vCPU, 16 GB RAM). That is below the production rule above, which is fine for a lab that never sees real load. Then create an RBD image and check its object size with rbd info, or stop one OSD and watch the cluster go degraded and recover. New to SSH? Our 30-day Linux lab plan covers logins, sudo and the firewall first.
Planning a lab or a proof of concept and not sure which plan fits? Message us on Telegram or email info@rscomputers-ks.com, and we will suggest a plan and quote the setup.
Frequently asked questions
What does Ceph stand for?
It is not an acronym. Ceph is short for cephalopod, the octopus and squid family, a theme that runs through release names such as Octopus, Squid and Tentacle. Sage Weil created it and first presented it in 2006.
Is Ceph like S3?
Ceph's object gateway, RGW, speaks a large subset of the Amazon S3 API, so most S3 tools and SDKs work once you change the endpoint and keys. The difference is that you run the storage yourself, and you should test the specific API calls your software depends on.
Is Ceph faster than ZFS?
On a single server, no. ZFS writes to local disks while Ceph waits for copies on other machines across the network, so ZFS has lower latency. They solve different problems: ZFS protects data inside one box, while Ceph keeps data available when a whole server fails and grows by adding servers.
Is 10GbE enough for Ceph?
It is the minimum Ceph's hardware guide recommends, and it suits small HDD clusters and modest SSD ones. NVMe nodes can fill a 10 Gb/s link on their own, so the same guide suggests 25 Gb/s for substantial workloads and 100 Gb/s for dense nodes.
Is Ceph a NAS?
Not by itself. CephFS is a shared file system that Linux clients mount directly. To serve ordinary NAS clients you add a gateway: NFS through NFS-Ganesha, or SMB, which Tentacle made easier with a manager module that creates Samba shares on CephFS.
Is Ceph better than RAID?
They work at different levels. RAID protects against a failed drive inside one server, while Ceph keeps copies or erasure-coded chunks on different servers, so it also survives a dead motherboard or a lost rack. The usual advice is one OSD per raw drive with no RAID underneath, letting Ceph handle redundancy.