Deploy Rook-Ceph on Kubernetes: Block and Object Storage from Scratch
Quick answer
Turn a handful of raw disks into a self-healing storage cluster. Install the Rook operator with Helm, stand up a CephCluster, verify OSDs with the toolbox, then serve both RBD block volumes and S3-compatible object storage — every command runnable, step by step.
- Step 1: Confirm You Have Raw Devices
- Step 2: Install the Rook Operator with Helm
- Step 3: Apply the CephCluster
- Step 4: Watch the Cluster Come Up
- Step 5: Install and Use the Toolbox
advanced · 90 min
Before you begin
- A multi-node Kubernetes cluster (3+ worker nodes) with cluster-admin access
- At least one unformatted raw block device (extra disk) per storage node — no partitions, no filesystem, no LVM
- kubectl configured for the cluster
- Helm v3 installed
- Comfortable with PVs, PVCs, and StorageClasses
Most people meet Ceph as a black box behind a managed service and never see how the pieces fit. Rook changes that: it's a Kubernetes operator that runs a full Ceph cluster inside your cluster, turning raw disks attached to your nodes into replicated block, file, and object storage — all driven by custom resources instead of the sprawling Ceph CLI. You write YAML; the operator handles mons, OSDs, managers, and the CSI drivers.
In this tutorial you'll build that from nothing: install the Rook operator with Helm, hand it some raw disks via a CephCluster resource, confirm the OSDs came up with the Ceph toolbox, then carve out two kinds of storage on top — an RBD-backed StorageClass for ReadWriteOnce block volumes, and a CephObjectStore with an S3 bucket you can actually put objects into. By the end you'll have a self-healing storage layer that other workloads can consume like any cloud provider's disks.
If you're still deciding whether Ceph is the right engine for your cluster versus lighter options, read Rook-Ceph vs Longhorn vs OpenEBS first — Rook-Ceph is the heavyweight: most capable, most resource-hungry, and the right pick when you need block, file, and object storage from one system.
What You'll Build
- The Rook operator installed with Helm into the
rook-cephnamespace - A CephCluster that consumes raw devices across your nodes as OSDs
- The Rook toolbox pod for running
ceph statusandceph osd status - A CephBlockPool + StorageClass (
rook-ceph.rbd.csi.ceph.com) serving RBD block volumes - A test PVC + pod that mounts a Ceph-backed volume
- A CephObjectStore + ObjectBucketClaim exposing an S3-compatible endpoint you can write to
Step 1: Confirm You Have Raw Devices
This is the step people skip, and it's the one that breaks everything. Rook creates OSDs only from raw, unformatted block devices. A disk with an existing partition table, filesystem, or LVM signature is silently ignored.
On each storage node, check what's available:
lsblk -fYou want a device (e.g. /dev/sdb or /dev/nvme1n1) whose FSTYPE column is empty and that has no child partitions. If your extra disk already has leftover data from a previous install, wipe it (destructive — be sure of the device name):
sudo sgdisk --zap-all /dev/sdb
sudo wipefs -a /dev/sdbNote the device name — it's usually consistent across nodes (/dev/sdb), which lets you use a deviceFilter later.
Step 2: Install the Rook Operator with Helm
Add the Rook chart repo and install the operator. The operator itself stores no data — it's the controller that watches your Ceph custom resources and reconciles the actual daemons.
1helm repo add rook-release https://charts.rook.io/release
2helm repo update
3
4helm install rook-ceph rook-release/rook-ceph \
5 --namespace rook-ceph \
6 --create-namespace \
7 --version v1.20.7Wait for the operator to be running:
kubectl -n rook-ceph rollout status deploy/rook-ceph-operator
kubectl -n rook-ceph get podsYou should see a single rook-ceph-operator-* pod in Running. No Ceph daemons exist yet — the operator is idle until you give it a CephCluster.
Step 3: Apply the CephCluster
This is the resource that tells Rook to build a Ceph cluster. useAllNodes: true schedules OSDs across every eligible node; useAllDevices: false plus a deviceFilter keeps Rook off your OS disk and targets only the extra disks.
1apiVersion: ceph.rook.io/v1
2kind: CephCluster
3metadata:
4 name: rook-ceph
5 namespace: rook-ceph
6spec:
7 cephVersion:
8 image: quay.io/ceph/ceph:v20.2.1
9 dataDirHostPath: /var/lib/rook
10 mon:
11 count: 3
12 allowMultiplePerNode: false
13 mgr:
14 count: 2
15 dashboard:
16 enabled: true
17 storage:
18 useAllNodes: true
19 useAllDevices: false
20 deviceFilter: "^sdb"Save it as cluster.yaml and apply:
kubectl apply -f cluster.yamlThe deviceFilter: "^sdb" is a regex matching device names — adjust it to whatever lsblk showed you (e.g. ^nvme1n1). If your disk names differ per node, set useAllNodes: false and provide an explicit per-node nodes: list instead (Rook ignores the nodes: list while useAllNodes: true, so you must flip it to false for per-node config to take effect).
Step 4: Watch the Cluster Come Up
This takes several minutes. The operator provisions mons first, then the manager, then discovers devices and creates one OSD per raw disk.
kubectl -n rook-ceph get pods -wYou're waiting for this rough shape:
# 3 mons, 2 mgrs, and one OSD per disk
kubectl -n rook-ceph get pods -l app=rook-ceph-mon
kubectl -n rook-ceph get pods -l app=rook-ceph-osdThe key signal is rook-ceph-osd-0, rook-ceph-osd-1, ... reaching Running. If no OSD pods appear after ~10 minutes, jump to Common Issues — it's almost always a device-detection problem.
Check the cluster resource's own view of health:
kubectl -n rook-ceph get cephcluster rook-cephThe PHASE should progress to Ready and HEALTH to HEALTH_OK (or HEALTH_WARN on a fresh 3-node cluster — that's fine while it settles).
Step 5: Install and Use the Toolbox
The toolbox is a pod with the Ceph CLI baked in, wired to your cluster's credentials. It's how you inspect Ceph directly.
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.20.1/deploy/examples/toolbox.yaml
kubectl -n rook-ceph rollout status deploy/rook-ceph-toolsNow open a shell and check status:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- bashInside the toolbox:
ceph status
ceph osd status
ceph osd treeceph status should report health: HEALTH_OK, your mon quorum (3 daemons), and OSDs up and in. ceph osd status lists each OSD with its capacity and node. This is your ground truth — if Ceph is healthy here, the Kubernetes layer will work.
Step 6: Create a Block Pool and StorageClass
Now the payoff: expose Ceph RBD as a Kubernetes StorageClass. First a CephBlockPool (a replicated pool with 3 copies), then a StorageClass that points the Ceph CSI RBD driver at it.
1apiVersion: ceph.rook.io/v1
2kind: CephBlockPool
3metadata:
4 name: replicapool
5 namespace: rook-ceph
6spec:
7 failureDomain: host
8 replicated:
9 size: 3
10---
11apiVersion: storage.k8s.io/v1
12kind: StorageClass
13metadata:
14 name: rook-ceph-block
15provisioner: rook-ceph.rbd.csi.ceph.com
16parameters:
17 clusterID: rook-ceph
18 pool: replicapool
19 imageFormat: "2"
20 imageFeatures: layering
21 csi.storage.k8s.io/fstype: ext4
22 # CSI credential secrets — REQUIRED, or PVCs hang in Pending and pods fail to mount.
23 # These secrets are created by the Rook operator in the operator namespace.
24 csi.storage.k8s.io/provisioner-secret-name: rook-csi-rbd-provisioner
25 csi.storage.k8s.io/provisioner-secret-namespace: rook-ceph
26 csi.storage.k8s.io/controller-expand-secret-name: rook-csi-rbd-provisioner
27 csi.storage.k8s.io/controller-expand-secret-namespace: rook-ceph
28 csi.storage.k8s.io/node-stage-secret-name: rook-csi-rbd-node
29 csi.storage.k8s.io/node-stage-secret-namespace: rook-ceph
30reclaimPolicy: Delete
31allowVolumeExpansion: trueSave as blockpool.yaml and apply:
kubectl apply -f blockpool.yaml
kubectl get storageclassreplicated.size: 3 with failureDomain: host means Ceph keeps three copies of every object, each on a different node — which is exactly why you need 3+ nodes. allowVolumeExpansion: true lets you grow PVCs later by editing their requested size.
Step 7: Provision a PVC and Mount It
Prove the StorageClass works end to end with a PVC and a pod that writes to it.
1apiVersion: v1
2kind: PersistentVolumeClaim
3metadata:
4 name: ceph-block-test
5spec:
6 accessModes:
7 - ReadWriteOnce
8 storageClassName: rook-ceph-block
9 resources:
10 requests:
11 storage: 1Gi
12---
13apiVersion: v1
14kind: Pod
15metadata:
16 name: ceph-block-writer
17spec:
18 containers:
19 - name: writer
20 image: busybox:1.36
21 command: ["sh", "-c", "echo rook-ceph-works > /data/hello.txt && sleep 3600"]
22 volumeMounts:
23 - name: vol
24 mountPath: /data
25 volumes:
26 - name: vol
27 persistentVolumeClaim:
28 claimName: ceph-block-testkubectl apply -f pvc-test.yaml
kubectl get pvc ceph-block-test # should reach Bound
kubectl get pod ceph-block-writer # should reach Running
kubectl exec ceph-block-writer -- cat /data/hello.txtA Bound PVC and rook-ceph-works printed back means the CSI driver provisioned an RBD image, mapped it, formatted it ext4, and mounted it into the pod. If the PVC hangs in Pending, see fix unbound PersistentVolumeClaims.
Step 8: Stand Up S3-Compatible Object Storage
Block storage is one half. The other reason people run Ceph is a self-hosted S3. A CephObjectStore deploys RADOS Gateway (RGW); an ObjectBucketClaim (OBC) then provisions a bucket and hands you credentials in a Secret.
1apiVersion: ceph.rook.io/v1
2kind: CephObjectStore
3metadata:
4 name: my-store
5 namespace: rook-ceph
6spec:
7 metadataPool:
8 failureDomain: host
9 replicated:
10 size: 3
11 dataPool:
12 failureDomain: host
13 replicated:
14 size: 3
15 gateway:
16 port: 80
17 instances: 1kubectl apply -f objectstore.yaml
kubectl -n rook-ceph get pods -l app=rook-ceph-rgwOnce the RGW pod is Running, create a StorageClass for buckets and an ObjectBucketClaim:
1apiVersion: storage.k8s.io/v1
2kind: StorageClass
3metadata:
4 name: rook-ceph-bucket
5provisioner: rook-ceph.ceph.rook.io/bucket
6reclaimPolicy: Delete
7parameters:
8 objectStoreName: my-store
9 objectStoreNamespace: rook-ceph
10---
11apiVersion: objectbucket.io/v1alpha1
12kind: ObjectBucketClaim
13metadata:
14 name: my-bucket
15spec:
16 generateBucketName: rook-demo
17 storageClassName: rook-ceph-bucketkubectl apply -f obc.yaml
kubectl get objectbucketclaim my-bucket # PHASE should be BoundThe OBC creates a ConfigMap (bucket name, host) and a Secret (access + secret keys), both named my-bucket.
Step 9: Write to the Bucket with an S3 Client
Pull the connection details the OBC generated and test with the AWS CLI from a throwaway pod.
export AWS_ACCESS_KEY_ID=$(kubectl get secret my-bucket -o jsonpath='{.data.AWS_ACCESS_KEY_ID}' | base64 -d)
export AWS_SECRET_ACCESS_KEY=$(kubectl get secret my-bucket -o jsonpath='{.data.AWS_SECRET_ACCESS_KEY}' | base64 -d)
export BUCKET_NAME=$(kubectl get cm my-bucket -o jsonpath='{.data.BUCKET_NAME}')The RGW service inside the cluster is rook-ceph-rgw-my-store.rook-ceph.svc. Run the CLI from a pod so you can reach it:
1kubectl run s3-test --rm -it --restart=Never \
2 --image=amazon/aws-cli:2.15.30 \
3 --env="AWS_ACCESS_KEY_ID=$AWS_ACCESS_KEY_ID" \
4 --env="AWS_SECRET_ACCESS_KEY=$AWS_SECRET_ACCESS_KEY" \
5 --command -- sh -c "
6 echo 'hello from ceph rgw' > /tmp/obj.txt &&
7 aws --endpoint-url http://rook-ceph-rgw-my-store.rook-ceph.svc:80 s3 cp /tmp/obj.txt s3://$BUCKET_NAME/obj.txt &&
8 aws --endpoint-url http://rook-ceph-rgw-my-store.rook-ceph.svc:80 s3 ls s3://$BUCKET_NAME/
9 "Seeing obj.txt in the listing means the full object path works: OBC provisioned the bucket, RGW served the S3 API, and your credentials authenticated.
Step 10: Check Placement Group and Cluster Health
Back in the toolbox, confirm everything is healthy and the placement groups are active+clean:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph pg stat
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph dfceph pg stat should show all PGs active+clean. ceph df shows raw vs usable capacity — remember 3x replication means usable space is roughly one-third of raw. If you see HEALTH_WARN about pool size or too few PGs, that's the expected shape on a minimal cluster; see Common Issues.
Common Issues
- No OSD pods are created — the top cause is that your disks aren't actually raw. Re-run
lsblk -fon each node; anyFSTYPEor existing partition means Rook skips the device. Wipe withsgdisk --zap-all+wipefs -a(Step 1), then delete and recreate theCephCluster. Also check thedeviceFilterregex matches your real device names. - PVC stuck in
Pending— usually the CSI provisioner can't reach a healthy cluster, or theStorageClass/pool doesn't exist yet. Confirmceph statusisHEALTH_OKin the toolbox and the pool from Step 6 is present. Full triage in fix unbound PersistentVolumeClaims. HEALTH_WARN: Degraded data redundancyor PGs stuckundersized—replicated.size: 3withfailureDomain: hostneeds three separate nodes to place three copies. On fewer nodes, either add nodes or setsize: 2(andmin_size: 1) for a lab — never do that in production.- Mon quorum never forms / mons crash-loop — check that
dataDirHostPath(/var/lib/rook) is writable and not shared between clusters. Leftover state from a previous Rook install here corrupts a fresh deploy; clean it before re-creating the cluster. - RGW pod
CrashLoopBackOff— the object store's metadata/data pools couldn't be created, again usually a node-count vs replication mismatch. Checkkubectl -n rook-ceph logs -l app=rook-ceph-rgw.
Frequently Asked Questions
Why does Rook ignore my extra disk?
Because it isn't raw. Rook (via Ceph's ceph-volume) only claims devices with no partition table, filesystem, or LVM signature. A disk that was ever formatted — even briefly — carries metadata that makes Rook skip it. Confirm with lsblk -f (the FSTYPE column must be empty), then wipe with sgdisk --zap-all /dev/sdX && wipefs -a /dev/sdX before re-applying the CephCluster.
How many nodes and disks do I actually need?
For anything resembling production, three nodes each with at least one dedicated disk, so the default 3x replication can place each copy on a different host. You can run a single-node lab with replicated.size: 1 and allowMultiplePerNode: true, but you lose all redundancy — a disk or node failure loses data. Ceph's whole value is the replication you'd be turning off.
What's the difference between the CephBlockPool and the StorageClass?
The CephBlockPool is a Ceph-side construct: a RADOS pool with a replication policy that defines how data is stored. The StorageClass is the Kubernetes-side binding that points the Ceph CSI RBD driver at that pool so PVCs can provision volumes from it. You need both — the pool holds the data, the StorageClass exposes it to Kubernetes.
Can I run block, file, and object storage from one Ceph cluster?
Yes — that's Rook-Ceph's main advantage over lighter tools. The same CephCluster backs RBD block volumes (this tutorial), a CephFilesystem for ReadWriteMany shared file storage, and a CephObjectStore for S3. One storage layer, three access modes. If you only ever need ReadWriteOnce block volumes, that flexibility is overkill and a simpler tool may serve you better.
Is it safe to delete the CephCluster to start over?
Deleting the CephCluster resource removes the daemons but, by design, does not wipe your disks or the dataDirHostPath — Rook guards against accidental data loss. To truly start clean you must set the cleanup confirmation annotation (see Tear Down) or manually wipe the disks and /var/lib/rook on each node. Skipping that is the most common reason a "fresh" reinstall fails.
Tear Down
1# Remove workloads and storage abstractions first
2kubectl delete pod ceph-block-writer s3-test --ignore-not-found
3kubectl delete pvc ceph-block-test --ignore-not-found
4kubectl delete objectbucketclaim my-bucket --ignore-not-found
5kubectl delete cephobjectstore my-store -n rook-ceph --ignore-not-found
6kubectl delete cephblockpool replicapool -n rook-ceph --ignore-not-found
7kubectl delete storageclass rook-ceph-block rook-ceph-bucket --ignore-not-found
8
9# Tell Rook it's allowed to wipe the disks, then delete the cluster
10kubectl -n rook-ceph patch cephcluster rook-ceph --type merge \
11 -p '{"spec":{"cleanupPolicy":{"confirmation":"yes-really-destroy-data"}}}'
12kubectl -n rook-ceph delete cephcluster rook-ceph
13
14# Remove the toolbox and the operator
15kubectl -n rook-ceph delete deploy rook-ceph-tools --ignore-not-found
16helm uninstall rook-ceph -n rook-ceph
17kubectl delete namespace rook-cephThe cleanupPolicy annotation launches a job that zaps the OSD disks so they're raw again. If you skip it, manually run sgdisk --zap-all + wipefs -a on each device and delete /var/lib/rook on every node before any future Rook install — otherwise leftover state will break the next deploy.
Official References
- Rook Documentation — Ceph Quickstart — Official install and cluster-bring-up workflow this builds on
- Rook — Helm Operator Chart — Values reference for the operator chart
- Rook — Block Storage (RBD) — CephBlockPool + StorageClass details
- Rook — Object Storage — CephObjectStore and ObjectBucketClaim
- Ceph Documentation — Upstream Ceph concepts: OSDs, mons, PGs, CRUSH, and pools
We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.
Struggling with this in production?
We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.