Kubernetes

Deploy Rook-Ceph on Kubernetes: Block and Object Storage from Scratch

Advanced90 min to complete13 min readJuly 2, 2026

Quick answer

Turn a handful of raw disks into a self-healing storage cluster. Install the Rook operator with Helm, stand up a CephCluster, verify OSDs with the toolbox, then serve both RBD block volumes and S3-compatible object storage — every command runnable, step by step.

advanced · 90 min

Before you begin

  • A multi-node Kubernetes cluster (3+ worker nodes) with cluster-admin access
  • At least one unformatted raw block device (extra disk) per storage node — no partitions, no filesystem, no LVM
  • kubectl configured for the cluster
  • Helm v3 installed
  • Comfortable with PVs, PVCs, and StorageClasses
Kubernetes
Rook
Ceph
Storage
CSI
Object Storage
Platform Engineering

Most people meet Ceph as a black box behind a managed service and never see how the pieces fit. Rook changes that: it's a Kubernetes operator that runs a full Ceph cluster inside your cluster, turning raw disks attached to your nodes into replicated block, file, and object storage — all driven by custom resources instead of the sprawling Ceph CLI. You write YAML; the operator handles mons, OSDs, managers, and the CSI drivers.

In this tutorial you'll build that from nothing: install the Rook operator with Helm, hand it some raw disks via a CephCluster resource, confirm the OSDs came up with the Ceph toolbox, then carve out two kinds of storage on top — an RBD-backed StorageClass for ReadWriteOnce block volumes, and a CephObjectStore with an S3 bucket you can actually put objects into. By the end you'll have a self-healing storage layer that other workloads can consume like any cloud provider's disks.

If you're still deciding whether Ceph is the right engine for your cluster versus lighter options, read Rook-Ceph vs Longhorn vs OpenEBS first — Rook-Ceph is the heavyweight: most capable, most resource-hungry, and the right pick when you need block, file, and object storage from one system.

What You'll Build

  • The Rook operator installed with Helm into the rook-ceph namespace
  • A CephCluster that consumes raw devices across your nodes as OSDs
  • The Rook toolbox pod for running ceph status and ceph osd status
  • A CephBlockPool + StorageClass (rook-ceph.rbd.csi.ceph.com) serving RBD block volumes
  • A test PVC + pod that mounts a Ceph-backed volume
  • A CephObjectStore + ObjectBucketClaim exposing an S3-compatible endpoint you can write to

Step 1: Confirm You Have Raw Devices

This is the step people skip, and it's the one that breaks everything. Rook creates OSDs only from raw, unformatted block devices. A disk with an existing partition table, filesystem, or LVM signature is silently ignored.

On each storage node, check what's available:

bash
lsblk -f

You want a device (e.g. /dev/sdb or /dev/nvme1n1) whose FSTYPE column is empty and that has no child partitions. If your extra disk already has leftover data from a previous install, wipe it (destructive — be sure of the device name):

bash
sudo sgdisk --zap-all /dev/sdb
sudo wipefs -a /dev/sdb

Note the device name — it's usually consistent across nodes (/dev/sdb), which lets you use a deviceFilter later.

Step 2: Install the Rook Operator with Helm

Add the Rook chart repo and install the operator. The operator itself stores no data — it's the controller that watches your Ceph custom resources and reconciles the actual daemons.

bash
1helm repo add rook-release https://charts.rook.io/release
2helm repo update
3
4helm install rook-ceph rook-release/rook-ceph \
5  --namespace rook-ceph \
6  --create-namespace \
7  --version v1.20.7

Wait for the operator to be running:

bash
kubectl -n rook-ceph rollout status deploy/rook-ceph-operator
kubectl -n rook-ceph get pods

You should see a single rook-ceph-operator-* pod in Running. No Ceph daemons exist yet — the operator is idle until you give it a CephCluster.

Step 3: Apply the CephCluster

This is the resource that tells Rook to build a Ceph cluster. useAllNodes: true schedules OSDs across every eligible node; useAllDevices: false plus a deviceFilter keeps Rook off your OS disk and targets only the extra disks.

yaml
1apiVersion: ceph.rook.io/v1
2kind: CephCluster
3metadata:
4  name: rook-ceph
5  namespace: rook-ceph
6spec:
7  cephVersion:
8    image: quay.io/ceph/ceph:v20.2.1
9  dataDirHostPath: /var/lib/rook
10  mon:
11    count: 3
12    allowMultiplePerNode: false
13  mgr:
14    count: 2
15  dashboard:
16    enabled: true
17  storage:
18    useAllNodes: true
19    useAllDevices: false
20    deviceFilter: "^sdb"

Save it as cluster.yaml and apply:

bash
kubectl apply -f cluster.yaml

The deviceFilter: "^sdb" is a regex matching device names — adjust it to whatever lsblk showed you (e.g. ^nvme1n1). If your disk names differ per node, set useAllNodes: false and provide an explicit per-node nodes: list instead (Rook ignores the nodes: list while useAllNodes: true, so you must flip it to false for per-node config to take effect).

Step 4: Watch the Cluster Come Up

This takes several minutes. The operator provisions mons first, then the manager, then discovers devices and creates one OSD per raw disk.

bash
kubectl -n rook-ceph get pods -w

You're waiting for this rough shape:

bash
# 3 mons, 2 mgrs, and one OSD per disk
kubectl -n rook-ceph get pods -l app=rook-ceph-mon
kubectl -n rook-ceph get pods -l app=rook-ceph-osd

The key signal is rook-ceph-osd-0, rook-ceph-osd-1, ... reaching Running. If no OSD pods appear after ~10 minutes, jump to Common Issues — it's almost always a device-detection problem.

Check the cluster resource's own view of health:

bash
kubectl -n rook-ceph get cephcluster rook-ceph

The PHASE should progress to Ready and HEALTH to HEALTH_OK (or HEALTH_WARN on a fresh 3-node cluster — that's fine while it settles).

Step 5: Install and Use the Toolbox

The toolbox is a pod with the Ceph CLI baked in, wired to your cluster's credentials. It's how you inspect Ceph directly.

bash
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.20.1/deploy/examples/toolbox.yaml
kubectl -n rook-ceph rollout status deploy/rook-ceph-tools

Now open a shell and check status:

bash
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- bash

Inside the toolbox:

bash
ceph status
ceph osd status
ceph osd tree

ceph status should report health: HEALTH_OK, your mon quorum (3 daemons), and OSDs up and in. ceph osd status lists each OSD with its capacity and node. This is your ground truth — if Ceph is healthy here, the Kubernetes layer will work.

Step 6: Create a Block Pool and StorageClass

Now the payoff: expose Ceph RBD as a Kubernetes StorageClass. First a CephBlockPool (a replicated pool with 3 copies), then a StorageClass that points the Ceph CSI RBD driver at it.

yaml
1apiVersion: ceph.rook.io/v1
2kind: CephBlockPool
3metadata:
4  name: replicapool
5  namespace: rook-ceph
6spec:
7  failureDomain: host
8  replicated:
9    size: 3
10---
11apiVersion: storage.k8s.io/v1
12kind: StorageClass
13metadata:
14  name: rook-ceph-block
15provisioner: rook-ceph.rbd.csi.ceph.com
16parameters:
17  clusterID: rook-ceph
18  pool: replicapool
19  imageFormat: "2"
20  imageFeatures: layering
21  csi.storage.k8s.io/fstype: ext4
22  # CSI credential secrets — REQUIRED, or PVCs hang in Pending and pods fail to mount.
23  # These secrets are created by the Rook operator in the operator namespace.
24  csi.storage.k8s.io/provisioner-secret-name: rook-csi-rbd-provisioner
25  csi.storage.k8s.io/provisioner-secret-namespace: rook-ceph
26  csi.storage.k8s.io/controller-expand-secret-name: rook-csi-rbd-provisioner
27  csi.storage.k8s.io/controller-expand-secret-namespace: rook-ceph
28  csi.storage.k8s.io/node-stage-secret-name: rook-csi-rbd-node
29  csi.storage.k8s.io/node-stage-secret-namespace: rook-ceph
30reclaimPolicy: Delete
31allowVolumeExpansion: true

Save as blockpool.yaml and apply:

bash
kubectl apply -f blockpool.yaml
kubectl get storageclass

replicated.size: 3 with failureDomain: host means Ceph keeps three copies of every object, each on a different node — which is exactly why you need 3+ nodes. allowVolumeExpansion: true lets you grow PVCs later by editing their requested size.

Step 7: Provision a PVC and Mount It

Prove the StorageClass works end to end with a PVC and a pod that writes to it.

yaml
1apiVersion: v1
2kind: PersistentVolumeClaim
3metadata:
4  name: ceph-block-test
5spec:
6  accessModes:
7    - ReadWriteOnce
8  storageClassName: rook-ceph-block
9  resources:
10    requests:
11      storage: 1Gi
12---
13apiVersion: v1
14kind: Pod
15metadata:
16  name: ceph-block-writer
17spec:
18  containers:
19    - name: writer
20      image: busybox:1.36
21      command: ["sh", "-c", "echo rook-ceph-works > /data/hello.txt && sleep 3600"]
22      volumeMounts:
23        - name: vol
24          mountPath: /data
25  volumes:
26    - name: vol
27      persistentVolumeClaim:
28        claimName: ceph-block-test
bash
kubectl apply -f pvc-test.yaml
kubectl get pvc ceph-block-test         # should reach Bound
kubectl get pod ceph-block-writer       # should reach Running
kubectl exec ceph-block-writer -- cat /data/hello.txt

A Bound PVC and rook-ceph-works printed back means the CSI driver provisioned an RBD image, mapped it, formatted it ext4, and mounted it into the pod. If the PVC hangs in Pending, see fix unbound PersistentVolumeClaims.

Step 8: Stand Up S3-Compatible Object Storage

Block storage is one half. The other reason people run Ceph is a self-hosted S3. A CephObjectStore deploys RADOS Gateway (RGW); an ObjectBucketClaim (OBC) then provisions a bucket and hands you credentials in a Secret.

yaml
1apiVersion: ceph.rook.io/v1
2kind: CephObjectStore
3metadata:
4  name: my-store
5  namespace: rook-ceph
6spec:
7  metadataPool:
8    failureDomain: host
9    replicated:
10      size: 3
11  dataPool:
12    failureDomain: host
13    replicated:
14      size: 3
15  gateway:
16    port: 80
17    instances: 1
bash
kubectl apply -f objectstore.yaml
kubectl -n rook-ceph get pods -l app=rook-ceph-rgw

Once the RGW pod is Running, create a StorageClass for buckets and an ObjectBucketClaim:

yaml
1apiVersion: storage.k8s.io/v1
2kind: StorageClass
3metadata:
4  name: rook-ceph-bucket
5provisioner: rook-ceph.ceph.rook.io/bucket
6reclaimPolicy: Delete
7parameters:
8  objectStoreName: my-store
9  objectStoreNamespace: rook-ceph
10---
11apiVersion: objectbucket.io/v1alpha1
12kind: ObjectBucketClaim
13metadata:
14  name: my-bucket
15spec:
16  generateBucketName: rook-demo
17  storageClassName: rook-ceph-bucket
bash
kubectl apply -f obc.yaml
kubectl get objectbucketclaim my-bucket    # PHASE should be Bound

The OBC creates a ConfigMap (bucket name, host) and a Secret (access + secret keys), both named my-bucket.

Step 9: Write to the Bucket with an S3 Client

Pull the connection details the OBC generated and test with the AWS CLI from a throwaway pod.

bash
export AWS_ACCESS_KEY_ID=$(kubectl get secret my-bucket -o jsonpath='{.data.AWS_ACCESS_KEY_ID}' | base64 -d)
export AWS_SECRET_ACCESS_KEY=$(kubectl get secret my-bucket -o jsonpath='{.data.AWS_SECRET_ACCESS_KEY}' | base64 -d)
export BUCKET_NAME=$(kubectl get cm my-bucket -o jsonpath='{.data.BUCKET_NAME}')

The RGW service inside the cluster is rook-ceph-rgw-my-store.rook-ceph.svc. Run the CLI from a pod so you can reach it:

bash
1kubectl run s3-test --rm -it --restart=Never \
2  --image=amazon/aws-cli:2.15.30 \
3  --env="AWS_ACCESS_KEY_ID=$AWS_ACCESS_KEY_ID" \
4  --env="AWS_SECRET_ACCESS_KEY=$AWS_SECRET_ACCESS_KEY" \
5  --command -- sh -c "
6    echo 'hello from ceph rgw' > /tmp/obj.txt &&
7    aws --endpoint-url http://rook-ceph-rgw-my-store.rook-ceph.svc:80 s3 cp /tmp/obj.txt s3://$BUCKET_NAME/obj.txt &&
8    aws --endpoint-url http://rook-ceph-rgw-my-store.rook-ceph.svc:80 s3 ls s3://$BUCKET_NAME/
9  "

Seeing obj.txt in the listing means the full object path works: OBC provisioned the bucket, RGW served the S3 API, and your credentials authenticated.

Step 10: Check Placement Group and Cluster Health

Back in the toolbox, confirm everything is healthy and the placement groups are active+clean:

bash
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph pg stat
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph df

ceph pg stat should show all PGs active+clean. ceph df shows raw vs usable capacity — remember 3x replication means usable space is roughly one-third of raw. If you see HEALTH_WARN about pool size or too few PGs, that's the expected shape on a minimal cluster; see Common Issues.

Common Issues

  • No OSD pods are created — the top cause is that your disks aren't actually raw. Re-run lsblk -f on each node; any FSTYPE or existing partition means Rook skips the device. Wipe with sgdisk --zap-all + wipefs -a (Step 1), then delete and recreate the CephCluster. Also check the deviceFilter regex matches your real device names.
  • PVC stuck in Pending — usually the CSI provisioner can't reach a healthy cluster, or the StorageClass/pool doesn't exist yet. Confirm ceph status is HEALTH_OK in the toolbox and the pool from Step 6 is present. Full triage in fix unbound PersistentVolumeClaims.
  • HEALTH_WARN: Degraded data redundancy or PGs stuck undersized — replicated.size: 3 with failureDomain: host needs three separate nodes to place three copies. On fewer nodes, either add nodes or set size: 2 (and min_size: 1) for a lab — never do that in production.
  • Mon quorum never forms / mons crash-loop — check that dataDirHostPath (/var/lib/rook) is writable and not shared between clusters. Leftover state from a previous Rook install here corrupts a fresh deploy; clean it before re-creating the cluster.
  • RGW pod CrashLoopBackOff — the object store's metadata/data pools couldn't be created, again usually a node-count vs replication mismatch. Check kubectl -n rook-ceph logs -l app=rook-ceph-rgw.

Frequently Asked Questions

Why does Rook ignore my extra disk?

Because it isn't raw. Rook (via Ceph's ceph-volume) only claims devices with no partition table, filesystem, or LVM signature. A disk that was ever formatted — even briefly — carries metadata that makes Rook skip it. Confirm with lsblk -f (the FSTYPE column must be empty), then wipe with sgdisk --zap-all /dev/sdX && wipefs -a /dev/sdX before re-applying the CephCluster.

How many nodes and disks do I actually need?

For anything resembling production, three nodes each with at least one dedicated disk, so the default 3x replication can place each copy on a different host. You can run a single-node lab with replicated.size: 1 and allowMultiplePerNode: true, but you lose all redundancy — a disk or node failure loses data. Ceph's whole value is the replication you'd be turning off.

What's the difference between the CephBlockPool and the StorageClass?

The CephBlockPool is a Ceph-side construct: a RADOS pool with a replication policy that defines how data is stored. The StorageClass is the Kubernetes-side binding that points the Ceph CSI RBD driver at that pool so PVCs can provision volumes from it. You need both — the pool holds the data, the StorageClass exposes it to Kubernetes.

Can I run block, file, and object storage from one Ceph cluster?

Yes — that's Rook-Ceph's main advantage over lighter tools. The same CephCluster backs RBD block volumes (this tutorial), a CephFilesystem for ReadWriteMany shared file storage, and a CephObjectStore for S3. One storage layer, three access modes. If you only ever need ReadWriteOnce block volumes, that flexibility is overkill and a simpler tool may serve you better.

Is it safe to delete the CephCluster to start over?

Deleting the CephCluster resource removes the daemons but, by design, does not wipe your disks or the dataDirHostPath — Rook guards against accidental data loss. To truly start clean you must set the cleanup confirmation annotation (see Tear Down) or manually wipe the disks and /var/lib/rook on each node. Skipping that is the most common reason a "fresh" reinstall fails.

Tear Down

bash
1# Remove workloads and storage abstractions first
2kubectl delete pod ceph-block-writer s3-test --ignore-not-found
3kubectl delete pvc ceph-block-test --ignore-not-found
4kubectl delete objectbucketclaim my-bucket --ignore-not-found
5kubectl delete cephobjectstore my-store -n rook-ceph --ignore-not-found
6kubectl delete cephblockpool replicapool -n rook-ceph --ignore-not-found
7kubectl delete storageclass rook-ceph-block rook-ceph-bucket --ignore-not-found
8
9# Tell Rook it's allowed to wipe the disks, then delete the cluster
10kubectl -n rook-ceph patch cephcluster rook-ceph --type merge \
11  -p '{"spec":{"cleanupPolicy":{"confirmation":"yes-really-destroy-data"}}}'
12kubectl -n rook-ceph delete cephcluster rook-ceph
13
14# Remove the toolbox and the operator
15kubectl -n rook-ceph delete deploy rook-ceph-tools --ignore-not-found
16helm uninstall rook-ceph -n rook-ceph
17kubectl delete namespace rook-ceph

The cleanupPolicy annotation launches a job that zaps the OSD disks so they're raw again. If you skip it, manually run sgdisk --zap-all + wipefs -a on each device and delete /var/lib/rook on every node before any future Rook install — otherwise leftover state will break the next deploy.

Official References

We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.

Struggling with this in production?

We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.