Install Longhorn on Kubernetes with S3 Backups and Disaster Recovery
Quick answer
Stand up Longhorn as a distributed block-storage layer on your cluster, provision a replicated PVC, wire an S3 backup target, take snapshots and scheduled backups with RecurringJobs, and restore a volume — the full backup-and-DR loop, end to end.
- Step 1: Install the Node Prerequisites
- Step 2: Install Longhorn with Helm
- Step 3: Access the Longhorn UI
- Step 4: Create a StorageClass and Provision a PVC
- Step 5: Configure the S3 Backup Target
intermediate · 60 min
Before you begin
- A multi-node cluster (3+ nodes so replicas can spread across failure domains)
- kubectl configured and Helm 3 installed
- open-iscsi installed and running on every node
- An S3-compatible bucket plus access key and secret (AWS S3 or self-hosted MinIO)
Longhorn is a cloud-native distributed block storage system for Kubernetes, originally built by Rancher and now a CNCF project. It turns the local disks on your nodes into replicated, synchronously-mirrored volumes, and — the part most people actually adopt it for — it bakes snapshots and off-cluster backups to S3 straight into the storage layer. No external SAN, no separate backup operator for the volumes themselves.
If you're still deciding which storage layer to run, I compared the trade-offs in Rook/Ceph vs Longhorn vs OpenEBS. The short version: Longhorn is the pragmatic choice when you want replicated block storage with a friendly UI and built-in backup/restore, and you don't need Ceph's object/file storage breadth. This tutorial takes that decision as made and walks the full operational loop — install, provision, back up, and restore.
By the end you'll have a working StorageClass, a stateful workload writing to a Longhorn volume, an S3 backup target, scheduled backups running on a cron, and a volume you've restored from a backup. That last step is the one that matters: a backup you've never restored is a hope, not a plan.
What You'll Build
- Longhorn installed into
longhorn-systemvia Helm, with the environment prerequisites verified - A replicated PVC (default 3 synchronous replicas) backing a test workload
- An S3 backup target configured with an AWS-key Secret
- A snapshot (in-cluster, fast) and a backup (off-cluster, to S3)
- A RecurringJob that snapshots and backs up on a schedule
- A restore of that backup into a new volume — the disaster-recovery payoff
Step 1: Install the Node Prerequisites
Longhorn attaches volumes over iSCSI, so every node that will run Longhorn-backed pods needs open-iscsi installed and the iscsid service running. On Debian/Ubuntu nodes:
sudo apt-get update && sudo apt-get install -y open-iscsi
sudo systemctl enable --now iscsidOn RHEL/Fedora derivatives it's iscsi-initiator-utils. Longhorn also needs a filesystem that supports extended attributes (ext4 or xfs) on the data path, and NFSv4 client support on nodes only if you plan to use ReadWriteMany volumes (more on that later).
Rather than eyeball every node, use Longhorn's CLI, longhornctl, to check the whole cluster. The old scripts/environment_check.sh was retired — modern Longhorn ships longhornctl (the longhorn/cli release), and its check preflight command flags missing packages, kernel modules, and multipath conflicts across every node:
1# Download longhornctl for your platform (Linux amd64 shown) and install it
2curl -sSfL -o longhornctl \
3 https://github.com/longhorn/cli/releases/download/v1.12.0/longhornctl-linux-amd64
4sudo install longhornctl /usr/local/bin/longhornctl
5
6# Run the preflight check against your cluster (add --enable-spdk if you plan to use the V2 engine)
7longhornctl check preflightFix anything it reports as an error before moving on — a missing open-iscsi on even one node causes volumes scheduled there to fail to attach. (longhornctl install preflight can install the missing dependencies for you if you'd rather not touch each node by hand.)
Step 2: Install Longhorn with Helm
Add the chart repo and install into a dedicated longhorn-system namespace. The namespace name is a convention Longhorn's own tooling assumes, so keep it.
1helm repo add longhorn https://charts.longhorn.io
2helm repo update
3
4helm install longhorn longhorn/longhorn \
5 --namespace longhorn-system \
6 --create-namespace \
7 --version 1.12.1Watch the components come up. Longhorn runs a manager DaemonSet (one pod per node), a driver deployer, CSI plugins, and the UI:
kubectl -n longhorn-system get pods --watchGive it a couple of minutes. When longhorn-manager, longhorn-driver-deployer, the csi-* pods, and longhorn-ui are all Running, the install is healthy. Confirm the CSI driver and the default StorageClass registered:
kubectl get storageclass
# longhorn (default) driver.longhorn.io Delete Immediate ...The chart ships a longhorn StorageClass and marks it the cluster default. That's usually what you want, but I'll show a custom one next so the parameters are explicit rather than magic.
Step 3: Access the Longhorn UI
The UI is where snapshots, backups, and volume health are easiest to reason about. For a quick look, port-forward it:
kubectl -n longhorn-system port-forward svc/longhorn-frontend 8080:80Open http://localhost:8080. You'll see the dashboard with your nodes, their available scheduling capacity, and (once you create volumes) each volume's replica placement. For anything beyond a lab, expose it behind an Ingress with authentication — the UI has no auth of its own and grants full control over your storage, so never expose it unprotected.
Step 4: Create a StorageClass and Provision a PVC
You can use the default longhorn class, but defining your own makes the replica count and reclaim behavior explicit. Save this as longhorn-sc.yaml:
1apiVersion: storage.k8s.io/v1
2kind: StorageClass
3metadata:
4 name: longhorn-replicated
5provisioner: driver.longhorn.io
6allowVolumeExpansion: true
7reclaimPolicy: Delete
8volumeBindingMode: Immediate
9parameters:
10 numberOfReplicas: "3"
11 staleReplicaTimeout: "30"
12 fsType: "ext4"numberOfReplicas: "3" is the important line: Longhorn keeps three synchronous copies of every volume, ideally on three different nodes, so a single node loss doesn't lose data. (If StorageClasses, reclaim policies, and binding modes are fuzzy, my PVCs & StorageClasses tutorial covers the fundamentals.)
Apply it and provision a PVC plus a workload that writes to it:
1apiVersion: v1
2kind: PersistentVolumeClaim
3metadata:
4 name: data-vol
5spec:
6 accessModes: ["ReadWriteOnce"]
7 storageClassName: longhorn-replicated
8 resources:
9 requests:
10 storage: 2Gi
11---
12apiVersion: apps/v1
13kind: Deployment
14metadata:
15 name: writer
16spec:
17 replicas: 1
18 selector:
19 matchLabels: { app: writer }
20 template:
21 metadata:
22 labels: { app: writer }
23 spec:
24 containers:
25 - name: writer
26 image: busybox:1.36
27 command: ["sh", "-c", "while true; do date >> /data/log.txt; sleep 5; done"]
28 volumeMounts:
29 - name: data
30 mountPath: /data
31 volumes:
32 - name: data
33 persistentVolumeClaim:
34 claimName: data-volkubectl apply -f longhorn-sc.yaml
kubectl apply -f writer.yaml
kubectl get pvc data-vol # should reach Bound
kubectl exec deploy/writer -- tail -n 3 /data/log.txtNote the ReadWriteOnce access mode — Longhorn volumes are block devices, so they're RWO (one node mounts them at a time) by default. ReadWriteMany needs the NFS-based share-manager, covered in the FAQ.
Step 5: Configure the S3 Backup Target
Backups go off-cluster to S3. Longhorn needs two things: the backup target URL and a Secret holding the credentials. Create the Secret in longhorn-system — the key names are Longhorn-specific and must match exactly:
kubectl -n longhorn-system create secret generic aws-backup-secret \
--from-literal=AWS_ACCESS_KEY_ID='<your-access-key>' \
--from-literal=AWS_SECRET_ACCESS_KEY='<your-secret-key>'For self-hosted MinIO, add --from-literal=AWS_ENDPOINTS='https://minio.example.com:9000' to that Secret.
Now point Longhorn at the bucket. The URL format is s3://<bucket>@<region>/<optional-prefix> — the @region part is required and a common source of "invalid backup target" errors.
Heads up — the old way is gone. Through Longhorn 1.7 you set this by patching the settings.longhorn.io resources backup-target and backup-target-credential-secret. Longhorn 1.8.0 introduced multiple backupstores and made the backup target a first-class BackupTarget custom resource; a default BackupTarget is created automatically on install and upgrade, and backup-target no longer appears in the settings reference. Configure the default target declaratively instead:
1apiVersion: longhorn.io/v1beta2
2kind: BackupTarget
3metadata:
4 name: default
5 namespace: longhorn-system
6spec:
7 backupTargetURL: "s3://my-longhorn-backups@us-east-1/"
8 credentialSecret: "aws-backup-secret"
9 pollInterval: "5m0s"kubectl apply -f backuptarget.yaml
kubectl -n longhorn-system get backuptarget defaultYou can reach the same result three other ways, all equivalent:
- At install time via Helm, set
defaultBackupStore.backupTarget,defaultBackupStore.backupTargetCredentialSecret, anddefaultBackupStore.pollIntervalin your values. - Declaratively via a ConfigMap named
longhorn-default-resourceinlonghorn-system(keysbackup-target,backup-target-credential-secret,backupstore-poll-interval) — handy for GitOps. - In the UI, under Backup and Restore → Backup Targets, create or edit a target with the S3
URLandCredential Secret.
Whichever you pick, open the UI's Backup and Restore → Backup Targets view to confirm the default target reports as available (green, no error). If it shows a connection error, re-check the @region, the bucket name, and the Secret's key spelling — those three cover the vast majority of failures.
Step 6: Take a Snapshot and a Backup
A snapshot is an in-cluster, point-in-time copy of the volume's state — fast, local, and the building block Longhorn backs up from. A backup is that snapshot pushed to S3, so it survives losing the whole cluster.
In the UI, go to Volume, click your data-vol volume, and use Take Snapshot. Then click that snapshot and choose Backup to push it to S3. Watch it appear under the Backup tab once the upload completes.
You can also drive this from the CLI. Longhorn exposes volumes and backups as CRDs, so you can create a Backup custom resource, but the cleanest scriptable path is the snapshot-then-backup CRD pair or the RecurringJob in the next step. For a one-off manual backup, the UI is genuinely the fastest route — this is one place the graphical tool earns its keep.
Verify the backup landed by listing objects in your bucket:
aws s3 ls s3://my-longhorn-backups/backupstore/ --recursive | headYou'll see Longhorn's backupstore layout with volume and backup metadata blocks.
Step 7: Schedule Backups with a RecurringJob
Manual backups don't survive contact with reality. A RecurringJob runs snapshots and backups on a cron schedule and prunes old ones via a retain count. Save as recurring-backup.yaml:
1apiVersion: longhorn.io/v1beta2
2kind: RecurringJob
3metadata:
4 name: daily-backup
5 namespace: longhorn-system
6spec:
7 cron: "0 2 * * *" # 02:00 every day
8 task: "backup" # "snapshot" for in-cluster only; "backup" pushes to S3
9 groups:
10 - default # applies to all volumes in the "default" recurring-job group
11 retain: 7 # keep the 7 most recent backups
12 concurrency: 2kubectl apply -f recurring-backup.yaml
kubectl -n longhorn-system get recurringjobsThe groups: [default] line means this job applies to every volume that belongs to the default recurring-job group — and Longhorn adds new volumes to default automatically unless you opt out. To bind a specific volume to a named group instead, add a recurring-job-group.longhorn.io/<group>: enabled label to the volume. A common pattern is a frequent snapshot job (hourly, cheap, local) plus a daily backup job (off-cluster, for DR).
Step 8: Restore a Backup (Disaster Recovery)
Here's the payoff. Simulate data loss and recover from S3.
First, delete the workload and its volume to prove the restore is real, not the original data:
kubectl delete deploy writer
kubectl delete pvc data-vol # with reclaimPolicy Delete, the Longhorn volume is removedNow restore. In the UI, go to the Backup tab, select the backup of data-vol, and click Restore. Give the restored volume a name (e.g. data-vol-restored). Longhorn pulls the data back from S3 and creates a new volume.
To use it from a pod, create a PV/PVC pair pointing at the restored Longhorn volume. The simplest approach: from the restored volume in the UI, use Create PV/PVC, which generates both objects for you. Then mount that PVC and confirm your data survived:
kubectl exec deploy/writer-restored -- tail -n 5 /data/log.txt
# the timestamps from before the "disaster" are still thereFor a true cross-cluster DR posture, Longhorn also supports DR volumes: in a second cluster pointed at the same S3 backup target, you create a volume from a backup with the DR flag set. It stays in a standby state, continuously incrementally restoring from the latest backups, and can be activated (promoted to a normal read-write volume) in minutes if the primary cluster dies. That's the model for real regional failover — one active cluster writing backups, a standby cluster warming DR volumes from them.
Common Issues
- Volume stuck in
Attaching/ pod can't mount —open-iscsiis missing oriscsidisn't running on the node where the pod landed. Install it and restart the service; re-runlonghornctl check preflightfrom Step 1. invalid backup targetin the UI — almost always the URL format. It must bes3://bucket@region/, with the@regionpresent, and the credential Secret name must be set in thedefaultBackupTarget'sspec.credentialSecret.- Backup fails with access denied — the Secret key names must be exactly
AWS_ACCESS_KEY_IDandAWS_SECRET_ACCESS_KEY(andAWS_ENDPOINTSfor MinIO). A typo here fails silently as an auth error. - Only 1 replica scheduled, "Degraded" volume — on a single-node cluster (or one with taints), Longhorn can't spread three replicas across distinct nodes. You need 3+ schedulable nodes, or lower
numberOfReplicasfor a lab. - RWX PVC won't bind — ReadWriteMany requires the NFS-based share-manager and NFSv4 client packages on nodes; plain block volumes are RWO only.
Frequently Asked Questions
How many replicas does Longhorn keep, and are they synchronous?
By default Longhorn keeps 3 replicas per volume and writes to them synchronously — every write is acknowledged only after it lands on all healthy replicas, so there's no data loss on a single-node failure. You set the count per StorageClass with numberOfReplicas, and Longhorn tries to place each replica on a distinct node and disk for real fault isolation.
What's the difference between a Longhorn snapshot and a backup?
A snapshot is an in-cluster, point-in-time copy of the volume stored on the same nodes — fast to take and restore, but it dies with the cluster. A backup is a snapshot exported to your S3 (or NFS) backup target, so it survives losing the entire cluster and is what disaster recovery relies on. Backups are incremental after the first full one, uploading only changed blocks.
Can Longhorn provide ReadWriteMany volumes?
Yes, but not natively as block storage. Longhorn volumes are block devices and therefore ReadWriteOnce by default. For ReadWriteMany, Longhorn provisions an NFS-based share-manager pod that fronts the volume, and your nodes need NFSv4 client support. It works well for shared-config use cases but adds an NFS hop, so don't reach for RWX unless the workload genuinely needs concurrent multi-node access.
Should I use Longhorn's built-in backups or a tool like Velero?
They solve different layers and pair well. Longhorn backs up the volume data (block-level, incremental, to S3). Velero backs up Kubernetes resources — Deployments, Services, ConfigMaps, namespaces — and can orchestrate volume snapshots via CSI. For full application recovery you often want both: Velero to recreate the objects and Longhorn to restore the data. I cover the resource side in Velero backup and disaster recovery.
Is the V2 (SPDK) data engine ready for production?
The V2 data engine, built on SPDK for higher performance on NVMe, reached General Availability in Longhorn v1.12.0 (2026). It's production-ready but opt-in — the V1 engine is still the default, and it's what this tutorial uses. Stick with V1 unless you specifically need V2's NVMe performance and have reviewed its feature-parity notes, since a few V1 capabilities are still catching up on V2.
Tear Down
1# Remove the test workload and volume
2kubectl delete deploy writer writer-restored --ignore-not-found
3kubectl delete pvc data-vol data-vol-restored --ignore-not-found
4
5# Longhorn blocks uninstall unless you explicitly confirm deletion,
6# which guards against wiping storage by accident:
7kubectl -n longhorn-system patch settings.longhorn.io deleting-confirmation-flag \
8 --type=merge -p '{"value":"true"}'
9
10# Uninstall the Helm release
11helm uninstall longhorn -n longhorn-system
12
13# Remove any lingering CRDs and the namespace
14kubectl delete crd -l app.kubernetes.io/name=longhorn
15kubectl delete namespace longhorn-system --ignore-not-foundDeleting Longhorn removes its volumes but does not touch your S3 backups — those remain in the bucket and can be restored into a fresh install pointed at the same backup target.
Official References
- Longhorn v1.12.0 — Quick Installation — Prerequisites and Helm/kubectl install paths
- Longhorn v1.12.0 — Command Line Tool (longhornctl) — Installing
longhornctland running thecheck preflightenvironment check - Longhorn v1.12.0 — Set Backup Target — The
defaultBackupTarget resource, S3/NFS URL formats, and credential Secret setup - Longhorn v1.12.0 — Recurring Snapshots and Backups — RecurringJob CRD, groups, cron, and retention
- Longhorn v1.12.0 — Disaster Recovery Volumes — Cross-cluster DR volumes and activation
- Longhorn v1.12.0 — V1 and V2 Volume Feature Support — Feature parity for the GA V2 (SPDK) data engine
We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.
Struggling with this in production?
We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.