Kubernetes

Install Longhorn on Kubernetes with S3 Backups and Disaster Recovery

Intermediate60 min to complete13 min readJuly 2, 2026Updated August 19, 2026

Quick answer

Stand up Longhorn as a distributed block-storage layer on your cluster, provision a replicated PVC, wire an S3 backup target, take snapshots and scheduled backups with RecurringJobs, and restore a volume — the full backup-and-DR loop, end to end.

intermediate · 60 min

Before you begin

  • A multi-node cluster (3+ nodes so replicas can spread across failure domains)
  • kubectl configured and Helm 3 installed
  • open-iscsi installed and running on every node
  • An S3-compatible bucket plus access key and secret (AWS S3 or self-hosted MinIO)
Kubernetes
Longhorn
Storage
Backups
Disaster Recovery
S3
Helm

Longhorn is a cloud-native distributed block storage system for Kubernetes, originally built by Rancher and now a CNCF project. It turns the local disks on your nodes into replicated, synchronously-mirrored volumes, and — the part most people actually adopt it for — it bakes snapshots and off-cluster backups to S3 straight into the storage layer. No external SAN, no separate backup operator for the volumes themselves.

If you're still deciding which storage layer to run, I compared the trade-offs in Rook/Ceph vs Longhorn vs OpenEBS. The short version: Longhorn is the pragmatic choice when you want replicated block storage with a friendly UI and built-in backup/restore, and you don't need Ceph's object/file storage breadth. This tutorial takes that decision as made and walks the full operational loop — install, provision, back up, and restore.

By the end you'll have a working StorageClass, a stateful workload writing to a Longhorn volume, an S3 backup target, scheduled backups running on a cron, and a volume you've restored from a backup. That last step is the one that matters: a backup you've never restored is a hope, not a plan.

What You'll Build

  • Longhorn installed into longhorn-system via Helm, with the environment prerequisites verified
  • A replicated PVC (default 3 synchronous replicas) backing a test workload
  • An S3 backup target configured with an AWS-key Secret
  • A snapshot (in-cluster, fast) and a backup (off-cluster, to S3)
  • A RecurringJob that snapshots and backs up on a schedule
  • A restore of that backup into a new volume — the disaster-recovery payoff

Step 1: Install the Node Prerequisites

Longhorn attaches volumes over iSCSI, so every node that will run Longhorn-backed pods needs open-iscsi installed and the iscsid service running. On Debian/Ubuntu nodes:

bash
sudo apt-get update && sudo apt-get install -y open-iscsi
sudo systemctl enable --now iscsid

On RHEL/Fedora derivatives it's iscsi-initiator-utils. Longhorn also needs a filesystem that supports extended attributes (ext4 or xfs) on the data path, and NFSv4 client support on nodes only if you plan to use ReadWriteMany volumes (more on that later).

Rather than eyeball every node, use Longhorn's CLI, longhornctl, to check the whole cluster. The old scripts/environment_check.sh was retired — modern Longhorn ships longhornctl (the longhorn/cli release), and its check preflight command flags missing packages, kernel modules, and multipath conflicts across every node:

bash
1# Download longhornctl for your platform (Linux amd64 shown) and install it
2curl -sSfL -o longhornctl \
3  https://github.com/longhorn/cli/releases/download/v1.12.0/longhornctl-linux-amd64
4sudo install longhornctl /usr/local/bin/longhornctl
5
6# Run the preflight check against your cluster (add --enable-spdk if you plan to use the V2 engine)
7longhornctl check preflight

Fix anything it reports as an error before moving on — a missing open-iscsi on even one node causes volumes scheduled there to fail to attach. (longhornctl install preflight can install the missing dependencies for you if you'd rather not touch each node by hand.)

Step 2: Install Longhorn with Helm

Add the chart repo and install into a dedicated longhorn-system namespace. The namespace name is a convention Longhorn's own tooling assumes, so keep it.

bash
1helm repo add longhorn https://charts.longhorn.io
2helm repo update
3
4helm install longhorn longhorn/longhorn \
5  --namespace longhorn-system \
6  --create-namespace \
7  --version 1.12.1

Watch the components come up. Longhorn runs a manager DaemonSet (one pod per node), a driver deployer, CSI plugins, and the UI:

bash
kubectl -n longhorn-system get pods --watch

Give it a couple of minutes. When longhorn-manager, longhorn-driver-deployer, the csi-* pods, and longhorn-ui are all Running, the install is healthy. Confirm the CSI driver and the default StorageClass registered:

bash
kubectl get storageclass
# longhorn (default)   driver.longhorn.io   Delete   Immediate   ...

The chart ships a longhorn StorageClass and marks it the cluster default. That's usually what you want, but I'll show a custom one next so the parameters are explicit rather than magic.

Step 3: Access the Longhorn UI

The UI is where snapshots, backups, and volume health are easiest to reason about. For a quick look, port-forward it:

bash
kubectl -n longhorn-system port-forward svc/longhorn-frontend 8080:80

Open http://localhost:8080. You'll see the dashboard with your nodes, their available scheduling capacity, and (once you create volumes) each volume's replica placement. For anything beyond a lab, expose it behind an Ingress with authentication — the UI has no auth of its own and grants full control over your storage, so never expose it unprotected.

Step 4: Create a StorageClass and Provision a PVC

You can use the default longhorn class, but defining your own makes the replica count and reclaim behavior explicit. Save this as longhorn-sc.yaml:

yaml
1apiVersion: storage.k8s.io/v1
2kind: StorageClass
3metadata:
4  name: longhorn-replicated
5provisioner: driver.longhorn.io
6allowVolumeExpansion: true
7reclaimPolicy: Delete
8volumeBindingMode: Immediate
9parameters:
10  numberOfReplicas: "3"
11  staleReplicaTimeout: "30"
12  fsType: "ext4"

numberOfReplicas: "3" is the important line: Longhorn keeps three synchronous copies of every volume, ideally on three different nodes, so a single node loss doesn't lose data. (If StorageClasses, reclaim policies, and binding modes are fuzzy, my PVCs & StorageClasses tutorial covers the fundamentals.)

Apply it and provision a PVC plus a workload that writes to it:

yaml
1apiVersion: v1
2kind: PersistentVolumeClaim
3metadata:
4  name: data-vol
5spec:
6  accessModes: ["ReadWriteOnce"]
7  storageClassName: longhorn-replicated
8  resources:
9    requests:
10      storage: 2Gi
11---
12apiVersion: apps/v1
13kind: Deployment
14metadata:
15  name: writer
16spec:
17  replicas: 1
18  selector:
19    matchLabels: { app: writer }
20  template:
21    metadata:
22      labels: { app: writer }
23    spec:
24      containers:
25        - name: writer
26          image: busybox:1.36
27          command: ["sh", "-c", "while true; do date >> /data/log.txt; sleep 5; done"]
28          volumeMounts:
29            - name: data
30              mountPath: /data
31      volumes:
32        - name: data
33          persistentVolumeClaim:
34            claimName: data-vol
bash
kubectl apply -f longhorn-sc.yaml
kubectl apply -f writer.yaml
kubectl get pvc data-vol          # should reach Bound
kubectl exec deploy/writer -- tail -n 3 /data/log.txt

Note the ReadWriteOnce access mode — Longhorn volumes are block devices, so they're RWO (one node mounts them at a time) by default. ReadWriteMany needs the NFS-based share-manager, covered in the FAQ.

Step 5: Configure the S3 Backup Target

Backups go off-cluster to S3. Longhorn needs two things: the backup target URL and a Secret holding the credentials. Create the Secret in longhorn-system — the key names are Longhorn-specific and must match exactly:

bash
kubectl -n longhorn-system create secret generic aws-backup-secret \
  --from-literal=AWS_ACCESS_KEY_ID='<your-access-key>' \
  --from-literal=AWS_SECRET_ACCESS_KEY='<your-secret-key>'

For self-hosted MinIO, add --from-literal=AWS_ENDPOINTS='https://minio.example.com:9000' to that Secret.

Now point Longhorn at the bucket. The URL format is s3://<bucket>@<region>/<optional-prefix> — the @region part is required and a common source of "invalid backup target" errors.

Heads up — the old way is gone. Through Longhorn 1.7 you set this by patching the settings.longhorn.io resources backup-target and backup-target-credential-secret. Longhorn 1.8.0 introduced multiple backupstores and made the backup target a first-class BackupTarget custom resource; a default BackupTarget is created automatically on install and upgrade, and backup-target no longer appears in the settings reference. Configure the default target declaratively instead:

yaml
1apiVersion: longhorn.io/v1beta2
2kind: BackupTarget
3metadata:
4  name: default
5  namespace: longhorn-system
6spec:
7  backupTargetURL: "s3://my-longhorn-backups@us-east-1/"
8  credentialSecret: "aws-backup-secret"
9  pollInterval: "5m0s"
bash
kubectl apply -f backuptarget.yaml
kubectl -n longhorn-system get backuptarget default

You can reach the same result three other ways, all equivalent:

  • At install time via Helm, set defaultBackupStore.backupTarget, defaultBackupStore.backupTargetCredentialSecret, and defaultBackupStore.pollInterval in your values.
  • Declaratively via a ConfigMap named longhorn-default-resource in longhorn-system (keys backup-target, backup-target-credential-secret, backupstore-poll-interval) — handy for GitOps.
  • In the UI, under Backup and Restore → Backup Targets, create or edit a target with the S3 URL and Credential Secret.

Whichever you pick, open the UI's Backup and Restore → Backup Targets view to confirm the default target reports as available (green, no error). If it shows a connection error, re-check the @region, the bucket name, and the Secret's key spelling — those three cover the vast majority of failures.

Step 6: Take a Snapshot and a Backup

A snapshot is an in-cluster, point-in-time copy of the volume's state — fast, local, and the building block Longhorn backs up from. A backup is that snapshot pushed to S3, so it survives losing the whole cluster.

In the UI, go to Volume, click your data-vol volume, and use Take Snapshot. Then click that snapshot and choose Backup to push it to S3. Watch it appear under the Backup tab once the upload completes.

You can also drive this from the CLI. Longhorn exposes volumes and backups as CRDs, so you can create a Backup custom resource, but the cleanest scriptable path is the snapshot-then-backup CRD pair or the RecurringJob in the next step. For a one-off manual backup, the UI is genuinely the fastest route — this is one place the graphical tool earns its keep.

Verify the backup landed by listing objects in your bucket:

bash
aws s3 ls s3://my-longhorn-backups/backupstore/ --recursive | head

You'll see Longhorn's backupstore layout with volume and backup metadata blocks.

Step 7: Schedule Backups with a RecurringJob

Manual backups don't survive contact with reality. A RecurringJob runs snapshots and backups on a cron schedule and prunes old ones via a retain count. Save as recurring-backup.yaml:

yaml
1apiVersion: longhorn.io/v1beta2
2kind: RecurringJob
3metadata:
4  name: daily-backup
5  namespace: longhorn-system
6spec:
7  cron: "0 2 * * *"      # 02:00 every day
8  task: "backup"          # "snapshot" for in-cluster only; "backup" pushes to S3
9  groups:
10    - default             # applies to all volumes in the "default" recurring-job group
11  retain: 7               # keep the 7 most recent backups
12  concurrency: 2
bash
kubectl apply -f recurring-backup.yaml
kubectl -n longhorn-system get recurringjobs

The groups: [default] line means this job applies to every volume that belongs to the default recurring-job group — and Longhorn adds new volumes to default automatically unless you opt out. To bind a specific volume to a named group instead, add a recurring-job-group.longhorn.io/<group>: enabled label to the volume. A common pattern is a frequent snapshot job (hourly, cheap, local) plus a daily backup job (off-cluster, for DR).

Step 8: Restore a Backup (Disaster Recovery)

Here's the payoff. Simulate data loss and recover from S3.

First, delete the workload and its volume to prove the restore is real, not the original data:

bash
kubectl delete deploy writer
kubectl delete pvc data-vol      # with reclaimPolicy Delete, the Longhorn volume is removed

Now restore. In the UI, go to the Backup tab, select the backup of data-vol, and click Restore. Give the restored volume a name (e.g. data-vol-restored). Longhorn pulls the data back from S3 and creates a new volume.

To use it from a pod, create a PV/PVC pair pointing at the restored Longhorn volume. The simplest approach: from the restored volume in the UI, use Create PV/PVC, which generates both objects for you. Then mount that PVC and confirm your data survived:

bash
kubectl exec deploy/writer-restored -- tail -n 5 /data/log.txt
# the timestamps from before the "disaster" are still there

For a true cross-cluster DR posture, Longhorn also supports DR volumes: in a second cluster pointed at the same S3 backup target, you create a volume from a backup with the DR flag set. It stays in a standby state, continuously incrementally restoring from the latest backups, and can be activated (promoted to a normal read-write volume) in minutes if the primary cluster dies. That's the model for real regional failover — one active cluster writing backups, a standby cluster warming DR volumes from them.

Common Issues

  • Volume stuck in Attaching / pod can't mount — open-iscsi is missing or iscsid isn't running on the node where the pod landed. Install it and restart the service; re-run longhornctl check preflight from Step 1.
  • invalid backup target in the UI — almost always the URL format. It must be s3://bucket@region/, with the @region present, and the credential Secret name must be set in the default BackupTarget's spec.credentialSecret.
  • Backup fails with access denied — the Secret key names must be exactly AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (and AWS_ENDPOINTS for MinIO). A typo here fails silently as an auth error.
  • Only 1 replica scheduled, "Degraded" volume — on a single-node cluster (or one with taints), Longhorn can't spread three replicas across distinct nodes. You need 3+ schedulable nodes, or lower numberOfReplicas for a lab.
  • RWX PVC won't bind — ReadWriteMany requires the NFS-based share-manager and NFSv4 client packages on nodes; plain block volumes are RWO only.

Frequently Asked Questions

How many replicas does Longhorn keep, and are they synchronous?

By default Longhorn keeps 3 replicas per volume and writes to them synchronously — every write is acknowledged only after it lands on all healthy replicas, so there's no data loss on a single-node failure. You set the count per StorageClass with numberOfReplicas, and Longhorn tries to place each replica on a distinct node and disk for real fault isolation.

What's the difference between a Longhorn snapshot and a backup?

A snapshot is an in-cluster, point-in-time copy of the volume stored on the same nodes — fast to take and restore, but it dies with the cluster. A backup is a snapshot exported to your S3 (or NFS) backup target, so it survives losing the entire cluster and is what disaster recovery relies on. Backups are incremental after the first full one, uploading only changed blocks.

Can Longhorn provide ReadWriteMany volumes?

Yes, but not natively as block storage. Longhorn volumes are block devices and therefore ReadWriteOnce by default. For ReadWriteMany, Longhorn provisions an NFS-based share-manager pod that fronts the volume, and your nodes need NFSv4 client support. It works well for shared-config use cases but adds an NFS hop, so don't reach for RWX unless the workload genuinely needs concurrent multi-node access.

Should I use Longhorn's built-in backups or a tool like Velero?

They solve different layers and pair well. Longhorn backs up the volume data (block-level, incremental, to S3). Velero backs up Kubernetes resources — Deployments, Services, ConfigMaps, namespaces — and can orchestrate volume snapshots via CSI. For full application recovery you often want both: Velero to recreate the objects and Longhorn to restore the data. I cover the resource side in Velero backup and disaster recovery.

Is the V2 (SPDK) data engine ready for production?

The V2 data engine, built on SPDK for higher performance on NVMe, reached General Availability in Longhorn v1.12.0 (2026). It's production-ready but opt-in — the V1 engine is still the default, and it's what this tutorial uses. Stick with V1 unless you specifically need V2's NVMe performance and have reviewed its feature-parity notes, since a few V1 capabilities are still catching up on V2.

Tear Down

bash
1# Remove the test workload and volume
2kubectl delete deploy writer writer-restored --ignore-not-found
3kubectl delete pvc data-vol data-vol-restored --ignore-not-found
4
5# Longhorn blocks uninstall unless you explicitly confirm deletion,
6# which guards against wiping storage by accident:
7kubectl -n longhorn-system patch settings.longhorn.io deleting-confirmation-flag \
8  --type=merge -p '{"value":"true"}'
9
10# Uninstall the Helm release
11helm uninstall longhorn -n longhorn-system
12
13# Remove any lingering CRDs and the namespace
14kubectl delete crd -l app.kubernetes.io/name=longhorn
15kubectl delete namespace longhorn-system --ignore-not-found

Deleting Longhorn removes its volumes but does not touch your S3 backups — those remain in the bucket and can be restored into a fresh install pointed at the same backup target.

Official References

We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.

Struggling with this in production?

We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.