Kubernetes
12 min readOctober 7, 2026

etcd Backup and Restore: Operating Kubernetes' Source of Truth

Part ofKubernetes
CO
Coding Protocols Team
Platform Engineering
etcd Backup and Restore: Operating Kubernetes' Source of Truth

Quick answer

Every Deployment, Secret, and ConfigMap your cluster knows about lives in etcd — not in the apiserver, not in etcd's cache, in etcd itself. Velero backs up what the API server can see; it does not back up etcd. Here's how snapshot save/restore actually works, who controls etcd on your cluster, and why these two backup strategies protect completely different failure modes.

12 min read · Kubernetes

Ask most platform engineers what etcd is for and you'll get "it's Kubernetes' database." True, but it undersells the point: etcd is not a place Kubernetes state lives — it is the only place. The API server holds nothing of its own. Every Deployment spec, every Secret, every ConfigMap, every lease and endpoint and RBAC binding exists because etcd says it exists. If etcd is gone and unrecoverable, your cluster's control plane has nothing to serve, regardless of how healthy your nodes and pods look at that exact moment.

That makes etcd backup a different problem from the "back up my applications" backup most teams already have solved with Velero. Velero backs up Kubernetes objects by calling the API server — it has no idea etcd exists underneath. An etcd snapshot backs up the API server's entire database directly, bypassing the API server. The two are not redundant, and neither one substitutes for the other.


What etcd Actually Stores

etcd is a distributed, strongly-consistent key-value store. Kubernetes uses a single etcd cluster (typically 3 or 5 members for quorum) as the backing store for every object the API server manages. When you run kubectl apply, the API server validates the request, then writes the result to etcd. When you run kubectl get, in the common case the API server is reading from its in-memory watch cache, which etcd itself populated at startup and keeps in sync via etcd's watch mechanism.

This is why "etcd is slow" and "the API server is slow" are usually the same incident, and why etcd running out of disk, losing quorum, or corrupting its data file is a control-plane-wide outage — not a degradation.


Who Actually Controls etcd

This is the detail that gets blurred in a lot of writing about Kubernetes backup, and it matters enormously for what you can and can't do:

Self-managed clusters (kubeadm, kOps, on-prem, Talos, k3s): You run etcd. You have shell access to the nodes, the TLS certificates etcd uses, and full control over snapshot timing, retention, and restore. Everything in this post applies directly.

Managed Kubernetes (EKS, GKE, AKS): The control plane — including etcd — is run by the cloud provider, inside infrastructure you cannot SSH into. You have no etcdctl access, no certificates, and no ability to take your own etcd snapshot. AWS, Google, and Microsoft each run their own internal etcd backup and recovery process for their managed control planes, but it is not exposed to you as a feature, and you cannot invoke it on demand.

If you're on EKS/GKE/AKS, this post is still worth understanding — it's exactly what the provider is doing on your behalf, and it explains why a managed-Kubernetes outage is rarely "we lost etcd" (the provider's own etcd backup/restore handles that) and why your own backup strategy on managed Kubernetes should be Velero, not etcd snapshots.


Taking a Snapshot

On a self-managed cluster, run etcdctl snapshot save from a machine with network access to an etcd member and the cluster's client TLS certificates — usually a control plane node itself:

bash
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-snapshot-$(date +%Y%m%d%H%M).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

On a kubeadm cluster, those certificate paths are the defaults — confirm yours with kubectl -n kube-system describe pod etcd-<node-name> if you're on a different distribution. snapshot save talks to a running etcd member over its client port and streams a point-in-time snapshot to disk; it does not require stopping etcd or taking the cluster offline.

Verify a snapshot is valid before you trust it:

bash
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-snapshot-20261007.db --write-out=table

Get It Off the Node

A snapshot sitting on the same control plane node it was taken from protects you against almost nothing — the most common things that destroy etcd (disk corruption, an accidental rm -rf, the node itself dying) take the snapshot with it. Ship snapshots to object storage immediately after taking them:

bash
aws s3 cp /backup/etcd-snapshot-$(date +%Y%m%d%H%M).db \
  s3://my-etcd-backups/ --sse aws:kms

A cron job on each control plane node (or a CronJob if your etcd runs as static pods you can exec into) covers the "take and ship" half. Hourly snapshots with a short retention and daily snapshots with 30-day retention is a reasonable default — etcd snapshots are small (typically tens to low hundreds of MB for most clusters) and cheap to keep.


Restoring From a Snapshot

This is the part that trips people up, for two reasons: the tool changed, and the process has more steps than "run one restore command."

The tool: etcdctl snapshot restore is deprecated as of etcd 3.5 and removed starting in etcd 3.6. The current, correct tool is etcdutl — a separate binary scoped specifically to operations that work directly on etcd's data files (restore, defragmentation offline, data migration), as opposed to etcdctl, which talks to a running etcd cluster over the network. Restoring a snapshot is inherently an offline, file-level operation, which is why it moved to etcdutl.

The process, run once per cluster member you're restoring:

bash
etcdutl snapshot restore /backup/etcd-snapshot-20261007.db \
  --name etcd-1 \
  --data-dir /var/lib/etcd-restored \
  --initial-cluster etcd-1=https://10.0.1.10:2380,etcd-2=https://10.0.1.11:2380,etcd-3=https://10.0.1.12:2380 \
  --initial-cluster-token etcd-cluster-restore \
  --initial-advertise-peer-urls https://10.0.1.10:2380

--name and --initial-advertise-peer-urls change per member; --initial-cluster lists every member identically on every host. etcdutl writes a fresh data directory from the snapshot — it doesn't touch the running etcd process, because there typically isn't one at this point.

Then, for each member:

  1. Stop the kube-apiserver and the existing etcd process on that node (on kubeadm, this means moving or removing the etcd and kube-apiserver static pod manifests out of /etc/kubernetes/manifests/ so the kubelet stops them).
  2. Replace the old etcd data directory with the one etcdutl just produced.
  3. Start etcd again pointing at the restored data directory, then start the kube-apiserver again.

Do this for every member before considering the cluster recovered — a partial restore (some members on the old, divergent data; some on the restored snapshot) is a quorum problem waiting to happen, not a working cluster.

For Kubernetes specifically, add --bump-revision and --mark-compacted to the restore command. The API server's watch cache tracks etcd's internal revision counter; restoring a snapshot resets that counter backward, which can make the API server's cache believe it has already seen revisions it hasn't. Bumping the revision past anything clients might have already observed avoids serving stale or inconsistent watch results right after a restore.

What a restore actually gets you back: the cluster state as of the snapshot — every object definition, every Secret, every RBAC rule, as they existed at snapshot time. Anything created or changed between the snapshot and the failure is gone. This is the same point-in-time trade-off every backup system has; the fix is the same, too — take snapshots often enough that the gap is tolerable.


Kubernetes Production Readiness Checklist

The pre-launch checks we run before calling a cluster production-ready — probes, resources, RBAC, upgrades, and backups. Plain Markdown you can commit to your repo.

Free. Instant download. You'll also get the occasional deep-dive from the newsletter — unsubscribe anytime.

Defragmentation

etcd doesn't reclaim disk space from deleted keys automatically in the way you'd expect — space freed by compaction stays allocated to etcd's database file, just unused, until you defragment:

bash
etcdctl defrag --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

Two things matter operationally here. First, defragmenting a member blocks reads and writes on that member while it rebuilds its on-disk state — on a 3-member cluster, that member briefly can't serve requests, though the cluster as a whole keeps working through the other two. Second, a defrag request is not replicated — it only ever applies to the member you targeted. Run it against each member one at a time, with a health check between each, rather than targeting --cluster or all endpoints simultaneously; defragmenting every member at once is how a routine maintenance task turns into a quorum outage.

Clusters that see heavy write churn (lots of ConfigMap/Secret updates, high-frequency status updates from controllers) fragment faster and benefit from a scheduled defrag — monthly is a reasonable starting cadence, more often if etcd_mvcc_db_total_size_in_bytes keeps climbing relative to actual key count.


etcd Backup vs Velero: Two Different Recovery Scopes

This is worth stating plainly because it's the single most common point of confusion:

etcd snapshotVelero
Backs upThe API server's entire databaseKubernetes objects, via the API
RestoresThe whole control plane's state at onceSelected namespaces/resources, selectively
PVC dataNot included — etcd stores object metadata, not volume contentsIncluded, via CSI snapshots or Kopia
GranularityAll-or-nothing per snapshotPer-namespace, per-resource-type, cross-cluster
Available on EKS/GKE/AKSNo — managed by the providerYes — this is your actual DR tool on managed Kubernetes

An etcd restore is the right tool when the control plane itself is the thing that broke — a corrupted etcd data file, a botched upgrade, an accidental kubectl delete of something cluster-critical across many namespaces at once. Velero is the right tool for everything else: migrating a namespace to a new cluster, recovering one accidentally deleted application, or any disaster recovery scenario on managed Kubernetes, where etcd snapshots simply aren't available to you. Most production setups that have the option to do both should do both — they're not competing for the same backup budget, they protect different things.


Monitoring etcd Health

A handful of metrics (exposed on etcd's own /metrics endpoint, typically scraped from the control plane) catch problems before they become an incident:

promql
1# Leader changes — frequent churn indicates network issues or resource contention
2increase(etcd_server_leader_changes_seen_total[1h])
3
4# Disk fsync latency — etcd is extremely sensitive to slow disks
5histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m]))
6
7# Database size relative to quota (default quota is 2GB)
8etcd_mvcc_db_total_size_in_bytes / etcd_server_quota_backend_bytes_bytes

Alert on fsync p99 latency climbing past 10-25ms (etcd was designed for low-latency local SSDs, not network-attached storage with unpredictable latency) and on database size approaching the quota — once etcd hits its backend quota it goes into an alarm state and rejects writes cluster-wide until you compact and defragment.


Quick Reference

bash
1# Take a snapshot
2ETCDCTL_API=3 etcdctl snapshot save backup.db \
3  --endpoints=https://127.0.0.1:2379 \
4  --cacert=<ca> --cert=<cert> --key=<key>
5
6# Verify it
7etcdctl snapshot status backup.db --write-out=table
8
9# Restore (per member, offline — note etcdutl, not etcdctl)
10etcdutl snapshot restore backup.db \
11  --name <member-name> \
12  --data-dir <restored-data-dir> \
13  --initial-cluster <name1=url1,name2=url2,...> \
14  --initial-cluster-token <token> \
15  --initial-advertise-peer-urls <this-member-peer-url> \
16  --bump-revision 1000000000 --mark-compacted
17
18# Defragment one member at a time
19etcdctl defrag --endpoints=<single-member> --cacert=<ca> --cert=<cert> --key=<key>
ScenarioTool
Routine backupetcdctl snapshot save
Verify a backupetcdctl snapshot status
Restore the control planeetcdutl snapshot restore (not etcdctl)
Reclaim disk spaceetcdctl defrag, one member at a time
Recover an application/namespaceVelero, not etcd

Frequently Asked Questions

Can I take an etcd snapshot on EKS, GKE, or AKS?

No. On all three, etcd runs inside the cloud provider's managed control plane, which you have no shell or certificate access to. The provider runs its own internal backup/recovery process for its control planes, but it isn't a feature you invoke. Your own backup responsibility on managed Kubernetes is the application layer — Velero — not etcd.

Why does etcdctl snapshot restore fail or warn about deprecation?

Because it's deprecated since etcd 3.5 and removed in 3.6 — the correct tool is now etcdutl snapshot restore, which does the same job but is scoped specifically to offline operations on etcd's data files rather than live-cluster operations.

Does restoring an etcd snapshot bring back my PersistentVolume data?

No. etcd stores the metadata about your PersistentVolumeClaims and PersistentVolumes — the objects, their bindings, their status — not the actual bytes written to the underlying disk. Volume data recovery is a storage-layer concern (CSI snapshots, Velero with Kopia/Restic), entirely separate from etcd.

How often should I take etcd snapshots?

Hourly is a common baseline for production clusters, with daily snapshots retained longer (30 days is reasonable) and hourly ones kept for a shorter window. The right frequency is whatever makes the gap between "last snapshot" and "the moment things broke" tolerable for your change velocity — a cluster where RBAC and cluster-scoped resources change constantly needs tighter snapshots than one that's mostly stable.

What happens if I restore only some etcd members and not others?

You get a cluster in a state no healthy etcd cluster should ever be in — some members serving the restored (older) data, others still on diverged data from before the restore, with no way for Raft consensus to reconcile the difference cleanly. Always restore every member from the same snapshot before bringing the cluster back up; a partial restore risks data loss and quorum failures that are considerably harder to fix than redoing the restore correctly the first time.


For the application-layer backup strategy that complements etcd on self-managed clusters — and replaces it entirely on EKS/GKE/AKS — see Kubernetes Disaster Recovery: Backup and Restore with Velero. For the upgrade process that depends on etcd being healthy before you touch the control plane, see Kubernetes Cluster Upgrades: Zero-Downtime Strategy for Production.

Running self-managed Kubernetes and need a real disaster recovery plan for your control plane, not just your applications? Talk to us at Coding Protocols — we help platform teams design and test etcd backup strategies before they're needed under incident pressure.

Official References

Was this article helpful?

Be the first to rate this article

Related Topics

etcd
Kubernetes
Backup
Disaster Recovery
Platform Engineering
EKS

Found this useful? Share it.

Practice this

Related tools

Read Next

Want this running in production, not just on paper?

We're a hands-on DevOps consultancy — Kubernetes, CI/CD, and cloud infrastructure.

Explore Our Services