Kubernetes
8 min readJuly 2, 2026Updated August 19, 2026

What Is a Kubernetes Operator? The Custom-Controller Pattern Explained

AJ
Ajeet Yadav
Platform & Cloud Engineer
What Is a Kubernetes Operator? The Custom-Controller Pattern Explained

Quick answer

A Kubernetes operator is a custom controller plus one or more CRDs that extends the Kubernetes API to automate an application's full lifecycle — install, upgrade, backup, failover. It encodes a human operator's runbook as software that continuously reconciles desired state.

8 min read · Kubernetes

A Kubernetes operator is a custom controller paired with one or more Custom Resource Definitions (CRDs) that extends the Kubernetes API to automate a specific application's full lifecycle — installation, configuration, upgrades, backups, scaling, and failure recovery. Instead of running those tasks by hand, you declare the desired state of your application as a custom resource, and the operator continuously reconciles reality toward that declaration. In short: an operator encodes a human operator's runbook as software that never sleeps.

The pattern was introduced by CoreOS in 2016 and has become the standard way to run stateful, complex workloads — databases, message brokers, monitoring stacks — on Kubernetes. If you've ever installed something with a CRD like PostgresCluster or PrometheusRule, you've used an operator.


How an operator works

An operator has exactly two moving parts:

  1. A custom controller — a long-running process (usually a pod in your cluster) that watches for changes to particular resources and acts on them.
  2. One or more CRDs — schema definitions that teach the Kubernetes API server about new resource kinds, such as EtcdCluster or KafkaTopic.

The controller runs a reconcile loop, which is a direct application of control theory — the same feedback principle as a thermostat. The loop repeats three steps forever:

  • Observe — read the current, actual state of the world (pods, secrets, external systems).
  • Diff — compare that actual state against the desired state declared in the custom resource's spec.
  • Act — take whatever concrete actions close the gap: create a StatefulSet, run a backup job, trigger a failover, patch a config.

Crucially, reconciliation is level-triggered, not edge-triggered. The controller doesn't react to a single "create" event and forget; it repeatedly drives toward the target state, so it self-heals after crashes, missed events, or manual drift. A simplified reconcile function looks like this:

go
1func (r *DatabaseReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
2    // 1. OBSERVE — fetch the desired state from the custom resource
3    var db v1alpha1.Database
4    if err := r.Get(ctx, req.NamespacedName, &db); err != nil {
5        return ctrl.Result{}, client.IgnoreNotFound(err)
6    }
7
8    // 2. DIFF — is the backing StatefulSet what the spec asks for?
9    desired := buildStatefulSet(&db) // replicas, version, storage from db.Spec
10    var current appsv1.StatefulSet
11    err := r.Get(ctx, statefulSetKey(&db), &current)
12
13    // 3. ACT — create or update to converge toward desired state
14    if apierrors.IsNotFound(err) {
15        return ctrl.Result{}, r.Create(ctx, desired)
16    }
17    if !equalSpec(current, desired) {
18        return ctrl.Result{}, r.Update(ctx, desired)
19    }
20
21    // Nothing to do; requeue to keep watching for drift.
22    return ctrl.Result{RequeueAfter: time.Minute}, nil
23}

That's the whole idea. The genius is that the same loop that installs the app also upgrades it, recovers it, and keeps it healthy — because every one of those is just another diff to reconcile.

What a CRD is

A Custom Resource Definition is how you extend the Kubernetes API with your own object types. Out of the box, Kubernetes knows about Pod, Deployment, Service, and so on. A CRD registers a new kind — say Database — with its own schema, validation rules, and API group. Once applied, kubectl get databases works exactly like kubectl get pods, and the API server stores and validates your custom objects natively.

yaml
1apiVersion: acme.io/v1alpha1
2kind: Database
3metadata:
4  name: orders-db
5spec:
6  engine: postgres
7  version: "16.2"
8  replicas: 3
9  storageGB: 100
10  backupSchedule: "0 2 * * *"

The CRD defines the vocabulary (what fields are valid), and the operator's controller supplies the behavior (what those fields actually do). A CRD with no controller is just an inert object in etcd — it's the controller that makes a custom resource mean something. Together they form the operator.

The operator pattern: encoding operational knowledge

The reason operators matter isn't the plumbing — it's what they let you capture. Running a production database involves a mountain of day-2 operational knowledge: how to take a consistent backup, how to promote a replica when the primary dies, how to perform a rolling version upgrade without data loss, how to resize storage safely. Traditionally that knowledge lived in runbooks, wiki pages, and the heads of a few senior engineers.

The operator pattern turns that runbook into executable software. A mature database operator doesn't just deploy Postgres — it knows to quiesce writes before a backup, to fence the old primary during failover, to bump one replica at a time on upgrade. The operator is the on-call engineer's expertise, running continuously and identically every time.

This is why the ecosystem is full of vendor-authored operators. Projects like the External Secrets Operator sync secrets from external vaults, and the NVIDIA GPU Operator automates the entire driver, toolkit, and device-plugin stack on GPU nodes — deep, product-specific operational knowledge that would be tedious and error-prone to run by hand.

Kubernetes Production Readiness Checklist

The pre-launch checks we run before calling a cluster production-ready — probes, resources, RBAC, upgrades, and backups. Plain Markdown you can commit to your repo.

Free. Instant download. You'll also get the occasional deep-dive from the newsletter — unsubscribe anytime.

When you need an operator (and when you don't)

Operators are powerful, but they're not free — you're now maintaining a stateful controller with its own bugs, RBAC, and upgrade path. Reach for one when:

  • Your app is stateful and lifecycle-heavy. Databases, message queues, and storage systems have complex backup, failover, and upgrade logic that benefits enormously from automation.
  • You're distributing software to others. If you ship a product that runs on Kubernetes, an operator gives your users a clean, declarative install-and-manage experience.
  • Day-2 operations dominate. When the hard part isn't deploying but keeping it running correctly, that operational knowledge is exactly what an operator captures.

You almost certainly don't need to write an operator when:

  • Your app is stateless. A Deployment plus an HorizontalPodAutoscaler already gives you self-healing and scaling. Adding an operator is pure overhead.
  • Templating is enough. If you just need to parameterize and install YAML, use Helm or Kustomize. For turning a bundle of resources into a simple custom API without writing controller code, tools like kro fit better — see kro vs Crossplane vs Helm for where each sits.
  • An operator already exists. Don't rebuild the Prometheus or Postgres operator. Adopt the community one.

The rule of thumb: write an operator when you have real operational logic to encode and reconcile — not merely to deploy resources.

Common misconceptions

"An operator is just a Helm chart." No. Helm renders and applies YAML once per install or upgrade, then stops. An operator runs continuously, watching and correcting drift forever. Different altitude, different job.

"CRD equals operator." A CRD is only the schema. Without a controller reconciling those resources, nothing happens. The operator is the CRD plus the controller.

"Operators are only for databases." Databases are the classic example, but operators automate certificates, secrets, GPU drivers, service meshes, backups, and cost controls. Any workload with meaningful lifecycle logic is a candidate.

"You must write an operator in Go." Go with controller-runtime is the most common path, but you can build operators in Python, Rust, or Java, and frameworks like the Operator SDK support Ansible and Helm-based operators too.

Frequently Asked Questions

What is the difference between a Kubernetes operator and a controller?

Every operator is a controller, but not every controller is an operator. A controller is any process running a reconcile loop against Kubernetes resources — the built-in Deployment and Node controllers are examples. An operator is specifically a custom controller paired with CRDs that automates a particular application's domain-specific lifecycle. "Operator" describes intent and packaging; "controller" describes the underlying mechanism.

Is a CRD required to build an operator?

Almost always, yes. The operator pattern is defined by extending the API with custom resources that users declare, and CRDs are how you register those resources. In theory you could write a controller that only reconciles built-in resources, but without a CRD you have a controller, not an operator in the conventional sense.

Do I need to know Go to write an operator?

No, though Go is the dominant choice because controller-runtime and Kubebuilder are Go libraries. The Operator SDK also supports Ansible-based and Helm-based operators for simpler cases, and community frameworks let you build controllers in Python, Rust, and other languages. For serious custom logic, most teams still land on Go.

Are Kubernetes operators secure to run in production?

They can be, with care. An operator runs with a ServiceAccount and RBAC permissions — often broad ones — so a compromised or buggy operator has significant blast radius. Scope its RBAC to the least privilege it needs, pin and verify images, run it in a dedicated namespace, and prefer well-maintained operators over abandoned ones.

When should I not use an operator?

Skip the operator when your workload is stateless and self-heals with a Deployment and HPA, or when you only need to template and install manifests — Helm or Kustomize is simpler. Also skip building your own if a mature operator already exists for your software. Operators earn their keep only when there's genuine day-2 operational logic to reconcile.

See also

Trying to decide whether an operator is the right abstraction — or whether Helm, kro, or a plain Deployment would serve you better? Talk to us at Coding Protocols — we help platform teams design the right level of automation without over-engineering it.

Official References

Was this article helpful?

Be the first to rate this article

Related Topics

Kubernetes
Operators
Custom Controllers
CRD
Reconciliation
Platform Engineering

Found this useful? Share it.

Practice this

Related tools

Read Next