Build a Kubernetes Operator from Scratch with Kubebuilder and controller-runtime
Quick answer
Go past 'what is an operator' and actually build one. Scaffold a CRD with Kubebuilder, write a reconcile loop in Go that creates and owns a Deployment, mirror status back, and watch it self-heal — the real controller-runtime mechanics, step by step.
- Step 1: Scaffold the Project
- Step 2: Create the API
- Step 3: Define the API Types
- Step 4: Implement the Reconcile Loop
- Step 5: Wire Up the Watch
advanced · 75 min
Before you begin
- Go 1.22+ installed
- Kubebuilder v4 installed (book.kubebuilder.io)
- A cluster you can use (kind or minikube is fine) and kubectl configured
- Comfortable reading Go; basic understanding of CRDs and Deployments
Most "Kubernetes operator" content stops at the concept: a custom resource plus a control loop that reconciles it. That's the what. This tutorial is the how — you'll build a working operator in Go with Kubebuilder and controller-runtime, the same toolchain the real ecosystem operators are built on.
The operator manages a custom WebApp resource. A user writes a few lines of YAML — an image and a replica count — and your controller creates and continuously owns a Deployment for it, mirrors the ready-replica count back into the resource's status, and self-heals if the Deployment is deleted or changed. By the end you'll have touched every core mechanic: the CRD, the reconcile loop, owner references and garbage collection, the status subresource, and the watch that ties it together.
This is the controller you don't have to write if you use a tool like kro — see kro vs Crossplane vs Helm — but understanding it is what lets you reach for (or skip) those tools deliberately.
What You'll Build
- A
WebAppCRD (webapp.example.com/v1alpha1) with a typed spec and a status subresource - A reconcile loop that create-or-updates a
Deploymentfrom the spec - Owner references so the Deployment is garbage-collected when the
WebAppis deleted - A watch on owned Deployments so status stays current and the controller self-heals drift
- Generated CRD manifests and RBAC from Go markers
Step 1: Scaffold the Project
Create an empty directory and initialize the project. The --repo flag is your Go module path.
mkdir webapp-operator && cd webapp-operator
kubebuilder init --domain example.com --repo example.com/webapp-operatorThis generates the project skeleton: cmd/main.go (the manager entrypoint), a Makefile, config/ (Kustomize manifests), and the PROJECT file that tracks your APIs.
Step 2: Create the API
kubebuilder create api --group webapp --version v1alpha1 --kind WebAppAnswer y to both prompts ("Create Resource" and "Create Controller"). Kubebuilder scaffolds two files you'll edit:
api/v1alpha1/webapp_types.go— the Go structs for your CRDinternal/controller/webapp_controller.go— the reconcile loop
(In Kubebuilder v4, controllers live under internal/controller/, not the old controllers/.)
Step 3: Define the API Types
Open api/v1alpha1/webapp_types.go and replace the scaffolded WebAppSpec and WebAppStatus with real fields. The marker comments drive code generation — they become CRD schema, validation, and the status subresource.
1// WebAppSpec defines the desired state of WebApp.
2type WebAppSpec struct {
3 // Image is the container image to run.
4 Image string `json:"image"`
5
6 // Replicas is the desired number of pods.
7 // +kubebuilder:validation:Minimum=1
8 // +kubebuilder:default=1
9 Replicas int32 `json:"replicas,omitempty"`
10}
11
12// WebAppStatus defines the observed state of WebApp.
13type WebAppStatus struct {
14 // ReadyReplicas mirrors the managed Deployment's ready replica count.
15 ReadyReplicas int32 `json:"readyReplicas,omitempty"`
16}
17
18//+kubebuilder:object:root=true
19//+kubebuilder:subresource:status
20//+kubebuilder:printcolumn:name="Image",type=string,JSONPath=`.spec.image`
21//+kubebuilder:printcolumn:name="Ready",type=integer,JSONPath=`.status.readyReplicas`
22
23// WebApp is the Schema for the webapps API.
24type WebApp struct {
25 metav1.TypeMeta `json:",inline"`
26 metav1.ObjectMeta `json:"metadata,omitempty"`
27
28 Spec WebAppSpec `json:"spec,omitempty"`
29 Status WebAppStatus `json:"status,omitempty"`
30}Two markers matter most here:
//+kubebuilder:subresource:statusmakesstatusa subresource, so updates tostatusdon't bump the resource's main generation and can't be set bykubectl apply. This is the correct separation between desired state (spec, user-owned) and observed state (status, controller-owned).//+kubebuilder:validation:Minimum=1and//+kubebuilder:default=1become OpenAPI schema in the CRD — the API server rejectsreplicas: 0and defaults an omitted value, before your controller ever runs.
Step 4: Implement the Reconcile Loop
This is the heart of the operator. Open internal/controller/webapp_controller.go and replace the Reconcile method. The whole job: read the WebApp, make the cluster match it, report what you observed.
1package controller
2
3import (
4 "context"
5
6 appsv1 "k8s.io/api/apps/v1"
7 corev1 "k8s.io/api/core/v1"
8 metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
9 ctrl "sigs.k8s.io/controller-runtime"
10 "sigs.k8s.io/controller-runtime/pkg/client"
11 "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
12 logf "sigs.k8s.io/controller-runtime/pkg/log"
13
14 webappv1alpha1 "example.com/webapp-operator/api/v1alpha1"
15)
16
17//+kubebuilder:rbac:groups=webapp.example.com,resources=webapps,verbs=get;list;watch;create;update;patch;delete
18//+kubebuilder:rbac:groups=webapp.example.com,resources=webapps/status,verbs=get;update;patch
19//+kubebuilder:rbac:groups=webapp.example.com,resources=webapps/finalizers,verbs=update
20//+kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
21
22func (r *WebAppReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
23 logger := logf.FromContext(ctx)
24
25 // 1. Fetch the WebApp. If it's gone, there's nothing to do —
26 // owner references will garbage-collect the Deployment for us.
27 var app webappv1alpha1.WebApp
28 if err := r.Get(ctx, req.NamespacedName, &app); err != nil {
29 return ctrl.Result{}, client.IgnoreNotFound(err)
30 }
31
32 // 2. Declare the Deployment we want, then create-or-update it idempotently.
33 deploy := &appsv1.Deployment{
34 ObjectMeta: metav1.ObjectMeta{Name: app.Name, Namespace: app.Namespace},
35 }
36 op, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error {
37 labels := map[string]string{"app": app.Name}
38 replicas := app.Spec.Replicas
39
40 deploy.Spec.Replicas = &replicas
41 deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels}
42 deploy.Spec.Template.ObjectMeta.Labels = labels
43 deploy.Spec.Template.Spec.Containers = []corev1.Container{{
44 Name: "web",
45 Image: app.Spec.Image,
46 Ports: []corev1.ContainerPort{{ContainerPort: 80}},
47 }}
48
49 // Own the Deployment so it's deleted (and re-reconciled) with the WebApp.
50 return controllerutil.SetControllerReference(&app, deploy, r.Scheme)
51 })
52 if err != nil {
53 return ctrl.Result{}, err
54 }
55 if op != controllerutil.OperationResultNone {
56 logger.Info("reconciled deployment", "operation", op)
57 }
58
59 // 3. Report observed state. Update the STATUS subresource only.
60 app.Status.ReadyReplicas = deploy.Status.ReadyReplicas
61 if err := r.Status().Update(ctx, &app); err != nil {
62 return ctrl.Result{}, err
63 }
64
65 return ctrl.Result{}, nil
66}A few things that make this correct, not just compiling:
client.IgnoreNotFound— a deletedWebApptriggers one final reconcile whereGetreturnsNotFound. You don't treat that as an error; you let Kubernetes' garbage collector remove the owned Deployment via the owner reference.CreateOrUpdateis idempotent — itGets the Deployment, runs your mutate function, and creates or updates as needed. Reconcile must be safe to run any number of times; this pattern guarantees it.SetControllerReference— stamps theWebAppas the Deployment's owner. That single line is what enables cascading delete and the self-healing watch in the next step.r.Status().Update— writes only the status subresource, never the spec. Controllers report status; users own spec.
Step 5: Wire Up the Watch
Find SetupWithManager in the same file. For declares the primary type; Owns is the important addition — it watches Deployments this controller owns and re-queues the parent WebApp whenever one changes.
func (r *WebAppReconciler) SetupWithManager(mgr ctrl.Manager) error {
return ctrl.NewControllerManagedBy(mgr).
For(&webappv1alpha1.WebApp{}).
Owns(&appsv1.Deployment{}).
Complete(r)
}Owns(&appsv1.Deployment{}) is why two things work without any extra code: when the Deployment's pods become ready, the controller re-reconciles and updates status.readyReplicas; and if someone deletes the Deployment, the controller is triggered and recreates it. That's the self-healing control loop, for free.
Step 6: Generate Manifests and RBAC
Your Go markers are the source of truth. Regenerate the deepcopy methods, the CRD, and the RBAC role:
make generate # regenerates zz_generated.deepcopy.go from the types
make manifests # regenerates the CRD YAML and the RBAC ClusterRole from markersmake manifests reads those //+kubebuilder:rbac markers and writes a least-privilege ClusterRole into config/rbac/ — so the operator's permissions are derived from the code that actually needs them. (If RBAC is new to you, see the Kubernetes RBAC & Security tutorial.)
Step 7: Install the CRD and Run Locally
The fastest dev loop runs the controller outside the cluster, against your kubeconfig:
make install # installs the WebApp CRD into the cluster
make run # runs the controller locally, streaming logs to your terminalmake run blocks and logs. Leave it running and open a second terminal for the next step.
Step 8: Create a WebApp and Watch It Reconcile
1kubectl apply -f - <<'EOF'
2apiVersion: webapp.example.com/v1alpha1
3kind: WebApp
4metadata:
5 name: demo
6spec:
7 image: nginx:1.27
8 replicas: 3
9EOFNow watch the operator do its job:
kubectl get webapp demo # your custom resource, with the printer columns
kubectl get deployment demo # the Deployment your controller created
kubectl get webapp demo -o jsonpath='{.status.readyReplicas}' # status, mirrored backWithin a few seconds status.readyReplicas climbs to 3 as the pods come up — pushed there by the Owns watch firing on each Deployment status change.
Step 9: Prove the Control Loop
This is the part that makes it an operator and not a one-shot script.
1# Self-healing: delete the managed Deployment and watch it come back
2kubectl delete deployment demo
3kubectl get deployment demo # recreated almost immediately
4
5# Drift correction: scale the Deployment by hand
6kubectl scale deployment demo --replicas=1
7kubectl get deployment demo # reconciled back to 3 (the WebApp spec wins)
8
9# Declarative change: edit the WebApp, not the Deployment
10kubectl patch webapp demo --type=merge -p '{"spec":{"replicas":5}}'
11
12# Garbage collection: delete the WebApp, the Deployment goes with it
13kubectl delete webapp demo
14kubectl get deployment demo # NotFound — removed by owner referenceThe Deployment heals because the Owns watch re-queues the WebApp; it scales back because reconcile always rewrites the Deployment to match spec; it's garbage-collected because of the owner reference you set in Step 4.
Step 10 (Optional): Deploy the Operator to the Cluster
Running in-cluster is the same controller, packaged as an image and run as a Deployment with the generated RBAC:
make docker-build docker-push IMG=<registry>/webapp-operator:v0.1.0
make deploy IMG=<registry>/webapp-operator:v0.1.0make deploy applies the CRD, the RBAC ClusterRole/Binding, and the controller Deployment into the *-system namespace.
Common Issues
no matches for kind "WebApp"— you forgotmake install, or edited the types without rerunningmake manifests. Regenerate and reinstall.status.readyReplicasnever updates — you droppedOwns(&appsv1.Deployment{}), so the controller isn't watching the Deployment and never re-reconciles when it becomes ready.forbiddenerrors in-cluster — your//+kubebuilder:rbacmarkers don't cover a resource you touch. Add the marker,make manifests, andmake deployagain.- Deleting the WebApp leaves the Deployment behind —
SetControllerReferenceisn't being called (or returns an error you're swallowing). Owner references are what drive garbage collection.
Frequently Asked Questions
What's the difference between Kubebuilder and controller-runtime?
controller-runtime is the Go library that provides the manager, clients, caches, and the Reconciler interface. Kubebuilder is the scaffolding tool and project framework built on top of it — it generates the CRD types, the controller skeleton, the Makefile, and the deployment manifests so you don't wire all of that by hand. You write controller-runtime code inside a Kubebuilder project.
Do I need a status subresource?
You should use one for any non-trivial operator. Without //+kubebuilder:subresource:status, a kubectl apply of the whole object can clobber controller-written status, and status updates bump the resource's metadata.generation. The subresource cleanly separates user-owned spec from controller-owned status and gives you a dedicated /status endpoint via r.Status().Update().
Why use CreateOrUpdate instead of just Create?
Reconcile runs repeatedly — on every change, on resyncs, and after restarts. A bare Create fails the second time with AlreadyExists. controllerutil.CreateOrUpdate reads the current object, applies your mutate function, and creates or patches as needed, making the loop idempotent — the core requirement of any reconciler.
How does the operator self-heal a deleted Deployment?
Through the Owns(&appsv1.Deployment{}) watch plus owner references. When you set the WebApp as the Deployment's controller owner, controller-runtime watches that Deployment; any change (including deletion) re-queues the owning WebApp, and the next reconcile recreates or repairs it. You never write delete-detection logic yourself.
Is this the same as writing an operator with the Operator SDK?
Largely, yes. The Operator SDK's Go-based operators use Kubebuilder and controller-runtime underneath, so the reconcile code you wrote here is portable. The SDK adds extra paths (Helm- and Ansible-based operators, OLM packaging) on top of the same controller-runtime core.
Tear Down
1# If you deployed in-cluster:
2make undeploy
3
4# Remove the CRD (and any remaining WebApps):
5make uninstall
6
7# Local kind cluster, if you spun one up just for this:
8kind delete clusterOfficial References
- Kubebuilder Book — Quick Start — Official scaffolding workflow, project layout, and the tutorial this builds on
- Kubebuilder Book — Architecture — How the manager, controller, cache, and clients fit together
- controller-runtime GoDoc — API reference for the manager,
Reconciler,client, andcontrollerutil - Kubernetes — Operator pattern — The concept and ecosystem context
- Kubernetes — Custom Resources — CRDs, subresources, and when to extend the API
We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.
Struggling with this in production?
We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.