DevOps & Platform
13 min readJuly 2, 2026Updated September 19, 2026

The Terraform Kubernetes-Provider Trap: Where to Draw the Line

AJ
Ajeet Yadav
Platform & Cloud Engineer
The Terraform Kubernetes-Provider Trap: Where to Draw the Line

Quick answer

Using Terraform's kubernetes, helm, and kubernetes_manifest providers to run in-cluster workloads feels tidy on day one and hurts every day after. Terraform is the right tool for the cluster and the cloud around it — not for the apps inside it. Here's where the line belongs.

13 min read · DevOps & Platform

Terraform is the best tool I know for building a Kubernetes cluster and the cloud around it — the VPC, the node groups, the IAM, the load balancers, the DNS. So there is an obvious temptation, once the cluster exists, to keep going: add the kubernetes provider, the helm provider, maybe kubernetes_manifest, and manage the Deployments and Helm releases and custom resources inside the cluster from the same codebase. One tool, one apply, one mental model. Tidy.

It's a trap. Not because those providers are badly written — they work — but because Terraform's execution model is fundamentally the wrong shape for in-cluster workloads. Terraform is a plan/apply tool that reconciles at apply time against a state file. Kubernetes workloads live under continuous controllers that mutate objects every minute of every day. Bolt one onto the other and you get plan-time credential failures, perpetual spurious diffs, a blast radius that couples your churny app manifests to your long-lived cloud state, and CRD races that Terraform's graph can't reason about.

My stance, stated plainly: draw the line at the cluster boundary. Terraform owns the cloud, the cluster, IAM, and a thin, pinned bootstrap layer. Everything that runs inside the cluster — apps, their Helm releases, their custom resources — belongs to a continuous reconciler like Argo CD or Flux. Below is the case, failure mode by failure mode.


The chicken-and-egg: provider config resolved before the cluster exists

The kubernetes and helm providers need an endpoint and credentials. The natural instinct is to feed them from the cluster resource you're creating in the same configuration:

hcl
1resource "aws_eks_cluster" "this" {
2  name = "prod"
3  # ...
4}
5
6provider "kubernetes" {
7  host                   = aws_eks_cluster.this.endpoint
8  cluster_ca_certificate = base64decode(aws_eks_cluster.this.certificate_authority[0].data)
9  token                  = data.aws_eks_cluster_auth.this.token
10}
11
12resource "helm_release" "app" {
13  name       = "checkout"
14  repository = "https://charts.example.com"
15  chart      = "checkout"
16  # deployed into a cluster that doesn't exist during the first plan
17}

Terraform configures providers at plan time, before any resource in the run has been created. On a fresh workspace the cluster endpoint and CA are unknown values, the auth token can't be fetched, and the provider comes up misconfigured or pointed at nothing. Sometimes you get lucky and a two-phase apply papers over it; more often you get Kubernetes cluster unreachable or a plan that can't even build. This is the same class of ordering hazard I dissected in the Karpenter IAM deadlock — a single graph that has to both create an authority and authenticate against it in one pass is inherently fragile. Splitting cluster creation and in-cluster resources into separate states is the usual escape hatch, which is already a hint that they don't belong together.

kubernetes_manifest needs a live API server at plan time

The kubernetes_manifest resource is worse, because it doesn't just need credentials at apply time — it needs to talk to a live cluster API server during plan. It does a server-side dry-run to compute the diff, which means the API server must be reachable and the resource's CRD must already be installed before you can even plan.

hcl
resource "kubernetes_manifest" "cert" {
  manifest = yamldecode(file("${path.module}/certificate.yaml"))
}

The consequences are concrete and painful:

  • terraform plan in CI fails whenever the runner can't reach the cluster API — a private endpoint, a broken kubeconfig, a network blip. Your plan, the thing that's supposed to be a safe read-only preview, now has a hard dependency on cluster reachability.
  • A fresh workspace can't plan at all if the cluster doesn't exist yet, because there's no API server to dry-run against.
  • CRD-then-CR in one run breaks, because the manifest for the custom resource is validated against a CRD that hasn't been applied yet.

A plan that requires the very thing it might be creating to already be up and reachable is not a preview — it's a runtime dependency masquerading as one.

Reconciliation mismatch: Terraform corrects drift once; Kubernetes drifts constantly

This is the deepest problem, and it's philosophical, not fixable with a flag. Terraform reconciles at apply time only. It compares state to reality when you run it, converges once, and then does nothing until the next run. Kubernetes is the opposite: it's controllers all the way down, mutating objects continuously.

An HPA rewrites spec.replicas seconds after you set it. Admission webhooks inject sidecars and labels. Operators patch the resources they own. So a Deployment you manage with the kubernetes provider produces a diff every single plan:

hcl
1resource "kubernetes_deployment" "api" {
2  metadata { name = "api" }
3  spec {
4    replicas = 2   # the HPA already scaled this to 7
5    # ...
6  }
7}

Now you're stuck between two bad options: let Terraform fight the HPA and stomp replicas back to 2 on every apply, or paper over it with a growing pile of lifecycle { ignore_changes = [...] } blocks until Terraform is no longer meaningfully managing the resource at all. Neither is management; both are noise.

GitOps controllers are built for exactly this. Argo CD and Flux run a continuous control loop, understand server-side apply and field ownership, and can be told to ignore fields that other controllers own. Drift correction is their entire job, not an apply-time side effect. If you want the shape of that setup, I walked through it in Argo CD in production and the complete guide to Argo CD GitOps.

Blast radius: one state file for infra and apps is a coupling you'll regret

Put the VPC, the EKS cluster, the node groups and forty Helm releases in one state, and every trivial app change now drags the entire infrastructure graph through plan and apply. The costs compound:

  • Every app deploy locks the infra state. A one-line image bump contends for the same state lock as a network change. Deploys serialize behind each other and behind infra work.
  • Plans get slow. Terraform refreshes the whole graph — every cloud resource and every Helm release — to tell you one chart changed.
  • Blast radius balloons. A botched provider upgrade, a corrupted state entry, or a bad terraform apply while iterating on an app can now damage the state that also owns your production VPC and cluster. The things that change hourly (apps) should not share a failure domain with the things that change quarterly (network, cluster).

Long-lived infrastructure and high-churn workloads have opposite change frequencies and opposite risk profiles. Coupling them in one state optimizes for neither. Even the good pattern — Terraform for infra — argues for keeping app churn out; see Terraform for EKS infrastructure as code for how much belongs in that layer already.

Terraform Day-2 Operations Checklist

State hygiene, drift, imports, policy checks, and upgrade routines — everything after `terraform apply` works. Plain Markdown, commit it to your repo.

Free. Instant download. You'll also get the occasional deep-dive from the newsletter — unsubscribe anytime.

CRDs and CRs in the same run: a race Terraform can't see

Installing an operator's CRDs and then creating custom resources of those kinds in one apply is a genuine ordering problem, and Terraform's dependency graph doesn't understand it. A CRD isn't usable the instant the object is created — the API server has to establish it (the Established condition) and register the new API endpoints. Terraform sees "CRD resource created" and immediately tries to create the CR, which races the API server's registration:

hcl
1resource "helm_release" "cert_manager" {
2  name  = "cert-manager"
3  chart = "cert-manager"
4  # installs CRDs like Certificate, ClusterIssuer
5}
6
7resource "kubernetes_manifest" "issuer" {
8  manifest   = yamldecode(file("issuer.yaml"))  # kind: ClusterIssuer
9  depends_on = [helm_release.cert_manager]        # not enough
10}

depends_on guarantees ordering of Terraform operations, not readiness of the API surface. The CR apply frequently fails with "no matches for kind ClusterIssuer" because the kind isn't registered yet, and your fix becomes targeted applies, retries, or splitting the run in two. GitOps controllers handle this natively with retries and health-based sync waves — they expect the API surface to appear asynchronously and wait for it.

Day-2 app config is a bad fit for plan/apply

Rolling out a new image tag, tuning replica counts, flipping a feature-flag ConfigMap — this is the daily rhythm of running applications, and it wants a fast, low-ceremony, auto-reconciling loop. The plan/apply cycle is deliberately heavyweight: generate a plan, review it, acquire a lock, apply. That's exactly right for infrastructure you touch occasionally and want to scrutinize. It's friction for the tenth image bump of the day, and the friction pushes people toward -target and out-of-band kubectl, which quietly desyncs state from reality — the worst of both worlds.


When Terraform inside the cluster IS the right call: the bootstrap layer

I'm not arguing Terraform should never touch the cluster's API. There's a real, bounded job it does well: the bootstrap layer. Right after the cluster is created, a small set of cluster-level singletons has to exist before GitOps can take over — including the GitOps controller itself, which obviously can't deploy itself via GitOps. Terraform is a fine tool for that thin slice:

hcl
1# Bootstrap only: the handful of singletons GitOps needs to exist first.
2resource "helm_release" "argocd" {
3  name             = "argocd"
4  namespace        = "argocd"
5  create_namespace = true
6  repository       = "https://argoproj.github.io/argo-helm"
7  chart            = "argo-cd"
8  version          = "7.7.11"   # pinned, always
9}
10
11resource "helm_release" "ingress_nginx" {
12  name       = "ingress-nginx"
13  namespace  = "ingress-nginx"
14  repository = "https://kubernetes.github.io/ingress-nginx"
15  chart      = "ingress-nginx"
16  version    = "4.11.3"
17}

(ingress-nginx here is just illustrative of a bootstrap singleton — note that the ingress-nginx project entered retirement in 2026, so for new clusters you'd more likely bootstrap a Gateway API-based controller. The pattern is identical whatever the chart.)

The mandate for this layer is narrow: the ingress controller, the GitOps controller (Argo CD/Flux), the autoscaler (cluster-autoscaler or Karpenter), and maybe a couple of foundational operators. Alongside it, the platform-provisioning that's genuinely Terraform's domain also fits: IRSA/IAM roles and their bindings, StorageClass objects, and the namespaces your platform components live in. The governing pattern is one sentence:

Terraform builds the cluster and installs the GitOps controller; the GitOps controller deploys everything else.

Keep this layer pinned to exact chart versions, keep it minimal, and — importantly — keep it in a separate state or workspace from the cluster itself, so bootstrap churn doesn't contend with cluster-lifecycle operations. The moment bootstrap starts growing into "well, this one app is easier to just Helm-release here," push back. That's the trap reasserting itself.

The recommendation: draw the line, then hand off

The clean division of responsibility:

  • Terraform owns: cloud resources (VPC, subnets, security groups), the cluster, node groups/Karpenter, IAM and IRSA, DNS and load balancers, and a thin pinned bootstrap (GitOps controller, ingress, autoscaler, storage classes, namespaces) — ideally in its own state.
  • GitOps owns: every application, every app Helm release, every custom resource, and all day-2 config. Argo CD or Flux reconciles it continuously from Git.

The handoff point is deliberate: Terraform's last act inside the cluster is installing Argo CD and pointing it at your config repo (an app-of-apps or ApplicationSet root). From there, workloads flow through Git, not through terraform apply.

If your philosophy is instead "the Kubernetes API should be the single control plane for everything, including cloud infrastructure," that's a coherent position too — but the tool for it is Crossplane, not Terraform's Kubernetes providers. Crossplane turns cloud resources into Kubernetes objects reconciled by in-cluster controllers, which is a genuine continuous-reconciliation model rather than plan/apply bolted onto a cluster. I compared that approach in Crossplane: infrastructure as code on Kubernetes. Just don't try to get Crossplane's continuous model out of Terraform's providers — you'll get the ceremony of Terraform with none of the reconciliation.


Frequently Asked Questions

Is it ever OK to use the Terraform helm provider?

Yes, for a thin bootstrap layer only: installing the GitOps controller, ingress controller, autoscaler, and a couple of foundational operators right after cluster creation is legitimate, since those singletons must exist before GitOps can take over. Pin exact chart versions, keep the list short, and isolate it in its own state. The problem isn't the helm provider itself — it's using it for the dozens of application releases that belong in Argo CD or Flux.

Why does terraform plan fail when I use kubernetes_manifest?

Because kubernetes_manifest runs a server-side dry-run against the live cluster API during plan to compute its diff, so the API server must be reachable and the resource's CRD must already exist. That fails outright on a fresh workspace with no cluster yet, or a CI runner that can't reach a private endpoint. It also breaks when a CRD and a custom resource of that kind are applied in the same run, since the CR is validated against a CRD that isn't installed yet.

Should I manage Kubernetes Deployments with Terraform?

I'd strongly advise against it outside static platform bootstrap. Terraform reconciles only at apply time, while Kubernetes controllers mutate objects continuously — an HPA rewriting replicas, webhooks injecting sidecars, operators patching resources. That causes a spurious diff on every plan, forcing you to fight controllers or bury fields in ignore_changes. Deployments belong to a continuous reconciler — use Argo CD or Flux and keep app manifests in Git.

How do I split Terraform and GitOps responsibilities?

Draw the line at the cluster boundary. Terraform owns the cloud (VPC, IAM, DNS, load balancers), the cluster and node pools, and a thin pinned bootstrap layer whose last act installs the GitOps controller and points it at your config repo. From there, the GitOps controller owns every application, Helm release, custom resource, and day-2 change, reconciling continuously from Git. Keep bootstrap in its own state so app churn never contends with cluster-lifecycle operations.

Is Crossplane a better fit than Terraform for Kubernetes-native infrastructure?

Yes, if your goal is a single Kubernetes-native control plane that continuously reconciles cloud infrastructure like a Deployment — that's real continuous reconciliation, not plan/apply strapped onto a cluster. Terraform's Kubernetes providers can't do that; they still run one-shot applies. Many teams run both, Terraform for the cluster and bootstrap, Crossplane for app-facing infrastructure APIs. The mistake is expecting kubernetes/helm providers to behave like a control plane.


The pattern that has held up for me across many clusters is boring on purpose: Terraform draws the cloud and the cluster, installs the GitOps controller as its final in-cluster act, and stops at the API boundary; Argo CD or Flux takes it from there and never gives it back. If you're standing up that layer, Terraform for EKS infrastructure as code covers the infra side, Argo CD in production covers the handoff, and Crossplane on Kubernetes covers the alternative philosophy if you want the cluster to be your one control plane.

Not sure where your line between Terraform and GitOps should sit? Talk to us at Coding Protocols — we help platform teams split infrastructure and workloads cleanly so plans stay fast and deploys stay safe.

Official References

Was this article helpful?

Be the first to rate this article

Related Topics

Terraform
Kubernetes
GitOps
Infrastructure as Code
Platform Engineering
Helm
Argo CD

Found this useful? Share it.

Practice this

Related tools

Read Next