AWS

Build the Cheapest Container Setup on AWS With Terraform

Intermediate40 min to complete12 min readOctober 2, 2026

Quick answer

A real, working Terraform configuration for running a Fargate Spot container on AWS with no ALB and no NAT Gateway — API Gateway, Cloud Map, and a dual-stack subnet with IPv6-only egress, built step by step and including the fix for the one gotcha that breaks a naive attempt.

intermediate · 40 min

Before you begin

  • An AWS account with credentials configured for Terraform (AWS CLI or environment variables)
  • Terraform 1.5 or later
  • Basic familiarity with VPCs, ECS, and Terraform syntax
  • Comfort reading HCL — this tutorial builds a real multi-resource configuration, not a single-file toy
AWS
Terraform
Fargate
Containers
Cost Optimization
Networking
ECS

The write-up on skipping the ALB and NAT Gateway for a small Fargate container covers the architecture and the cost math. This is the hands-on companion: a complete, working Terraform configuration for it, built resource by resource, including the fix for the one gotcha that will otherwise leave you staring at "no target endpoints found" for an hour.

By the end you'll have a single ARM64 Fargate Spot task, discovered by AWS Cloud Map, reachable through an API Gateway HTTP API — no load balancer, no NAT Gateway, costing a few dollars a month.

What You'll Build

  • A VPC with a dual-stack subnet: a normal private IPv4 range for in-VPC service discovery, plus IPv6 with a default route out through an egress-only internet gateway — no NAT Gateway anywhere
  • An ECS cluster running Fargate Spot, with a single ARM64 task
  • Cloud Map private DNS service discovery registering the task automatically as it starts and stops
  • An API Gateway HTTP API with a VPC Link, routing public requests straight to the task through Cloud Map — no ALB, no target group
  • The SRV-record fix Cloud Map needs for this integration to work at all, applied from the start instead of discovered via a bare, unhelpful 500 error
Rendering diagram…

Step 1: A Dual-Stack VPC and Subnet

The subnet needs to do two different jobs: give the task a private IPv4 address (Cloud Map and the VPC Link only resolve IPv4 targets — this is gotcha #1 from the write-up, and building the subnet dual-stack from the start sidesteps it entirely), and route the task's actual internet-bound traffic out over IPv6, with no NAT Gateway involved.

hcl
1terraform {
2  required_providers {
3    aws = {
4      source  = "hashicorp/aws"
5      version = "~> 5.0"
6    }
7  }
8}
9
10variable "region" {
11  default = "us-east-1"
12}
13
14provider "aws" {
15  region = var.region
16}
17
18data "aws_availability_zones" "available" {
19  state = "available"
20}
21
22resource "aws_vpc" "main" {
23  cidr_block                       = "10.0.0.0/24"
24  assign_generated_ipv6_cidr_block = true
25  enable_dns_support                = true
26  enable_dns_hostnames              = true
27
28  tags = { Name = "cheap-container" }
29}
30
31resource "aws_subnet" "main" {
32  vpc_id                          = aws_vpc.main.id
33  availability_zone                = data.aws_availability_zones.available.names[0]
34  cidr_block                       = "10.0.0.0/26"
35  ipv6_cidr_block                  = cidrsubnet(aws_vpc.main.ipv6_cidr_block, 8, 0)
36  assign_ipv6_address_on_creation  = true
37
38  tags = { Name = "cheap-container" }
39}

Now the egress-only internet gateway and the route table. Notice what's not here: there's no IPv4 route to anywhere. The private IPv4 address every task gets is reachable inside the VPC only — exactly what Cloud Map and the VPC Link need, and nothing more.

hcl
1resource "aws_egress_only_internet_gateway" "main" {
2  vpc_id = aws_vpc.main.id
3}
4
5resource "aws_route_table" "main" {
6  vpc_id = aws_vpc.main.id
7}
8
9resource "aws_route" "ipv6_egress" {
10  route_table_id              = aws_route_table.main.id
11  destination_ipv6_cidr_block = "::/0"
12  egress_only_gateway_id      = aws_egress_only_internet_gateway.main.id
13}
14
15resource "aws_route_table_association" "main" {
16  subnet_id      = aws_subnet.main.id
17  route_table_id = aws_route_table.main.id
18}

Step 2: Security Groups

Two groups: one for the API Gateway VPC Link's elastic network interfaces, one for the Fargate task. Using the dedicated rule resources (rather than inline ingress/egress blocks) means the two groups can reference each other without a circular-dependency error.

hcl
1resource "aws_security_group" "vpc_link" {
2  name   = "cheap-container-vpc-link"
3  vpc_id = aws_vpc.main.id
4}
5
6resource "aws_security_group" "task" {
7  name   = "cheap-container-task"
8  vpc_id = aws_vpc.main.id
9}
10
11resource "aws_vpc_security_group_egress_rule" "vpc_link_to_task" {
12  security_group_id            = aws_security_group.vpc_link.id
13  referenced_security_group_id = aws_security_group.task.id
14  from_port                    = 8080
15  to_port                      = 8080
16  ip_protocol                  = "tcp"
17}
18
19resource "aws_vpc_security_group_ingress_rule" "task_from_vpc_link" {
20  security_group_id            = aws_security_group.task.id
21  referenced_security_group_id = aws_security_group.vpc_link.id
22  from_port                    = 8080
23  to_port                      = 8080
24  ip_protocol                  = "tcp"
25}
26
27# The task's actual internet access — pulling the image, reaching any
28# IPv6-capable dependency — goes out over IPv6 only, matching the route
29# table from Step 1.
30resource "aws_vpc_security_group_egress_rule" "task_ipv6_egress" {
31  security_group_id = aws_security_group.task.id
32  cidr_ipv6          = "::/0"
33  ip_protocol        = "-1"
34}

Step 3: Cloud Map Service Discovery

A private DNS namespace scoped to the VPC, and a service inside it that the ECS service will register the running task with automatically.

hcl
1resource "aws_service_discovery_private_dns_namespace" "main" {
2  name = "internal"
3  vpc  = aws_vpc.main.id
4}
5
6resource "aws_service_discovery_service" "app" {
7  name = "app"
8
9  dns_config {
10    namespace_id = aws_service_discovery_private_dns_namespace.main.id
11
12    # SRV, not A — confirmed directly against AWS's API Gateway private
13    # integration docs: "If you use Amazon ECS to populate entries in AWS
14    # Cloud Map, you must configure your Amazon ECS task to use SRV
15    # records." An A record only carries an IP; API Gateway's VPC Link
16    # needs the port too, which only SRV records carry. Get this wrong and
17    # the API returns a bare 500 with no useful error message anywhere.
18    dns_records {
19      ttl  = 10
20      type = "SRV"
21    }
22
23    routing_policy = "MULTIVALUE"
24  }
25
26  # health_check_custom_config, not health_check_config — the latter is
27  # Route 53 health checks against a public endpoint, which doesn't apply
28  # to a private, VPC-only task. This is ECS-managed health instead: ECS
29  # tells Cloud Map directly when the task is healthy or gone.
30  health_check_custom_config {
31    failure_threshold = 1
32  }
33}

Step 4: The ECS Cluster, on Fargate Spot

hcl
1resource "aws_ecs_cluster" "main" {
2  name = "cheap-container"
3}
4
5resource "aws_ecs_cluster_capacity_providers" "main" {
6  cluster_name       = aws_ecs_cluster.main.name
7  capacity_providers = ["FARGATE_SPOT", "FARGATE"]
8
9  default_capacity_provider_strategy {
10    capacity_provider = "FARGATE_SPOT"
11    weight             = 1
12    base               = 0
13  }
14}
15
16resource "aws_iam_role" "execution" {
17  name = "cheap-container-execution"
18
19  assume_role_policy = jsonencode({
20    Version = "2012-10-17"
21    Statement = [{
22      Action    = "sts:AssumeRole"
23      Effect    = "Allow"
24      Principal = { Service = "ecs-tasks.amazonaws.com" }
25    }]
26  })
27}
28
29resource "aws_iam_role_policy_attachment" "execution" {
30  role       = aws_iam_role.execution.name
31  policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
32}

Step 5: The Task Definition and Service

ARM64, 0.25 vCPU / 0.5 GB — the exact size the cost numbers in the write-up are based on. This version deliberately has no logConfiguration — that's gotcha #2. We'll add CloudWatch logging back in as an explicit, priced opt-in at the end, rather than let the container silently time out trying to reach it.

hcl
1resource "aws_ecs_task_definition" "app" {
2  family                   = "cheap-container"
3  requires_compatibilities = ["FARGATE"]
4  network_mode              = "awsvpc"
5  cpu                        = 256
6  memory                     = 512
7  execution_role_arn        = aws_iam_role.execution.arn
8
9  runtime_platform {
10    cpu_architecture        = "ARM64"
11    operating_system_family = "LINUX"
12  }
13
14  container_definitions = jsonencode([
15    {
16      name      = "app"
17      image     = "docker.io/nginxinc/nginx-unprivileged:stable-alpine"
18      essential = true
19      portMappings = [
20        { containerPort = 8080, protocol = "tcp" }
21      ]
22    }
23  ])
24}
25
26resource "aws_ecs_service" "app" {
27  name            = "app"
28  cluster         = aws_ecs_cluster.main.id
29  task_definition = aws_ecs_task_definition.app.arn
30  desired_count   = 1
31
32  capacity_provider_strategy {
33    capacity_provider = "FARGATE_SPOT"
34    weight             = 1
35    base               = 1
36  }
37
38  network_configuration {
39    subnets          = [aws_subnet.main.id]
40    security_groups  = [aws_security_group.task.id]
41    assign_public_ip = false
42  }
43
44  service_registries {
45    registry_arn   = aws_service_discovery_service.app.arn
46    container_name = "app"
47    container_port = 8080
48  }
49
50  depends_on = [aws_ecs_cluster_capacity_providers.main]
51}

nginx-unprivileged listens on 8080 without needing root — a reasonable stand-in for whatever your own image is. Swap the image line for your own ECR or Docker Hub image and the rest of this doesn't change.

This is where Cloud Map and API Gateway actually connect. integration_uri takes the Cloud Map service's ARN directly — API Gateway resolves the current task's address itself at request time.

hcl
1resource "aws_apigatewayv2_vpc_link" "main" {
2  name               = "cheap-container"
3  security_group_ids = [aws_security_group.vpc_link.id]
4  subnet_ids         = [aws_subnet.main.id]
5}
6
7resource "aws_apigatewayv2_api" "main" {
8  name          = "cheap-container"
9  protocol_type = "HTTP"
10}
11
12resource "aws_apigatewayv2_integration" "app" {
13  api_id             = aws_apigatewayv2_api.main.id
14  integration_type   = "HTTP_PROXY"
15  integration_uri    = aws_service_discovery_service.app.arn
16  integration_method = "ANY"
17  connection_type    = "VPC_LINK"
18  connection_id      = aws_apigatewayv2_vpc_link.main.id
19}
20
21resource "aws_apigatewayv2_route" "app" {
22  api_id    = aws_apigatewayv2_api.main.id
23  route_key = "ANY /{proxy+}"
24  target    = "integrations/${aws_apigatewayv2_integration.app.id}"
25}
26
27resource "aws_apigatewayv2_stage" "main" {
28  api_id      = aws_apigatewayv2_api.main.id
29  name        = "$default"
30  auto_deploy = true
31}
32
33output "api_url" {
34  value = aws_apigatewayv2_api.main.api_endpoint
35}

Step 7: Apply and Test

bash
terraform init
terraform apply

This takes a few minutes — mostly the VPC Link, which provisions its own elastic network interfaces in the background. Once it's done:

bash
curl "$(terraform output -raw api_url)/"

You should get nginx's default welcome page back, served by a Fargate Spot task that has no load balancer and no NAT Gateway anywhere in its path. On a fresh apply, give it a beat — the first request or two can come back as a 503 for 30-60 seconds while the VPC Link catches up with everything else, even after ECS already reports the task healthy. See Common Issues below if it doesn't clear on its own.

Optional: Add CloudWatch Logging (the $7/Month Trade-off)

Left as-is, this container has no shipped logs — the awslogs driver can't reach CloudWatch over an IPv6-only egress path, so adding a logConfiguration block to Step 5's task definition without anything else would just make the task time out trying to log. If you want logs, the fix is a CloudWatch Logs VPC interface endpoint, reachable over the same private IPv4 path Cloud Map already uses:

hcl
1resource "aws_cloudwatch_log_group" "app" {
2  name              = "/ecs/cheap-container"
3  retention_in_days = 14
4}
5
6resource "aws_security_group" "vpc_endpoints" {
7  name   = "cheap-container-endpoints"
8  vpc_id = aws_vpc.main.id
9}
10
11resource "aws_vpc_security_group_ingress_rule" "endpoints_from_task" {
12  security_group_id            = aws_security_group.vpc_endpoints.id
13  referenced_security_group_id = aws_security_group.task.id
14  from_port                    = 443
15  to_port                      = 443
16  ip_protocol                  = "tcp"
17}
18
19resource "aws_vpc_endpoint" "logs" {
20  vpc_id              = aws_vpc.main.id
21  service_name        = "com.amazonaws.${var.region}.logs"
22  vpc_endpoint_type    = "Interface"
23  subnet_ids           = [aws_subnet.main.id]
24  security_group_ids   = [aws_security_group.vpc_endpoints.id]
25  private_dns_enabled  = true
26}

Then add the matching logConfiguration to the container definition in Step 5:

hcl
1logConfiguration = {
2  logDriver = "awslogs"
3  options = {
4    "awslogs-group"         = aws_cloudwatch_log_group.app.name
5    "awslogs-region"        = var.region
6    "awslogs-stream-prefix" = "app"
7  }
8}

That endpoint runs about $0.01/hour per AZ (roughly $7/month for one AZ) plus a small per-GB processing charge — real money relative to the rest of this setup, which is the entire point of calling it out rather than silently adding it.

Common Issues

  • Right after terraform apply finishes, curl returns a 503 even though ECS reports the service healthy. This is propagation lag, not a bug — on a fresh deploy we saw this take 30-60 seconds to clear after aws ecs describe-services already showed runningCount: 1 and rolloutState: COMPLETED, and Cloud Map's discover-instances already showed the instance as HEALTHY with a correct AWS_INSTANCE_PORT. The VPC Link/API Gateway layer takes a little longer to catch up than ECS and Cloud Map do. Just retry after a short wait; if it's still 503 after a couple of minutes, something else is wrong and it's worth re-checking the earlier steps.
  • terraform apply succeeds, everything shows healthy, and curl just returns a bare {"message":"Internal Server Error"} with no other detail anywhere. This is the SRV-vs-A mistake from Step 3 — if dns_records.type is "A" instead of "SRV", Cloud Map never gets an AWS_INSTANCE_PORT attribute, and API Gateway has no way to know what port to connect to. Confirm with aws servicediscovery discover-instances --namespace-name internal --service-name app — if AWS_INSTANCE_PORT is missing from the instance's attributes, that's the bug. Fixing dns_records.type forces a replacement of the Cloud Map service; if the old one won't delete because of "Service contains registered instances," deregister the task manually first: aws servicediscovery deregister-instance --service-id <old-id> --instance-id <task-id>, then re-apply. The ECS service also needs a fresh deployment to pick up the new registration, which happens automatically but takes a minute or two.
  • curl returns "no target endpoints found." Almost always means the subnet ended up IPv6-only somewhere, or the task's security group doesn't allow the VPC Link's security group in on the container port. Confirm the task actually got a private IPv4 address: aws ecs describe-tasks and check the attachments block for a privateIPv4Address.
  • The service never reaches RUNNING. Check aws ecs describe-services for the events list — a Spot capacity shortfall in that AZ is the most common cause this early. Retrying, or adding "FARGATE" as a fallback in capacity_provider_strategy, both work.
  • The container can't pull its image. If you swapped in a private ECR image, confirm ECR's public endpoint path applies (this setup has no NAT Gateway, so IPv4-only private resources won't be reachable at all) — see the write-up for exactly what IPv6-only egress can and can't reach.
  • terraform destroy hangs on the VPC Link or the subnet. The VPC Link's own ENIs can take a few minutes to release after the link itself is deleted — terraform destroy will eventually succeed; re-running it after a short wait clears any transient DependencyViolation.

Frequently Asked Questions

Can I run more than one container behind this setup?

Yes — add another Cloud Map service, task definition, and ECS service, then another aws_apigatewayv2_integration and aws_apigatewayv2_route pointing at a different path. They can all share the same VPC Link and subnet.

Does desired_count = 1 mean there's no redundancy?

Correct, and that's a deliberate trade-off this architecture makes for the price — see the write-up's discussion of what you're giving up against a real ALB/NAT-backed setup. Bumping desired_count and capacity_provider_strategy's weights gives you more replicas, but Cloud Map's DiscoverInstances call doesn't health-check the way an ALB target group does, so more replicas alone isn't the same safety net.

Why nginx-unprivileged instead of the regular nginx image?

Fargate tasks don't run as root by default in many hardened configurations, and the regular nginx image's default config needs port 80, which needs root. nginx-unprivileged listens on 8080 as a non-root user out of the box — one less thing to fight with for a tutorial that's about networking, not container hardening.

Official References

For the architecture rationale, the verified cost breakdown, and why this isn't necessarily a production-ready pattern on its own, read the full write-up.

We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.

Struggling with this in production?

We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.