Build the Cheapest Container Setup on AWS With Terraform
Quick answer
A real, working Terraform configuration for running a Fargate Spot container on AWS with no ALB and no NAT Gateway — API Gateway, Cloud Map, and a dual-stack subnet with IPv6-only egress, built step by step and including the fix for the one gotcha that breaks a naive attempt.
- Step 1: A Dual-Stack VPC and Subnet
- Step 2: Security Groups
- Step 3: Cloud Map Service Discovery
- Step 4: The ECS Cluster, on Fargate Spot
- Step 5: The Task Definition and Service
intermediate · 40 min
Before you begin
- An AWS account with credentials configured for Terraform (AWS CLI or environment variables)
- Terraform 1.5 or later
- Basic familiarity with VPCs, ECS, and Terraform syntax
- Comfort reading HCL — this tutorial builds a real multi-resource configuration, not a single-file toy
The write-up on skipping the ALB and NAT Gateway for a small Fargate container covers the architecture and the cost math. This is the hands-on companion: a complete, working Terraform configuration for it, built resource by resource, including the fix for the one gotcha that will otherwise leave you staring at "no target endpoints found" for an hour.
By the end you'll have a single ARM64 Fargate Spot task, discovered by AWS Cloud Map, reachable through an API Gateway HTTP API — no load balancer, no NAT Gateway, costing a few dollars a month.
What You'll Build
- A VPC with a dual-stack subnet: a normal private IPv4 range for in-VPC service discovery, plus IPv6 with a default route out through an egress-only internet gateway — no NAT Gateway anywhere
- An ECS cluster running Fargate Spot, with a single ARM64 task
- Cloud Map private DNS service discovery registering the task automatically as it starts and stops
- An API Gateway HTTP API with a VPC Link, routing public requests straight to the task through Cloud Map — no ALB, no target group
- The SRV-record fix Cloud Map needs for this integration to work at all, applied from the start instead of discovered via a bare, unhelpful 500 error
Step 1: A Dual-Stack VPC and Subnet
The subnet needs to do two different jobs: give the task a private IPv4 address (Cloud Map and the VPC Link only resolve IPv4 targets — this is gotcha #1 from the write-up, and building the subnet dual-stack from the start sidesteps it entirely), and route the task's actual internet-bound traffic out over IPv6, with no NAT Gateway involved.
1terraform {
2 required_providers {
3 aws = {
4 source = "hashicorp/aws"
5 version = "~> 5.0"
6 }
7 }
8}
9
10variable "region" {
11 default = "us-east-1"
12}
13
14provider "aws" {
15 region = var.region
16}
17
18data "aws_availability_zones" "available" {
19 state = "available"
20}
21
22resource "aws_vpc" "main" {
23 cidr_block = "10.0.0.0/24"
24 assign_generated_ipv6_cidr_block = true
25 enable_dns_support = true
26 enable_dns_hostnames = true
27
28 tags = { Name = "cheap-container" }
29}
30
31resource "aws_subnet" "main" {
32 vpc_id = aws_vpc.main.id
33 availability_zone = data.aws_availability_zones.available.names[0]
34 cidr_block = "10.0.0.0/26"
35 ipv6_cidr_block = cidrsubnet(aws_vpc.main.ipv6_cidr_block, 8, 0)
36 assign_ipv6_address_on_creation = true
37
38 tags = { Name = "cheap-container" }
39}Now the egress-only internet gateway and the route table. Notice what's not here: there's no IPv4 route to anywhere. The private IPv4 address every task gets is reachable inside the VPC only — exactly what Cloud Map and the VPC Link need, and nothing more.
1resource "aws_egress_only_internet_gateway" "main" {
2 vpc_id = aws_vpc.main.id
3}
4
5resource "aws_route_table" "main" {
6 vpc_id = aws_vpc.main.id
7}
8
9resource "aws_route" "ipv6_egress" {
10 route_table_id = aws_route_table.main.id
11 destination_ipv6_cidr_block = "::/0"
12 egress_only_gateway_id = aws_egress_only_internet_gateway.main.id
13}
14
15resource "aws_route_table_association" "main" {
16 subnet_id = aws_subnet.main.id
17 route_table_id = aws_route_table.main.id
18}Step 2: Security Groups
Two groups: one for the API Gateway VPC Link's elastic network interfaces, one for the Fargate task. Using the dedicated rule resources (rather than inline ingress/egress blocks) means the two groups can reference each other without a circular-dependency error.
1resource "aws_security_group" "vpc_link" {
2 name = "cheap-container-vpc-link"
3 vpc_id = aws_vpc.main.id
4}
5
6resource "aws_security_group" "task" {
7 name = "cheap-container-task"
8 vpc_id = aws_vpc.main.id
9}
10
11resource "aws_vpc_security_group_egress_rule" "vpc_link_to_task" {
12 security_group_id = aws_security_group.vpc_link.id
13 referenced_security_group_id = aws_security_group.task.id
14 from_port = 8080
15 to_port = 8080
16 ip_protocol = "tcp"
17}
18
19resource "aws_vpc_security_group_ingress_rule" "task_from_vpc_link" {
20 security_group_id = aws_security_group.task.id
21 referenced_security_group_id = aws_security_group.vpc_link.id
22 from_port = 8080
23 to_port = 8080
24 ip_protocol = "tcp"
25}
26
27# The task's actual internet access — pulling the image, reaching any
28# IPv6-capable dependency — goes out over IPv6 only, matching the route
29# table from Step 1.
30resource "aws_vpc_security_group_egress_rule" "task_ipv6_egress" {
31 security_group_id = aws_security_group.task.id
32 cidr_ipv6 = "::/0"
33 ip_protocol = "-1"
34}Step 3: Cloud Map Service Discovery
A private DNS namespace scoped to the VPC, and a service inside it that the ECS service will register the running task with automatically.
1resource "aws_service_discovery_private_dns_namespace" "main" {
2 name = "internal"
3 vpc = aws_vpc.main.id
4}
5
6resource "aws_service_discovery_service" "app" {
7 name = "app"
8
9 dns_config {
10 namespace_id = aws_service_discovery_private_dns_namespace.main.id
11
12 # SRV, not A — confirmed directly against AWS's API Gateway private
13 # integration docs: "If you use Amazon ECS to populate entries in AWS
14 # Cloud Map, you must configure your Amazon ECS task to use SRV
15 # records." An A record only carries an IP; API Gateway's VPC Link
16 # needs the port too, which only SRV records carry. Get this wrong and
17 # the API returns a bare 500 with no useful error message anywhere.
18 dns_records {
19 ttl = 10
20 type = "SRV"
21 }
22
23 routing_policy = "MULTIVALUE"
24 }
25
26 # health_check_custom_config, not health_check_config — the latter is
27 # Route 53 health checks against a public endpoint, which doesn't apply
28 # to a private, VPC-only task. This is ECS-managed health instead: ECS
29 # tells Cloud Map directly when the task is healthy or gone.
30 health_check_custom_config {
31 failure_threshold = 1
32 }
33}Step 4: The ECS Cluster, on Fargate Spot
1resource "aws_ecs_cluster" "main" {
2 name = "cheap-container"
3}
4
5resource "aws_ecs_cluster_capacity_providers" "main" {
6 cluster_name = aws_ecs_cluster.main.name
7 capacity_providers = ["FARGATE_SPOT", "FARGATE"]
8
9 default_capacity_provider_strategy {
10 capacity_provider = "FARGATE_SPOT"
11 weight = 1
12 base = 0
13 }
14}
15
16resource "aws_iam_role" "execution" {
17 name = "cheap-container-execution"
18
19 assume_role_policy = jsonencode({
20 Version = "2012-10-17"
21 Statement = [{
22 Action = "sts:AssumeRole"
23 Effect = "Allow"
24 Principal = { Service = "ecs-tasks.amazonaws.com" }
25 }]
26 })
27}
28
29resource "aws_iam_role_policy_attachment" "execution" {
30 role = aws_iam_role.execution.name
31 policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
32}Step 5: The Task Definition and Service
ARM64, 0.25 vCPU / 0.5 GB — the exact size the cost numbers in the write-up are based on. This version deliberately has no logConfiguration — that's gotcha #2. We'll add CloudWatch logging back in as an explicit, priced opt-in at the end, rather than let the container silently time out trying to reach it.
1resource "aws_ecs_task_definition" "app" {
2 family = "cheap-container"
3 requires_compatibilities = ["FARGATE"]
4 network_mode = "awsvpc"
5 cpu = 256
6 memory = 512
7 execution_role_arn = aws_iam_role.execution.arn
8
9 runtime_platform {
10 cpu_architecture = "ARM64"
11 operating_system_family = "LINUX"
12 }
13
14 container_definitions = jsonencode([
15 {
16 name = "app"
17 image = "docker.io/nginxinc/nginx-unprivileged:stable-alpine"
18 essential = true
19 portMappings = [
20 { containerPort = 8080, protocol = "tcp" }
21 ]
22 }
23 ])
24}
25
26resource "aws_ecs_service" "app" {
27 name = "app"
28 cluster = aws_ecs_cluster.main.id
29 task_definition = aws_ecs_task_definition.app.arn
30 desired_count = 1
31
32 capacity_provider_strategy {
33 capacity_provider = "FARGATE_SPOT"
34 weight = 1
35 base = 1
36 }
37
38 network_configuration {
39 subnets = [aws_subnet.main.id]
40 security_groups = [aws_security_group.task.id]
41 assign_public_ip = false
42 }
43
44 service_registries {
45 registry_arn = aws_service_discovery_service.app.arn
46 container_name = "app"
47 container_port = 8080
48 }
49
50 depends_on = [aws_ecs_cluster_capacity_providers.main]
51}nginx-unprivileged listens on 8080 without needing root — a reasonable stand-in for whatever your own image is. Swap the image line for your own ECR or Docker Hub image and the rest of this doesn't change.
Step 6: API Gateway HTTP API, Routed Through a VPC Link
This is where Cloud Map and API Gateway actually connect. integration_uri takes the Cloud Map service's ARN directly — API Gateway resolves the current task's address itself at request time.
1resource "aws_apigatewayv2_vpc_link" "main" {
2 name = "cheap-container"
3 security_group_ids = [aws_security_group.vpc_link.id]
4 subnet_ids = [aws_subnet.main.id]
5}
6
7resource "aws_apigatewayv2_api" "main" {
8 name = "cheap-container"
9 protocol_type = "HTTP"
10}
11
12resource "aws_apigatewayv2_integration" "app" {
13 api_id = aws_apigatewayv2_api.main.id
14 integration_type = "HTTP_PROXY"
15 integration_uri = aws_service_discovery_service.app.arn
16 integration_method = "ANY"
17 connection_type = "VPC_LINK"
18 connection_id = aws_apigatewayv2_vpc_link.main.id
19}
20
21resource "aws_apigatewayv2_route" "app" {
22 api_id = aws_apigatewayv2_api.main.id
23 route_key = "ANY /{proxy+}"
24 target = "integrations/${aws_apigatewayv2_integration.app.id}"
25}
26
27resource "aws_apigatewayv2_stage" "main" {
28 api_id = aws_apigatewayv2_api.main.id
29 name = "$default"
30 auto_deploy = true
31}
32
33output "api_url" {
34 value = aws_apigatewayv2_api.main.api_endpoint
35}Step 7: Apply and Test
terraform init
terraform applyThis takes a few minutes — mostly the VPC Link, which provisions its own elastic network interfaces in the background. Once it's done:
curl "$(terraform output -raw api_url)/"You should get nginx's default welcome page back, served by a Fargate Spot task that has no load balancer and no NAT Gateway anywhere in its path. On a fresh apply, give it a beat — the first request or two can come back as a 503 for 30-60 seconds while the VPC Link catches up with everything else, even after ECS already reports the task healthy. See Common Issues below if it doesn't clear on its own.
Optional: Add CloudWatch Logging (the $7/Month Trade-off)
Left as-is, this container has no shipped logs — the awslogs driver can't reach CloudWatch over an IPv6-only egress path, so adding a logConfiguration block to Step 5's task definition without anything else would just make the task time out trying to log. If you want logs, the fix is a CloudWatch Logs VPC interface endpoint, reachable over the same private IPv4 path Cloud Map already uses:
1resource "aws_cloudwatch_log_group" "app" {
2 name = "/ecs/cheap-container"
3 retention_in_days = 14
4}
5
6resource "aws_security_group" "vpc_endpoints" {
7 name = "cheap-container-endpoints"
8 vpc_id = aws_vpc.main.id
9}
10
11resource "aws_vpc_security_group_ingress_rule" "endpoints_from_task" {
12 security_group_id = aws_security_group.vpc_endpoints.id
13 referenced_security_group_id = aws_security_group.task.id
14 from_port = 443
15 to_port = 443
16 ip_protocol = "tcp"
17}
18
19resource "aws_vpc_endpoint" "logs" {
20 vpc_id = aws_vpc.main.id
21 service_name = "com.amazonaws.${var.region}.logs"
22 vpc_endpoint_type = "Interface"
23 subnet_ids = [aws_subnet.main.id]
24 security_group_ids = [aws_security_group.vpc_endpoints.id]
25 private_dns_enabled = true
26}Then add the matching logConfiguration to the container definition in Step 5:
1logConfiguration = {
2 logDriver = "awslogs"
3 options = {
4 "awslogs-group" = aws_cloudwatch_log_group.app.name
5 "awslogs-region" = var.region
6 "awslogs-stream-prefix" = "app"
7 }
8}That endpoint runs about $0.01/hour per AZ (roughly $7/month for one AZ) plus a small per-GB processing charge — real money relative to the rest of this setup, which is the entire point of calling it out rather than silently adding it.
Common Issues
- Right after
terraform applyfinishes,curlreturns a 503 even though ECS reports the service healthy. This is propagation lag, not a bug — on a fresh deploy we saw this take 30-60 seconds to clear afteraws ecs describe-servicesalready showedrunningCount: 1androlloutState: COMPLETED, and Cloud Map'sdiscover-instancesalready showed the instance asHEALTHYwith a correctAWS_INSTANCE_PORT. The VPC Link/API Gateway layer takes a little longer to catch up than ECS and Cloud Map do. Just retry after a short wait; if it's still 503 after a couple of minutes, something else is wrong and it's worth re-checking the earlier steps. terraform applysucceeds, everything shows healthy, andcurljust returns a bare{"message":"Internal Server Error"}with no other detail anywhere. This is the SRV-vs-A mistake from Step 3 — ifdns_records.typeis"A"instead of"SRV", Cloud Map never gets anAWS_INSTANCE_PORTattribute, and API Gateway has no way to know what port to connect to. Confirm withaws servicediscovery discover-instances --namespace-name internal --service-name app— ifAWS_INSTANCE_PORTis missing from the instance's attributes, that's the bug. Fixingdns_records.typeforces a replacement of the Cloud Map service; if the old one won't delete because of "Service contains registered instances," deregister the task manually first:aws servicediscovery deregister-instance --service-id <old-id> --instance-id <task-id>, then re-apply. The ECS service also needs a fresh deployment to pick up the new registration, which happens automatically but takes a minute or two.curlreturns "no target endpoints found." Almost always means the subnet ended up IPv6-only somewhere, or the task's security group doesn't allow the VPC Link's security group in on the container port. Confirm the task actually got a private IPv4 address:aws ecs describe-tasksand check theattachmentsblock for aprivateIPv4Address.- The service never reaches
RUNNING. Checkaws ecs describe-servicesfor theeventslist — a Spot capacity shortfall in that AZ is the most common cause this early. Retrying, or adding"FARGATE"as a fallback incapacity_provider_strategy, both work. - The container can't pull its image. If you swapped in a private ECR image, confirm ECR's public endpoint path applies (this setup has no NAT Gateway, so IPv4-only private resources won't be reachable at all) — see the write-up for exactly what IPv6-only egress can and can't reach.
terraform destroyhangs on the VPC Link or the subnet. The VPC Link's own ENIs can take a few minutes to release after the link itself is deleted —terraform destroywill eventually succeed; re-running it after a short wait clears any transientDependencyViolation.
Frequently Asked Questions
Can I run more than one container behind this setup?
Yes — add another Cloud Map service, task definition, and ECS service, then another aws_apigatewayv2_integration and aws_apigatewayv2_route pointing at a different path. They can all share the same VPC Link and subnet.
Does desired_count = 1 mean there's no redundancy?
Correct, and that's a deliberate trade-off this architecture makes for the price — see the write-up's discussion of what you're giving up against a real ALB/NAT-backed setup. Bumping desired_count and capacity_provider_strategy's weights gives you more replicas, but Cloud Map's DiscoverInstances call doesn't health-check the way an ALB target group does, so more replicas alone isn't the same safety net.
Why nginx-unprivileged instead of the regular nginx image?
Fargate tasks don't run as root by default in many hardened configurations, and the regular nginx image's default config needs port 80, which needs root. nginx-unprivileged listens on 8080 as a non-root user out of the box — one less thing to fight with for a tutorial that's about networking, not container hardening.
Official References
- aws_apigatewayv2_integration — HTTP_PROXY integrations, including Cloud Map as a valid
integration_uritarget - aws_service_discovery_service —
health_check_custom_configvs.health_check_config - aws_egress_only_internet_gateway — IPv6-only outbound routing
- aws_ecs_cluster_capacity_providers — enabling Fargate Spot on a cluster
For the architecture rationale, the verified cost breakdown, and why this isn't necessarily a production-ready pattern on its own, read the full write-up.
We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.
Struggling with this in production?
We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.