Cut Your Observability Bill with an OpenTelemetry Collector: Tail Sampling and Cardinality Control
Quick answer
Your observability bill scales with spans, active series, and GB ingested — not with how much of that data you actually query. This tutorial deploys an OpenTelemetry Collector gateway that keeps every error and slow trace, samples the boring ones, and strips high-cardinality labels before anything hits your backend.
- Step 1: Add the Helm Repo and Inspect the Chart
- Step 2: Understand the Gateway Topology
- Step 3: Write the Gateway Config — Traces
- Step 4: Add the Metrics Pipeline — Kill High Cardinality
- Step 5: Define Exporters and Assemble the Pipelines
intermediate · 70 min
Before you begin
- A Kubernetes cluster (kind or minikube is fine) with kubectl and Helm configured
- An app emitting OTLP traces and metrics (the OpenTelemetry demo works well)
- A tracing/metrics backend to export to (or just the debug exporter for verification)
- Basic familiarity with the OpenTelemetry Collector — receivers, processors, exporters
Observability vendors bill on volume: spans ingested, active metric time series, and gigabytes stored. The uncomfortable truth is that most of that data is never queried. You keep 100% of your traces so you can find the 0.1% that failed, and you carry high-cardinality metric labels like user_id and request_id that explode your active-series count without ever appearing in a dashboard. The bill grows linearly with traffic; the value does not.
The fix isn't sampling blindly at the SDK — that throws away the errors you most need. The fix is a gateway OpenTelemetry Collector that sees the full picture and makes intelligent decisions before export: keep every failed and slow trace, probabilistically sample the healthy ones, and drop label dimensions nobody queries. This tutorial builds exactly that pipeline. For the why behind the economics — especially once AI and LLM workloads enter the mix — read the true cost of observing AI workloads.
This assumes you already have a Collector running the basics. If you don't, the OpenTelemetry Collector on Kubernetes guide covers the agent/gateway topology and OTLP plumbing this tutorial builds on.
What You'll Build
- A gateway Collector deployed via the
open-telemetry/opentelemetry-collectorHelm chart as a single-instance Deployment - A
tail_samplingprocessor that keeps all errors and slow traces, and probabilistically samples the rest - A metrics pipeline that drops high-cardinality labels (
request_id,user_id) and filters out unused series memory_limiter+batchprocessors ordered correctly so the Collector protects itself and exports efficiently- Dual export: your real backend plus the
debugexporter, so you can watch the cost-control decisions happen
Step 1: Add the Helm Repo and Inspect the Chart
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update
helm show values open-telemetry/opentelemetry-collector | head -n 40The chart supports several deployment modes. We want mode: deployment (not daemonset) with a single replica — this is the critical constraint for tail sampling, which we'll explain in Step 3.
Step 2: Understand the Gateway Topology
The recommended production layout is two tiers:
- Agents (a DaemonSet, one per node) receive OTLP from your apps, add node/pod resource attributes, and forward everything to the gateway.
- Gateway (a Deployment) is where the expensive, stateful processors live — tail sampling, cardinality reduction, batching for export.
This tutorial focuses on the gateway. Your apps (or agents) send OTLP to it; the gateway decides what survives and exports the survivors. Keeping the cost-control logic in one tier means one place to reason about, and — importantly — tail sampling needs all spans of a trace to arrive at the same Collector instance.
Step 3: Write the Gateway Config — Traces
Create gateway-values.yaml. Start with the traces pipeline. The key processor is tail_sampling: unlike head sampling (which decides at the start of a trace, before you know if it errored), tail sampling buffers spans and decides once the trace is complete.
1mode: deployment
2replicaCount: 1 # single instance — see the note below
3
4image:
5 repository: otel/opentelemetry-collector-contrib # tail_sampling lives in contrib
6
7config:
8 receivers:
9 otlp:
10 protocols:
11 grpc:
12 endpoint: 0.0.0.0:4317
13 http:
14 endpoint: 0.0.0.0:4318
15
16 processors:
17 memory_limiter:
18 check_interval: 1s
19 limit_percentage: 80
20 spike_limit_percentage: 25
21
22 tail_sampling:
23 decision_wait: 10s # how long to buffer a trace before deciding
24 num_traces: 50000 # max traces kept in memory at once
25 expected_new_traces_per_sec: 1000
26 policies:
27 # 1. Always keep traces with an error status.
28 - name: keep-errors
29 type: status_code
30 status_code:
31 status_codes: [ERROR]
32 # 2. Always keep slow traces (>= 500ms end to end).
33 - name: keep-slow
34 type: latency
35 latency:
36 threshold_ms: 500
37 # 3. Sample 10% of everything else.
38 - name: sample-the-rest
39 type: probabilistic
40 probabilistic:
41 sampling_percentage: 10
42
43 batch:
44 send_batch_size: 8192
45 timeout: 5sThree things about tail_sampling that trip people up:
- Policies are OR-ed. A trace is kept if any policy says keep. So errors and slow traces are always retained; the
probabilisticpolicy only governs the traces the first two didn't already catch. decision_waitmust exceed your trace duration. If a trace takes 8s and you setdecision_wait: 2s, the Collector decides before the trace is complete and may drop the error span that arrives late. Set it comfortably above your p99 latency.- Single instance is mandatory here. Tail sampling requires every span of a trace to reach the same Collector. Two replicas behind a normal Service will split a trace across both, and each sees a partial trace. Step 8 covers how to scale this correctly.
Step 4: Add the Metrics Pipeline — Kill High Cardinality
Active series are the metrics equivalent of spans: every unique combination of label values is one billable series. A http_requests_total metric with a user_id label and 100,000 users is 100,000 series — for a single metric. You almost never query by user_id; you query by route and status. Drop the label, keep the signal.
Add these to the processors block in the same file:
1 # Delete high-cardinality attributes from every metric.
2 transform/drop-labels:
3 metric_statements:
4 - context: datapoint
5 statements:
6 - delete_key(attributes, "request_id")
7 - delete_key(attributes, "user_id")
8 - delete_key(attributes, "session_id")
9
10 # Drop entire metrics you never query (e.g. noisy Go runtime internals).
11 filter/drop-unused:
12 metrics:
13 metric:
14 - 'name == "process_runtime_go_gc_pause_ns"'
15 - 'IsMatch(name, "^rpc_client_.*_bucket$")'transform uses OTTL (the OpenTelemetry Transformation Language). delete_key(attributes, "user_id") runs per datapoint and collapses what used to be thousands of series into a handful. filter/drop-unused removes whole metrics by name — check your backend's "series by metric name" view first so you drop things nobody dashboards on, not something load-bearing.
Step 5: Define Exporters and Assemble the Pipelines
Finish gateway-values.yaml with exporters and the service block that wires everything together. The debug exporter prints to the Collector's logs so you can verify decisions without touching your backend.
1 exporters:
2 # Your real backend. Swap the endpoint/headers for your vendor or Tempo/Prometheus.
3 otlphttp/backend:
4 endpoint: https://otlp.your-backend.example.com
5 headers:
6 authorization: "Bearer ${env:BACKEND_TOKEN}"
7
8 # Prints a summary of what survived — remove once you've verified.
9 debug:
10 verbosity: normal
11
12 service:
13 pipelines:
14 traces:
15 receivers: [otlp]
16 processors: [memory_limiter, tail_sampling, batch]
17 exporters: [otlphttp/backend, debug]
18 metrics:
19 receivers: [otlp]
20 processors: [memory_limiter, transform/drop-labels, filter/drop-unused, batch]
21 exporters: [otlphttp/backend, debug]Processor order matters and is not cosmetic:
memory_limitergoes first so it can reject data and shed load before any expensive work happens — it protects the Collector from OOMing under a traffic spike.tail_sampling(traces) and the cardinality processors (metrics) run next, while data is still per-item and cheap to inspect.batchgoes last, right before export, so it groups the survivors into efficient payloads. Batching before sampling would waste work grouping spans you're about to drop.
Step 6: Deploy the Gateway
Store the backend token as a secret and reference it, rather than baking it into values:
1kubectl create namespace otel
2
3kubectl create secret generic otel-backend \
4 --namespace otel \
5 --from-literal=token='YOUR_BACKEND_TOKEN'
6
7helm install otel-gateway open-telemetry/opentelemetry-collector \
8 --namespace otel \
9 --values gateway-values.yaml \
10 --set-string 'extraEnvs[0].name=BACKEND_TOKEN' \
11 --set-string 'extraEnvs[0].valueFrom.secretKeyRef.name=otel-backend' \
12 --set-string 'extraEnvs[0].valueFrom.secretKeyRef.key=token'Confirm it came up and the config parsed cleanly:
kubectl -n otel rollout status deployment/otel-gateway-opentelemetry-collector
kubectl -n otel logs deployment/otel-gateway-opentelemetry-collector | grep -i "everything is ready"Step 7: Point Traffic at the Gateway and Verify
Send OTLP to the gateway's Service on port 4317 (gRPC) or 4318 (HTTP). For the OpenTelemetry demo or any instrumented app, set the exporter endpoint:
# In your app / agent deployment:
# OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-gateway-opentelemetry-collector.otel.svc:4318Now watch the debug exporter to confirm the cost-control logic is firing:
kubectl -n otel logs -f deployment/otel-gateway-opentelemetry-collectorGenerate some load, including a few failing and slow requests. You should see error and slow traces consistently exported while healthy ones appear at roughly 10%. On the metrics side, inspect a datapoint's attributes in the debug output — user_id and request_id should be gone. Once you trust it, remove debug from the exporters lists (it's verbose and adds overhead) and helm upgrade.
Step 8: Scale Without Breaking Tail Sampling
A single gateway replica eventually becomes a bottleneck. You cannot just bump replicaCount — that splits traces across instances and breaks sampling decisions. The correct pattern is a two-tier setup: a small layer of Collectors running only the load_balancing exporter, which hashes on trace ID and routes every span of a given trace to the same backend gateway.
1# On the routing tier (not the gateway) — routes by trace ID:
2exporters:
3 load_balancing:
4 routing_key: traceID
5 protocol:
6 otlp:
7 tls:
8 insecure: true
9 resolver:
10 k8s:
11 service: otel-gateway-opentelemetry-collector.otelWith routing_key: traceID, you can now run many gateway replicas: each receives complete traces, so tail_sampling makes correct decisions. This is the single most common tail-sampling production mistake — scaling the sampling tier directly instead of putting a trace-ID-aware load balancer in front of it.
Common Issues
- Error traces missing from the backend —
decision_waitis shorter than your trace duration, so the Collector decided before the error span arrived. Raisedecision_waitabove your p99 end-to-end latency. - Traces look incomplete / spans missing — you scaled
replicaCountabove 1 without aload_balancingtier. Spans of one trace are landing on different instances. Go back to Step 8. - Collector OOMKilled under load —
memory_limiterisn't first in the pipeline, ornum_tracesis too high for the pod's memory.tail_samplingbuffers traces in RAM; lowernum_tracesor raise the memory request. - A dashboard broke after deploy — you dropped a label or metric you actually query in
transform/drop-labelsorfilter/drop-unused. Check your backend's query history before dropping, and reintroduce the needed dimension. unknown type: tail_sampling— you're running the core Collector image, not contrib. Setimage.repository: otel/opentelemetry-collector-contrib.
Frequently Asked Questions
How is tail sampling different from head sampling, and why does it save money without losing errors?
Head sampling decides at the start of a trace — before you know whether it errored or was slow — so to keep all errors you'd have to keep everything. Tail sampling buffers the whole trace and decides after it completes, so it can keep 100% of errors and slow traces while sampling the healthy majority at, say, 10%. You cut ingested span volume dramatically while retaining exactly the traces you'd actually investigate.
Why must the tail sampling Collector run as a single instance?
tail_sampling needs every span of a trace to arrive at the same Collector so it can evaluate the complete trace. With multiple replicas behind an ordinary Service, spans of one trace get distributed across instances, and each sees only a fragment — so the error span might land on a different replica than the one making the keep/drop decision. To scale, put a load_balancing exporter tier with routing_key: traceID in front, as shown in Step 8.
How does dropping metric labels reduce cost?
Backends bill on active time series, and each unique combination of label values is a separate series. A single metric with a user_id label and 100,000 users produces 100,000 series. Since you almost never query by user_id, deleting that attribute with transform's delete_key collapses those into a handful of series while preserving the dimensions you actually chart, like route and status_code.
What value should I use for decision_wait?
Set it comfortably above your p99 end-to-end trace duration — 10s is a reasonable starting point for typical web services. Too low and the Collector decides before slow or late-arriving spans land, dropping the very traces you wanted to keep. Too high and the Collector buffers more traces in memory, raising the risk of hitting num_traces and increasing memory pressure. Tune it against your real latency distribution.
Can I apply cost controls at the SDK instead of the Collector?
You can head-sample or drop attributes in the SDK, but you lose the two things that make the Collector approach powerful: it sees complete traces (so it can tail-sample on error/latency), and it's a central control point you can retune with a config change instead of redeploying every service. Keep SDK-level sampling conservative (or off) and concentrate cost decisions in the gateway.
Tear Down
1# Remove the gateway release:
2helm uninstall otel-gateway --namespace otel
3
4# Remove the secret and namespace:
5kubectl delete secret otel-backend --namespace otel
6kubectl delete namespace otel
7
8# Local kind cluster, if you spun one up just for this:
9kind delete clusterOfficial References
- OpenTelemetry Collector — Configuration — Receivers, processors, exporters, and the service pipeline model
- tail_sampling processor — Full list of policies,
decision_wait, andnum_tracessemantics - Scaling the Collector / trace-ID load balancing — Why tail sampling needs the
load_balancingexporter to scale - transform processor & OTTL — Editing attributes and metrics with the OpenTelemetry Transformation Language
- opentelemetry-collector Helm chart — Deployment modes, values reference, and the agent/gateway pattern
We built Podscape to simplify Kubernetes workflows like this — logs, events, and cluster state in one interface, without switching tools.
Struggling with this in production?
We help teams fix these exact issues. Our engineers have deployed these patterns across production environments at scale.