Loading...

Thanos vs Cortex vs Mimir: how to choose

All three solve the same Prometheus limitations: a single server does not scale horizontally, does not retain data cheaply for long periods, and gives you no global view across clusters. They differ mainly in architecture and in how much of it you have to run.

Thanos attaches a sidecar to existing Prometheus servers, ships blocks to object storage, and answers queries by fanning out across sidecars and store gateways. It is additive — Prometheus keeps scraping exactly as before — which makes it the least disruptive to adopt.

Cortex and Mimir take remote-written samples into a purpose-built distributed system. Mimir is Grafana's fork of Cortex and is where most of that lineage's development has gone, with a strong focus on making a large deployment operable. Both are more capable at very large scale and more to operate than Thanos.

Decision matrix: which one fits your situation

Your situationUseWhy
Existing Prometheus, want long retention and a global viewThanosSidecar model is additive; Prometheus keeps working as it does today.
Very large scale, many tenantsMimirBuilt for horizontal scale and multi-tenancy from the start.
Already invested in CortexCortexMigration has cost and Mimir's advantages may not justify it yet.
Small team, minimal moving partsThanosFewer components to run and reason about.
Want commercial support and Grafana alignmentMimirActively developed by Grafana Labs with a clear upgrade story.
Just need longer retention on one PrometheusNeitherBigger disk and longer local retention may be enough. Do not adopt a distributed system for one server.

Object storage cost is the real economics

All three store blocks in object storage, which is why long retention becomes affordable. The line item people forget is request cost — queries over long ranges fetch many objects, and at scale the API request charges can rival storage. Caching layers exist in all three specifically for this and are not optional at volume.

Before adopting any of them, check whether the actual requirement is long retention of everything or long retention of a subset. Downsampling and recording rules on a single Prometheus solve a surprising number of these problems at a fraction of the complexity, and a distributed metrics system is a large thing to run for a requirement that a bigger disk would have met.

Frequently asked questions

Is Mimir just a rebranded Cortex?

It began as a fork and has diverged substantially, with significant work on scalability and operational simplicity. If you are choosing fresh today, Mimir is generally the more actively developed path. If you already run Cortex successfully, that is not on its own a reason to migrate.

Can I use these without changing my Prometheus setup?

Thanos, largely yes — add a sidecar and keep scraping as you do. Cortex and Mimir expect remote write, which is a Prometheus configuration change and a different data path, though not a large one. Thanos is the lower-friction adoption; the others may fit better at scale.

Do I need one of these at all?

Often not. If you have a handful of clusters and need weeks rather than years of retention, a well-sized Prometheus per cluster with Grafana querying multiple datasources is simpler and adequate. Adopt these when you genuinely need a unified query across clusters, retention measured in months or years, or multi-tenancy with isolation.

How do they handle high availability?

All support running Prometheus in HA pairs, but they deduplicate at different points. Thanos merges replicas at query time via a replica label, so both copies are stored. Cortex and Mimir elect a leader in the distributor's HA tracker and drop the other replica's samples on the write path, so only one is stored. Test a replica failure during evaluation — this difference is a common source of confusing gaps and double-counting in production.

Need this managed for you, not just automated?

We're also a hands-on DevOps consultancy — Kubernetes, CI/CD, and cloud infrastructure.

Explore Our Services