The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Observability

Metric Cardinality In Observability: Strategies

A deep dive into how modern TSDBs handle the label explosion problem and the trade-offs between storage- query speed- and data granularity
by Netdata Team · August 10, 2025

It’s a story familiar to any SRE or DevOps engineer. You add a seemingly innocuous label to a key metric—user_id, request_id, container_id—to gain deeper insight. Suddenly, your monitoring bill skyrockets, your Prometheus_TSDB instance starts gasping for memory, and dashboards slow to a crawl. You have just triggered a label_explosion, the single biggest challenge in modern metrics-based observability: metrics_cardinality.

High cardinality isn’t an edge case; it’s the new normal in a world of microservices, containers, and complex user interactions. As the number of unique time series grows into the millions or even billions, it places immense pressure on monitoring systems, impacting storage_costs, query performance, and the fundamental ability to scale.

There is no silver bullet for this problem, but two dominant strategies have emerged in the observability ecosystem: aggressive metric_compression and manual rollup_strategy. This guide will compare these two approaches, exploring how tools like Prometheus and Grafana_Mimir implement them, and discuss the critical trade-offs you make with each.

What is High-Cardinality and Why Is It a Problem?

In a time-series database (TSDB), a “time series” is a unique combination of a metric name and a set of key-value pairs called labels. Cardinality refers to the number of these unique combinations.

  • Low Cardinality: http_requests_total{method="GET", status="200", job="api-server"} This metric has a small, finite number of possible label combinations. Its cardinality is low and predictable.

  • High Cardinality: http_requests_total{..., client_ip="1.2.3.4", user_id="u-5678"} By adding labels with many unique values (IP addresses, user IDs, container IDs, session IDs), you create a combinatorial explosion. If you have 10,000 users making requests from 5,000 IP addresses, this single metric can generate millions of unique time series.

This label_explosion creates severe problems for most modern TSDBs:

  1. Massive Index Size: The biggest issue isn’t the raw data points; it’s the index. Every unique time series requires an entry in the database’s index so it can be found quickly. A massive index consumes vast amounts of RAM and disk space, driving up observability_cost.
  2. Slow Ingestion: The system struggles to process and index millions of new series arriving via remote_write or scraping, leading to ingestion delays and dropped data.
  3. Slow Queries: Queries that need to aggregate data across millions of series (e.g., sum(http_requests_total)) become painfully slow. The TSDB has to find every series in the index, load its data, and then perform the aggregation, consuming significant CPU and memory.

Strategy 1: Aggressive Compression (The Prometheus & Mimir Model)

Modern TSDBs like the Prometheus_TSDB and its horizontally-scalable implementations (Grafana_Mimir, Cortex, Thanos) don’t try to prevent high cardinality. Instead, they are engineered to manage its impact through sophisticated compression techniques.

In-Memory Index and Head Block

When new data arrives, it’s written to an in-memory “head block.” This is where active time series are indexed for fast writes and reads. High cardinality churns this head block rapidly, as new series are constantly being created. This leads to high memory usage, one of the first signs of a tsdb_scaling problem.

Persistent Storage and Chunk Compression

Periodically, the in-memory data is flushed to disk in immutable blocks, typically covering a two-hour window (tsdb_block_size). This is where metric_compression works its magic:

  • Data Point Compression: Timestamps and values are compressed using techniques like delta encoding and XOR, as pioneered by GorillaDB. This chunk_compression is incredibly efficient for the raw data points themselves.
  • Index Compression: The labels and series metadata are also heavily compressed using dictionaries and other methods to reduce the on-disk footprint.

However, while these techniques drastically reduce storage_costs, they do not solve the fundamental problem. The index, even when compressed on disk, is still logically massive. When you run a query, the system still has to decompress and process all that metadata to find the relevant series, which is why query performance degrades as cardinality grows. Compression makes storing the data feasible, but it doesn’t make querying it fast.

Strategy 2: Manual Roll-ups and Recording Rules (The Pre-aggregation Fix)

This is the traditional SRE approach to controlling cardinality. If you can’t afford to store and query the raw high_cardinality_metrics, you don’t. You pre-aggregate the data into a new, lower-cardinality metric.

The Power of Recording Rules

The primary tool for this in the Prometheus ecosystem is recording rules. These are queries that run at a regular interval (e.g., every minute), with the result saved as a new time series.

For example, imagine you have a high-cardinality metric for HTTP requests that includes a user_id label. You can create a rollup_strategy with a recording rule that removes this high-cardinality label and saves the aggregated total. This creates a new, much lower-cardinality metric.

Downsampling and Metrics Retention

This rollup_strategy is often paired with tiered metrics_retention policies. You might:

  1. Keep the raw, high-cardinality data for a short period (e.g., 24-48 hours) for fine-grained debugging.
  2. Drop the high-cardinality labels after that period.
  3. Keep the rolled-up, low-cardinality metrics for much longer (e.g., 13+ months) for long-term trending.

Systems like Grafana_Mimir and cortex_rollup provide dedicated components to manage this downsampling process efficiently at scale.

The Downsides of Manual Roll-ups

While effective at controlling costs and improving query performance, this approach has significant drawbacks:

  • Loss of Granularity: You’ve permanently thrown away the details. If an incident occurs and you need to know which user_id was causing a spike three days ago, you can’t. The raw data is gone. Exemplars can link back to raw traces, but they don’t solve the problem of being unable to query the metric itself.
  • Manual Toil: This process is incredibly brittle. Every time a new service introduces a high-cardinality metric, an SRE needs to manually write, test, and deploy a new recording rule. It’s a constant, error-prone battle to keep cardinality in check.
  • Inflexibility: You must decide what questions you will want to ask in the future, today. If you didn’t create a roll-up by_customer_id, you will never be able to answer questions about a specific customer’s experience from your long-term metrics.

A Third Way: Adaptive, Real-Time Solutions

The choice between raw, expensive data and cheap, aggregated data is a false dichotomy. Modern observability platforms are moving towards a model of adaptive_metrics that aims to provide the best of both worlds.

Netdata, for example, challenges this trade-off by moving intelligence to the edge. The Netdata Agent, running on each node, collects thousands of metrics every second. It can store this high-fidelity data locally for a configurable period, giving you raw, granular data for immediate debugging without any of the remote_write or ingestion bottlenecks.

When this data is streamed to a central backend, it can be intelligently downsampled over time without requiring manual recording rules. The system understands that as you “zoom out” on a chart from the last hour to the last month, you are interested in trends, not individual data points. This adaptive approach means:

  • You retain raw data for debugging when you need it most—in the recent past.
  • You get efficient, long-term storage for trends without manual configuration.
  • You avoid the “pre-aggregation trap”, maintaining the flexibility to explore your data in new ways without being limited by decisions you made months ago.

Other strategies like metric_deduplication, used by Mimir and Thanos, are also crucial for tsdb_scaling in high-availability setups, but they solve a different problem than cardinality itself.

Conclusion: Know Your Trade-Offs

There is no perfect solution for metrics_cardinality. The right strategy depends on your specific needs and budget.

  • The Compression-First Approach (Prometheus): This model gives you full detail but puts the burden on your query engine and your wallet. It’s powerful if you can afford the hardware and tolerate slower queries at scale.
  • The Roll-up Strategy (Mimir, Cortex): This model prioritizes query performance and storage_costs but sacrifices data granularity and requires significant manual SRE effort to maintain. It’s a practical choice for large-scale, cost-sensitive environments.
  • The Adaptive, Edge-First Approach (Netdata): This emerging model offers a compelling alternative, providing both high-fidelity raw data and long-term trends without the manual overhead of roll-ups, fundamentally changing the observability_cost equation.

Understanding these trade-offs is the first step toward building an observability stack that is not only powerful but also sustainable. As your systems grow, the way you handle cardinality will be the single most important factor determining the success of your monitoring strategy.

Experience a new approach to metrics without the cardinality headache. Try Netdata for free today.