The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-monitoring-maturity-model ▌

Operations Guides

Consul monitoring maturity model: from survival to expert

Consul runs three independent subsystems that must all be healthy for the cluster to function: Raft consensus, Serf gossip, and the catalog state machine with its anti-entropy sync pipeline. Most teams monitor one or two well and discover the third during an incident.

The model is cumulative. Each level adds signals to the previous one; you cannot skip to Mature without the Operational baseline in place. A team with composite-pattern detection at Expert but no client-agent RPC monitoring at Operational is blind to the most common silent failure mode: the catalog drifting from reality while every server-side metric looks healthy.

Use this as both an audit of current coverage and a prioritized roadmap. Signals within each level are ordered roughly by incident frequency. Treat the boundaries as maturity gates: a team can be Mature on servers and Survival on clients, and the weakest tier wins.

flowchart TD
  L1["Survival
leader, gossip, process, disk"] L2["Operational
commitTime, lastContact, RPC failures, peer count"] L3["Mature
goroutines, gossip queues, catalog churn, cert expiry, GC pauses"] L4["Expert
log divergence, anti-entropy timing, blocking queries, cross-DC RPC"] L1 --> L2 --> L3 --> L4

Level 1: Survival

If any of these fail, the cluster is either down or about to be. A team with only these signals will detect total outages but miss every silent degradation and most partial failures.

Signals to cover:

  • Leader exists. Query /v1/status/leader and assert a non-empty response. Without a leader, all writes fail: service registrations, KV writes, session creation, and health check state updates.
  • All server nodes alive in gossip. Query /v1/agent/members, filter for Tags.role == consul, and assert Status == 1 for each. A failed server in a 3-server cluster is one failure from quorum loss.
  • Agent process up on every node. A dead agent means no health checks execute locally, and the catalog shows last-known state until gossip marks the node failed.
  • HTTP API responsive. A 200 from /v1/status/leader within a tight timeout distinguishes “Consul is slow” from “Consul is gone.”
  • Server data directory not full. Raft cannot persist log entries when the volume fills; the server will crash or fall behind and lose quorum.

Gotchas:

  • /v1/status/leader can return a stale result from a server that has not yet processed the leader change. Cross-check with multiple servers during incidents.
  • A server can be alive in gossip while being non-functional for Raft. Gossip and Raft are independent protocols; one being healthy does not imply the other.
  • Brief leaderless periods of 1 to 5 seconds during rolling restarts are normal. Use a sustained-absence window of 15 to 30 seconds for paging thresholds, or you will alert on every upgrade.

Level 2: Operational

This is where a competent team running Consul in production should live. The signals here catch performance degradation, partial failures, and the conditions that lead to leader elections.

Signals to add:

  • Raft commit time (consul.raft.commitTime). End-to-end write latency and the single best indicator of Raft health. Captures disk write latency, network replication latency, and FSM apply time in one number. Leader-only.
  • Raft last contact (consul.raft.leader.lastContact). Time since the leader last heard from each follower. Growing values indicate replication lag and approaching election timeouts. Leader-only.
  • Leader election count (consul.raft.state.leader). This is a counter incremented each time this server becomes leader; alert on increments outside maintenance. More than 2 transitions per 10 minutes outside a maintenance window indicates a systemic problem.
  • DNS query latency. User-facing latency of service discovery for any consumer using Consul DNS on port 8600.
  • Health check distribution. Passing, warning, and critical counts across the catalog. A sudden spike in critical checks usually indicates a shared downstream dependency failure.
  • Server memory RSS. Tracks resource consumption and growth trends. Correlate with catalog size and goroutine count.
  • Client agent RPC failure rate (consul.client.rpc.failed). Elevated rates mean agents cannot push state to the catalog. This is the single most under-monitored signal in Consul deployments.
  • File descriptor usage. Approaching the limit causes connection failures, gossip instability, and Raft timeouts. Default OS ulimit of 1024 is too low; servers should run with at least 65536.
  • Raft peer count (consul.raft.peers or /v1/operator/raft/configuration). Must match expected cluster size. A drop below quorum is a page.

Gotchas:

  • Many Raft metrics are leader-only: commitTime, kvs.apply, catalog.register. When leadership changes, your monitoring must follow the leader or you will have gaps in the time series.
  • Blocking queries inflate HTTP latency metrics. Filter requests with ?wait= or ?index= parameters when analyzing API performance, or your p99 numbers will be meaningless.
  • consul.client.rpc.failed is a counter. Alert on the rate of change, not the absolute value.
  • Autopilot’s dead server cleanup can remove a failed server automatically, causing the peer count to drop even when a replacement is being provisioned. This is by design but confusing if unexpected.

Level 3: Mature

Full coverage for a professional SRE team. These signals catch slow leaks, capacity trends, and the secondary effects that compound primary failures into outages.

Signals to add:

  • Goroutine count (consul.runtime.num_goroutines). Monotonic growth indicates a leak. Each blocking query, each gRPC stream, and each watch holds a goroutine for its lifetime.
  • Gossip queue depth (consul.serf.queue_Event, consul.serf.queue_Intent, consul.serf.queue_Query). Sustained non-zero values mean the agent is receiving gossip faster than it can process it, causing delayed failure detection and false positives.
  • Catalog registration and deregistration rate (consul.catalog.register, consul.catalog.deregister). High churn drives Raft load, anti-entropy overhead, and downstream consumer updates.
  • KV operation latency (consul.kvs.apply). Write path for KV specifically. If KV latency is high but general Raft metrics are normal, the issue is KV-specific: large values, deep key trees, or transaction contention.
  • HTTP API latency per endpoint. Catches single-subsystem bottlenecks masked by aggregate averages.
  • xDS stream count (consul.xds.server.streams). Connect control-plane health; should match your Envoy proxy count. Each stream holds a goroutine and a file descriptor on the server.
  • Certificate expiration time for both leaf and CA root. Cliff-edge failure: everything works until the exact moment it does not.
  • ACL resolution latency (consul.acl.ResolveToken). Every authenticated API request pays this cost. High latency with a high cache-miss ratio means the token cache is undersized.
  • consul members -wan membership state per DC. A remote DC disappearing from the WAN pool breaks prepared-query failover and cross-DC service lookups.
  • Raft snapshot size. Tracks state growth. Large snapshots slow recovery, compete for disk I/O during creation, and can temporarily double memory usage.
  • Disk I/O latency on server volumes. Track await and %util, not just throughput. This is the leading indicator for Raft issues and the most common source of leader-instability incidents.
  • Go GC pause duration (consul.runtime.gc_pause_ns, a nanosecond summary). Stop-the-world pauses affect Raft timing. Pauses approaching the election timeout will trigger elections.
  • Session operation timing and invalidation (consul.session.apply and consul.session_ttl.invalidate are timers; track application-side lock loss for rate). Spikes mean distributed locks releasing across the cluster, which can trigger application-level leader elections and cache flushes.

Gotchas:

  • Disk I/O latency is the most common root cause of leader instability. Teams provision acceptable CPU and memory, then put the Raft data directory on EBS gp2 with exhaustible burst credits, shared NFS, or spinning disks. The first visible symptom is leader elections, by which point the cluster is already degraded. Page on disk write latency, not on downstream Raft symptoms.
  • Go’s garbage collector does not immediately return memory to the OS. runtime.alloc_bytes (live heap) and VmRSS (OS-level resident set) can diverge significantly. Investigate monotonic RSS growth, not absolute ratios.
  • Certificate expiration monitoring without renewal success tracking is a countdown clock with no early warning. Track both. A Vault CA backend outage will not surface as expiry-time degradation until hours later, when certificates start failing to rotate.
  • Gossip and Raft are independent. A server can be alive in gossip while having corrupt Raft state, applying entries incorrectly and diverging from the leader. This only manifests when that server wins an election.
  • Composite patterns matter more than individual thresholds at this level. Leader churn, health-check thundering herds, blocking-query amplification, and gossip/Raft divergence each have distinctive multi-signal signatures. Build alerts that fire on the combination, not on any single metric.

Level 4: Expert

Signals that experienced operators add after their third or fourth major incident. They catch subtle internal state and provide the 30-minute warning that would have averted the last outage.

Signals to add:

  • Raft log index divergence across servers (consul.raft.last_index and consul.raft.applied_index gauges on all servers). Indices should be nearly identical across the peer set. Large divergence means a follower is falling behind or has inconsistent state, and is a candidate for serving corrupt data if it wins an election.
  • Anti-entropy sync timing per agent. Detects agents that cannot sync even though they are alive in gossip. The catalog drifts silently while every server metric looks healthy.
  • Blocking-query count and distribution. Identifies watch accumulation before it becomes a goroutine leak. Each consul-template instance and each watch holds a goroutine and a connection; leaks accumulate for weeks before cascading.
  • Prepared-query failover event rate. Detects when consumers are silently being served from remote DCs. Receiving results does not mean local health; it may mean prepared-query failover is masking a complete DC service failure.
  • Cross-DC RPC latency. WAN gossip being healthy does not guarantee that cross-DC RPC is fast enough for application timeouts.
  • Certificate renewal success rate. Not just expiry time. Track CSR acceptance, signing latency, and Envoy acknowledgment separately.
  • Catalog size by service. Which services contribute most to catalog bloat, driving snapshot size, memory, and query latency.
  • Raft compaction timing and success. If snapshot creation fails, the log grows indefinitely.
  • Agent event handler execution time. Custom handlers blocking the agent main loop.
  • Network partition detection via cross-referenced member lists. Query consul members from multiple servers and diff. Asymmetric partitions, where server A sees B but B does not see A, are common and invisible to leader-centric metrics.

Gotchas:

  • Pairwise Raft network latency matters, not just leader-to-follower. Measure between all server pairs. Asymmetric partitions are the failure mode most likely to cause confusing split-brain symptoms in a 5-server cluster.
  • These signals rarely fire in steady state. Their value is during incidents and capacity planning. Do not expect them to populate dashboards with activity; expect them to provide critical context when everything else is on fire.
  • Consul’s telemetry uses time-window aggregation, and rapid transients under a second may not appear in metrics sampled per 10 seconds. Per-second collection matters here more than at any other level.

Where most teams get stuck

  • Disk I/O ignored until it causes elections. The leading indicator is write latency (await), not throughput or utilization. This should be page-level on server volumes. If you only learn about disk pressure from consul.raft.commitTime spikes, you are reacting too late.
  • Client-to-server RPC health unmonitored. Server-side monitoring misses the pipeline from client agents to the catalog. The catalog silently goes stale. Monitor consul.client.rpc.failed on every client agent and alert on sustained non-zero rates.
  • Blocking query accumulation treated as normal load. Each watch and each consul-template instance holds a goroutine and a connection. Leaks accumulate for weeks before cascading into FD exhaustion or goroutine saturation. Track goroutine count trends and correlate with expected blocking-query load.
  • Composite pattern detection absent. Individual thresholds miss multi-signal failure modes. Teams with 100 metrics but no composite alerts miss these until they escalate.
  • Version-specific behavior untracked. Consul 1.19 removed state-store metrics with the doubled consul.consul prefix; use the single-prefix names such as consul.state.config_entries. The telemetry.disable_compat_1.9 option was removed in Consul 1.13 after the 1.9-style consul.http metrics had been deprecated. Adjust monitoring assumptions on every upgrade, and treat the version-specific upgrade notes as required reading for whoever owns the dashboards.

How Netdata helps

Netdata’s Consul collector scrapes the agent telemetry endpoint at per-second resolution, which matters because the most expensive Consul failures (Raft timing violations, brief gossip partitions) produce transients that 10- or 15-second pollers miss.

  • Per-second Raft commit time correlated with disk I/O latency shows the causal chain from disk await degradation to commit-time spikes to leader elections in a single view, instead of triangulating across three dashboards after the fact.
  • Goroutine count and file descriptor usage trended together surfaces slow blocking-query leaks before they become outages. Anomaly detection flags monotonic growth that static thresholds miss.
  • Client-agent RPC failure rate collected on every node, not just servers closes the most common blind spot. Each client agent reports its own consul.client.rpc.failed without depending on a central scraper that only talks to servers.
  • Leader-only metrics follow the leader automatically. Each node reports its own telemetry, so consul.raft.commitTime and consul.kvs.apply appear on whichever server is currently leader, with no gaps during transitions.
  • ML anomaly scoring across Raft, gossip, and RPC signals catches the multi-signal signatures of leader churn, thundering herds, and gossip/Raft divergence that single-metric thresholds cannot.