The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-service-zero-healthy-instances ▌

Operations Guides

Consul service has zero healthy instances: discovery returns nothing

A consumer calls /v1/health/service/<name>?passing=true and gets []. DNS lookups for <name>.service.consul return nothing. Every downstream consumer treats the service as gone. This is total unavailability for that one service, even if the rest of the Consul cluster is healthy.

Two consumer surfaces are affected at once. API consumers see the empty array directly. DNS consumers (port 8600) get a negative response, and clients may cache it. Recovery from zero means both fixing the checks and flushing the negative caches downstream.

There are two distinct flavors of this incident and the fix paths diverge sharply. Either the instances are still registered in the catalog but every check is critical, or the instances have been deregistered entirely (often by DeregisterCriticalServiceAfter) and no longer appear in /v1/catalog/service/<name>. Settle which one you are in first: re-registering instances is a different operation from clearing a check.

There is also a runway you probably missed. Zero is the bottom of a slope. A service dropping from N instances to one or two is the “traffic is about to overwhelm the survivors” signal. Alerting only on zero catches the cliff, not the approach.

What this means

?passing=true is a strict filter. An instance appears in the result only when every check attached to it is passing. That includes the auto-generated serfHealth check on the node. If serfHealth is critical because the agent went down, every service on that node disappears from ?passing=true results even if the service-level checks were fine.

Discovery has two surfaces that share the same catalog but read it differently. The HTTP API gives you control over consistency mode (stale, default, consistent). DNS uses stale consistency by default, so a freshly corrected catalog may still serve stale DNS responses, and downstream resolvers (dnsmasq, systemd-resolved, the OS) may cache them further. A “fixed in Consul but still failing in the app” gap is almost always DNS caching.

Newly registered checks start in the critical state by default. A freshly registered instance will not appear in ?passing=true results until the first successful check execution lands. On rolling restarts or mass re-registrations, this creates a brief zero-healthy window that is normal, not a bug.

flowchart TD
  A["?passing=true returns empty"] --> B["Query /v1/catalog/service/"]
  B --> C{"Instances in catalog?"}
  C -->|Yes| D["Checks critical, instances still registered"]
  C -->|No| E["Instances deregistered"]
  D --> F["Read /v1/health/state/critical for Output"]
  D --> G["Check serfHealth per node"]
  E --> H["Check DeregisterCriticalServiceAfter timer"]
  E --> I["Verify agents hosting instances are alive"]
  F --> J["Treat as service or check failure"]
  G --> K["Treat as node or agent failure"]
  H --> L["Treat as registration lifecycle issue"]

Common causes

CauseWhat it looks likeFirst thing to check
Bad deploy failing checksAll instances of the new version go critical within one check interval of the rollout; Output fields share an error patternCompare deploy timestamp to first critical transition
Host resource exhaustionInstances on the same host or AZ fail together; host-level serfHealth may also go criticalCross-reference host CPU, memory, and fd usage with affected node list
Over-strict check after config changeCount drops to zero immediately after a check definition change; Output shows timeouts or refused connections that were not failing beforeDiff the recent check configuration
Agents hosting instances went offlineserfHealth on the owning nodes goes critical, then instances disappear after the deregister windowCross-reference consul members with the service’s owning nodes
DeregisterCriticalServiceAfter drove count to zeroCatalog has fewer instances than expected; service may be missing from /v1/catalog/service entirelyCompare /v1/health/service/<name> (any state) to /v1/catalog/service/<name> counts
Stale read showing more healthy than existSome clients see healthy instances, others see none; reads from followers disagreeQuery each server with consistent mode and compare results

Quick checks

Run these read-only commands first. They distinguish “checks critical” from “instances deregistered” and tell you whether you are fighting a real outage or a stale-read illusion.

# Is it really zero? Compare healthy vs total.
SVC=myservice
echo "passing:  $(curl -s "http://127.0.0.1:8500/v1/health/service/$SVC?passing=true" | jq 'length')"
echo "any:      $(curl -s "http://127.0.0.1:8500/v1/health/service/$SVC" | jq 'length')"
echo "catalog:  $(curl -s "http://127.0.0.1:8500/v1/catalog/service/$SVC" | jq 'length')"

# What do the check Output fields say? This is the diagnostic text.
curl -s "http://127.0.0.1:8500/v1/health/service/$SVC" \
  | jq '.[].Checks[] | {CheckID, Status, Output}'

# serfHealth per owning node. If these are critical, the problem is the
# node or agent, not the service check.
curl -s "http://127.0.0.1:8500/v1/health/service/$SVC" \
  | jq -r '.[] | "\(.Node): serfHealth=\(.Checks[] | select(.CheckID=="serfHealth") | .Status)"'

# Is the disagreement a stale-read artifact? Query every server in consistent mode.
for s in server1 server2 server3; do
  printf "%s: " "$s"
  curl -s "http://$s:8500/v1/health/service/$SVC?passing=true&consistent=true" | jq 'length'
done

# DNS path, direct, bypassing downstream resolvers.
dig @127.0.0.1 -p 8600 "$SVC.service.consul" SRV

# Cluster still has a leader? Writes (re-registration) need this.
curl -s http://127.0.0.1:8500/v1/status/leader

# Any registration/deregister churn visible in metrics?
curl -s "http://127.0.0.1:8500/v1/agent/metrics?format=prometheus" \
  | grep -Ei 'catalog.*register'

If passing is zero but catalog is non-zero, instances exist but all checks are critical. If catalog is also zero or smaller than the expected fleet, instances have been deregistered and the issue is registration lifecycle, not check health.

How to diagnose it

  1. Settle the “critical vs deregistered” question first. Compare /v1/health/service/$SVC, /v1/health/service/$SVC?passing=true, and /v1/catalog/service/$SVC counts as shown above. The shape of the gap determines which branch you are in.
  2. Read the Output field on every critical check. The status tells you nothing actionable; the output tells you whether you have a refused connection, a timeout, an expired certificate, or a check script that exited non-zero. This is the single highest-signal field in the incident.
  3. Cross-reference with the owning nodes’ serfHealth. If serfHealth is critical on the nodes hosting the service, the root cause is node or agent health, not the service. Fixing the check is wasted effort until the agent is back.
  4. Look for a deploy-shaped cliff. If all instances went critical within the same check interval of a deploy, the new version is failing the check. The check is doing its job.
  5. Look for an AZ-shaped cliff. If the affected nodes cluster in one AZ or subnet, suspect an infrastructure event. Cross-reference host-level signals.
  6. Rule out stale reads. Query each server directly in consistent mode. If they disagree, you have a catalog consistency problem, not a service health problem. Stale reads can show more healthy instances than actually exist, which is the dangerous direction: clients route to dead endpoints.
  7. Check the DeregisterCriticalServiceAfter value on the service registration. With a value below the one-minute minimum, which Consul clamps, a check that stays critical for a few minutes can silently deregister the instance. There is no automatic re-registration; you have to put it back.
  8. Verify the DNS path separately. Even after the catalog is correct, DNS clients may be holding negative responses. Confirm with a direct dig against the agent before chasing application reports.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
/v1/health/service/<name>?passing=true count (API-derived)Directly tracks the runway to zeroDrop below service-specific minimum, or any single-step drop greater than one instance
Critical checks for the service (API-derived from /v1/health/service/<name>)The leading indicator before DeregisterCriticalServiceAfter triggersSustained non-zero critical count on a service with short deregister window
consul.catalog.register / consul.catalog.deregister latency-timer sample ratesReveals registration lifecycle churn independent of check statusDeregister rate spiking without a corresponding deploy or scale-down
consul.client.rpc.failed on client agentsTells you if agents hosting instances can push state to serversSustained non-zero on agents that own the affected service
consul.serf.member.failed (counter)Signals gossip failure transitions that set serfHealth and gate every service on the nodeAny node hosting a critical service also showing suspect or failed
consul.dns.stale_queriesTells you DNS answers are stale, the inverse signal of catalog freshnessSustained non-zero during a “fixed but still failing” report
consul.raft.commitTimeGates whether re-registration writes can land at allSustained above 100ms means re-registration will be slow

Fixes

Pick the branch by diagnosis, not by habit. Restarting the agent is rarely the right first move and can make things worse by re-triggering the critical-on-startup window.

If checks are critical but instances are still registered

The instances are visible to Consul. The fix is on the service or its dependencies.

  • Read Output on every critical check. Refused connection usually means the service process is down. Timeout usually means the service is saturated or stuck. TLS errors usually mean certificate expiry or CA mismatch.
  • If a deploy shipped a version that fails its check, roll back. Do not loosen the check to make the deploy green; that hides the regression and trains the team to ignore the signal.
  • If host resources are exhausted (CPU pinned, fd limit hit, memory pressure), the fix is on the host, not on Consul. The check is correctly reporting that the service cannot do its job.
  • If the check criteria themselves are wrong after a config change, fix the check definition. Bring the new criteria to the service owner first; silent check loosening is how services stop being monitored in practice.

If instances have been deregistered

The catalog no longer has them. Recovery requires re-registration, and whatever caused the deregistration must be fixed first or the new registration will deregister again.

  • Confirm DeregisterCriticalServiceAfter is the mechanism. If your fleet relies on the agent’s local service definitions, restarting the agent (or sending it a SIGHUP that re-reads its config) will re-register from local state. If registrations were done via the HTTP API, you have to re-register explicitly.
  • Verify the underlying check now passes before re-registering, otherwise you are starting the deregister countdown again.
  • If you want a wider safety margin during incidents, raise DeregisterCriticalServiceAfter. Short values make outages quieter and recoveries more manual. Long values keep instances around but let dead entries accumulate. Pick per service, not globally.

If the disagreement is a stale-read problem

This is the most dangerous variant because some clients see healthy instances that are actually dead.

  • Query each server directly with consistent=true. If they disagree, you have a catalog consistency problem on the server side, not a discovery problem on the client side.
  • After the catalog is consistent, flush downstream DNS caches. Negative responses can persist for the resolver’s negative TTL, which on some operating systems defaults to several minutes.
  • For application clients using the HTTP API, ensure they are using ?passing=true and an appropriate consistency mode. Without ?passing=true, the API returns all instances including critical ones, which can route traffic to dead endpoints.

Prevention

  • Alert on the runway, not the cliff. Per-service thresholds on passing instance count, with the threshold above the minimum needed for survival, gives you lead time. Alerting only on zero gives you none.
  • Treat DeregisterCriticalServiceAfter as a deliberate trade-off. Short values keep the catalog clean during real crashes but turn every check storm into a mass deregistration. Long values keep instances around but let dead entries accumulate. Pick per service, not globally.
  • Validate checks against real failures periodically. A check that never goes critical during a real outage is worse than no check, because it suppresses the signal. Inject a failure and measure detection latency.
  • Track the gap between /v1/health/service/$SVC?passing=true and /v1/catalog/service/$SVC over time. A growing gap means checks are staying critical long enough to be the steady state, which is usually a config or capacity issue.
  • For DNS consumers, document the negative-caching behavior. Knowing the cache TTL up front saves hours of “we fixed it but the app is still failing” debugging.

How Netdata helps

  • Per-second per-service health instance counts let you see the runway to zero before you hit it, and see recovery the moment it starts, instead of waiting for a poll interval.
  • Correlating passing/critical instance counts with consul.client.rpc.failed on the owning agents tells you immediately whether zero is a service problem or an agent-to-server pipeline problem.
  • Correlating instance count drops with consul.catalog.deregister rates distinguishes “checks critical” from “instances gone” without manual API calls.
  • ML anomaly detection on per-service check transition rates surfaces the flapping-before-the-storm pattern, where checks oscillate and then collapse, earlier than static thresholds.
  • DNS stale query metrics alongside catalog metrics show the gap between “fixed in Consul” and “fixed for clients.”
  • Host-level CPU, memory, file descriptor, and disk signals on the same timeline as Consul metrics let you confirm or rule out host resource exhaustion without switching tools.