The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-ping-health-check-trap ▌

Operations Guides

Traefik /ping returns 200 while everything is broken: the health-check trap

Traefik’s /ping endpoint answers one narrow question: “Is the Traefik process alive enough to answer this request?” It does not answer “Can Traefik correctly route production traffic?”

That distinction matters because Traefik is both a data-plane proxy and a control-plane configuration reconciler. The process can stay alive while its Docker socket or Kubernetes API watch has failed, its routing table is stale, every backend is failing health checks, or a certificate is approaching expiry. In all of those states, /ping can still return 200 OK.

Treat /ping as a liveness probe, not a health verdict. A useful Traefik health model needs separate signals for process availability, configuration freshness, routing behavior, backend reachability, TLS state, and resource saturation.

What /ping actually answers

The endpoint must be explicitly enabled, commonly through ping: {} in static configuration or the equivalent --ping option. It is usually exposed on the dashboard or management entrypoint, commonly port 8080, rather than on the public 80 and 443 entrypoints.

# Check the ping endpoint on the management entrypoint
curl -sS -o /dev/null -w '%{http_code}\n' http://localhost:8080/ping

A 200 response tells you that:

  • The Traefik process exists.
  • The management entrypoint accepted the connection.
  • Traefik generated an HTTP response.

It does not tell you that Traefik can accept traffic on public entrypoints, select a valid router, reach a provider, forward to a healthy backend, or serve an unexpired certificate.

During graceful shutdown, ping.terminatingStatusCode can make /ping return a non-200 status, which lets an orchestrator stop sending new work to the process. That still does not make it a proxy readiness check.

Why liveness is not health

Traefik’s request path has several independently failing stages. /ping bypasses almost all of them.

flowchart LR
  P[/ping request/] --> MP[Management entrypoint]
  MP --> ALIVE[Process response]
  T[Client request] --> EP[Public entrypoint]
  EP --> R[Router match]
  R --> MW[Middleware chain]
  MW --> LB[Service load balancer]
  LB --> B[Backend server]
  CFG[Configuration provider] -. supplies routes .-> R
  ACME[Certificate resolver] -. supplies TLS state .-> EP

The /ping path proves only the top branch. A client request depends on the lower branch and on two background systems: the provider watchers that keep routes current and the certificate resolver that keeps TLS credentials valid.

This is why the endpoint is dangerous as the only health input. A load balancer or Kubernetes probe can keep sending traffic to an instance that is alive but operationally wrong.

What a 200 can hide

Hidden conditionWhat users seeBetter signal
Provider desyncNew routes return 404; removed backends keep receiving traffictraefik_config_last_reload_success, entrypoint 404 rate
Stale routing tableRequests arrive, but no current router matchestraefik_entrypoint_requests_total{code="404"}
Backend pool collapseTraefik returns 503 for the servicetraefik_service_server_up and service 503 rate
Backend application failure502, 503, or 504 depending on the failuretraefik_service_requests_total{code=~"5.."}
Retry amplificationBackend load and latency rise while some clients still see successtraefik_service_retries_total
Certificate renewal failureClients begin rejecting HTTPS as expiry approachestraefik_tls_certs_not_after plus Traefik ACME logs
File descriptor exhaustionNew connections fail abruptlyprocess_open_fds / process_max_fds
Process memory or goroutine growthLatency degrades before an OOM restartgo_goroutines, process_resident_memory_bytes

The first row is especially insidious. When a provider loses connectivity, Traefik retains its last-known configuration instead of flushing routes. Existing routes keep working, so ordinary traffic can look normal while every deployment after the disconnection is invisible.

Provider connectivity and configuration freshness

A live process is not necessarily reconciling current configuration. Traefik may be unable to reach the Kubernetes API, Docker socket, Consul, etcd, or another provider while continuing to serve the last successful routing table.

# Inspect configuration reload metrics
curl -s http://localhost:8080/metrics | grep traefik_config_

Watch traefik_config_last_reload_success as a timestamp and traefik_config_reloads_total as activity. If the timestamp stops advancing while the platform is actively deploying or scaling services, suspect provider desync.

Do not alert on age alone for a static file-based deployment. A configuration unchanged for two weeks may be completely healthy. Correlate the timestamp with deployment activity, Kubernetes events, or expected provider churn.

Distinguish between:

  • No reload attempts: the provider watcher may be disconnected or stopped.
  • Reload attempts without a newer successful timestamp: updates may be failing or not applied.
  • Frequent reloads: a busy environment may be causing a rebuild storm and unnecessary CPU pressure.

Traefik v3 does not provide a reliable traefik_config_reloads_failure_total series for this purpose. Infer failure from reload activity, the success timestamp, logs, and provider behavior.

Route correctness is separate from process health

Traefik generates an entrypoint-level 404 when a request matches no router. That is different from a backend returning a 404 after a service was selected.

A rising rate in traefik_entrypoint_requests_total{code="404"} can mean:

  • A provider update was missed.
  • An annotation or label was silently ignored.
  • A router rule is wrong.
  • Clients are requesting stale hostnames or paths.
  • External scanners are enumerating paths.

When a route that should exist is missing, compare the intended configuration with what Traefik actually loaded:

# Inspect routers currently loaded by Traefik
curl -s http://localhost:8080/api/http/routers

Keep this API restricted to trusted networks. It exposes routing and backend topology and must not be reachable from the public internet.

Backend health is separate from backend application health

traefik_service_server_up{service, url} reports whether a backend server is passing Traefik’s configured health check. When every URL for a service drops to 0, Traefik has no healthy backend and returns 503.

Two limitations matter.

First, the metric exists only for services with Traefik health checks enabled. If the series is absent, that means “not monitored by Traefik health checks,” not “healthy.”

Second, a health endpoint can pass while real application paths fail. A backend may return 200 from /health while its database dependency is down and /api returns errors. Cross-check server state against service error rates:

  • traefik_service_server_up == 0 for all service URLs points to backend pool collapse.
  • traefik_service_requests_total{code="503"} confirms Traefik has nowhere to send requests.
  • traefik_service_requests_total{code="502"} indicates Traefik could not get a valid backend response.
  • traefik_service_requests_total{code="504"} indicates the backend exceeded the configured response time.
  • Rising traefik_service_retries_total can show intermittent failures and traffic amplification before clients consistently see errors.

Retries deserve attention because they can mask user-facing failure while multiplying backend load. A successful final response after two failed attempts is still three attempts against a struggling backend.

TLS state is invisible to /ping

Certificate renewal runs in the background. Traefik can keep serving HTTPS with the current certificate while renewal repeatedly fails. /ping has no visibility into that condition.

# Inspect certificate expiry metrics
curl -s http://localhost:8080/metrics | grep traefik_tls_certs_not_after

traefik_tls_certs_not_after exposes certificate expiry as a Unix timestamp. For Let’s Encrypt certificates, renewal is normally attempted well before expiry. A certificate inside the renewal window is a signal to verify that automation is still working; a certificate close to expiry means renewal has likely been failing for some time.

No Prometheus metric directly explains an ACME failure. Correlate approaching expiry with Traefik logs for renewal errors, DNS provider problems, challenge reachability, storage issues, or rate limiting. For externally visible hostnames, pair the internal metric with a synthetic TLS check that validates the certificate clients actually receive.

Resource exhaustion can bypass /ping

A reverse proxy can fail at the edge even when its management endpoint still answers.

The most abrupt case is file descriptor exhaustion. Traefik uses descriptors for client connections, backend connections, provider connections, logs, and other process I/O. When process_open_fds approaches process_max_fds, new work can fail with little graceful degradation.

# Inspect process file descriptor metrics
curl -s http://localhost:8080/metrics | grep -E 'process_(open|max)_fds'

Memory and goroutines are slower-moving signals. Rising go_goroutines and process_resident_memory_bytes, disconnected from traffic growth, can indicate hung backend connections or a goroutine leak. The system may stay responsive until the process hits a container limit and is killed.

A /ping request requires very little of the proxy pipeline, so it can remain successful close to the point where accepting or proxying real connections begins to fail.

A practical health model

Do not replace /ping with one larger synthetic request and declare the problem solved. Build layered checks, each with a precise meaning.

  • Process liveness: /ping answers whether the process can respond. Use it for restart or termination decisions, not for full traffic health.
  • Public entrypoint acceptance: Verify that the real 80 and 443 listeners accept connections. A management port can work while a public entrypoint is unavailable.
  • Configuration freshness: Track traefik_config_last_reload_success and compare it with expected deployment activity.
  • Route correctness: Watch entrypoint 404s and inspect /api/http/routers when intended routes disappear.
  • Backend capacity: Track traefik_service_server_up where health checks are configured, and use service 5xx rates where they are not.
  • Client-visible behavior: Monitor 502, 503, and 504 separately because they imply different root causes.
  • TLS validity: Track traefik_tls_certs_not_after and use external TLS probes for production hostnames.
  • Resource headroom: Track file descriptors, memory, goroutines, CPU, and connection growth.
  • Synthetic path checks: For a small number of representative routes, send a request through the normal public path and validate the expected response. This tests more of the chain but does not replace per-service metrics.

The result should let you say precisely which layer failed. “Ping is green” does not do that.

Signals to watch in production

SignalWhy it mattersWarning sign
process_start_time_secondsDetects unexpected process restartsRepeated starts or a scrape target disappearing
traefik_entrypoint_requests_totalShows whether public traffic is arrivingSustained drop from baseline or rising entrypoint 404s
traefik_config_last_reload_successShows configuration freshnessTimestamp frozen during active deployments
traefik_config_reloads_totalShows provider-driven change activityFlat during churn, or excessively frequent reloads
traefik_service_server_upShows backend health-check stateOne or more servers at 0, especially all URLs at 0
traefik_service_requests_total by codeShows backend and proxy-generated response behaviorSustained 5xx, especially a growing 503 share
traefik_service_retries_totalReveals intermittent failure and amplified loadRetry rate rising with latency or 5xx
traefik_service_request_duration_secondsShows service latency from Traefik’s perspectivePer-service percentile deviation from baseline
traefik_tls_certs_not_afterShows remaining certificate lifetimeExpiry approaching without successful renewal
process_open_fds / process_max_fdsMeasures descriptor headroomHigh sustained ratio or continuing growth
go_goroutinesReveals connection or goroutine leaksGrowth unrelated to traffic
process_resident_memory_bytesShows memory pressure before an OOM killPersistent growth toward the container limit

Alert thresholds need deployment context. A static file-provider deployment, a low-traffic staging service, and a high-churn Kubernetes ingress controller do not share a definition of abnormal configuration age, latency, or error rate.

Common misuses

  • Using /ping as the only load-balancer health check. This keeps a stale or backend-isolated instance in rotation.
  • Checking /ping on the wrong port. The endpoint is usually on the management entrypoint, not the public traffic entrypoints.
  • Assuming absent health metrics mean healthy backends. traefik_service_server_up is absent when service health checks are not configured.
  • Aggregating all 5xx responses together. 502, 503, and 504 point to different failure mechanisms.
  • Alerting only on certificate expiry. By the time expiry is close, renewal may already have failed repeatedly.
  • Treating config age as universally bad. Static configurations may legitimately go unchanged for long periods; correlate age with expected changes.
  • Trusting backend health checks without error-rate correlation. A lightweight health path can pass while real requests fail.

How Netdata helps

  • Netdata can place /ping availability beside public request rates so a green management check is not mistaken for successful traffic delivery.
  • Correlating traefik_config_last_reload_success with entrypoint 404s helps identify provider desync without waiting for a developer to report a missing route.
  • Per-service 502, 503, and 504 trends preserve the distinction between invalid backend responses, exhausted backend pools, and timeouts.
  • Combining traefik_service_server_up, retries, and service latency exposes cascading backend failure before every backend reaches zero.
  • Certificate expiry, process restarts, file descriptor usage, memory, and goroutine trends provide the process and resource context that /ping omits.
  • Anomaly views help compare current behavior with each service’s baseline instead of applying one generic threshold to heterogeneous routes.