The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-503-service-unavailable ▌

Operations Guides

Traefik 503 Service Unavailable: no healthy backends left in the pool

Traefik is returning 503 Service Unavailable for every request to a service. Clients are down. Traefik itself looks fine: the process is up, /ping returns 200, other services route normally. This is the backend pool collapse failure mode: Traefik matched a router to a service, but every backend in that service’s load balancer pool has been marked down by health checks, so Traefik has nowhere to send the request.

A Traefik 503 is not a Traefik failure. It is Traefik correctly reporting that its upstream pool is empty. The investigation belongs one hop upstream, at the backends or at the health check configuration itself.

This article covers how to confirm the pool collapse, how to distinguish a genuine backend outage from the three common impostors (post-deploy warm-up, a service intentionally scaled to zero, and a misconfigured health check), and how to stop it from paging you for the wrong reason.

What this means

Traefik’s request pipeline ends at the service load balancer, which maintains a pool of backend servers. Health-checker goroutines probe each backend on a configurable interval (the Traefik default is 30 seconds, with a 5 second timeout; healthy means a 2XX or 3XX response, or whatever status you configured). When a backend fails its checks, Traefik removes it from rotation. When every backend for a service is out of rotation, Traefik short-circuits and returns 503 without proxying anything.

Three properties matter for diagnosis:

  1. Traefik stays healthy throughout. /ping returns 200 because it only checks process liveness. It says nothing about backend health. See the health-check trap guide for why this bites operators.
  2. The 503 appears at the service level, not the entrypoint level. You will see it in traefik_service_requests_total{code="503"} for the affected service, while entrypoint metrics keep counting the requests as arriving normally.
  3. Do not confuse it with its neighbors. Per the Traefik FAQ: 404 means no router matched at all; 502 means Traefik contacted the backend and got an invalid response; 504 means the backend accepted the request but did not respond in time. 503 specifically means the pool is empty. Each of these has a different root cause and a different fix.

The defining metric is traefik_service_server_up{service, url}, a 1/0 gauge per backend URL. When it reads 0 for every URL in a service, the service is down. One critical caveat: this series only exists for services that have Traefik health checks enabled. If health checks are not configured, the series is absent, and absence means unmonitored, not healthy.

Common causes

CauseWhat it looks likeFirst thing to check
Genuine backend outageAll traefik_service_server_up went to 0 together, backends are actually down or crashedBackend process/pod status and application logs
Post-deploy warm-up window503s for 10-30s right after a rollout; backends pass readiness but app is still initializingCorrelate 503 window with deploy timestamp
Service scaled to zeroService intentionally has no replicas; Traefik has no servers (or all marked down)Desired replica count, autoscaler state
Misconfigured health-check pathBackends are up and serving real traffic, but health check returns 404/non-2XXcurl the health-check path on a backend directly
Dependency failure (DB, cache)All backends fail health checks simultaneously because the check depends on a shared resourceBackend health endpoint response body
Health check differs from traffic pathBackends “healthy” but 5xx, or the reverse: check path fails while app worksCompare traefik_service_server_up with real 5xx rate
Cascading pool collapseBackends marked down gradually, retries spike first, then all downtraefik_service_retries_total trend before the 503s
Autoscaler scale-downPool shrinks below what traffic requires, survivors overload and fail checksScaling events in the window before the incident

Quick checks

All read-only. Run these before changing anything. The examples assume Traefik’s metrics and API are exposed on :8080; adjust to your metrics/API entrypoint.

# 1. Confirm the pool state: every backend URL at 0 for the service
curl -s http://localhost:8080/metrics | grep traefik_service_server_up

# 2. Confirm the 503s are service-level and which service
curl -s http://localhost:8080/metrics | grep 'traefik_service_requests_total' | grep 'code="503"'

# 3. Check retries in the lead-up (cascade pattern shows retries spiking first)
curl -s http://localhost:8080/metrics | grep traefik_service_retries_total

# 4. See what Traefik's API thinks the service looks like (backend URLs, status)
curl -s http://localhost:8080/api/http/services | head -100

# 5. Sanity check that Traefik itself is alive (expected: 200, tells you nothing about backends)
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8080/ping

Then probe a backend directly, the same way the health checker would:

# 6. Hit the configured health-check path on a backend, bypassing Traefik
curl -s -o /dev/null -w '%{http_code}\n' http://<backend-ip>:<port>/<healthcheck-path>

# 7. Hit the real application path on the same backend for comparison
curl -s -o /dev/null -w '%{http_code}\n' http://<backend-ip>:<port>/<real-path>

Checks 6 and 7 together separate “backend is dead” from “health check is lying.” If 6 returns 404 or 500 while 7 returns 200, the backends are fine and the health check is misconfigured.

How to diagnose it

Work through these in order. Each step eliminates a class of causes.

  1. Confirm the pool collapse. From check 1: is traefik_service_server_up 0 for every URL in the service? If the series is entirely absent, the service has no health checks configured and the 503 is coming from somewhere else (for example, a service explicitly configured with no servers). Re-read the metric absence caveat before proceeding.

  2. Establish the timeline. Did all backends go to 0 at the same instant, or one by one? Simultaneous failure points to a shared dependency (database, cache, DNS) or a config change. Gradual decline, especially with traefik_service_retries_total rising first, is the cascading failure pattern: fewer healthy backends take more load, fail, and concentrate load further.

  3. Correlate with deployment activity. Was there a rollout, scale event, or config change in the minutes before? A 503 window of 10 to 30 seconds immediately after a deploy is the warm-up edge case: the backend passes readiness, Traefik adds it to the pool, but the application is still initializing connection pools and caches, and the health check flaps. This is expected behavior, not an incident.

  4. Check for intentional scale-to-zero. If the service is supposed to be dormant (batch service, off-hours tool, preview environment), zero healthy backends may be the correct state. The alert condition needs a traffic floor, not just a health state.

  5. Probe the backend directly (checks 6 and 7). Dead backend: fix the backend. Alive backend with a failing health-check path: fix the health check. This fork determines everything downstream.

  6. If backends are alive but marked down, inspect the health check config. Common misconfigurations: wrong path (app has no /health, check returns 404), wrong port, expected status code not matching what the app returns, or a check path that depends on a downstream dependency that flaps independently of the app.

  7. If backends are genuinely down, move upstream. Application logs, resource exhaustion, database connectivity. Traefik’s job is done at this point; it told you the truth.

flowchart TD
  A[503 on service] --> B{traefik_service_server_up
all URLs = 0?} B -->|series absent| C[No health checks configured
investigate service config] B -->|all zero| D{Backends alive
on direct probe?} D -->|no| E{Simultaneous
or gradual?} E -->|simultaneous| F[Shared dependency failure
DB / cache / DNS] E -->|gradual + retries rising| G[Cascading collapse
fix initial backend failure] D -->|yes| H{Health-check path
returns 2XX?} H -->|no| I[Misconfigured health check
path / port / expected status] H -->|yes| J[Check recent deploy or scale event
warm-up window or scale-to-zero]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_service_server_up{service, url}The defining signal: per-backend health state0 for every URL in a service; also any single URL at 0 beyond two check intervals
traefik_service_requests_total{code="503"}Confirms client-visible impact per serviceAny sustained 503s on a service that normally has backends
traefik_service_retries_totalEarly warning of cascade; retries spike before the pool emptiesRetry rate above 5% of request rate sustained
traefik_service_request_duration_secondsRising latency on remaining backends during partial pool lossp95 climbing while backends drop out one by one
Backend application health (external to Traefik)Traefik only sees what the health check seesDirect probe failures, crash loops, resource saturation

On the first row: one backend down reduces capacity and resilience. All backends down is the outage. Alert differently on the two conditions, and add a traffic floor to the all-down alert so a legitimately idle or scaled-to-zero service does not page anyone.

The v3 issue that renders the service label as an opaque hash instead of service@provider remains open as traefik/traefik#10734. If you see hash labels, map them via the /api/http/services output rather than assuming the issue is fixed in your minor.

Fixes

Genuine backend outage

Fix the backend, not Traefik. Roll back the bad deploy, restore the shared dependency, or replace the crashed instances. Once backends pass health checks again, Traefik returns them to rotation immediately. There is no warm-up or ramp-up mechanism: a recovered backend gets its full traffic share instantly, which can re-kill a fragile recovery. If that happens, investigate the backend’s ability to handle cold load rather than blaming the proxy.

Post-deploy warm-up window

Do not “fix” this with alert tuning alone; fix the deploy. Give the backend health check an interval and threshold that matches the application’s real startup time, and use startup probes in Kubernetes so Traefik’s provider does not see the pod as ready before the app actually is. Brief 5xx spikes right after a rollout are expected; do not alert on them.

Scaled to zero

If the service is legitimately dormant, the 503 is correct behavior. Suppress the all-backends-down alert for that service, or gate it on observed request traffic so it only fires when someone is actually being served errors.

Misconfigured health check

Correct the path, port, scheme, or expected status so the check reflects the application. Longer term, make the health check meaningful: Traefik’s fallback behavior without an explicit check only detects connection-level failure, and a check that hits a trivial /healthz while real traffic needs a database will lie to you in the other direction (all backends “up”, real requests 5xx). Cross-referencing traefik_service_server_up against the actual 5xx rate catches both directions of lie.

Cascading collapse

Break the feedback loop first: reduce or disable the retry middleware on the affected service so surviving backends stop absorbing amplified traffic. Then address whatever degraded the first backend. Retries masking errors on the dashboard while tripling backend load is how a partial failure becomes a total one.

Prevention

  • Configure explicit health checks on every production service. Without them, traefik_service_server_up does not exist for that service and you are blind to pool state until clients report errors.
  • Make health checks representative. The check path should exercise what real traffic needs, or at least fail when the app cannot serve. Otherwise you trade 503s for silent 5xx.
  • Align health-check timing with deploy behavior. Check interval and thresholds should tolerate real startup time so rollouts do not flap the pool.
  • Alert on partial pool loss, not just total loss. One backend down in a two-backend pool is 50% capacity gone and zero redundancy. That is a ticket, not a page, but it must not be invisible.
  • Watch retries as a leading indicator. Retry rate relative to request rate is the earliest signal of a cascade in progress.
  • Put a traffic floor on the all-down alert. Scale-to-zero and dormant services are healthy at zero backends. Page only when zero backends coincides with real request traffic.

How Netdata helps

  • Netdata charts traefik_service_server_up per backend URL, so you see the pool draining in real time and can tell gradual cascade from simultaneous failure at a glance.
  • Per-service response code breakdown puts the 503 rate next to retries and latency on one dashboard, which is exactly the correlation that separates a cascade from a config mistake.
  • Comparing service-level 5xx against entry-level metrics shows the errors are service-scoped, confirming Traefik itself is healthy and the problem is upstream.
  • Netdata’s process and Go runtime collectors cover the Traefik side of the house (FDs, goroutines, memory), so you can rule out proxy-level resource issues quickly and focus on the backends.
  • Anomaly detection on per-service request rates helps distinguish a real outage from a legitimately idle service, which is the difference between a page and a non-event.