The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-circuit-breaker ▌

Operations Guides

Traefik circuit breaker: shedding load from a failing backend

A backend is degrading. It is not fully down, so health checks still pass or flap, and Traefik keeps sending it full traffic. If you also have the retry middleware attached, every failed request comes back two or three more times, and the retry amplification loop finishes what the original failure started. The circuit breaker middleware exists for exactly this situation: it watches error and latency ratios on a router, and when they cross a threshold you define, it stops forwarding requests and answers with a fast 503 until the backend recovers.

The operational problem is that the breaker is nearly invisible. Traefik exposes no per-middleware Prometheus metrics, so there is no breaker-state gauge to alert on. An open breaker looks, at a glance, like a backend pool collapse: clients get 503s. And a badly tuned breaker is worse than none at all: it either never trips, or it flaps between open and closed and turns a steady degradation into an intermittent one that is much harder to debug.

What this means

The circuit breaker is a middleware. It sits in the middleware chain of a router and continuously evaluates an expression over the traffic it observes. The expression is built from three measurement functions:

  • NetworkErrorRatio(): ratio of requests that failed at the network level (connection refused, reset, timeout before a response).
  • ResponseCodeRatio(from, to, dividedByFrom, dividedByTo): ratio of responses in one status code range relative to another, for example ResponseCodeRatio(500, 600, 0, 600) for 5xx as a fraction of all responses.
  • LatencyAtQuantileMS(quantile): request latency at a given quantile in milliseconds. The quantile must be a float with a trailing .0, so LatencyAtQuantileMS(50.0) is valid and LatencyAtQuantileMS(50) is not.

You combine these with comparison operators (>, >=, <, <=, ==, !=) and logical operators && and ||. The OR keyword is not supported; writing OR in an expression fails with a parse error at load time. Use ||.

The breaker has three states:

stateDiagram-v2
  [*] --> Closed
  Closed --> Open: expression true at checkPeriod
  Open --> Recovering: fallbackDuration elapsed
  Recovering --> Closed: recoveryDuration elapsed
  Recovering --> Open: expression still true
  Open: clients get fast 503
  Recovering: linear ramp of probes
  • Closed: normal operation. All traffic flows. The expression is evaluated every checkPeriod (default 100ms).
  • Open: the expression evaluated true. The breaker short-circuits the chain and returns 503 (configurable via responseCode) immediately, without touching the backend. This lasts fallbackDuration (default 10s).
  • Recovering: after fallbackDuration, the breaker lets a linearly increasing number of probe requests through over recoveryDuration (default 10s). If the expression is satisfied again during recovery, it re-opens. If recovery completes, it closes.

Two properties of this design matter operationally. First, each router gets its own independent instance of the breaker, even if multiple routers reference the same middleware definition. One router can be open while another, using the “same” breaker, is closed. Second, the breaker only observes what happens after its own position in the middleware chain. Anything a preceding middleware short-circuits (auth rejections, rate limit 429s) is invisible to it.

The defaults are reasonable starting points, but the expression itself has no default. You must write it, and the quality of the whole mechanism depends on that expression matching your backend’s real failure signature.

Configuring it

A minimal dynamic configuration (file provider YAML):

http:
  middlewares:
    api-breaker:
      circuitBreaker:
        expression: "NetworkErrorRatio() > 0.30"
        checkPeriod: 100ms
        fallbackDuration: 10s
        recoveryDuration: 10s

Common expression patterns:

  • Connection-level failure: NetworkErrorRatio() > 0.30. Trips when 30% of requests fail before a response.
  • Application-level failure: ResponseCodeRatio(500, 600, 0, 600) > 0.30. Trips when 30% of all responses are 5xx.
  • Latency degradation: LatencyAtQuantileMS(50.0) > 500. Trips when the median request takes over 500ms.
  • Combined: NetworkErrorRatio() > 0.30 || ResponseCodeRatio(500, 600, 0, 600) > 0.50.

Ordering with retry. Retries add load; the breaker removes it. They complement each other only if ordered correctly. Place the retry middleware first and the circuit breaker after it in the chain. If the retry middleware sits after the breaker, the breaker can trip on the first failure before a retry has a chance to succeed, which defeats the retry entirely. With retry first, the breaker sees the final outcome of each request after retries are exhausted, which is the signal you actually want to trip on.

Known limitations to design around:

  • There is no fallback service. An open breaker returns a static status code (503 by default). It cannot redirect traffic to a standby backend or a maintenance page upstream. Requests for fallback routing have been declined by the maintainers.
  • In current stable releases, there is no minimum request count. A ratio-based expression can trip on the first handful of requests after a quiet period. A RequestThreshold() expression function that addresses this has been merged into the master branch but is not yet in a stable release. Until then, low-traffic routers need conservative ratios or latency-based expressions, which are less jumpy on small samples.
  • The official documentation does not document the sliding-window size. Traefik’s circuit-breaker dependency implements 10 one-second rolling buckets for counters and a 10-second rolling histogram, so treat ratios as roughly the last 10 seconds and validate against your own traffic rate.

Common failure modes of the breaker itself

SymptomWhat it looks likeFirst thing to check
Breaker never tripsBackend degraded for minutes, no 503 spike, retries climbingExpression threshold too strict for real traffic; && chain requiring two conditions that never coincide
Breaker flaps503s come in pulses roughly fallbackDuration + recoveryDuration apartThreshold right at the backend’s steady-state error rate; recovery probes trip it again immediately
Trips on healthy backend503s at low traffic times, backend fineRatio expression on a low-volume router: two failed requests out of four is 50%. No minimum request count in stable
Trips before retries helpBreaker opens on transient single-attempt failuresMiddleware order wrong: retry placed after the breaker
Breaker on one route onlyOne router 503s, sibling routers to the same backend finePer-router instance behavior; this is by design, not a bug
Expression rejected at loadBreaker never active, config error in logsOR instead of `

Quick checks

All read-only. These assume the metrics and API endpoints are reachable on the Traefik instance (typically port 8080 for the dashboard/API entrypoint; adjust for your deployment).

# 1. Is the 503 rate elevated on a specific service?
curl -s http://localhost:8080/metrics | grep 'traefik_service_requests_total' | grep 'code="503"'

# 2. Are the backends actually down, or is something upstream of them refusing traffic?
curl -s http://localhost:8080/metrics | grep traefik_service_server_up

# 3. Are retries running at the same time? (retry amplification signal)
curl -s http://localhost:8080/metrics | grep traefik_service_retries_total

# 4. Is latency rising on the affected service?
curl -s http://localhost:8080/metrics | grep traefik_service_request_duration_seconds

# 5. Is the breaker middleware actually loaded, and with what expression?
curl -s http://localhost:8080/api/rawdata | grep -i -A5 circuitbreaker

Check 2 is the discriminator. The breaker is a middleware; it does not change backend health state. If 503s are spiking while traefik_service_server_up still shows 1 for the service’s backends, the 503s are coming from the middleware layer, not from health-check eviction. That is either your circuit breaker working as intended or, if you have not configured one, another short-circuiting middleware. Note that traefik_service_server_up only exists for services with health checks enabled; absence of the series means unmonitored, not healthy.

How to diagnose a 503 spike

  1. Confirm the 503s are real and scoped. Look at traefik_service_requests_total{code="503"} per service. A single service means a service-level cause; many services at once points at a shared dependency or at Traefik itself.

  2. Check backend health state. If traefik_service_server_up is 0 for every URL in the service, this is backend pool collapse, not the breaker. Follow Traefik 503 Service Unavailable: no healthy backends left in the pool.

  3. If backends are up, suspect the middleware layer. Verify a circuit breaker is attached to the router (check 5 above) and evaluate its expression by hand against recent traffic: what fraction of recent requests were network errors or 5xx?

  4. Check the timing pattern. An open breaker produces a distinctive signature: a block of 503s lasting roughly fallbackDuration, then a recovery window, then either normal traffic or another block. Flapping at this cadence means the threshold is too close to the backend’s steady-state error rate.

  5. Check retries. If traefik_service_retries_total is climbing in step with the 503s, you have both mechanisms active. Verify the middleware order (retry first) and consider whether retries are feeding the failure the breaker is trying to shed.

  6. Check the logs. Breaker state transitions are not exposed as metrics. State changes and recovery decisions are emitted at DEBUG; fallback activation also emits a WARN. Traefik’s default log level is ERROR, so raise the level temporarily to see these transitions.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_service_requests_total{code="503"} per serviceThe only direct observable of breaker output (and of pool collapse)Sustained rate, or periodic bursts at breaker-timing cadence
traefik_service_server_up per serviceSeparates “breaker open” (still 1) from “no healthy backends” (0)503s with all URLs at 1: middleware-level shedding
traefik_service_retries_total per serviceRetry amplification coexists with and can trigger the breakerRetries and 503s rising together
traefik_service_request_duration_secondsA LatencyAtQuantileMS-based breaker trips on this before errors appearp50/p95 approaching your expression threshold
traefik_entrypoint_requests_total{code="503"}Entrypoint-level view; compare with service-level to see where 503s originateEntrypoint 503s exceeding the sum of service 503s

Because there is no breaker-state metric, breaker-aware alerting has to be composite: alert on 503 rate combined with backends still up. That combination almost always means intentional shedding (breaker, rate limiter, or another short-circuit) rather than infrastructure failure, and it should be routed and severity-tagged differently than a genuine pool collapse. Breaker 503s are the mechanism working; they deserve a ticket and a dashboard annotation, not a page.

Tuning guidance

Threshold never trips. The expression does not match the backend’s real failure mode. A backend that hangs responds too slowly rather than erroring, so NetworkErrorRatio() stays at zero and only a latency expression will catch it. Match the expression to the failure signature you actually see in an incident, not the one you expect.

Threshold flaps. The trip point sits inside the backend’s normal error band. Raise the ratio, lengthen fallbackDuration so the backend gets real time to recover, and lengthen recoveryDuration so the probe ramp is gentler. A breaker that re-opens on its first probe burst is telling you the backend needs minutes, not seconds.

Trips on low traffic. With no minimum request count in stable releases, small samples make ratios violent. For quiet routers, prefer a latency-based expression, or a higher ratio combined with && against a second condition so a single bad request cannot open the circuit.

503s surprise clients. The breaker returns a bare 503 by default. You can change responseCode, but you cannot return a body or redirect. If clients need a graceful degradation path, implement it at the client or in a higher layer, not in Traefik.

Prevention

  • Write the expression from incident data. Take the error ratio and latency quantiles from your last real backend incident and set the threshold just outside that range. An expression invented in a vacuum will misfire.
  • Fix the middleware order once, in review. Retry first, breaker after. Make it a lint or code-review rule for your dynamic config.
  • Baseline per service. Universal thresholds across heterogeneous backends mask problems. A compute-heavy API and a static file service need different ratios and quantiles.
  • Load-test the breaker before you need it. Trip it deliberately in staging with fault injection and confirm the 503 cadence, the recovery ramp, and that your dashboards show the composite signature.
  • Revisit expressions after upgrades. The expression language is stable, but new functions land (for example RequestThreshold() on master). As of v3.7.12 it is not in a stable release; recheck the docs for your version before assuming a function exists.

How Netdata helps

  • Netdata collects Traefik’s Prometheus endpoint and charts traefik_service_requests_total by response code per service, so a 503 spike from an open breaker is visible at per-second resolution instead of being averaged away.
  • Plotting traefik_service_server_up alongside the 503 rate makes the key discriminator automatic: 503s with backends still up means middleware-level shedding; 503s with backends at 0 means pool collapse.
  • Correlating traefik_service_retries_total with the 503 and latency charts shows whether retry amplification is feeding the failure the breaker is shedding, which is the escalation path in Traefik cascading backend failure.
  • Latency histograms from traefik_service_request_duration_seconds let you watch the quantile your LatencyAtQuantileMS expression trips on, so you can see a threshold approaching before the breaker opens.
  • Because Traefik exposes no breaker-state metric, anomaly detection on the 503 rate is the practical early warning: the periodic open/recover cadence of a flapping breaker shows up as a repeating anomaly pattern rather than a flat threshold breach.