The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-retry-amplification ▌

Operations Guides

Traefik retry amplification: how retries turn a slow backend into an outage

A backend service starts degrading. Responses get slow, some connections fail. Traefik’s retry middleware re-sends the failed requests, usually to other backends in the pool. Backend load doubles or triples. The struggling service falls further behind, which produces more failures, which produces more retries. Within minutes, a partial degradation becomes a complete outage, and Traefik is doing a significant share of the damage while trying to help.

The worst part: this loop can be invisible on the dashboards your team actually watches. If a request fails twice but succeeds on the third attempt, the client got a 200. Your success-rate graphs look fine while the backend absorbs 3x the real client traffic and slides toward collapse.

This guide covers how the loop forms, how to confirm you are in one, and how to break it safely.

What this means

The retry middleware exists to mask transient backend failures. A connection reset, a brief network blip, a backend restarting mid-request: retrying against another server in the load-balancer pool turns a client-visible error into a success. That is useful behavior when failures are rare and uncorrelated.

It becomes dangerous when failures are correlated with load. The retry budget is per-request: every failed attempt can be re-sent, and each re-send consumes backend capacity. When the backend’s problem is that it is overloaded (database lock contention, connection pool exhaustion, a slow downstream dependency), adding more requests makes the root cause worse. The retry mechanism assumes failures are random; under saturation they are causal, and every retry feeds the cause.

Two aggravating factors make this worse in practice:

  • Health checks may still pass. Health check paths often differ from real traffic paths. A backend can answer /health in 5ms while /api/v1/data times out. Traefik keeps the backend in rotation and keeps retrying into it. See Traefik health checks pass but requests fail.
  • The default retry triggers are transport failures, with optional status-code ranges. By default, status is empty and disableRetryOnNetworkError is false, so an unserved response can be retried; an ordinary HTTP 500 is not retried unless you add it to status. A degrading backend still produces timeouts, resets, and hung connections, so the default trigger overlaps the overload failure mode.
flowchart TD
  A[Backend degrades: slow responses, transport errors] --> B[Retry middleware re-sends failed requests]
  B --> C[Backend load doubles or triples]
  C --> D[More timeouts and connection failures]
  D --> B
  C --> E[Fewer effective backends: pool capacity drops]
  E --> F[Survivors get original traffic plus retry traffic]
  F --> D
  D --> G[All backends saturated: 502/503/504 to clients]

Common causes

The loop always needs two things: a backend that is degrading, and retries that multiply the load. The initial degradation is the root cause; the retry config determines how fast it escalates.

CauseWhat it looks likeFirst thing to check
Backend resource saturation (DB locks, connection pool exhaustion)Latency climbs steadily, then transport errors appear; retries and latency rise togetherBackend CPU, memory, DB connection counts, slow query log
Downstream dependency failure cascading upstreamMultiple backends slow simultaneously because they share a database or cacheThe shared dependency’s health, not the backends themselves
Retry attempts configured too high for the workloadtraefik_service_retries_total rate approaching or exceeding request rateThe retry middleware config: attempts count and which routers use it
Traffic spike against insufficient headroomSudden request-rate increase preceding the latency climbtraefik_service_requests_total rate vs. baseline
Rolling deployment in progressRetries elevated but latency normal; self-resolvingWhether a deployment is actually running (do not confuse this with amplification)
Non-idempotent requests being retriedDuplicate records, double charges, repeated side effects reported downstreamWhether retryNonIdempotentMethod is enabled and which methods the retried routes accept

The distinguishing test: if retries are high but latency is low, you are looking at instance flapping (common during rolling updates, usually not dangerous). If retries and latency are both rising together, you are in amplification territory.

Quick checks

All read-only. Run against Traefik’s metrics endpoint (adjust host and port to your deployment).

# Retry counters per service
curl -s http://localhost:8080/metrics | grep traefik_service_retries_total

# Request counters per service, with status codes
curl -s http://localhost:8080/metrics | grep traefik_service_requests_total

# Backend health status (only present for services with health checks enabled)
curl -s http://localhost:8080/metrics | grep traefik_service_server_up

# Service latency histograms
curl -s http://localhost:8080/metrics | grep traefik_service_request_duration_seconds

The critical computation is the retry-to-request ratio. Take two samples of traefik_service_retries_total and traefik_service_requests_total 60 seconds apart, compute the per-second rates for the affected service, and divide:

  • Retries/requests under 1%: occasional blips. Not an incident by itself.
  • Retries/requests over 5% sustained: backends are unhealthy and retries are amplifying load. Investigate now.
  • Retries/requests approaching or exceeding 1.0: every request is being retried at least once. You are deep in the loop and Traefik is sending roughly double the client traffic to the backends. Treat as urgent.

Also check whether traefik_service_server_up shows all backends as 1 while 5xx rates and retries climb. That combination means the health checks are passing but real traffic is failing, which is the classic setup for this incident.

How to diagnose it

  1. Identify the affected service. Find which service label on traefik_service_retries_total is spiking. Do not start with aggregate dashboards; you need the per-service view.

  2. Confirm the amplification signature. For that service, verify all three signals moving together: retry rate rising sharply, traefik_service_request_duration_seconds p95/p99 climbing, and 5xx rate moderate and rising. All three together is the pattern. Retries alone, with flat latency, is flapping, not amplification.

  3. Compute the retry ratio as described above. This tells you severity and how much extra load Traefik is adding.

  4. Check backend health signals. Look at traefik_service_server_up per backend URL. Declining or flapping values mean the pool is shrinking and survivors are absorbing original plus retry traffic. All values at 1 with high 5xx means health checks are misleading.

  5. Find the original failure. The retry loop is a symptom. Check the backend’s own resources: CPU, memory, database connection pool, lock contention, downstream dependencies. A shared dependency (database, cache) failing is a common trigger because it degrades every backend at once.

  6. Check what changed. A slow code deployment, a database query regression, an autoscaler scale-down, or a sudden traffic spike. The trigger explains why the backend degraded before the first retry was ever sent.

  7. Check request methods on the retried routes. If the affected routes accept POST/PUT and retries are firing, assume duplicate side effects have occurred and plan reconciliation.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_service_retries_total (rate)The amplification mechanism itself; counts each retry per serviceRate >5% of request rate sustained; approaching request rate is severe
traefik_service_request_duration_seconds (p95/p99)Shows the degradation the retries are responding to and worseningRising together with retries is the amplification signature
traefik_service_requests_total{code=~"5.."}Distinguishes Traefik-generated errors (502/503/504) from backend pass-throughModerate and rising alongside retries; 503 means pool collapse has started
traefik_service_server_upPool capacity; each backend dropping out concentrates load on survivorsAny backend at 0, flapping values, or declining count over time
Retry-to-request ratioThe single best top-level indicator for this failure modeSustained >5%; >50% means amplification is actively escalating

Note that traefik_service_server_up only exists for services with health checks configured. Absence of the series means unmonitored, not healthy.

Fixes

Break the loop first: reduce retry attempts

The first response is to cut retries on the affected service. Reduce the retry middleware’s attempts value (or remove the middleware from the router chain entirely) via dynamic configuration. This stops Traefik from multiplying load while you fix the actual problem.

Tradeoff: you lose masking of genuinely transient failures, so some clients will see errors they would not have seen before. During an active amplification incident, that is the right trade. A few visible errors are far better than a total outage.

Do not restart Traefik as a first move. It does not fix the backend, and it drops all in-flight connections.

Fix the root cause at the backend

Retries are never the root cause. Work the backend problem:

  • Database or connection pool exhaustion: kill blocking queries, raise pool limits, or shed load.
  • Bad deployment: roll back. This is the fastest fix when the degradation correlates with a release.
  • Shared dependency down: failing over or restoring the database/cache recovers all backends at once.
  • Insufficient capacity: scale the backend pool out. But note that adding backends during an active retry storm means new instances join a pool receiving amplified traffic; cut retries first, then scale.

Right-size the retry configuration

After the incident, revisit the retry policy rather than restoring it blindly:

  • Keep attempts low. Each additional attempt is another full request against the backend pool. For most services, one or two retries is the ceiling of what is safe under load.
  • Scope retries to idempotent routes. Retrying GET is usually safe. By default the middleware excludes non-idempotent POST, LOCK, and PATCH; only set retryNonIdempotentMethod: true deliberately, because it can duplicate side effects. Apply retry middlewares only to routers serving idempotent traffic.
  • Do not stack retries. If clients also retry aggressively on timeout, and the client timeout is shorter than Traefik’s full retry sequence, client-side retries multiply on top of Traefik’s. Coordinate timeouts so the client gives up after Traefik, not before.

Prevention

  • Alert on the retry-to-request ratio, not just error rates. Retries mask errors from clients, so a success-rate dashboard will not warn you. The ratio is the leading indicator: it rises before 5xx does.
  • Dashboard the amplification triad. Retry rate, service latency, and 5xx rate on the same per-service graph. The visual correlation is the fastest way to recognize the pattern at 3 a.m.
  • Make health checks meaningful. If health checks hit a lightweight path while real traffic stresses the database, Traefik will keep routing and retrying into functionally impaired backends. Health endpoints should exercise the dependencies that matter.
  • Document a per-service retry budget. Decide which services get retries at all, how many attempts, and which methods. Default-off for non-idempotent routes.
  • Load-test the failure mode. Deliberately degrade a backend in staging with retries enabled and watch the ratio. Teams that have seen the loop once in a controlled setting recognize it instantly in production.

How Netdata helps

  • Per-service retry rates out of the box: Netdata charts traefik_service_retries_total per service, so the spiking service is visible immediately rather than buried in an aggregate.
  • The amplification triad on one screen: retries, traefik_service_request_duration_seconds percentiles, and 5xx rates per service are correlated on the same dashboard, which is exactly the comparison this diagnosis requires.
  • Retry-to-request ratio alerting: alert on the ratio crossing thresholds (for example >5% sustained) so you are paged on the leading indicator, not on the 503s that arrive after pool collapse.
  • Backend pool visibility: traefik_service_server_up per backend URL shows the pool shrinking in real time as the cascade progresses.
  • ML anomaly detection on retry counters: catches retry rates deviating from baseline even when they have not yet crossed a static threshold, which matters because rolling deployments make naive thresholds noisy.