The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik
TRAEFIK · OPERATIONS PLAYBOOK

Traefik is a proxy and a config reconciler at once — and it fails silently at the seam between them

A Go edge router that rebuilds its routing table live from providers it watches, terminates TLS, load-balances to backends it health-checks, and renews its own certificates. When a provider watch dies, a backend pool collapses, a file-descriptor limit is hit, or an ACME renewal quietly fails, Traefik itself stays up and /ping keeps returning 200. We trace how that design behaves under load and what to do when it breaks.

"

Traefik's dynamic configuration gets you to production in minutes, then hands you a set of failure modes that stay invisible precisely because the process never crashes.

The defaults work. Until a provider watch (the Kubernetes API, the Docker socket, Consul, a file) dies and Traefik silently keeps serving its last-known config — new services never appear, removed backends keep taking traffic, and /ping still returns 200. Until every backend for a service fails its health check and each request comes back 503 Service Unavailable. Until the file-descriptor limit — often the container default of 1024 — is hit and new connections are refused instantly with too many open files. Until an ACME renewal quietly fails for 53 days and the certificate expires under live traffic. Until a slow backend triggers retries that amplify the load and finish it off.

These guides are written for engineers who already run Traefik, not for people learning what a reverse proxy is. The goal is the mental model of how the router actually behaves under load, the failure patterns that keep recurring, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident.

How Traefik actually runs in production

Traefik is not just a proxy. It is a request pipeline — entrypoint, router, middleware chain, service load-balancer, backend — fed by provider watchers that continuously rebuild the routing table, wrapped in a Go runtime that terminates TLS and manages certificates. Most production failures live between these layers, not inside any one of them.

01
entrypoints / clients
TCP/UDP listeners bound to ports. Each accepts connections and spawns a goroutine per connection; TLS is terminated here. Every client, backend, and provider connection costs a file descriptor. <code>traefik_open_connections</code> tracks the entrypoint count; <code>process_open_fds</code> is the real ceiling.
ENTRY
02
routers
Rule-based matchers evaluating Host, Path, Headers, and SNI in priority order. A request that matches no router returns a 404 at the entrypoint level (<code>traefik_entrypoint_requests_total{code="404"}</code>) — a routing fault, not a backend one. Routers are rebuilt on every configuration change.
ROUTE
03
middleware chain
An ordered pipeline: auth, rate limiting, compression, buffering, circuit breakers, retries. Each middleware can short-circuit the chain and each adds CPU. There are no per-middleware metrics — you infer their cost from the gap between entrypoint and service latency.
MIDDLEWARE
04
services / load balancer
Backend pools with weighted round-robin and health checks. <code>traefik_service_server_up</code> is 0/1 per backend — but only exists when health checks are configured. When every backend is down, the service returns <code>503</code> while Traefik itself stays perfectly healthy.
SERVICE
05
backends + connection pool
Go <code>net/http</code> transports pool connections to each backend. Slow or hung backends leave connections in ESTABLISHED, CLOSE_WAIT, or TIME_WAIT; poor reuse burns ephemeral ports. This layer produces most 502/504 errors — the backend, not Traefik, is the fault.
BACKEND
06
providers + config aggregator
Background watchers poll Kubernetes, Docker, Consul, etcd, or files and feed a configuration aggregator that atomically swaps the routing table. On disconnect Traefik retains the old config and retries — it does not flush routes. <code>traefik_config_last_reload_success</code> is the freshness signal.
PROVIDER
07
ACME / certificate resolver
A background process issues and renews Let's Encrypt certificates via HTTP-01, TLS-ALPN-01, or DNS-01 challenges, storing them in <code>acme.json</code> or a KV store. There is no renewal-failure metric — <code>traefik_tls_certs_not_after</code> approaching now is the only in-band signal that renewal has broken.
TLS
08
Go runtime (goroutines, heap, GC)
One goroutine per connection plus watchers and health checkers; the heap holds the routing table and buffers. A goroutine leak from hung backends grows <code>go_goroutines</code> and RSS in lockstep until an instant OOM kill. TLS handshakes dominate CPU.
RUNTIME

Why this matters: 'Traefik is returning errors' or 'my service isn't reachable' can come from a backend that is unreachable (502), a service with no healthy backends (503), a backend that is merely slow (504), a request that matched no router (404), a frozen provider serving stale config, an expired certificate, a file-descriptor cliff, a goroutine leak, or retry amplification. The symptom rhymes but each layer has a different signal — and a different fix.

The failures you'll actually see

Most Traefik incidents fall into a small set of recurring patterns. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The backend blackout

Every backend for a service fails its health check, Traefik has nowhere to send the request, and each one comes back 503 Service Unavailable. traefik_service_server_up reads 0 for every URL in the service. Traefik itself is perfectly healthy — /ping returns 200, the process is fine — which is exactly why teams look in the wrong place first. Common after a bad deploy, a shared-dependency failure, or a health-check path that doesn't match reality.

  • 503 Service Unavailable on every request to the service
  • traefik_service_server_up = 0 for all URLs in the service
  • Traefik process healthy, /ping returning 200
  • Backends may pass their own health endpoint while the app path fails
Investigate
CRITICAL

The bad-gateway storm

Traefik reaches the backend but gets an invalid response or a reset connection, returning 502 Bad Gateway. If retry middleware is enabled, each failure is re-sent to other backends, doubling or tripling load on an already-degraded pool — Traefik effectively DDoSes the backend while trying to help. Retries and latency climb together; when retries per second exceed requests per second, a partial failure is turning into a total one.

  • 502 Bad Gateway at the service level
  • traefik_service_retries_total spiking with rising latency
  • Backends flapping (server_up alternating 0/1)
  • CLOSE_WAIT sockets or ephemeral ports climbing
Investigate
CRITICAL

The file-descriptor cliff

Traefik sits at the edge and every connection costs a file descriptor. When process_open_fds reaches process_max_fds — often the container default of 1024 — new connections are refused instantly with too many open files. Existing connections keep working, so throughput looks fine while new users are dropped. This is a cliff-edge with zero graceful degradation, and the single most common cause of 'mysterious' production outages.

  • accept4: too many open files in the logs
  • process_open_fds / process_max_fds above 95% and rising
  • New connections refused while existing ones continue
  • process_max_fds stuck at the 1024 default
Investigate
IMMINENT

The silent config drift

A provider watch loses connectivity — the Kubernetes API becomes unreachable, the Docker socket is unmounted, an RBAC token expires — and Traefik keeps serving its last-known-good configuration. Established routes work, so nothing fires. But new services are never routed and removed backends keep receiving traffic. The blast radius grows with every deploy, and it's usually discovered hours later when a developer reports 'my new service isn't reachable'.

  • traefik_config_last_reload_success frozen while deploys happen
  • traefik_config_reloads_total flat (or reloads failing)
  • Entrypoint 404s rising on routes that exist in the provider
  • /ping healthy, existing routes unaffected
Investigate
IMMINENT

The certificate cliff

ACME renewal has been failing silently — a blocked HTTP-01 port, a rotated DNS token, a hit rate limit, or an acme.json permission change. Traefik keeps serving the current certificate until it expires, then every HTTPS client is rejected. There is no renewal-failure metric; by the time a certificate is 7 days from expiry, renewal has been failing for roughly 53 days. The failure is invisible until users report TLS errors.

  • traefik_tls_certs_not_after approaching the current time
  • ACME errors in the logs (challenge failed, rate limited, lock)
  • Certificate serial not changing across days
  • A trickle of TLS errors from strict clients before full expiry
Investigate
ACTIVE

Routing into the void

Requests arrive for a hostname or path that matches no router and Traefik returns a 404 at the entrypoint level — its own 404, not the backend's. Either the route was never created (a silently-ignored annotation or a wrong prefix), or the config went stale, or someone is scanning. Teams waste time investigating the application when the fault is in Traefik's routing table; the two 404s look identical in a log but live in different layers.

  • traefik_entrypoint_requests_total{code="404"} elevated or rising
  • Routes missing from /api/http/routers that exist in the provider
  • 404s appearing in step with a deployment
  • Annotations or labels present but silently not applied
Investigate
Choosing a tool

Best Traefik Monitoring Tools (10 Ranked for 2026)

A ranked review of the tools teams actually shortlist here, what each one is genuinely good at, and how the pricing behaves as you scale.

Traefik monitoring maturity levels

Traefik observability works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much your edge matters. Most production deployments should land at the second level.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: is Traefik alive and passing traffic? You will not learn what broke, but you will learn that something broke before users do. Survival is enough for dev instances and non-critical edges.

  • Process alive / scrape reachable If the metrics target is down, the proxy is down — no traffic passes.
  • Entrypoint request rate > 0 traefik_entrypoint_requests_total flatlining means traffic isn't arriving or can't be accepted.
  • 5xx / 503 presence Any 503 means a service has no healthy backends left.
  • File descriptor ratio process_open_fds / process_max_fds — the cliff-edge that drops new connections.
  • TLS certificate expiry traefik_tls_certs_not_after within 14 days means renewal has likely already failed.
  • Process memory (RSS) vs limit Go GC hides pressure until the OOM kill is instant.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production deployments should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose an incident on its own: backend failures, stale config, cert expiry, resource saturation, routing gaps.

  • Per-service 5xx by code 502 vs 503 vs 504 have different root causes; never lump them.
  • Backend server up per service traefik_service_server_up = 0 across a service is an all-backends-down 503.
  • Per-service latency p95 / p99 Establish per-service baselines; one threshold cannot fit every backend.
  • Config reload freshness traefik_config_last_reload_success age is the only signal for silent provider desync.
  • Open connections per entrypoint Rising without traffic growth is a leak or backends backing up.
  • File descriptor ratio Alert at 80%, page at 95% and still rising.
  • Goroutine count (baseline deviation) go_goroutines climbing off traffic is the earliest leak signal.
  • Retry rate traefik_service_retries_total ratio to requests predicts an amplification cascade.
  • Entrypoint 404 rate No router matched — a routing fault, distinct from a backend 404.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. Memory trending toward OOM, connection pools degrading, retries starting to amplify, a certificate quietly failing to renew, config churn eating CPU. None of these page you on day one. They become page-out incidents on day thirty.

  • Go heap and GC pause times go_memstats_heap_inuse_bytes trend and go_gc_duration_seconds p99.
  • CPU utilisation Attribute it: TLS handshakes, compression, or config rebuilds.
  • Retry-to-request ratio The amplification early warning, not the raw retry count.
  • TLS version distribution Sustained TLS 1.0/1.1 is legacy clients or a downgrade attack.
  • TIME_WAIT / CLOSE_WAIT counts Ephemeral-port exhaustion and connection leaks at the OS layer.
  • Config reload rate Above ~1/s is a rebuild storm burning CPU on churn.
  • Entrypoint vs service latency The gap is Traefik's own middleware/TLS overhead.
  • ACME renewal tracking Alert when renewal has failed for 2 days, not at 7 days to expiry.

Level 4: Expert

Reactive instrumentation after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. Per-router cardinality, provider-watcher liveness, ACME rate-limit consumption, per-instance config drift, kernel accept-queue drops. Most teams never need every signal here. Add the ones your incident history says you do.

  • Per-router metrics (addRoutersLabels) High-cardinality; only where you need per-route attribution.
  • Provider watcher liveness A watcher can panic and stop while the process stays up — absence, not failure.
  • DNS resolution time for backends Slow service-name resolution adds invisible latency in K8s/Docker.
  • ACME rate-limit consumption Certs issued this week vs the Let's Encrypt weekly ceiling.
  • Per-instance config consistency Compare reload timestamps and router hashes across replicas.
  • Connection reuse ratio Reused vs total connections — pool efficiency and port pressure.
  • Dashboard / API exposure probes Is /api/rawdata reachable from outside the intended network?
  • acme.json integrity + kernel accept queue Storage corruption and SYN/accept-queue drops under load.

Operating mistakes worth avoiding

The traps Traefik teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after an incident.

Trusting /ping as a health indicator

<code>/ping</code> returns 200 if the process is alive — period. It does not check provider connectivity, route validity, backend reachability, or certificate state. A Traefik instance can be 'healthy' by <code>/ping</code> while routing to dead backends with expired certificates on stale configuration. Always supplement liveness with metric-based signals: backend health, config freshness, and certificate expiry.

Not monitoring configuration provider health

The single most common gap. Teams watch request handling obsessively but never check whether the provider is still delivering updates. When the Docker socket goes away or the Kubernetes API becomes unreachable, Traefik serves stale config for hours and <code>/ping</code> stays green. Alert on <code>traefik_config_last_reload_success</code> age — and remember a dead watcher stops incrementing <code>traefik_config_reloads_total</code>, so absence, not failure, is the signal.

Running with the default file descriptor limit

Traefik handles every incoming connection, yet containers often inherit the default 1024 FD limit. That is fine for staging and catastrophic under production load — the failure is a sudden cliff of <code>too many open files</code> with new connections refused and existing ones fine, which is why it's so often misdiagnosed. Raise both the OS <code>ulimit</code> and the container/systemd limit to 65536 or higher.

Enabling retries without watching the retry rate

Retry middleware masks backend instability — a request that failed twice and succeeded on the third try shows as 'success' — while quietly tripling backend load. When a backend is already degrading, that extra load pushes it over. Make the ratio of <code>traefik_service_retries_total</code> to request rate a top-level metric, and know that non-idempotent retries (POST, PUT) can cause duplicate side effects.

Lumping 502, 503, and 504 together as '5xx'

A 502 means Traefik reached the backend but got garbage; a 503 means no backends are available at all; a 504 means the backend was reachable but too slow. They have completely different root causes and fixes. Aggregating them as 'errors' leads straight to the wrong diagnosis — and comparing entrypoint-level with service-level 5xx tells you whether Traefik or the backend generated them.

Checking certificate expiry but not renewal health

Most teams alert when a certificate is 7 days from expiry. By then ACME renewal has been failing for roughly 53 days (a 90-day Let's Encrypt cert renews at 30 days out). Alert when renewal has failed for two consecutive days, not when expiry is imminent — and pair the <code>traefik_tls_certs_not_after</code> metric with an external synthetic probe of the real production hostname, since the metric tracks dormant and staging certs too.

Confusing entrypoint 404s with backend 404s

An entrypoint-level 404 means 'no router matched' — a Traefik configuration fault. A service-level 404 means 'the backend returned 404' — an application concern. They look identical in a raw log but require completely different investigation paths. Teams that don't split <code>traefik_entrypoint_requests_total{code="404"}</code> from service 404s waste an incident looking at the wrong layer.

Uncoordinated timeouts across the stack

The timeout chain client → load balancer → Traefik → backend must have each layer's timeout longer than the one behind it. The classic break is a cloud LB idle timeout longer than Traefik's <code>respondingTimeouts.idleTimeout</code>: Traefik closes an idle keep-alive connection the LB still reuses, producing intermittent 502s that are notoriously hard to reproduce. Coordinate <code>respondingTimeouts</code>, <code>forwardingTimeouts</code>, and transport settings with the rest of the stack.

Traefik runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up Traefik monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing.