The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-traffic-drop-to-zero ▌

Operations Guides

Traefik traffic dropped to zero: entrypoint request rate flatlined

The request rate on an entrypoint that normally serves traffic has fallen to zero, or close enough. Clients are timing out or getting connection errors. Sometimes Traefik’s process is still running and /ping still returns 200, which makes this worse: your health checks are green while no traffic is flowing.

There are two fundamentally different families of root cause, and telling them apart early is the whole game. Either traffic is not reaching Traefik at all (DNS, cloud load balancer, firewall, network partition upstream of the proxy), or traffic is arriving at the host but Traefik cannot accept it (a dead listener, file descriptor exhaustion, a port-bind failure after restart). The checks below are ordered to split those two families within the first few minutes.

A note on detection: a flatlined entrypoint is an anomaly-detection problem, not a static-threshold problem. rate(traefik_entrypoint_requests_total[5m]) == 0 will also fire at 3 a.m. on a genuinely quiet service. What makes this an incident is the deviation from the established baseline for that entrypoint at that time of day, sustained for more than a few minutes. If you are reading this because an alert fired, confirm the baseline deviation before assuming an outage.

What this means

traefik_entrypoint_requests_total is a counter with labels {code, method, protocol, entrypoint}. It increments once per HTTP request accepted at the entrypoint listener. If the rate is zero, one of three things is true:

  1. No packets arrive. Something upstream of Traefik (DNS record, cloud LB target health, security group, firewall, upstream router) is sending traffic elsewhere or nowhere. Traefik is fine; the path to it is not.
  2. Packets arrive but nothing accepts them. The entrypoint listener is gone or broken. The process can stay alive and healthy-looking in this state. There is a long-running upstream bug (traefik/traefik#8841, open since 2022 and still reported on recent v3.x releases) where a failed setsockopt on an accepted connection kills the listener’s accept loop, leaving set tcp 10.0.0.1:443: setsockopt: invalid argument in the logs and a dead entrypoint behind a living process.
  3. Traefik is alive but saturated. File descriptor exhaustion is the classic case: at the FD limit, 100% of new connections fail instantly while existing connections keep working. From the outside this looks like intermittent or total connection refusal. From the metrics it looks like a request-rate collapse with traefik_open_connections pinned at a plateau.

The less common direction matters too: if you came here because of a sudden spike rather than a drop, that is a different problem (traffic flood or DDoS), and the diagnostics below do not apply.

flowchart TD
  A[Entrypoint request rate flatlined] --> B{Process alive and port listening?}
  B -->|No| C[Process dead or bind failed]
  B -->|Yes| D{SYNs arriving at the host?}
  D -->|No| E[Upstream problem: DNS, cloud LB, firewall, partition]
  D -->|Yes| F{FD ratio near limit?}
  F -->|Yes| G[FD exhaustion: new connections refused]
  F -->|No| H{setsockopt errors in logs?}
  H -->|Yes| I[Dead accept loop, listener zombie: restart Traefik]
  H -->|No| J[Check open connections, TLS, and recent deploys]

Common causes

CauseWhat it looks likeFirst thing to check
Cloud LB or upstream health check failingLB targets marked unhealthy, traffic drained; Traefik itself fineLB target health in the cloud console; TCP connect to the entrypoint from outside the LB
DNS or upstream network partitionRate fell to zero instantly across all entrypoints at onceResolve the public hostname, trace the path, check upstream provider status
Firewall or security group changeConnections time out rather than refuse; often follows an infra changess -tn state syn-recv on the host while a client retries
FD exhaustionprocess_open_fds / process_max_fds at or near 1; existing connections work, new ones failprocess_open_fds vs process_max_fds on the metrics endpoint
Listener accept loop dead (setsockopt bug)Process alive, /ping 200, port may still appear listening, but zero accepts; setsockopt: invalid argument in logsgrep setsockopt on Traefik logs
Port-bind failure after restartProcess restarted recently and never re-bound the entrypoint; “listen tcp :443: bind: address already in use” style errorsprocess_start_time_seconds, then logs from startup
Unexpected graceful shutdownLogs show “Stopping server gracefully” with no operator actionTraefik logs around the rate drop
Legitimate traffic migrationRate dropped here but rose on another instance or entrypointCompare traefik_entrypoint_requests_total across all instances and entrypoints

Quick checks

Run these from the Traefik host or pod. All are read-only. Port 8080 below assumes the default internal entrypoint where /ping and /metrics live; adjust for your configuration. Adjust the log path in check 7 as well (or use docker logs / journalctl -u traefik).

# 1. Is the process alive, and when did it start?
pgrep -af traefik
grep 'Max open files' /proc/$(pgrep -f traefik | head -1)/limits

# 2. Is the entrypoint port actually listening?
ss -tlnp | grep -E ':(80|443)\b'

# 3. Are connection attempts arriving? (SYN received but not answered, or nothing at all)
ss -tn state syn-recv | head
ss -s

# 4. FD usage: current vs limit
ls /proc/$(pgrep -f traefik | head -1)/fd | wc -l

# 5. Metrics: request rate, open connections, FD ratio, reload freshness
curl -s http://localhost:8080/metrics | grep -E 'traefik_entrypoint_requests_total|traefik_open_connections|process_open_fds|process_max_fds|traefik_config_last_reload_success'

# 6. Ping (liveness only - remember /ping says nothing about listeners or routing)
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8080/ping

# 7. The silent-listener bug signature and shutdown signatures
grep -iE 'setsockopt|stopping server|graceful' /var/log/traefik/traefik.log | tail -20

# 8. TCP connect from an outside host, bypassing any LB
nc -zv <traefik-host-ip> 443

Interpretation shortcuts:

  • Port not listening + recent start time: bind failure or a crash-looping process that keeps losing the port race. Check startup logs for the bind error.
  • Port listening + zero SYNs arriving: the problem is upstream of the host. Stop debugging Traefik.
  • SYNs arriving + FD ratio near 1: FD exhaustion. New connections cannot be accepted.
  • Port listening + SYNs arriving + FDs fine + setsockopt in logs: the accept loop is dead. Only a restart recovers the listener.
  • Everything normal on this instance: check whether another instance took the traffic (failover, DNS change, LB weight change). A flatline plus a matching spike elsewhere is a migration, not an outage.

How to diagnose it

  1. Confirm the anomaly is real. Compare the current rate against the same time-of-day baseline for that entrypoint. Nights, weekends, and planned maintenance windows flatline entrypoints legitimately. Also confirm you are summing correctly: a label change or a new instance name in the query can look like a traffic drop.

  2. Split upstream vs local. From a host outside your network path (or at least outside the LB), attempt a TCP connect to the entrypoint. Simultaneously watch ss -tn state syn-recv on the Traefik host. If nothing arrives, the fault is DNS, the cloud LB (target health, listener rules), a security group, or a network partition. Escalate there; Traefik is a victim, not a cause.

  3. Check the FD ratio. Compute process_open_fds / process_max_fds. Above 0.95, you are at the cliff edge: existing keep-alive connections continue to serve (which is why some traffic may still flow) but new connections are refused instantly. If process_max_fds is 1024, that is the default container limit and it is too low for a production edge proxy.

  4. Look for the dead-listener signature. Search Traefik logs for setsockopt: invalid argument or setsockopt: operation not supported. The first is the long-standing bug in which a failed setsockopt on an accepted connection propagates up and kills the listener’s accept loop while the process keeps running (issue #8841). The second was caused by MPTCP support in v3.4.2 through v3.4.4 and was fixed in v3.4.5 by removing MPTCP; if you are on those versions, upgrade. The v3.4.5 change removes the MPTCP path only; the separate setsockopt: invalid argument issue remains open and was not fixed by that release.

  5. Rule out an unexpected shutdown or restart. Look for “Stopping server gracefully” without an operator action, and check process_start_time_seconds for a recent restart. There are upstream reports of spontaneous graceful stops; in at least one reported case an upgrade resolved it. Also check whether a container healthcheck or orchestrator event (eviction, node drain, OOM kill of a sidecar) triggered it.

  6. Check for provider desync as a contributing factor. Provider disconnection does not normally zero out an entrypoint (stale routes keep serving), but if the rate drop coincided with a mass route change, check traefik_config_last_reload_success age and entrypoint-level 404s. If traffic is arriving but every request gets a Traefik-generated 404, you want Traefik 404 not found: requests arriving with no matching router instead.

  7. If everything local is clean, widen the blast radius check. Compare request rates across all Traefik instances and all entrypoints. Traffic that vanished here and appeared elsewhere is a routing or LB decision upstream. Traffic that vanished everywhere is DNS, a shared LB, or a genuine demand change.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
rate(traefik_entrypoint_requests_total[5m]) by entrypointThe flatline itselfSustained >50% drop from the same-time-of-day baseline without a known cause
process_open_fds / process_max_fdsFD exhaustion zeroes new connections before any other signal movesRatio >0.80 trending up; >0.95 is cliff edge
traefik_open_connections by entrypointPlateau at the limit while request rate falls means connections cannot be accepted or closedFlat ceiling coinciding with the rate drop
process_start_time_secondsCatches restarts and crash loops behind a flatlineRecent change matching the rate drop
traefik_config_last_reload_success ageStale config after provider loss can strand routesTimestamp frozen while deployments are happening
traefik_entrypoint_requests_total{code="404"}Distinguishes “no traffic” from “traffic arriving but no router matched”Rising 404 share alongside falling total rate
Scrape target up / /pingBaseline liveness; absence of metrics is itself a signalTarget down, or /ping non-200

Fixes

Upstream reachability (DNS, LB, firewall)

Restore the path, not the proxy. Traefik is healthy in this scenario; restarting it accomplishes nothing. Re-point or repair DNS, fix the LB health check or target group, or roll back the firewall/security-group change. The common trap: LB health checks that target an entrypoint port Traefik stopped serving, or a timeout mismatch between the cloud LB idle timeout and Traefik’s respondingTimeouts.idleTimeout causing the LB to reuse dead connections. Align those timeouts so the LB’s idle timeout is shorter than Traefik’s.

FD exhaustion

Raise the limit and restart to clear the pressure. The immediate relief is a restart, which closes all FDs, but it drops every in-flight connection, so treat it as a deliberate action, not a reflex. Then fix the limit: 1024 (the common container default) is not viable for an edge proxy. Set a high nofile ulimit in the container spec, compose file, or systemd unit. Afterward, investigate why FDs grew: long-lived WebSocket/gRPC connections, a connection leak (traefik_open_connections growing without matching request-rate growth), or keep-alive misconfiguration.

Dead accept loop (setsockopt bug)

Restart is the only reliable recovery. There is no confirmed fix upstream as of this writing; restart Traefik and the listener recovers. Because the process stays alive and /ping stays green, add detection that survives a “healthy” process: alert on the entrypoint rate anomaly itself, and grep logs for setsockopt as a confirmation signal. If you are on v3.4.2 through v3.4.4, upgrade to v3.4.5 or later to eliminate the MPTCP variant. For the separate invalid-argument issue, no confirmed upstream workaround beyond restart plus external dead-listener detection exists; issue #8841 remains open.

Bind failure after restart

Find what owns the port. Check startup logs for the bind error, then ss -tlnp to see what holds the port. Common cases: a previous Traefik process not fully terminated, two replicas scheduled to the same host port, or missing capabilities in a hardened container. Fix the conflict and let the process rebind cleanly.

Unexpected graceful shutdown

Identify the trigger before restarting blindly. Check orchestrator events (evictions, node drains, healthcheck failures) and any external process supervisor. If logs show a graceful stop with no external signal, note the version and consider upgrading, since at least one reported instance of this pattern was resolved by an upgrade.

Prevention

  • Alert on baseline deviation, not zero. Static rate == 0 alerts false-fire on quiet services. Alert on a sustained drop versus the same-time-of-day baseline, combined with a traffic floor where one exists. Corroborate with traefik_open_connections and the FD ratio so a single metric cannot page you alone.
  • Monitor the FD ratio with headroom. Page at >0.95 sustained and rising; ticket at >0.80. Keep steady state below 70% of process_max_fds.
  • Raise FD limits deliberately. Never run a production Traefik at the 1024 default.
  • Log-scan for setsockopt. It is the earliest reliable indicator of the dead-listener bug; the process will not tell you otherwise.
  • Do not trust /ping. It proves the process is alive, nothing more. It does not check listeners, routing, providers, or certificates.
  • Watch per-instance rates in HA. A flatline on one replica with healthy siblings is an instance-level fault; a flatline everywhere is upstream. Compare traefik_config_last_reload_success across replicas to catch drift.

How Netdata helps

  • Per-second entrypoint throughput makes the exact moment of the flatline visible, which lets you align it with deploys, orchestrator events, and log lines instead of guessing at the timeline.
  • FD ratio tracking (process_open_fds vs process_max_fds) alongside open connections shows the exhaustion cliff forming before the request rate collapses, turning a page into a ticket.
  • Process restart detection via start-time changes correlates a traffic drop with a crash loop or bind failure without manual log archaeology.
  • Anomaly detection on request rates handles the “what is normal varies per deployment” problem: it flags deviations from learned baselines per entrypoint rather than relying on static thresholds.
  • Cross-signal correlation in one view (request rate, open connections, FDs, config reload freshness, 404 share) is what separates the six causes in the table above quickly, instead of checking each in isolation.