The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-backend-connection-pool ▌

Operations Guides

Traefik backend connection pool: keep-alive, MaxIdleConnsPerHost, and reuse

Every request Traefik proxies needs a TCP connection to a backend. Whether that connection is freshly dialed or reused from a pool determines how much per-request overhead you pay, and it is one of the least monitored parts of the proxy. A misconfigured pool shows up as elevated latency on every request (TCP and possibly TLS handshake costs per request), as growing TIME_WAIT socket counts marching toward ephemeral port exhaustion, or as sporadic 502s with no backend outage to explain them.

The core problem is a default. Traefik’s backend transports are built on Go’s http.Transport, and the Go standard library defaults MaxIdleConnsPerHost to 2. That value was chosen for HTTP clients, not for a reverse proxy fanning thousands of requests per second into a handful of backends. Traefik’s global static transport defaults to 200 idle connections per host, however; the Go value applies when a named dynamic ServersTransport omits maxIdleConnsPerHost.

What the pool is and why it matters

Traefik’s load balancer holds a pool of net/http transports, one per backend target, and each transport maintains its own pool of idle keep-alive connections. When a request arrives for a service, the load balancer picks a backend server, then asks that server’s transport for a connection. If an idle connection exists in the pool, it is reused immediately. If not, a new connection is dialed.

The economics:

  • Reused connection: near-zero setup cost. The request goes out on an established TCP stream.
  • New connection: full TCP handshake (one round trip minimum), plus a TLS handshake if the backend scheme is HTTPS, plus ephemeral port allocation on the Traefik side.

At low traffic, this barely matters. At production rates, a pool that fails to reuse connections turns every request into a connection-establishment exercise. Each closed connection then sits in TIME_WAIT for roughly 60 seconds, holding an ephemeral port. This is the classic failure: everything looks fine on dashboards, then Traefik starts returning 502s because the kernel cannot assign a local port for a new backend connection.

How the pool works

flowchart LR
  REQ[Request matched to service] --> LB[Load balancer picks backend]
  LB --> POOL{Idle connection in pool?}
  POOL -->|yes| REUSE[Reuse connection]
  POOL -->|no| DIAL[Dial new connection]
  REUSE --> RESP[Response received]
  DIAL --> RESP
  RESP --> CLOSEHDR{Connection close requested?}
  CLOSEHDR -->|no| IDLE[Return to idle pool]
  CLOSEHDR -->|yes| CLOSE[Close socket - TIME_WAIT]
  IDLE --> EVICT{idleConnTimeout expired?}
  EVICT -->|yes| CLOSE
  EVICT -->|no| IDLE

A connection leaves the idle pool in one of three ways:

  1. Reuse: a new request takes it.
  2. Local eviction: Traefik’s forwardingTimeouts.idleConnTimeout expires and closes it.
  3. Remote death: the backend, a NAT gateway, or a stateful firewall closes it first. Traefik does not learn about this until it tries to reuse the connection, at which point the write fails.

Case 3 is the dangerous one. The connection looks healthy in the pool, gets handed to a real request, and the request dies.

The defaults and where they bite

Traefik exposes the pool through serversTransport, configured globally in static config or per-service via a dynamic ServersTransport resource.

SettingDefaultWhat it controls
maxIdleConnsPerHostglobal static default: 200; dynamic resource default: 0, which falls back to 2Idle keep-alive connections retained per backend host
forwardingTimeouts.idleConnTimeout90s (zero means no limit)How long an idle pooled connection is kept before Traefik closes it
forwardingTimeouts.dialTimeout30sTime allowed to establish a new backend connection

Three behaviors trip operators up:

  • Zero is not unlimited. Setting maxIdleConnsPerHost: 0 does not mean “no cap”. It falls back to the default of 2. To disable connection reuse entirely, set it to -1. Note that -1 also makes idleConnTimeout meaningless, because there are no idle connections left to time out, and it maximizes churn: expect a large TIME_WAIT population and higher per-request latency.
  • A value of 2 is per host, per transport, and only caps retention. Under concurrent load, Go’s transport opens as many connections as in-flight requests demand; maxIdleConnsPerHost only governs how many are kept afterward. With 2 retained, a burst of 50 concurrent requests to one backend opens 50 connections, then closes 48 of them when the burst drains. The next burst pays full dial cost again and adds 48 more sockets to TIME_WAIT.
  • Kubernetes CRD users: Traefik v3.4 through v3.5.2 set the ServersTransport CRD validation minimum for maxIdleConnsPerHost to 0, which made -1 impossible to express via CRD. PR #12077 corrected the minimum to -1 in v3.5.3 and later v3 releases.

Migration history matters here too: in Traefik v1.x the global default was 200 idle connections per host. Traefik v2 moved the global setting under serversTransport and retained its 200 default, but a named dynamic ServersTransport that omits the field falls back to Go’s stdlib value of 2. Verify which transport each service actually uses.

Where pool problems show up in production

Ephemeral port exhaustion and TIME_WAIT growth

When connections are created and destroyed per request instead of reused, each one leaves a socket in TIME_WAIT for about 60 seconds. The steady-state TIME_WAIT population is roughly 60 times the new-connection rate. Common drivers:

  • A named dynamic ServersTransport whose omitted maxIdleConnsPerHost falls back to 2 on a high-throughput service.
  • Backends responding with Connection: close, which Traefik honors. No pool setting overrides this; the backend forces one connection per request.
  • maxIdleConnsPerHost: -1 set deliberately (for example, to work around dead-connection 502s) without accounting for the churn cost.

The symptom sequence is distinctive: backends pass health checks, but proxied requests start failing with 502s, and the kernel logs report “cannot assign requested address”. For the full diagnostic path on that symptom, see Traefik 502 Bad Gateway.

Dead pooled connections and sporadic 502s

NAT gateways and stateful firewalls between Traefik and its backends track connection state and silently expire idle entries. The backend itself may also close keep-alive connections on its own idle timer. Traefik’s pool still holds the socket, considers it reusable, writes a request into it, and gets a reset. The result is a low-rate, seemingly random 502 pattern that no health check catches.

The governing rule: Traefik’s idleConnTimeout must be shorter than every idle timeout in the path. That includes the backend’s own keep-alive timeout and any intermediary NAT or firewall idle timeout. If the backend closes connections at 60s and Traefik keeps them for 90s, every connection aged 60-90s in the pool is a 502 waiting to happen. If an intermediary expires state at 75s, same story. When you cannot learn or change the intermediary’s timeout, shortening idleConnTimeout below it is the fix. A retry middleware can mask the residual race, but timeout ordering is the actual repair.

The same mismatch exists on the client side: if a cloud load balancer in front of Traefik has a longer idle timeout than Traefik’s respondingTimeouts.idleTimeout, the LB reuses connections Traefik already closed, producing the same class of intermittent errors. Order the whole chain, not just one link.

File descriptors

Every pooled idle connection holds a file descriptor on the Traefik process, in addition to the two FDs per active proxied request. Raising maxIdleConnsPerHost raises steady-state FD usage. That is usually the right trade, but it interacts with the process FD limit: if you are running with the container default of 1024, aggressive pooling and moderate concurrency will collide. Watch process_open_fds / process_max_fds when changing pool sizes. See Traefik file descriptor monitoring for the headroom math.

Checking pool behavior from the OS

Traefik does not expose connection pool metrics directly, so pool state has to be observed at the socket layer. These checks are read-only and safe to run on a live host.

# Connection states for the Traefik process
ss -tnp | grep traefik | awk '{print $1}' | sort | uniq -c

# TIME_WAIT sockets toward a specific backend port
ss -tn state time-wait | grep :<backend_port> | wc -l

# System-wide socket summary
ss -s

# Connection states from /proc (01=ESTABLISHED, 06=TIME_WAIT, 08=CLOSE_WAIT, 0A=LISTEN)
cat /proc/$(pgrep traefik)/net/tcp | awk 'NR>1 {print $4}' | sort | uniq -c

What you are looking for:

  • High TIME_WAIT with high request rate: connections are churning instead of being reused. Check maxIdleConnsPerHost and whether backends send Connection: close.
  • Growing CLOSE_WAIT: the backend closed the connection and Traefik has not. A growing CLOSE_WAIT count is a leak indicator.
  • TIME_WAIT approaching the ephemeral port range: exhaustion is near. The short-term mitigations are widening net.ipv4.ip_local_port_range and enabling net.ipv4.tcp_tw_reuse (which only applies to outbound connections, so it does help the proxy-to-backend direction). Check the effective value with sysctl net.ipv4.tcp_tw_reuse. Both are kernel-level changes: review them against your environment’s networking constraints before applying, and treat them as buying time while you fix the pool.

Tuning guidance

  • Set maxIdleConnsPerHost deliberately. For a service with meaningful concurrency, 2 is wrong. Size it to roughly the steady-state concurrency per backend so connections survive between bursts instead of being torn down. The cost is FDs and backend-side connection slots, not meaningful memory or CPU.
  • Order the idle timeouts. idleConnTimeout on Traefik must be strictly lower than the backend’s keep-alive timeout and any NAT/firewall idle timeout in between. If you cannot verify the intermediary, shorten Traefik’s side.
  • Fix backends that send Connection: close. That header defeats pooling entirely. If it comes from a legacy application or an intermediary, that is the thing to change; no Traefik setting compensates for it.
  • Avoid -1 as a first resort. Disabling reuse silences dead-connection 502s by removing reuse, but it converts them into churn, latency, and port pressure. Prefer timeout ordering plus retries.
  • Do not forget cold start. Traefik has no pool warmup. After a restart, every backend connection is new, so the first wave of traffic pays full dial and TLS cost and can hammer fragile backends. Plan backend capacity for post-restart connection storms.

Signals to watch in production

SignalWhy it mattersWarning sign
Reuse ratio (reused vs total connections)The single best efficiency measure of the pool; requires socket-level observation since Traefik exposes no pool metricFalling reuse at stable request rate
TIME_WAIT count (OS level, ss -tn state time-wait)Direct measure of connection churn; predicts ephemeral port exhaustionSustained growth, or exceeding ~50% of the port range
CLOSE_WAIT count (OS level)Connections the backend closed that Traefik has not releasedMonotonic growth
process_open_fds / process_max_fdsIdle pooled connections consume FDs; pool tuning interacts with the FD limitRatio above 80%
traefik_service_requests_total{code="502"}Sporadic low-rate 502s with healthy backends are the classic dead-pooled-connection signatureIntermittent 502s with traefik_service_server_up all green
traefik_service_request_duration_secondsPer-request dial/TLS cost inflates latency when reuse failsLatency floor rises after a restart or config change and never drops
traefik_service_retries_totalRetries triggered by dead pooled connections amplify backend loadRetry rate rising alongside sporadic 502s

How Netdata helps

  • Per-second socket and TCP-state visibility at the host level, so TIME_WAIT and CLOSE_WAIT growth is visible as a trend, not discovered at port exhaustion.
  • Correlation of Traefik’s service-level 5xx breakdown (traefik_service_requests_total by code) with socket churn, which separates “backend is failing” from “pool is failing”: healthy backends plus rising TIME_WAIT plus sporadic 502s points at reuse, not the application.
  • process_open_fds against process_max_fds tracking, so raising maxIdleConnsPerHost does not quietly trade a churn problem for an FD cliff.
  • Service latency histograms (traefik_service_request_duration_seconds) per second, making the latency floor shift from lost reuse measurable against the pre-change baseline.
  • Retry rate (traefik_service_retries_total) alongside error rate, surfacing the amplification that dead pooled connections cause when a retry middleware is in the chain.