The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / proxysql / proxysql-server-connections-delayed ▌

Operations Guides

ProxySQL Server_Connections_delayed above zero: queries waiting on the backend pool

When Server_Connections_delayed in ProxySQL’s stats_mysql_global table climbs above zero, queries are waiting for a backend connection that was not immediately available. This is not an error counter. It is a pressure signal: ProxySQL’s backend connection pool could not instantly satisfy a connection request, so the requesting session had to wait. A brief blip during a traffic burst is normal. A sustained increase means the pool is undersized, multiplexing has degraded, or backends are disappearing from rotation faster than the pool can adapt.

The counter is cumulative since ProxySQL start. To assess current pressure, you need the rate of change, not the absolute value. A value of 10,000 accumulated over weeks is noise. A delta of 500 in five minutes is the signal you need to investigate.

What this means

ProxySQL’s backend connection pool is the multiplexing core. When a client sends a query, ProxySQL borrows a backend connection, routes the query, returns the result, then returns the connection to the pool. When no free connection is available and the pool cannot create a new one, the query must wait. That wait increments Server_Connections_delayed.

A critical nuance: ConnFree == 0 alone does not mean saturation. ProxySQL may still create new connections up to the backend’s max_connections limit (configured in mysql_servers). True saturation means ProxySQL cannot create new connections and has no free ones. The symptom is ConnERR increasing (ProxySQL tried to open a new connection and was rejected) or ConnPool_get_conn_failure climbing.

The companion metric to watch alongside this is Server_Connections_aborted. When backends close connections unexpectedly, often because the backend MySQL’s wait_timeout fires on an idle pooled connection that ProxySQL still thinks is valid, the next query on that connection hits an error. ProxySQL handles this by reconnecting, but the aborted connection reduces effective pool size and can contribute to delayed connections in the next burst.

flowchart TD
    A[Query arrives] --> B{Free backend connection?}
    B -- Yes --> C[Borrow, execute, return to pool]
    B -- No --> D{Can create new connection?}
    D -- Yes --> E[Create, execute, return to pool]
    D -- No, at max_connections --> F[Query waits]
    F --> G[Server_Connections_delayed++]
    G --> H{Connection frees before timeout?}
    H -- Yes --> C
    H -- No --> I[max_connect_timeouts++]
    I --> J[Error 9001 to client]

Common causes

CauseWhat it looks likeFirst thing to check
Pool at max_connectionsConnFree is zero, ConnUsed at limit, ConnERR stableCompare ConnUsed to max_connections in runtime_mysql_servers
Multiplexing collapsehostgroup_locked high relative to connected, Active_Transactions elevatedCheck Client_Connections_hostgroup_locked ratio
Backend SHUNNEDOne or more backends SHUNNED in stats_mysql_connection_pool, ConnERR rising on that backendCheck monitor check results and backend status
Backend dropping pooled connectionsServer_Connections_aborted rising, no corresponding backend issueCompare backend wait_timeout to ProxySQL mysql-wait_timeout
Prepared statements pinning connectionsStmt_Server_Active_Total high, multiplexing ratio degradedCheck application prepared statement usage

Quick checks

All commands connect to the ProxySQL admin interface on port 6032 and run read-only queries. Replace -u admin -padmin with your configured admin credentials.

# Check the delayed and aborted counters
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global \
      WHERE Variable_Name IN ('Server_Connections_delayed','Server_Connections_aborted', \
      'Server_Connections_connected','Server_Connections_created');"
# Check per-backend pool state
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT hostgroup, srv_host, srv_port, status, ConnUsed, ConnFree, ConnOK, ConnERR \
      FROM stats_mysql_connection_pool ORDER BY hostgroup, srv_host;"
# Cross-reference pool usage with configured limits
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT cp.hostgroup, cp.srv_host, cp.srv_port, cp.ConnUsed, cp.ConnFree, rs.max_connections \
      FROM stats_mysql_connection_pool cp \
      JOIN runtime_mysql_servers rs \
      ON cp.hostgroup = rs.hostgroup_id AND cp.srv_host = rs.hostname AND cp.srv_port = rs.port;"
# Check pool operation outcomes
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global \
      WHERE Variable_Name IN ('ConnPool_get_conn_success','ConnPool_get_conn_failure', \
      'ConnPool_get_conn_immediate');"
# Check multiplexing health
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global \
      WHERE Variable_Name IN ('Client_Connections_connected','Client_Connections_hostgroup_locked', \
      'Active_Transactions');"
# Check backend status across hostgroups
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT hostgroup_id, hostname, port, status, max_connections FROM runtime_mysql_servers;"
# Check for connect timeout errors
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global \
      WHERE Variable_Name = 'max_connect_timeouts';"

How to diagnose it

  1. Establish whether the counter is actively increasing. Take two readings of Server_Connections_delayed a few seconds apart. If the value is not changing, the pressure was transient and may have resolved. If it is climbing, proceed.

  2. Check per-backend pool saturation. Join stats_mysql_connection_pool with runtime_mysql_servers to compare ConnUsed against max_connections. If ConnUsed equals or approaches max_connections and ConnFree is zero, the pool for that backend is saturated.

  3. Check ConnPool_get_conn_failure. This metric is the most direct indicator of pool starvation. It counts how often ProxySQL tried to get a backend connection from the pool and failed at that instant. A rising rate confirms the pool cannot satisfy demand. Note that ConnPool_get_conn_failure does not directly translate to client-facing errors. The MySQL thread retries according to mysql-connect_retries_on_failure (default 10). A high failure count can accumulate in seconds without client impact if retries eventually succeed.

  4. Check backend status. Look for SHUNNED backends. When a backend is SHUNNED, its connections are not available for new queries, effectively reducing pool capacity for that hostgroup. Even if ConnFree shows free connections at other times, during a shunning window the pool cannot provide connections for that hostgroup.

  5. Check multiplexing health. If Client_Connections_hostgroup_locked is approaching Client_Connections_connected, multiplexing has collapsed. Each client session pins a dedicated backend connection, making the effective pool size equal to the client count rather than the configured max_connections. Check Active_Transactions to see if long-running transactions are the cause.

  6. Check Server_Connections_aborted. If this is rising alongside Server_Connections_delayed, backends are closing connections out from under ProxySQL. The most common cause is the backend MySQL’s wait_timeout firing on idle pooled connections. ProxySQL keeps the connection in its pool thinking it is valid, but the next query that tries to use it gets an error.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Server_Connections_delayed (rate)Direct evidence of pool pressureSustained rate above zero
ConnPool_get_conn_failure (rate)ProxySQL tried and failed to get a connectionRising rate correlates with delayed
ConnFree per backendIdle connections available for reuseZero with ConnUsed at max_connections
ConnUsed / max_connectionsPool utilization per backendAbove 80% sustained
Server_Connections_aborted (rate)Backends closing connections unexpectedlyRising rate, especially with delayed
Client_Connections_hostgroup_locked ratioMultiplexing degradationAbove 50% of connected
Active_TransactionsConnections pinned by transactionsHigh relative to client count
Backend statusSHUNNED reduces effective poolAny backend not ONLINE
max_connect_timeoutsQueries waiting beyond timeout limitAny sustained increase

Fixes

Pool at max_connections

If ConnUsed is at max_connections and backends are otherwise healthy, the cap is too low for the workload. Increase max_connections in mysql_servers for the affected backend.

UPDATE mysql_servers SET max_connections = <new_value>
  WHERE hostgroup_id = <hg> AND hostname = '<host>';
LOAD MYSQL SERVERS TO RUNTIME;
SAVE MYSQL SERVERS TO DISK;

The sum of max_connections across all ProxySQL instances for a given backend must not exceed what the backend MySQL can handle. ProxySQL shares the backend’s connection capacity with direct admin connections, replication threads, monitoring tools, and any non-proxy applications. Leave headroom.

Also check whether the backend MySQL itself is rejecting connections. If ConnERR is rising and Server_Connections_aborted is climbing, the backend may have hit its own max_connections limit.

Multiplexing collapse

If hostgroup_locked is high relative to connected, the application is disabling multiplexing. Common causes: ORMs setting session variables (SET NAMES, SET sql_mode, SET time_zone) on every connection, long-running transactions, GET_LOCK(), temporary tables, user-defined variables, or prepared statements in some configurations.

Short-term: increase max_connections on backends to absorb the 1:1 mapping. Long-term: identify which SET commands are breaking multiplexing. Move session variable initialization to server-side defaults or ProxySQL’s mysql-init_connect so they do not need per-session SET commands. Check stats_mysql_processlist for sessions with disabled multiplexing and the reason.

Backend SHUNNED reducing pool capacity

When a backend is SHUNNED, its connections are unavailable for new queries. If enough backends in a hostgroup are SHUNNED, the remaining backends absorb all traffic and may saturate their own pools. Check the monitor check results to understand why the backend was shunned. Common causes: connection errors exceeding mysql-shun_on_failures (default 5), or replication lag exceeding the configured max_replication_lag.

If shunning is transient and self-corrects within mysql-shun_recovery_time_sec (default 10 seconds), the pool pressure is temporary. If shunning is persistent or flapping, investigate the underlying backend health or adjust monitor thresholds.

Backend MySQL wait_timeout closing pooled connections

If the backend MySQL’s wait_timeout is lower than ProxySQL’s mysql-wait_timeout (default 28800000ms, approximately 8 hours), the backend may close idle connections that ProxySQL still tracks as valid. The next query on that connection fails, increments Server_Connections_aborted, and forces ProxySQL to create a new connection. During bursts, this churn contributes to pool pressure.

Prevention

  • Track ConnUsed / max_connections per backend as a trend. The leading indicators are ConnFree trending down and ConnPool_get_conn_failure appearing before Server_Connections_delayed rises. Aim for at least 20% pool capacity free during peak.
  • Monitor the multiplexing ratio. Client_Connections_connected / Server_Connections_connected should be well above 1:1. A declining ratio means the effective pool is shrinking even without traffic growth.
  • Alert on ConnPool_get_conn_failure rate, not just ConnFree. This is the most direct indicator that ProxySQL tried and failed to get a connection. It surfaces pressure before delayed connections become visible.
  • Align timeout settings. Ensure the backend MySQL’s wait_timeout is not silently closing pooled connections. Set mysql-connection_max_age_ms if needed.
  • Size max_connections with all consumers in mind. The sum of ProxySQL’s per-backend max_connections across all instances must leave room for direct connections, replication threads, and monitoring tools on the backend MySQL.
  • Watch for connection churn after deployments. Application restarts create connection storms. Stagger restarts or configure connection warming if available in your version.

How Netdata helps

Netdata’s ProxySQL collector surfaces these signals at per-second resolution, which matters because pool pressure can be transient and invisible to minute-level polling:

  • Server_Connections_delayed as a rate, not just a cumulative counter, so you can see when pressure is actively building versus historical accumulation.
  • ConnPool_get_conn_failure correlated with per-backend ConnUsed and ConnFree, letting you see which specific backend is saturating before the delayed counter confirms it.
  • Server_Connections_aborted trended alongside delayed connections, surfacing the backend wait_timeout interaction without a separate investigation.
  • Multiplexing ratio (hostgroup_locked / connected) trended over time, so you can catch the gradual multiplexing collapse that turns into pool starvation weeks later.
  • ML anomaly detection on pool operation rates, catching the early deviation from baseline that precedes a visible saturation event.