The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / proxysql / proxysql-how-it-works-in-production ▌

Operations Guides

How ProxySQL actually works in production: a mental model for operators

When ProxySQL fails in production, the layer where it fails determines what you see: connection exhaustion at the frontend pool, mis-routing in the query processor, pool starvation from multiplexing collapse, or false health decisions in the monitor module. Understanding these layers is prerequisite to debugging any of them.

This article covers the request path from client connection to backend response, the components that make routing and pooling decisions, and the failure modes characteristic of each. The ProxySQL runbooks in this section build on the terminology and component relationships described here.

What it is and why it matters

ProxySQL fully parses the MySQL wire protocol on the frontend, understands every query, and makes per-query decisions about routing, caching, and connection management. It is not a TCP relay. It speaks the MySQL wire protocol only and does not support PostgreSQL.

This is what enables:

  • Connection multiplexing: N application connections served by M backend connections, where M is much smaller than N. The primary reason most teams deploy ProxySQL.
  • Query routing: queries sent to different hostgroups based on rule matching. Read/write splitting is the most common application.
  • Query caching: result sets cached in process memory with TTL-based expiration. No invalidation on data change; it is a time-bounded stale-read cache.
  • Backend health monitoring: automatic detection of backend failures, replication lag, and role changes via background probe threads.

Each has characteristic failure modes: multiplexing collapses when application behavior pins connections, query rules mis-route writes to read replicas, the cache churns with high-cardinality queries, and the monitor module makes false decisions with aggressive thresholds. The internal structure is what lets you tell these apart under pressure.

How it works

The diagram shows the path a query takes through ProxySQL, from client connection to result return. The monitor module operates independently, probing backends in the background.

flowchart TD
    C[Client] -->|port 6033| FP[Frontend pool]
    FP -->|auth: mysql_users| QP[Query processor + rules]
    QP -->|cache hit| QC[Query cache, TTL-based]
    QP -->|miss or no cache rule| MX[Multiplexing engine]
    QC -->|serve result| C
    MX -->|borrow connection| BP[Backend pool, per hostgroup]
    BP -->|route query| BE[MySQL backends]
    BE -->|result set| BP
    BP -->|return connection| MX
    MX -->|return result| C
    MON[Monitor module] -.->|health probes| BE

Frontend connection pool

Client connections arrive on the data-plane port (default 6033). Each client establishes a MySQL session with ProxySQL itself, not with any backend. ProxySQL authenticates the client against its own mysql_users table, independently of backend credentials. The session carries state: user, schema, transaction status, session variables, autocommit mode.

The frontend pool has a global connection limit (mysql-max_connections, default 2048). Per-user limits can also be set in mysql_users. When either limit is reached, new connections are rejected.

Query processor and mysql_query_rules

Every incoming query is parsed and normalized into a digest. The query processor evaluates the query against mysql_query_rules, an ordered list matched by rule_id in ascending order. Rules can match on regex patterns (match_digest on the normalized digest, or match_pattern on the full query text), schema, user, client address, and flagIN/flagOUT chains.

The first matching rule with apply=1 terminates evaluation and determines: target hostgroup, whether to cache (cache_ttl), whether to mirror, query timeout, delay, and query rewriting via replace_pattern. If no rule matches, the query falls to the user’s default_hostgroup.

Complex regex in match_pattern is evaluated on every matching query and is more expensive than match_digest, which operates on the shorter normalized form. A poorly designed rule chain is a latency tax on every query; under sufficient load, it can saturate worker threads.

TTL query cache

If a matching query rule sets cache_ttl > 0, ProxySQL checks its in-memory query cache before routing to a backend. The cache is keyed by query digest, user, and schema. A cache hit returns the result immediately without touching any backend.

Limitations: invalidation is TTL-based only; there is no API to purge entries when underlying data changes. Cache size is controlled by mysql-query_cache_size_MB (default 256 MB). Limits are soft, with a background purge thread; official documentation does not describe LRU eviction. The query cache is incompatible with prepared statements.

Multiplexing engine

This is the critical abstraction that makes ProxySQL valuable. When a client sends a query that is not cached, the multiplexing engine borrows a backend connection from the pool, routes the query, receives the result, returns it to the client, and returns the backend connection to the pool for reuse by another client.

Multiplexing breaks when session state exists on the connection. Conditions that disable multiplexing include:

  • Open transactions (BEGIN, START TRANSACTION, until COMMIT/ROLLBACK)
  • SET commands that change session variables
  • Temporary tables
  • LOCK TABLES
  • User-defined variables (@var)
  • GET_LOCK()
  • Prepared statements

When multiplexing is disabled for a session, a backend connection is pinned to that client for the duration. This is visible as Client_Connections_hostgroup_locked in stats_mysql_global. If many sessions are pinned, the multiplexing ratio degrades toward 1:1, eliminating the pooling benefit and potentially exhausting backend connections.

Backend connection pool, per hostgroup

Backend connections are organized by hostgroup. Each hostgroup is a set of MySQL backends that serve the same class of queries. In a read/write split deployment, a writer hostgroup contains the primary and a reader hostgroup contains replicas.

Each backend entry in mysql_servers has a max_connections limit: ProxySQL’s self-imposed ceiling on connections to that backend. This is independent of MySQL’s own max_connections setting. The two interact: the sum of ProxySQL’s max_connections across all ProxySQL instances, plus connections from other consumers (direct admin connections, monitoring tools, replication threads), must not exceed the backend MySQL’s actual max_connections.

Backend server status is tracked per hostgroup:

StatusBehavior
ONLINEReceiving traffic
SHUNNEDTemporarily avoided due to connection errors or replication lag; self-recovering
OFFLINE_SOFTDraining; no new connections, existing ones kept
OFFLINE_HARDRemoved; existing connections killed immediately

Monitor module

The monitor module runs on its own background threads and continuously probes backends:

  • Connect checks: can ProxySQL establish a TCP connection to the backend?
  • Ping checks: is the backend responsive to mysql_ping()?
  • Read-only checks: is @@read_only set? Used for automatic read/write role detection.
  • Replication lag checks: how far behind is this replica? When lag exceeds max_replication_lag, the backend is shunned.
  • Group Replication and Galera checks: node state for InnoDB Cluster, Group Replication, and Galera/PXC topologies.

Monitor results drive backend status transitions. The monitor uses its own credentials (mysql-monitor_username, mysql-monitor_password), separate from application credentials. If monitor credentials expire, health checks fail and ProxySQL shuns backends it cannot verify.

Admin interface and three-layer config

ProxySQL exposes a MySQL-protocol admin interface (default port 6032) for configuration and stats retrieval. This is the control plane, separate from the data plane.

Configuration exists in three layers:

  • MEMORY: staging area. Changes made via the admin interface land here first.
  • RUNTIME: active configuration. LOAD ... TO RUNTIME activates MEMORY changes.
  • DISK: persistent SQLite storage. SAVE ... TO DISK persists changes across restarts.

A change is not live until loaded to RUNTIME. A change is not durable until saved to DISK. On restart, ProxySQL loads from DISK. Config drift between layers is a common operational hazard: an operator changes config, applies it to RUNTIME, forgets to SAVE to DISK, and the next restart reverts the fix.

Thread pool

Worker threads (mysql-threads, default 4) handle all client connections and query processing using non-blocking I/O (epoll). Each thread manages many connections via an event loop. The thread count is a hard ceiling on parallelism, set at startup and cannot be changed at runtime without restart.

If worker threads saturate (complex regex evaluation, high connection count, TLS overhead), queries queue in the event loop. Latency increases uniformly across all queries regardless of backend. Adding more backends does not help when the proxy itself is the bottleneck.

Where it shows up in production

Deployment patterns and their operational concerns:

  • Standalone: single instance. Simple but a single point of failure unless a VIP is managed externally.
  • ProxySQL Cluster: multiple instances synchronizing configuration via proxysql_servers. Cluster sync is tracked in stats_proxysql_servers_checksums. Introduces split-brain risk during network partitions.
  • Sidecar: ProxySQL runs on the same host as the application. Reduces network latency but creates resource contention (CPU, memory, file descriptors).
  • Galera or Group Replication-aware: special hostgroup tables (mysql_galera_hostgroups, mysql_group_replication_hostgroups) automate writer/reader role routing based on cluster state.

Tradeoffs and when to use it

ProxySQL adds a stateful component to your database path. The benefits come with costs.

Added latency: every query passes through parsing, rule matching, and connection borrowing. ProxySQL’s overhead is typically under 1 ms per query, but complex regex rules or large result sets increase this. Under worker thread saturation, latency is added uniformly to every query.

Configuration complexity: the three-layer config model, the query rule chain, and the multiplexing engine all introduce state that must be understood. Config drift, rule mis-routing, and multiplexing collapse are failure modes that do not exist with direct connections.

Multiplexing assumptions: capacity models for ProxySQL typically assume a healthy multiplexing ratio (10:1 or better). If application behavior silently disables multiplexing (ORMs setting session variables, prepared statements, long transactions), the proxy degrades to a 1:1 mapping with overhead. The capacity plan becomes fiction.

Connection storms after restart: the backend connection pool starts empty. The first burst of queries creates backend connections synchronously, which can overwhelm backend MySQL’s connection handling.

Signals to watch in production

SignalWhy it mattersWarning sign
Backend status per hostgroup (ONLINE, SHUNNED, OFFLINE_SOFT, OFFLINE_HARD)Determines whether queries can reach backends at allZero ONLINE backends in any hostgroup with active traffic
Client_Connections_hostgroup_locked / Client_Connections_connectedMeasures multiplexing effectivenessRatio above 0.5 sustained means multiplexing is collapsing
ConnPool_get_conn_failureDirect indicator that queries cannot get backend connectionsAny sustained increase means pool starvation
Query_Cache_count_GET_OK / Query_Cache_count_GETCache hit ratio; declining ratio means more backend loadHit rate dropping while Questions rate is stable
MySQL_Monitor_WorkersWhether health check threads are runningZero sustained means monitoring is down, status decisions are stale
ProxySQL_UptimeContext for cold-start conditions; gates alerts that should not fire during warmupRecent restart means pools are cold and stats tables are reset
Query_Processor_time_nsecCPU time in the query processorElevated with idle backends indicates regex rules are the bottleneck

The preceding names are the documented stats_mysql_global variables.

How Netdata helps

  • Per-second collection of stats_mysql_global counters reveals multiplexing degradation (locked connections rising relative to Client_Connections_connected) before it becomes a backend connection outage. Polling at minute granularity misses the transition.
  • Correlating backend status transitions with monitor check failure rates (connect, ping, read-only, replication lag) distinguishes real backend failures from monitor false positives. A backend flapping between ONLINE and SHUNNED with healthy direct connections points to monitor threshold tuning, not a backend problem.
  • Per-backend pool metrics (ConnUsed, ConnFree, ConnERR from stats_mysql_connection_pool) show pool pressure building on specific backends before ConnPool_get_conn_failure starts climbing.
  • Memory subsystem metrics (jemalloc_resident, jemalloc_allocated, query_digest_memory in stats_memory_metrics) catch slow memory growth that leads to OOM kills, and distinguish useful allocation from fragmentation.
  • Process-level signals (per-core CPU, file descriptor counts) catch thread saturation and FD exhaustion that ProxySQL does not expose in its own stats tables. FD exhaustion is a binary cliff: every connection (client plus backend plus monitor) consumes one FD, and exhaustion causes immediate total failure.