The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / proxysql / proxysql-query-rule-cpu-overload ▌

Operations Guides

ProxySQL query rule CPU overload: expensive regex saturating the worker threads

ProxySQL host CPU is pegged. Client query latency is climbing. But the MySQL backends are idle: ConnFree is greater than zero across the pool, backend ping latency is normal, and there are no slow queries on the database side. The proxy itself is the bottleneck.

When the proxy’s worker threads are saturated, every query slows down uniformly regardless of which backend it targets or how complex the SQL is. The symptom looks like a backend problem from the application’s perspective, but the databases are fine.

The root cause is almost always an expensive regular expression in match_pattern or replace_pattern on a high-traffic mysql_query_rules entry. Every query passes through the rule chain sequentially. A regex that takes 100 microseconds to evaluate costs 100 microseconds on every query, every time. At thousands of queries per second across four worker threads (the mysql-threads default), the CPU budget is consumed by pattern matching, leaving no cycles for query routing and result forwarding.

Query Processor mechanics

ProxySQL’s Query Processor parses every incoming SQL statement and walks it through the ordered mysql_query_rules table. Each rule can match on match_pattern (regex against the full query text), match_digest (regex against the normalized digest), schemaname, username, client_addr, or flagIN/flagOUT chains. Rules are evaluated in rule_id order. The first matching rule with apply=1 terminates evaluation. Rules without apply=1 allow evaluation to continue.

The regex engine is configurable via mysql-query_processor_regex. The default is PCRE (value 1). RE2 (value 2) is also available and guarantees linear-time matching, meaning it cannot exhibit catastrophic backtracking. However, RE2 cannot apply both CASELESS and GLOBAL modifiers simultaneously via re_modifiers, which limits some use cases.

Worker threads (mysql-threads, default 4, maximum 255) handle all client connections and query processing via non-blocking event loops. Each thread manages many connections, but if a thread blocks on CPU-intensive regex evaluation, every connection on that thread stalls. Because mysql-threads is a startup parameter, changing it requires LOAD MYSQL VARIABLES TO RUNTIME, SAVE MYSQL VARIABLES TO DISK, and a full ProxySQL restart.

The aggregate time spent inside the Query Processor is tracked as Query_Processor_time_nsec in stats_mysql_global. This counter is only populated when mysql-stats_time_query_processor is set to true. It is disabled by defaultbecause the timing measurement itself adds approximately 0.3 microseconds of latency per query. If you have never enabled it, Query_Processor_time_nsec will be zero regardless of actual processing cost.

Per-rule timing is not available in upstream ProxySQL. You must infer the culprit by correlating rule ordering, hit counts, and pattern complexity.

The diagnostic triad:

  • ProxySQL process CPU is saturated (all worker threads near 100%)
  • Backend pool has free connections (ConnFree > 0, backends not saturated)
  • Client latency is uniformly elevated (not specific to one hostgroup or digest)
flowchart TD
    A[High client latency] --> B{ConnFree > 0?}
    B -->|Yes, backends idle| C{ProxySQL CPU near 100%?}
    B -->|No, pool empty| D[Backend pool starvation]
    C -->|Yes| E{Query_Processor_time_nsec elevated?}
    C -->|No| F[Check TLS, connection churn, or network]
    E -->|Yes| G[Query rule CPU overload]
    E -->|Not tracked| H[Enable mysql-stats_time_query_processor]

Common causes

CauseWhat it looks likeFirst thing to check
Catastrophic backtracking in PCRE regexCPU spikes after a query pattern change; latency is proportional to query text lengthReview match_pattern for nested quantifiers like (a+)+ or (a|a)*
Broad match_pattern on high-traffic queriesAll queries slow down; hits on the suspect rule are very highCheck stats_mysql_query_rules.hits against rule ordering
Long rule chain without early terminationMany rules with apply=0; every query walks most of the chainCount active rules and check which have apply=1
replace_pattern with complex substitutionCPU spike correlates with a rule that rewrites query textReview rules where replace_pattern is non-null
New rule deployed via config changeLatency spike correlates with a recent LOAD MYSQL QUERY RULES TO RUNTIMECheck deployment history and rule table version

Quick checks

These are read-only queries against the admin interface (default port 6032). They do not modify configuration.

# Check ProxySQL process CPU
ps -p $(pidof proxysql) -o pid,pcpu,pmem,rss
# Per-thread CPU to detect worker thread saturation
ps -L -p $(pidof proxysql) -o pid,lwp,pcpu
# Check worker thread count and key counters
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global \
      WHERE Variable_Name IN ('MySQL_Thread_Workers','Questions','Query_Processor_time_nsec','Active_Transactions');"
# Verify backends are idle (ConnFree > 0 confirms not a pool issue)
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT hostgroup, srv_host, srv_port, status, ConnUsed, ConnFree, Latency_us FROM stats_mysql_connection_pool;"
# Check which rules are matching and how often (hits reset on rule reload)
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT rule_id, hits FROM stats_mysql_query_rules ORDER BY rule_id;"
# Review active regex engine and key variables
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT variable_name, variable_value FROM global_variables \
      WHERE variable_name IN ('mysql-query_processor_regex','mysql-threads','mysql-stats_time_query_processor');"
# Review all runtime query rules for pattern complexity
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SELECT rule_id, active, match_pattern, match_digest, replace_pattern, destination_hostgroup, apply \
      FROM runtime_mysql_query_rules ORDER BY rule_id;"

How to diagnose it

  1. Confirm the pattern. Verify that ProxySQL process CPU is near 100% across all worker threads (not just one), backend connections are free (ConnFree > 0), and client latency is elevated. This distinguishes query rule CPU overload from backend connection pool exhaustion, backend query slowness, or network latency.

  2. Check whether query processor timing is tracked. If Query_Processor_time_nsec is zero and you have never set mysql-stats_time_query_processor=true, the counter is simply not populated. Enable it to get a baseline. Note the 0.3 microsecond per-query overhead.

  3. List all active rules and identify hot paths. Pull stats_mysql_query_rules.hits to see which rules are evaluated most frequently. A rule with millions of hits is on the hot path. Any regex cost on that rule is multiplied by the hit count.

  4. Review regex patterns for complexity. Look for long patterns, nested quantifiers, alternation with overlap, and patterns that match against match_pattern (full query text) rather than match_digest (normalized digest text, which is shorter and faster to match).

  5. Binary search by disabling rules in groups. Since per-rule timing is not available, disable the bottom half of the rule chain (active=0), load to runtime, and observe whether CPU drops. Repeat until you isolate the specific rule. Warning: disabling rules changes routing behavior. Test off-hours or on a staging instance first.

  6. Test the suspect regex in isolation. Extract the match_pattern from the suspect rule and run it against representative query strings using a PCRE test harness. Measure evaluation time. A pattern that takes more than a few microseconds on a typical query string is the problem.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ProxySQL process CPU (per-thread)Worker thread CPU is the hard ceiling on throughputAll worker threads near 100% sustained
Query_Processor_time_nsecAggregate time spent in the Query ProcessorRate increasing disproportionate to Questions rate
ConnFree per backendDistinguishes proxy-side CPU from backend pool exhaustionConnFree > 0 while latency is high = proxy bottleneck
stats_mysql_query_rules.hitsShows which rules are on the hot pathHigh-hit rule with complex match_pattern
MySQL_Thread_WorkersConfirms the hard parallelism ceilingCPU saturated with default 4 threads
Client query latency (from histogram)Uniform latency increase signals proxy-side delaystats_mysql_commands_counters histogram shifts toward higher buckets across all command types

Fixes

Emergency: disable the suspect rule

If the proxy is causing visible client impact, disable the rule immediately and load to runtime.

# Disable the suspect rule and activate immediately
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "UPDATE mysql_query_rules SET active=0 WHERE rule_id=<rule_id>; LOAD MYSQL QUERY RULES TO RUNTIME;"

This is safe and reversible. If CPU drops immediately, you have confirmed the rule was the cause. Save to disk only after confirming the fix.

Switch regex engine from PCRE to RE2

RE2 guarantees linear-time matching and cannot exhibit catastrophic backtracking. If your re_modifiers do not require both CASELESS and GLOBAL simultaneously, switching is low-risk.

# Switch to RE2 regex engine
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SET mysql-query_processor_regex=2; LOAD MYSQL VARIABLES TO RUNTIME;"

Test in a staging environment first. RE2 does not support backreferences or lookahead. If any existing rule uses unsupported syntax, that rule will stop matching after the switch. Check ProxySQL logs for regex compilation errors.

Replace match_pattern with match_digest

match_digest matches against the normalized query digest, where literal values are replaced with ?. For routing rules that do not need to inspect literal values, this is significantly faster because the digest text is shorter and more predictable than the full query.

# Example: rewrite a rule to use match_digest instead of match_pattern
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "UPDATE mysql_query_rules SET match_pattern=NULL, match_digest='^SELECT.*FROM users' WHERE rule_id=<rule_id>; LOAD MYSQL QUERY RULES TO RUNTIME;"

Rules that need to match specific literal values cannot use match_digest, since WHERE id = 42 normalizes to WHERE id = ?.

Restructure the rule chain with flagIN/flagOUT

Every query walks the chain from rule_id 1 until it hits a match with apply=1. Use flagIN/flagOUT chaining to create rule pipelines: a query matched by a rule with flagOUT=N only evaluates rules with flagIN=N next. This skips irrelevant rules entirely.

For static routing by username and schemaname (no regex needed), consider mysql_query_rules_fast_routing. This table provides hash-based O(1) routing without any regex evaluation.

Increase mysql-threads (requires restart)

If the rule chain is already optimized and the proxy is still CPU-bound, increasing mysql-threads raises the parallelism ceiling.

# Increase thread count (requires restart to take effect)
mysql -u admin -padmin -h 127.0.0.1 -P 6032 \
  -e "SET mysql-threads=8; LOAD MYSQL VARIABLES TO RUNTIME; SAVE MYSQL VARIABLES TO DISK;"
# Then restart ProxySQL

mysql-threads above 16 can degrade throughput on high-core-count systems due to context switching overhead. Test before committing to high thread counts in production. This is a capacity fix, not a rule-complexity fix: if the regex itself is pathological, more threads only delays the saturation point.

Prevention

  • Review every new query rule for regex complexity before loading to runtime. Test the pattern against representative query strings and measure evaluation time.
  • Prefer match_digest over match_pattern unless the rule specifically needs to inspect literal values in the query text.
  • Keep the rule chain short. Every query walks from rule_id 1 until it hits apply=1. Use apply=1 aggressively to terminate evaluation early.
  • Enable mysql-stats_time_query_processor=true permanently on production instances. The 0.3 microsecond per-query overhead is negligible compared to the cost of flying blind during a CPU overload incident.
  • Consider RE2 as the default regex engine (mysql-query_processor_regex=2) if no rules require PCRE-specific features.
  • Audit rule changes as configuration events. Track LOAD MYSQL QUERY RULES TO RUNTIME in your change management system. A rule addition is the most common trigger for this failure pattern.

How Netdata helps

  • Per-second CPU metrics on the ProxySQL host reveal thread saturation as it develops, not minutes after. Per-core breakdown shows whether all worker threads are pegged or only one is hot-spotting.
  • Query_Processor_time_nsec trend correlates directly with rule complexity. A step change after a rule deployment is the clearest signal that a new regex is expensive.
  • ConnFree per backend distinguishes proxy-side CPU saturation from backend pool exhaustion in a single dashboard. Free backends combined with high proxy CPU points to this pattern.
  • stats_mysql_query_rules.hits collected over time shows which rules are on the hot path, so you can prioritize optimization toward high-traffic rules rather than guessing.