The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / proxysql / proxysql-query-rules-order-apply ▌

Operations Guides

ProxySQL query rule order and apply=1: how rule chains silently mis-route

ProxySQL query rules look deceptively simple: write a regex, point it at a hostgroup, done. But mysql_query_rules is an ordered chain, and the chain’s behavior depends on three interacting variables that most operators never think about together: rule_id order, the apply flag, and flagIN/flagOUT chaining. Get any of these wrong and queries route to the wrong backend silently, with no error.

The mis-route may not manifest until weeks after the configuration change. A rule chain can work correctly by accident, because a later rule happens to set the same hostgroup the operator intended. Then someone adds a new rule, the last match changes, and writes start landing on read-only replicas.

What it is and why it matters

ProxySQL routes every query through the Query Processor, which evaluates the incoming SQL against the mysql_query_rules table. Each rule has:

  • A rule_id that determines evaluation order (ascending)
  • Match criteria (match_digest, match_pattern, schemaname, username, client_addr, and others)
  • A destination_hostgroup to route matching queries to
  • An apply flag that controls whether evaluation stops after a match
  • Optional flagIN/flagOUT values for rule chaining

The critical behavior: rules are evaluated in rule_id order, and the first match with apply=1 terminates evaluation. If apply=0 (the default when not explicitly set), the rule’s destination_hostgroup and other settings are applied, but evaluation continues to subsequent rules. Any later matching rule overrides the earlier rule’s settings.

A rule without apply=1 is not a terminal rule. It is a tentative rule whose routing decision can be overridden by anything later in the chain.

How it works

The evaluation algorithm:

  1. ProxySQL receives a query from the client session.
  2. Starting from the lowest rule_id where active=1 and flagIN=0, the query is tested against each rule’s match criteria.
  3. If a rule matches, its settings (destination_hostgroup, cache_ttl, timeout, etc.) are applied to the query.
  4. If apply=1, evaluation stops. The query uses this rule’s settings.
  5. If apply=0, evaluation continues. Any subsequent matching rule overrides the previous settings.
  6. If no rule matches, the query falls through to the user’s default_hostgroup from mysql_users.
flowchart TD
    A[Incoming query] --> B{Match rule at
current rule_id?} B -- No --> C{More active rules
in chain?} C -- Yes --> B C -- No --> D[Route to
default_hostgroup] B -- Yes --> E[Apply destination_hostgroup,
cache_ttl, timeout, etc.] E --> F{apply = 1?} F -- Yes --> G[Route to
destination_hostgroup] F -- No --> H{flagOUT set?} H -- Yes --> I[Jump to rules with
matching flagIN] I --> B H -- No --> C

apply=0 is not an error, but it is usually wrong

When a rule matches with apply=0, the rule’s destination_hostgroup is recorded, but the query continues through the chain. If a later rule also matches (with a broader regex, for example), that later rule’s destination_hostgroup wins. The first rule effectively did nothing.

This is the classic silent mis-route. It happens most often when:

  • A read-routing rule is added with a broad pattern like ^SELECT and apply=0
  • A later, more general rule sends everything to the writer hostgroup
  • The read rule “works” in testing because the test query also matches the writer rule and the writer happens to accept reads

flagIN/flagOUT chaining

flagIN and flagOUT create rule pipelines. By default, flagIN=0 for all queries entering the chain. When a rule with flagOUT=N matches, evaluation jumps to rules with flagIN=N. This allows multi-stage routing: first classify the query type, then apply hostgroup-specific rules.

The mysql-query_processor_iterations variable (default 0) controls whether the query processor can loop back to the beginning of the rule set. If set greater than 0, a matching rule can restart processing from rule_id 1, up to the iteration limit. This is rarely used and makes the chain harder to reason about.

With the default mysql-query_processor_iterations=0, flagOUT selects later matching flagIN rules and the chain progresses in rule_id order. Setting mysql-query_processor_iterations above 0 explicitly allows a matching rule to restart from the beginning of the rule set for up to that many iterations.

Invalid regex fails silently

When LOAD MYSQL QUERY RULES TO RUNTIME compiles regex patterns, a malformed pattern does not produce an error. The rule is loaded, marked active, but never matches anything. The only evidence is stats_mysql_query_rules.hits staying at zero for that rule.

A typo in a critical write-routing rule can pass a config reload with no warning. Queries that should match the rule silently fall through to the default hostgroup, which may be a reader.

stats_mysql_query_rules resets on every reload

The hits counter in stats_mysql_query_rules resets to zero every time query rules are loaded to runtime. This means:

  • You cannot accumulate hit counts across rule changes
  • If you LOAD rules to fix a mis-route, you lose the baseline that would confirm the fix
  • Monitoring systems that sample hits will show a brief zero spike after every config change, which must be distinguished from a genuinely dead rule

Where it shows up in production

The deferred mis-route

This is the most common and most dangerous pattern:

  1. Operator writes a read/write split rule set. The write rule (^INSERT, ^UPDATE, ^DELETE) routes to hostgroup 0. The read rule (^SELECT) routes to hostgroup 1.
  2. The write rule has apply=1. The read rule does not.
  3. During testing, SELECT queries land on hostgroup 1 (correct). The operator declares victory.
  4. Months later, someone adds a new rule with a higher rule_id that catches SELECT FOR UPDATE and routes it to hostgroup 0 (the writer). This is intentional.
  5. But the broad pattern also matches regular SELECT queries. Since the read rule has apply=0, those queries now override to hostgroup 0. All reads silently move to the writer.

No error is generated. ProxySQL is doing exactly what its rules say. The only symptoms are increased load on the writer and reduced load on read replicas.

Writes hitting read-only replicas

If a write-routing rule has apply=0 and a later rule routes the same query to a read-only replica, the application receives MySQL error 1290 (“The MySQL server is running with the –read-only option”). This error appears in stats_mysql_errors with the reader hostgroup as the source.

If replicas are not configured as read-only, the writes succeed on a non-authoritative backend. This causes silent data divergence with no error anywhere. The only detection is comparing data between primary and replica, or noticing that rows written via ProxySQL are missing from the primary.

transaction_persistent bypass

When a user has transaction_persistent=1 in mysql_users and a transaction is active, ProxySQL pins all subsequent queries in that transaction to the hostgroup where the transaction started, regardless of query rules. This is by design (a transaction must go to a single backend), but it can mask rule mis-configuration: queries inside transactions always route “correctly” because rules are bypassed entirely. The mis-route only surfaces for autocommit queries outside transactions.

Cluster checksum divergence

In a ProxySQL Cluster, different nodes can have different mysql_query_rules loaded if sync is lagging or a change was applied to only one node. The stats_proxysql_servers_checksums table reveals this. If nodes have different rules, the same query may route differently depending on which ProxySQL instance the client connects to. This produces intermittent symptoms that are extremely difficult to reproduce in testing, because the test client may hit a different proxy node than the one exhibiting the problem.

Common misuses

  • Missing apply=1 on terminal rules: Every rule that sets a destination_hostgroup and is meant to be the final routing decision should have apply=1. Leaving it unset (defaults to 0) creates a deferred mis-route that tests fine until a later rule change shifts the last match.
  • Broad regex early in the chain: A rule with rule_id=1 and pattern ^SELECT will match before more specific rules. If it has apply=1, it shadows everything. If it has apply=0, it sets a tentative destination that later rules override. Either way, specificity is lost.
  • Using flagIN/flagOUT without documentation: Multi-stage chains are powerful but extremely hard to debug. If you use them, document the pipeline explicitly: which flagOUT values exist, which rules have matching flagIN, and what each stage decides.
  • Assuming stats_mysql_query_rules persists: Hit counters reset on every LOAD MYSQL QUERY RULES TO RUNTIME. Do not rely on them for trend analysis across config changes. Export and persist hit counts externally if you need history.
  • Passing the string ‘NULL’ instead of SQL NULL for username: In automation (Ansible, scripts), ensure that NULL is passed as SQL NULL, not serialized as the string 'NULL'. The string 'NULL' causes ProxySQL to match a user literally named NULL rather than treating the field as a wildcard; SQL NULL remains the wildcard form in the current rule-matching implementation.

How to check your rule chain

These are read-only diagnostic queries and are safe to run against the admin interface in production:

-- Inspect rule order and apply flags as ProxySQL sees them at runtime
SELECT rule_id, active, match_digest, match_pattern,
       username, destination_hostgroup, apply,
       flagIN, flagOUT
FROM runtime_mysql_query_rules
ORDER BY rule_id;
-- Check which rules are actually being hit (resets on every LOAD TO RUNTIME)
SELECT rule_id, hits FROM stats_mysql_query_rules ORDER BY rule_id;
-- Find write statements routed to unexpected hostgroups
SELECT hostgroup, schemaname, username, digest_text, count_star
FROM stats_mysql_query_digest
WHERE digest_text LIKE 'INSERT%'
   OR digest_text LIKE 'UPDATE%'
   OR digest_text LIKE 'DELETE%'
ORDER BY count_star DESC;

The first query shows the chain as ProxySQL sees it at runtime. Note the table is runtime_mysql_query_rules, not mysql_query_rules (the staging table). The second shows which rules are actually matching traffic. A critical rule with zero hits is either dead (invalid regex) or shadowed by an earlier rule. The third catches the symptom: write digests landing in a reader hostgroup.

Signals to watch in production

SignalWhy it mattersWarning sign
stats_mysql_query_rules.hits per ruleReveals which rules match traffic and which are deadCritical rule (e.g., write routing) with zero or declining hits
stats_mysql_query_digest hostgroup per digestShows where each query type actually landsWrite digests (INSERT, UPDATE, DELETE) appearing in a reader hostgroup
stats_mysql_errors error code 1290Read-only replica rejecting writes routed to itError 1290 appearing from a reader hostgroup backend
Backend query distribution per hostgroupUnexpected shift in traffic balanceWriter hostgroup absorbing read traffic, or reader hostgroup receiving writes
stats_proxysql_servers_checksums (cluster)Config divergence between ProxySQL nodesChecksum mismatch on mysql_query_rules module between peers
Servers_table_versionInternal MySQL servers-table revision in stats_mysql_globalAn unexpected increment confirms LOAD MYSQL SERVERS TO RUNTIME ran

How Netdata helps

  • Per-second metric collection on Questions, Slow_queries, and backend connection pool metrics reveals routing shifts as they happen, not minutes later during the next polling interval.
  • Backend traffic distribution: correlating per-hostgroup query counts and connection usage shows when traffic shifts from readers to the writer (or vice versa) without an explicit rule review.
  • Error rate baselines: if stats_mysql_errors data is exported through Netdata, new MySQL error codes (such as 1290) appearing after a rule reload are surfaced even at low rates that coarser polling would miss.
  • Cluster checksum monitoring: in ProxySQL Cluster deployments, checksum divergence on mysql_query_rules between nodes is a leading indicator of split-brain routing, detectable before clients report inconsistent behavior.