The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / proxysql / proxysql-query-cache-memory-growth ▌

Operations Guides

ProxySQL query cache memory growth: high-cardinality caching toward OOM

Query_Cache_Memory_bytes sits at the mysql-query_cache_size_MB ceiling, the purge rate is rising, and hit rate is dropping despite a full cache. This is the high-cardinality caching pattern: the cache is churning through unique keys that never repeat, and the soft limit cannot keep RSS bounded because it only triggers eviction of expired entries.

The worst case is OOM kill. The kernel terminates ProxySQL, clients disconnect, and on restart the cache is empty, stats tables reset, and the cycle starts again. If you are mid-incident, run PROXYSQL FLUSH QUERY CACHE for immediate relief. Then fix the root cause before memory grows back.

What this means

The query cache stores result sets in ProxySQL’s process memory. Each unique combination of query, user, and schema becomes a separate cache entry. When cached queries include high-cardinality values (UUIDs, timestamps, session tokens, per-request identifiers), each invocation produces a cache entry that is written once and never read again.

mysql-query_cache_size_MB is a soft limit (default 256 MB). The purging thread uses it as a threshold to trigger eviction, but the cache only evicts entries whose TTL has expired. If cache_ttl is long and the query stream is diverse, entries accumulate faster than they expire. Memory grows beyond the configured ceiling.

The signature pattern is three signals moving together:

  • Query_Cache_Memory_bytes at or near mysql-query_cache_size_MB.
  • Query_Cache_Purged rate rising (entries being evicted).
  • Hit rate (Query_Cache_count_GET_OK / Query_Cache_count_GET) declining despite the cache being full.

The cache is doing work (writes, evictions, lookups) without providing value (hits). Each cache lookup adds overhead to every matched query, and RSS grows from entries the soft limit cannot reclaim.

flowchart TD
    A["cache_ttl set on query rule"] --> B["High-cardinality query matches"]
    B --> C["Each call creates unique cache key"]
    C --> D["Cache fills with write-once entries"]
    D --> E["mysql-query_cache_size_MB soft limit reached"]
    E --> F["Purge evicts only expired entries"]
    F --> G{"TTL still valid?"}
    G -->|Yes| H["Memory exceeds soft limit"]
    G -->|No| C
    H --> I["RSS grows, hit rate stays low"]
    I --> J["OOM kill"]

Common causes

CauseWhat it looks likeFirst thing to check
High-cardinality queries cachedQuery_Cache_Entries high, Query_Cache_count_SET rising fast, hit rate near 0%stats_mysql_query_digest for queries matching cache rules that have many unique parameter sets
Excessively long cache_ttlMemory growing steadily, purge rate low until the TTL window starts expiringruntime_mysql_query_rules WHERE cache_ttl > 0
mysql-query_cache_size_MB too large for host memoryRSS close to system limit, other processes competing/proc/<pid>/status VmRSS vs available system memory
Memory fragmentation amplifying cache growthRSS growing faster than Query_Cache_Memory_bytes, jemalloc_resident well above jemalloc_allocatedstats_memory_metrics jemalloc columns

Quick checks

All commands connect to the admin interface (default port 6032). Replace credentials with your actual admin user and password.

# Check all query cache counters
mysql -u <admin_user> -p -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM stats_mysql_global WHERE Variable_Name LIKE 'Query_Cache%';"

# Check which query rules have caching enabled and their TTLs
mysql -u <admin_user> -p -h 127.0.0.1 -P 6032 \
  -e "SELECT rule_id, match_digest, match_pattern, cache_ttl FROM runtime_mysql_query_rules WHERE cache_ttl > 0;"

# Check the configured soft limit
mysql -u <admin_user> -p -h 127.0.0.1 -P 6032 \
  -e "SELECT Variable_Name, Variable_Value FROM global_variables WHERE Variable_Name = 'mysql-query_cache_size_MB';"

# Check detailed memory breakdown including jemalloc fragmentation
mysql -u <admin_user> -p -h 127.0.0.1 -P 6032 \
  -e "SELECT * FROM stats_memory_metrics;"

# Check OS-level RSS, virtual size, and swap usage
cat /proc/$(pidof proxysql)/status | grep -E 'VmRSS|VmSize|VmSwap'

# Check top queries by frequency to spot high-cardinality patterns
mysql -u <admin_user> -p -h 127.0.0.1 -P 6032 \
  -e "SELECT hostgroup, schemaname, digest_text, count_star FROM stats_mysql_query_digest ORDER BY count_star DESC LIMIT 20;"

How to diagnose it

  1. Confirm cache saturation. Check Query_Cache_Memory_bytes against the configured limit. Calculate the utilization ratio: Query_Cache_Memory_bytes / (mysql-query_cache_size_MB * 1048576). If this ratio is at or above 1.0, the cache is at its soft limit.

  2. Calculate the hit rate. Divide Query_Cache_count_GET_OK by Query_Cache_count_GET. A hit rate below 20% on a full cache means the cache is churning, not serving. Both counters are cumulative since ProxySQL start, so use deltas between two readings if the process has been running for a long time.

  3. Check the purge rate. Query_Cache_Purged is a cumulative counter. If it is climbing fast relative to uptime, entries are being evicted at high volume. Compare the purge rate against Query_Cache_count_SET (writes). If writes and purges are both high but hits are low, the cache is a revolving door.

  4. Identify cached query rules. Run the second quick check command to list all rules with cache_ttl > 0. Note the TTL value (in milliseconds) and the match_digest or match_pattern for each.

  5. Cross-reference with query digests. For each cached rule, check stats_mysql_query_digest to see how many unique parameter sets match that rule. A query like SELECT * FROM sessions WHERE token = ? with count_star in the millions is a high-cardinality candidate. Each unique token value creates a separate cache entry.

  6. Check RSS vs cache memory. Compare jemalloc_resident (or OS-level VmRSS) against Query_Cache_Memory_bytes. If RSS is significantly larger than cache memory, fragmentation or other memory consumers (connection buffers, prepared statement cache) are amplifying the problem. Check stats_memory_metrics for the full breakdown.

  7. Assess OOM proximity. Check /proc/<pid>/status for VmRSS against available system memory. Also check VmSwap. If swap is non-zero, performance has already degraded and OOM is approaching.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Query_Cache_Memory_bytesCurrent cache memory usage vs soft limitApproaching or exceeding mysql-query_cache_size_MB
Query_Cache_count_GET_OK / Query_Cache_count_GETHit rate measures cache effectivenessDeclining despite cache being full indicates churn
Query_Cache_PurgedCumulative count of evicted entriesRising rate signals eviction pressure
Query_Cache_EntriesCurrent number of cached entriesGrowing without proportional hit rate means high cardinality
Query_Cache_count_SETCache write rateMuch higher than GET_OK rate means write-once, never-read
jemalloc_residentActual process RSS from allocatorGrowing faster than Query_Cache_Memory_bytes signals fragmentation
ProxySQL_UptimeContext for cumulative countersRecent restart resets all cumulative metrics, making ratios unreliable

Fixes

Immediate relief: flush the cache

PROXYSQL FLUSH QUERY CACHE;

This immediately frees all query cache memory. It is safe to run in production. The trade-off: the cache is cold after the flush, so hit rate drops to zero temporarily and backend load increases until the cache warms back up. If the underlying query rules are unchanged, memory will grow back. This buys time, not a permanent fix.

Shorten cache_ttl

Reduce the TTL on rules matching high-cardinality queries so entries expire sooner. This limits the working set size. For example, if a rule currently sets cache_ttl = 300000 (5 minutes), reducing it to 30000 (30 seconds) means entries are eligible for purging much sooner.

UPDATE mysql_query_rules SET cache_ttl = 30000 WHERE rule_id = <id>;
LOAD MYSQL QUERY RULES TO RUNTIME;
SAVE MYSQL QUERY RULES TO DISK;

Trade-off: shorter TTL means more cache misses and more backend load for queries that do repeat. Use this when the query mix has some repeatable queries mixed with high-cardinality ones.

Stop caching high-cardinality queries entirely

Remove cache_ttl from rules that match queries with UUIDs, timestamps, session tokens, or other per-request unique values. These queries produce cache entries that are never reused. The cache lookup overhead is pure cost with no benefit.

UPDATE mysql_query_rules SET cache_ttl = NULL WHERE rule_id = <id>;
LOAD MYSQL QUERY RULES TO RUNTIME;
SAVE MYSQL QUERY RULES TO DISK;

This is the correct permanent fix when a cached query’s parameter space is effectively unbounded. Even with a short TTL, each second of traffic floods the cache with single-use entries.

Prevent caching empty results

If cached queries frequently return empty result sets, each empty result still consumes a cache entry. Setting cache_empty_result to 0 on the query rule prevents caching these entries. The column is present in ProxySQL v2.0+ mysql_query_rules schemas.

UPDATE mysql_query_rules SET cache_empty_result = 0 WHERE rule_id = <id>;
LOAD MYSQL QUERY RULES TO RUNTIME;
SAVE MYSQL QUERY RULES TO DISK;

Useful when a high-cardinality query often returns no rows (for example, existence checks against a sparse table).

Reduce mysql-query_cache_size_MB

If the soft limit is set high relative to available system memory, lower it.

UPDATE global_variables SET Variable_Value = 128
  WHERE Variable_Name = 'mysql-query_cache_size_MB';
LOAD MYSQL VARIABLES TO RUNTIME;
SAVE MYSQL VARIABLES TO DISK;

This reduces the ceiling for cache growth. Trade-off: less cache capacity means more evictions and potentially lower hit rates for queries that do benefit from caching. This is a mitigation, not a root-cause fix.

Reset the query digest table separately

The digest table is a separate memory consumer from the query cache. If it is also growing (many unique query digests accumulating), it contributes to RSS independently.

SELECT * FROM stats_mysql_query_digest_reset;

Reading from stats_mysql_query_digest_reset returns the current contents and clears the table; this is the documented and version-portable reset path. Current ProxySQL source also handles TRUNCATE TABLE stats.stats_mysql_query_digest, but prefer _reset unless you have verified that form on your deployed version.

These two subsystems (query cache and query digests) are independent. Flushing one does not free memory from the other.

Prevention

  • Audit query rules before enabling cache_ttl. Before adding caching to a rule, check the cardinality of queries it matches. A query that includes a UUID, timestamp, or per-request token in its WHERE clause will produce a unique cache entry for every invocation.
  • Monitor hit rate as a ratio, not absolute memory. A full cache with a near-zero hit rate is worse than an empty cache. The cache adds lookup overhead to every matched query while providing no offload to backends.
  • Alert on the combination, not individual signals. Query_Cache_Memory_bytes near the limit alone is normal for a well-utilized cache. It becomes a problem only when combined with declining hit rate and rising purge rate. Alert on all three together.
  • Watch RSS divergence. Track jemalloc_resident against Query_Cache_Memory_bytes. If RSS grows much faster than cache memory, fragmentation is amplifying the problem. A restart clears fragmented memory but is not a sustainable fix.
  • Set mysql-query_cache_size_MB conservatively. The soft limit does not prevent RSS from exceeding it. Leave headroom for connection buffers, query digest storage, and other memory consumers in the same process.

How Netdata helps

  • Per-second collection of Query_Cache_Memory_bytes, Query_Cache_Entries, and Query_Cache_Purged makes cache growth and eviction rates visible at high resolution, catching the gradual RSS climb before it reaches OOM.
  • The hit rate ratio (Query_Cache_count_GET_OK / Query_Cache_count_GET) can be visualized alongside cache memory usage, making churn immediately visible: memory climbing while hit rate falls.
  • jemalloc_resident and jemalloc_allocated from stats_memory_metrics let you separate cache growth from allocator fragmentation. RSS diverging from Query_Cache_Memory_bytes signals fragmentation, not just cache pressure.
  • Correlating cache metrics with Questions rate and backend pool utilization shows whether cache churn is driving backend load spikes as misses pass through to MySQL.
  • ProxySQL_Uptime gating prevents false alerts during cold start, when cache counters reset and ratios are temporarily meaningless.