The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-cache-full ▌

Operations Guides

uWSGI cache subsystem: hit ratio drops and 'full' insert failures

Response times are creeping upward on cached endpoints. Application logs show nothing. Worker busy ratio is normal. The cause is likely silent: uWSGI cache misses are increasing, and each miss forces the worker to compute or fetch the response instead of serving from shared memory.

This article covers uWSGI’s built-in cache subsystem (cache2), not external caches. If caches[] is absent from the stats server output, the uWSGI cache is not enabled and these diagnostics do not apply.

The stats server exposes cache metrics in the caches[] array. Two symptoms dominate: a rising miss rate with steady traffic (cache churn or insufficient capacity), and a non-zero full counter (insert operations failing because the cache has no free slots). Both degrade response time, but they have different root causes and fixes.

What it means

The uWSGI cache is an in-process key-value store backed by shared memory mapped across all workers. It is configured via cache2 in your uWSGI config. Each cache has a fixed number of slots (items) and a fixed block size (blocksize). When the cache is full, new insert attempts fail silently or with a warning, depending on configuration.

Three counters tell you almost everything:

FieldMeaning
hitsCumulative count of successful cache lookups (monotonic)
missCumulative count of failed lookups (key not found)
fullCumulative count of insert operations that failed because the cache had no free slot

The hit ratio is hits / (hits + miss). A drop with stable traffic means the cache is too small for the working set, or items are expiring before re-request.

The full counter is the more direct signal. Any non-zero value means the cache rejected an insert. Each rejected insert is a future cache miss for that key.

flowchart TD
    A[Cache miss rate rising] --> B{full counter greater than 0?}
    B -->|Yes| C[Cache capacity exhausted]
    B -->|No| D{items near max_items?}
    D -->|Yes| E[Undersized cache or churn]
    D -->|No| F{blocksize too small?}
    F -->|Possibly| G[Silent insert failures]
    F -->|No| H[TTL expiry too aggressive]
    C --> I[Increase items or enable purge_lru]
    E --> I
    G --> J[Increase blocksize]
    H --> K[Review expires values]

Common causes

CauseWhat it looks likeFirst thing to check
Cache undersized for working setitems at or near max_items; full non-zero and rising; hit ratio low even after restartCompare items to max_items and check the full counter trend
TTL expiry too aggressiveHit ratio drops after restart and never recovers; full is zero; items well below max_itemsCheck expires values in cache_set calls or routing config
Blocksize too small for valuesitems stays at 0 or very low despite insert attempts; full may be zero; no error loggedVerify blocksize exceeds your largest cached value
Master process not running (no sweeper)TTL-based expiry never fires; cache fills with expired entries never cleaned; full rises over timeConfirm master = true in config
purge_lru initialization bug (fixed in 2.0.24)purge_lru=1 set but eviction does not work; cache fills and stays fullCheck uWSGI version; upgrade to 2.0.24+ if using purge_lru

Quick checks

# Dump all cache stats from the stats server
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.caches[]'

# Focused view of key cache metrics
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.caches[] | {name, items, max_items, hits, miss, full}'

# Calculate current hit ratio
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.caches[] | .hits as $h | .miss as $m | {name, hit_ratio: ($h / ($h + $m))}'

# Check whether items is near capacity
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.caches[] | {name, items, max_items, utilization: (.items / .max_items)}'

# Look for the DANGER full-cache warning in uWSGI logs
journalctl -u uwsgi --no-pager | grep "DANGER.*cache.*FULL"

# Check uWSGI version for known cache bugs
uwsgi --version

# Verify master process is running (required for cache sweeper)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.pid'

Adjust the stats socket address (127.0.0.1:9191) to match your deployment. For UNIX socket stats, use uwsgi --connect-and-read /path/to/stats.sock.

How to diagnose it

  1. Confirm the cache is enabled. Check your uWSGI config for cache2 directives. If the application uses Redis or Memcached internally, the uWSGI cache metrics are irrelevant. Verify by checking whether caches[] appears in the stats output.

  2. Check the full counter trend. Take two readings a few minutes apart. If full is increasing, inserts are actively failing. If it is zero, the cache has room but misses are still rising (pointing to TTL or churn).

  3. Check items versus max_items. If items is at or near max_items, the cache is full. The first slot of a cache is reserved internally, so a cache configured with items=1000 holds only 999; items at 999 with max_items of 1000 means the cache is at capacity.

  4. Check hit ratio over time. If hit ratio is high immediately after a restart but degrades over hours, TTL expiry or eviction is too aggressive relative to the access pattern. If hit ratio is low from the start, the cache is undersized or the working set does not fit.

  5. Check for silent insert failures. If items stays near zero despite the application attempting inserts, blocksize may be too small for the values being stored. A single-block cache stores at most blocksize - 1 bytes per item. Compare your actual response body sizes against the configured blocksize.

  6. Check whether the master process is running. The cache sweeper thread, responsible for TTL-based expiration, only runs when the master process is enabled. Without it, expired items are never cleaned up and the cache fills permanently with stale entries. Confirm master = true in your config.

  7. Check the uWSGI version if using purge_lru. purge_lru was added in 2.0.4; 2.0.7 fixed an off-by-one corruption bug in cache LRU mode and 2.0.24 fixed purge_lru cache initialization. If purge_lru=1 is set but eviction does not seem to work, upgrade to 2.0.24 or newer.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
caches[].full (rate)Each increment is a failed insert; that key will miss on next lookupAny non-zero rate
caches[].miss (rate)Rising miss rate means more requests bypass the cache and hit the backendRate increasing while traffic is stable
caches[].items vs max_itemsShows how close the cache is to capacityitems within 1 of max_items
Hit ratio hits / (hits + miss)Primary cache efficiency metricSustained drop below baseline
Worker avg_rtCache misses show up as slower responses; correlation confirms cache impactavg_rt rising in step with miss rate
uWSGI versionKnown bugs affected purge_lru and bitmap mode in older 2.0.x releasesRunning pre-2.0.24 with purge_lru enabled

Fixes

Cache is undersized

Increase the items value in your cache2 configuration. Each item consumes blocksize bytes of shared memory, so the total footprint is items * blocksize. Verify the host has enough RAM for the larger allocation.

If you cannot increase items due to memory constraints, enable LRU eviction with purge_lru=1 (available since uWSGI 2.0.4). This evicts the least recently accessed item when the cache is full, making room for new inserts.

When purge_lru=1 is active, the expires argument on cache_set calls is ignored and eviction is purely access-based (official uWSGI docs). If you need TTL-based expiry alongside LRU eviction, do not rely on purge_lru; size the cache to hold the full working set instead.

Blocksize is too small

Increase blocksize to accommodate your largest cached value. For example, with blocksize=65536, the largest storable item is at most 65535 bytes. If your cached responses are larger, inserts fail silently with no error.

If values vary widely in size and most are small, consider bitmap mode (bitmap=1), which lets an item span multiple contiguous blocks and is production-ready since uWSGI 2.0.2. If you combine bitmap=1 with purge_lru=1, use uWSGI 2.0.24 or newer: 2.0.7 fixed an off-by-one corruption bug in cache LRU mode and 2.0.24 fixed purge_lru cache initialization.

TTL expiry too aggressive

Review the expires values passed to cache_set or configured in routing rules. If items expire before the next request for the same key, the cache provides no benefit. Increase expires to match or exceed the typical inter-request interval for each cached resource.

If you are using purge_lru=1, TTL expiry is not honored (see the caveat above). If TTL-based expiry is required, do not use purge_lru; size the cache to hold the full working set instead.

Master process not running

Enable master = true in your uWSGI config. Without the master process, the cache sweeper thread never starts, and TTL-based expiration does not work. Expired items accumulate until the cache is full of stale data with no mechanism to reclaim slots.

Suppressing full-cache log warnings

If you have consciously decided to let the cache run full (for example, with purge_lru handling eviction), the repeated *** DANGER cache "<name>" is FULL !!! *** log lines on every insert can flood your logs. Use ignore_full (added in uWSGI 2.0.4) to suppress these warnings only when you have an eviction strategy in place. Otherwise it hides a real capacity problem.

Prevention

  • Size the cache to the working set, not to a round number. Measure how many unique keys the application references within the TTL window. Set items to at least 120% of that count to allow for growth.
  • Match blocksize to actual value sizes. Audit your largest cached responses and set blocksize with headroom. Silent insert failures from blocksize mismatch are the hardest to diagnose because they produce no error.
  • Monitor the full counter as a rate, not as an absolute. Any non-zero rate means inserts are being rejected. Alert on it.
  • Track hit ratio over time, not just point-in-time. A slowly declining hit ratio over days or weeks indicates the working set is growing past cache capacity.
  • Verify the master process is enabled if you rely on TTL expiry. Without the sweeper thread, the cache fills with expired entries and never recovers.
  • Keep uWSGI current. The cache subsystem has received fixes in recent 2.0.x releases. Verify which bugs affect your version before relying on purge_lru or bitmap mode.

How Netdata helps

Netdata collects uWSGI cache metrics from the stats server at per-second resolution. Correlate these signals during investigation:

  • Cache hit and miss rates over time. Reveals whether a declining hit ratio is gradual (working set growth) or sudden (code change or traffic shift).
  • The full counter rate. The most direct signal of capacity exhaustion. Per-second collection catches burst insert failures that coarser polling intervals miss.
  • Miss rate vs. worker avg_rt. Confirms cache misses are the cause of response time degradation, not a downstream dependency.
  • items relative to max_items. Shows proximity to capacity before inserts start failing.
  • Anomaly detection on hit ratio. Surfaces slow working-set drift over weeks that static thresholds would miss.