The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-lru-nuked-vs-expired ▌

Operations Guides

Varnish n_lru_nuked vs n_expired: healthy eviction or an undersized cache

Every Varnish cache evicts objects. The operational question is whether those evictions are healthy housekeeping or evidence that storage is too small for the working set. Two counters hold the answer. MAIN.n_expired tracks objects that aged out past their TTL. MAIN.n_lru_nuked tracks objects forcibly evicted from storage because a new object needed the space.

The distinction matters because nuking is not inherently a problem. A right-sized cache continuously nukes the unpopular tail of its working set, which is correct LRU behavior. The problem arrives when nuking removes objects that would still serve hits, driving cache hit rate down and pushing more traffic to backends.

What the two counters mean

Both counters are available in varnishstat. On Varnish 4.x and later they carry the MAIN. prefix. On Varnish 3.x the prefix is absent.

CounterDefinition
MAIN.n_expiredObjects removed from cache because they reached their TTL and aged out naturally
MAIN.n_lru_nukedObjects forcefully evicted from storage to make room for a new object
MAIN.n_lru_limitedTimes more storage space was needed but the nuke_limit was reached before enough space was freed
MAIN.n_lru_movedMove operations on the LRU list, where accessed objects are promoted toward the head

n_expired is the healthy baseline. Every object with a finite TTL eventually expires. If your cache only ever increments n_expired and never touches n_lru_nuked, storage is large enough to hold the entire working set through each object’s full TTL. That is ideal but not always cost-effective. For most workloads, some nuking is expected and normal.

n_lru_nuked increments only when storage is full and Varnish must evict a live, not-yet-expired object to insert a new one. The LRU algorithm selects the least recently used object from the tail of the list. This keeps the cache populated with the hottest content, but degrades hit rate when the working set exceeds the allocated storage.

n_lru_limited is the escalation signal. When storage is under enough pressure that Varnish would need to nuke more objects than the nuke_limit parameter allows (default 50 in Varnish 6.0 LTS), the fetch fails and the client receives a 503. This counter tells you nuking has progressed from a cache efficiency problem to an availability problem.

n_lru_moved is not an eviction counter. It tracks how often objects are repositioned on the LRU list because they were accessed. High values are normal cache activity.

How eviction works

When a client request results in a cache miss, Varnish fetches the object from the backend and stores it. The storage allocator checks whether enough free space exists in the configured stevedore (typically malloc or file).

If free space exists, the object is inserted. If not, Varnish walks the LRU list from the tail, evicting objects one by one until enough space is freed. Each evicted object increments n_lru_nuked. This happens synchronously during the fetch path: the worker thread inserting the new object performs the nuking.

Objects expire through a separate path. The expiry thread removes objects whose TTL, grace, and keep timers have all elapsed. Each removal increments n_expired. This happens asynchronously and does not block request processing.

These two counters measure fundamentally different exit paths from the cache:

flowchart TD
    A[Object in cache] --> B{Why removed?}
    B -->|TTL elapsed| C[n_expired]
    B -->|Storage full, new fetch| D[n_lru_nuked]
    C --> E[Healthy eviction]
    D --> F{Hit rate declining?}
    F -->|No| G[Healthy: LRU pruning tail]
    F -->|Yes| H[Undersized: working set exceeds storage]

A race condition in the nuking path is worth knowing about. A worker thread can request LRU nuking to free space, but a competing thread can claim the freed space before the first thread uses it. This means n_lru_nuked can increment even when the total free space across the stevedore would appear sufficient. The counter reflects nuking attempts, not a clean capacity calculation.

The discriminator ratio

A single counter value tells you nothing in isolation. The ratio that separates TTL-bound eviction from storage-bound eviction is:

n_lru_nuked / (n_lru_nuked + n_expired)

  • Below 0.5: the cache is TTL-bound. Most objects leave the cache by expiring naturally. Storage is adequate for the working set.
  • Above 0.5: the cache is storage-bound. More objects leave by forceful eviction than by natural expiry. The working set does not fit in the allocated storage.
# Read both counters in one shot
varnishstat -1 -f MAIN.n_lru_nuked -f MAIN.n_expired

Take two readings several seconds apart and compute the deltas. The ratio of the deltas is more informative than the ratio of cumulative totals, because cumulative totals include the initial cache warmup period where n_lru_nuked was zero.

Hit rate is the decisive signal

The ratio tells you whether the cache is storage-bound, but not whether that matters. A cache can be storage-bound and still perform well. The decisive signal is whether the cache hit rate is declining.

Moderate nuking with a stable hit rate is healthy. The LRU is evicting the unpopular tail: objects cached once, rarely re-requested, and unlikely to generate hits. The hot objects stay. This is the LRU doing exactly what it should.

Nuking with a declining hit rate is the problem. Objects that would still serve hits are evicted before they are re-requested. Each evicted popular object becomes a cache miss on its next request, generating a backend fetch. The backend fetch retrieves the same object that was just nuked, inserts it (nuking something else), and the cycle repeats. This is cache thrashing.

# Watch nuking and cache outcomes together
varnishstat -1 -f MAIN.n_lru_nuked -f MAIN.n_expired -f MAIN.cache_hit -f MAIN.cache_miss -f MAIN.cache_hitpass

If n_lru_nuked is climbing and cache_miss is climbing while cache_hit stays flat or drops, the cache is undersized. The working set is larger than what the stevedore can hold.

Another useful correlation: check whether SMA.{name}.g_space is near zero when nuking is active. For malloc storage, use SMA.s0.g_space. For file storage, use SMF.s0.g_space. If storage shows meaningful free space but nuking is still happening, the issue may be malloc fragmentation rather than genuine capacity shortage. Fragmentation causes g_space to report bytes that cannot be allocated as contiguous blocks.

When nuking escalates to 503s

Varnish caps the number of objects a single fetch can nuke via the nuke_limit parameter. The default is 50 in Varnish 6.0 LTS. If a fetch needs to free space and hits this limit before enough space is freed, the fetch fails.

Two consequences follow:

  1. MAIN.n_lru_limited increments. This counter is the hard signal that storage pressure has progressed beyond cache inefficiency into request failure.

  2. The client receives a 503 response. With streaming enabled (the default in Varnish 4.x+), Varnish may begin client delivery in parallel with the backend fetch, so the failure can manifest as a truncated response or a “transfer closed with outstanding read data remaining” error rather than a clean 503.

A version history note: in Varnish versions before 4.1.7, the nuke_limit parameter was not enforced. The fix introduced in 4.1.7-beta1 began honoring the limit, which caused new 503 errors on caches with heavy nuking. If you see n_lru_limited incrementing after an upgrade from an older version, the parameter change is likely the cause. The workaround is to raise nuke_limit, but the real fix is more storage or a smaller working set.

Where it shows up in production

Several real-world patterns produce nuking:

  • Traffic growth without storage resize. The working set grew as the site added content or traffic, but the -s malloc,SIZE allocation was not updated. Nuking begins gradually and accelerates as the working set exceeds capacity.

  • Large objects entering the cache. A few large responses such as video segments, API payloads, or PDFs can consume disproportionate storage. Use varnishtop -I ObjHeader:Content-Length to identify the heaviest objects by response size. Note that chunked responses do not carry a Content-Length header.

  • Object overhead with many small objects. Each cached object carries metadata (objecthead, objectcore, object structs) beyond its body bytes. A cache holding many small objects may nuke aggressively because per-object overhead consumes storage faster than body size alone would suggest.

  • Cache warmup after restart. After a child process restart, the cache is empty and fills rapidly. Nuking is zero initially while storage is empty, then spikes once storage fills. This is transient and should subside as the cache reaches steady state.

  • Shortlived objects and transient storage. Objects with TTL below the shortlived parameter threshold go to transient storage rather than the main stevedore. The default is 10 seconds. If many short-TTL objects are created, the main stevedore may show different nuking patterns than expected. Transient storage is unbounded by default and grows independently.

Signals to watch in production

SignalWhy it mattersWarning sign
MAIN.n_lru_nuked rateStorage is full and objects are being forcibly evictedSustained nonzero rate exceeding n_expired rate
MAIN.n_expired rateHealthy TTL-based eviction baselineNear zero while n_lru_nuked climbs: almost no natural expiry, all removals are forced
n_lru_nuked / (n_lru_nuked + n_expired)Discriminates TTL-bound from storage-bound evictionAbove 0.5 means storage-bound
MAIN.cache_hit rateWhether evicted objects are being re-requestedDeclining while n_lru_nuked climbs means cache thrashing
SMA.{name}.g_spaceFree space in the stevedoreNear zero confirms storage is full
SMA.{name}.c_failAllocation failures after evictionNonzero means even nuking could not free usable space
MAIN.n_lru_limitedFetch failures because nuke_limit was hitAny increment means 503s from storage pressure
MAIN.n_objectCurrent object count in cacheDeclining during nuke storms confirms active eviction

How Netdata helps

Netdata’s Varnish collector surfaces per-second rates for both n_lru_nuked and n_expired, so you can read the storage-bound ratio as a live signal rather than computing deltas manually.

  • Rate computation: Netdata derives rates from cumulative counters automatically. You see nukes-per-second and expirations-per-second without manual delta math.
  • Hit rate correlation: the same dashboard shows n_lru_nuked alongside cache_hit, cache_miss, and backend_req, confirming whether nuking is driving miss rate up or the LRU is cleanly pruning the tail.
  • Storage utilization context: SMA.*.g_bytes and g_space appear alongside eviction counters, showing whether the stevedore is genuinely full or fragmentation is the real issue.
  • n_lru_limited visibility: the 503-producing escalation path is monitored as a distinct signal. Alert on any increment before users see errors.