The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-cache-stampede-thundering-herd ▌

Operations Guides

Varnish cache stampede: a popular object expires and the herd hits the backend

A popular cached object hits its TTL and expires. In the next second, hundreds of concurrent requests for that object all miss simultaneously. Varnish forwards all of them to the backend, which was sized for the small fraction of traffic that normally leaks through, not a synchronized burst of identical requests. The backend slows. Worker threads pile up waiting for responses. If the backend cannot recover quickly, thread exhaustion follows and sessions start dropping.

The signature is a sudden, sharp spike in MAIN.cache_miss and MAIN.backend_req that mirror each other, often correlating with a deployment, a purge, or synchronized TTL expiry across a group of objects.

Varnish has a built-in defense: request coalescing. When multiple clients request the same uncached URL, Varnish sends one request to the backend and puts the rest to sleep on a “busy object,” waking them when the fetch completes. This works when the herd is asking for the same object. But when hundreds of unique URLs expire at the same time, coalescing cannot help. The number of concurrent backend fetches equals the unique-URL cardinality of the expiring set.

The fix is not more threads. More threads paper over the symptom while the backend drowns. The fix is grace mode, staggered TTLs, and soft-purge: mechanisms that either prevent the synchronized miss entirely or serve stale content while the cache refills.

What this means

The cascade:

  1. A popular object’s TTL expires, or it is invalidated by a ban or purge.
  2. Concurrent requests that would have been cache hits all become misses.
  3. Each miss generates a backend fetch. Request coalescing collapses requests for the same URL, but different URLs generate different fetches.
  4. The backend receives a spike of concurrent requests, often exceeding its capacity.
  5. Backend response time increases, holding worker threads longer.
  6. The thread pool fills, the queue grows, and sessions or requests are dropped.

The critical distinction from other failure patterns: the cache miss spike happens first, before backend degradation. If the backend degraded first and then hit rate dropped, you have a different problem. In a stampede, the cache is the trigger, not the backend.

flowchart TD
    A[Popular object expires] --> B[Concurrent requests all miss]
    B --> C{Coalescing applicable?}
    C -->|Same URL, few unique| D[One backend fetch, others wait]
    C -->|Many unique URLs| E[Many concurrent backend fetches]
    D --> F[Backend load manageable]
    E --> G[Backend load spike]
    G --> H[Backend response time rises]
    H --> I[Worker threads held longer]
    I --> J[thread_queue_len rises]
    J --> K[sess_dropped / req_dropped]

Common causes

CauseWhat it looks likeFirst thing to check
Synchronized TTL expiryMany objects with identical TTLs expire at the same momentvarnishlog -i TTL for objects expiring in clusters
Mass invalidation (ban/purge storm)MAIN.bans_added or MAIN.n_purges spike, followed immediately by miss spikevarnishadm ban.list and application purge logs
VCL change causing cache key changeHit rate drops to near zero after VCL reload, all requests missvarnishadm vcl.list for recent changes
Cold cache after restartMGT.child_start recently incremented, hit rate at 0%MAIN.uptime compared to MGT.uptime
Hit-for-pass amplificationcache_hitpass high, coalescing not engagingVCL for return(pass) or beresp.uncacheable with zero TTL

Request coalescing only works for cacheable objects. If VCL returns pass for a request, or if a response is marked beresp.uncacheable = true with a TTL of zero, every concurrent request for that URL bypasses coalescing and hits the backend independently. A single URL that should be cached but is accidentally passed can amplify a stampede from one backend fetch to hundreds.

Quick checks

All commands below are read-only and safe to run on a production Varnish instance:

# Check miss rate and backend request rate together
varnishstat -1 -f MAIN.cache_hit -f MAIN.cache_miss -f MAIN.cache_hitpass -f MAIN.backend_req

# Check request coalescing activity
varnishstat -1 -f MAIN.busy_sleep -f MAIN.busy_wakeup -f MAIN.busy_killed

# Check thread pool saturation
varnishstat -1 -f MAIN.threads -f MAIN.thread_queue_len -f MAIN.threads_limited

# Check if grace is already serving stale content
varnishstat -1 -f MAIN.cache_hit_grace

# Check for recent ban/purge activity
varnishstat -1 -f MAIN.bans_added -f MAIN.n_purges

# Check backend health under load
varnishadm backend.list -p

# Check object expiry patterns
varnishlog -i TTL -g request | head -100

# Check session/request drops (cascade confirmation)
varnishstat -1 -f MAIN.sess_dropped -f MAIN.req_dropped

How to diagnose it

  1. Confirm the miss spike mirrors the backend_req spike. If cache_miss and backend_req spike together at the same ratio, the cache is the trigger. Take two readings 5 to 10 seconds apart to compute the rate.

  2. Check request coalescing counters. Rising busy_sleep with matching busy_wakeup means coalescing is working: requests for the same URL are collapsed into a single fetch. If busy_sleep is low but backend_req is high, the herd is hitting many unique URLs and coalescing cannot help.

  3. Check for busy_killed. Any nonzero rate means requests timed out waiting for a busy object to be fetched. Clients received 503s because Varnish ran out of resources to manage the coalescing queue. The stampede has outgrown the coalescing safety net.

  4. Check temporal correlation. Did the spike start immediately after a VCL reload, a deployment, a purge operation, or a restart? Match the timestamp of the first miss spike to operational events. A stampede that starts at a predictable interval (every 5 minutes, every hour) points to synchronized TTL expiry.

  5. Check thread pool state. If thread_queue_len is rising, the stampede is cascading into thread starvation. The problem has moved from “backend is slow” to “users get nothing.” If sess_dropped or req_dropped is incrementing, the cascade has reached the user.

  6. Check whether grace is active. cache_hit_grace tracks hits served from stale objects. If this counter is rising during the spike, grace mode is mitigating the stampede. If it is zero and you have grace configured, the grace period may be too short, or the objects may not have had a previous cached version to serve as stale.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
MAIN.cache_miss rateDirect count of cache lookups finding nothingSudden spike not explained by traffic increase
MAIN.backend_req rateLoad Varnish places on backendsSpike mirroring cache_miss at approximately 1:1 ratio
MAIN.busy_sleep rateRequests waiting on coalesced fetchHigh rate with low busy_wakeup rate = fetches completing slowly
MAIN.busy_killed rateRequests killed from busy wait listAny nonzero value = clients got 503s during coalescing
MAIN.cache_hit_graceHits served from stale contentHigh during a stampede = grace is preventing backend overload
MAIN.thread_queue_lenWorker pool saturationSustained nonzero = stampede cascading into thread starvation
MAIN.sess_dropped / MAIN.req_droppedUsers getting nothingCascade endpoint; zero is the only acceptable sustained value
MAIN.bans_added / MAIN.n_purgesInvalidation activitySpike preceding miss spike = mass invalidation trigger

Fixes

Grace mode (stale-while-revalidate)

Grace is the primary defense. When an object’s TTL expires, grace allows Varnish to serve the stale object to clients while a single background fetch refreshes it. The herd gets stale content instantly. Only one request goes to the backend.

sub vcl_backend_response {
    set beresp.grace = 6h;
}

The default grace is 10 seconds (the default_grace parameter; 10s in every supported release), which is too short for a popular object with high request rates. A grace period of minutes to hours (depending on content staleness tolerance) means that even if the backend is slow or down, Varnish continues serving stale content and does not generate a stampede.

Tradeoff: longer grace means clients may see staler content. For rapidly-changing content, shorter grace is appropriate. For versioned assets or API responses with cache headers, longer grace is safe.

Soft-purge

A hard purge removes the object from cache immediately. The next request misses and fetches from the backend. If the object is popular, the stampede begins.

Soft-purge reduces the object’s TTL without removing it. The object stays available for grace serving while a background fetch refreshes it. Using the bundled purge VMOD (its functions may only be called from vcl_hit or vcl_miss, never vcl_recv):

import purge;

sub vcl_recv {
    if (req.method == "PURGE") {
        return (hash); # fall through to vcl_hit / vcl_miss
    }
}

sub vcl_hit {
    if (req.method == "PURGE") {
        purge.soft(ttl = 0s, grace = 6h, keep = 1h);
        return (synth(200, "Soft purge"));
    }
}

sub vcl_miss {
    if (req.method == "PURGE") {
        purge.soft(ttl = 0s, grace = 6h, keep = 1h);
        return (synth(200, "Soft purge"));
    }
}

This sets the object’s TTL to zero so it is immediately stale, but preserves its grace and keep values. Varnish serves it as stale content while refreshing in the background.

Staggered TTLs

If many objects share the same TTL, they all expire at the same moment. Adding jitter to each object’s TTL spreads expiry across a window, preventing synchronized misses.

The exact implementation depends on your VCL and available VMODs, but the principle is simple: 1000 objects with a nominal 5-minute TTL should expire anywhere between 5 and 6 minutes, not all at the 5-minute mark. Never let a large set of popular objects share a single expiry instant.

Fix hit-for-pass misconfiguration

If VCL is accidentally passing cacheable content, or if uncacheable responses have a zero TTL, request coalescing is disabled for those URLs and every concurrent request hits the backend independently.

For responses that should not be cached, set a non-zero TTL on the uncacheable marker:

sub vcl_backend_response {
    if (beresp.status >= 400) {
        set beresp.uncacheable = true;
        set beresp.ttl = 120s;
    }
}

This creates a hit-for-miss object: Varnish remembers the decision not to cache and serves it from cache for 120 seconds. During that window, request coalescing applies to the hit-for-miss object, preventing repeated backend fetches for the same uncacheable URL.

Prevention

  • Set grace on every cacheable object type. Default 10 seconds is insufficient for most production traffic. Use 5 to 30 minutes for most content, hours for static assets.
  • Use soft-purge instead of hard purge for popular objects. Reserve hard purge for content that must be immediately unavailable.
  • Add TTL jitter. Even 30 to 60 seconds of jitter across objects with the same nominal TTL prevents the herd from forming.
  • Watch coalescing counters proactively. busy_sleep, busy_wakeup, and busy_killed tell you whether coalescing is engaging and succeeding. Most teams do not watch these until after their first stampede.
  • Audit VCL for accidental pass. Any return(pass) in vcl_recv disables coalescing. Verify that pass is only used for genuinely uncacheable content, and that uncacheable responses have a non-zero TTL to create hit-for-miss objects.
  • Do not respond to a stampede by adding threads. More worker threads means more concurrent backend fetches and more backend load. The backend is already the bottleneck.

How Netdata helps

  • Per-second metric collection means the cache_miss to backend_req correlation is visible at the resolution needed to confirm a stampede pattern, not smoothed away by 60-second polling.
  • Correlating cache_miss, backend_req, busy_sleep, and thread_queue_len on a single timeline shows the full cascade: cache miss triggers backend spike, which triggers coalescing activity, which (if insufficient) triggers thread saturation.
  • cache_hit_grace alongside cache_miss reveals when grace is absorbing the stampede versus when misses are reaching the backend.
  • Anomaly detection on backend_req rate can flag the spike before thread saturation begins.
  • busy_killed as an alert signal catches the moment coalescing fails and users start seeing 503s.