The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-purge-vs-ban-xkey ▌

Operations Guides

Varnish purge vs ban vs xkey: choosing an invalidation method that scales

Varnish offers three mechanisms for invalidating cached objects: purge, ban, and the xkey VMOD (surrogate-key invalidation). Each has a fundamentally different cost model. Purge is O(1) and immediate but works only on a single exact hash. Bans are expression-based and flexible but accumulate in a list that every cache lookup must evaluate. xkey operates on secondary key indexes and avoids list growth entirely, but the open-source implementation has known scaling limits.

The choice matters most under load. An application that issues one ban per content update seems harmless until the ban list reaches tens of thousands of entries and every cache hit becomes an O(n) scan. A CMS that purges by individual URL works fine with dozens of pages but breaks when it needs to invalidate every object associated with a template, author, or category. This article covers the cost model of each method at the cache-object level, why per-URL bans are the anti-pattern that drives teams toward xkey, and which Varnish counters reveal when your invalidation strategy is hurting cache performance.

How each invalidation method works

Purge: exact hash, immediate removal

Purge, triggered by return(purge) in vcl_recv when a PURGE request arrives, removes the single object matching the request’s hash. The hash is composed of whatever vcl_hash assembles, typically the Host header plus the URL. Purge discards the object and all its Vary variants immediately. It is O(1): one hash lookup, one object removal, no lingering state.

Purge is the right tool when you know the exact URL or hash components of the object to invalidate. It leaves no trace, adds no overhead to future lookups, and does not depend on any background thread to complete.

The limitation is scope. If you need to invalidate every object derived from a particular database record, and those objects live under different URLs, purge requires you to enumerate every URL and issue a separate PURGE request for each one.

Ban: expression-based, lazy, accumulates

A ban adds an expression to an in-memory ban list. The expression can match object metadata (obj.* variables such as obj.http.x-tag) or request properties (req.* variables such as req.url or req.http.host). Bans are evaluated lazily: when a ban is inserted, Varnish does not scan existing cached objects. Instead, every cache lookup tests the found object against all outstanding bans.

The ban lurker is a background thread that walks the object store testing objects against obj.* bans. When the lurker finds a match, it evicts the object. This is asynchronous and spreads the cost across idle periods. Bans using req.* variables cannot be processed by the lurker because req.* only has meaning in the context of a live client request. Those bans are checked synchronously on every cache hit and persist in the list until every cached object has been tested against them through the lookup path.

The cost model is O(n x m) where n is the number of objects and m is the number of outstanding bans. The ban list grows until the lurker has verified that every object older than the oldest ban has been tested. If bans are added faster than the lurker processes them, the list grows without bound.

xkey: surrogate-key, tag-based

The xkey VMOD (vmod_xkey from varnish-modules) provides surrogate-key invalidation. Objects are associated with one or more secondary keys (tags) set via response headers. When you call xkey.purge("tag"), all objects sharing that key are invalidated. xkey.softpurge() marks matching objects as expired but extends their grace and keep timers, so stale content is still served while the backend refreshes.

xkey avoids the ban list entirely. It maintains an index of secondary keys to objects. Invalidation is a lookup in that index, not a linear scan of the ban list. This means invalidation cost does not grow with the number of outstanding invalidations.

The tradeoff is operational. xkey is a VMOD, not a core Varnish feature. It must be imported and configured in VCL. Keys are registered when objects are inserted, so objects cached before import xkey; was present in VCL cannot be purged via xkey. The open-source implementation also has documented performance limitations under high insert or purge rates.

flowchart TD
    A["Object needs invalidation"] --> B{"Method?"}
    B --> C["Purge: exact hash"]
    B --> D["Ban: expression"]
    B --> E["xkey: surrogate key"]
    C --> C1["Immediate, O(1)
No list, no lurker"] D --> D1["obj.* bans: lurker
processes asynchronously"] D --> D2["req.* bans: checked on
every cache hit, lurker skips"] D1 --> D3["List grows until all
objects tested - O(n x m)"] D2 --> D3 E --> E1["Key index lookup
All tagged objects removed
No list growth"]

The per-URL ban anti-pattern

The most common scaling failure with Varnish invalidation is per-URL bans. An application or CMS issues a ban matching a specific host and URL path, using req.* variables:

ban("req.http.host == ... && req.url == ...")

This combines the worst properties of both other methods. It targets a single exact URL, which is exactly what purge does in O(1). But it uses req.* variables, so the ban lurker cannot process it. The ban sits in the list and is tested on every cache lookup until every object in the cache has been checked against it. If the application issues one such ban per content update and updates happen frequently, the ban list grows continuously and every cache hit pays the price.

The fix is straightforward. For exact-URL invalidation, use purge. For group-based invalidation (all objects associated with a category, template, or dependency), use xkey surrogate keys. Reserve bans for genuine pattern-matching use cases, and prefer obj.* variables so the lurker can process them asynchronously.

When to use which method

MethodBest forCost modelWatch out for
PurgeSingle exact-URL invalidationO(1), immediate, no listMust enumerate every URL for group invalidation
Ban (obj.*)Pattern-based invalidation the lurker can processO(n x m), list accumulates but lurker keeps it boundedComplex regex slows evaluation; lurker can fall behind under load
Ban (req.*)Invalidation that depends on request propertiesO(n x m), lurker cannot helpStays in list until every object tested; per-URL req.* bans are the anti-pattern
xkey purgeGroup or tag-based invalidation via surrogate keysIndex lookup, no list growthVMOD dependency; open-source version has scaling limits

A practical decision flow:

  • One URL, known exactly. Use purge (return(purge)).
  • Multiple related URLs sharing a tag. Use xkey with surrogate keys set on the backend response.
  • Genuine pattern match across unknown URLs. Use ban with obj.* variables only, and monitor the ban list.
  • Never. Per-URL bans using req.*. This is always the wrong tool.

xkey scaling limits and maintenance mode

The varnish-modules repository describes its VMODs as “feature-complete and maintained to stay compatible with new Varnish releases and to fix bugs.” The xkey VMOD is documented to have known scalability problems under high insert or purge rates.

The specific scaling problems are worth understanding if you depend on xkey:

  • Locking contention on busy sites. xkey piggybacks on the expiry data structure’s mutexes. Under high insert rates or frequent purges, lock contention can degrade performance on sites with heavy cache churn.
  • Objects inserted before xkey import cannot be purged. If Varnish started with VCL that did not include import xkey;, objects cached during that period have no secondary key index entries and are invisible to xkey purge. This is a real problem when config management deploys xkey VCL after Varnish has already started with boilerplate configuration.
  • Persisted cache reindexing on restart. With MSE (Massive Storage Engine, Varnish Enterprise), restarting Varnish with persisted objects requires xkey to reindex every object one by one. For caches with millions of objects, this means millions of disk operations on restart before xkey is functional.

For open-source Varnish Cache, there is currently no direct replacement for xkey’s tag-based invalidation. Teams that need group invalidation on open-source Varnish have two practical options: use bans with obj.* variables and carefully monitor the ban list, or restructure the application to issue individual purges for each URL in the group.

Signals to watch in production

SignalWhy it mattersWarning sign
MAIN.n_purges rateTracks purge operation volume.Sustained rate greater than 10x baseline can indicate an application bug or unauthorized invalidation.
MAIN.bans (gauge)Current ban count. Large values mean every cache lookup is doing more work.Growing trend, especially above 1000.
MAIN.bans_added rateBan insertion rate. If this consistently exceeds bans_deleted, the lurker is falling behind.bans_added rate much higher than bans_deleted rate.
MAIN.bans_deleted rateBan removal rate. Should track bans_added in steady state.Significantly lower than bans_added.
MAIN.bans_lurker_contentionLurker yielded to lookups. Indicates the lurker is competing with request processing and losing.Any sustained nonzero rate with growing ban list.
MAIN.bans_lurker_obj_killedObjects evicted by the lurker. Should roughly correspond to ban injection rate.Near zero with growing ban list, meaning the lurker has stalled.

Note: MAIN.bans_obj_killed counts objects killed by bans during lookup; MAIN.bans_lurker_obj_killed counts objects killed by the lurker. Both counters exist. When invalidation volume spikes unexpectedly, identify the source:

# Check current ban list size and contents
varnishadm ban.list

# Monitor for ban activity in real time
varnishlog -q 'VCL_Log ~ "ban" or CLI ~ "ban"'

# Monitor for PURGE method requests
varnishlog -q 'ReqMethod eq "PURGE"'

If your VCL handles PURGE or BAN methods, verify that it enforces ACL checks on client.ip, not on X-Forwarded-For headers, which can be spoofed. Without ACL enforcement, any network actor can invalidate cache content.

How Netdata helps

  • Per-second MAIN.n_purges and MAIN.bans_added rates. A sudden spike in either counter, correlated with a drop in MAIN.cache_hit and a rise in MAIN.backend_req, pinpoints mass invalidation as the cause of a backend load surge rather than a VCL change or TTL expiry.
  • MAIN.bans as a gauge with trend tracking. Anomaly detection flags sustained growth in the ban list before cache lookup latency degrades visibly to users.
  • MAIN.bans_lurker_contention correlation. When the lurker is losing lock contention, the ban list grows and cache hit latency increases. Seeing both signals on the same dashboard makes the connection immediate.
  • Invalidation-to-hit-rate overlay. Overlaying purge and ban counters against cache hit ratio and backend request rate reveals whether your invalidation strategy is the root cause of a hit rate drop, as opposed to a VCL change, storage pressure, or TTL expiration.
  • MAIN.n_obj_purged tracking. When purge volume is high, this counter confirms how many objects were actually removed, distinguishing a broad purge from a narrow one and helping estimate the cache warmup cost that follows.