The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-pass-vs-miss ▌

Operations Guides

Varnish pass vs miss: why s_pass and cache_miss are not the same thing

A Varnish cache_miss and an s_pass both result in a backend fetch, and that surface similarity is where the confusion starts. Teams see backend request rates climbing, glance at the counters, and reach for the wrong fix. A miss is a cache lookup that found nothing and will try to cache the response. A pass is a deliberate VCL decision to bypass the cache entirely via return(pass). The difference matters because the two paths have different storage behavior, different tuning levers, and different operational consequences.

The core distinction

Miss (cache_miss). A cache lookup was performed, no matching object was found, and Varnish fetches from the backend with the intent to cache the result. The response body goes to your configured cache storage (malloc or file). If the response is cacheable, it becomes a cached object and the next request for the same key produces a cache_hit.

Pass (s_pass). VCL decided this request should bypass the cache entirely. The request goes through vcl_pass, then to the backend. The response body goes to transient storage, which is unbounded by default. The response is never cached, regardless of what headers the backend sends. Every subsequent request for the same key repeats the full round trip.

The practical consequence: if you lump these together and see “high miss rates,” you might try to fix it by increasing cache size or extending TTLs. But if the traffic is actually passes (intentional or not), those fixes do nothing. You need to either fix the VCL logic causing unintentional passes, or accept that this traffic is correctly uncacheable and plan backend capacity accordingly.

How it works

Every client request enters vcl_recv, where VCL makes the routing decision. The request either takes the pass path or proceeds to vcl_hash for a cache lookup.

flowchart TD
    A["Client request"] --> B{"vcl_recv decision"}
    B -->|"return(pass)"| C["s_pass: bypass cache"]
    B -->|"default or return(hash)"| D{"Cache lookup result"}
    D -->|"valid object"| E["cache_hit"]
    D -->|"nothing found"| F["cache_miss: fetch, will cache"]
    D -->|"hit-for-miss obj"| G["cache_hitmiss: refetch, no coalescing"]
    D -->|"hit-for-pass obj"| H["cache_hitpass: short-circuit to pass"]
    F --> I["Body to cache storage"]
    C --> J["Body to transient storage"]
    H --> J
    G --> J

The pass path

When vcl_recv returns pass, the request skips the cache lookup entirely. It proceeds to vcl_pass, then to the backend. The backend response is delivered to the client through transient storage but never inserted into the cache. The s_pass counter increments.

Common legitimate reasons to pass:

  • Authenticated or personalized content where caching would serve the wrong user
  • Non-idempotent methods (POST, PUT, DELETE), which the built-in VCL passes by default
  • Content with Set-Cookie headers that VCL has decided not to strip

The problem path: a VCL bug or overly broad cookie logic causes requests that should be cached to pass instead. You see elevated s_pass, backend request rates climb, and hit ratio suffers. But the standard hit rate formula may not reveal this clearly, because passes are excluded from the denominator.

The miss path

When vcl_recv does not return pass, the request proceeds to vcl_hash to compute the cache key, then performs a cache lookup. If no matching object exists, cache_miss increments. Varnish fetches from the backend with the intent to cache. The response body goes to configured cache storage (SMA.s0 for malloc, SMF.* for file-backed). If the backend response is cacheable, the object is stored and served as a cache_hit on the next request.

Hit-for-miss and hit-for-pass: cached decisions not to cache

Varnish caches the decision that certain responses are uncacheable, so it does not waste time attempting cache lookups for them on subsequent requests. These cached decisions are called hit-for-miss and hit-for-pass objects.

When a backend response is uncacheable (TTL <= 0, Set-Cookie, Cache-Control: no-store, Vary: *, and similar), the built-in VCL in Varnish 5.0+ creates a hit-for-miss object. This object tells Varnish: for the next request with this cache key, skip the cache lookup, fetch from the backend, and do not coalesce concurrent requests. The cache_hitmiss counter tracks how many requests hit these objects.

The cache_hitmiss counter itself arrived in Varnish 5.2, one release after the hit-for-miss behavior landed in 5.0.

Before Varnish 5.0, the default behavior created hit-for-pass objects instead. The cache_hitpass counter tracks hits on these. The behavioral difference matters: hit-for-pass preserves conditional request headers (If-Modified-Since, If-None-Match), so the backend can return a lightweight 304 response. Hit-for-miss strips conditional headers and fetches the full body every time.

Starting in Varnish 5.1, you can explicitly create hit-for-pass objects using return(pass(DURATION)) in vcl_backend_response. This gives you the conditional-request efficiency of hit-for-pass when you know the backend supports it.

Pipe: complete bypass

return(pipe) in VCL is a third path that bypasses Varnish’s object machinery entirely. The request and response are forwarded byte-for-byte between client and backend. Pipe is used for WebSocket upgrades and other non-HTTP-caching traffic. Piped traffic does not consume transient storage the way pass does; the data flows through socket buffers. But large pipe bodies still consume memory.

Where it shows up in production

Unintentional pass storms

The most common operational problem is VCL that unintentionally passes traffic. A typical pattern: VCL strips cookies before the cache lookup, but the stripping logic has a bug that catches requests it should not. Or an application deploy adds Set-Cookie to responses that were previously cacheable, triggering hit-for-miss objects that suppress caching for their TTL window.

Symptoms: s_pass or cache_hitmiss elevated, backend_req rate climbing, hit ratio declining. The hit rate formula cache_hit / (cache_hit + cache_miss + cache_hitpass) may understate the problem because passes are not in the denominator at all.

Transient storage growth

Every pass body goes to transient storage. Transient storage is malloc-backed and unbounded by default; an unsized malloc storage reports SMA.Transient.g_space as 0, because the counter is only maintained for sized storages. If a significant fraction of your traffic is pass traffic, transient storage grows without limit.

This is the silent OOM path. The configured cache storage (say, -s malloc,8G) may be well within limits, but process RSS climbs past it as transient storage accumulates pass bodies. Eventually the OOM killer fires and Varnish dies with no warning from the cache storage counters.

You can assign transient to a sized storage — the -s Transient=malloc,1G syntax predates the 6.x line and works in every supported release: This caps the growth but means pass traffic can fail if the cap is hit.

Hit-for-miss TTL trapping

Hit-for-miss objects have their own TTL, commonly 120s in many versions. During this window, Varnish will not attempt to cache the response, even if the application fixes the caching headers. If an application deploy temporarily sends Set-Cookie on cacheable responses, you are stuck with hit-for-miss objects until their TTLs expire. A purge (return(purge) in VCL or an HTTP PURGE request) drops the specific object immediately; a ban only adds a rule that matches objects lazily, as lookups or the ban lurker reach them, so it is the slower way to clear a hit-for-miss window.

Tradeoffs

Pass is correct for authenticated content where the cache key cannot disambiguage users, non-idempotent methods, and responses that are truly per-user. In these cases, elevated s_pass is expected. The operational concern is backend capacity planning, not cache tuning.

Pass is a bug when content that should be cached is being passed due to VCL logic errors, application header changes, or cookie stripping that is too aggressive or not aggressive enough. The fix is always in VCL or the application, not in cache sizing or TTL tuning.

Hit-for-miss vs hit-for-pass

The built-in VCL uses hit-for-miss for uncacheable responses in Varnish 5.0+. This is the right default for most workloads because it prevents request coalescing on uncacheable content. However, for large objects that change infrequently and where the backend supports conditional requests, explicit hit-for-pass via return(pass(DURATION)) is more efficient: it preserves conditional headers and can produce 304 responses, avoiding unnecessary body transfers.

Signals to watch in production

SignalWhy it mattersWarning sign
MAIN.s_pass rateCounts requests taking the pass pathElevated rate when you expect caching
MAIN.cache_miss rateCounts lookups that found nothing, will fetch with intent to cacheSpike correlates with backend load increase
MAIN.cache_hitpass rateCounts hits on cached pass decisionsUnexpectedly high means content marked uncacheable that should not be
MAIN.cache_hitmiss rateCounts hits on cached miss decisionsElevated means application sending uncacheable headers on cacheable content
SMA.Transient.g_bytesMemory consumed by pass, hit-for-pass, and hit-for-miss bodiesMonotonic growth indicates pass traffic accumulating without bound
MAIN.backend_req rateTotal requests forwarded to backendsRatio to client_req reveals cache effectiveness including pass traffic
MAIN.n_objectcoreObject count including hit-for-miss, hit-for-pass, and busy objectsDisproportionately high relative to n_object indicates many cached non-cache decisions

Reading the hit rate formula correctly

The standard hit rate formula is:

cache_hit / (cache_hit + cache_miss + cache_hitpass)

This formula excludes s_pass entirely. A high pass rate can mask low effective cache efficiency because passes do not appear in either the numerator or the denominator. If 40% of your traffic is pass traffic and your hit rate on the remaining 60% is 90%, your formula reports 90%, but your backend is still handling 46% of all client requests (10% misses plus 40% passes).

To get the full picture, compute backend request rate as a fraction of client request rate: backend_req / client_req. This captures the actual load ratio regardless of how that load is distributed across miss, pass, and hit-for-miss paths.

How Netdata helps

  • Per-second counter collection lets you see s_pass, cache_miss, cache_hitpass, and cache_hitmiss rates change in real time, not at 10-second or 60-second resolution where pass storms appear as step functions.
  • Correlating s_pass with backend_req on a single dashboard distinguishes “more misses because of cold cache” from “more passes because of a VCL change” without switching tools.
  • Transient storage tracking via SMA.Transient.g_bytes alongside process RSS gives early warning before pass-driven memory growth triggers an OOM kill.
  • ML anomaly detection on hit rate, pass rate, and backend request rate surfaces the moment a VCL deploy or application change shifts traffic from the miss path to the pass path, even if absolute values have not crossed a static threshold.
  • n_objectcore relative to n_object as a derived view reveals when hit-for-miss or hit-for-pass objects are accumulating, indicating that the application is making previously cacheable content uncacheable.