The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-grace-masking-backend-failure ▌

Operations Guides

Varnish grace masking a backend outage: the ticking-clock incident

Your origin servers went down 12 minutes ago. The dashboard shows a 97% cache hit rate, zero 503s, and normal response times.

This is grace masking. Varnish is serving cached objects past their TTL because no healthy backend can refresh them. Clients see stale 200s. Monitoring sees healthy cache traffic. The outage stays invisible until grace expires on enough popular objects, at which point 503s cascade.

Grace buys time proportional to the shortest grace TTL across your working set. When that runway runs out, the hit rate can drop from 97% to near zero in minutes. The impact is deferred, not prevented.

What this means

Grace mode lets Varnish serve a stale object while dispatching a background fetch to refresh it. When the backend is unreachable, the background fetch fails and Varnish keeps serving the stale object until its grace period expires. From the client side, the response is a valid 200 with an elevated Age header. From the monitoring side, the request counts as a cache_hit.

In Varnish 7.x, the builtin vcl_hit delivers an object when obj.ttl >= 0s (fresh) or when obj.ttl + obj.grace > 0s (stale but within grace). When TTL has expired but grace has not, the object is delivered as a grace hit and a background fetch is dispatched. If the backend is healthy, the refresh succeeds silently. If the backend is sick, the refresh fails and Varnish keeps serving the stale object, consuming the grace runway with each request.

Grace hits are included in the cache_hit counter, so standard client-facing signals stay green:

  • Cache hit ratio holds steady or drops only slightly.
  • Client 503 rate stays at zero.
  • Response latency stays low, because Varnish is serving from memory.

Only signals that measure backend state independently reveal the truth.

There is a second mechanism that can shorten the runway. When a background fetch receives a 5xx response from the backend (as opposed to a connection timeout, which fails without a response), the builtin vcl_backend_response may store that error response, overwriting the graced object in cache. For the TTL of that error object (commonly 120 seconds, depending on default_ttl and response headers), Varnish serves the error instead of the stale 200. This creates a 503 window embedded inside the grace period, appearing before you would expect grace to expire.

flowchart TD
    A[Backend outage] --> B[Probes mark backends sick]
    B --> C[Grace serves stale objects]
    C --> D[Client metrics: hit rate stable, zero 503s]
    C --> E[Backend signals: VBE happy drops, backend_unhealthy climbs]
    D --> F{Backend recovers before grace expires?}
    F -->|Yes| G[Background fetches refresh objects]
    F -->|No| H[Grace expires on popular objects]
    H --> I[503 cascade begins]

Common causes

The backend outage itself can have many root causes. What matters operationally is that grace masks all of them equally, and the fix for the masking problem is the same regardless of the underlying failure.

CauseWhat it looks likeFirst thing to check
Backend application crash or deploy failureVBE.*.happy drops to zero across all backends; backend_unhealthy incrementsvarnishadm backend.list -p for probe details
Network partition between Varnish and backendbackend_fail increments (TCP connect failures); probes time outNetwork connectivity from Varnish host to backend port
Backend overload causing probe timeoutsBackend slow but eventually responds; probes fail intermittentlyBackend response time and load metrics
5xx overwrite during outage503s for specific objects while grace should still cover themWhether 5xx background fetch responses overwrite graced objects

Quick checks

Run these read-only commands to confirm whether grace is masking a backend outage.

# Check backend health with probe details
varnishadm backend.list -p

# Check grace-serving activity (V7+)
varnishstat -1 -f MAIN.cache_hit_grace

# Check backend sick counters (connections not attempted)
varnishstat -1 -f MAIN.backend_unhealthy -f MAIN.backend_fail

# Check per-backend probe success counts
varnishstat -1 -f 'VBE.*.happy'

# Check object expiry rate (declining rate signals objects surviving via grace)
varnishstat -1 -f MAIN.n_expired

# Check background fetch thread failures (V7+)
varnishstat -1 -f MAIN.bgfetch_no_thread

# Check synthetic response rate (503s generated by Varnish)
varnishstat -1 -f MAIN.s_synth

# Inspect object TTL and grace values to estimate runway
varnishlog -i TTL

How to diagnose it

  1. Confirm backend health independently. Run varnishadm backend.list -p and check whether backends show probe failures. Look at VBE.*.happy values: if happy is below the probe threshold, the backend is sick. Check MAIN.backend_unhealthy to see how many connections Varnish skipped because backends were marked unhealthy. A quiet backend_fail counter does not mean the backend is fine. backend_unhealthy counts connections that were never attempted, so a completely offline backend shows zero backend_fail but climbing backend_unhealthy.

  2. Check if grace is actively serving. On Varnish 7+, check MAIN.cache_hit_grace. If this counter is climbing while backends are sick, grace is masking the outage. Each grace hit is also counted in cache_hit, which is why hit ratio looks normal.

  3. Check the natural expiry rate. MAIN.n_expired counts objects expiring from cache by TTL. During a grace-masking event, this rate declines because objects that would normally expire are kept alive by grace serving. A declining n_expired rate alongside sick backends is a tell.

  4. Estimate the grace runway. Run varnishlog -i TTL to inspect the TTL and grace values of objects being served. The output shows remaining TTL, grace, and keep for each object. The shortest grace values across your popular objects define how long the runway lasts. If most objects have a 30-second grace and the backend has been down for 25 minutes, you are already past the cliff for those objects.

  5. Check for 5xx overwrite. If you see 503s for specific objects while other objects still serve via grace, check whether background fetches receiving 5xx responses are overwriting graced objects. Look for patterns where 503s appear in bursts for specific URLs, then stop when the error object TTL expires and grace serving resumes.

  6. Verify VCL grace configuration. Check whether your VCL sets beresp.grace explicitly or relies on the default_grace parameter (default: 10 seconds). A 10-second grace runway masks very little. Production setups typically set grace to several minutes or more depending on content type and staleness tolerance.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
VBE.*.happyPer-backend probe success count within the windowDrops below probe threshold means backend is sick
MAIN.backend_unhealthyConnections not attempted because backend is sickSustained rate above zero means backends marked sick
MAIN.cache_hit_grace (V7+)Objects served from grace rather than fresh cacheClimbing while backends are sick means grace masking active
MAIN.n_expiredObjects expiring naturally from cache by TTLDeclining rate means objects surviving past TTL via grace
MAIN.backend_failTCP connection failures to backendsSustained rate above zero means backend unreachable
MAIN.bgfetch_no_thread (V7+)Background fetches that failed for lack of threadsRate above zero means thread starvation preventing refreshes
MAIN.s_synthSynthetic responses, typically 503 error pagesSpiking means grace exhausted and Varnish is generating errors

Fixes

Add return(abandon) for 5xx background fetches

The single most important VCL change to prevent the 5xx overwrite problem:

sub vcl_backend_response {
    if (beresp.status >= 500 && bereq.is_bgfetch) {
        return (abandon);
    }
}

When a background fetch gets a 5xx from the backend, return (abandon) discards the response without storing it. The graced object stays in cache and Varnish continues serving it. Without this, the error response can overwrite the graced object, forcing Varnish to serve the error for that object’s TTL.

This is the standard pattern recommended in the Varnish grace documentation.

Configure dynamic grace based on backend health

Set req.grace in vcl_recv to use a short grace when the backend is healthy and a longer grace when it is sick:

sub vcl_recv {
    if (req.backend.healthy) {
        set req.grace = 10s;
    } else {
        set req.grace = 24h;
    }
}

When the backend is healthy, a short grace is enough for background fetches to complete. When the backend is sick, the full grace value applies, giving Varnish a longer runway. The grace runway extends automatically when backends fail.

The tradeoff: longer grace means serving staler content during transient backend issues. Set the sick-backend grace based on the maximum staleness your application can tolerate.

Extend grace during an active incident

If a backend outage is detected and the estimated grace runway is shorter than the expected recovery time, reload VCL with extended grace values.

Use varnishadm vcl.load and varnishadm vcl.use to load and activate the new VCL without dropping connections. The grace extension applies to objects fetched after the VCL change. Objects already in cache retain their original grace values, so this extends the runway for future requests but does not retroactively save objects whose grace has already expired.

Set appropriate default grace values

The default_grace parameter defaults to 10 seconds. Set beresp.grace explicitly in vcl_backend_response based on your application’s staleness tolerance. Common production values range from 2 minutes to several hours, depending on content type. Static assets can tolerate long grace periods. Personalized or time-sensitive content needs shorter grace or none at all.

Prevention

  • Monitor backend health independently of client-facing metrics. If you only alert on hit ratio, 503 rate, and response latency, you will not detect a grace-masking event until it is too late. Alert on VBE.*.happy and MAIN.backend_unhealthy directly.
  • Alert on cache_hit_grace climbing (V7+). Grace serving is normal in small amounts during background refreshes. A sustained climb while backends are healthy can indicate intermittent backend issues. A climb while backends are sick is the ticking clock.
  • Include the abandon pattern in all production VCL. return (abandon) for 5xx background fetches prevents the overwrite that creates embedded 503 windows during outages.
  • Track grace runway as a capacity metric. Know the distribution of grace TTLs across your working set. If your shortest grace values are 30 seconds, a backend outage longer than 30 seconds will produce user-visible failures. Size grace values to cover realistic backend recovery times.
  • Audit health probe configuration. A probe testing a static endpoint that always returns 200 will report healthy even when the application is broken. Ensure probes test real application health. Check the probe threshold, window, and interval to understand how quickly Varnish detects failures.

How Netdata helps

Netdata collects the backend-independent signals that reveal grace masking, with per-second resolution.

  • Backend health. Netdata monitors VBE.*.happy per backend and MAIN.backend_unhealthy rate, so you see backends go sick immediately, independent of what the client-facing hit ratio shows.
  • Grace serving. The MAIN.cache_hit_grace counter (V7+) is collected natively. A spike in grace hits while backends are sick is the earliest warning that the clock is ticking.
  • Expiry rate correlation. MAIN.n_expired is tracked alongside MAIN.n_lru_nuked, letting you distinguish natural TTL expiry from grace-extended survival. A declining n_expired rate during a backend outage confirms grace masking.
  • Background fetch failures. MAIN.bgfetch_no_thread and MAIN.fetch_failed are monitored, catching the thread starvation that prevents grace refreshes from completing.
  • Anomaly detection. ML-based anomaly detection flags the correlation of sick backends plus stable hit ratio plus rising grace hits as abnormal, even when no individual static threshold is breached.