The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-vcl-fail ▌

Operations Guides

Varnish vcl_fail: VCL runtime errors during request processing

MAIN.vcl_fail is incrementing. The VCL loaded and compiled successfully at vcl.load time, but something is breaking at runtime while real requests flow through the VCL subroutines.

The vcl_fail counter (Varnish 5.1 and later) counts failures that prevented VCL from completing. Each increment means a request could not finish its normal VCL lifecycle. On the client side, the request is diverted to vcl_synth with a 503 status. On the backend side, the fetch fails. In both cases the user gets an error response instead of the content they asked for.

The vcl_fail counter is a summary. It tells you VCL broke, but not why. The root cause lives in the shared memory log (VSL) under the VCL_Error tag, in workspace overflow counters, or in the specific VMOD that called VRT_fail(). Correlating vcl_fail with s_synth, ws_*_overflow, and recent VCL changes is the diagnostic path.

What this means

VCL executes through a sequence of subroutines during request processing: vcl_recv, vcl_hash, cache lookup, vcl_hit / vcl_miss / vcl_pass, then optionally vcl_backend_fetch / vcl_backend_response, and finally vcl_deliver. At any point, the VCL can hit a runtime error condition that prevents completion.

return(fail) can appear in any VCL subroutine, including inside VMODs. When a failure occurs on the client side (in vcl_recv, vcl_hash, vcl_hit, vcl_miss, vcl_pass, vcl_deliver, vcl_pipe, or vcl_purge), the request is sent to vcl_synth with a 503 status and reason “VCL failed”. When a failure occurs on the backend side (in vcl_backend_fetch, vcl_backend_response, or vcl_backend_error), the fetch fails. If return(fail) is called from within vcl_synth itself, Varnish sends a minimal 500 response and forces the connection closed.

flowchart TD
    A["VCL execution during request"] --> B{"Runtime error condition?"}
    B -- No --> C["Normal request flow"]
    B -- Yes --> D["MAIN.vcl_fail increments"]
    D --> E["Client side: diverted to vcl_synth 503"]
    D --> F["Backend side: fetch fails"]
    D --> G{"Severe VCL or VMOD bug?"}
    G -- Yes --> H["Child process panic"]
    G -- No --> I["Synthetic error to client"]
    E --> J["MAIN.s_synth increments"]

A vcl_fail increment is distinct from a VCL compilation error. Compilation errors happen at vcl.load time and prevent the VCL from being activated at all. Runtime VCL failures happen against already-loaded, successfully compiled VCL.

It is also distinct from backend_fail (TCP connection to backend failed) and fetch_failed (backend connected but the fetch broke). Those are backend communication problems. The backend might be perfectly healthy while VCL is failing.

Common causes

CauseWhat it looks likeFirst thing to check
Workspace exhaustionws_client_overflow or ws_backend_overflow incrementing alongside vcl_fail. Large Cookie headers, many synthetic headers added by VCL.varnishstat -1 -f 'MAIN.ws_*_overflow'
VMOD failureA VMOD called VRT_fail() or hit an internal error. VCL_Error entries in the log reference the VMOD.varnishlog -q 'VCL_Error' -g request
Illegal VCL operationVCL attempted something not allowed in the current subroutine context, or req.restarts exceeded max_restarts.varnishlog -q 'VCL_Error' -g request
ESI processing failureesi_errors incrementing. ESI includes failing, sometimes cascading into VCL failures on sub-requests.varnishstat -1 -f MAIN.esi_errors
Child panic (escalation)MGT.child_panic incrementing. Severe VCL or VMOD bugs escalate beyond vcl_fail to a full child crash.varnishstat -1 -f MGT.child_panic

Quick checks

Run these read-only checks to establish the scope of the problem.

# Check vcl_fail rate (take two readings 10 seconds apart, compute delta)
varnishstat -1 -f MAIN.vcl_fail

# Check for workspace overflows (the most common vcl_fail cause)
varnishstat -1 -f 'MAIN.ws_*_overflow' -f MAIN.losthdr

# Check synthetic response rate (vcl_fail produces synthetic responses)
varnishstat -1 -f MAIN.s_synth

# Check for child panics (severe escalation)
varnishstat -1 -f MGT.child_panic -f MGT.child_died -f MGT.child_start

# Check VCL state and loaded versions
varnishadm vcl.list

# Check loaded VCL count (multiple VCLs may indicate failed reloads)
varnishstat -1 -f MAIN.n_vcl -f MAIN.n_vcl_avail -f MAIN.n_vcl_discard

# Search the shared memory log for VCL_Error entries
varnishlog -q 'VCL_Error' -g request

# Check ESI errors if Edge Side Includes are in use
varnishstat -1 -f MAIN.esi_errors -f MAIN.esi_warnings

How to diagnose it

  1. Confirm vcl_fail is actually incrementing. Take two varnishstat readings 10 seconds apart. Compute the delta. If the rate is zero, the failures may have been transient or already resolved.

  2. Check workspace overflow counters. Run varnishstat -1 -f 'MAIN.ws_*_overflow'. If ws_client_overflow or ws_backend_overflow are incrementing, workspace exhaustion is the cause. This is the most common trigger for vcl_fail.

  3. Search for VCL_Error log entries. Run varnishlog -q 'VCL_Error' -g request. The VCL_Error tag records the specific error string when return(fail) is called with an argument, or when a VMOD calls VRT_fail(). The error string tells you exactly what broke.

  4. Correlate with s_synth. Run varnishstat -1 -f MAIN.s_synth. If s_synth is incrementing at a similar rate to vcl_fail, the failures are producing synthetic 503 responses to clients.

  5. Check for recent VCL changes. Run varnishadm vcl.list. Look at timestamps. A recently activated VCL is the most likely culprit. If the VCL status is “available” (not “active”), it is not serving traffic and cannot cause runtime failures.

  6. Check for child panics. Run varnishstat -1 -f MGT.child_panic. If panics are incrementing alongside vcl_fail, the VCL or VMOD bug is severe enough to crash the child process. Check for _.panic files in /var/lib/varnish/.

  7. Identify the failing subroutine or VMOD. Use varnishlog -g request -q 'VCL_Error' and examine the surrounding transaction context. The log entry shows which VCL subroutine was executing when the failure occurred, and if a VMOD was involved, which VMOD function failed.

  8. Check if max_restarts is the trigger. If VCL uses return(restart) and req.restarts exceeds max_restarts, control passes to vcl_synth with “Too many restarts”. Search for restart-related VCL_Error entries.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
MAIN.vcl_failDirect counter of VCL runtime failuresAny nonzero sustained rate
MAIN.s_synthSynthetic responses generated by Varnish, including those from vcl_failSpike correlating with vcl_fail
MAIN.ws_client_overflowClient workspace exhausted during request processingAny nonzero rate
MAIN.ws_backend_overflowBackend workspace exhausted during fetch processingAny nonzero rate
MAIN.losthdrHTTP headers dropped due to limit exceededAny nonzero rate
MGT.child_panicSevere VCL/VMOD bug escalating to child crashAny increment
MAIN.esi_errorsESI parse errors that may cascade into VCL failuresAny nonzero rate
MAIN.n_vclNumber of loaded VCLs (failed reloads accumulate)Unexpected growth

Fixes

Workspace exhaustion

If ws_client_overflow or ws_backend_overflow are incrementing, the per-request workspace is too small for the headers, cookies, or VCL-added data in your traffic. Large Cookie headers are the most common cause.

# Check current workspace sizes
varnishadm param.show workspace_client
varnishadm param.show workspace_backend

# Increase workspace (delayed effect, not immediately active)
varnishadm param.set workspace_client 96k
varnishadm param.set workspace_backend 96k

The current defaults are 96k for both. Increasing to 128k or more may resolve Cookie-driven overflows. The tradeoff is memory: the effective cost scales with concurrent worker contexts and configured pools, so estimate from thread_pool_max * thread_pools and verify RSS after the change.

Make the change persistent by adding -p workspace_client=96k -p workspace_backend=96k to the Varnish startup parameters.

VMOD failure

If VCL_Error log entries point to a specific VMOD, the VMOD is calling VRT_fail() or encountering an internal error. Common scenarios: a VMOD making external calls (DNS, Redis, HTTP) hitting a timeout or connection error, or a VMOD receiving unexpected input data.

  1. Identify the VMOD from the VCL_Error message.
  2. Check if the VMOD call is wrapped in error handling in VCL. If not, add one.
  3. If the VMOD itself has a bug, check for an updated version.
  4. As a stopgap, remove the VMOD call from the hot path and reload VCL.

Too many restarts

If VCL uses return(restart) and req.restarts exceeds max_restarts, the restart loop triggers vcl_fail.

varnishadm param.show max_restarts

The fix is to fix the VCL logic causing excessive restarts, not to raise max_restarts. A restart loop usually indicates a logic error where vcl_recv keeps redirecting back to itself.

If esi_errors is incrementing alongside vcl_fail, Edge Side Includes are breaking. ESI processing increases workspace pressure: each ESI sub-request runs its own VCL cycle, multiplying workspace and thread consumption.

See the related guide on Varnish ESI errors for detailed troubleshooting.

Severe bugs escalating to child panic

If MGT.child_panic is incrementing, the VCL or VMOD bug is crashing the child process entirely. The child restarts automatically under management process supervision, but each restart loses the entire cache.

  1. Check for panic files: ls /var/lib/varnish/*/_.panic
  2. Read the panic message to identify the failing component.
  3. Roll back to the previous VCL if a recent change triggered it: varnishadm vcl.use <previous_vcl_name>
  4. If a VMOD is causing the panic, remove the VMOD call from VCL and reload.

Prevention

  • Monitor vcl_fail continuously. It should be zero in steady state. Alert on sustained nonzero rate once MAIN.uptime > 300 (past warmup).

  • Monitor workspace overflow counters. ws_client_overflow and ws_backend_overflow are leading indicators. If they start incrementing before vcl_fail, you have early warning.

  • Test VCL changes before activation. Use varnishd -C -f your.vcl to compile-check VCL before loading. This catches compilation errors but not runtime errors. For runtime testing, load the new VCL without making it active, then use varnishadm vcl.use to activate after verification.

  • Track VCL versions. Use varnishadm vcl.list regularly. Accumulating loaded VCLs can indicate failed reload attempts or old VCLs not being discarded. Each loaded VCL consumes resources.

  • Size workspace for your traffic. If your application sets large Cookie headers, proactively increase workspace_client above the default. Monitor losthdr as well: if headers are being dropped, workspace pressure is already building.

  • Wrap VMOD calls with error handling. Any VMOD that can fail (network calls, external lookups) should have its return value checked in VCL.

How Netdata helps

  • Per-second resolution. Netdata collects MAIN.vcl_fail at one-second resolution, so you can pinpoint the exact moment failures start and correlate with deployments or traffic changes.

  • Workspace correlation. When vcl_fail spikes, overlay ws_client_overflow, ws_backend_overflow, and losthdr in the same dashboard. If workspace counters lead vcl_fail by even a few seconds, you have the root cause.

  • s_synth correlation. Overlaying MAIN.s_synth against vcl_fail confirms whether failures are reaching clients as synthetic 503 responses or whether vcl_synth is handling them gracefully.

  • Child panic escalation. Netdata monitors MGT.child_panic and MGT.child_died alongside MAIN.uptime vs MGT.uptime. If vcl_fail escalates to child panics, you see the crash-restart cycle immediately.

  • VCL lifecycle tracking. Changes in MAIN.n_vcl and MAIN.n_vcl_avail appear alongside vcl_fail rate. If a VCL reload precedes a spike, the new VCL is the likely trigger.