The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-child-panic ▌

Operations Guides

Varnish child panic: Child died signal, core dumps, and the crash loop

Child (NNNN) died signal=N means the Varnish child (worker) process crashed. The management process supervises and restarts it automatically, so a single crash is self-recovering. But each restart empties the cache, and if the child keeps crashing, caching drops to zero and backends absorb full uncached traffic.

The immediate question: one-off crash or crash loop? If the child recovers and the cache warms back up, it is a TICKET. If MAIN.uptime never stabilizes and stays far below MGT.uptime, the child is dying before the cache can warm. That is a PAGE.

What this means

Varnish runs two processes. The management process (root-owned) handles VCL compilation, CLI access, and child supervision. The child process drops privileges and handles all cache operations: accepting connections, executing VCL, serving objects, fetching from backends.

When the child hits an unrecoverable error such as an assertion failure or a segfault, Varnish calls its internal panic handler. The child logs a panic message with a stack trace and terminates. The management process detects the death via a signal and restarts the child. The listening socket stays open during the restart window, but no requests are served until the child is back.

Each restart empties the cache completely. All MAIN.* counters reset to zero. Cache hit rate drops to zero until the cache warms back up, which takes minutes to hours depending on traffic volume and object TTLs.

The critical distinction:

  • Single crash: MGT.child_panic or MGT.child_died increments once. The child restarts, MAIN.uptime eventually exceeds your warmup window and stabilizes. Cache hit rate recovers. This is a TICKET: investigate the root cause, but service has recovered.
  • Crash loop: MGT.child_panic or MGT.child_died increments repeatedly. MAIN.uptime never stabilizes because the child keeps dying before the cache warms. MAIN.uptime stays far below MGT.uptime. Cache hit rate stays near zero. Backends receive full uncached traffic. This is a PAGE.

Management counters (MGT.*) do not reset on child restart. MGT.uptime continues counting through restarts, while MAIN.uptime resets each time. MAIN.uptime far below MGT.uptime is the definitive signal of recent or repeated restarts.

flowchart TD
    A["Child running\nMAIN.uptime climbing"] --> B["Child crashes\nsignal or panic"]
    B --> C["MGT detects death\nrestarts child"]
    C --> D["Cache emptied\nMAIN.uptime resets"]
    D --> E{"Child stays up?"}
    E -->|"Yes: uptime recovers"| F["Single crash\nself-recovered - TICKET"]
    E -->|"No: crashes again"| B

Common causes

CauseWhat it looks likeFirst thing to check
OOM killchild_died increments, no _.panic file, kernel logs show oom-killerdmesg | grep -i oom
VCL bugchild_panic increments after VCL reload, panic trace references VCL subroutinesvarnishadm panic.show
VMOD bugchild_panic with stack trace in VMOD code, crashes correlate with specific trafficvarnishadm panic.show
Storage corruptionchild_panic with storage allocator in trace, crashes under memory pressure or fragmentationvarnishadm panic.show, check storage config
Version-specific bugchild_panic with assertion failure in known code pathCheck version against known issues

Quick checks

# Management counters for crash, restart, and dump events
varnishstat -1 -f 'MGT.child_panic' -f 'MGT.child_died' -f 'MGT.child_dump' -f 'MGT.child_start'
# Compare child uptime to management uptime - if MAIN is far below MGT, recent restart(s)
varnishstat -1 -f MAIN.uptime -f MGT.uptime
# Last panic message with full stack trace
varnishadm panic.show
# JSON output: varnishadm panic.show -j
# System logs for Varnish restart messages
journalctl -u varnish --since "1 hour ago"
# Kernel OOM killer activity
dmesg | grep -i oom
# Panic files and core dumps in the working directory
ls -la /var/lib/varnish/*/
# Loaded VCL versions and timestamps
varnishadm vcl.list
# Child process memory usage (newest varnishd PID is the child)
ps -p $(pgrep -n varnishd) -o rss= -o vsz=

How to diagnose it

  1. Confirm crash vs crash loop. Check MGT.child_panic, MGT.child_died, and MGT.child_start. A single increment of child_panic or child_died is one crash. Repeated increments mean a loop. Confirm by comparing MAIN.uptime to MGT.uptime: if MAIN.uptime is 30 seconds against MGT.uptime of several hours, the child restarted recently. If MAIN.uptime never climbs past your warmup window, the child is in a crash loop.

  2. Read the panic trace. Run varnishadm panic.show. This returns the signal number, the assertion that failed (if any), and the C stack trace. If the CLI has no panic to show, check for _.panic files in the working directory: ls /var/lib/varnish/*/_.panic. No panic message but child_died present suggests an external kill (OOM killer) rather than an internal Varnish panic.

  3. Check for OOM kills. Run dmesg | grep -i oom and look for entries referencing the varnishd child process. The OOM killer sends SIGKILL, which appears as child_died without a corresponding child_panic. Compare process RSS against system memory and configured storage size. Transient storage growth beyond configured limits is a common OOM path.

  4. Correlate with recent changes. Run varnishadm vcl.list and check timestamps. A VCL reload that preceded the first crash is a strong suspect. If the crash started after a VMOD upgrade or a new VMOD was loaded, the VMOD’s C code is the likely culprit. Check journalctl -u varnish for the exact timing of the first crash relative to deployments.

  5. Examine the stack trace for the failing subsystem. The panic trace shows the C call stack. Function names indicate the subsystem: vbf_fetch_thread points to fetch handling, VRE_match to regex or ban evaluation, HSH_Lookup to cache lookup, storage allocator functions to the object store. This narrows the root cause significantly.

  6. Check for stack overflow. If the panic trace includes THIS PROBABLY IS A STACK OVERFLOW - check thread_pool_stack parameter, the crash was caused by thread stack exhaustion. A reasonable first step is adding 128k to the current thread_pool_stack value.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
MGT.child_panicChild hit an internal panic (assertion failure)Any increment
MGT.child_diedChild died from a signalAny increment
MGT.child_dumpChild produced a core dumpAny increment (confirms core dump was written)
MGT.child_startChild was startedValue greater than 1 means at least one restart
MAIN.uptime vs MGT.uptimeRatio reveals restart frequencyMAIN.uptime stuck below 300s while MGT.uptime climbs
MAIN.cache_hitCache effectiveness after restartNear zero sustained (cache not warming)
Process RSSMemory consumption vs configured storageRSS exceeding configured storage plus overhead budget
dmesg OOM entriesKernel killed the processEntries referencing varnishd

Fixes

OOM kill

If dmesg shows the OOM killer targeting the child, Varnish is consuming more memory than the system allows. The most common cause is allocating all system RAM to cache storage (for example, -s malloc,32G on a 32GB machine). The OS, Varnish process overhead, thread stacks, workspace memory, and transient storage all need memory beyond the configured cache size.

Reserve 20-30% of system RAM for non-storage use. Reduce -s malloc,SIZE if needed. Also check transient storage: SMA.Transient.g_bytes can grow without bound because transient objects have no size limit by default. Uncacheable responses, pass traffic, and hit-for-pass objects all consume transient storage.

If the OOM pattern is intermittent and correlates with traffic spikes, cap transient storage by assigning it to a sized storage segment.

VCL bug

If panic.show shows a stack trace through VCL subroutines or the VCL runtime, the compiled VCL triggered an assertion or memory error. Check varnishadm vcl.list for the most recent VCL change. If the crash started after a reload, roll back to the previous VCL:

varnishadm vcl.list

Identify the active VCL, then switch to a known-good version:

# Non-disruptive: switches active VCL immediately
varnishadm vcl.use <label>

If no previous version is available, load the last known-good VCL file and activate it. Review the panic stack trace for the specific subroutine and function that failed. Complex VCL with inline C, deeply nested ESI, or heavy regex operations in vcl_recv or vcl_hash are common triggers.

VMOD bug

If the panic trace references functions in a VMOD (custom C module), the VMOD’s code segfaulted or triggered an assertion. VMOD bugs are harder to diagnose because the stack trace points to compiled C, not VCL.

The immediate fix is to remove or replace the failing VMOD. If the VMOD provides critical functionality, check for an updated version or file a bug with the panic trace attached. If the crash correlates with specific request patterns (particular URLs, headers, or user agents), blocking that traffic temporarily can stop the crash loop while a fix is developed.

Stack overflow

If the panic message includes THIS PROBABLY IS A STACK OVERFLOW, increase the thread_pool_stack parameter:

# Check current value
varnishadm param.show thread_pool_stack

# Set new value (runtime change, affects newly created threads only)
varnishadm param.set thread_pool_stack <new_value>

Existing threads keep their old stack size. A Varnish restart applies the new value to all threads.

Version-specific bug

Some child panics are caused by bugs in specific Varnish versions. Known examples:

  • Varnish 7.0.0 had a VRE_match() assertion failure during ban evaluation, triggered by specific request patterns hitting ban evaluation during cache lookup (GitHub issue #3714). The panic trace showed ban_evaluate through BAN_CheckObject through HSH_Lookup.
  • Varnish 9.0.0 was affected by CVE-2026-40394, a workspace overflow panic in HTTP/2 session setup, fixed in 9.0.1.

If your panic trace matches a known version-specific bug, upgrade to the nearest patched release. Before upgrading, review the release notes for breaking changes, especially VCL syntax or parameter defaults.

Storage corruption

If the panic trace references storage allocator functions (malloc, SMA, or SMF internals), the storage backend may be corrupted. This is rare but can occur under severe memory fragmentation or filesystem issues with file-backed storage.

The immediate fix is to restart Varnish with a clean storage state. For malloc storage, this is automatic: a restart starts with an empty cache. For file-backed storage, check whether the underlying filesystem is healthy. If corruption recurs, switch to malloc storage temporarily to isolate whether the issue is storage-backend-specific.

Prevention

  • Monitor MGT.child_panic and MGT.child_died. These counters do not reset on child restart, so any increment is a permanent record of a crash. Alert on any nonzero value.
  • Monitor MAIN.uptime against MGT.uptime. A widening gap is the leading indicator of a crash loop. If MAIN.uptime resets while MGT.uptime keeps climbing, the child restarted.
  • Enable core dumps. On systemd-managed deployments, add LimitCORE=infinity to the [Service] section of the Varnish unit file. Setting ulimit -c unlimited in a shell is not sufficient when systemd starts varnishd. Verify that the kernel’s core_pattern does not redirect core files to a pipeline that discards them.
  • Reserve memory headroom. Do not allocate all system RAM to cache storage. Keep 20-30% for the OS, Varnish process overhead, thread stacks, workspace, and transient storage.
  • Test VCL changes before production. Use varnishd -C -f <vcl_file> to compile-check VCL before loading it. Load new VCL alongside the existing one, verify it serves traffic correctly, then switch with vcl.use. Keep the previous VCL loaded for instant rollback.
  • Track Varnish version against known issues. Subscribe to security advisories. Versions past end of life do not receive patches for crash-inducing bugs.

How Netdata helps

  • Per-second MGT.child_panic and MGT.child_died collection shows the exact moment a crash occurred and whether the child recovered or looped, without waiting for a manual varnishstat poll.
  • MAIN.uptime and MGT.uptime as continuous gauges make the crash loop pattern immediately visible: MAIN.uptime resetting to zero while MGT.uptime keeps climbing is definitive.
  • Cache hit rate correlation shows the operational impact: hit rate dropping to zero on each restart and failing to recover in a loop confirms backends are taking full traffic.
  • Process RSS alongside configured storage size provides early warning before the OOM killer fires, especially for unbounded transient storage growth.
  • Correlating Varnish child crashes with system-level signals (memory pressure, dmesg OOM events, cgroup limits) in a single timeline eliminates the gap between “Varnish restarted” and “why did it restart.”