The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-hot-restart-race-draining ▌

Operations Guides

Envoy hot restart races: draining, epoch churn, and dropped connections

You deployed a new Envoy binary or triggered a reload that invoked a hot restart. Seconds later, clients report connection resets, your 5xx rate ticks up, and server.hot_restart_epoch is climbing on your dashboard. The new process was supposed to inherit listen sockets gracefully while the old one drained.

Hot restart is Envoy’s mechanism for zero-downtime binary upgrades and certain reloads. A new process launches, coordinates with the old one over a Unix domain socket, takes over the listen sockets, and the old process enters a drain sequence. Both processes run simultaneously during the handoff. That coexistence is where the races live.

The symptoms are specific: dropped connections during rollover, server.hot_restart_epoch incrementing in rapid succession (a crash loop that looks like a hot restart), or anomalous stats values after a version upgrade. The root causes are mechanical: file descriptor exhaustion from the dual-process window, the previous epoch dying before the handoff completes, a concurrency decrease that drops accept-queue connections, or stats schema changes that corrupt shared state.

What this means

Hot restart is a two-process coordination protocol. The old process (epoch N) is serving traffic. A new process (epoch N+1) starts and connects to the old process over a Unix domain socket. The new process sends a restart RPC. The old process hands off its listen sockets, indexed by worker, and transitions to server.state = 1 (DRAINING). The new process binds the inherited sockets, loads its initial xDS configuration, warms its clusters and listeners, and transitions to server.state = 0 (LIVE).

sequenceDiagram
    participant Old as Old (epoch N)
    participant UDS as UDS
    participant New as New (epoch N+1)

    Old->>UDS: listening for restart RPC
    New->>UDS: connect, send restart RPC
    UDS->>Old: forward RPC
    Old->>Old: state = DRAINING
    Old->>New: hand off listen sockets
    Note over Old,New: Both running: FDs double, race window
    New->>New: bind sockets, load xDS
    New->>New: state = LIVE
    Old->>New: transfer stats
    Old->>Old: drain connections
    Old->>Old: parent_connections to 0
    Old->>Old: shutdown

The dangerous window is between “old process starts draining” and “new process is fully LIVE.” During this period:

  • Both processes hold open file descriptors. Total FD usage briefly doubles. If baseline FD usage exceeds 50% of the limit, the dual-process window can push past the FD ceiling and cause silent connection refusal.
  • The old process is draining: it stops accepting new connections on some listeners and sends Connection: close on HTTP/1.1 connections and GOAWAY on HTTP/2 streams. It still holds in-flight connections open until they complete or the drain time expires.
  • The new process may still be initializing (waiting for xDS, warming clusters). If it has not reached LIVE, it is not accepting new connections on all listeners either.

If the old process finishes draining before the new process reaches LIVE, or if the old process crashes mid-drain, there is a gap where neither process serves traffic. Connections in the kernel’s accept queue or in-flight on the old process can be dropped.

On Linux, Envoy defaults to SO_REUSEPORT sockets, which lets both processes bind the same port simultaneously. During hot restart, sockets are passed by worker index so the kernel distributes new connections correctly. But if --concurrency decreases between epochs, some workers in the old process have no counterpart in the new process, and connections queued on those orphaned workers can be dropped.

Hot restart is enabled by default and is not supported on Windows. In Istio sidecar mode, the pilot-agent starts Envoy with --disable-hot-restart, so this mechanism does not apply unless you have overridden that flag.

Common causes

CauseWhat it looks likeFirst thing to check
Crash loop via hot restartserver.hot_restart_epoch increments multiple times per minute; server.uptime keeps resettingEnvoy stderr logs for crash reason, assert failures, or initialization errors
FD exhaustion in dual-process windowConnection refusal during restart; FD count at or near ulimit`ls /proc//fd
Previous epoch dies before handoffNew process fails to initialize; hot restart assert or RPC failure in logsserver.parent_connections drops to 0 abruptly; check parent exit code
Concurrency decreaseIntermittent connection drops after a concurrency reductionCompare --concurrency between old and new process startup flags
Stats corruption after version upgradeAnomalous gauge values, missing counters, or impossible stats after binary upgradeCompare stats output before and after; check version compatibility
High-frequency restart raceBoth parent and child die simultaneously at very short restart intervalsRestart trigger frequency; look for sub-second intervals

Quick checks

These commands are read-only and safe to run during an incident. Adjust the admin port for your deployment (default 9901, or 15000 in Istio sidecar mode).

# Check process state, epoch, and parent connections
curl -s http://localhost:9901/stats | grep -E 'server\.(state|hot_restart_epoch|live|parent_connections)'

# Check readiness: 200 means LIVE, 503 means draining or initializing
curl -s -o /dev/null -w "%{http_code}\n" http://localhost:9901/ready

# Get server info including epoch and uptime
curl -s http://localhost:9901/server_info | python3 -m json.tool

# Check FD usage on the Envoy process
ENVOY_PID=$(pgrep -x envoy | head -1)
ls /proc/$ENVOY_PID/fd | wc -l
grep 'Max open files' /proc/$ENVOY_PID/limits

# Check total connections across both processes
curl -s http://localhost:9901/stats | grep 'server.total_connections'

# Check hot restart version compatibility
curl -s http://localhost:9901/hot_restart_version

# Count Envoy processes (2 during hot restart, 1 otherwise)
pgrep -x envoy | wc -l

How to diagnose it

  1. Confirm a hot restart is in progress or recently completed. Check server.hot_restart_epoch. If it has incremented in the last few minutes, a hot restart occurred. Check pgrep -x envoy. If two processes are running, the handoff is still in progress.

  2. Determine whether the old process is still draining. Check server.parent_connections. A non-zero value means the old process still holds connections and has not finished draining. If server.parent_connections stays non-zero longer than --drain-time-s (default 600 seconds), the drain is stuck.

  3. Check for a crash loop. If server.hot_restart_epoch increments more than once per minute, the new process is crashing and the supervisor is restarting it. Look at Envoy stderr logs for the crash reason. Common causes: bad xDS config that fails validation, incompatible binary, or an assert failure in the hot restart RPC path.

  4. Check FD pressure. During the dual-process window, FD usage doubles. Run ls /proc/<pid>/fd | wc -l against both PIDs and compare against the limit from /proc/<pid>/limits. If usage is above 80% of the limit, FD exhaustion is the likely cause of connection drops.

  5. Verify the new process reached LIVE. Check server.state on the new process (should be 0). If it is stuck at 2 or 3 (PRE_INITIALIZING or INITIALIZING), the new process is waiting for xDS configuration. The old process eventually shuts down after --parent-shutdown-time-s (default 900 seconds). If the new process has not reached LIVE by then, you get a gap with no serving process.

  6. Check for stats anomalies after upgrades. If the hot restart crossed a version boundary, compare stats output before and after. Look for missing counters, impossible gauge values (negative connection counts, gauges that should be monotonic going backwards), or counters that reset unexpectedly. The stats transfer between old and new processes can produce corrupt values when the stat schema differs between versions.

  7. Check restart trigger frequency. If you use SIGHUP or a restart script, look at how frequently it fires. The hot restart protocol has a known race at very short restart intervals (Envoy issue #550). Restarts triggered within roughly 250ms of each other can kill both parent and child simultaneously.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
server.hot_restart_epochTracks restart count; rapid incrementing indicates a crash loopMore than 1 increment per minute
server.stateLifecycle state (0=LIVE, 1=DRAINING, 2=PRE_INITIALIZING, 3=INITIALIZING)Non-zero sustained outside planned restarts
server.parent_connectionsConnections still held by old process during drainNon-zero longer than --drain-time-s
server.total_connectionsTotal across both processes; proxy for FD usageSudden drop during restart window
server.live1 if process is not draining0 on new process after old process exits
FD count (/proc/<pid>/fd)FD exhaustion causes silent connection refusalAbove 50% of limit at baseline (doubles during restart)
server.uptimeResetting indicates a new process startedResetting repeatedly without planned deployment

Fixes

Crash loop via hot restart

If the new process crashes immediately after starting, the hot restart protocol cycles. Each failed epoch leaves behind state and compounds the problem.

Check Envoy stderr logs for the crash reason. If the crash is from bad xDS config, fix the config or roll back the control plane change. If the crash is from a binary incompatibility, verify version compatibility using --hot-restart-version on the new binary and compare against GET /hot_restart_version on the running process. If the versions are incompatible, perform a cold restart (stop the old process, start the new one) instead of a hot restart.

For environments where the parent process may die before the child initializes, consider --skip-hot-restart-on-no-parent. Without this flag, the child terminates if the parent is gone, which can create a loop if the parent is unstable.

FD exhaustion during dual-process window

The simplest fix is to raise the FD limit (ulimit -n or the container security context). For production proxies, an FD limit of at least 65536 is reasonable. The deeper fix is to reduce baseline FD usage so the dual-process window does not push past the limit: shorten idle timeouts to close stale connections faster, verify connection pooling is enabled, and check for FD leaks from access log files or excessive health check connections.

Keep baseline FD usage below 50% of the limit when hot restart is in use. The restart temporarily doubles it.

Previous epoch dies before handoff completes

If the old process crashes or is killed (OOM, signal, orchestrator eviction) before the new process has sent its restart RPC and received the listen sockets, the hot restart fails. The new process either asserts or falls back depending on flags.

Check the old process exit code and logs. If it was OOM-killed, address memory pressure. If it was killed by the orchestrator because the termination grace period was too short, increase the grace period to exceed --drain-time-s plus a buffer.

--skip-hot-restart-on-no-parent allows the new process to fall back to a normal startup if the parent is gone, instead of terminating. This trades the zero-downtime property for resilience.

Concurrency decrease drops connections

If --concurrency decreases between epochs, the old process has more workers than the new one. Connections queued on workers that have no counterpart in the new process are dropped because no worker inherits that socket.

This is a known limitation of the socket-handoff mechanism. To avoid it, do not decrease concurrency during a hot restart. If you must reduce worker count, do it as a separate cold restart after the hot restart completes, or accept the connection drops during the transition.

Stats corruption after version upgrade

When the old and new processes run different Envoy versions, the stats schema may differ. The stats transfer between processes can produce corrupted or missing values. Watch for anomalous stats after upgrades, because the stat schema can change across versions.

Envoy transfers stats between hot-restart processes as protobuf messages over the Unix domain socket rather than via a fixed-size shared memory region (change merged in PR #5910, shipped in Envoy 1.11.0). This eliminated the shared-memory size mismatch crash that occurred when upgrading across major versions. However, schema differences between versions can still produce anomalous values during the transfer.

The safest approach for major version upgrades is to skip hot restart entirely and perform a cold restart. For minor version upgrades within the same hot restart compatibility version (check --hot-restart-version), hot restart should be safe. Monitor stats output for anomalies after any upgrade that crosses a version boundary.

--skip-hot-restart-parent-stats disables stats import from the parent process entirely. This prevents corruption but means counters reset to zero on each restart, losing continuity.

Note: a cold restart (SIGTERM or /quitquitquit) does not perform graceful draining the way hot restart does. TCP resets occur on shutdown rather than the orderly Connection: close and GOAWAY sequence. Plan for client-visible disruption if you choose this path.

Prevention

  • Rate-limit restart triggers. The hot restart protocol races at very short intervals. Ensure your restart mechanism (SIGHUP handler, deployment script, supervisor) does not trigger more than one restart per second, and ideally no more than one per several seconds.
  • Size FD limits for the dual-process window. Calculate your limit as roughly baseline_peak_FD * 2.5 to account for the doubling during restart plus headroom.
  • Verify hot restart compatibility before upgrading. Compare --hot-restart-version of the new binary against the running process via the admin endpoint. If they differ, plan a cold restart.
  • Keep --parent-shutdown-time-s larger than --drain-time-s. The parent must survive long enough for the new process to reach LIVE. Default values (600s drain, 900s parent shutdown) provide a 300-second buffer. If your xDS initialization is slow, increase the parent shutdown time.
  • Use --skip-hot-restart-on-no-parent in unstable environments. If the parent process is prone to being killed (short Kubernetes termination grace periods, aggressive OOM killer), this flag prevents the child from terminating when the parent disappears.
  • Monitor stats after version upgrades. Compare key gauges and counters before and after any binary upgrade. If values look wrong, restart cold to reset the stats subsystem.
  • In Istio sidecar mode, do not send SIGHUP to Envoy. Istio starts Envoy with --disable-hot-restart. SIGHUP is not a supported restart mechanism in this mode. Use the pod lifecycle instead.

How Netdata helps

  • Per-second server.hot_restart_epoch collection detects crash loops within seconds of the first failed epoch, before the loop compounds.
  • Correlating server.state transitions with connection count changes across the restart window identifies whether drops are from the drain sequence, FD exhaustion, or a gap between processes.
  • server.parent_connections tracking distinguishes a stuck drain from a failed handoff by showing whether the old process is still draining or has exited prematurely.
  • FD utilization monitoring against the configured limit catches the dual-process doubling before it causes silent connection refusal.
  • Anomaly detection on stats values after a version upgrade surfaces corrupted or impossible gauge values that would otherwise go unnoticed until they trigger a false alert elsewhere.
  • Combining server.uptime resets with epoch increments in a single timeline distinguishes a crash-loop-via-hot-restart from a planned rolling deployment.