The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-how-it-works-in-production ▌

Operations Guides

How Envoy actually works in production: a mental model for operators

Envoy is a multi-threaded, event-driven L4/L7 proxy written in C++. Its architecture directly shapes what you see in metrics, access logs, and user-visible behavior during incidents. If you do not know that worker threads share nothing in the hot path, aggregate CPU utilization will mislead you. If you do not know that circuit breaker 503s are indistinguishable from upstream-generated 503s at the counter level, you will blame the wrong component. If you do not know that Envoy keeps serving traffic on stale configuration after an xDS disconnect, you will miss the slow-burn failure that surfaces hours later.

This article covers the abstractions operators need before debugging: the threading model, the request path through filter chains, per-worker connection pools, and the self-protection mechanisms (circuit breakers, outlier detection, overload manager). Once these are internalized, the failure patterns and the metrics that expose them become predictable.

Threading model

The threading model is the single most important thing to understand about Envoy. It determines how CPU is consumed, how connections are owned, how configuration reaches workers, and why certain metrics must be read at per-worker granularity rather than in aggregate.

Envoy uses three thread categories:

  • Main thread: Process lifecycle, xDS communication with the control plane, stats flushing, the admin interface, active health checking, and hot restart coordination. It does not process data plane traffic. A slow stats flush or a busy admin endpoint does not compete with request processing.
  • Worker threads: Each worker runs its own libevent event loop and accepts downstream connections via SO_REUSEPORT. Once a worker accepts a connection, it owns that connection for its entire lifetime. Workers share nothing in the hot path. This is how Envoy scales near-linearly, and it is why per-worker reasoning matters when diagnosing latency or saturation.
  • File flush thread: Asynchronous access log writes, keeping logging I/O off the data plane entirely.

The number of worker threads is controlled by --concurrency (default: number of hardware threads on the machine).

flowchart TD
    CP["xDS control plane
CDS, EDS, LDS, RDS, SDS"] MT["Main thread
xDS client, stats flush,
admin, health checks, hot restart"] WT["Worker threads
Each: libevent loop, thread-local
cluster mgr, own conn pools"] FF["File flush thread
async access log writes"] DC["Downstream clients"] US["Upstream hosts"] CP -->|"gRPC or REST"| MT MT -->|"RCU pointer swaps"| WT DC -->|"SO_REUSEPORT
worker owns conn for life"| WT WT -->|"upstream requests"| US WT -->|"access logs"| FF

Request path and self-protection

The request path

A downstream connection arrives at a listener socket. The kernel balances accepts across workers via SO_REUSEPORT. One worker wins and owns the connection for its lifetime. The connection then passes through:

  1. Listener filters (L4): Optional pre-processing such as TLS inspector to detect SNI before routing decisions.
  2. Network filter chain (L4): Includes the transport socket for TLS termination. The HTTP connection manager (HCM) is the terminal network filter for HTTP traffic.
  3. HTTP connection manager: Decodes the protocol (HTTP/1.1, HTTP/2, HTTP/3) and manages the downstream stream lifecycle.
  4. HTTP filter chain (L7): Sequential processing through configured filters: external authorization, rate limiting, Lua or Wasm extensions, compression, and finally the router filter.
  5. Router filter: Resolves the route against the route table, selects the target upstream cluster, applies load balancing policy, and initiates the upstream request.
  6. Connection pool: The router requests a connection from the per-worker, per-host, per-protocol pool. If no connection is available and the pool has not reached its limit, a new one is established. Otherwise the request queues.
  7. Upstream: The request is sent to the upstream host. The response flows back through the filter chain in reverse order.

Thread-local cluster manager and connection pools

Each worker maintains its own copy of cluster state through the thread-local cluster manager: connection pools, load balancing weights, and host health. The main thread receives configuration updates via xDS and distributes them to workers via RCU-style pointer swaps through thread-local storage slots. Configuration changes are lock-free in the hot path but eventually consistent across workers.

Connection pools are per-worker, per-upstream-host, and per-protocol (HTTP/1.1, HTTP/2, HTTP/3, TCP). A cluster with 10 hosts across 8 workers has up to 80 independent connection pools. Per-cluster aggregate stats can hide per-host or per-worker hotspots behind healthy averages.

HTTP/2 and HTTP/3 connection pools multiplex multiple concurrent streams over fewer TCP connections. upstream_cx_active can be low while request concurrency is high. Connection counts are not a throughput proxy for multiplexed protocols. Track upstream_rq_active for concurrent in-flight requests instead.

Circuit breakers

Circuit breakers are configured per-cluster and per-priority (default and high). They cap maximum connections, pending requests, concurrent requests for multiplexed protocols, and retries. When a limit is exceeded, Envoy fast-fails the request locally with a 503 and response flag UO. The request never reaches the upstream.

This is a deliberate protection mechanism, not a bug. Envoy sheds load early to prevent cascading failure. Workers share circuit breaker state via eventually consistent counters, so brief races between threads can allow limits to be marginally exceeded.

Check which circuit breakers are open via the stats endpoint:

curl -s localhost:9901/stats | grep circuit_breakers

Outlier detection vs active health checks

These are independent systems with different signal sources:

  • Active health checks run on the main thread. They send synthetic probes to upstream hosts at configured intervals. A host that fails health checks is removed from load balancing rotation.
  • Outlier detection is passive. It observes real traffic outcomes and ejects hosts based on consecutive 5xx errors, success rate thresholds, or failure percentage.

A host can pass active health checks but be ejected by outlier detection because real traffic is failing. Conversely, a newly failed host that has not yet received traffic will not be caught by outlier detection until health checks mark it unhealthy. Both systems contribute to membership_healthy, membership_degraded, and membership_excluded gauges.

When the healthy host percentage drops below the panic threshold (default 50%), Envoy enters panic mode and load-balances across all hosts including unhealthy ones. This is intentional. The alternative is concentrating all traffic on the few remaining healthy hosts until they also fail.

Overload manager

The overload manager monitors Envoy’s own resource consumption, primarily heap size. When configured thresholds are crossed, it triggers actions: stop accepting connections, stop accepting requests, disable HTTP keepalive, reduce timeouts, shrink heap.

If the overload manager is not configured, Envoy has no self-protection against memory exhaustion. It will consume memory until the container OOM-kills it. The overload manager depends on max_heap_size_bytes being set correctly. If it is not set or set too high, the monitor never fires.

Deployment variants and what they change

The deployment variant changes which signals matter and what “normal” looks like:

DeploymentWhat changesFocus areas
Sidecar (Istio or service mesh)Thousands of instances. Admin port typically 15000, not 9901. Per-pod resource limits interact directly with overload manager.xDS control plane load, fleet-wide correlation, per-pod memory pressure
Edge or gatewayFewer instances, higher per-instance connection counts. TLS termination dominates CPU.Connection management, TLS handshake rates, rate limiting
Front proxy (non-mesh)Often static configuration with no xDS dependency.Upstream health, throughput, keepalive tuning
HTTP/2 or gRPC upstreamsLong-lived streams. A single HTTP/2 connection can carry hundreds of concurrent streams. Stream-level contention invisible in connection metrics.upstream_rq_active, stream-level latency percentiles
HTTP/1.1 upstreamsEach request needs its own connection unless keepalive is configured. Connection churn adds repeated TCP and TLS handshake latency.Connection reuse rates, upstream_cx_connect_fail, keepalive timeouts

Failure modes built into the design

The same architecture that makes Envoy fast creates specific, predictable failure patterns.

Connection pool exhaustion cascade. Upstream slows down. Connections are held longer. The pool fills. Pending requests queue. The queue overflows. Envoy returns 503 with flag UO. The system looks broken, but it is protecting the upstream from additional load.

Memory pressure spiral. Large request or response bodies combined with buffering filters cause heap growth. If the overload manager is configured, it degrades service to survive. If it is not, the process is OOM-killed.

xDS stale configuration. The control plane disconnects. Envoy keeps serving with last-known-good config. Existing traffic works. New endpoints, route changes, and certificate rotations are invisible. The failure surfaces hours later when stale endpoints receive traffic or new services return 503 with flag NR. Check control_plane.connected_state and compare version_info in the config dump against what the control plane reports as current.

Retry amplification. Aggressive retry policy meets partial upstream failure. Retries add load to the already-degraded upstream. More failures trigger more retries. The upstream collapses under 2-3x normal load, accelerated by the mechanism designed to mask the failure.

Stats cardinality explosion. Dynamic route names or per-request metadata in stat tags fill the shared-memory stats region. New stats silently fail to register. Monitoring develops blind spots with no error or warning.

Hot worker. One worker thread saturates at 100% CPU (expensive filter, TLS handshake storm, lock contention) while others idle. Requests assigned to that worker experience high latency. Aggregate process CPU looks moderate. Only per-thread analysis or watchdog_miss counters expose the imbalance.

Outlier detection mass ejection. Aggressive outlier detection settings eject hosts during a correlated failure (network issue, shared dependency). Remaining hosts overload and get ejected. The panic threshold triggers. Envoy routes to all hosts including unhealthy ones by design.

Hot restart FD exhaustion. The new process starts before the old one finishes draining. Both processes hold file descriptors simultaneously. FD usage briefly doubles. If baseline FD usage is above 50% of ulimit -n, hot restart can trigger FD exhaustion.

Signals to watch

SignalWhy it mattersWarning sign
server.stateProcess lifecycle state (LIVE, DRAINING, INITIALIZING)Non-LIVE during steady-state operation
cluster.<name>.membership_healthy / membership_totalPer-cluster availability ratioDropping below 50% triggers panic mode
cluster.<name>.upstream_rq_pending_activeLeading indicator before 503s from overflowSustained nonzero value
cluster.<name>.circuit_breakers.<priority>.*_openSelf-protection mechanism activeAny gauge transitioning to 1
upstream_rq_total vs downstream_rq_total ratioRetry amplification detectionRatio above 1.5 sustained
server.watchdog_missWorker thread blocked past watchdog timeoutAny nonzero value
control_plane.connected_statexDS connectivity, stale config risk0 sustained beyond reconnect interval
cluster.<name>.outlier_detection.ejections_activePassive health ejections from real trafficApproaching membership_total

How Netdata helps

Netdata’s Envoy collector scrapes the admin stats endpoint at per-second granularity. The operational value is in correlation across metrics that aggregate dashboards miss:

  • Watch upstream_rq_pending_active and circuit_breakers.*_open alongside upstream 503 counters to separate circuit breaker rejections from upstream errors in real time.
  • Correlate server.memory_allocated with overload manager action gauges to confirm whether Envoy is shedding load due to its own memory pressure or listener limits.
  • Track the upstream_rq_total to downstream_rq_total ratio per cluster to surface retry amplification hidden by aggregate error rates.
  • Netdata’s anomaly detection on per-cluster latency histograms and membership ratios catches gradual degradation before fixed thresholds fire.
  • server.state, control_plane.connected_state, and cluster_manager.warming_clusters together distinguish planned drain from stuck initialization during hot restarts and xDS reconnects.