The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nginx / nginx-monitoring-maturity-model

Operations Guides

NGINX monitoring maturity model: from survival to expert

nginx exposes exactly seven scalars through stub_status. Latency distributions, upstream health, cache efficiency, and kernel-level drops live in access logs, error logs, or OS counters. Teams that collect only the stub_status numbers assume they have visibility. They do not.

This article defines four monitoring maturity levels. Level 1 tells you if nginx is alive. Level 2 tells you if it is healthy. Level 3 gives you leading indicators of saturation. Level 4 exposes the blind spots that only appear after repeated incidents. Use these levels to audit your current coverage and decide which signals to add next.

The levels are cumulative. Do not build Level 4 dashboards until Level 1 is automated and paging. A team that cannot reliably detect a dead master or a 5xx spike will not benefit from tracking per-worker connection imbalance. Start at the bottom and move up.

flowchart TD
    L1["Level 1: Survival"] --> L2["Level 2: Operational"]
    L2 --> L3["Level 3: Mature"]
    L3 --> L4["Level 4: Expert"]

Level 1 — Survival

The absolute minimum. If any of these signals fail, users are already affected or nginx is down entirely.

  • Master and worker process liveness. The master reads configuration and spawns workers; workers handle all traffic. Verify the master with kill -0 $(cat /var/run/nginx.pid) and count workers via pgrep -c -P $(cat /var/run/nginx.pid). A dead master means no new workers will spawn. A worker crash loop is visible in dmesg as OOM kills or segfaults.
  • Listening port responsiveness. Confirm that nginx holds listening sockets with ss -tlnp | grep nginx. A successful TCP handshake only means the kernel accepted the connection into the backlog. If all workers are blocked, the HTTP request still stalls, so treat this as a necessary but insufficient check.
  • Active connections. The Active connections: line from stub_status is a point-in-time capacity gauge. It includes keepalive idle connections, so interpret it against the configured worker_connections limit.
  • HTTP 5xx response rate. Derived from access log $status. Sustained rates above 1 percent warrant investigation; above 5 percent with traffic indicate an active outage. Distinguish 502 (upstream down or refusing connections), 504 (upstream too slow), and 503 (rate limiting or no available upstreams).
  • Error log [emerg] and [crit] rate. Check recent lines with grep -cE '\[(emerg|alert|crit)\]'. A failed reload logs [emerg], but the previous configuration stays active. Correlate with process state and 5xx rate before paging.

Level 2 — Operational

A competent production team monitors throughput, latency, and resource saturation. These signals separate an internal tool from a public-facing service under load.

  • Requests per second. Derived from the cumulative requests counter in stub_status. A sudden drop indicates upstream failure or network partition; a spike may signal abuse or a retry storm from a client.
  • Connection state breakdown. stub_status splits active connections into Reading, Writing, and Waiting. High Writing with low request rate indicates slow upstreams or slow clients. Sustained Reading above 20 percent of active connections suggests a slowloris-style attack or pathological network degradation.
  • Dropped connections (accepts - handled delta). If the cumulative accepts counter exceeds handled, nginx is dropping connections. This happens when worker_connections or file descriptor limits are reached. Track the rate of gap growth, not the absolute cumulative difference.
  • Latency percentiles. Log $request_time and $upstream_response_time to compute p50, p95, and p99. $request_time includes time spent sending the response to the client; a slow mobile client can inflate it even when upstream is fast. Always compare it with $upstream_response_time to isolate backend latency.
  • HTTP status code distribution. Break out 4xx and 5xx. Monitor 499 specifically; it is an nginx-specific code meaning the client closed the connection before the response completed. A 499 spike is the canary for user impatience before official timeouts trigger 5xx.
  • File descriptor utilization per worker. Each client connection, upstream socket, log file, and temp file consumes an FD. Check the count under /proc/$pid/fd against the Max open files limit in /proc/$pid/limits. Default OS limits are often 1024, which is dangerously low for a reverse proxy where each request can consume multiple handles.
  • Worker CPU and RSS. Per-worker CPU via ps -o %cpu= and RSS via /proc/$pid/status VmRSS. One worker at 100 percent while others idle indicates load imbalance, often from reuseport hash collisions when traffic comes from a small number of source IPs.
  • Error log rate by severity. Baseline the rate of [error] and [warn] messages. During an incident, error log volume can spike 100-1000x, which itself generates disk I/O pressure that slows the event loop.

Level 3 — Mature

Full coverage with leading indicators and upstream visibility. At this level you stop reacting to outages and start catching saturation before it drops connections.

  • Upstream connect and header time. $upstream_connect_time reveals TCP and TLS handshake cost. A value of 0.000 indicates keepalive pool reuse. $upstream_header_time isolates backend processing time from response body transfer. High connect time with low header time means the network or pool is the problem, not the application.
  • Per-upstream-server metrics. Parse $upstream_addr from access logs to aggregate latency and error rates per backend. One degraded server can hide inside a healthy average. Open-source nginx has no native per-upstream API, so access log parsing is the only option.
  • Cache hit rate and status distribution. Log $upstream_cache_status to track HIT, MISS, STALE, EXPIRED, and BYPASS. A sudden STALE spike can mask an upstream outage; a sudden MISS spike may indicate a cache purge, zone exhaustion, or mass TTL expiration.
  • Listen socket backlog and TcpExtListenOverflows. Use ss -tlnp to read Recv-Q (current queue depth) and Send-Q (configured backlog). Monitor nstat -az TcpExtListenOverflows; any nonzero increasing rate means the kernel is dropping connections before nginx sees them. This produces zero evidence in nginx logs.
  • Worker connection slot utilization. Calculate active_connections / (worker_connections * worker_processes). The default worker_connections is 512, not 1024. For reverse proxy, effective capacity is roughly half because each request uses two connection slots. Account for keepalive idle connections in your headroom.
  • Reload frequency and success. Track reconfiguring messages and configuration file .* test failed in error logs. A failed reload leaves the old configuration active, creating silent configuration drift. Frequent reloads without worker_shutdown_timeout cause old workers to accumulate on long-lived connections.
  • SSL session cache hit rate. Log $ssl_session_reused; "r" means the session was reused. Low hit rates drive CPU saturation from full TLS handshakes. Size the shared cache for connection rate multiplied by ssl_session_timeout.
  • Rate limit rejection rate. The default rejection status is 503, not 429. Monitor error log entries containing limiting requests and verify that legitimate traffic is not blocked.
  • Request and response size distributions. Log $request_length and $body_bytes_sent to detect shifts that affect buffer usage, compression CPU, and bandwidth.
  • Disk space on log partitions. Under incident conditions, log volume growth can fill the partition, causing writes to block the event loop and increasing latency for all requests.

Level 4 — Expert

These signals are added after the third or fourth major incident reveals a blind spot. They explain why aggregate metrics looked green while users suffered.

  • Per-worker load imbalance. Compare CPU and active connection counts across individual workers. With reuseport, a small number of client IPs can hash to a single worker, creating a hot spot that aggregate metrics miss.
  • Kernel conntrack table utilization. If the host runs iptables or nftables with connection tracking, compare conntrack -C against /proc/sys/net/netfilter/nf_conntrack_max. A full table drops packets silently and mimics upstream connectivity failures.
  • Per-location-block latency and error rate. Use custom log variables or separate access logs per location block to isolate which routes degrade. Open-source nginx provides no native per-location metrics.
  • Rate limit zone effective capacity. Estimate unique keys against zone size; limit_req_zone entries consume roughly 128 bytes each. Zone exhaustion produces could not allocate node errors and silently disables rate limiting for new clients.
  • Old worker accumulation after reloads. Count total worker processes against worker_processes. Old workers from previous reloads linger indefinitely on WebSocket, gRPC streaming, or long-polling connections unless worker_shutdown_timeout is set.
  • Temp file creation rate. When proxy or client body buffers overflow, nginx writes to disk. There is no native metric for this. Watch for persistent gaps between $request_time and $upstream_response_time, or monitor disk I/O on the mount point used by proxy_temp_path.
  • DNS resolution latency. For variable-based proxy_pass with the resolver directive, resolution failures async-wait for resolver_timeout (default 30s) and then return 502. Log resolver errors separately from upstream errors.
  • $upstream_status versus $status discrepancy. An upstream may return 500 while nginx serves 200 from stale cache or a custom error_page. Logging both reveals when failures are masked by nginx-layer resilience.
  • TIME_WAIT and ephemeral port utilization. Without effective keepalive reuse, every proxied request opens and closes a TCP connection to upstream. Check ss -tan state time-wait per destination to detect port exhaustion before connect() failed (99) appears.
  • TLS version and cipher distribution shifts. Log $ssl_protocol and $ssl_cipher to detect downgrade attacks or client population changes that affect CPU load.
  • Stale cache serving rate. A sustained high STALE rate with no corresponding MISS spike suggests upstream is down and proxy_cache_use_stale is hiding the failure from standard error metrics. Verify upstream health independently before celebrating the resilience.

How Netdata helps

Netdata automates the collection and correlation of signals across all four levels on a single node.

  • It scrapes stub_status and derives requests per second, connection state breakdown, and dropped connections without manual log parsing.
  • It correlates nginx access log metrics, including latency percentiles, response codes, and upstream times, with OS-level signals like TcpExtListenOverflows, per-worker CPU, and file descriptor usage.
  • It tracks per-process file descriptor consumption alongside configured worker_connections, surfacing when the OS limit becomes the real bottleneck before connection slots max out.
  • It monitors error log severity rates and surfaces spikes alongside upstream status code changes or cache status anomalies.
  • It visualizes per-worker CPU and memory RSS imbalance, helping identify reuseport hot spots or stuck old workers after frequent reloads.
The Netdata solution

Web server monitoring with Netdata

Netdata monitors NGINX with per-second request, connection, and latency metrics plus ML anomaly detection. Correlate connection and file-descriptor exhaustion, upstream cascade failures, buffer spill, and TLS CPU with the host signals behind them.