The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nginx / nginx-limit-req-burst-tuning

Operations Guides

NGINX limit_req burst and nodelay tuning: rate limiting without blocking real users

Most production nginx rate limiting configs fall into two camps: no burst at all, which rejects legitimate traffic during harmless spikes, or burst without nodelay, which queues real users into artificial delays that mimic upstream slowness. Neither is what you want. The limit_req module implements a leaky bucket at millisecond granularity, and the interaction between burst and nodelay determines whether a request is delayed, rejected, or forwarded immediately. Understanding that interaction, and sizing the shared memory zone to match your traffic profile, is the difference between rate limiting that protects upstreams and rate limiting that creates incidents during normal user behavior.

What it is and why it matters

limit_req_zone defines a shared memory zone and a sustained request rate. A matching limit_req directive inside a location, server, or http block enforces it. Without both pieces, nothing happens. The zone tracks state per key, typically $binary_remote_addr, and uses a leaky bucket to decide the fate of each request.

The default configuration, limit_req zone=one with no burst argument, rejects any request that arrives before the bucket has leaked enough capacity. At a rate of 10r/s, the bucket leaks one slot every 100 milliseconds. A second request arriving 10 ms after the first is rejected with a 503. This is correct for brute-force protection but catastrophic for legitimate browser bursts, API batching, or password managers that submit multiple forms rapidly.

Adding burst changes the behavior from immediate rejection to queuing, but without nodelay that queue translates into artificial latency. The Nth excess request in a burst=20 queue at 10r/s waits up to two seconds before nginx forwards it to the upstream. If your upstream response time is 50 ms but nginx delays the request by two seconds, users experience a timeout. Clients that time out generate 499s; retries amplify load.

The goal is absorbing short legitimate bursts without delaying them, while still enforcing a hard ceiling against sustained abuse. That is what nodelay does, and why the zone size matters more than most operators assume.

How it works

The leaky bucket operates at millisecond granularity. A rate of 10r/s means one request every 100 ms, not an average of ten requests over a one-second window. When a request arrives, nginx checks whether the bucket has accumulated at least one slot since the last request. If yes, the request passes immediately. If not, nginx looks at the burst parameter.

Without burst, the request is rejected immediately. The default rejection status is 503, unless you override it with limit_req_status 429.

With burst=N alone, excess requests enter a FIFO queue. They are released to the upstream at the configured rate, one by one. The queue length is N. The Nth excess request in a burst queue waits up to N multiplied by (1/rate) before reaching the upstream. This queuing happens inside nginx; the upstream sees none of the delay until the request is finally forwarded. During that wait, the client connection is held open. If the client timeout is shorter than the queue delay, the connection closes and nginx logs a 499.

With burst=N nodelay, nginx allocates a pool of N burst slots per key. When a request arrives above the base rate, nginx immediately forwards it and marks one slot as consumed. No queuing delay is introduced. Slots are freed back into the pool at the configured rate. As long as a free slot exists, the request passes through without delay. When all N slots are occupied, subsequent requests are rejected immediately. This preserves the hard rate cap while eliminating artificial latency for bursty traffic.

The delay parameter offers a middle ground. A configuration like limit_req zone=one burst=12 delay=8 forwards the first eight burst requests without delay, then enforces the rate limit for the remaining four before the hard rejection threshold. It is useful when you want to allow a small spike but still apply backpressure before the full burst ceiling.

nodelay is semantically equivalent to delay=infinity while still respecting the burst ceiling.

Burst is per-key and per-zone, not per-location. If two location blocks reference the same limit_req_zone, they share the same burst pool for each key. A burst of ten consumed by one location leaves zero available for the other. Order of arrival determines allocation.

Zone sizing determines how many unique keys nginx can track. On 64-bit systems, each entry consumes roughly 128 bytes. A 10 MB zone tracks approximately 80,000 unique IPs. Most guides quote the 32-bit figure, which is roughly double. If you size the zone for 16,000 IPs per megabyte on a 64-bit production server, you will exhaust the zone.

When a limit_req_zone fills, nginx evicts the least recently used entries to make room for new keys. If it cannot free sufficient space, the request receives a 503. This means a sudden influx of previously unseen IPs can evict tracked keys and cause collateral damage to existing clients.

flowchart TD
    A[Request arrives] --> B{Key exists in zone?}
    B -->|No| C[Allocate node]
    C -->|Zone full| D[LRU evict oldest]
    D -->|Still full| E[Reject 503]
    B -->|Yes| F[Check bucket]
    F -->|Within rate| G[Forward immediately]
    F -->|Over rate| H{Burst slots free?}
    H -->|Yes| I{nodelay set?}
    I -->|Yes| J[Forward now
Consume slot] I -->|No| K[Queue request
Release at rate] H -->|No| L[Reject 503] J --> M[Free slot at leak rate] K --> M

Where it shows up in production

API endpoints are the most common site of misconfiguration. A mobile app that batches three requests on launch, or a single-page application that fetches data in parallel, can trigger a burst limit instantly. Without nodelay, those requests queue and the app times out.

Authentication endpoints suffer the same problem. Password managers and security-conscious users may submit login forms multiple times in quick succession. Immediate rejection teaches them the site is broken; queuing without nodelay teaches them the site is slow.

Browser behavior also matters. Modern browsers open six or more parallel connections per host to load assets. If you apply limit_req to static asset locations without nodelay, the browser stalls waiting for CSS or JavaScript while the bucket leaks.

Shared zones across locations create hidden coupling. If /api and /webhook both reference the same zone, a burst of webhooks can consume the entire burst quota and cause API requests to be rejected. Operators often assume each location gets its own burst pool.

NAT and proxy aggregation silently break per-IP limits. A corporate gateway or a caching proxy may represent dozens of legitimate users with a single IP address. Per-IP rate limiting keyed to $binary_remote_addr treats all of those users as one entity. When the shared zone is undersized, the LRU eviction amplifies the problem by rotating tracked IPs under pressure, destabilizing limits for everyone.

Tradeoffs and common misuses

Burst without nodelay is appropriate only when you want absolute traffic smoothing. The queue enforces a perfectly uniform rate to the upstream, but it adds latency under load. That latency can cascade into client timeouts, 499 errors, and retry storms that increase load instead of reducing it.

Burst with nodelay eliminates artificial delay but enforces a hard cap at rate plus burst. Once the burst pool is exhausted, requests are rejected immediately. This is the right choice for most web applications because it absorbs legitimate spikes without punishing the user, while still protecting the upstream from sustained overload.

Delay parameter is useful when you want to allow a small unconditional burst but still apply backpressure before the ceiling. It sits between the two extremes.

Leaving the default 503 status mixes rate limiting with upstream outage signals. The default limit_req_status is 503. If your alerting treats 503 as “all backends are down,” rate-limited legitimate users will trigger a false incident page. Set limit_req_status 429 to separate capacity enforcement from backend health.

Undersizing the zone is the most dangerous misconfiguration. When limit_req_zone fills, LRU eviction begins. A flash crowd of new IPs can evict established tracking keys, causing nginx to reject requests from clients that were previously within their limit. If eviction cannot free enough space, enforcement fails entirely and requests receive 503. The error log shows could not allocate node, but by then the limiter has already stopped protecting you.

Assuming per-location isolation leads to quota starvation. Because burst is per-key across all locations sharing a zone, one high-traffic path can steal the burst allowance from another.

Signals to watch in production

SignalWhy it mattersWarning sign
Rate limit rejection rateMeasures enforcement activitySudden spike indicates attack, flash crowd, or limits set too low
$limit_req_statusDistinguishes PASSED, DELAYED, and REJECTEDHigh DELAYED means burst is queuing without nodelay; unexpected REJECTED means limits are too tight
Zone allocation errorsZone exhaustion breaks enforcement for new keysAny could not allocate node in the error log is critical
Unique key count vs zone sizeProactive capacity planningUnique clients approaching 80% of estimated zone capacity
Request time minus upstream response timeDetects artificial queuing delayGap growing under load indicates burst without nodelay

How Netdata helps

  • Correlate 503 or 429 spikes with stub_status request rate and connection state breakdown to determine whether rate limiting or upstream failure is the source.
  • Monitor the nginx error log for limiting requests and could not allocate node to catch zone exhaustion before it disables enforcement.
  • Track the gap between $request_time and $upstream_response_time to detect artificial delay introduced by queued bursts.
  • Alert on 5xx rate anomalies that coincide with rate limit zone saturation, separating legitimate enforcement from unintended rejection of real users.
The Netdata solution

Web server monitoring with Netdata

Netdata monitors NGINX with per-second request, connection, and latency metrics plus ML anomaly detection. Correlate connection and file-descriptor exhaustion, upstream cascade failures, buffer spill, and TLS CPU with the host signals behind them.