The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-listen-queue-full ▌

Operations Guides

uWSGI listen queue full: the backlog overflow that drops connections silently

Clients report connection timeouts or resets. Your load balancer shows 502s or 504s on requests to the uWSGI backend. The uWSGI master is running, workers are alive, the stats server responds, and application logs show no errors. This is listen queue overflow: the kernel accept backlog on the uWSGI listening socket is full, and the kernel is silently dropping new connections before accept().

The listen queue is the kernel-level buffer between completed TCP handshakes and available workers. When all workers are busy and this buffer fills, the kernel silently drops incoming connections. No uWSGI log entry is written. No uWSGI error counter increments. The stats server runs in the master process and stays responsive during complete worker starvation, so health checks that query the stats endpoint still pass. Health check endpoints handled by workers may also pass if they are lightweight enough to be accepted between drops. Meanwhile, real users cannot connect.

What this means

The uWSGI listen queue is the kernel-level accept() backlog on the socket the master process binds. When a connection arrives, the kernel completes the TCP handshake and places the connection in this queue. When a worker calls accept(), it pulls the next connection and processes it. If no worker calls accept() because all workers are busy, connections accumulate in the queue. Once the queue reaches its maximum size, the kernel drops new connections.

The critical problem is observability. uWSGI’s own listen_queue stats field is broken on standard Linux. For TCP sockets, it relies on TCP_INFO behavior that varies across kernel versions. For UNIX sockets, it requires a non-standard kernel ioctl. The load field in the stats JSON is identical to listen_queue (not average latency, despite its name). The listen_queue_errors field exists in the JSON output but is dead code: it is never incremented anywhere in the uWSGI source. All three fields almost always read 0 regardless of actual backlog.

The only reliable measurements come from outside uWSGI:

  • ss -ltn shows Recv-Q (current queue depth) and Send-Q (the configured backlog limit) for each listening socket.
  • TcpExtListenOverflows from /proc/net/netstat counts the number of times the kernel dropped a connection because the accept queue was full. This counter is system-wide, not per-socket.
flowchart TD
    A[New connection arrives] --> B{Worker idle?}
    B -- yes --> C[accept and process]
    B -- no --> D[Queue in kernel backlog]
    D --> E{Backlog full?}
    E -- no --> F[Wait for a worker]
    F --> B
    E -- yes --> G[Kernel drops connection]
    G --> H[Client sees timeout or RST]
    G --> I[No uWSGI log entry]
    G --> J[TcpExtListenOverflows increments]

When the backlog is full, the kernel does not warn uWSGI. With default kernel settings (net.ipv4.tcp_abort_on_overflow=0), the final ACK of the TCP handshake is silently dropped. The client sees a connection timeout, not an explicit error. If tcp_abort_on_overflow is set to 1, the kernel sends a RST instead and the client sees a connection reset. In both cases, uWSGI logs nothing because the connection never reached accept().

On some configurations, uWSGI may log the message *** uWSGI listen queue of socket "..." (fd: X) full !!! (N/M) *** where N is the current queue depth and M is the configured max. The message is emitted only in verbose mode (--verbose), and only when the master’s periodic queue check (roughly once per second) finds the queue full, so it is not emitted for every dropped connection and may not appear at all depending on timing and kernel measurement.

Common causes

CauseWhat it looks likeFirst thing to check
Worker pool starvationAll workers in busy status, throughput collapsing, avg_rt risingWorker busy ratio and per-worker uri in stats
Backlog too smallRecv-Q hits Send-Q ceiling during brief traffic spikes, drops clear when traffic subsides--listen value and net.core.somaxconn
Stuck workers without harakiriOne or more workers permanently busy, request count frozen, no harakiri configuredPer-worker request count delta and harakiri config
CLOSE-WAIT socket accumulationFile descriptor count climbing, connections piling up in CLOSE-WAIT statess -tan state close-wait or lsof on the uWSGI process
Graceful reload capacity gapThroughput drops to zero during deployment, queue fills before new workers are readyDeployment timestamps correlated with queue depth

Quick checks

Run these read-only commands during the incident. They are safe and non-disruptive.

# Check current accept queue depth and backlog limit (TCP socket)
# Recv-Q = current queue depth, Send-Q = configured backlog
# Replace :8000 with your uWSGI port
ss -ltn 'sport = :8000'

# For UNIX socket
ss -lxn | grep uwsgi

# Check kernel-level listen queue overflow counter (system-wide, not per-socket)
nstat -az TcpExtListenOverflows

# Check the kernel's somaxconn cap
cat /proc/sys/net/core/somaxconn

# Check worker busy ratio from uWSGI stats
uwsgi --connect-and-read 127.0.0.1:9191 | jq '([.workers[] | select(.status == "busy")] | length) as $busy | ([.workers[] | select(.pid > 0 and .status != "cheap")] | length) as $alive | if $alive > 0 then ($busy / $alive * 100) else 0 end'

# Check accepting worker count (alive, not cheaped, accepting connections)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | select(.pid > 0 and .status != "cheap" and .accepting == 1)] | length'

# Check what busy workers are doing (which URIs are stuck)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.status == "busy") | {id, uri, avg_rt}'

# Check for CLOSE-WAIT sockets (system-wide count; filter by PID with -p, requires root)
ss -tan state close-wait | wc -l

Do not rely on the listen_queue and listen_queue_errors fields in uWSGI stats. Both are unreliable or dead on standard Linux and almost always read 0 regardless of actual backlog. Use ss and nstat instead.

How to diagnose it

  1. Confirm the kernel is dropping connections. Run nstat -az TcpExtListenOverflows twice, a few seconds apart. If the counter is rising, the kernel is actively dropping connections because an accept queue is full somewhere on the host. This counter is system-wide, so if multiple services share the host, correlate with per-socket queue depth to attribute the drops to uWSGI.

  2. Check the socket queue depth. Run ss -ltn 'sport = :PORT' (for TCP) or ss -lxn (for UNIX sockets). If Recv-Q is non-zero and sustained, workers are not accepting fast enough. If Recv-Q equals Send-Q, the queue is full and connections are being dropped right now.

  3. Check worker saturation. Pull the stats JSON and compute the busy ratio. If 100% of non-cheaped workers are busy, the queue is filling because no worker can call accept(). Check which URIs the busy workers are processing to identify the slow endpoint.

  4. Check the effective backlog limit. On Linux, uWSGI reads net.core.somaxconn at bind time and refuses to start if --listen exceeds it, logging Listen queue size is greater than the system max net.core.somaxconn (N). Keep the two aligned: raise somaxconn first, then set --listen. (The kernel itself clamps the listen() backlog to somaxconn; it does not round the accept backlog to a power of two — that rounding historically applied to the SYN queue table, not the accept queue.)

  5. Check for stuck workers. If harakiri is not configured, workers that hang on blocking I/O stay stuck forever. Check whether any worker’s request count has stopped increasing while its status remains busy. Each stuck worker permanently reduces capacity by one slot.

  6. Rule out the reload window. If the incident correlates with a deployment, the graceful reload may have drained all old workers before new workers finished loading the application. Check whether last_spawn timestamps in the stats JSON align with the outage window.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Socket Recv-Q (ss -ltn)Current accept queue depth, measured externallySustained non-zero means workers cannot keep up
Socket Send-Q (ss -ltn)The effective backlog limit on the socketShould match your --listen value (clamped by somaxconn)
TcpExtListenOverflows (nstat)Kernel drop counter for accept queue overflowAny non-zero rate means connections are being dropped now
Worker busy ratio (uWSGI stats)How close the pool is to exhaustionSustained 100% means the queue is filling
Accepting worker count (uWSGI stats)Workers actively able to call accept()Approaching zero means imminent queue fill
avg_rt per worker (uWSGI stats)Application latency trend (exponential moving average, not cumulative)Rising trend approaching harakiri means workers are about to die
net.core.somaxconnKernel cap on backlogLower than --listen: uWSGI refuses to bind on Linux
CLOSE-WAIT socket countLeaked connections consuming file descriptor slotsGrowing count indicates uWSGI is not closing sockets after remote disconnect

Fixes

Increase the backlog

Raise --listen in the uWSGI configuration and ensure net.core.somaxconn is at least as high. The default --listen is 100. The default somaxconn varies: 128 on older distributions, 4096 on kernel 5.4 and newer. Both may be too low for production traffic.

Changing --listen requires restarting uWSGI; it is a bind-time setting, not reloadable.

[uwsgi]
listen = 1024
# Set somaxconn at runtime (non-persistent across reboots; add to sysctl.conf for persistence)
sysctl -w net.core.somaxconn=1024

somaxconn changes apply only to sockets created after the change. Restart uWSGI to pick up the new value. In containers, net.core.somaxconn may not be writable without specific capabilities depending on the runtime.

A larger backlog lets the kernel buffer brief traffic spikes without dropping connections. It does not fix the underlying problem if workers are consistently saturated. It buys time.

Tradeoff: a very large backlog can mask sustained capacity problems. Connections sit in the queue for seconds, latency increases, and clients time out before a worker picks them up. Size the backlog for burst absorption, not for chronic underprovisioning.

Add worker capacity

If the busy ratio is consistently above 80% during normal traffic, the worker pool is undersized. Increase --processes or adjust the cheaper subsystem range. The degradation curve is cliff-edge: below 100% utilization, latency is roughly flat. At 100%, latency jumps non-linearly because every additional request queues in the kernel backlog.

Before adding workers, check available memory. Each worker is a full process copy. Adding workers without memory headroom triggers swapping, which makes every worker slower and can worsen the saturation.

Enable harakiri

If workers are getting stuck on blocking calls and harakiri is not configured, enable it. Without harakiri, a single hung request permanently consumes a worker slot. Over time, stuck workers accumulate until the pool is exhausted and the listen queue fills.

[uwsgi]
harakiri = 30
harakiri-verbose = true

Set the timeout to 2-3x your expected maximum legitimate request duration. The harakiri-verbose option (Linux only) logs the killed worker’s blocked syscall and wait channel by reading /proc/<pid>/syscall and /proc/<pid>/wchan at kill time, alongside uWSGI’s standard in-flight request dump.

Harakiri does not fix the underlying hang. It kills the stuck worker and respawns it, temporarily restoring capacity. Use it as a safety net while you investigate the root cause.

Fix the slow endpoint

If all busy workers show the same URI, a single endpoint is the bottleneck. Common culprits: database queries without timeouts, external API calls without client-side timeouts, DNS resolution hanging, or lock contention. Check per-worker uri and running_time in the stats JSON.

As an emergency measure during an incident, consider blocking the slow endpoint at the load balancer to free workers for other traffic while you investigate.

Use chain reload

Standard graceful reload (SIGHUP) drains all old workers simultaneously. During the transition, capacity drops while new workers initialize and load the application. Use --chain-reload to cycle workers one at a time, maintaining partial capacity throughout the reload.

Prevention

  • Monitor TcpExtListenOverflows continuously. This kernel counter is the only reliable signal that connections are being dropped due to accept queue overflow. Alert on any non-zero rate of change. The counter is monotonic and system-wide; track the delta between polling intervals.
  • Monitor socket queue depth with ss. Alert on sustained non-zero Recv-Q on the uWSGI listening socket. This catches the problem before the queue fills and drops begin.
  • Do not trust uWSGI listen_queue, load, or listen_queue_errors stats fields. All three are unreliable or dead code on standard Linux. They almost always read 0 regardless of actual backlog.
  • Keep --listen and net.core.somaxconn aligned. On Linux, uWSGI refuses to start when --listen exceeds somaxconn. Check both after configuration changes and after container image rebuilds, where somaxconn defaults may differ from the host.
  • Track worker busy ratio as a leading indicator. Sustained busy ratio above 80% during normal traffic means the next traffic spike will fill the queue.
  • Configure harakiri. Stuck workers without a timeout permanently reduce capacity. The absence of harakiri is itself a monitoring blind spot because harakiri_count always reads 0.
  • Do not use the stats endpoint or master PID as a health check. The master and stats server remain responsive during complete worker starvation. Health checks must go through the worker pool to detect this failure mode.

How Netdata helps

Netdata surfaces the signals that matter for this failure mode and correlates them at per-second resolution:

  • TcpExtListenOverflows is collected from /proc/net/netstat, giving you a direct kernel-level count of accept queue drops without manual nstat polling.
  • Socket queue metrics (Recv-Q and Send-Q on listening sockets) are collected per-second, so you can see the queue filling before it overflows.
  • uWSGI worker metrics (busy ratio, accepting worker count, avg_rt, per-worker status) are collected directly from the stats server, letting you correlate worker saturation with kernel-level drops on the same timeline.
  • ML anomaly detection flags unusual changes in queue depth or worker busy ratio even when absolute values have not crossed a static threshold.
  • Correlation across layers (kernel TCP counters, uWSGI stats, application latency) on a single timeline shortens the diagnostic path from “clients see connection timeouts” to “workers are saturated on this endpoint.”