The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-listen-queue-stats-unreliable ▌

Operations Guides

uWSGI listen_queue always zero: why the stats field is broken on Linux

You open the uWSGI stats server JSON during a traffic spike. listen_queue reads 0. load reads 0. listen_queue_errors reads 0. But nginx is returning 502s, clients are seeing connection refused, and all workers are busy.

The fields do not measure what you think. On standard Linux, listen_queue is broken. load is identical to listen_queue (the source code has a TODO comment admitting this). listen_queue_errors is dead code that is never incremented. All three read 0 regardless of actual socket backlog pressure.

What this means

Three fields in the uWSGI stats JSON are not trustworthy on standard Linux:

FieldWhat it claims to beWhat it actually isTrustworthy?
listen_queueSocket backlog depthTCP tcpi_unacked (unacknowledged segments) or requires a non-standard ioctlNo
loadImplied latency or load metricIdentical to listen_queue, same valueNo
listen_queue_errorsQueue overflow countDead code, never incrementedNo

If you built alerts or dashboards on any of these fields, they will never fire and will always show zero. This is not a configuration problem. It is a known limitation in how uWSGI populates these fields on Linux.

The real accept queue depth is a kernel data structure not exposed through uWSGI’s measurement path. You need external tools to see it.

Why the field is broken on Linux

uWSGI’s master process attempts to measure the listen queue depth periodically. The implementation differs between TCP and UNIX sockets, and both paths have problems on standard Linux.

flowchart LR
    A["TCP listen socket"] --> B["uWSGI master check"]
    B --> C["getsockopt TCP_INFO"]
    C --> D["tcpi_unacked
unacked segments"] D --> E["listen_queue in stats
almost always 0"] A --> F["Kernel accept queue"] F --> G["NOT exposed via
TCP_INFO"] G --> H["Invisible to uWSGI"] F --> I["ss Recv-Q
actual depth"] F --> J["nstat TcpExtListenOverflows
drop counter"]

TCP sockets: tcpi_unacked is not the accept queue

For TCP listening sockets, uWSGI calls getsockopt(fd, IPPROTO_TCP, TCP_INFO, ...) and reads tcpi_unacked from the returned struct tcp_info, in the get_tcp_info() helper in core/master.c.

tcpi_unacked is a TCP congestion control metric. It counts unacknowledged segments in flight, not connections waiting in the accept queue. On a healthy server with no packet loss, this value is almost always 0.

The accept queue (connections that completed the TCP handshake but have not yet been pulled by accept()) is a separate kernel data structure. It is not exposed through TCP_INFO.

The max_queue field is similarly wrong. It is set to tcpi_sacked (SACKed segments), another congestion control metric unrelated to the configured listen backlog. There is a known IPv6 bug where max_queue always returns 0 (unbit/uwsgi issue #2577).

This implementation has not changed across any uWSGI 2.0.x release (the same tcpi_unacked/tcpi_sacked logic is present in 2.0.12 and in current 2.0.x source).

UNIX sockets: non-standard ioctl

For UNIX domain sockets, uWSGI uses a custom ioctl called SIOBKLGQ (ioctl number 0x8908). It is not part of mainstream Linux kernels: the code is compiled only when uWSGI is built with the UNBIT flag (a patch set some distributions ship), and the ioctl itself comes from a non-upstream kernel patch. On standard builds and kernels it is unavailable, so the queue measurement for UNIX sockets is never populated.

The load field is not latency

The top-level load field in the stats JSON is set to the same value as listen_queue. Despite its name suggesting latency or system load, it is not. The uWSGI source code contains a TODO comment acknowledging that this field does not measure what its name implies: in core/master.c, uwsgi.shared->load = backlog; sits directly under the comment // TODO load should be something more advanced based on different values, and listen_queue is set to the same backlog value.

listen_queue_errors is dead code

The listen_queue_errors field appears in the stats JSON output but no code path in the uWSGI source increments this counter. A full-source grep for its backing field (backlog_errors) finds only reads and serialization, never a write. It will always read 0.

Quick checks

Run these to confirm whether your listen_queue field is actually broken versus your system genuinely having no queue pressure:

# Check the uWSGI stats fields
uwsgi --connect-and-read 127.0.0.1:9191 | jq '{listen_queue, load, listen_queue_errors}'

# Check the real accept queue depth (TCP)
# Recv-Q = current queue depth, Send-Q = configured backlog limit
ss -ltn 'sport = :8000'

# Check the real accept queue depth (UNIX socket)
ss -lxn 'src /run/uwsgi/app.sock'

# Check kernel-level listen overflows (system-wide)
nstat -az TcpExtListenOverflows TcpExtListenDrops

If ss shows a non-zero Recv-Q while uWSGI stats shows listen_queue: 0, the stats field is confirmed broken on your system. If nstat shows non-zero TcpExtListenOverflows while listen_queue_errors is 0, that confirms the errors field is dead too.

How to measure the real backlog

ss: current queue depth

ss reads directly from kernel netlink and shows the actual accept queue depth:

# TCP socket - check Recv-Q for depth, Send-Q for backlog limit
ss -ltn 'sport = :8000'
# Output columns: State Recv-Q Send-Q Local Address:Port Peer Address:Port
# Recv-Q > 0 means connections are waiting for accept()
# Send-Q shows the effective backlog (min of --listen and somaxconn)

# UNIX socket
ss -lxn 'src /run/uwsgi/app.sock'

Recv-Q should be 0 in steady state. Any sustained non-zero value means workers cannot accept connections fast enough. Send-Q on a LISTEN socket shows the configured backlog, which is the effective minimum of uWSGI’s --listen value and the kernel’s net.core.somaxconn.

TcpExtListenOverflows: overflow counter

When the accept queue is full and the kernel drops a connection, it increments TcpExtListenOverflows and TcpExtListenDrops:

# Current values (system-wide, not per-socket)
nstat -az TcpExtListenOverflows TcpExtListenDrops

# These are cumulative counters. Track the rate of change.
# Any non-zero rate means connections are being actively dropped.

These counters are system-wide, not per-socket. On multi-service hosts, correlate with the ss output for uWSGI’s specific socket to attribute drops correctly.

somaxconn: the silent cap

The kernel caps the listen backlog at net.core.somaxconn. If you set --listen 1024 but somaxconn is 128 (the kernel default before Linux 5.4; 4096 since 5.4), the effective backlog is the lower of the two. uWSGI does not warn about this truncation.

# Check current somaxconn
cat /proc/sys/net/core/somaxconn

# Increase it (runtime change, not persisted across reboots)
# Persist via sysctl.conf or a systemd sysctl snippet
sysctl -w net.core.somaxconn=1024

Always increase somaxconn before or alongside increasing --listen. Otherwise the larger --listen value is silently ignored.

The “listen queue full” log message

uWSGI emits a log line when it detects a full listen queue:

*** uWSGI listen queue of socket "NAME" (fd: N) full !!! (QUEUE/MAX_QUEUE) ***

This message comes from the master process’s master_check_listen_queue() in core/master.c, which logs when queue > 0 && queue >= max_queue; both values are derived from tcpi_unacked and tcpi_sacked, not from the actual accept queue. So the log fires on a TCP congestion condition, not necessarily a full accept queue. It is emitted at verbose logging level (only with --log-verbose / verbose logging enabled).

If you see this log message, investigate TCP health (retransmissions, congestion). But the absence of this message does not mean the accept queue is not full. The real signal is TcpExtListenOverflows.

What to alert on instead

SignalSourceWhat it tells youAlert threshold
Accept queue depthss Recv-Q on the LISTEN socketConnections waiting for accept()Sustained > 0
Listen overflowsnstat TcpExtListenOverflowsConnections dropped by kernel (active outage)Any non-zero rate
Listen dropsnstat TcpExtListenDropsConnections dropped for any reasonAny non-zero rate
Worker busy ratiouWSGI stats (count status == "busy" / alive workers)Approaching capacity cliffSustained >= 80%
Accepting worker countuWSGI stats (count pid > 0 AND accepting == 1 AND status != "cheap")Available serving capacityZero for > 60s = critical

The first two signals replace the broken listen_queue and listen_queue_errors fields. They come from the kernel, not from uWSGI’s measurement path.

If you were previously alerting on listen_queue > 0 or listen_queue_errors > 0, those alerts will never fire. Replace them with ss-based queue depth monitoring and TcpExtListenOverflows rate alerting.

Prevention

  • Do not trust listen_queue, load, or listen_queue_errors in the uWSGI stats JSON on standard Linux. They are known-broken and will not improve without a source code change.
  • Monitor the accept queue externally with ss for depth and nstat for overflows. This is the only reliable signal for socket backlog pressure.
  • Set somaxconn before increasing --listen. The kernel silently truncates the backlog to somaxconn. A mismatch between --listen and somaxconn gives you a false sense of burst capacity.
  • Correlate queue depth with worker busy ratio. A growing Recv-Q with all workers busy means worker pool starvation. A growing Recv-Q with idle workers means accept contention (consider --thunder-lock).
  • Document this for your team. The most common failure mode is a new engineer building a dashboard on listen_queue, seeing it always at zero, and assuming the queue is healthy during an incident.

How Netdata helps

Netdata’s uWSGI collector reads the stats server JSON and surfaces worker-level metrics: busy ratio, accepting worker count, harakiri rate, avg_rt, exceptions, respawn rate, and per-worker RSS. For the broken listen_queue fields, Netdata’s Linux network monitoring provides the signals that fill the gap:

  • TCP accept queue depth from netlink, showing per-socket Recv-Q at per-second resolution.
  • TcpExtListenOverflows and TcpExtListenDrops from kernel counters, tracked as rates for overflow detection.
  • Worker busy ratio and accepting worker count from the uWSGI stats server, correlated against queue depth to distinguish starvation from accept contention.
  • Harakiri rate and avg_rt trends that precede queue buildup, giving earlier warning than the queue itself.

Seeing the kernel-level queue signals and uWSGI worker metrics in a single timeline is what distinguishes “listen_queue is zero, everything looks fine” from “Recv-Q is growing, all workers are busy, harakiri count just started rising.”