The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-listen-backlog-somaxconn ▌

Operations Guides

uWSGI listen backlog and net.core.somaxconn: sizing the connection queue

The connection backlog is the buffer between arriving TCP connections and your uWSGI workers. It is set by two independent values that must agree: uWSGI’s --listen option and the kernel’s net.core.somaxconn. The effective backlog is the smaller of the two. If either is too small, the queue fills during brief traffic spikes or downstream slowdowns, and the kernel starts dropping connections silently.

The defaults are both low. uWSGI ships with --listen 100. Linux kernels before 5.4 default net.core.somaxconn to 128; kernels 5.4 and later raised the default to 4096. On any service handling hundreds of requests per second, a 200-millisecond downstream hiccup fills a 100-slot queue in under a second. The result is intermittent client-side connection errors (timeouts or “connection refused”) with nothing in the uWSGI logs.

What it is and why it matters

When a client connects to uWSGI’s listening socket, the kernel completes the TCP handshake and places the connection in the socket’s accept queue (the listen backlog). The connection sits there until a uWSGI worker calls accept() to pull it out and begin processing. If all workers are busy, connections accumulate in this queue. If the queue is full, the kernel drops new connections with no application-level log entry.

Two limits control the queue size:

  • uWSGI --listen (or listen in config files): The backlog value uWSGI passes to the kernel’s listen() syscall when creating the socket. The uWSGI default is 100.
  • net.core.somaxconn: A kernel sysctl that sets an upper bound on the backlog any socket can request. The kernel silently caps the listen() backlog to this value.

The effective backlog is min(--listen, net.core.somaxconn). Setting --listen 1024 on a host where somaxconn is still 128 gives you an effective backlog of 128. On current uWSGI releases, if --listen exceeds somaxconn, uWSGI refuses to start entirely, printing an error (“Listen queue size is greater than the system max net.core.somaxconn”) and exiting. The check has been present since at least uWSGI 1.9.17, with the same exit behavior (uwsgi_nuclear_blast()); no released version we checked warns and continues.

This check applies to both TCP and UNIX domain sockets. Both values must be raised together. If you only raise --listen, uWSGI will not start.

The backlog is a cliff-edge resource. Below the limit, connections queue and eventually get served. At the limit, connections are dropped with no error counter in uWSGI and no warning. The only evidence is client-side errors or kernel-level overflow counters.

How it works

The interaction happens in three stages: socket creation, the kernel queue, and overflow handling.

Socket creation: the startup check

When uWSGI starts, the master process calls socket() followed by listen(fd, backlog) where backlog is the value from --listen. The kernel compares this to /proc/sys/net/core/somaxconn and silently truncates the backlog to the smaller value. On modern uWSGI, the master also explicitly checks somaxconn against --listen before calling listen(). If --listen > somaxconn, uWSGI prints an error and exits rather than silently running with a truncated backlog.

This means you cannot set --listen higher than somaxconn on any current uWSGI release.

flowchart TD
    A["uWSGI --listen
default: 100"] --> C["Effective backlog =
min(listen, somaxconn)"] B["net.core.somaxconn
default: 128 or 4096"] --> C C --> D["Kernel accept queue depth limit"] D --> E{"Queue depth < limit?"} E -- Yes --> F["Connection queued
waits for worker accept()"] E -- No --> G["Connection dropped
client sees timeout or RST"] F --> H["Worker calls accept()"] H --> I["Request processing begins"]

The kernel accept queue

Once the socket is listening, every new TCP connection the kernel accepts goes into the accept queue. The queue depth rises when connections arrive faster than workers can accept() them, and falls when workers catch up. In steady state with adequate capacity, the queue depth should be zero or near-zero. Any sustained non-zero depth means workers cannot keep up.

For UNIX domain sockets, the same mechanism applies. The kernel maintains an analogous queue, and somaxconn limits it the same way.

Overflow: silent connection drops

When the queue is full, the kernel’s behavior depends on the socket type and, for TCP, the net.ipv4.tcp_abort_on_overflow sysctl:

  • TCP sockets (default, tcp_abort_on_overflow=0): The kernel silently drops the final ACK of the three-way handshake. No RST is sent. The client retransmits and eventually times out or retries the connection. This is the most common production scenario.
  • TCP sockets (tcp_abort_on_overflow=1): The kernel sends RST after the final ACK. The client sees ECONNREFUSED immediately.
  • UNIX domain sockets: The kernel returns ECONNREFUSED to the connecting process immediately.

In all cases, uWSGI has no visibility into the drop. There is no log line, no exception, and no error counter. The uWSGI listen_queue_errors stats field exists in the JSON output but is not populated: it is read from uwsgi.shared->backlog_errors, which nothing in the uWSGI source ever increments.

The only way to detect overflow is at the kernel level.

Where it shows up in production

Behind nginx or a reverse proxy

When nginx sits in front of uWSGI (the most common deployment), nginx manages its own upstream connection pool. Under normal conditions, nginx opens connections to uWSGI as needed and reuses them. The uWSGI backlog sees only the connections nginx actually opens, not the full client load.

This changes under pressure. If uWSGI workers slow down (database latency, downstream API degradation), nginx’s upstream connections take longer to return. nginx opens more connections to handle incoming client requests. These connections queue in the uWSGI backlog. If the backlog is small, nginx gets connection failures and returns 502 to clients. The error appears in nginx logs (upstream failed (113: No route to host) or connect() failed (111: Connection refused)), not uWSGI logs.

Containers and Kubernetes

This isolation is a kernel property of network namespaces, not a Docker feature: net.core.somaxconn lives in the network namespace, and when a namespace is created its value is initialized to the SOMAXCONN constant. Any runtime that gives a container its own network namespace (Docker bridge mode, containerd, CRI-O, podman) behaves the same way; only --network host (or equivalent) shares the host’s namespace and therefore the host’s value.

Docker containers using bridge networking do not inherit the host’s somaxconn. Each container gets its own network namespace, and somaxconn is initialized to the kernel’s SOMAXCONN constant (4096 on kernel 5.4+, 128 on older kernels). Changing somaxconn on the host has no effect inside containers. Containers using --network host share the host’s network namespace and thus the host’s somaxconn.

To set it per-container:

# Docker: set somaxconn at container creation (match or exceed your --listen value)
docker run --sysctl net.core.somaxconn=1024 ...

In Kubernetes, net.core.somaxconn is classified as an unsafe sysctl. You must enable it on the kubelet with --allowed-unsafe-sysctls 'net.core.somaxconn' and then set it in the pod spec under securityContext.sysctls.

Graceful reloads

During a graceful reload, old workers finish current requests while new workers start. There is a window with reduced accepting capacity. If the application has a slow import phase, this window can be long enough for the backlog to fill. A backlog of 100 connections offers almost no cushion during this period. A backlog of 1024 absorbs the burst.

The startup warning

When uWSGI starts with the default --listen 100, the startup output includes an informational line: “your server socket listen backlog is limited to 100 connections.” This is not an error. It is a hint that the default may be too small for your traffic.

Sizing the connection queue

Production recommendation

For most production deployments, --listen 1024 with net.core.somaxconn set to at least 1024 provides enough queue depth to absorb brief spikes: downstream hiccups of 200-500 milliseconds, reload windows, traffic bursts from cache misses or retry storms.

For higher-traffic services, 4096 is reasonable, especially on kernel 5.4+ where somaxconn already defaults to 4096. The practical upper limit is 65535 (the backlog value is a 16-bit integer). Values above a few thousand provide diminishing returns unless your traffic legitimately sustains thousands of queued connections, which usually indicates a capacity problem rather than a queue-sizing problem.

Setting both values

# Set the kernel sysctl (non-persistent until added to sysctl.conf or /etc/sysctl.d/)
sysctl -w net.core.somaxconn=1024

# Verify the running value
cat /proc/sys/net/core/somaxconn
# Output: 1024

# Then configure uWSGI (INI format)
# listen = 1024

# Or on the command line
uwsgi --listen 1024 ...

To make the sysctl change persistent across reboots, add it to /etc/sysctl.conf or a file under /etc/sysctl.d/.

Verifying the running value

You cannot trust the listen_queue field in the uWSGI stats JSON. The field almost always reads 0 regardless of actual backlog. The master samples it with TCP_INFO for TCP sockets (kernel-version-dependent fields) and, for UNIX sockets, with the non-standard SIOBKLGQ ioctl that only exists on the unbit-patched kernel; on standard kernels the UNIX socket value stays 0.

Use the ss command to measure the actual queue depth externally:

# TCP socket: check Recv-Q (current queue depth) and Send-Q (backlog limit)
ss -ltn 'sport = :8000'

# UNIX socket: same columns
ss -lxn | grep uwsgi

In the output, Recv-Q is the current number of connections waiting in the accept queue. Send-Q is the configured backlog limit. In steady state, Recv-Q should be 0. Any sustained non-zero value means workers cannot accept fast enough.

To check for overflow events at the kernel level:

# System-wide overflow and drop counters (monotonic, check rate of change)
nstat -az TcpExtListenOverflows TcpExtListenDrops

These counters are system-wide, not per-socket. On multi-service hosts, correlate them with the per-socket ss output to attribute drops to uWSGI specifically.

Tradeoffs and sizing considerations

A larger backlog is not a substitute for adequate worker capacity. The backlog absorbs transient spikes; it does not fix sustained overload. If workers cannot keep up at current traffic levels, the queue fills regardless of size, and connections are eventually dropped. The difference is whether you have seconds of cushion or fractions of a second.

Backlog sizeWhat it handlesWhen it is insufficient
100 (uWSGI default)Brief microspikes on low-traffic servicesAny service behind nginx above 50 req/s, or any reload with slow startup
1024200-500ms downstream hiccups, reload windows, moderate burstsSustained overload lasting more than 1-2 seconds at high traffic
4096Longer degradation windows, high-traffic servicesSustained overload where workers are fundamentally undersized

The key decision is not the exact number but ensuring both values agree and are large enough for your traffic pattern. A mismatch (high --listen, low somaxconn) produces a startup failure on modern uWSGI or a silently truncated backlog on older versions.

Signals to watch in production

SignalWhy it mattersWarning sign
ss -ltn Recv-Q on the listening socketCurrent queue depth. Non-zero means workers are not keeping up.Sustained non-zero value, or value approaching Send-Q (the backlog limit)
nstat -az TcpExtListenOverflowsKernel-level overflow counter. Each increment is one or more dropped connections.Any non-zero rate of change in production
Worker busy ratioWhen this reaches 100%, every additional request queues in the backlog.Sustained 80% or above indicates limited headroom
Average response time (avg_rt)Rising response time means workers hold connections longer, causing the queue to fill faster.Sustained increase above 2x baseline
Accepting worker countFewer accepting workers means less capacity to drain the queue.Drop below expected minimum

How Netdata helps

  • Per-second collection of TCP socket queue depth lets you watch Recv-Q climb toward Send-Q in real time, before overflow occurs.
  • The TcpExtListenOverflows and TcpExtListenDrops kernel counters are collected natively, providing direct evidence of silent connection drops that uWSGI itself cannot report.
  • Worker busy ratio and accepting worker count from the uWSGI stats server are correlated alongside kernel socket metrics, so you can see the relationship between worker saturation and queue depth in a single view.
  • Average response time (avg_rt) trends are tracked per worker, letting you correlate latency increases with queue buildup.
  • Host-level kernel parameters like net.core.somaxconn are visible alongside runtime metrics, making it easy to confirm that the effective backlog matches your intent.