The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-write-read-errors ▌

Operations Guides

uWSGI write and read errors: broken pipes and clients that disconnect

write_errors and read_errors are per-core counters in the uWSGI stats server. write_errors increments when a socket write fails during response delivery, almost always because the client disconnected before the response finished (broken pipe). read_errors increments when the connection is lost during request body reading. Some of both are normal: users navigate away, mobile clients switch networks, browsers cancel pending requests. A sustained rise in write_errors, especially when correlated with rising response time or worker respawns, points to a real problem.

The diagnostic work is correlating write_errors with response time, respawn rate, and harakiri count to separate benign disconnects from server-side trouble. Do not treat every spike as an incident, and do not ignore a sustained rise because “clients disconnect all the time.”

What this means

When uWSGI calls write() or writev() to send response bytes to the client and the socket is already closed, the kernel delivers SIGPIPE to the process (or returns EPIPE if SIGPIPE is blocked). uWSGI intercepts this and increments the per-core write_errors counter. The same mechanism produces “write error” or “IOError: write error” in application error trackers like Sentry when the Python layer tries to write to the dead socket.

read_errors follows the same pattern on the input side: the client connection was lost while uWSGI was still reading the request body. This typically happens with slow uploads over unstable networks, proxy timeouts firing during large POST bodies, or clients that crash mid-request.

Both counters live in the stats server JSON under each worker’s cores array:

  • workers[].cores[].write_errors (per-core, monotonic)
  • workers[].cores[].read_errors (per-core, monotonic)

Both are suppressed entirely if uWSGI is started with --stats-no-cores. If your monitoring shows these metrics as zero or absent, verify that flag is not set before assuming there are no errors. Similarly, if the metrics subsystem is enabled, --metrics-no-cores is the corresponding option that suppresses per-core metrics from that collection path.

Common causes

CauseWhat it looks likeFirst thing to check
Normal client behaviorLow, steady write_errors rate with normal response times and no respawn correlationCompare rate against historical baseline
Slow responseswrite_errors rising alongside avg_rt; clients give up before response arrivesCheck downstream dependency latency
Proxy timeout mismatchwrite_errors spike when responses exceed the proxy upstream timeout; proxy closes the connectionCompare nginx uwsgi_read_timeout against your p95 response time
evil-reload-on-rss killswrite_errors spike aligned with respawn_count increases; workers killed via SIGKILL mid-responseCheck if --evil-reload-on-rss is configured
Harakiri mid-responsewrite_errors aligned with harakiri_count increases; worker killed after timeoutCheck which endpoints are timing out

Quick checks

# Sum write_errors across all workers and cores
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].cores[].write_errors] | add'

# Sum read_errors across all workers and cores
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].cores[].read_errors] | add'

# Per-worker write_errors and read_errors breakdown
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | {id: .id, write_errors: ([.cores[].write_errors] | add), read_errors: ([.cores[].read_errors] | add)}'

# Verify cores array is present (not suppressed by --stats-no-cores)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[0].cores | length'

# Check per-worker avg_rt (reported in microseconds)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.pid > 0) | {id, avg_rt_us: .avg_rt}'

# Check respawn and harakiri counts for correlation with write_errors
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | {id, respawn_count, harakiri_count}]'

# Check worker RSS against evil-reload-on-rss threshold
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.pid > 0) | {id, rss_mb: (.rss / 1048576)}'

These commands are read-only and safe to run at any time. The stats server does not require HTTP mode; uwsgi --connect-and-read works with the default raw JSON socket. If your stats server uses a UNIX socket, replace the address with the socket path.

How to diagnose it

  1. Establish the baseline. Poll write_errors and read_errors twice, 60 seconds apart, and compute the per-second rate. Some disconnects are always present. The question is whether the rate has changed significantly from your normal baseline.

  2. Correlate with response time. If write_errors are rising, check avg_rt in the same window. Rising write_errors with normal avg_rt suggests client-side behavior. Rising write_errors with rising avg_rt means your responses are slow enough that clients (or proxies) give up before they arrive.

  3. Correlate with respawn rate. If write_errors spikes align with respawn_count increases, workers are being killed mid-response. Subtract harakiri_count from respawn_count: if the difference matches the write_errors spike, check for --evil-reload-on-rss. If harakiri_count tracks the spike, workers are timing out on slow requests.

  4. Check the proxy timeout boundary. If nginx (or another reverse proxy) sits in front of uWSGI, its upstream timeout controls when the proxy abandons the connection. If nginx’s uwsgi_read_timeout is shorter than your p95 response time, nginx closes the upstream socket before uWSGI finishes writing, and uWSGI logs a write error.

  5. Check log format variables for per-request detail. If you need per-request granularity, add %(werr), %(rerr), and %(ioerr) to your uWSGI log format. These variables expose write errors, read errors, and their sum for each individual request, letting you see which endpoints are affected. They have been available since uWSGI 1.9.21.

flowchart td
    A["write_errors rate rising"] --> B{"avg_rt also rising?"}
    B -->|Yes| C["Slow responses: clients or proxy give up"]
    B -->|No| D{"respawn_count also rising?"}
    D -->|Yes, tracks harakiri| E["Harakiri killing workers mid-response"]
    D -->|Yes, exceeds harakiri| F["evil-reload-on-rss killing workers"]
    D -->|No| G{"Proxy timeout < p95?"}
    G -->|Yes| H["Proxy closing connections early"]
    G -->|No| I["Likely normal client disconnects"]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
write_errors rate (per-core sum)Primary indicator of client disconnects during responseSustained increase above baseline
read_errors rate (per-core sum)Connection loss during request body readSustained increase above baseline
avg_rt (EMA, microseconds)Slow responses cause clients to disconnectRising in lockstep with write_errors
harakiri_count (delta)Harakiri kills workers mid-response, generating write_errorsNon-zero rate correlating with write_errors
respawn_count (delta)evil-reload-on-rss kills generate respawns and write_errorsRespawns exceeding harakiri-driven respawns
Worker RSSMemory growth triggers evil-reload-on-rss thresholdsRSS approaching configured threshold
TX bytes per workerTruncated responses produce less TX than expectedTX per request ratio dropping

Fixes

Slow responses causing client disconnects

If avg_rt is rising alongside write_errors, the root cause is upstream. Your application is taking too long, and clients or proxies disconnect before the response arrives. Fix the downstream dependency (database query, external API call, cache miss path) rather than suppressing the symptom. If you cannot reduce response time immediately, increase the proxy upstream timeout to match your acceptable latency ceiling.

Proxy timeout mismatch

Align nginx’s uwsgi_read_timeout with your actual response time distribution. If your p95 is 5 seconds and uwsgi_read_timeout is 60 seconds, there is no mismatch. If your p95 is 30 seconds and the timeout is 10 seconds, every slow request generates a write error on the uWSGI side.

The nginx directive uwsgi_ignore_client_abort on; tells nginx not to tear down the upstream connection when the client disconnects. This prevents uWSGI from seeing a write error, but the tradeoff is that uWSGI continues processing a request whose response will never be delivered. This wastes worker capacity on abandoned requests. Use it only when you understand this cost.

evil-reload-on-rss killing workers mid-response

--evil-reload-on-rss sends SIGKILL to workers that exceed the RSS threshold, with no grace period. If the worker was mid-response, the client sees a truncated or broken response and uWSGI logs a write error. The fix is to replace --evil-reload-on-rss with --reload-on-rss, which triggers a graceful exit: the worker finishes the current request, then exits. If you must keep --evil-reload-on-rss for hard memory limits, monitor write_errors alongside respawn rates and accept that some clients will see broken responses during recycling.

Harakiri killing workers mid-response

When harakiri fires on a worker that has already started writing its response, the SIGKILL truncates the response and increments write_errors. This is expected behavior if the request genuinely exceeded the timeout. The fix is not to silence the write_errors but to address why the request was slow enough to hit harakiri. If the endpoint legitimately needs more time than the global harakiri, use the per-route harakiri action for that specific path.

Suppressing noise from normal disconnects

If you have confirmed that write_errors are from normal client behavior and the noise is polluting logs or error trackers, three uWSGI options work together to suppress it:

  • --ignore-write-errors true: suppresses uWSGI’s own log messages about write and writev errors
  • --ignore-sigpipe true: suppresses SIGPIPE log messages
  • --disable-write-exception true: prevents Python IOError/OSError exceptions from being raised on write failures

All three are needed because they address different layers. --ignore-write-errors and --ignore-sigpipe stop uWSGI’s C-level logging, but the Python application may still receive an IOError when it tries to write to the closed socket. --disable-write-exception prevents that exception from propagating to application code and error trackers.

The --write-errors-tolerance option sets the maximum number of allowed write errors per request (default: no tolerance). When the per-request write-error count exceeds it - and only if --disable-write-exception is not set - the Python layer raises IOError: write error; it does not restart the worker or log a warning by itself. Regardless of the mechanism, this option does not substitute for diagnosing the root cause of sustained write_errors.

read_errors

read_errors are typically less actionable than write_errors. They indicate the client connection was lost during request body reading, which is almost always client-side: network instability, client crash, or a proxy closing the connection during a slow upload. A sustained rise in read_errors without a corresponding rise in write_errors may indicate network problems between the proxy and uWSGI, or proxy misconfiguration dropping connections during large request bodies.

In the uWSGI source, read_errors is incremented only in the request-body read path (core/reader.c); failures during header parsing or protocol negotiation do not increment it.

Prevention

  • Monitor write_errors as a rate, not an absolute count. The counters are monotonic. Alert on sustained rate increase above baseline, not on the raw value.
  • Correlate write_errors with avg_rt and respawn_count. A write_errors spike without corresponding latency or respawn changes is likely benign client behavior.
  • Prefer --reload-on-rss over --evil-reload-on-rss. Graceful recycling avoids mid-response kills that generate write_errors.
  • Align proxy upstream timeouts with your response time distribution. Review nginx uwsgi_read_timeout whenever you deploy endpoints with different latency profiles.
  • Verify --stats-no-cores is not set. If it is, you lose all per-core metrics including write_errors and read_errors.
  • Add %(werr) and %(rerr) to your log format. Per-request error counts in access logs let you identify which endpoints are most affected by client disconnects.

How Netdata helps

Netdata collects uWSGI stats server output per second and correlates write_errors and read_errors with worker state.

  • Per-second rate charts for write_errors and read_errors show the start and duration of spikes without manual polling.
  • Correlation views overlay write_errors with avg_rt, respawn_count, and harakiri_count on a single timeline to separate client-driven disconnects from server-side problems.
  • Anomaly detection on write_errors rate flags deviations from the learned baseline, useful when normal disconnect volume is high enough to mask new patterns.
  • Per-worker charts show whether write_errors are concentrated on one worker or distributed across all workers.
  • RSS charts alongside write_errors and respawn data confirm or rule out evil-reload-on-rss as the cause.