The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-harakiri-not-configured ▌

Operations Guides

uWSGI harakiri not configured: stuck workers with no timeout and no recovery

uWSGI workers are vanishing one at a time. The master process is alive, the stats server responds, and harakiri_count reads zero across every worker. By every metric you thought mattered, the service looks healthy. Then you notice that half the workers have been in busy status for the last ten minutes without processing a single new request. Requests are timing out at the proxy. The listen queue is filling. The application is effectively down, and nothing in uWSGI is attempting to recover it.

The root cause is the absence of a harakiri configuration. Default uWSGI has no request timeout. The harakiri value defaults to 0, which disables the per-request watchdog entirely. A single hung request permanently consumes a worker slot. There is no kill, no respawn, and no counter increment. Workers accumulate in a stuck state until the pool is exhausted.

The primary signal operators learn to watch, harakiri_count, stays at zero. That zero is not health. It is a blind spot: the counter only increments when harakiri fires, and harakiri never fires when it is not configured.

What this means

The harakiri timer is uWSGI’s per-request watchdog. When configured, the master process tracks how long each worker has spent on its current request. If the request exceeds the configured timeout, the master sends SIGKILL to that worker and forks a replacement. Each kill increments harakiri_count and respawn_count. The worker slot is recycled within seconds, and the service self-heals. harakiri_count itself persists: it is a monotonic per-worker counter that the master does not reset when the worker respawns, so it keeps accumulating across the worker’s lifespans.

Without harakiri, none of this happens. A worker that calls a database query with no application-level timeout, hits an infinite loop, or blocks on a deadlocked resource stays in that state indefinitely. The worker is still alive from the OS perspective, still appears in the process table, and still reports status: "busy" in the stats server. But it will never accept another request as long as it lives.

The cascade from a single stuck worker to total outage is mechanical:

flowchart TD
    A["Worker accepts request"] --> B["Application hangs"]
    B --> C["Worker stuck in busy"]
    C --> D{"Harakiri configured?"}
    D -->|No| E["No timeout fires"]
    D -->|Yes| F["Master kills worker, respawns"]
    E --> G["Worker slot lost permanently"]
    G --> H["More workers get stuck over time"]
    H --> I["Pool exhausted, service down"]

The harakiri_count metric is useless when harakiri is not configured. It will read zero forever regardless of how many workers are stuck. Monitoring this counter as a health signal without verifying that the configuration is armed creates false confidence. For a deeper treatment of the worker lifecycle and how stuck workers fit into the broader failure pattern catalogue, see How uWSGI actually works in production.

Common causes

CauseWhat it looks likeFirst thing to check
Blocking database queryOne or more workers stuck on a single URI; database shows long-running queries or lock waitsDownstream database lock and slow query state
External API call without timeoutWorkers stuck after a downstream call; no exceptions loggedDownstream service health and client-side timeout settings
Infinite loop or regex backtrackingWorker CPU pinned at 100 percent; request count frozen/proc/<pid>/syscall and /proc/<pid>/wchan for the stuck PID
DNS resolution hangMultiple workers stuck simultaneously, often after DNS changesResolver configuration and test name resolution directly
Deadlock on shared resourceWorkers stuck with low CPU usage; often follows a recent deployApplication-level lock state and database lock tables

Quick checks

These commands assume a stats server enabled with --stats on a TCP socket at 127.0.0.1:9191. Adjust the address for your deployment. If the stats server uses a UNIX socket, substitute uwsgi --connect-and-read /path/to/stats.sock.

# Check whether harakiri is configured (path varies by deployment)
# Note: this only catches config files, not CLI flags or environment variables.
grep -ri harakiri /etc/uwsgi/

# Count workers by status (busy vs idle)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | select(.pid > 0 and .status != "cheap") | .status] | group_by(.) | map({status: .[0], count: length})'

# Check harakiri_count across all workers (will be zero if harakiri is not configured)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].harakiri_count] | add'

# Detect stuck requests: elapsed time of in-flight requests per worker
# request start is exposed as req_info.request_start (epoch seconds), only when in_request == 1
uwsgi --connect-and-read 127.0.0.1:9191 | jq --argjson now "$(date +%s)" '[.workers[] | select(.pid > 0) | .id as $wid | .cores[] | select(.in_request == 1) | {worker: $wid, core: .id, age_seconds: ($now - .req_info.request_start)}]'

# Show what each busy worker is serving
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.status == "busy") | {id: .id, uri: (.uri // "unknown")}'

# Check total request throughput (sum across workers, compare between polls)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].requests] | add'

# Check what a stuck worker is blocked on (Linux, replace PID)
cat /proc/<pid>/syscall
cat /proc/<pid>/wchan

# Check listen queue depth (Recv-Q on a LISTEN socket shows current accept queue length)
ss -ltn 'sport = :8000'

The cores[] array is suppressed if uWSGI is started with --stats-no-cores. If your stats output does not include cores, you cannot detect stuck request ages from the stats server alone. In that case, rely on per-worker request count deltas and OS-level process inspection.

How to diagnose it

  1. Verify the configuration. Check the uWSGI config file for a harakiri directive. If it is absent or set to 0, the watchdog is disabled. This is the root cause of the missing recovery.
  2. Count accepting workers. Use the stats server to count workers where pid > 0 and status != "cheap" and accepting == 1. Compare against the expected minimum. Each stuck worker is one fewer accepting worker.
  3. Identify stuck workers. Look for workers in busy status whose request count has not changed between two consecutive polls. A worker whose requests counter is frozen while its status is busy is stuck.
  4. Check request age. If cores are available, compute the elapsed time of each in-flight request using req_info.request_start. Any request older than a few seconds for a typical web endpoint is suspect.
  5. Inspect the blocked state. For each stuck worker PID, read /proc/<pid>/syscall and /proc/<pid>/wchan. A worker blocked on read or poll is likely waiting on a downstream dependency. A worker showing no syscall (CPU-bound) may be in an infinite loop.
  6. Check downstream dependencies. The application code is the proximate cause, but the root cause is usually downstream: database locks, dead external APIs, DNS failures, or network partitions.

For the broader pattern of worker pool exhaustion and how it relates to the listen queue, see uWSGI worker pool starvation: the silent outage where every worker is busy.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
harakiri_count (delta)Detects requests exceeding the timeout ceilingNonzero delta means hung requests are being killed; flat zero means either health or misconfiguration
Worker busy ratioShows pool saturation before exhaustionSustained increase with declining throughput indicates stuck workers
Stuck request age (req_info.request_start)The only direct signal for hung requests when harakiri is not configuredRequest age exceeding expected duration without completion
Accepting worker countTracks effective serving capacityDeclining count with stable traffic means workers are being lost to hangs
Request throughput (delta)Confirms real user impactThroughput dropping while incoming traffic is stable means workers cannot complete requests
Respawn count (delta)Distinguishes recycling from stuck workersZero respawn delta with rising busy ratio means no recovery mechanism is firing

Fixes

Configure harakiri

Set harakiri to approximately 2 to 3 times your maximum legitimate request duration. If your slowest normal endpoint takes 10 seconds, set harakiri = 30. This gives legitimate requests headroom while ensuring stuck workers are killed and replaced.

Always enable harakiri-verbose alongside it. When harakiri fires, this flag causes uWSGI to log the blocked syscall and wait channel by reading /proc/<pid>/syscall and /proc/<pid>/wchan (Linux only). Without verbose logging, a harakiri event produces a kill with no diagnostic context.

[uwsgi]
harakiri = 30
harakiri-verbose = true

The shortcut flag -t also sets the harakiri timeout.

The reliable harakiri mode requires master = true. Without the master process, harakiri falls back to a raw SIGALRM-based timer that the uWSGI documentation describes as unreliable. Most production deployments already run with a master, but verify it.

Consider graceful harakiri on uWSGI 2.0.22+

Standard harakiri sends SIGKILL immediately. The worker has no chance to flush buffers, which breaks tracing libraries (Sentry, DataDog, OpenTelemetry) that need to emit span data on exit. uWSGI 2.0.22 introduced the sub-options that support this two-stage kill (confirmed by the 2.0.22 changelog):

  • harakiri-graceful-timeout: gives the worker a grace period to shut down before SIGKILL.
  • harakiri-graceful-signal: the signal sent first (default SIGTERM).
  • harakiri-queue-threshold: only triggers harakiri when the listen queue crosses a threshold, avoiding false kills during brief spikes.

These options are additive and backwards-compatible. If you do not set them, harakiri behaves exactly as before: immediate SIGKILL, no questions asked.

Add application-level timeouts

Harakiri is a last-resort safety net, not a substitute for proper timeout handling in application code. Every downstream call (database queries, HTTP requests, cache lookups) should have its own timeout configured at the library or framework level. This prevents requests from hanging in the first place and reduces the frequency of harakiri-triggered respawns.

Understand async mode limitations

Harakiri is per-process, not per-coroutine. In gevent or asyncio mode, the harakiri counter resets every time a new request kicks in within the same worker. This means harakiri may not fire for an individual slow coroutine if other requests are being multiplexed in the same process. It detects when the entire process is stuck, not when one greenlet is slow. If you run async workers, you cannot rely on harakiri alone for per-request timeout enforcement.

Do not rely on http-timeout for request timeouts

http-timeout is a legacy option (“set internal http socket timeout”) that is not wired to any per-request behavior in current uWSGI releases; it does not set a per-request processing timeout. Independent of it, uWSGI continues processing a request after the client has disconnected, wasting a worker on work no one will see. Harakiri is the mechanism that bounds that work.

Prevention

  • Always configure harakiri. Treat it as a mandatory production setting. The default of 0 disables the watchdog.
  • Set harakiri-verbose. Without it, harakiri events are opaque kills with no diagnostic trail.
  • Monitor stuck request age as a compensating signal. When harakiri is configured, harakiri_count detects kills. Request age detection via req_info.request_start catches requests that are slow but have not yet hit the timeout, giving earlier warning.
  • Audit configurations during deployment reviews. Check for the presence of harakiri in every uWSGI config, including Emperor/Vassal configs where individual vassals may override parent settings.
  • Verify master mode. Harakiri’s reliable mode depends on master = true.
  • Set per-call timeouts in application code. Every database query, HTTP call, and external dependency interaction should have a timeout.

For a structured audit of which signals your monitoring stack should be tracking, see uWSGI monitoring checklist: the signals every production app server needs.

How Netdata helps

  • Per-second worker status tracking catches the transition from idle to stuck within seconds, rather than waiting for a coarse polling interval to reveal that a worker has not changed state.
  • Stuck request age detection correlates in_request flags with req_info.request_start timestamps, surfacing individual hung requests even before harakiri would fire.
  • Busy ratio trends show the gradual capacity erosion that precedes pool exhaustion.
  • Throughput correlation confirms whether a busy ratio increase is causing real user impact (throughput dropping) or is a transient spike that self-resolves.
  • Harakiri count as a configuration signal distinguishes between “harakiri_count is zero because the system is healthy” and “harakiri_count is zero because the watchdog is not armed,” by correlating the counter with the known configuration state.
  • Downstream dependency metrics displayed alongside uWSGI worker state shortens root-cause analysis by showing whether database latency, external API response times, or DNS resolution are driving the worker stalls.