The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-avg-rt-not-cumulative ▌

Operations Guides

uWSGI avg_rt is not a real average: why the latency number lies

uWSGI exposes a per-worker field called avg_rt in its stats server JSON. Most monitoring tools label it “average response time” and graph it as a latency indicator. It is not a cumulative or lifetime average.

The field is updated with the formula (old_avg_rt + current_request_time) / 2, an exponential moving average with a smoothing factor of 0.5. The most recent request contributes 50% of the displayed value. The request before that contributes 25%. By the seventh request back, the contribution is under 1%.

A single slow request can double the displayed avg_rt. A burst of fast requests can erase evidence of that slow request within a few samples. The metric reacts quickly to recent changes but tells you nothing about the distribution of latency across all requests served. It is a trend signal, not a latency SLI.

If your latency SLO is based on avg_rt, or your dashboards show it alongside p95/p99 as if they are comparable, you are looking at different things. This article covers the formula, where it diverges from operator expectations in production, and what to measure instead.

What avg_rt actually is

The avg_rt field appears in the uWSGI stats server JSON under each worker object as workers[].avg_rt. It is an integer representing microseconds.

uWSGI’s Metrics subsystem registers it as the gauge metric worker.<N>.avg_response_time, plus a core.avg_response_time aggregate (source: core/metrics.c, UWSGI_METRIC_GAUGE bound to uwsgi.workers[i].avg_response_time), which is what the carbon and statsd metric exporters pick up.

The value is per-worker, not aggregate. To get a fleet-level number, you must average across workers yourself, which introduces a second averaging step that further obscures the distribution.

You can read it directly from the stats server:

# Read avg_rt for all non-cheap workers (TCP stats socket)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | select(.pid > 0 and .status != "cheap")] | if length > 0 then (map(.avg_rt) | add / length) else 0 end'

The field and the (old + current) / 2 update rule have existed at least since uWSGI 1.9.x and are unchanged through the current 2.0.x line (the same expression appears in core/utils.c in 1.9.21 and in current source).

This is a long-known limitation: DataDog’s uwsgi-dogstatsd plugin tracked it in issue DataDog/uwsgi-dogstatsd#7 (“Wrong avg_response_time metric”, opened February 2016, still open), which points to upstream PR unbit/uwsgi#730 (also left open).

How the formula works

The update rule is:

avg_rt = (old_avg_rt + current_request_time) / 2

Every time a worker finishes a request, uWSGI takes the previous avg_rt value and the time the request took, sums them, and divides by two. The most recent request has exactly 50% influence on the new value.

flowchart LR
    R7["Request N-6 (~0.8%)"] --> AVG["avg_rt"]
    R6["Request N-5 (~1.6%)"] --> AVG
    R5["Request N-4 (~3.1%)"] --> AVG
    R4["Request N-3 (~6.3%)"] --> AVG
    R3["Request N-2 (~12.5%)"] --> AVG
    R2["Request N-1 (~25%)"] --> AVG
    R1["Request N (~50%)"] --> AVG

The weighting cascade is:

Requests agoWeight in current avg_rt
0 (most recent)50.0%
125.0%
212.5%
36.3%
43.1%
51.6%
60.8%
7+under 0.4% each

After roughly 7 requests, contributions from older measurements are negligible. The effective window is the last 5-7 requests, not the worker’s lifetime.

Compare this to a true cumulative average, which weights every request equally:

true_avg = sum(all_request_times) / count(all_requests)

A true cumulative average with 10,000 requests served would barely move on a single slow request (1/10,000 weight). The EMA moves by 50% on that same request.

Where it misleads you in production

Single slow request skews the number. If a worker has been serving requests at 10ms each, avg_rt sits around 10,000 (microseconds). One request takes 500ms (500,000us). The new avg_rt becomes (10000 + 500000) / 2 = 255000, or roughly 255ms. The displayed value jumps 25x from a single outlier. Anyone watching the dashboard sees a latency spike that, in a cumulative average, would be invisible.

Fast requests erase history. After that slow request, the next fast request (10ms) brings avg_rt to (255000 + 10000) / 2 = 132500 (~132ms). The next: 71250 (~71ms). Within 5-6 fast requests, avg_rt is back near baseline. If your polling interval is 10-15 seconds (common for stats server scrapes), you may never see the spike if enough fast requests arrived between polls to wash it out.

Aggregate averaging compounds the problem. If your monitoring averages avg_rt across N workers, you get the mean of N independent short-window EMAs. Workers serving different endpoints at different latencies produce wildly different avg_rt values, and the aggregate mean of EMAs is even less meaningful than the individual values.

Cold start inflates the number. When a worker is freshly spawned, its first few requests include cache misses, connection pool warmup, and lazy imports. These slow first requests set a high avg_rt that decays as faster requests follow. Alerting on avg_rt crossing a threshold will trigger false positives during cold start.

gevent mode includes I/O wait. In async (gevent/asyncio) mode, avg_rt reflects wall-clock time including time the worker spent waiting on I/O while the event loop handled other greenlets. A worker multiplexing many concurrent requests will show high avg_rt because each request’s wall-clock duration includes time spent yielding. The number does not indicate CPU time or actual processing time for a single request.

In async/gevent mode, uWSGI records start_of_request when the connection is accepted and end_of_request when the async request finishes, so avg_rt measures wall-clock time from accept to completion; time the request spent yielded to other greenlets is included, and the value is not per-greenlet CPU time.

avg_rt survives respawns. Source inspection confirms that avg_response_time is not reset when a worker is respawned: the respawn path resets only delta_requests (plus per-core harakiri timers) and carries a comment that worker counters must not be reset on reload. A worker that was harakiri-killed after serving a very slow request therefore carries its high avg_rt over to the replacement, and the restart itself will not produce a drop. Do not interpret changes in avg_rt around respawn events without cross-referencing respawn_count.

What to use instead

For latency SLIs and SLOs, use percentile-based metrics from access logs or application-level instrumentation. uWSGI’s avg_rt cannot give you this.

Per-endpoint p50/p95/p99 from access logs. If you are behind nginx, the access log records $request_time per request. Parse it, group by endpoint, and compute percentiles. This gives you the actual distribution, including tail latency that avg_rt cannot represent.

running_time / requests for cumulative average. uWSGI exposes workers[].running_time (total cumulative processing time in microseconds) and workers[].requests (total request count). Dividing the two gives the true cumulative average per worker:

# True cumulative average per worker (workers with 0 requests will show null)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.pid > 0) | {id: .id, true_avg_ms: (.running_time / .requests / 1000)}'

This is a real average, not an EMA. It weights every request equally. However, it still does not give you percentiles, and it mixes all endpoints together. Like avg_rt and requests, running_time is not reset on respawn (only delta_requests is), so the ratio is cumulative across the worker slot’s lifetime; treat it as a per-slot average, not a per-process one.

Application-level metrics. Instrument your WSGI or Rack application to emit per-request timing with endpoint labels to your metrics backend (Prometheus, StatsD, OpenTelemetry). This is the only way to get per-endpoint latency distributions with proper percentiles. uWSGI’s stats server was designed for operational diagnostics (worker state, harakiri counts, queue depth), not for SLI reporting.

When avg_rt is still useful

avg_rt is a fast-reacting trend signal that tells you whether latency is changing in the short term. Used correctly:

  • Trend detection. A sustained upward slope in avg_rt across multiple workers suggests a systemic slowdown (downstream dependency, resource contention). The EMA’s responsiveness makes it good for detecting the onset of degradation.
  • Correlation with other signals. avg_rt rising alongside worker busy ratio approaching 100% indicates capacity exhaustion. avg_rt rising alongside harakiri count indicates requests approaching the kill threshold. avg_rt rising on one worker only suggests that worker has a specific problem.
  • Quick triage. During an incident, glancing at avg_rt per worker can tell you which workers are slow, even if the absolute number is not reliable as a latency measurement.

Never use it as the sole basis for a latency alert or SLO. Use it as a supporting signal alongside throughput, error rates, and harakiri counts.

Signals to watch in production

SignalWhy it mattersWarning sign
Per-endpoint p95/p99 (from logs)True tail latency that avg_rt cannot representSustained increase above SLO threshold
Worker busy ratioSaturation indicator independent of latencySustained above 80%, or any sustained 100%
Harakiri rate (delta)Requests exceeding the kill timeoutAny sustained non-zero rate
avg_rt per worker (trend only)Fast-reacting indicator that latency is changingSustained upward slope across multiple workers
running_time / requestsTrue cumulative average per workerSignificant divergence from avg_rt confirms EMA volatility
Throughput (delta requests)Whether requests are completing at allSudden drop without corresponding traffic decrease

How Netdata helps

Netdata collects uWSGI stats server data at per-second resolution, which matters for avg_rt specifically because the EMA window is only 5-7 requests. At 10-second polling intervals, the metric can spike and recover between samples, making the volatility invisible.

  • Per-worker avg_rt alongside harakiri count and busy ratio lets you see whether a latency spike corresponds to a worker being killed and respawned, or to genuine application slowdown.
  • Correlating avg_rt with downstream dependency metrics (database query time, Redis latency, external API response time) shortens diagnosis when avg_rt trends upward. The root cause is almost always downstream, not in uWSGI itself.
  • Per-second worker busy ratio catches brief saturation events that avg_rt’s EMA washes out. Workers hitting 100% busy for 3 seconds then recovering may show only a minor blip in avg_rt, but busy ratio at 1-second resolution shows the full event.
  • Respawn count tracking lets you distinguish avg_rt changes caused by worker recycling (max-requests, harakiri) from genuine latency shifts.
  • Throughput (delta requests) at per-second resolution provides the denominator that avg_rt lacks. If avg_rt spikes but throughput is stable, the spike may be a single slow request. If avg_rt spikes and throughput drops, the system is genuinely degraded.