The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-sv-idle-zero ▌

Operations Guides

PgBouncer sv_idle at zero: no headroom and one slow query from a cascade

Your PgBouncer dashboards look green. cl_waiting is zero, maxwait is zero, no clients are queuing, no errors in the log. But SHOW POOLS tells a different story: sv_idle is 0 and sv_active equals pool_size. Every server connection in the pool is checked out. Nothing is waiting yet, but nothing is available either.

This is the “looks green, is actually yellow” state, and it is one of the most dangerous steady states a connection pooler can sit in. The next request that arrives while all connections are busy queues immediately. There is no buffer, no graceful degradation. PgBouncer’s saturation curve is cliff-edge: below 100% utilization, assignment latency is effectively zero; at 100%, latency jumps to unbounded FIFO queuing.

If your monitoring only watches cl_waiting, you will never see this state. The first signal you get is the incident itself: waiters appear, maxwait climbs, application timeouts fire, retries multiply the queue. This article is about catching the condition before that happens, figuring out why headroom disappeared, and getting it back.

What this means

Each PgBouncer pool (one per (database, user) pair) holds server connections in several states: sv_active (executing for a client), sv_idle (connected to PostgreSQL, immediately reusable), sv_used (idle but used at least once, still considered good), sv_tested (running server_check_query or server_reset_query), and sv_login (authenticating). The sum across these states is bounded by pool_size, plus whatever the reserve pool adds.

sv_idle = 0 with sv_active = pool_size means every slot is checked out and the ready reserve is empty. In transaction pooling mode, sv_idle is not waste: it is the inventory that absorbs the next burst. When inventory hits zero, the system has no shock absorber left. A single slow query, one lock wait, one idle-in-transaction session, and the pool tips from “fully utilized” to “queuing” with no warning in between.

flowchart TD
  A[sv_idle trending toward zero over days] --> B[sv_idle = 0, sv_active = pool_size]
  B --> C{Next request arrives}
  C -->|a connection frees in time| D[assigned instantly, state persists]
  C -->|all connections busy| E[cl_waiting = 1, maxwait starts]
  E --> F[application timeout fires]
  F --> G[client retries, new waiter joins queue]
  G --> H[queue grows faster than it drains: cascade]
  D --> B

The important distinction: this is not yet pool exhaustion. It is the precondition for pool exhaustion. Pool exhaustion is the cascade after the cliff. Zero headroom is standing at the edge of it.

Common causes

CauseWhat it looks likeFirst thing to check
Organic traffic growthsv_idle / pool_size declining slowly over days or weeks, avg_query_time flatTrend of peak sv_active against pool_size over the last month
Backend queries getting sloweravg_query_time and avg_xact_time elevated above baseline, connections held longerSHOW STATS_AVERAGES for query time vs baseline; PostgreSQL-side slow query sources
Idle-in-transaction sessionsavg_xact_time much larger than avg_query_time (10x or more)pg_stat_activity on PostgreSQL for idle in transaction state
Long-running analytical or batch queriesOne or a few sv_active connections with very old request_timeSHOW SERVERS for active connections sorted by request_time
Session pooling modesv_active tracks connected clients rather than concurrent work; headroom maps to session count, not query turnoverSHOW DATABASES for pool_mode; zero idle headroom may be inherent to the mode when client count sits at pool_size
Pool simply undersizedZero headroom at every peak, brief cl_waiting blips already appearingPeak demand vs pool_size; see pool utilization high
Reserve pool masking undersizingTotal server connections exceed pool_size, log shows “taking connection from reserve_pool”Compare sum of sv_* per pool against pool_size from SHOW DATABASES

Quick checks

All commands run against the PgBouncer admin console. Adjust host, port, and user for your deployment.

# Per-pool snapshot: the state you are diagnosing
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW POOLS;"

Look at sv_active, sv_idle, cl_waiting, and maxwait per pool. You are confirming sv_idle = 0, sv_active = pool_size, cl_waiting = 0. Note which (database, user) pools are affected; it is often one pool, not all.

# Configured pool size per database
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW DATABASES;"

Compare pool_size here against the sv_active counts from SHOW POOLS. If total server connections exceed pool_size, the reserve pool is being drawn and the base pool is chronically undersized.

# Wait time and query time averages: is the backend the reason connections are held?
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW STATS_AVERAGES;"

Compare avg_wait_time (should still be near zero in this state) against avg_query_time and avg_xact_time. If avg_xact_time is many times avg_query_time, idle-in-transaction is eating your headroom.

# Which server connections have been checked out the longest
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW SERVERS;"

Look for active connections with old request_time. Those are the connections holding pool slots. The link column ties each one back to a client.

# Rule out administrative state before treating this as an incident
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW DATABASES;" | awk -F'|' '$(NF-1)==1 || $NF==1'

In SHOW DATABASES output the last two columns are paused and disabled; this prints any database where either is set. A paused database distorts every other signal, and alerting on pool state during maintenance is a classic false positive.

How to diagnose it

  1. Confirm the state per pool. From SHOW POOLS, identify which pools have sv_idle = 0 and sv_active = pool_size with cl_waiting = 0. A single saturated pool coexisting with healthy pools points to a workload or sizing problem scoped to one (database, user) pair, not a global event.

  2. Establish whether this is new or chronic. If you have historical per-pool metrics, look at the sv_idle / pool_size ratio over days to weeks. A slow decline toward zero is organic growth or gradual backend degradation. A sudden drop to zero is a specific event: a deploy, a new query pattern, a batch job.

  3. Attribute the hold time. Check avg_query_time vs avg_xact_time in SHOW STATS_AVERAGES. If both are elevated proportionally, the backend is slower and connections are held longer per unit of work. If avg_xact_time dwarfs avg_query_time, the connections are being held while doing nothing: idle-in-transaction. Confirm on the PostgreSQL side with pg_stat_activity filtered on state = 'idle in transaction'.

  4. Find the specific holders. In SHOW SERVERS, sort active connections by request_time. One or two connections checked out for minutes while the rest turn over in milliseconds means a small number of heavy queries is consuming a large share of pool capacity. Five analytical queries in a pool of 20 is 25% of capacity gone to background work.

  5. Check pool mode. In session pooling mode, each connected client holds a server connection until disconnect, so sv_active tracks connected clients rather than concurrent work; low or zero sv_idle can therefore be the steady state while every session slot is in use. After sessions disconnect, server_idle_timeout (default 600s) governs how long their former server connections remain idle. The headroom levers are different there: you size for concurrent sessions, not for transaction turnover.

  6. Check for reserve pool draw. If the sum of all sv_* states exceeds pool_size, the reserve pool is active. Reserve connections only get created after a client has waited longer than reserve_pool_timeout (default 5s), so sustained reserve usage means waiters already happened and the base pool is undersized.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
sv_idle / pool_size per poolThe headroom ratio itself; the early warning this article exists forTrending toward zero over days; sustained zero at any time
sv_active / pool_size per poolPool utilization; at 100% the next request must queueAbove 85% sustained; see the utilization guide
cl_waiting per poolConfirms whether the cliff has been reachedAny sustained non-zero value; in the zero-headroom state it is still zero, which is the trap
maxwaitAge of the oldest waiter; user-facing pain once queuing startsApproaching application-side timeouts; see the maxwait guide
avg_wait_timeRolling average of queuing delay; catches sub-polling-interval waits a snapshot missesCreeping above zero during peaks when it used to be zero; see the avg_wait_time guide
avg_query_time / avg_xact_timeTells you why connections are held: slow backend vs idle-in-transactionSustained 2x baseline deviation; large xact/query gap
Reserve pool draw (sum of sv_* > pool_size)Overflow capacity in use means the base pool is undersizedAny sustained activation, not just brief spikes
paused / disabled from SHOW DATABASESContext that suppresses false positives1 on any database: treat other signals as maintenance artifacts

The single most valuable addition most deployments are missing: alert on the sv_idle / pool_size ratio, not just on cl_waiting. A reasonable rule of thumb is to keep at least roughly 20% of pool_size idle at peak. That buffer absorbs bursts and keeps you off the cliff edge. Treat a sustained drop below that as a capacity ticket, not a page.

Fixes

Kill the specific holder, if there is one

If diagnosis shows one long-running query or one idle-in-transaction session is consuming the slots, canceling it on the PostgreSQL side (pg_cancel_backend, or pg_terminate_backend for a stuck session) frees the connection back to the pool immediately. Coordinate with the application team first: terminating a backend that PgBouncer is using causes the attached client to receive an error. This is expected behavior, but it is user-visible.

Fix idle-in-transaction in the application

If avg_xact_time >> avg_query_time, the fix is in application code: transactions that stay open across HTTP calls, batch loops, or user think-time hold server connections while doing nothing. Shorter transactions return connections faster and restore headroom without changing any PgBouncer setting. As a backstop, PostgreSQL’s idle_in_transaction_session_timeout can auto-terminate these sessions, at the cost of errors for the offending clients.

Address backend slowness

If avg_query_time is elevated, the pool is a symptom. The connections are held longer because PostgreSQL is slower: missing indexes, lock contention, I/O saturation. Fixing the backend restores pool turnover. Increasing pool_size in this situation only sends more concurrent load to an already struggling database.

Increase pool_size, with the PostgreSQL budget in mind

If demand genuinely exceeds capacity and the backend is healthy, raise default_pool_size or the per-database pool_size (a RELOAD applies it). Two constraints:

  • The sum of all pool sizes across all PgBouncer instances targeting a PostgreSQL server must stay comfortably under that server’s max_connections, leaving room for superuser access, replication, and direct connections. Keep PgBouncer’s total potential draw under about 80% of max_connections.
  • Bigger pools shift the bottleneck to PostgreSQL. Validate the backend can handle the additional concurrent queries before you grant them.

Session mode: resize for sessions, not turnover

In session pooling mode, sv_idle = 0 under full client load is structural. The fix is either sizing the pool to peak concurrent sessions, or evaluating whether the workload can move to transaction mode (after auditing for session-dependent features like prepared statements, temp tables, SET variables, and advisory locks).

Prevention

  • Alert on headroom, not just queuing. Add an alert on sv_idle / pool_size sustained near zero, and on the ratio’s multi-day trend. This is the signal that fires before users notice anything.
  • Trend peak utilization weekly. Track peak sv_active / pool_size per pool and extrapolate. Linear growth gives you a runway estimate: (1.0 - current_ratio) / weekly_growth in weeks until persistent queuing.
  • Always pair wait time with query time. avg_wait_time near zero with rising avg_xact_time means you are consuming headroom silently. By the time wait time moves, you are already at the cliff.
  • Exclude maintenance states. Every pool alert must check paused and disabled from SHOW DATABASES and suppress during administrative operations.
  • Watch reserve pool usage as a sizing defect. Brief reserve draws during spikes are the design. Regular draws mean raise the base pool.
  • Do not “reclaim” idle connections. High sv_idle in transaction mode is the ready reserve. Shrinking pool_size because idle connections look wasteful is how you manufacture the zero-headroom state.
  • Audit application transaction hygiene. Idle-in-transaction is the most common silent headroom killer in transaction mode. Catch it with the avg_xact_time vs avg_query_time gap before it shows up as saturation.

How Netdata helps

  • Netdata collects SHOW POOLS per pool continuously, so sv_idle, sv_active, and cl_waiting are time series rather than snapshots. The multi-day drift of sv_idle toward zero, the pattern point-in-time checks miss, becomes a visible trend line.
  • Because SHOW POOLS is a snapshot, brief queuing events between scrapes are easy to miss. Netdata pairs cl_waiting with avg_wait_time from SHOW STATS, so wait time injected between polls still shows up in the average.
  • The diagnostic pivot in this article, avg_xact_time vs avg_query_time, is a direct chart correlation: both come from the same stats source, so the idle-in-transaction gap is visible without cross-referencing tools.
  • Per-pool breakdowns keep one saturated (database, user) pool from hiding inside a healthy aggregate, which is exactly how zero headroom usually presents.
  • Alerting on the sv_idle / pool_size ratio and on sv_active / pool_size crossing 85% gives you the capacity ticket while the state is still yellow, instead of the page after cl_waiting turns red. See the monitoring checklist for the full signal set.