The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-reserve-pool-activation ▌

Operations Guides

PgBouncer reserve pool activation: overflow capacity that hides an undersized pool

PgBouncer’s reserve pool is overflow capacity: extra server connections beyond pool_size that PgBouncer may open when clients have waited too long. Used as designed, it absorbs a short traffic spike and goes quiet. Used as a crutch, it hides a chronically undersized pool for months, until the day both base pool and reserve are exhausted and the queuing cliff is steeper than it would have been otherwise.

This article covers the activation mechanism, how to detect reserve usage, how to tell healthy burst absorption from chronic undersizing, and what to change when it is the latter.

For the broader saturation picture, see PgBouncer pool utilization high: sv_active approaching pool_size before clients queue. For the pooling model, see How PgBouncer actually works in production: a mental model for operators.

What the reserve pool is

Two settings control it:

  • reserve_pool_size: how many additional server connections a pool may open beyond its pool_size. Default is 0, so the reserve pool is disabled entirely. If you have never set this, nothing in this article is happening on your system.
  • reserve_pool_timeout: how long a client must wait in the queue before PgBouncer may draw from the reserve. Default is 5 seconds. This is a global setting, not per-database.

The total server connection ceiling for a pool is pool_size + reserve_pool_size, capped further by max_db_connections (per database) and max_user_connections (per user). If max_db_connections is lower than pool_size + reserve_pool_size, the reserve connections you think you have may never be created.

A version note that matters for monitoring: in PgBouncer 1.24.0, the per-database parameter was renamed from reserve_pool to reserve_pool_size, and the SHOW DATABASES output column changed with it. The old name is still accepted as a config alias. If you parse SHOW DATABASES by column name (or use an exporter that does), check which column your version emits; tooling written against the old name silently stopped reporting this value on 1.24.0 and later. Also in 1.24.0, reserve_pool_size became settable per user, in addition to globally and per database.

How activation works

The trigger is not pool fullness. The trigger is a waiting client. The sequence:

  1. A client sends a query and no server connection is free, so the client enters the pool’s FIFO wait queue.
  2. The client waits. While the wait is shorter than reserve_pool_timeout (default 5 seconds), PgBouncer does nothing special.
  3. Once the client has waited longer than reserve_pool_timeout, PgBouncer may open an additional server connection from the reserve, up to reserve_pool_size extra connections, subject to the max_db_connections and max_user_connections caps.
  4. When load subsides, those extra connections return to normal lifecycle rules and the pool shrinks back toward pool_size.
flowchart TD
  A[Client sends query] --> B{Free server connection in pool?}
  B -- yes --> C[Assign and execute]
  B -- no --> D[Client enters FIFO wait queue]
  D --> E{Waited longer than reserve_pool_timeout?}
  E -- no --> D
  E -- yes --> F{reserve_pool_size > 0 and reserve slots free?}
  F -- yes --> G[Take connection from reserve_pool]
  G --> C
  F -- no --> H[Keep waiting]
  H --> I{Wait exceeds query_wait_timeout?}
  I -- no --> D
  I -- yes --> J[Client disconnected with error]

Two consequences follow:

  • Activation implies waiting already happened. Reserve connections only appear after at least one client has been blocked for more than reserve_pool_timeout. If you see reserve usage, you also had cl_waiting > 0 and maxwait above the timeout. The reserve pool does not prevent queuing; it responds to it.
  • The reserve is a delay, not a cure. When pool_size + reserve_pool_size is fully in use, clients queue again and eventually hit query_wait_timeout (default 120 seconds) and are disconnected. The reserve buys headroom for a spike; it does not change the shape of the failure when demand keeps climbing.

How to detect reserve pool activation

There is no dedicated “reserve pool in use” counter. Detection is a comparison plus a log signal.

# 1. Per-pool server connection totals
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW POOLS;"

# 2. Configured pool_size per database
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW DATABASES;"

# 3. Confirm the reserve settings actually in effect
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -Atc "SHOW CONFIG;" | grep -E "reserve_pool|pool_size"

# 4. Log evidence of reserve draws
grep -ic "taking connection from reserve_pool" /var/log/pgbouncer/pgbouncer.log

For each pool, sum sv_active + sv_idle + sv_used + sv_tested + sv_login from SHOW POOLS. Compare it with that pool’s effective pool_size: SHOW DATABASES shows the database override/default, while SHOW USERS shows a user-level pool_size override when one is set. If the total exceeds the effective base size, the excess came from the reserve. Equivalently, an sv_active / pool_size ratio above 100% is only possible with reserve connections in play.

PgBouncer logs the lowercase warning taking connection from reserve_pool each time it opens a reserve connection. The trigger checks the age of the first (oldest) waiter, not every queued client. A large pileup can therefore still produce many lines if PgBouncer repeatedly opens reserve slots, and log volume can become a secondary problem during an incident.

One caveat on PgBouncer 1.24.0 or later: setting default_pool_size to 0 means an unlimited base pool. In that mode PgBouncer bypasses the normal base-pool/reserve-pool branch, so reserve_pool_size does not create a distinct overflow tier; only database and user connection caps still apply.

Brief activation vs sustained activation

This is the decision the whole article exists for.

PatternWhat it meansWhat to do
Reserve connections appear for seconds to under a minute during a traffic spike, then disappearIntended use. The reserve did its job.Nothing, beyond noting the spike.
Reserve in use for more than about 5 minutes continuouslyBase pool_size is too small for current sustained load.Plan a pool_size increase.
Reserve drawn at roughly the same times every day (peak hours, batch windows)Chronic undersizing. The reserve is part of your steady-state capacity now.Treat reserve usage as your real pool_size and resize accordingly.
Reserve fully consumed AND cl_waiting still growingBoth base and reserve are exhausted.Active incident. See the pool exhaustion guide.

Brief activation is informational. Reserve connections in use for more than 5 minutes sustained is a ticket-level signal that the pool is undersized.

Why sustained reserve usage is dangerous

The reserve pool masks the real number. If your dashboards and capacity reviews look at sv_active / pool_size and mentally cap at 100%, a pool that routinely runs at 130% of pool_size (because the reserve quietly covers the gap) never shows up as a sizing problem. Growth continues, the reserve absorbs it, and the first visible symptom arrives only when demand exceeds pool_size + reserve_pool_size, at which point:

The queuing behavior at that point is cliff-edge, and the cliff is taller than it would have been for a correctly sized base pool, because you deferred the resize decision for as long as the reserve held.

Pooling mode changes the risk calculus. In session mode, server connections are held for the entire client session, so reserve connections are returned slowly and reserve exhaustion is much more likely under the same client count. In transaction mode, connections turn over per transaction, so a small reserve stretches much further. If you run session mode with a reserve pool, treat any reserve activation as more serious than the same event in transaction mode.

Signals to watch in production

SignalWhy it mattersWarning sign
Total server connections per pool vs pool_sizeDirect detection of reserve usage: total above pool_size means reserve is activeTotal exceeds pool_size for more than a minute
sv_active / pool_size ratioA ratio above 100% is only possible with reserve connectionsRatio above 100% sustained; ratio persistently 85-100% predicts reserve use
cl_waiting per poolReserve activation requires a waiting client, so queue depth precedes and accompanies itAny sustained non-zero value
maxwait (SHOW POOLS)The oldest waiter’s age; must exceed reserve_pool_timeout (default 5s) before the reserve kicks inValues repeatedly crossing 5s; approaching application timeouts
avg_wait_time (SHOW STATS / SHOW STATS_AVERAGES)The queuing latency PgBouncer injects, which is what the reserve exists to capSustained values well above baseline even while the reserve absorbs load
taking connection from reserve_pool log linesFrequency of reserve draws; the only per-event recordAny steady rate outside known spikes; sudden flood during incidents
paused / disabled (SHOW DATABASES)Context: during an administrative pause, queuing and reserve behavior look alarming but are expectedCheck before escalating anything

What to do when reserve usage is sustained

The fix is to size the base pool for the load you actually have, not to grow the reserve.

  1. Measure real demand. Look at peak sv_active including reserve connections over a representative week. That peak, plus burst headroom, is your target pool_size. Utilization below 70% is healthy, 70-85% watch closely, above 85% act.
  2. Check the PostgreSQL budget first. Every additional server connection consumes a max_connections slot on the backend. Sum pool_size (plus reserve) across all pools and all PgBouncer instances targeting the same PostgreSQL, and keep the total comfortably below max_connections minus reserved superuser slots. Raising pool_size past what PostgreSQL can serve trades queuing at the pooler for login failures at the backend.
  3. Check max_db_connections and max_user_connections. If these caps are below your intended pool_size + reserve_pool_size, raising pool_size alone does nothing.
  4. Increase pool_size and RELOAD. Pool size changes apply on reload; no restart needed.
  5. Verify latency attribution before and after. Compare avg_wait_time against avg_query_time. If avg_query_time is high, the backend is slow and a bigger pool mostly buys you more concurrent slow queries; fix the queries first. If avg_query_time is fine and avg_wait_time carries the latency, the resize is the right call. See PgBouncer avg_wait_time high: the latency the pool itself is injecting.
  6. Keep the reserve small and in place. After resizing, the reserve returns to its intended role: absorbing the next genuine spike, not carrying daily peak.

Do not respond to sustained reserve usage by increasing reserve_pool_size. That deepens the masking effect and makes the eventual cliff worse.

How Netdata helps

  • Netdata’s PgBouncer collector queries the admin console continuously, so per-pool sv_active, sv_idle, and the other server connection states are captured as time series. “Total connections above pool_size” becomes a trend, not a snapshot you had to be watching.
  • Per-pool breakdown matters here: one pool living in its reserve while nine others idle looks fine in any aggregate. Netdata charts split by pool so the chronic offender stands out.
  • Correlating cl_waiting, maxwait, and avg_wait_time on the same dashboard shows the full activation chain: clients waited, maxwait crossed reserve_pool_timeout, the reserve opened, and wait time either recovered (healthy spike) or stayed elevated (undersized pool).
  • Baseline views of avg_query_time next to avg_wait_time let you confirm whether reserve events are a sizing problem (wait high, query time normal) or a backend problem (both high) before you touch pool_size.
  • Alerting on duration, not just presence, of reserve usage maps directly to the brief-versus-sustained decision: a few minutes of reserve draw during a deploy should not page anyone; 30 minutes every evening should open a ticket.