The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-postgresql-max-connections ▌

Operations Guides

PgBouncer and PostgreSQL max_connections: when pool_size outruns the backend limit

PgBouncer exists to reduce the number of connections PostgreSQL has to hold, so it is easy to assume that adding PgBouncer makes the connection limit problem go away. It does not. PgBouncer is itself a consumer of PostgreSQL connection slots, and its worst-case demand is arithmetic you control in config files: one pool per (database, user) pair, each pool allowed to open pool_size server connections, plus reserve_pool_size overflow, multiplied by every PgBouncer instance pointing at the same backend.

When that worst-case sum exceeds what PostgreSQL can accept, the failure does not show up at steady state. It shows up during the burst: a deploy, a retry storm, a cold start, a batch job. PgBouncer tries to open new server connections, PostgreSQL refuses them with FATAL: sorry, too many clients already, and the pool drains while clients queue. This is a static configuration audit, not a runtime mystery.

This article gives you the arithmetic, the audit procedure, and the signals that tell you the mismatch is about to bite.

The arithmetic that matters

The sum of (pool_size + reserve_pool_size) across every pool, on every PgBouncer instance targeting one PostgreSQL, must stay under max_connections minus superuser_reserved_connections, minus anything else that connects directly.

Written out:

PostgreSQL budget  =  max_connections
                    - superuser_reserved_connections
                    - direct application clients
                    - replication connections
                    - admin / ops access headroom

PgBouncer demand   =  SUM over all instances of:
                        SUM over all (database, user) pools of:
                          pool_size + reserve_pool_size

Requirement        =  PgBouncer demand <= ~80% of max_connections

The ~80% target leaves room for everything that is not PgBouncer: psql sessions during incidents, monitoring, migrations, logical replication, and the superuser connections you will need when the backend is full. If PgBouncer demand can reach 100% of max_connections, then at the exact moment you most need to connect and fix things, you cannot get in.

Three properties of PgBouncer make this sum easy to get wrong:

  • Pools multiply by (database, user) pairs, not by database. pool_size is per pool, and a pool exists per unique (database, user) combination. Four users against two databases at default_pool_size of 20 is up to 160 server connections, not 40.
  • Per-database and per-user overrides hide. A database stanza can override pool_size, and max_db_connections / max_user_connections cap totals across pools. The effective ceiling per instance is not just default_pool_size from SHOW CONFIG.
  • Multiple instances multiply demand silently. Each PgBouncer process, including so_reuseport siblings and sidecars, has independent pools. Two instances each sized “safely” at 60% of max_connections are collectively at 120%.
flowchart TD
  subgraph pgb1[PgBouncer instance 1]
    p1a[pool db1/app1: pool_size + reserve]
    p1b[pool db1/app2: pool_size + reserve]
    p1c[pool db2/app1: pool_size + reserve]
  end
  subgraph pgb2[PgBouncer instance 2]
    p2a[pool db1/app1: pool_size + reserve]
    p2b[pool db2/app2: pool_size + reserve]
  end
  budget[PostgreSQL max_connections]
  reserved[minus superuser_reserved_connections]
  direct[minus direct clients and replication]
  headroom[~20 percent safety headroom]
  p1a --> budget
  p1b --> budget
  p1c --> budget
  p2a --> budget
  p2b --> budget
  budget --> reserved --> direct --> headroom

Every pool on every instance draws from the same budget. The audit question is whether the sum of the left side fits inside what remains after the right side is subtracted.

What the failure looks like

The failure mode is “backend connection failure with PostgreSQL max_connections reached as the root cause”:

  1. A burst hits: deploy, cold start, retry storm, or a batch job. Many pools want to grow toward pool_size simultaneously.
  2. PostgreSQL hits max_connections and refuses new logins. Clients connecting directly to PostgreSQL see FATAL: sorry, too many clients already. PgBouncer server connections fail with S: login failed in the PgBouncer log. The PgBouncer log records this as server login failed: <severity> <PostgreSQL error message> or closing because: <PostgreSQL error message>; the exact format varies by code path and version.
  3. Inside PgBouncer, sv_login rises or fluctuates (connections attempting and failing), sv_idle declines, and total server connections stop growing even though demand is unmet.
  4. Existing server connections keep working until they expire (server_lifetime) or error out, so the pool drains gradually rather than dying instantly.
  5. cl_waiting grows, maxwait climbs, and clients start hitting query_wait_timeout (default 120s) or their own application timeouts and retry, which deepens the queue.

The distinguishing feature versus ordinary pool exhaustion: in ordinary exhaustion, total server connections are stable at pool_size and all are busy. In this failure, the total server connection count is declining or stuck below pool_size while PostgreSQL reports itself full. PgBouncer has headroom on paper and cannot spend it.

One nasty property: this failure is triggered by the burst, not by steady state. Everything looks fine at 40% utilization for months, right up until the day every pool wants its full pool_size at once. That is why it has to be caught by the static audit, not by watching dashboards.

How the mismatch creeps in

  • New users or databases added without re-running the sum. Each new (database, user) pair is a whole new pool worth of potential connections.
  • A second PgBouncer instance deployed for HA or rollout, sized against the full max_connections as if it were alone.
  • reserve_pool_size treated as free. Reserve connections are real PostgreSQL connections. They must be in the sum.
  • pool_size set to 0 to mean “use the default” out of habit. Since PgBouncer 1.24.0, a pool_size of 0 is documented to mean unlimited, which removes the per-pool ceiling entirely; older versions treated 0 differently. Check your version before assuming 0 is safe. Specifically, PgBouncer 1.24.0 changed default_pool_size = 0 to mean unlimited (previously it meant “use the default”). Per-database pool_size = 0 follows the same semantic on 1.24.0+.
  • max_db_connections and max_user_connections left at 0 (unlimited), so nothing caps runaway demand across pools for a hot database or user.
  • PostgreSQL max_connections lowered during a “right-sizing” exercise while PgBouncer configs were left alone.

Auditing the arithmetic

This is a read-only procedure. Run it per PostgreSQL instance, and re-run it after any change to users, databases, PgBouncer instance count, or pool configuration.

1. Get the PostgreSQL budget.

# On PostgreSQL itself
psql -Atc "SHOW max_connections;"
psql -Atc "SHOW superuser_reserved_connections;"

# Current usage for context
psql -Atc "SELECT count(*) FROM pg_stat_activity;"

2. Get per-database pool configuration from every PgBouncer instance.

# Run against each PgBouncer instance's admin console
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW DATABASES;"

SHOW DATABASES gives you per-database pool_size, reserve_pool_size, and max_connections (the per-database cap, 0 means unlimited). This is the authoritative view because it includes per-database overrides of default_pool_size.

3. Enumerate the actual pools.

psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW POOLS;"

Each row is one (database, user) pool. Caveat: pools are created on demand. SHOW POOLS shows pairs that have connected; a configured user that has not connected yet has no pool row but will create one under load. For the worst case, count the (database, user) combinations your applications can use, not just the ones currently visible.

4. Compute worst-case demand per instance.

For each instance, for each possible pool: min(pool_size + reserve_pool_size, max_db_connections if set, max_user_connections if set). Sum across pools. Then sum across all instances targeting this PostgreSQL.

5. Compare against the budget.

demand <= 0.8 * max_connections
and demand <= max_connections - superuser_reserved_connections - direct_clients - replication_slots_in_use

If either check fails, you are one burst away from login refusals. Also sanity-check SHOW CONFIG for default_pool_size: if it is 0 and you run PgBouncer 1.24.0 or newer, that may mean unlimited, and the per-pool ceiling is gone. Confirmed: since 1.24.0, default_pool_size = 0 means unlimited connections per pool (PR #1227).

Signals to watch

The audit is the primary defense, but these runtime signals tell you the ceiling is being approached or has been hit:

SignalWhy it mattersWarning sign
PostgreSQL connection count vs max_connectionsThe actual budget consumptionSustained above 80% of max_connections
sv_login per pool (SHOW POOLS)Server connections stuck or failing in the login phasePersistently > 0 with cl_waiting growing
Total server connections vs pool_sizeWhether PgBouncer can reach its configured ceilingTotal declining or pinned below pool_size during a burst
cl_waiting and maxwait (SHOW POOLS)Client impact once server connections cannot be establishedcl_waiting > 0 sustained, maxwait > 15s
PgBouncer log: S: login failedDirect evidence of backend login refusals (log-only, no SHOW counter exists)Any occurrence during load
PostgreSQL log: FATAL: sorry, too many clients alreadyThe limit itself being hitAny occurrence
Reserve pool activationOverflow being drawn, meaning base pools are at ceilingUsage sustained beyond brief spikes

PgBouncer has no error counters in its SHOW commands. The login refusal itself is only visible in logs, so the metric signals (sv_login, cl_waiting, pool totals) are your leading indicators and the logs are confirmation.

Keeping demand under the limit

Once the audit shows over-subscription, the levers, roughly in order of preference:

  • Shrink pool_size to match real concurrency. Most pools need far fewer server connections than their defaults suggest. Use sv_active history: if a pool never exceeds 8 active connections, a pool_size of 20 is pure risk, not capacity.
  • Set max_db_connections and max_user_connections. These cap total server connections per database and per user across all pools, turning unbounded multiplication into a bounded one. They are 0 (unlimited) by default.
  • Consolidate users. Every extra (database, user) pair is a pool. Fewer application roles means fewer pools and a smaller worst case.
  • Raise PostgreSQL max_connections only deliberately. Each connection costs backend memory, and a high limit invites the very pile-up PgBouncer exists to prevent. Raising the limit to make room for oversized pools is usually the wrong direction.
  • Treat reserve_pool_size as part of the budget, not as slack. If reserve usage is sustained, grow the base pool within budget or reduce demand; do not let the reserve mask chronic undersizing.
  • Never rely on pool_size = 0 on 1.24.0+. Unlimited pools make the entire audit meaningless. Confirmed: 1.24.0 changed default_pool_size = 0 to mean unlimited (PR #1227), so a pool_size of 0 no longer means “fall back to the default”.

If you are currently in the incident (PostgreSQL full, PgBouncer login failures), the fastest relief is freeing slots on the PostgreSQL side: terminate idle or runaway direct connections, and reduce demand by pausing non-critical database stanzas in PgBouncer (PAUSE dbname) rather than restarting anything. A PgBouncer restart during this failure makes it worse: every pool re-establishes connections at once, producing a login storm against a backend that is already full.

How Netdata helps

  • PostgreSQL connection utilization against max_connections as a first-class chart, so budget consumption is visible before the FATAL errors start, per backend.
  • PgBouncer pool state collection from the admin console (sv_active, sv_idle, sv_login, cl_waiting, maxwait per pool), letting you see when total server connections stall below pool_size while the backend is full: the signature of this failure.
  • Correlation between PgBouncer sv_login spikes and PostgreSQL connection saturation on one dashboard, which turns “is the pooler broken or is the database full?” into a glance instead of a two-terminal investigation.
  • Per-pool breakdowns so one hot (database, user) pair does not hide inside a healthy aggregate.
  • Alerting on approach, not just arrival: warn when PostgreSQL connection usage crosses a sustained percentage of max_connections, while there is still time to run the audit and shrink pools.