The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-how-it-works-in-production ▌

Operations Guides

How PgBouncer actually works in production: a mental model for operators

Most PgBouncer incidents are the predictable consequence of a handful of internal mechanisms interacting under load: a single-threaded event loop, per-pool FIFO wait queues, fixed-size socket buffers, a small set of server connection states, and a pool mode that decides when server connections change hands. Hold these in your head and the alert thresholds stop being arbitrary numbers.

This article is the model, not the triage. It explains what PgBouncer is doing internally so that when cl_waiting spikes or sv_login won’t drain, you already know which part of the machine is hurting and why.

What it is and why it matters

PgBouncer is a single-threaded, event-driven connection multiplexer between application clients and PostgreSQL. Its entire purpose is to share a small pool of server connections across many client connections, because PostgreSQL backends are expensive and scarce (bounded by max_connections).

The consequence that bites: PgBouncer is itself a consumer of the scarce resource it manages. Every server connection it holds is one PostgreSQL slot. The sum of pool_size across all pools, times the number of PgBouncer instances pointing at the same backend, must fit inside max_connections with room left for superuser access, replication, and direct admin connections. A common planning rule is to keep PgBouncer’s total potential demand under 80% of the backend’s connection limit.

How it works

The event loop: one thread, one core

All I/O runs on a single-threaded libevent loop: every client socket, every server socket, every DNS lookup, every admin command. One thread, one CPU core. This is both the strength and the constraint.

The strength is overhead. An idle connection costs roughly 2KB, so a PgBouncer handling thousands of clients sits at tens of megabytes of RSS and typically under 5% of one core. The constraint is that anything slow and synchronous blocks everything: all pools, all clients, the admin console itself. TLS handshakes at high connection churn, SCRAM authentication storms, and pathological logging are the usual suspects. CPU saturation shows up on a single core while system-wide CPU looks idle, so per-process CPU is the only meaningful view.

Multi-core scaling means running multiple PgBouncer processes sharing a port via so_reuseport. Each process has independent pools and independent stats, so monitoring must aggregate across processes.

Pools: one per (database, user)

PgBouncer does not have “a pool”. It has one pool per unique (database, user) pair, each with its own server connections, its own wait queue, and its own pool_size. This is why aggregate dashboards lie: one saturated pool can be drowning clients while nine others idle, and the average looks green.

Server connection states

Within a pool, every server connection is in exactly one state, and the state names in SHOW POOLS map directly to where the connection is in its lifecycle:

StateMeaning
sv_activeExecuting a query or transaction for a client
sv_idleConnected to PostgreSQL, available for immediate reuse
sv_usedIdle, but unchecked for longer than server_check_delay; needs a health check before reuse
sv_testedCurrently running server_reset_query or server_check_query (server_check_query sends an empty query by default on 1.25+)
sv_loginAuthenticating with PostgreSQL right now
sv_active_cancel / sv_being_canceledForwarding or completing a query cancel request

The lifecycle reads: login -> active -> used -> tested -> idle -> back to active. Healthy pools are mostly sv_idle with bursts of sv_active. A pool stuck at high sv_login with low sv_active is failing to establish backend connections. Connections piling up in sv_used mean the check pipeline is slow.

The wait queue: where pain is measured

When all server connections in a pool are busy, clients enter a FIFO queue (cl_waiting). The age of the oldest waiter is exposed as maxwait (plus maxwait_us) in SHOW POOLS. This is the single most useful measurement PgBouncer gives you: queueing delay in wall-clock seconds, per pool, right now. If no server connection frees up within query_wait_timeout (default 120s), the client is disconnected with an error.

Two nuances matter. First, maxwait counts from when the query was sent, not when the client connected; a session-mode client can sit connected for hours with maxwait at zero. Second, if your application’s own timeout is shorter than query_wait_timeout, the application gives up and retries long before PgBouncer ejects the waiter, and each retry adds a new waiter. That retry amplification is how a slow patch becomes a full pool exhaustion cascade.

If reserve_pool_size > 0 (default 0, disabled), PgBouncer opens extra server connections beyond pool_size once a client has waited longer than reserve_pool_timeout (default 5s). This is overflow capacity for spikes. Sustained reserve usage means the base pool is undersized, not that the feature is working.

Socket buffers: where backpressure lives

Each connection has a pair of socket buffers sized by pkt_buf (default 4096 bytes). Data flows client sbuf -> parse/route -> server sbuf -> PostgreSQL, and back. When a result set exceeds pkt_buf, PgBouncer streams it in chunks, pausing the server read when the client write buffer fills. Nothing is zero-copy; every byte passes through user space. This is why very large result sets degrade pool capacity even when queries are “fast”: the server connection stays busy while PgBouncer dribbles rows to a slow client.

flowchart LR
  subgraph clients["Application clients"]
    c1["client 1"]
    c2["client 2"]
    cn["client N"]
  end
  subgraph pgb["PgBouncer (single thread, one core)"]
    q["FIFO wait queue
(cl_waiting, maxwait)"] pool["pool per (database, user)"] s1["sv_active"] s2["sv_idle / used / tested"] s3["sv_login"] pool --> s1 pool --> s2 pool --> s3 end pg[("PostgreSQL
max_connections slots")] c1 --> pool c2 --> pool cn --> q q -->|pool full| pool s1 --> pg s2 --> pg s3 --> pg

Pool modes: when server connections change hands

The pool mode is the most consequential configuration choice, because it defines when a server connection returns to the pool:

  • Session (the default): the client holds its server connection for the entire session. Turnover is low, multiplexing benefit is minimal, and on return PgBouncer runs server_reset_query (default DISCARD ALL) to scrub session state. That reset runs on every return and occupies the server connection while it executes, but avg_query_time counts client queries rather than this reset.
  • Transaction: the server connection is returned after each transaction. Turnover is high and the multiplexing ratio is where the real wins are, but everything session-scoped breaks or silently misbehaves: temp tables, SET variables, advisory locks, LISTEN/NOTIFY, and protocol-level named prepared statements (unless you enable max_prepared_statements, added for transaction and statement pooling in PgBouncer 1.21).
  • Statement: returned after each statement. Breaks multi-statement transactions outright. Rarely appropriate.

Resources PgBouncer competes for

  • File descriptors: one per client connection, one per server connection, plus listening sockets, the log file, pipe FDs, and admin sockets. The FD ceiling is a hard wall: at the OS limit, PgBouncer cannot accept new connections at all. max_client_conn must be set with margin below the ulimit after accounting for all non-client FDs; a classic failure is max_client_conn = 10000 against a 1024 FD limit.
  • PostgreSQL slots: as above, pool demand must fit inside max_connections.
  • CPU: one core. Fine until TLS termination or auth churn makes the loop the bottleneck.
  • Memory: roughly 2KB per idle connection, more with full buffers and substantially more with TLS state.

Where it shows up in production

The failure archetypes all fall out of the machinery above:

  1. Pool exhaustion cascade. All server connections busy, cl_waiting grows, maxwait climbs past application timeouts, retries deepen the queue. The most common PgBouncer incident, and almost always rooted in either slow backend queries (avg_query_time up) or connections held too long (idle-in-transaction: avg_xact_time much larger than avg_query_time).
  2. Client connection limit. max_client_conn reached; new connections refused immediately with no more connections allowed (max_client_conn). No queueing, just refusal. Frequently the FD limit in disguise.
  3. FD exhaustion. Same wall, one layer down. An exhausted listener logs an accept() failure and suspends the pooler; new server-side socket() calls also fail, so new connections and replacements stop rather than crashing the process.
  4. Backend unreachable. PostgreSQL down, network partition, or stale DNS. Signature: sv_login elevated, total server connections declining, cl_waiting growing. PgBouncer’s own DNS cache (check SHOW DNS_HOSTS) can keep it dialing a dead primary until the TTL expires.
  5. Pool mode mismatch. Application relies on session state in transaction mode. Every PgBouncer metric looks healthy; only application error logs (“prepared statement does not exist”, missing temp tables) reveal it. The most insidious failure because it impersonates an application bug.
  6. Connection leak. Sessions that never release (long idle transactions in session mode, abandoned clients) exhaust the pool under light load.
  7. Administrative state. PAUSE stops new query routing while existing transactions finish (cl_waiting spikes, sv_active drains to zero); DISABLE refuses new clients while existing ones keep working. Looks exactly like an outage on the metrics, so every PgBouncer alert should check paused/disabled from SHOW DATABASES before escalating.

Tradeoffs and when this matters

The central tradeoff is pooling efficiency versus session semantics. Transaction mode gives you the multiplexing ratio that justifies running PgBouncer at all, at the cost of every session-scoped feature. Session mode is safe for anything but barely reduces connection counts. Teams switch to transaction mode for efficiency and discover the semantic cost during the next incident, because PgBouncer has no metric for “your session state just vanished”.

The second tradeoff is the single thread. You get very low per-connection overhead and no lock contention, and in exchange you accept that one stalled synchronous operation freezes every pool at once, and that scaling past one core means multiple processes with independent state.

Third: server_reset_query correctness versus latency. DISCARD ALL on every session-mode return is what makes connection reuse safe; it is also real backend work that avg_query_time does not isolate, so prolonged connection turnover can coexist with ordinary-looking query averages.

Finally, the instrumentation contract: SHOW commands give you snapshots and aggregates, and nothing else. There are zero error counters in the admin console. Authentication failures, connection refusals, and timeout events exist only in the log file. Stats are rolling averages that smooth sub-period spikes, cumulative counters reset on restart, and column positions in SHOW STATS shift between versions, so reference columns by name.

Signals to watch in production

Per-pool, always; aggregates hide the drowning pool.

SignalWhy it mattersWarning sign
cl_waiting (SHOW POOLS)Clients blocked on a server connection; the primary saturation signalAny value sustained > 60s, database not paused
maxwait / maxwait_us (SHOW POOLS)Age of the oldest waiter; user-facing pain in seconds> 5s impacting; > 15s likely causing app failures
sv_active vs pool_sizePool utilization; the leading indicator before queueing starts> 85% sustained; 100% means the next request queues
avg_wait_time (SHOW STATS_AVERAGES)Latency PgBouncer itself injects, distinct from database time> 100ms sustained
avg_query_time vs avg_xact_timeSeparates “backend slow” from “connections held idle in transaction”avg_xact_time » avg_query_time
sv_login (SHOW POOLS)Backend connection establishment healthSustained > 0 with sv_idle draining
used_clients vs max_client_connProximity to the hard client refusal wall> 80% sustained
Process CPU and FD countSingle-core saturation and the real connection ceilingCPU > 70% of one core; FDs > 80% of ulimit

The correlation that shortens almost every incident: avg_wait_time high with avg_query_time low means the pool is too small or connections are held too long. avg_wait_time high with avg_query_time high means PostgreSQL is the root cause and PgBouncer is the messenger. Check both before touching either.

How Netdata helps

  • Netdata’s PgBouncer collector queries the admin console directly and charts cl_waiting, maxwait, and the full sv_active/sv_idle/sv_used/sv_tested/sv_login breakdown per pool, so you see which (database, user) pool is saturated instead of an aggregate that hides it.
  • It plots avg_wait_time alongside avg_query_time and avg_xact_time, exactly the pairing needed to attribute latency to the pool, the backend, or idle-in-transaction clients.
  • Per-process CPU and file descriptor usage for the PgBouncer process are collected from the host at per-second resolution, so event-loop saturation and FD-approaching-ulimit show up on the same dashboard as pool metrics.
  • Because PgBouncer exposes no error counters, pairing the metrics with log-based alerts on no more connections allowed, query_wait_timeout, and login failure strings closes the blind spot the SHOW commands leave open.
  • Anomaly detection on cl_waiting and maxwait per pool catches slow-building saturation (growing transaction times eating headroom) before the queue crosses alert thresholds.