The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-database-paused ▌

Operations Guides

PgBouncer database paused or disabled: maintenance state that looks like an outage

cl_waiting is climbing, sv_active just dropped to zero, maxwait is ticking upward. The pattern looks identical to pool exhaustion or a backend failure. But if someone is running planned PostgreSQL maintenance with PAUSE, the metrics are behaving as designed: PAUSE stops new query routing, existing transactions finish, server connections close, and every pending client queues. From the client’s perspective, it looks like an outage. The difference is that it is planned and reversible with RESUME.

PgBouncer has three administrative states that alter traffic flow: PAUSE, DISABLE, and SUSPEND. Each produces a distinct metric signature, and all can trigger saturation alerts if your monitoring does not account for them. The paused and disabled columns in SHOW DATABASES, and the global state visible in SHOW STATE (added in PgBouncer 1.19.0), are the first things to check before escalating any PgBouncer saturation signal.

A PAUSE without a matching RESUME freezes traffic indefinitely. There is no auto-timeout.

What this means

PgBouncer’s administrative commands support controlled maintenance: zero-downtime PostgreSQL upgrades, host cutovers, connection draining before restarts. They are intentional features, not bugs. But because their metric signatures overlap with real failure modes, they are a common source of false-positive pages.

The playbook’s severity definitions for cl_waiting and maxwait both include an explicit condition: the database must NOT be paused or disabled. Without this condition, every planned maintenance window becomes an incident.

Three commands create maintenance states:

  • PAUSE [db]: Stops routing new queries to server connections. Existing transactions are allowed to complete. Once all server connections are released, sv_active drops to zero and cl_waiting spikes as pending clients queue. The command blocks until all server connections are released. RESUME [db] reverses it.

  • DISABLE db: Rejects all new client connections to the specified database. Existing client connections continue working and are NOT closed. This drains traffic by preventing new arrivals while letting in-flight work finish. ENABLE db reverses it.

  • SUSPEND: A global command (not per-database) that flushes all socket buffers and stops PgBouncer from listening for data on them. The command blocks until all buffers are empty. New client connections wait. This is the most severe administrative state because it stops data on client and server sockets; the admin pool is deliberately exempt so RESUME remains reachable. RESUME reverses it.

flowchart TD
    A["Alert: cl_waiting rising, sv_active dropping"] --> B{"SHOW DATABASES: paused = 1?"}
    B -- Yes --> C["Maintenance: PAUSE on this db"]
    B -- No --> D{"SHOW DATABASES: disabled = 1?"}
    D -- Yes --> E["Maintenance: DISABLE on this db"]
    D -- No --> F{"SHOW STATE available?"}
    F -- Yes --> G{"SHOW STATE = suspended?"}
    G -- Yes --> H["Maintenance: SUSPEND (global)"]
    G -- "active" --> I["Real incident: investigate further"]
    F -- "Pre-1.19.0: check logs" --> J["Grep SUSPEND in pgbouncer log"]
    C --> K["Run RESUME db when maintenance is done"]
    E --> L["Run ENABLE db to accept new clients"]
    H --> M["Run RESUME to unfreeze all I/O"]

The three maintenance states compared

StateScopeWhat stopsWhat keeps workingMetric signatureHow to reverse
PAUSE [db]Per-databaseNew query routing to server connectionsExisting transactions until they completesv_active drops to 0, cl_waiting spikes, sv_idle drops to 0RESUME [db]
DISABLE dbPer-databaseNew client connectionsAll existing client connectionsNew connection errors, existing clients unaffectedENABLE db
SUSPENDGlobal (all databases)Client and server socket I/O; buffers must flushThe admin console (the admin pool is exempt)New database work waits; SHOW STATE reports suspendedRESUME

The critical operational difference: PAUSE lets existing clients finish their work before the pool drains. DISABLE lets existing clients keep working indefinitely (new connections are rejected, connected clients can still query). SUSPEND freezes database I/O; the admin console remains available for RESUME.

Quick checks

Run these read-only checks before escalating any PgBouncer saturation alert. Adjust the socket path (-h) and port to match your deployment.

# Check paused/disabled flags per database (reference columns by name, not position)
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "SHOW DATABASES;"

# Check global PgBouncer state (requires 1.19.0+)
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "SHOW STATE;"

# See which pools have waiting clients and how long they have waited
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "SHOW POOLS;"

# Check admin console responsiveness (a hang is not expected during SUSPEND because the admin pool is exempt)
time psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "SHOW LISTS;" > /dev/null

# Look for recent administrative commands in the log
grep -i "PAUSE\|DISABLE\|SUSPEND\|RESUME\|ENABLE" /var/log/pgbouncer/pgbouncer.log | tail -20

The paused and disabled columns in SHOW DATABASES output use 1 for active and 0 for normal. Column positions are version-dependent: PgBouncer’s column layouts shift across releases as new fields are added. current_client_connections was added to SHOW USERS and SHOW DATABASES in PgBouncer 1.24.0.

On PgBouncer versions before 1.19.0, there is no SHOW STATE command. Detect a global SUSPEND from the log or SHOW STATE on supported versions; the admin console is not itself suspended.

How to diagnose it

  1. Check SHOW DATABASES first. Look at the paused and disabled columns for the database showing saturation. If paused = 1, someone ran PAUSE and has not yet run RESUME. If disabled = 1, someone ran DISABLE and has not yet run ENABLE.

  2. Check SHOW STATE for global states. If the output is suspended, PgBouncer is in a global SUSPEND. This is distinct from per-database PAUSE and does not show up in SHOW DATABASES flags.

  3. Check the log for context. Administrative commands are logged when issued through the admin console. Look for PAUSE, DISABLE, SUSPEND, RESUME, or ENABLE entries to determine who initiated the state and when.

  4. Cross-reference with your change calendar. If a PostgreSQL upgrade, host cutover, or connection draining procedure is in progress, the paused state is expected.

  5. If none of the above apply, it is a real incident. The metrics are telling you about a genuine pool exhaustion cascade, backend failure, or connection leak. Proceed with the normal diagnostic path: check avg_query_time for backend slowdown, check sv_login for backend connectivity, check SHOW SERVERS for stuck connections.

Metrics during maintenance vs real outage

Several metric patterns are identical whether PgBouncer is paused or experiencing a real failure. The table below maps the overlapping signals to the distinguishing check.

SignalDuring PAUSEDuring real outageHow to tell them apart
cl_waiting risingYes, all pending clients queueYes, same mechanismSHOW DATABASES: paused = 1
sv_active dropping to 0Yes, after transactions completePossible, if backend is downCheck sv_login: rising in backend failure, zero during PAUSE
sv_idle dropping to 0Yes, connections released during PAUSEPossible, if pool drainedSHOW DATABASES: paused = 1
maxwait increasingYes, waiters age in queueYes, sameSHOW DATABASES: paused = 1
avg_query_time elevatedNo, queries are not executingPossibly, if backend is slowDuring PAUSE, query time should be stable or dropping
New connections rejectedNo (during PAUSE); Yes (during DISABLE)Yes, if max_client_conn hit or FD exhaustionSHOW DATABASES: disabled = 1 vs SHOW LISTS: free_clients = 0
Admin console responsiveYesYes (unless event loop stalled)If admin console hangs, suspect SUSPEND or event loop stall

The key insight: during PAUSE, avg_query_time and avg_xact_time should not be elevated because queries are not running. If you see high cl_waiting but normal avg_query_time, and the database is paused, the pattern is consistent with maintenance. If you see high cl_waiting with elevated avg_query_time, the backend is the problem regardless of the paused state.

Exiting maintenance states

Each administrative state has a specific reversal command. Running the wrong one does nothing.

Resuming a paused database:

# Resume a specific database
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "RESUME mydb;"

If PAUSE was run without a database argument, the global pause state is active and RESUME without an argument clears it. Per-database pauses are cleared with RESUME <db>.

Enabling a disabled database:

# Re-enable client connections
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "ENABLE mydb;"

Resuming from SUSPEND:

# Resume all I/O globally
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "RESUME;"

Warning: send RESUME to an admin console session when PgBouncer is in SUSPEND. The admin pool is deliberately not suspended so the control plane remains reachable; if it is not reachable, diagnose that separately.

After running any reversal command, verify the state has cleared:

# Confirm paused = 0 and disabled = 0
psql -h /var/run/postgresql -p 6432 -U pgbouncer pgbouncer -c "SHOW DATABASES;"

Common pitfalls

No auto-timeout on PAUSE. A PAUSE without a matching RESUME freezes traffic indefinitely. There is no built-in timeout that automatically resumes a paused database. If the operator who ran PAUSE loses their session, forgets, or is interrupted, the database stays paused until someone manually runs RESUME.

A per-database pause and the global pause state are distinct. PAUSE mydb sets that database’s pause flag; PAUSE sets the global pause state. RESUME mydb clears only the database flag; RESUME clears the global state. Run both, in the order your procedure paused them, when both commands were used.

PAUSE in session pooling mode may never complete. The documented contract remains that PAUSE waits for every server connection to be released; in session mode that means waiting for the client to disconnect. This is current behavior, not a version-specific bug to work around by retrying PAUSE.

Concurrent PAUSE and RESUME can deadlock. If RESUME is issued while a PAUSE is still in progress (waiting for connections to drain), the RESUME may return immediately but the PAUSE command may never return, leaving the system in an inconsistent state. PgBouncer issue #715 reports this behavior and remains open; no fixed release is documented.

Cross-database queueing after PAUSE/RELOAD. On PgBouncer versions before 1.24.0, a RELOAD recycled TLS server connections across databases. 1.24.0 changed this so TLS connections are recycled only when TLS settings change; non-TLS configuration reloads therefore no longer cause that reconnect wave on unrelated databases.

RELOAD is not RESUME. Operators sometimes run RELOAD expecting it to resume traffic after a PAUSE. It does not. RELOAD re-reads the configuration file. RESUME is the only command that reverses a PAUSE.

Prevention

Make paused/disabled state a first-class signal in your monitoring. Every PgBouncer saturation alert for cl_waiting or maxwait must include a condition that checks whether the database is paused or disabled. If it is, suppress the alert. The playbook’s severity definitions already encode this: TICKET only fires if cl_waiting > 0 sustained AND the database is NOT paused/disabled.

Automate the PAUSE/RESUME lifecycle. If you use PAUSE for PostgreSQL maintenance (upgrades, restarts, host cutovers), wrap it in a script that:

  1. Runs PAUSE db.
  2. Performs the maintenance.
  3. Runs RESUME db.
  4. Verifies that SHOW DATABASES reports paused = 0.

If step 3 fails or is skipped, step 4 makes the problem visible immediately.

Document who can run administrative commands. PAUSE, DISABLE, SUSPEND, KILL, and RESUME are available to users listed in admin_users in PgBouncer’s configuration. Restrict this list to the smallest set of operators and automation service accounts. Use stats_users (read-only SHOW commands) for monitoring tools and dashboards.

Prefer DISABLE for connection draining. DISABLE is safer for scenarios where you want to stop new traffic without waiting for existing work to complete. Existing clients keep working, and there is no risk of PAUSE hanging on long-lived sessions in session pooling mode. Use PAUSE only when you need all server connections fully released (for example, before a host cutover where the old backend address will stop responding).

Monitoring integration

The paused and disabled state flags from SHOW DATABASES should be part of any PgBouncer alert condition. Alert rules for cl_waiting or maxwait should suppress when the database is paused or disabled.

Per-second collection of cl_waiting, sv_active, and maxwait lets you pinpoint the exact moment a PAUSE begins (sharp drop in sv_active, immediate spike in cl_waiting) and when RESUME takes effect (queue drains, sv_active returns to baseline).

Correlating the state flag with wait queue metrics in a single dashboard distinguishes planned maintenance from unexpected pool exhaustion without manual SHOW DATABASES queries.