The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-listen-notify-transaction-mode ▌

Operations Guides

PgBouncer LISTEN/NOTIFY not working: why pub/sub needs session pooling

Your application issues LISTEN job_events, the command succeeds, and PostgreSQL’s logs show NOTIFY firing on schedule. But the listener never receives anything. No error in the application. No error in PgBouncer. No metric anywhere that moves. The feature worked in staging, worked before you put PgBouncer in front of the database, and now it silently does nothing.

This is the pool mode mismatch failure pattern, and LISTEN/NOTIFY is its most confusing variant because the failure is completely silent. Unlike prepared statements (which at least produce “prepared statement does not exist” errors), a lost LISTEN registration produces no error at all. The notification is delivered to a backend connection your client no longer holds, or to whichever client happens to hold that connection next.

The root cause is almost always the same: the listener connection is going through a pool running in transaction pooling mode. The fix is routing listener connections through session pooling, or taking them off the pooled path entirely. The broader mental model is covered in How PgBouncer actually works in production; this article is narrowly about confirming and fixing the LISTEN/NOTIFY case.

What this means

PostgreSQL’s LISTEN/NOTIFY is session-level state. When a client runs LISTEN channel_name, the registration lives on that specific backend connection, the server process on the PostgreSQL side. Notifications for that channel are delivered to that backend, which forwards them to the client attached to it. The registration survives as long as the session survives.

Transaction pooling breaks the assumption underneath all of this. In transaction mode, PgBouncer holds the server connection only for the duration of one transaction, then returns it to the pool and may hand it to a completely different client. From the client’s perspective the connection looks continuous. From PostgreSQL’s perspective, the session state the client created, including the LISTEN registration, is now attached to a server connection that client no longer owns.

The sequence:

flowchart TD
    A[Client sends LISTEN via PgBouncer] --> B[PgBouncer assigns server conn S1]
    B --> C[LISTEN registered on backend S1]
    C --> D[Transaction ends - S1 returned to pool]
    D --> E[S1 reassigned to another client]
    E --> F[NOTIFY arrives at backend S1]
    F --> G[Delivered to whoever holds S1 - never to the listener]

Note the asymmetry that trips people up during testing: NOTIFY works fine through transaction pooling. It is a single statement with no session state behind it, so any backend connection can execute it. LISTEN does not. Operators test the publish path, see notifications flowing, and conclude pub/sub works. It does not. PgBouncer’s own SQL feature map documents this explicitly: LISTEN is supported with session pooling and never with transaction pooling, while NOTIFY works with both.

The second trap: this often “works” in development or under light load. If only one client is using the pool, PgBouncer may keep handing back the same server connection, and the registration appears to survive. The failure only shows up under concurrent load in production, which is why it survives testing.

Common causes

CauseWhat it looks likeFirst thing to check
Listener connected through a transaction-mode poolLISTEN succeeds, NOTIFY fires in PG logs, listener receives nothingSHOW CONFIG; and per-database pool_mode
Whole deployment switched to transaction mode without auditing app codeMultiple session-dependent features broken (LISTEN, advisory locks, temp tables, SET variables)Application error logs for “does not exist” or missing session state
Listener shares a connection pool with regular query trafficNotifications arrive intermittently, worse under loadWhether the listener uses its own connection or the shared app pool
ORM or client library transparently routes all connections through one DSNEverything points at the same PgBouncer database aliasThe listener’s connection string versus the rest of the app
Client-side pooler also reclaims connectionsSame symptom even in session mode if the app pool recycles the physical connectionClient library pool settings for connection lifetime and reuse

One cause that is NOT PgBouncer: if the NOTIFY runs inside a transaction that rolls back, the notification is never sent. That is standard PostgreSQL behavior, independent of pooling. Rule it out before blaming the pooler.

Quick checks

All read-only. Run them against the PgBouncer admin console.

# 1. Check the global pool mode
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW CONFIG;" | grep pool_mode
# 2. Check per-database pool settings (per-database overrides beat the global default)
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW DATABASES;"
# 3. Confirm pool turnover is happening on the listener's pool
#    High sv_active churn with the listener connected means the backend is being recycled
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW POOLS;"
# 4. Bypass PgBouncer entirely: run LISTEN/NOTIFY against PostgreSQL directly.
#    If this works (it will), the pooler path is the problem.
#    Do this in two interactive psql sessions, not psql -c: LISTEN only
#    receives notifications while the session stays open.
#    Session A: psql -h <postgres-host> -U <user> -d <db>
#      LISTEN test_chan;
#    Session B: NOTIFY test_chan, 'hello';
#    Session A should print the asynchronous notification.
# 5. Check what the listener connection actually looks like from PgBouncer's side
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW CLIENTS;"

How to diagnose it

  1. Confirm the symptom is a lost registration, not a missing NOTIFY. Check PostgreSQL’s logs or run a known-good listener connected directly to PostgreSQL (check 4 above). If a direct listener receives the notification and the pooled listener does not, you have confirmed the pooling path is dropping it.

  2. Identify which PgBouncer database entry the listener uses. Get the listener’s connection string from the application config. Map it to a [databases] entry in pgbouncer.ini.

  3. Check the effective pool_mode for that entry. Pool mode can be set globally, per database in the [databases] section, and per user in the [users] section. SHOW CONFIG; shows the global default, SHOW DATABASES; shows per-database settings, and SHOW USERS; shows per-user settings. Any one of these set to transaction (or statement) on the listener’s path is the bug.

  4. Rule out a client-side pool doing the same thing. Some client libraries run their own connection pool in front of PgBouncer and hand the listener a different physical connection after idle periods. If the listener uses a shared application pool, the registration can be lost client-side even with PgBouncer in session mode. The listener must hold one dedicated, long-lived connection end to end.

  5. Verify with a load test. Single-client testing can pass because the same backend gets reused. Reproduce with concurrent traffic on the pool and watch the notification stop arriving.

Metrics and signals to monitor

The hard truth: there is no PgBouncer metric for missed notifications. PgBouncer’s SHOW commands expose pool and traffic counters, not session-state loss, so detection is indirect.

SignalWhy it mattersWarning sign
pool_mode (SHOW CONFIG / SHOW DATABASES / SHOW USERS)The actual root cause; should be part of config-drift checkstransaction on a database entry serving a listener
Server assignment rate (avg_server_assignment_count) (introduced in 1.23.0)Pool turnover; a listener’s pool should show near-zero turnover in session modeHigh assignment rate on the pool the listener uses
SHOW CLIENTS connection age for the listenerThe listener should hold one old, stable connectionListener connection cycling (young connect_time repeatedly)
Application-side notification lagThe only true measure of the failureRising time between NOTIFY and handler execution

The reliable detector is an end-to-end canary: have the application timestamp when it sends a NOTIFY and when the listener processes it, and alert on the lag or on missed heartbeats. Infrastructure metrics will never see this failure.

Fixes

Route listeners through a session-mode database alias

The standard fix. Add a second entry in pgbouncer.ini pointing at the same PostgreSQL database but with session pooling:

[databases]
mydb = host=pg-primary dbname=mydb pool_mode=transaction
mydb_session = host=pg-primary dbname=mydb pool_mode=session

Point only the listener connection at mydb_session. Everything else keeps the efficiency of transaction pooling. Issue a RELOAD in the admin console to apply. Tradeoffs: the listener pins one server connection for its entire lifetime, consuming one slot from its pool and one PostgreSQL connection slot. Size the session-mode pool accordingly, typically small, since only listeners use it. If you have many listener processes, each holds a backend connection, which partially defeats pooling for that workload. That is the cost of correctness here.

The same approach works per user if you would rather separate by role than by database alias.

Bypass PgBouncer for the listener

Give each process that needs LISTEN/NOTIFY one dedicated connection straight to PostgreSQL, and route all other traffic through PgBouncer in transaction mode. This is the pattern several production queue systems document: pooled connections for inserts and job fetches, one raw connection for the notification listener.

Tradeoffs: the listener’s connection is not protected by the pooler, so you handle reconnection, failover, and backend slot accounting yourself. Keep the count of direct connections small and include them in your PostgreSQL max_connections budget.

Move the wakeup off LISTEN/NOTIFY entirely

If neither option fits (for example, a serverless runtime where no process can hold a long-lived connection), the durable alternative is to treat NOTIFY as a hint only: workers poll an outbox or queue table on a short interval, and notifications just make the poll happen sooner. A missed notification then costs latency, not correctness. This is an application change, so it is a redesign rather than a fix, but it removes the session-state dependency completely.

Prevention

  • Audit before switching pool modes. Session-dependent features (LISTEN/NOTIFY, advisory locks, temp tables, prepared statements, SET variables) must be inventoried before any database or user moves to transaction pooling. The failure is silent, so “deploy and watch” does not find it.
  • Test under concurrency. A single-client smoke test passes even with the wrong pool mode. LISTEN/NOTIFY verification must run while other clients churn the pool.
  • Treat pool_mode as a reviewed config surface. Per-database and per-user overrides in pgbouncer.ini should be visible in code review, and SHOW CONFIG / SHOW DATABASES output should be diffed after every RELOAD.
  • Run a notification canary. A periodic NOTIFY with an end-to-end lag measurement is the only alert that catches this class of regression, including regressions introduced by client library upgrades.

How Netdata helps

Netdata cannot see a missed notification directly, because no counter exists for it in PgBouncer. What it can do is make the surrounding state visible so you confirm the root cause in minutes instead of hours:

  • Pool state per (database, user): sv_active, sv_idle, and client counts per pool, so you can see that the listener’s pool is churning connections at transaction-mode rates.
  • Pool turnover signals: server assignment rate as a continuous time series, so “the backend connection keeps getting recycled” is a graph, not a theory.
  • Config and state visibility: collecting SHOW POOLS, SHOW DATABASES, and SHOW CONFIG over time means pool_mode changes show up as deploy-correlated events rather than mysteries.
  • Correlated saturation context: if a misdiagnosed “fix” (like moving everything to session mode) starts exhausting the pool, cl_waiting, maxwait, and avg_wait_time show the tradeoff immediately. See PgBouncer pool exhaustion for that failure mode.
  • Restart and reload correlation: stats resets and config reloads are visible against metric continuity, which matters because a RELOAD is exactly when a pool_mode regression gets introduced.

The honest limitation: the definitive detection of this bug remains an application-side notification-lag canary. Use Netdata for the infrastructure half of the correlation.