The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-no-more-connections-allowed ▌

Operations Guides

PgBouncer no more connections allowed (max_client_conn): the front door is full

Your application logs fill with connection errors, and every new connection attempt to PgBouncer fails immediately with:

ERROR: no more connections allowed (max_client_conn)

This is not pool exhaustion. The client never gets in the door. There is no queue, no wait, no query_wait_timeout. PgBouncer counts the client connection, sees it would exceed max_client_conn, and refuses it on the spot. Existing clients keep working; only new ones are turned away.

The confusing part is that the pool itself may be completely healthy. sv_active can be well below pool_size, cl_waiting can be zero, and avg_wait_time can be flat. PostgreSQL is fine, the pools are fine, and yet applications cannot connect. The bottleneck is the front door, not the backend.

What this means

max_client_conn is a hard ceiling on the number of simultaneous client connections PgBouncer will accept (default: 100). When used_clients reaches it, every additional connection attempt is rejected, logged by PgBouncer as no more connections allowed (max_client_conn).

Two properties make this failure mode distinct:

  • No queuing. In pool exhaustion, the client connects successfully and then waits in cl_waiting for a server connection. Here, the client is refused before any of that machinery engages. That is why your pool metrics look green during the incident.
  • No metric. PgBouncer has no SHOW counter for refused connections. The refusal exists only in the log file and in client-side errors. If you monitor only SHOW output, the first signal you get is applications failing.

Admin connections to the special pgbouncer database are exempt from max_client_conn, so you can still get in and diagnose while clients are being refused.

flowchart TD
  A[New client connection] --> B{used_clients at max_client_conn?}
  B -- Yes --> C[Rejected instantly: no more connections allowed]
  B -- No --> D{Server connection available in pool?}
  D -- Yes --> E[Query executes]
  D -- No --> F[Client joins cl_waiting queue]
  F --> G[query_wait_timeout if wait is too long]

Common causes

CauseWhat it looks likeFirst thing to check
Application connection leakused_clients climbs steadily over hours or days, never returns to baselineSHOW CLIENTS for connections with very old connect_time
Uncoordinated deploymentsused_clients steps up each time a new app version or instance rolls outClient count by source IP in SHOW CLIENTS vs expected instance count x pool size
max_client_conn exceeds the FD limitRefusals start well below the configured max_client_connPgBouncer startup log line reporting the FD limit and effective max_client_conn
Retry storm during an incidentSudden spike in used_clients while another failure (pool exhaustion, backend down) is in progressWhether cl_waiting or sv_login was elevated just before refusals began
Limit simply too smallSlow organic growth; utilization sits above 80% at peak for weeksused_clients / max_client_conn trend over time

Quick checks

All commands are read-only. Adjust host, port, and user for your environment.

# 1. Confirm refusals in the log (the only place they are recorded)
grep -c "no more connections allowed" /var/log/pgbouncer/pgbouncer.log
tail -200 /var/log/pgbouncer/pgbouncer.log | grep "no more connections allowed"

# 2. Check current client utilization and the configured limit
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW LISTS;"
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW CONFIG;" | grep max_client_conn

# 3. Verify the pools are actually healthy (distinguishes from pool exhaustion)
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW POOLS;"

# 4. Look for leaked client connections (old connect_time, idle state)
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -c "SHOW CLIENTS;"

# 5. Check the actual file descriptor situation
PGBPID=$(pgrep -f pgbouncer)
ls /proc/$PGBPID/fd | wc -l
grep "Max open files" /proc/$PGBPID/limits

Interpreting SHOW LISTS: used_clients is the count of active client objects. free_clients is the pre-allocated slab of client structures and drops toward zero as clients connect; when used_clients reaches max_client_conn, refusals begin. Treat used_clients / max_client_conn as the utilization signal, with free_clients as corroboration.

How to diagnose it

  1. Confirm the symptom class. Grep the log for no more connections allowed. If it is present, you are at the client ceiling. If instead you see query_wait_timeout or clients reporting long waits after connecting, you are dealing with pool exhaustion, a different incident with different fixes. See PgBouncer pool exhaustion.

  2. Check pool health to rule out compounding causes. Run SHOW POOLS. If sv_active is well below pool_size and cl_waiting is zero, the pools are healthy and the problem is purely the front door. If cl_waiting was high before refusals began, the client ceiling may be a secondary effect: applications timed out waiting, retried, opened new connections, and piled onto max_client_conn. Fix the pool problem first.

  3. Identify who holds the connections. Run SHOW CLIENTS and group by addr (source IP) and connect_time. Leaked connections show up as long-lived connections from application hosts that far outnumber the configured application pool size. A deployment problem shows up as more source IPs than you expect, or roughly double the expected connections per instance (old and new versions running side by side).

  4. Check the FD limit before assuming the config value is real. PgBouncer needs one FD per client connection, plus one per server connection, plus listening sockets, the log file, pipe FDs, and admin sockets. If the OS limit (Max open files in /proc/<pid>/limits, or LimitNOFILE under systemd) is lower than what max_client_conn requires, PgBouncer may lower the effective limit at startup or hit the FD wall first. Check the startup log: PgBouncer logs its detected FD limit and the effective max_client_conn near the top of the log after each start. If the logged value is lower than your configured value, the FD limit is the real ceiling.

  5. Check for other client limits. max_db_client_connections (per database) and max_user_client_connections (per user) also reject new clients when exceeded, and both default to 0 (unlimited). If they are set, check SHOW CONFIG to see whether one of them is the binding constraint. Per-database and per-user limits produce distinct messages: client connections exceeded (max_db_client_connections) and client connections exceeded (max_user_client_connections), respectively. Both differ from the global limit’s no more connections allowed (max_client_conn).

Metrics and signals to monitor

SignalWhy it mattersWarning sign
used_clients / max_client_conn (SHOW LISTS + SHOW CONFIG)Utilization ratio for the front door; failure is binary at 100%Above 80% sustained
free_clients (SHOW LISTS)Remaining pre-allocated client slots; corroborates the ratioBelow 10% of max_client_conn
Log: no more connections allowedThe only record of actual refusalsAny sustained occurrence
cl_waiting and maxwait (SHOW POOLS)Tells you whether a pool problem is driving the retry stormNon-zero before or during refusals
Process FD count vs Max open files (/proc/<pid>/)FDs can run out before max_client_conn doesAbove 80% of limit
SHOW CLIENTS oldest connect_timeLeak detection: connections older than any legitimate sessionConnections hours or days old

Fixes

Fix the application connection leak

The most common root cause. Connections are opened and never closed: a code path that skips close() on error, an ORM pool with no maximum lifetime, or a worker that leaks a connection per job.

  • Identify the offenders from SHOW CLIENTS by source IP and connection age, then fix the application code or pool configuration (maximum lifetime, idle timeout, pool size ceiling).
  • As a mitigation while the fix rolls out, client_idle_timeout can reclaim idle client connections. Understand what it will kill before enabling it: any client that legitimately sits idle longer than the timeout gets disconnected.
  • Do not restart PgBouncer to clear leaked connections as a first move. Restarting drops all server connections and triggers a thundering herd of reconnects against PostgreSQL, and the leak will refill the ceiling within hours anyway.

Raise max_client_conn (safely)

If demand is legitimate, raising the limit is correct, but only after the FD math checks out.

  • The FD budget is roughly: max_client_conn + total server connections + listening sockets + log FD + pipe FDs + admin sockets. Keep at least 20% headroom beyond that sum.
  • If the OS limit needs raising, set LimitNOFILE in the systemd unit (or the ulimit for the PgBouncer user). The FD limit is fixed at process start, so this change requires a PgBouncer restart to take effect.
  • The config change itself (max_client_conn in pgbouncer.ini) applies with RELOAD, but it cannot exceed what the FD limit allows.
  • Plan for rolling restarts with multiple PgBouncer processes on so_reuseport if a full restart is disruptive.

Coordinate deployments

If used_clients steps up with each deploy, the problem is arithmetic: application instances x connections per instance is not accounted for in max_client_conn. Compute the worst case (all current instances plus a full extra batch during a rolling deploy) and size for that, or cap per-instance pool sizes so the product fits.

Break a retry storm

If refusals are a symptom of pool exhaustion plus application retries, raising max_client_conn alone will not help; it just lets more clients queue. Fix the pool saturation first (slow queries, idle-in-transaction, undersized pool_size), and reduce application-side retry aggressiveness. See PgBouncer query_wait_timeout and PgBouncer maxwait high.

Prevention

  • Alert on the ratio, not the refusal. used_clients / max_client_conn above 80% sustained is a ticket; above 95% is urgent. The refusal itself is only visible in logs, so the ratio is your early warning.
  • Alert on FD headroom too. FD count above 80% of the limit catches the case where the real ceiling is lower than the configured one.
  • Verify the startup log after every restart. Confirm the logged effective max_client_conn matches what you configured. This catches silent FD capping immediately.
  • Budget connections as code. Every new application deployment should declare its connection footprint. Track max_client_conn as a capacity line item, not a set-and-forget config value. See PgBouncer capacity planning.
  • Ship PgBouncer logs somewhere queryable. Since refusals, auth failures, and timeouts exist only in the log, log aggregation is not optional for this service.

How Netdata helps

  • Netdata collects SHOW LISTS and SHOW CONFIG data, so used_clients against max_client_conn is graphed continuously, letting you see the slow leak or the deployment step-change instead of discovering it at 100%.
  • Pool metrics (sv_active, sv_idle, cl_waiting, maxwait) on the same dashboard let you confirm in seconds that this is a front-door problem and not pool exhaustion, the key diagnostic branch.
  • Host-level file descriptor collection alongside PgBouncer metrics exposes the FD-vs-config gap that silently caps max_client_conn.
  • Because PgBouncer exposes no refusal counter, correlating the utilization ratio crossing 100% with the moment application error rates spiked is how you reconstruct the incident timeline; per-second collection makes that correlation tight.
  • Anomaly detection on used_clients flags slow leaks that sit below static thresholds for weeks before they become an incident.