The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / pgbouncer / pgbouncer-max-client-conn-tuning ▌

Operations Guides

PgBouncer max_client_conn tuning: setting the client limit against real FD headroom

Applications fail with connection errors, but PgBouncer’s pools look healthy: sv_active is well below pool_size, cl_waiting is zero, and PostgreSQL is idle. The log tells the real story: no more connections allowed (max_client_conn) or accept failed: Too many open files. Clients are being refused before they ever reach a pool.

The usual root cause is a mismatch between two limits operators treat as one. max_client_conn is a configuration value. The OS file-descriptor limit (ulimit -n) is a hard kernel ceiling. PgBouncer cannot accept more client connections than it has file descriptors for, no matter what the config says. If you set max_client_conn = 10000 while the process runs with a 1024 FD limit, PgBouncer keeps 10000, logs the kernel limit and its own expected-FD-use estimate at startup, then can hit the FD ceiling under load and refuse connections long before that configured limit.

This guide covers how to compute the real limit, verify what the running process actually has, and close the gap.

What this means

Every connection PgBouncer manages costs file descriptors:

  • One FD per client connection.
  • One FD per server connection to PostgreSQL.
  • Listening sockets (TCP and Unix socket).
  • The log file.
  • Internal pipe FDs.
  • Admin console sockets.
  • DNS resolution sockets.

The effective client capacity is therefore not max_client_conn. It is:

effective_limit = min(max_client_conn, fd_limit - server_connections - overhead)

where overhead covers listening sockets, the log file, pipes, admin sockets, and DNS. A sizing rule of thumb that accounts for both:

fd_limit >= max_client_conn x 2 + 500

The “x 2” covers the server connection each active client will typically need, and the 500 covers non-connection FDs plus safety margin.

Two more facts that change the math:

  • Admin connections to the special pgbouncer database are exempt from max_client_conn. You can always get in to run SHOW commands even when clients are being refused.
  • PgBouncer does not clamp max_client_conn to the detected FD limit. At startup it logs the kernel soft/hard limits and a maximum expected FD use; compare those logs with /proc/<pid>/limits and SHOW CONFIG.
flowchart TD
  A[max_client_conn from config] --> C{Effective client capacity}
  B[OS FD limit from ulimit or systemd] --> D[Subtract server connections]
  D --> E[Subtract listen, log, pipe, admin, DNS FDs]
  E --> C
  C --> F[Clients accepted up to this point]
  F --> G[New connections refused: no more connections allowed]
  E -. exceeded .-> H[EMFILE: Too many open files; accept and new server sockets fail]

Common causes

CauseWhat it looks likeFirst thing to check
max_client_conn raised without raising the FD limitRefusals or FD errors start below the configured limitSHOW CONFIG value vs grep "Max open files" /proc/<pid>/limits
Default ulimit (1024) never changedPgBouncer refuses connections around a few hundred clients despite a four-digit config/proc/<pid>/limits on the running process
systemd unit missing LimitNOFILElimits.conf was edited but the service still runs with a low limit/proc/<pid>/limits, not ulimit -n in your shell
Application connection leakused_clients climbs steadily toward the real limit under normal loadSHOW CLIENTS for connections with very old connect_time
Retry storm after an incidentused_clients spikes during pool exhaustion as applications reconnectPgBouncer log for refusal bursts correlated with cl_waiting events

Quick checks

All read-only. Run against the admin console and the live process.

# 1. Find the PgBouncer PID
pgrep -f pgbouncer

# 2. Check the ACTUAL FD limit of the running process (not your shell's ulimit)
grep "Max open files" /proc/$(pgrep -f pgbouncer)/limits

# 3. Count currently open FDs
ls /proc/$(pgrep -f pgbouncer)/fd | wc -l

# 4. Check the runtime max_client_conn (as loaded, not as written in the ini)
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW CONFIG;" | grep max_client_conn

# 5. Check current client usage against the limit
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW LISTS;" | grep -E "used_clients|free_clients"

# 6. Look for refusal and FD errors in the log
grep -c "no more connections allowed" /var/log/pgbouncer/pgbouncer.log
grep -c "Too many open files" /var/log/pgbouncer/pgbouncer.log

# 7. Sum current server connections (they consume FDs too)
psql -h 127.0.0.1 -p 6432 -U pgbouncer pgbouncer -Atc "SHOW LISTS;" | grep -E "used_servers|free_servers"

Check 2 is the one people skip. Your interactive shell’s ulimit -n tells you nothing about a service started by systemd or an init script. Only /proc/<pid>/limits reflects what the process actually has.

How to diagnose it

  1. Establish the three numbers. From the quick checks: the loaded max_client_conn (from SHOW CONFIG), the process FD limit (from /proc/<pid>/limits), and the current FD count. If SHOW CONFIG disagrees with the ini file, a later reload or wrong config file is in effect; do not infer an FD-based clamp from a lower value.

  2. Compute the budget. Take the FD limit. Subtract total server connections (used_servers plus headroom for pool growth up to the sum of all pool_size values). Subtract roughly 500 for listening sockets, the log file, pipes, admin sockets, and DNS. What remains is your real client capacity. Compare it to max_client_conn.

  3. Check whether FDs or config is the binding constraint. If the open FD count is near the FD limit while used_clients is well below max_client_conn, FD exhaustion is refusing connections, not the config. This is the worse failure mode: EMFILE can also block new server connections. It is not normally a PgBouncer crash loop, but the listener can remain suspended until FDs free up.

  4. If the limit is real and being reached legitimately, find the consumers. Run SHOW CLIENTS and group by source address and connect_time. Connections with very old connect_time and no recent activity point to an application leak. A broad distribution of fresh connections points to genuine concurrency growth or a retry storm.

  5. Check for a retry cascade. If refusals coincide with cl_waiting > 0 and rising maxwait in SHOW POOLS, the sequence is pool exhaustion first, application timeouts and retries second, client-slot exhaustion third. Fix the pool problem; the client limit is a symptom. See the pool exhaustion guide linked below.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
used_clients / max_client_connProximity to the hard refusal wall>80% sustained; >95% is urgent
Open FD count vs /proc/<pid>/limitsThe real ceiling, which may be lower than the configFD usage >80% of limit
PgBouncer log: no more connections allowedActive refusal. Only visible in logs, not in any SHOW commandAny sustained rate
PgBouncer log: Too many open files / accept() failureFD exhaustion; new client and server sockets can stop being createdAny occurrence
used_clients trend over weeksCapacity runway for client slotsSteady climb without traffic explanation
SHOW CLIENTS old connect_time entriesLeaked connections consuming slots indefinitelyLong-lived idle connections accumulating

Fixes

Raise the FD limit to match, then set max_client_conn below it

This is the correct fix when you genuinely need more client capacity. Do both sides together; changing only one recreates the mismatch.

On systemd-managed hosts, set the limit on the unit, not in limits.conf. Systemd does not apply /etc/security/limits.conf to services:

# /etc/systemd/system/pgbouncer.service.d/override.conf
[Service]
LimitNOFILE=21000

Apply with systemctl daemon-reload and a restart of PgBouncer. A restart drops all server connections and clients reconnect at once, so plan it for a low-traffic window. Then size max_client_conn using the rule of thumb in reverse: with LimitNOFILE=21000, a max_client_conn of 10000 fits the fd_limit >= max_client_conn x 2 + 500 rule with room to spare.

On SysV init systems, set ulimit -n inside the init script or the service’s defaults file before the daemon starts. Changing the limit after the process is running has no effect; the FD ceiling is fixed at process start.

Lower max_client_conn to fit reality

If you cannot or do not want to raise the FD limit, bring max_client_conn down to what the FD budget supports and treat client-slot exhaustion as a capacity signal. A lower honest limit beats a higher fictional one: with an honest limit, used_clients / max_client_conn alerts mean something.

Fix the application side

If SHOW CLIENTS shows leaked connections, no limit change fixes the problem. Reduce application pool sizes, fix code paths that open connections without closing them, and account for the full fan-out: application pool size times instance count must fit inside max_client_conn with headroom. Load balancer health checks also consume client slots; include them in the budget.

Verify after every change

After any reload or restart, re-run checks 2, 4, and 5. Confirm /proc/<pid>/limits shows the new FD limit and SHOW CONFIG shows the intended max_client_conn. The process may keep a previous configuration after an ineffective reload, so “I changed the config” is not evidence of anything.

Prevention

  • Validate the pair together. Any change to max_client_conn must ship with a corresponding FD-limit change, verified on the running process via /proc/<pid>/limits.
  • Codify the formula. Put fd_limit >= max_client_conn x 2 + 500 in a comment next to the setting in your config management so the next person does not tune one side alone.
  • Alert on the ratio, not the wall. used_clients / max_client_conn > 80% sustained gives you runway. Waiting for refusal log lines means users already saw errors.
  • Alert on FD usage independently. FD count vs the process limit catches the case where server connections, not clients, consume the budget.
  • Account for multiplication. Multiple PgBouncer processes (so_reuseport) each get their own per-process FD limit, but application instances multiply client connections across all of them. Do the arithmetic per process and in aggregate.
  • Parse the log. Refusals and EMFILE exist only in the log. A monitoring setup that scrapes only SHOW output is blind to both.

How Netdata helps

  • Tracks used_clients against max_client_conn continuously, so you see the utilization ratio trending toward the refusal wall instead of discovering it from application errors.
  • Collects per-process file-descriptor usage from /proc, letting you correlate FD consumption with client and server connection counts on one timeline.
  • Correlates client-slot saturation with pool signals (cl_waiting, maxwait, sv_active) so you can tell “too many clients” apart from “pool exhaustion driving retries that create clients.”
  • Surfaces connection-count anomalies per database and user, which helps pinpoint whether growth is a leak (steady, one source) or a traffic event (broad, correlated).
  • Retains per-second history across restarts, making it visible when PgBouncer restarted with a reset FD limit.