The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postgres / postgres-prepared-statement-plan-cache

Operations Guides

PostgreSQL prepared statement plan cache: generic vs custom plan pitfalls

When an application repeats the same query with different parameters, prepared statements avoid repeated parsing. But PostgreSQL’s plan cache does not work the way many operators expect. The server does not automatically cache the first plan it generates. It executes the first five invocations with custom plans that see the actual bound values. At the sixth execution, the planner compares the average cost of those custom plans against a generic plan built with placeholders. If the generic plan looks cheaper, the session switches to it permanently. There is no automatic reversion.

This heuristic is session-scoped and hardcoded. For uniform data it works well. For skewed distributions, partitioned tables, or connection-pooled workloads, it causes sudden latency cliffs, memory bloat, or protocol errors. You need to know how the 5-execution switch, the plan_cache_mode setting (PostgreSQL 12+), and client driver behavior interact before these hit production.

What it is and why it matters

A prepared statement is a server-side object created with PREPARE name AS .... The client later executes it with EXECUTE name(parameters). The benefit is avoiding repeated parsing and reusing an execution plan. But PostgreSQL must trade off plan optimality against stability. A plan optimized for one parameter value may be terrible for another. The planner resolves this by testing custom plans first, then choosing a generic plan after five executions.

The generic plan relies on table statistics and averages without knowledge of the actual parameter values. When data is skewed, the generic plan can be off by orders of magnitude. Once the switch happens, the session is stuck with that plan until the prepared statement is deallocated or the connection closes. In long-lived connections or pooled environments, a bad plan persists until the session ends.

Drivers and ORMs often prepare statements transparently. You may never issue an explicit PREPARE yet still be exposed to the plan cache. Python’s psycopg and Go’s pgx both prepare automatically under certain thresholds, and any driver using the extended query protocol can trigger the server-side plan cache.

How it works

The logic lives in choose_custom_plan() in the PostgreSQL planner. When a prepared statement executes:

  1. Executions 1 through 5 always generate custom plans. The planner sees the literal values bound to each parameter and optimizes accordingly.
  2. At execution 6, the planner computes the average estimated cost of all prior custom plans and compares it to the estimated cost of a generic plan using $n placeholders.
  3. If the generic plan cost is less than or equal to the average custom plan cost, PostgreSQL caches the generic plan and uses it for all subsequent executions in that session.
  4. If the generic plan is more expensive, the session continues with custom plans. The exact re-evaluation frequency is internal, but the critical invariant is that once a generic plan is chosen, the session never switches back mid-session.

The plan_cache_mode GUC, available from PostgreSQL 12, overrides this heuristic entirely. Valid values are:

  • auto: the default, using the 5-execution heuristic.
  • force_custom_plan: always generate a custom plan per execution.
  • force_generic_plan: always use a generic plan from the first execution.

This setting is consulted at execution time, not at PREPARE time. Changing it affects existing prepared statements immediately. Check the current value with SHOW plan_cache_mode;.

You can confirm a generic plan is active by running EXPLAIN EXECUTE stmt_name(...). If the output shows $1, $2, and so on instead of literal values, the generic plan is in use.

There is no built-in re-evaluation. If the data distribution shifts after the generic plan is adopted, the session continues using the stale plan. The only operator-level escape is DEALLOCATE stmt_name followed by recreation of the prepared statement, or reconnecting to start a fresh session.

flowchart TD
    A[PREPARE statement] --> B[Executions 1-5]
    B --> C[Custom plan per execution
literal values visible] C --> D[Execution 6 and later] D --> E{Avg custom cost >=
generic plan cost?} E -->|Yes| F[Generic plan cached
$n placeholders] E -->|No| G[Custom plans continue] F --> H[Sticks until DEALLOCATE
or session ends] G --> H

Where it shows up in production

The skewed-data trap

The most dangerous manifestation is skewed data. Consider a table with a status column where 99% of rows are shipped and 1% are returned. A prepared statement SELECT * FROM orders WHERE status = $1 might execute its first five invocations with returned. The planner chooses an index scan. At execution six, the generic plan also favors an index scan because it is optimized for the rare value. When the application later binds shipped, the executor performs a random index lookup across nearly the entire table. Latency jumps from milliseconds to minutes.

Partitioned tables and memory bloat

Partitioned tables create a different problem. Because a generic plan is built without parameter values, PostgreSQL cannot prune partitions at plan time. Before PostgreSQL 15, a generic plan referenced all partitions and could consume gigabytes of memory per backend. With a connection pooler maintaining dozens or hundreds of backend connections, this multiplies into severe memory pressure. PostgreSQL 15 reduces this overhead by allowing partition pruning at execution time for generic plans, but the inability to prune at plan time remains.

Connection poolers and driver mismatches

Connection poolers add protocol-level breakage. Prepared statements are session-scoped. PgBouncer in pool_mode=transaction returns the backend to the pool after each commit or rollback. The next transaction on that connection starts fresh, and the prepared statement no longer exists. If the client driver tries to reuse the prepared statement name, the server responds with prepared statement "X" already exists or does not exist.

The Java PostgreSQL JDBC driver defaults to prepareThreshold=5, meaning it does not create a server-side prepared statement until the fifth execution on a single connection. In a transaction-pooled environment, the driver and the server can get out of sync. Disable server-side prepares with prepareThreshold=0. Some driver versions also issue server-side prepares for metadata queries when PgBouncer is present; in that case set preparedStatementCacheQueries=0 as well.

Tradeoffs and when to use it

When generic plans win

Generic plans eliminate planning overhead and are predictable. They work well for simple lookup queries on uniform data where the same plan is optimal regardless of parameter value. Short point queries on primary keys or uniformly distributed foreign keys are good candidates. If planning time dominates execution time, force_generic_plan can improve throughput.

When custom plans are safer

Custom plans are safer when data is skewed or when partition pruning matters, but they cost CPU. Every execution re-runs the planner. For complex queries on skewed columns, that overhead is usually smaller than the cost of a bad generic plan. Set plan_cache_mode = force_custom_plan at the database, user, or session level when you observe bimodal latency or when queries touch heavily skewed columns like tenant IDs, status fields, or rare event types.

Warning: force_custom_plan increases planning CPU overhead. Apply it selectively, not globally, unless you have proven the planner cost exceeds the execution penalty of bad generic plans.

Working around bad plan adoption

If a generic plan has been adopted and is performing badly, you cannot force a mid-session reversion. DEALLOCATE stmt_name followed by re-preparation resets the counter, but this requires application cooperation or session cycling.

For PgBouncer, the cleanest fix is to use pool_mode=session for databases that rely on prepared statements, though this reduces pooling efficiency. Alternatively, disable server-side prepares in the driver and let the client handle parameter binding without PREPARE.

For partitioned workloads, prefer PostgreSQL 15 or later to avoid generic-plan memory bloat. If you are on an older release and cannot upgrade, monitor backend memory closely and consider force_custom_plan to reduce the number of generic plans held in memory.

Signals to watch in production

SignalWhy it mattersWarning sign
pg_stat_statements.stddev_exec_time / mean_exec_timeReveals plan flapping or bimodal latency from skewed parameter valuesRatio greater than 1 on a prepared query that should be uniform
Application latency by connection ageGeneric plan adoption happens after the 5th execution in a sessionLatency jumps sharply after a connection has executed the same statement several times
Backend memory per connectionPre-PG15 generic plans for partitioned tables reference all partitionsBackend RSS grows with prepared statement use against many partitions
PgBouncer error rateTransaction pooling breaks session-scoped prepared statementsprepared statement already exists or does not exist errors spike

pg_stat_statements requires the extension to be loaded. It aggregates across sessions, so it will not show a clean 5-execution cutoff; pair it with application logs or connection-level tracing.

How Netdata helps

  • Compare pg_stat_statements latency percentiles to spot bimodal distributions from bad generic plans.
  • Monitor backend process memory with query throughput to catch partition-plan bloat before OOM.
  • Track PgBouncer connection state and error rates to detect prepared-statement mismatches.
  • Compare planning time versus execution time per query to justify force_custom_plan or force_generic_plan.
  • Overlay query latency charts with deployment markers to correlate driver-level prepare threshold changes.
The Netdata solution

PostgreSQL monitoring with Netdata

Netdata monitors PostgreSQL with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate connection saturation, lock waits, autovacuum progress, replication lag, and checkpoint I/O against the rest of your stack so you catch the incidents in these runbooks before they page anyone.