The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postgres / postgres-slow-queries-diagnosis

Operations Guides

PostgreSQL Slow Queries: Diagnosis From Log To Plan To Fix

A query that returned in 10 ms yesterday is now taking 8 seconds. No deploys, no schema changes. Before adding an index or restarting the database, determine whether the slowness is in the plan, the data, or the environment. This guide covers a three-layer workflow: log-based discovery with log_min_duration_statement, aggregate profiling with pg_stat_statements, and per-query execution plan capture with auto_explain and manual EXPLAIN (ANALYZE, BUFFERS).

flowchart TD
    A[Slow query reported] --> B[Check pg_stat_statements for mean_exec_time and stddev_exec_time]
    B --> C{High stddev relative to mean?}
    C -->|Yes| D[Plan flapping or parameter skew]
    C -->|No| E[Stable plan or system bottleneck]
    D --> F[Capture plan with auto_explain or EXPLAIN]
    E --> F
    F --> G{Estimated rows far from actual?}
    G -->|Yes| H[Stale statistics or skewed data]
    G -->|No| I[Bloat, locks, or cache pressure]

What This Means

A slow query is a symptom. PostgreSQL’s planner chooses a path based on statistics from ANALYZE, bound parameter values, and configuration such as work_mem. When these inputs change, the same query text can switch from a hash join to a nested loop and slow down by orders of magnitude.

log_min_duration_statement logs statements that exceed a millisecond threshold after execution completes. It will not log queries that hang indefinitely. pg_stat_statements normalizes queries by replacing literals with placeholders and groups them by queryid. It exposes mean_exec_time, stddev_exec_time, and calls, which show whether slowness is consistent or intermittent. It does not capture execution plans. auto_explain closes that gap by logging the actual plan, including timing and buffer usage, for queries that exceed its threshold.

queryid is computed from the post-parse-analysis representation. Two semantically different queries can collide into one entry due to a hash collision, though this is rare. Identical query text with different search_path contexts produces different queryid values. The hash is stable across minor versions on the same architecture with matching catalog metadata, but is not guaranteed stable across major versions or different architectures.

Common Causes

CauseWhat it looks likeFirst thing to check
Stale statistics or skewed dataEXPLAIN shows row estimates far from actual rows; sequential scan on a large table that should use an indexlast_autoanalyze in pg_stat_user_tables for tables in the query
Generic plan regression (prepared statements)Query is fast for the first 5 executions, then suddenly slows; partition pruning stops workingWhether the query uses prepared statements and the value of plan_cache_mode
Table bloat (dead tuples)Sequential scan or index scan touches mostly dead pages; n_dead_tup is growing while last_autovacuum is stalen_dead_tup / (n_live_tup + n_dead_tup) ratio in pg_stat_user_tables
Lock contentionQuery in pg_stat_activity shows wait_event_type = 'Lock'pg_locks for ungranted locks and the blocker PID
Buffer cache pressureEXPLAIN (ANALYZE, BUFFERS) shows high shared read=; cache hit ratio dropspg_stat_database blks_hit / (blks_hit + blks_read)

Quick Checks

Run these read-only checks before making any configuration changes.

-- Top queries by total execution time
SELECT queryid, query, calls, total_exec_time, mean_exec_time, stddev_exec_time
FROM pg_stat_statements
ORDER BY total_exec_time DESC LIMIT 10;
-- Sessions waiting on locks
SELECT pid, wait_event_type, wait_event, query
FROM pg_stat_activity
WHERE wait_event_type = 'Lock';
-- Dead tuple accumulation for tables in the slow query
-- Replace (...) with the target table names
SELECT schemaname, relname, n_live_tup, n_dead_tup,
       n_dead_tup::float / NULLIF(n_live_tup + n_dead_tup, 0) AS dead_ratio
FROM pg_stat_user_tables
WHERE relname IN (...);
-- Cache hit ratio for the database
SELECT datname,
       blks_hit::float / NULLIF(blks_hit + blks_read, 0) AS cache_hit_ratio
FROM pg_stat_database
WHERE datname = current_database();
-- Ungranted locks and blockers
SELECT l.locktype, l.relation::regclass, l.pid, l.mode, l.granted, a.query
FROM pg_locks l JOIN pg_stat_activity a ON l.pid = a.pid
WHERE NOT l.granted;

How To Diagnose It

  1. Confirm the database is the bottleneck. Use pg_stat_statements to find the query entry. Compare mean_exec_time and stddev_exec_time. If stddev_exec_time is high relative to mean_exec_time, the same query text is producing bimodal performance. This is the signature of plan flapping or parameter-sensitive skew.

  2. Determine whether the query is waiting or working. Query pg_stat_activity. If wait_event_type is 'Lock', the query is blocked. Find the blocker in pg_locks. Only terminate the blocker with pg_terminate_backend if you accept the risk of aborting that session’s work.

  3. Capture the execution plan. If the query is reproducible, run EXPLAIN (ANALYZE, BUFFERS).

    Warning: EXPLAIN (ANALYZE, BUFFERS) executes the query. For DML statements, wrap it in a transaction and roll back to avoid modifying data, or rely on auto_explain instead. Compare the planner’s estimated row counts to the actual row counts. If estimates are off by more than an order of magnitude, the planner is working with stale or insufficient statistics.

  4. Check statistics freshness. In pg_stat_user_tables, look at last_autoanalyze for every table referenced in the query. If the timestamp predates the last known bulk change, run ANALYZE on those tables and re-check the plan. For correlated columns, consider CREATE STATISTICS to improve cardinality estimates.

  5. Check for plan caching regression. PostgreSQL uses custom plans for the first 5 executions of a prepared statement, then evaluates whether a generic plan is competitive. For queries with skewed data distributions or partitioned tables, the generic plan can be catastrophically worse because it loses per-execution partition pruning. Set plan_cache_mode = force_custom_plan at the session level and re-run the query. If performance recovers, the generic plan is the culprit.

  6. Check for bloat and cache efficiency. In EXPLAIN (ANALYZE, BUFFERS) output, look at shared read= versus shared hit=. A plan dominated by shared read= points to cold cache or a working set larger than shared_buffers. High Heap Fetches in an Index Only Scan indicate dead tuples or a visibility map held back by long-running transactions. At the database level, a sustained drop in cache hit ratio in pg_stat_database confirms cache pressure.

  7. Correlate across time. If pg_stat_statements shows the queryid but you need the exact plan that was slow, use auto_explain logs if enabled. The queryid in pg_stat_statements can be correlated with log entries to link aggregate statistics to a specific plan shape.

Metrics & Signals To Monitor

SignalWhy it mattersWarning sign
pg_stat_statements.mean_exec_timeBaseline latency per query fingerprintP99 exceeds SLO or spikes relative to previous day
pg_stat_statements.stddev_exec_time / mean_exec_timePlan stability and data skewRatio greater than 1 suggests plan flapping
pg_stat_user_tables.n_dead_tup ratioBloat causing scans to read dead pagesRatio greater than 20% on active tables
pg_stat_database.blks_hit / (blks_hit + blks_read)Cache efficiencyBelow 95% for OLTP workloads
pg_stat_activity.wait_event_type = 'Lock'Time spent waiting, not executingSustained ungranted locks for more than 30 seconds
auto_explain log frequencyNew regressions appearingSudden increase in plans logged for previously fast queries

Fixes

Stale Statistics Or Skewed Data

Run ANALYZE on the affected tables. If the planner still underestimates cardinality, increase the per-column statistics target with ALTER TABLE ... ALTER COLUMN ... SET STATISTICS 500. For correlated columns, create extended statistics with CREATE STATISTICS. Tradeoff: higher targets increase ANALYZE duration and planner memory usage.

Generic Plan Regression

Set plan_cache_mode = force_custom_plan at the session or database level for queries that suffer from skewed parameters or partitioned tables. This forces parameter-specific planning on every execution. Tradeoff: planning overhead increases CPU consumption, especially for high-throughput queries.

Table Bloat

If autovacuum is not keeping up, tune per-table settings. Lower autovacuum_vacuum_scale_factor to 0.01 or 0.05 for large, high-churn tables. For severe existing bloat, use pg_repack to rebuild the table online.

Warning: pg_repack requires roughly 2x the disk space of the target table and a primary key or unique NOT NULL index.

Lock Contention

Terminate the blocker with pg_terminate_backend only if it is safe to do so. Set lock_timeout on DDL operations and idle_in_transaction_session_timeout globally to prevent future cascades. Fix application code that holds transactions open while calling external services.

Missing Or Inefficient Index

Add a partial or covering index only if pg_stat_statements shows the query pattern is frequent and the table is large enough to matter. Tradeoff: every additional index slows writes and increases vacuum I/O.

Prevention

  • Enable pg_stat_statements. It requires shared_preload_libraries and a server restart. This is the minimum viable observability for query performance.
  • Set log_min_duration_statement. Configure it to a threshold that captures your tail latency so you have a log trail for ad-hoc investigations.
  • Enable auto_explain. Add it to shared_preload_libraries and set auto_explain.log_min_duration to match your slow-query threshold. Be aware that auto_explain.log_analyze = on adds execution overhead.
  • Monitor stddev_exec_time. A rising standard deviation is often the first signal of a plan regression.
  • Tune autovacuum per table. Keep statistics and dead tuple maps current so the planner and vacuum keep up with the workload.
  • Set timeouts. Use statement_timeout to contain runaway queries and lock_timeout to prevent indefinite lock waits.

How Netdata Helps

  • Correlates PostgreSQL query latency with system-level CPU, disk I/O, and memory pressure to distinguish database-level slowness from resource starvation.
  • Tracks pg_stat_statements metrics over time to surface regressions without manual polling of cumulative counters.
  • Monitors cache hit ratio, dead tuple accumulation, and lock wait events alongside query throughput.
  • Alerts on checkpoint spikes and replication lag that can manifest as query slowdowns.
  • Provides historical context for auto_explain and log analysis by overlaying slow query periods with system resource charts.
The Netdata solution

PostgreSQL monitoring with Netdata

Netdata monitors PostgreSQL with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate connection saturation, lock waits, autovacuum progress, replication lag, and checkpoint I/O against the rest of your stack so you catch the incidents in these runbooks before they page anyone.