The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postgres / how-postgres-works-in-production

Operations Guides

How PostgreSQL Actually Works In Production: A Mental Model For Operators

PostgreSQL is a reliable relational database, but that reliability is not magic. It is the emergent behavior of five concrete subsystems that interact in specific, observable ways: the Write-Ahead Log, Multi-Version Concurrency Control, checkpoints, autovacuum, and the process-per-connection model. If you do not understand how these interact, you will misdiagnose table bloat as a missing index, interpret checkpoint I/O spikes as disk failures, and respond to connection exhaustion by raising max_connections until the OOM killer arrives.

This article traces the path of a committed write from the client through the backend, WAL, and data files. It explains why a DELETE does not free space, why more connections can mean less throughput, and why a forgotten replication slot can fill a primary’s disk even when the replica looks healthy. The goal is to give you a concrete mental model for reasoning through incidents at 3 a.m. without guessing.

Why These Five Abstractions Matter

Every production incident in PostgreSQL eventually touches WAL, MVCC, checkpoints, autovacuum, or the connection model. WAL determines whether you can recover from a crash. MVCC explains why tables grow after deletes and why long transactions cause bloat. Checkpoints explain why latency spikes on a predictable cadence. Autovacuum determines whether the database survives long enough to need a major version upgrade. The process-per-connection model dictates your memory ceiling and why connection pooling is architecturally mandatory rather than optional.

These details are the control levers you have, whether you run on EC2, in Kubernetes, or on a managed service that hides the filesystem.

How It Works

flowchart TD
    Client[Client connection] -->|SQL command| Backend[Backend process]
    Backend -->|read/modify pages| SharedBuffers[Shared buffers]
    Backend -->|append record| WAL[WAL buffer]
    WAL -->|flush| WALDisk[WAL files]
    SharedBuffers -->|checkpoint flush| DataFiles[Data files]
    Autovacuum[Autovacuum worker] -->|reclaim dead tuples| SharedBuffers
    Autovacuum -->|mark all-visible| VM[Visibility map]

WAL Is The Source Of Truth

The Write-Ahead Log is an append-only, sequential journal. Every change is written to WAL before it is applied to heap pages. In practice, WAL is the source of truth and the data files are a materialized cache. If the server crashes, recovery replays WAL from the last checkpoint forward to bring the data files to a consistent state.

The default wal_level is replica, which supports archiving and streaming replication. full_page_writes is on by default, writing full page images after every checkpoint to prevent torn pages. Turning fsync off risks unrecoverable corruption. If you need lower commit latency, use synchronous_commit = off or local instead, which still flushes WAL but may return before the disk acknowledge.

Checkpoints Bound Recovery Time

A checkpoint is a point where all dirty buffers are flushed to disk. Checkpoints begin every checkpoint_timeout seconds (default 5 minutes) or when max_wal_size is about to be exceeded (default 1 GB), whichever comes first. The background writer helps smooth I/O between checkpoints, but it does not eliminate the spike. checkpoint_completion_target defaults to 0.9, spreading the flush over most of the checkpoint interval.

If max_wal_size is too small for your write volume, checkpoints become forced and frequent. You will see I/O latency spikes on a fixed schedule. The tuning goal is to keep checkpoints_req rare, well under 10 percent of checkpoints_timed.

MVCC Creates Dead Tuples

Multi-Version Concurrency Control allows readers and writers to proceed without blocking each other by keeping multiple versions of rows in the table. Every UPDATE creates a new tuple version; every DELETE leaves a dead tuple behind. A transaction sees the snapshot that existed when it started, implemented through xmin and xmax transaction IDs in each row header.

Because old versions stay in the table until reclaimed, the free space from a DELETE is not immediately reusable for new rows of different sizes. A table can grow to many times its logical size while still reporting the same row count. Only VACUUM, VACUUM FULL, or pg_repack reclaims that space. The Visibility Map tracks which pages contain only all-visible tuples, enabling index-only scans and allowing autovacuum to skip clean pages. All access is page-oriented, and the page is the 8 KB I/O unit.

Autovacuum Is A First-Class Write Workload

Autovacuum is not optional cleanup. It is a background write workload that reclaims dead tuple space, updates the free space map, freezes transaction IDs, and maintains the visibility map. Workers launch when a table exceeds a threshold of dead tuples: autovacuum_vacuum_threshold (default 50 tuples) plus autovacuum_vacuum_scale_factor (default 0.2, or 20 percent of table rows).

Anti-wraparound vacuum fires when a table’s oldest XID exceeds autovacuum_freeze_max_age (default 200 million). Because the XID counter is 32-bit, wraparound is inevitable without freezing. PostgreSQL emits warnings when roughly 40 million XIDs remain, and refuses writes when approximately 3 million remain.

Long-running transactions block vacuum progress because a worker cannot remove dead tuples newer than the oldest active snapshot. Disabling autovacuum because it competes for I/O causes bloat that slows the system down far more than the vacuum itself. VACUUM also writes WAL block by block, so it is recoverable after a crash.

Connection-per-process Sets Your Memory Ceiling

Each client connection spawns a dedicated operating-system process consuming roughly 5-10 MB of resident memory even when idle. The default max_connections is 100. Context-switching dominates after roughly CPU count multiplied by 2 active connections, so raising max_connections to absorb load usually makes throughput worse.

Connection poolers such as PgBouncer multiplex many client connections over a small pool of backends. This avoids the per-process memory tax and is mandatory at scale. Memory pressure is amplified by work_mem, which is allocated per operation, not per session. A query with four hash joins can use four times work_mem simultaneously. Treat connection pooling and conservative work_mem sizing as part of the architecture, not as optimizations.

Where These Abstractions Show Up In Production

Table bloat after deletes. A developer runs a large DELETE and expects the table to shrink. Instead, query latency degrades because sequential scans traverse pages full of dead tuples. The fix is not to run VACUUM FULL during business hours, which acquires an exclusive lock, but to ensure autovacuum is tuned and not blocked by idle transactions.

Checkpoint I/O spikes. Every five minutes, latency jumps for ten to twenty seconds. pg_stat_bgwriter shows checkpoints_req approaching checkpoints_timed. The root cause is a max_wal_size that is too small for the write volume. Spreading checkpoints by raising max_wal_size and checkpoint_timeout smooths the spike.

Connection storms during deploys. A rolling deployment opens new connections before old ones close. Without a pooler, the backend count hits max_connections and the database rejects new clients. The fix is PgBouncer in transaction mode, not a larger max_connections that consumes memory and increases context switches.

Replication lag and WAL retention. Streaming replication ships WAL from primary to replica. A forgotten replication slot retains WAL indefinitely, filling the primary’s disk even though the replica looks healthy. Slots must be monitored explicitly; they do not clean themselves up when a consumer disappears.

Transaction ID wraparound emergencies. A database that has been running fine suddenly goes read-only. age(datfrozenxid) has been approaching the limit for months while autovacuum appeared active but was blocked by idle transactions or misconfigured thresholds. Recovery requires emergency VACUUM FREEZE and may take hours.

Lock contention versus internal bottlenecks. High CPU with low throughput and wait_event_type = 'Lock' in pg_stat_activity indicates application-level contention. If the wait event is LWLock, the bottleneck is internal PostgreSQL coordination, often from high WAL volume or aggressive parallel query launches.

Tradeoffs & Common Misuses

Durability versus latency. fsync = off is faster and absolutely unacceptable for production. synchronous_commit = off provides lower latency at the cost of a small data-loss window. synchronous_commit = remote_apply guarantees durability on a replica but adds network round-trip latency to every commit.

Autovacuum aggressiveness versus query performance. Aggressive settings prevent bloat but consume I/O and CPU continuously. Conservative settings preserve foreground query performance but allow dead tuples to accumulate. The correct balance is workload-dependent: high-churn OLTP needs frequent vacuum, while read-heavy systems can tolerate relaxed settings.

Index coverage versus write amplification. Each index speeds specific reads but slows every INSERT, UPDATE, and DELETE because the engine must maintain the index. Partial and covering indexes reduce write overhead, but only for queries that match the predicate.

Shared buffers sizing. The default shared_buffers of 128 MB is too small for production, but setting it above roughly 40 percent of RAM can starve the operating system page cache and lead to out-of-memory kills. A reasonable starting point is 25 to 40 percent of RAM for OLTP workloads.

Memory allocation: cache versus sort/hash. shared_buffers caches data pages, while work_mem allocates per-operation memory for sorts and hashes. Over-allocating work_mem causes OOM; under-allocating it spills to disk. Size work_mem conservatively and account for concurrent operations and parallel workers.

Signals To Watch In Production

SignalWhy it mattersWarning sign
Buffer cache hit ratioIndicates whether the working set fits in memoryBelow 95 percent for OLTP
Dead tuple ratioMeasures bloat and vacuum healthAbove 20 percent sustained
Transaction ID ageTracks wraparound riskage(datfrozenxid) above 500 million
Required versus timed checkpointsReveals checkpoint pressure from WAL volumecheckpoints_req above 10 percent of checkpoints_timed
Connection utilizationCapacity headroom before connection refusalAbove 90 percent of max_connections
Replication lagFailover readiness and WAL retention riskAbove 30 seconds on async replicas

How Netdata Helps

  • Correlate PostgreSQL query latency with system disk latency to confirm whether latency spikes are checkpoint I/O rather than storage failures.
  • Track per-table dead tuple ratios to spot vacuum lag before bloat degrades sequential scan and index performance.
  • Monitor replication lag, WAL directory size, and connection counts together to distinguish inactive replication slots from network saturation.
  • Alert on transaction ID age per database to provide weeks of warning before wraparound emergencies.
  • Visualize buffer cache hit ratio alongside operating system memory metrics to validate shared_buffers sizing without guessing.
The Netdata solution

PostgreSQL monitoring with Netdata

Netdata monitors PostgreSQL with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate connection saturation, lock waits, autovacuum progress, replication lag, and checkpoint I/O against the rest of your stack so you catch the incidents in these runbooks before they page anyone.