The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postgres / postgres-shared-buffers-tuning

Operations Guides

PostgreSQL shared_buffers tuning: 25% of RAM and why that rule of thumb breaks

The PostgreSQL documentation suggests 25% of RAM as a starting value for shared_buffers on dedicated servers and warns that values above 40% rarely improve performance. In production, operators routinely use 25% as a default and then encounter checkpoint storms, OOM kills, or degraded analytical query performance. The rule breaks because PostgreSQL does not coordinate with the Linux kernel page cache. The same page can exist in both shared_buffers and the OS page cache, so an oversized PostgreSQL cache starves the kernel of memory needed for sequential scans, WAL, and temporary files.

This guidance covers self-managed PostgreSQL on Linux. Managed services such as AWS Aurora use different defaults and failure modes, noted where relevant.

What shared_buffers is and why it matters

shared_buffers is PostgreSQL’s primary page cache. The default of 128 MB is inadequate for production. Pages read from disk load into this cache, and backends read from it directly. When a backend modifies a page, the change is written to shared_buffers and the page is marked dirty. Dirty pages flush to disk via the checkpointer or background writer. Changing shared_buffers requires a server restart; a reload is not sufficient.

Because shared_buffers caches table and index pages, its hit ratio indicates memory health. In OLTP workloads, aim for a cache hit ratio above 99%. If the working set exceeds shared_buffers, backends read from disk, adding milliseconds to queries that should complete in microseconds. However, increasing shared_buffers is not free. Each page requires management overhead, and the checkpointer must eventually flush every dirty page.

Check the current value and hit ratio:

SELECT name, setting, unit FROM pg_settings WHERE name = 'shared_buffers';
SELECT blks_hit, blks_read,
       blks_hit::float/(blks_hit+blks_read) AS ratio
FROM pg_stat_database
WHERE datname = current_database();

A low ratio alone does not mandate a larger shared_buffers. It may indicate the working set exceeds total RAM.

How the dual-cache architecture works

PostgreSQL does not inspect the OS page cache before storing a block in shared_buffers, and the kernel does not inspect shared_buffers before caching a file page. The same 8 KB page routinely exists in both caches simultaneously. This double buffering means total memory consumption for a single page can reach 16 KB or more.

flowchart TD
    query["Backend query"]
    disk["Disk storage"]
    sb["shared_buffers"]
    pc["OS page cache"]
    dirty["Dirty pages"]
    chk["Checkpointer"]

    disk -->|read on miss| sb
    disk -->|also cached by kernel| pc
    sb -->|query hit| query
    sb -.->|same page duplicated| pc
    query -->|writes| dirty
    dirty -->|flush| chk
    chk -->|checkpoint I/O spike| disk

When PostgreSQL evicts a clean page from shared_buffers, the kernel may still hold it in the page cache. A shared_buffers miss does not always mean a physical disk read. The OS page cache acts as a second line of defense, particularly for sequential scans that PostgreSQL may not keep in shared_buffers.

The downside emerges when shared_buffers is set too high. On a dedicated database server, RAM is finite. An oversized shared_buffers leaves less memory for the kernel page cache, slowing large sequential scans and increasing disk I/O. It also leaves less headroom for backend memory, connection overhead, and temporary files.

Checkpoint behavior compounds the problem. The checkpointer flushes dirty shared_buffers pages to disk at regular intervals. A larger cache allows more dirty pages to accumulate, so each checkpoint writes a larger volume. If storage cannot absorb the burst, query latency spikes until the flush completes. On Linux, PostgreSQL’s shared memory is allocated via mmap and is not swappable. If the sum of shared_buffers, connection overhead, and work_mem approaches available RAM, the kernel may OOM-kill backends under pressure.

Where the 25% rule breaks in production

The 25% guideline assumes a moderate-size OLTP server where the working set fits comfortably in RAM and queries access pages randomly. This assumption fails in several common scenarios.

OLAP and sequential scan workloads. Analytical queries scan large data ranges. The OS page cache is optimized for this access pattern because it caches underlying file blocks efficiently and evicts them under pressure. PostgreSQL’s shared_buffers uses a clock sweep algorithm that is less effective for massive sequential scans. For OLAP, values at or below 25% often perform better because the OS cache handles the scan load.

Large-memory systems. On servers with 512 GB or more of RAM, 25% yields 128 GB of shared_buffers or more. Beyond roughly 128 GB, the overhead of managing PostgreSQL’s internal buffer mapping table can decrease performance. Operators on very large systems often cap shared_buffers at 128-256 GB and let the OS page cache absorb the remainder.

Managed services with aggressive defaults. AWS Aurora PostgreSQL is reported to default shared_buffers to a large percentage of RAM. Combined with high connection counts and per-operation work_mem allocations, this can trigger OOM kills. Reducing shared_buffers to 25-40% may resolve these incidents. Always verify current defaults in your specific Aurora version and instance class.

Checkpoint amplification. Larger shared_buffers means more dirty pages can accumulate between checkpoints. When a checkpoint triggers, the checkpointer must flush all dirty buffers. Without adequate I/O bandwidth or tuning, this creates a latency spike that stalls queries.

Memory-constrained deployments. On instances with 16 GB or less of RAM, setting shared_buffers to 25% leaves only 12 GB for the OS, connection overhead, and work_mem. If max_connections is raised without pooling, or if work_mem is tuned aggressively, the remaining memory exhausts quickly. A smaller shared_buffers combined with greater reliance on the OS page cache often performs better and is less likely to trigger OOM kills.

Tradeoffs and when to deviate

Sizing shared_buffers is a tradeoff between cache locality and total system memory pressure. There is no single correct percentage.

OLTP vs. OLAP. For OLTP with random page access, shared_buffers in the 25-40% range is often appropriate. For OLAP, values at or below 25% are usually better because the OS page cache handles sequential scans more efficiently. If your workload mixes both, start at 25% and measure checkpoint duration and cache hit ratio before raising it.

Huge pages. On Linux, PostgreSQL can use explicit huge pages to reduce TLB misses for large shared_buffers allocations. The huge_pages parameter accepts off, try, or on. Setting it to try allows PostgreSQL to use them if available and fall back gracefully. Transparent Huge Pages (THP), however, are discouraged by the official PostgreSQL documentation because the kernel’s compaction process can cause latency spikes. Disable THP on database servers.

Calculate required huge pages before enabling:

# Check current huge page size and availability
grep Huge /proc/meminfo

# Estimate pages needed (shared_buffers in bytes / huge page size)
# For example, with shared_buffers = 64GB and Hugepagesize = 2048 kB:
# 64 * 1024*1024*1024 / (2048 * 1024) = 32768

If huge_pages = on, PostgreSQL will fail to start if the exact number of huge pages is not reserved in /proc/sys/vm/nr_hugepages.

Memory accounting. shared_buffers is not the only RAM consumer. The total footprint includes backend processes, autovacuum workers, maintenance_work_mem, wal_buffers, and per-query work_mem allocations. A safe heuristic is to ensure that shared_buffers plus expected concurrent memory usage stays below 80% of available RAM.

Setting vm.overcommit_memory = 2 on Linux makes the kernel strictly enforce overcommit limits and can prevent the OOM killer from targeting PostgreSQL backends. However, this is a system-wide, disruptive change that requires careful tuning of vm.overcommit_ratio or vm.overcommit_kbytes. If misconfigured, it can cause allocation failures for unrelated processes. Test in staging before applying in production.

Checkpoint tuning. If you increase shared_buffers, tune checkpoint behavior to avoid I/O storms. Set checkpoint_completion_target = 0.9 so the checkpointer spreads writes over most of the checkpoint interval. Increase max_wal_size so checkpoints are driven by timeout rather than WAL volume. Forced checkpoints produce I/O spikes.

Monitor checkpoint behavior:

SELECT checkpoints_timed, checkpoints_req,
       checkpoint_write_time, checkpoint_sync_time
FROM pg_stat_bgwriter;

Rising checkpoints_req relative to checkpoints_timed indicates WAL-driven checkpoints, which increases burst I/O.

Container and cgroup limits. In containers, shared_buffers must fit within the cgroup memory limit alongside the kernel page cache, connection overhead, and tmpfs usage. If the container lacks an adequate memory limit or swap configuration, the host OOM killer terminates the container during checkpoint spikes. Size shared_buffers to leave headroom for at least one checkpoint’s worth of dirty pages plus concurrent backend allocations.

Monitoring signals that indicate misconfiguration

When debugging production issues, correlate shared_buffers sizing with these signals:

  • Cache hit ratio sustained below 95% in OLTP. Check pg_stat_database. If disk latency is low and throughput is acceptable, the miss rate may be tolerable. If disk latency spikes, shared_buffers may be too small or the working set exceeds total RAM.
  • Checkpoint write time spikes. In Netdata, look for postgresql.checkpoint_write_time rising sharply at regular intervals. This indicates the checkpointer is flushing too many dirty pages at once. Either reduce shared_buffers or increase max_wal_size and checkpoint_completion_target.
  • System RAM exhaustion without obvious backend growth. If pg_stat_activity shows moderate connections but the OS reports high memory usage, check for double buffering. The kernel page cache plus shared_buffers may exceed physical RAM. Use free -m or /proc/meminfo to verify available memory and page cache size.
  • OOM kills targeting postgres backends. Check dmesg or the kernel log for OOM killer activity. If backends die while shared_buffers is above 40%, reduce it and verify work_mem and max_connections are not compounding the pressure.

Sizing workflow

  1. Establish baseline. Record current shared_buffers, cache hit ratio, and checkpoint frequency.
  2. Classify workload. OLTP favors 25-40%. OLAP favors 10-25%. Mixed workloads start at 25%.
  3. Check total memory pressure. Ensure shared_buffers + (max_connections * work_mem * average concurrent sorts) + maintenance_work_mem + OS overhead stays below 80% of RAM.
  4. Tune checkpoints. Set max_wal_size high enough to time out, and checkpoint_completion_target = 0.9.
  5. Restart and measure. Because shared_buffers requires a restart, schedule the change during a maintenance window. Compare checkpoint write times and query latency for at least 24 hours before further adjustments.
The Netdata solution

PostgreSQL monitoring with Netdata

Netdata monitors PostgreSQL with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate connection saturation, lock waits, autovacuum progress, replication lag, and checkpoint I/O against the rest of your stack so you catch the incidents in these runbooks before they page anyone.