The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postgres / postgres-fsync-disabled

Operations Guides

PostgreSQL fsync=off: why disabling it ruins durability

You are reviewing a PostgreSQL instance that crashed during a kernel update and now fails to start with checksum or WAL errors. Or you are benchmarking write throughput and a guide suggests turning off fsync to remove the “fsync bottleneck.” Maybe you are auditing configuration drift and found fsync = off in a postgresql.conf that was copied from an old developer environment.

PostgreSQL’s default is fsync = on in all supported versions. Disabling it can make bulk loads and heavy write workloads appear dramatically faster because the database stops waiting for the operating system to flush write-ahead log (WAL) records to durable storage. The cost is that an operating system crash or power loss can leave the data directory in an unrecoverable state. WAL is the source of truth; fsync = off means the database trusts the OS page cache with that truth. When that trust is broken by a power event, the resulting corruption may not be detected until you attempt to start the server or query a specific table.

What it is and why it matters

fsync controls whether PostgreSQL forces WAL and data file modifications to durable storage using system calls such as fsync() or fdatasync(). When fsync = on, PostgreSQL issues these calls according to the wal_sync_method parameter. On Linux, wal_sync_method defaults to fdatasync.

The guarantee is simple: when a transaction commits, its WAL records survive a power loss.

When fsync = off, PostgreSQL still writes WAL records to kernel buffers, but it never forces them to disk. The database assumes the OS will eventually flush the data, and it assumes that any crash will be recoverable. That assumption is wrong for OS-level crashes. The OS page cache is a volatile staging area. If the OS panics or the host loses power before the kernel page cache is written to disk, committed transactions evaporate. Worse, because data files may have been written without the corresponding WAL, the on-disk state can become internally inconsistent. The difference between losing a few seconds of transactions and losing the entire database is the difference between a recoverable incident and a restore-from-backup emergency.

How it works

Under normal operation, the commit path looks like this. A transaction generates WAL records in shared memory. At commit, the WAL writer (or the committing backend) writes those records to the current WAL file and issues a sync method call. Only after the sync succeeds does PostgreSQL acknowledge the commit to the client. If the server crashes after the commit returns, crash recovery replays WAL from the last checkpoint and brings the database to a consistent state.

With fsync = off, the sync call is skipped. WAL records may sit in the OS page cache for seconds or minutes. Data file modifications from checkpoints or background writer activity are also not forced to disk. Checkpoints lose their meaning as a recovery boundary because the writes they issued may never have reached the physical medium. The database appears to run normally because reads and writes succeed from the perspective of the PostgreSQL process. The corruption is only revealed when the OS cache is lost.

flowchart TD
    A[COMMIT received] --> B{fsync setting}
    B -->|on| C[wal_sync_method flushes WAL]
    C --> D[WAL on durable storage]
    D --> E[Acknowledge to client]
    B -->|off| F[WAL written to OS page cache]
    F --> E
    F --> G[Power loss or OS crash]
    G --> H[Unflushed WAL lost]
    H --> I[Corruption or silent data loss]

Because WAL is the source of truth and data files are treated as a cache, losing WAL means the data files cannot be reconciled. The checkpoint mechanism, crash recovery, and point-in-time recovery all depend on WAL being present and complete. fsync = off breaks that dependency.

There is a secondary interaction with full_page_writes. When fsync = off, partial page writes caused by an OS crash cannot be recovered using WAL full-page images, because the WAL itself may not have been flushed. The PostgreSQL documentation suggests that if you disable fsync, you might also want to disable full_page_writes. In practice, this means you are disabling two independent safety mechanisms at once.

Another related parameter is commit_delay, which batches commits to amortize the cost of a WAL flush. When fsync = off, there is almost no flush cost to amortize, so commit_delay provides no benefit.

Where it shows up in production

The most common production encounter with fsync = off is not a deliberate choice. It is a configuration that was copied from a benchmark environment, a developer workstation, or an old tuning guide and never reverted before going live.

Benchmarking and load testing. Disabling fsync is a legitimate way to measure CPU-bound query throughput without storage latency. The error is leaving the setting in place when the test ends.

Developer environments and containers. PostgreSQL configuration files from development environments sometimes disable fsync to reduce disk wear or speed up test suites. When these configurations are promoted to staging or production via configuration management templates, the setting comes along for the ride.

Misunderstanding safer alternatives. Teams looking to reduce commit latency sometimes conflate fsync = off with synchronous_commit = off. The former disables storage sync entirely. The latter simply returns the commit acknowledgment before the WAL flush completes, while still allowing the background WAL writer to flush to disk. They are not equivalent.

Storage migration or hardware testing. During storage benchmarking or migration validation, operators may disable fsync to isolate throughput numbers. If the instance is later promoted to a production role without reverting the change, the durability guarantee is gone.

Tradeoffs and common misuses

fsync = off is not a performance tuning knob for production OLTP. It is a durability bypass switch.

fsync = off versus synchronous_commit = off. This distinction is critical. synchronous_commit = off allows the database to return a commit acknowledgment before WAL is flushed, but PostgreSQL still writes WAL and still flushes it to disk via the walwriter process. A crash may lose transactions that committed in the last wal_writer_delay window, but the on-disk database remains logically consistent. Those transactions simply never happened, as far as recovery is concerned. This makes synchronous_commit = off an acceptable choice for non-critical batch loads or telemetry ingestion where losing a few seconds of data is acceptable. With fsync = off, the WAL itself may be incomplete or missing, so the database cannot recover to a consistent state at all.

fsync = off versus full_page_writes = off. Disabling full_page_writes reduces WAL volume by omitting full page images, but it still requires fsync to be on to protect against partial page writes during an OS crash. Disabling both removes both the page-image safety net and the WAL durability guarantee.

The data_sync_retry defense. By default, data_sync_retry is off. When it is off, any failure to flush modified data files causes PostgreSQL to emit a PANIC-level error and crash the entire instance. This behavior was introduced in response to kernel-level fsync bugs, where retrying fsync after a failure could falsely report success while data was lost. The PANIC forces human intervention rather than risking silent corruption. This parameter does not make fsync = off safe, but it illustrates how seriously PostgreSQL treats storage sync failures.

Signals to watch in production

SignalWhy it mattersWarning sign
fsync setting in pg_settingsDetermines whether WAL is forced to durable storagefsync = off in any production or staging instance
Confusion with synchronous_commitTeams often disable the wrong parameter when chasing latencyfsync = off set alongside comments about “group commit” or “commit latency”
commit_delay having no effectWhen fsync = off, there is almost no WAL flush cost to amortizeTuning commit_delay produces zero throughput change
PANIC logs after storage errorsdata_sync_retry = off (default) crashes the instance on flush failureAny PANIC that follows a storage sync or file flush failure
Corruption on startup after a crashWithout fsync, OS crashes leave the data directory inconsistentErrors about invalid WAL records, missing pages, or checksum failures following a reboot
full_page_writes = on with fsync = offFull page images cannot protect against torn pages if WAL itself is not durableMisconfiguration that disables sync while keeping page images enabled

How Netdata helps

Netdata collects PostgreSQL parameters and runtime metrics that surface configuration drift without manual pg_settings inspection.

  • Configuration audit. Netdata collects PostgreSQL parameters, making it easy to spot fsync = off or full_page_writes = off during routine fleet reviews.
  • WAL and checkpoint correlation. By monitoring WAL generation rates alongside disk flush latency, you can identify whether commit latency is actually bounded by storage sync or by other factors before considering any durability tradeoff.
  • Crash and PANIC detection. Netdata monitors PostgreSQL log severity. A spike in PANIC messages related to data_sync_retry or storage sync failures surfaces immediately.
  • Commit latency baselines. Tracking transaction commit times helps you measure the real impact of synchronous_commit adjustments, which is the safer first lever to pull when write latency is too high.
The Netdata solution

PostgreSQL monitoring with Netdata

Netdata monitors PostgreSQL with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate connection saturation, lock waits, autovacuum progress, replication lag, and checkpoint I/O against the rest of your stack so you catch the incidents in these runbooks before they page anyone.