The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-memory-vs-file-buffer ▌

Operations Guides

Fluentd memory vs file buffer: why the default buffer loses data on restart

You restarted Fluentd for a config change, or the OOM killer restarted it for you, or Kubernetes rescheduled the pod. The process came back healthy. Every metric looks normal. But downstream there is a gap in the logs covering the minutes before the restart, and no error anywhere explains it.

The explanation is almost always the same: the output was using the memory buffer, and every chunk that had not been flushed at the moment the process died was deleted with it. This is not a bug. It is the documented behavior of the memory buffer. The Fluentd troubleshooting documentation itself lists “change buffer type from memory to file” as a standard remediation for exactly this symptom.

This article covers how to confirm which buffer type each of your outputs is actually using, why the data is unrecoverable once the process is gone, and how to move to a file buffer with explicit limits.

What this means

Fluentd buffers events in chunks between the filter chain and the output plugin. Chunks move from staged (accumulating events) to queued (ready for flush) to flushed. Where those chunks live depends on the buffer type:

  • Memory buffer: chunks are Ruby objects in the process heap. Fast, but they exist only as long as the process does. When Fluentd shuts down, buffered logs that cannot be written quickly are deleted. On a SIGKILL, OOM kill, or crash, there is no shutdown path at all: the heap simply ceases to exist.
  • File buffer: chunks are written to disk as append-only binary files. They survive restarts and are reloaded on startup, then flushed. Delivery becomes at-least-once, so a brief duplication window after a restart is expected.

The trap is the default. The @type parameter in <buffer> is not mandatory. If you omit it, the output plugin may specify its own buffer plugin; otherwise Fluentd falls back to the memory buffer. A config that never mentions buffers can still be running memory buffers everywhere.

flowchart TD
  A[Events staged in buffer chunks] --> B{Process stops: crash, restart, OOM}
  B -->|"@type memory"| C[Chunks exist only in RAM]
  C --> D[All unflushed data lost]
  B -->|"@type file"| E[Chunks written to disk]
  E --> F[Reloaded on startup and flushed]
  F --> G[At-least-once: brief duplicates possible]

One nuance: flush_at_shutdown defaults to true for non-persistent buffers like memory and false for persistent buffers like file. So on a clean, graceful shutdown the memory buffer tries to flush before exit. That helps with planned restarts when the destination is healthy. It does nothing for crashes and OOM kills, and it does not save you when the destination is the reason you restarted, because the flush attempt will fail and the remaining chunks are deleted anyway.

Common causes

CauseWhat it looks likeFirst thing to check
Memory buffer in use (explicit or fallback default)Log gap in the destination matching the restart window, no Fluentd errorswith_config=true on the monitor API, look at each output’s buffer section
Plugin default overrode team intentSame gap, but the config file says nothing about buffers at allThe running config from the API, not the file on disk
OOM kill during backpressureGap plus a process restart nobody initiateddmesg for OOM entries, container restart count
Graceful restart while destination was downGap even though shutdown was clean, retry warnings in the log before restartFluentd log for retry and flush failures around the shutdown
File buffer configured but path lostGap despite @type file, usually in KubernetesWhether the buffer path is on a persistent volume or ephemeral container storage

Quick checks

All read-only.

# 1. Confirm the running buffer config for every output plugin
curl -s "http://localhost:24220/api/plugins.json?with_config=true" | \
  jq '.plugins[] | select(.plugin_category=="output")'

Look at each output’s buffer section. If there is no @type file, you are on the plugin default, which for a bare output plugin is memory. This shows the running config, which is what matters; the file on disk may differ after a partial reload.

# 2. Check the on-disk config for explicit buffer sections
grep -n -A5 "<buffer" /etc/fluent/fluentd.conf
# td-agent: /etc/td-agent/td-agent.conf

An absent <buffer> section, or one without @type, means the default decided for you.

# 3. Confirm the process actually restarted (and why)
systemctl status fluentd          # fluent-package; td-agent: systemctl status td-agent
dmesg | grep -i oom               # OOM kill evidence
kubectl get pods -l app=fluentd   # Kubernetes: check RESTARTS count
# 4. See whether chunks were waiting when the process died
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") | {id: .plugin_id, queue: .buffer_queue_length, bytes: .buffer_total_queued_size}'

A queue that was non-trivial before the restart is, roughly, the size of your loss if the buffer was memory-backed.

# 5. If a file buffer exists, check what is on disk
du -sh /var/log/fluent/buffer/    # or wherever your buffer path points

Chunk files present after a restart will be replayed. An empty buffer directory after a restart with a memory buffer is exactly what you would expect.

How to diagnose it

  1. Establish the timeline. Find the restart time (systemd status, pod restart timestamp, Fluentd log start). Find the gap window in the destination. If they align, buffer loss on restart is the working theory.
  2. Determine the restart type. Graceful (deploy, config change) versus hard (OOM kill, crash, node failure, pod eviction). SIGHUP is a reload, not a restart. Hard stops with a memory buffer lose everything queued, no exceptions. Graceful stops only flush what the destination accepts before shutdown.
  3. Verify the effective buffer type from the API, not the config file. ?with_config=true shows what is actually running. This catches the case where the file on disk was edited but a reload only partially applied, or where a plugin’s own default silently chose memory.
  4. Estimate the blast radius. The pre-restart buffer_total_queued_size for that output is the upper bound of lost bytes. If you were not capturing it, use input emit rate times the gap duration as a rough estimate.
  5. Rule out the lookalikes. A gap can also come from overflow_action throw_exception dropping events at a full buffer while Fluentd stayed up, or from a pos_file desync on the input side. The distinguishing feature here is that the gap brackets the restart itself and Fluentd metrics show no overflow or retry anomaly during it.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
buffer_total_queued_size per outputBytes at risk if the process dies on a memory bufferSustained growth; any high value with @type memory is pure exposure
buffer_queue_length per outputChunks waiting to flush; on memory buffer, this is the loss inventoryNon-zero and growing, especially before a planned restart
buffer_available_buffer_space_ratiosProximity to overflow, which compounds restart loss with drop lossBelow 20% and falling
write_count rateConfirms chunks are draining before you restartFlat while queue is non-zero: do not restart yet
retry_count and retry stateDestination failing; restarting into a down destination maximizes lossNon-zero with retry.next_time far in the future
Process RSS and restart countOOM kills are ungraceful by definition; each one on a memory buffer is a loss eventAny OOM entry in dmesg, any unexpected restart
buffer_oldest_timekeyAge of oldest buffered data; tells you how far back a loss would reachOldest timekey hours behind current time

Fixes

Confirm exposure, then restart deliberately (if you must)

Before any planned restart of a memory-buffered Fluentd, check buffer_queue_length and write_count. If the queue is non-zero, wait for it to drain or fix whatever is blocking the output first. A restart with a stalled output and a full memory buffer is a deliberate data loss event. For ungraceful events (OOM, crash), the data is already gone; the fix is making sure it cannot happen again.

Switch production outputs to the file buffer

Set the type and an explicit size limit per output:

<match **>
  @type forward
  <buffer>
    @type file
    path /var/log/fluent/buffer/forward
    total_limit_size 8GB
  </buffer>
  ...
</match>

Defaults if you do not set them: chunk_limit_size 256MB and total_limit_size 64GB for file buffers, versus 8MB and 512MB for memory. The 64GB default is almost never what you want. Set total_limit_size deliberately, size the filesystem to hold it with headroom, and keep the partition’s free space above roughly 2x the configured limit.

Two constraints from the file buffer documentation that bite in production:

  • Local disk only. Do not put the buffer path on remote filesystems (NFS, GlusterFS, HDFS). Major data loss has been observed with remote filesystems.
  • Path characters. On versions before v1.19.3, a path containing [ or ] could cause buffer chunks to be ignored after a restart (they are special characters in Dir.glob, and resume scanned the buffer directory with a glob); fixed in v1.19.3 (fluent/fluentd PR #5305), which resumes buffers with such paths correctly.

In Kubernetes, the buffer path must live on a persistent volume to survive pod replacement. A file buffer on ephemeral container storage is a memory buffer with extra steps: the container dies, the filesystem goes with it.

Handle the version-specific behaviors

  • v1.16.0 and later: corrupted chunk files found at startup are moved to a backup directory instead of silently deleted. On older versions they were just deleted, so a crash could lose file-buffered data too.
  • v1.19.0 and later: file buffer plugins evacuate chunk files when retry limits are exceeded instead of discarding the queue; they are written to ${root_dir}/buffer/${plugin_id}/ (root_dir is system_config.root_dir, or the default backup dir /tmp/fluent, overridable via FLUENT_BACKUP_DIR). The memory buffer does not support this.

Accept the tradeoff honestly

File buffers cost disk I/O and disk capacity, and restart replay produces a burst of flushes and possible brief duplicates. That is the price of durability, and the official recommendation for usual workloads is the file buffer for exactly this reason. The legitimate use for memory buffers is high-throughput paths where losing a minute of data on restart is acceptable, decided on purpose, per output.

Prevention

  • Explicit @type file on every production output. Never rely on the fallback default or a plugin’s own default; write the buffer section deliberately.
  • Explicit total_limit_size per output. The 64GB file default and 512MB memory default are both wrong for most deployments; size from your actual log rate and acceptable backlog.
  • Pre-restart queue check in the runbook. No restart proceeds while buffer_queue_length is non-zero or write_count is flat.
  • Alert on unexpected restarts. Process liveness with sustained-failure gating, plus restart count tracking, because each ungraceful restart is a potential loss event even after you move to file buffers.
  • Periodic config audit via the API. ?with_config=true in CI or a scheduled check, diffing effective buffer config against intent, catches plugins that default to memory and reloads that partially applied.
  • Filesystem capacity headroom. Buffer partition free space above 2x total configured limits, monitored, because a full disk converts your durable buffer back into a data loss path via overflow.

How Netdata helps

  • Netdata’s Fluentd collector polls the monitor agent API and charts buffer_queue_length, buffer_total_queued_size, and buffer_available_buffer_space_ratios per output plugin, so the size of the data at risk on any given restart is visible before you pull the trigger.
  • Correlating write_count rate against retry_count on the same dashboard tells you whether a restart would land on a draining pipeline or a stalled one.
  • Process RSS and restart events on the host side, viewed next to buffer queue growth, expose the OOM-kill-on-memory-buffer loop that produces repeated silent gaps.
  • After switching to file buffers, buffer_oldest_timekey trending back toward current time after a restart confirms the replay drained and the durability fix actually works.