The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-monitoring-maturity-model ▌

Operations Guides

Fluentd monitoring maturity model: from survival to expert

Most teams monitoring Fluentd stop at “is the process running?” and discover the gap during an incident, when the SIEM is missing the exact logs they need for a postmortem. Fluentd can be alive, responsive, and green on every dashboard while silently dropping events, accumulating a buffer that will overflow in forty minutes, or retrying into a backoff so deep the pipeline is effectively dead.

This is a maturity model for Fluentd monitoring: four levels, each with the specific signals, collection commands, and failure modes it catches. Use it as a self-assessment. Find the highest level where you have every signal covered, with alerting, in production. Everything above that is your roadmap.

The model is cumulative. Level 3 without Level 1 coverage is a monitoring illusion: sophisticated dashboards on top of a pipeline that can still die unobserved.

flowchart TD
  L1["Level 1: Survival
Process alive, emit errors,
retries, RSS, error logs"] L2["Level 2: Operational
monitor_agent, per-output buffers,
in/out rate balance, FD, CPU, disk"] L3["Level 3: Mature
pos lag, flush latency, oldest timekey,
stage vs queue, retry backoff, TLS expiry"] L4["Level 4: Expert
Ruby GC, per-thread CPU, UDP drops,
inotify, per-worker, config integrity"] L1 --> L2 --> L3 --> L4

Level 1: Survival

Goal: you know within minutes if Fluentd is dead, hung, or actively losing data. Five signals. Nothing here requires the monitor_agent API.

The signals

Process liveness. Not just “a PID exists”: in multi-worker mode the supervisor can stay alive while workers die, and systemctl status showing active only confirms the supervisor. Gate the page on sustained absence (more than 2 minutes) so transient restarts, rolling updates, and container rescheduling do not page you. In Kubernetes, CrashLoopBackOff shows up as repeated brief absences rather than one long one.

# Process liveness
systemctl is-active td-agent    # td-agent package
systemctl is-active fluentd     # fluent-package
pgrep -af fluentd               # generic; count workers vs configured N

Emit failures. The data-loss signal. Any failed emit means events were dropped and will never be delivered, most commonly because the buffer hit total_limit_size and the default overflow_action: throw_exception rejected them. Fluentd has no emit_error_count field in the monitor_agent API; watch the emit transaction failed and buffer overflow lines in the Fluentd log and the per-plugin counters (retry_count, drop_oldest_chunk_count).

retry_count per output. Any sustained non-zero value means a destination is failing and events are accumulating in buffers. This is a cumulative counter: treat any increment in a healthy system as a ticket.

Process RSS trend. Ruby memory grows and plateaus; a high-but-stable RSS is normal. A monotonically rising trend over hours is a leak, unbounded memory buffers, or tag explosion, and it ends in an OOM kill. In containers, alert at 80% of the cgroup limit. After an OOM, confirm with dmesg | grep -i oom.

# RSS trend (collect over time, not as a one-shot)
ps -o rss= -p $(pgrep -f fluentd | head -1) | awk '{print $1/1024 " MB"}'

Fluentd error-log check. Fluentd’s own logs are the last place teams look. Tail them for connection failures, TLS errors, and plugin exceptions:

# Error-log check (paths vary: /var/log/td-agent/ or /var/log/fluent/)
grep -iE "error|failed|retry|overflow" /var/log/td-agent/td-agent.log | tail -20

What Level 1 misses

A running process with a full buffer and a stalled output looks identical to a healthy one. Level 1 tells you the pipeline is broken; it does not tell you where, and it does not warn you before data loss starts.

Level 2: Operational

Goal: you know when the pipeline is backing up or retrying, per output, before the buffer overflows. This level requires the monitor_agent API.

Enable the API

<source>
  @type monitor_agent
  bind 127.0.0.1
  port 24220
</source>

Notes: the port auto-increments per worker (worker 0 = 24220, worker 1 = 24221). Bind to localhost or firewall it; the endpoint exposes internal state. Since v1.19.3 the monitor_agent security defaults changed: include_config and include_retry now default to false (and include_debug_info, new in v1.19.3, also defaults to false). If your dashboards or scripts relied on config or retry fields from older versions, verify they still return data after upgrading.

The signals

Per-output buffer_queue_length and buffer_total_queued_size. Queue length tells you the backlog in chunks; total queued bytes tells you the real footprint, since chunk sizes vary. Alert on sustained growth, not a static threshold: legitimate batch work creates temporary queues. Compare total bytes against total_limit_size and, for file buffers, against filesystem free space, whichever binds first.

# Per-output buffer state
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") |
      {id: .plugin_id, queue: .buffer_queue_length,
       bytes: .buffer_total_queued_size,
       avail_pct: .buffer_available_buffer_space_ratios}'

buffer_available_buffer_space_ratios. Percentage of configured buffer space remaining. Below 20% and still filling is a ticket; below 5% and filling, overflow is imminent. Estimate time-to-overflow as available_space / growth_rate.

Input vs output emit_records balance. The single most telling health metric of the pipeline. Compute both rates as deltas over time:

# Input and output record totals (derive rates from deltas)
curl -s http://localhost:24220/api/plugins.json | \
  jq '{in: [.plugins[] | select(.plugin_category=="input") | .emit_records // 0] | add,
       out: [.plugins[] | select(.plugin_category=="output") | .emit_records // 0] | add}'

The output/input ratio should approach 1.0 over a window of at least max(2 * flush_interval, 10 minutes). A sustained deficit means data loss or unbounded buffer growth. Caveats: on Fluentd versions before v1.19.0, input emit_records requires enable_input_metrics true in <system> or the counter is always 0 (the parameter was introduced in v1.14.0 with default false, and the default flipped to true in v1.19.0); verify on your version before trusting the ratio. Time-sliced outputs (daily S3 files) legitimately diverge; use buffer_oldest_timekey (Level 3) for those.

File descriptors. Fluentd holds one FD per tailed file, plus buffer chunk files, plus output connections. Alert above 75% of the soft limit:

# FD usage vs limit
ls /proc/$(pgrep -f fluentd | head -1)/fd | wc -l
grep "Max open files" /proc/$(pgrep -f fluentd | head -1)/limits

The default ulimit -n of 1024 is inadequate for production. FD exhaustion shows up as in_tail silently losing files or outputs failing with cryptic errors, not a clear message.

CPU per worker process. Due to the GVL, a single worker saturates one core: “100% CPU” on the process while the host looks idle means the pipeline is CPU-bound. Sustained usage above roughly 70% of one core caps throughput; plan for multi-worker or simpler parsers.

Buffer disk usage. For file-backed buffers, watch both the buffer directory size and the partition it lives on. Buffer files sharing a partition with system logs is a cascade waiting to happen.

What Level 2 misses

You see that the pipeline is degrading, but not why, and not how stale the data is. Level 2 also cannot see retry state depth, input-side lag, or slow-but-successful flushes.

Level 3: Mature

Goal: you understand where and why the pipeline is degrading, with leading indicators instead of cliff-edge alerts.

The signals

in_tail position lag. Compare the recorded position in the pos_file against actual file size. A gap that grows over time means ingestion is falling behind; an inode mismatch after rotation means Fluentd lost the file:

# pos_file vs reality (format: path<TAB>position<TAB>inode)
while IFS=$'\t' read -r filepath pos inode; do
  printf '%s pos=%s pos_inode=%s actual_size=%s actual_inode=%s\n' \
    "$filepath" "$pos" "$inode" \
    "$(stat -c%s "$filepath" 2>/dev/null || echo 0)" \
    "$(stat -c%i "$filepath" 2>/dev/null || echo MISSING)"
done < /var/log/td-agent/td-agent.pos

Also track tracked_file_count (v1.19.0+, a gauge of files currently tailed) and rotated_file_count (v1.14.1+) to confirm rotation is being detected on schedule.

Flush latency. flush_time_count / write_count gives average flush time in milliseconds. Rising average flush time is the earliest indicator of destination degradation, appearing before retries start. Track slow_flush_count (flushes over slow_flush_log_threshold, default 20s) as the outlier counter. Healthy: average flush time under 50% of flush_interval.

buffer_oldest_timekey. The age of the oldest buffered data. now - buffer_oldest_timekey beyond roughly 2 * flush_interval for non-time-sliced outputs means a severe delivery backlog. This is safer than rate comparisons for time-sliced outputs and bursty workloads.

Stage vs queue distinction. High buffer_stage_length with low buffer_queue_length is healthy batching. Low stage with high queue is backpressure. Teams that watch one “buffer usage” number miss this entirely. Rule of thumb: queue sustained above 5x stage means the output is struggling.

Retry backoff state. retry_count tells you errors happened; the retry object tells you how bad it is. retry.steps climbing with retry.next_time far in the future means the pipeline is effectively stalled even while “retrying”. If next_time is 30 minutes out, recovery will be slow even after the destination returns. Note the default change above: retry fields require include_retry true in the monitor_agent config on current versions.

TLS certificate expiry. Certificates expire at an exact moment and fail every output connection at once, as a cliff. Fluentd does not expose expiry through the API; check externally:

# Certificate expiry (paths from your config)
openssl x509 -enddate -noout -in /path/to/cert.pem

Alert at 30 days (plan), 7 days (act), 24 hours (page).

Confirmed-loss counters. drop_oldest_chunk_count incrementing means data was permanently discarded; write_secondary_count non-zero means the primary output failed past its retry limits and chunks went to the fallback. Both are ticket-worthy on any increment; sustained drops with a near-full buffer and a stalled output are a page.

Level 4: Expert

Goal: you catch silent failures, runtime-level contention, and configuration drift that standard metrics never surface. These signals typically get added after painful incidents.

Ruby GC behavior. GC pauses block all event processing. GC statistics are not cleanly exposed externally; infer pressure from RSS churn, context switches, and periodic throughput dips correlating with GC cycles. Major GC storms create a feedback loop: pauses cause flush timeouts, which cause retries, which allocate more objects. Tuning is environment-dependent; the official performance guide suggests RUBY_GC_HEAP_OLDOBJECT_LIMIT_FACTOR=1.2 for memory-constrained environments (default 2.0).

Per-thread CPU for GVL contention. One thread pegged while others idle confirms parsing or serialization is starving flush threads:

# Per-thread CPU
ps -T -p $(pgrep -f fluentd | head -1) -o spid,%cpu,comm

The fix is multi-worker mode or simpler parsers, not more flush threads: CPU-bound work does not parallelize across Ruby threads.

Kernel UDP drops. For UDP inputs (syslog, forward over UDP), the kernel drops packets before Fluentd ever sees them when socket buffers overflow. Check the drops column in /proc/net/udp. This is upstream data loss invisible to every Fluentd metric.

Inotify watch exhaustion. in_tail relies on inotify on Linux. If watches run out, it falls back to polling, which is far less responsive:

# Inotify limit
cat /proc/sys/fs/inotify/max_user_watches

Per-worker decomposition. Workers have independent buffers, queues, and event loops. Aggregate metrics mask a struggling worker. Query each worker’s monitor_agent port (24220 + worker_id) separately. Remember that in_tail does not support multi-worker and must be pinned with <worker N>, so only that worker’s port shows tail metrics.

Configuration integrity. A tampered config can silently redirect logs or disable collection. Track modification times and hashes against your deployment pipeline:

# Config integrity
stat -c '%Y' /etc/td-agent/td-agent.conf
sha256sum /etc/td-agent/td-agent.conf

Complement this with network-level checks: outbound connections from the Fluentd process to unexpected destinations (ss -tnp | grep fluentd), and inbound connections to forward ports from outside the sender allowlist.

Assessing yourself: quick checklist

  • Liveness with sustained-failure gating, not a raw PID check. Worker count verified against configuration.
  • Data-loss counters (emit_error_count, drop_oldest_chunk_count) alerting on any increment, not on thresholds.
  • Input/output rate balance computed from deltas, with input metrics confirmed non-zero (the enable_input_metrics trap on older versions).
  • Stage and queue tracked separately. One combined buffer number is a blind spot.
  • Retry state, not just retry count. next_time in the future is a stalled pipeline.
  • Host-level signals covered: RSS trend, FD vs ulimit, buffer partition free space. None of these come from the API.
  • Version assumptions verified. monitor_agent defaults and available fields changed across v1.14.x, v1.19.0, and later releases; confirm what your version actually returns.

How Netdata helps

Fluentd’s signals split across two planes: the monitor_agent API (buffer, retry, flush, and emit metrics) and the host (RSS, FD count, disk, per-thread CPU, kernel UDP drops, inotify). Correlating across that split is where most manual setups fall apart. Specifically:

  • Netdata collects Fluentd plugin metrics from the monitor_agent endpoint and keeps them as per-second time series, so rate derivation (emit_records, flush_time_count, retry_count) and spike detection between infrequent samples are handled for you.
  • Buffer gauges (buffer_queue_length, buffer_total_queued_size, buffer_available_buffer_space_ratios) charted next to process RSS and FD usage make the backpressure-cascade pattern visible in one view instead of three terminals.
  • Per-worker decomposition maps naturally to per-instance and per-dimension views, so a single struggling worker is not averaged away.
  • Anomaly detection on input/output rate balance catches slow divergence, the 5% sustained deficit that accumulates into millions of lost events, without a hand-tuned threshold.
  • Host-level signals the API never exposes (per-thread CPU, disk usage of the buffer partition, /proc/net/udp drops) are collected by the same agent, closing the Level 4 gap.