The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-config-reload-failed ▌

Operations Guides

Fluentd config reload failed: SIGHUP that partially applies

You edited the Fluentd config, sent SIGHUP, watched the log line saying the config reloaded, and moved on. Hours later you notice a new output never started receiving data, or an old filter is still dropping events you told it to keep. The process never crashed. No error fired. But the pipeline running in memory is not the pipeline in the config file.

This is the partial reload failure: some plugins reloaded with the new configuration, others are still running the old one, and one or more plugins you expected are absent from the running process. Because Fluentd stays up and keeps processing events through whatever plugins did load, this failure is silent unless you verify the loaded plugin inventory after every reload.

The defining check: the plugin set reported by the monitor agent at /api/plugins.json must exactly match the plugin set implied by the config file. If a plugin is missing after a reload, the reload did not fully apply, and you are running a hybrid of the old and new configuration.

What this means

On SIGHUP, Fluentd does not patch configuration in place. The supervisor kills the worker process and respawns it with the new config. Plugins that use shared server sockets (via server_helper) keep their listen sockets across the respawn, so there is no connection downtime, but everything else is rebuilt: buffers, filters, outputs, parsers. A CPU spike and a brief throughput dip during this window is normal.

The respawn can fail partway. A plugin that fails to initialize during the respawn does not necessarily kill the process. Depending on where the failure lands, you can end up with a worker that loaded most of the new config, kept some old state, or silently dropped a plugin that could not start. Fluentd logs may show a reload message even when individual plugin initialization failed.

Since v1.9 there is also SIGUSR2 graceful reload, which reloads config in-process without restarting the worker, with its own limitations (changes to <system> are ignored, and any plugin using class variables makes the whole reload fail with “Failed to reload config file: Unreloadable plugin”). Since v1.18, zero-downtime restart (SIGUSR2 to the supervisor) is the recommended approach. Which mechanism you used changes what “partially applied” looks like, but the verification step is the same: diff the running plugin inventory against the config.

Common causes

CauseWhat it looks likeFirst thing to check
Plugin failed to initialize during worker respawnPlugin present in config, absent from /api/plugins.json after SIGHUPFluentd log around the reload timestamp for plugin load or configure errors
Unreloadable plugin blocks in-process reload“Failed to reload config file: Unreloadable plugin” in logs after SIGUSR2; nothing changedWhether any plugin uses class variables; fluent-plugin-prometheus before v1.8.2 is a known case (USR2 reload support was added in v1.8.2, and v1.8.5 for in_prometheus_monitor)
Third-party plugin breaks core at load timeReload or restart fails after installing/upgrading a plugin; parser errors on previously valid configPlugins that require 'yajl/json_gem' replace Ruby’s JSON module and break Fluentd’s config parser (splunkhec pre-v2.1 is the documented case; fixed in fluent-plugin-splunkhec 2.1)
Signal delivered to wrong process or process groupSIGHUP kills a worker with SignalException instead of triggering a clean reload, especially in containersContainer init system; dumb-init forwarding signals to the whole process group (fixed by DUMB_INIT_SETSID=0)
Old worker state not cleaned upRepeated SIGHUPs produce “No such process” errors or leaked workers on older versions`ps aux
<worker> directive reload collisionConfig works at startup, fails on reload with “specified worker_id collisions is detected”Whether the config uses <worker N> sections; a known reload-only failure
<system> changes with SIGUSR2New <system> settings silently not appliedGraceful reload ignores <system> changes by design; a restart is required

Quick checks

These are all read-only.

# 1. Full plugin inventory the worker actually loaded
curl -s http://localhost:24220/api/plugins.json | jq '.plugins[] | {id: .plugin_id, type: .type, category: .plugin_category}'

# 2. How many plugins of each category are running
curl -s http://localhost:24220/api/plugins.json | jq '[.plugins[].plugin_category] | group_by(.) | map({(.[0]): length}) | add'

# 3. What config the process thinks it has
curl -s http://localhost:24220/api/config.json | jq .

# 4. Process tree: supervisor plus expected worker count, nothing else
ps aux | grep '[f]luentd'

# 5. Reload-related log lines around the time you sent the signal
grep -iE "reload|sighup|unreloadable|failed to (configure|start)" /var/log/td-agent/td-agent.log | tail -30
# fluent-package: /var/log/fluent/fluentd.log

# 6. Plugin load errors since the reload
journalctl -u td-agent --since "30 minutes ago" | grep -iE "(load_plugin|LoadError|uncaught|cannot load)"

Notes on these:

  • Port 24220 assumes a single worker. In multi-worker mode, query each worker’s port (24220, 24221, …). Each worker has an independent plugin inventory.
  • /api/plugins.json requires <source> @type monitor_agent </source> in the config. If it was never configured, you have no in-band way to verify plugin state, which is itself a gap worth fixing.
  • Since v1.19.3, the config and retry fields are no longer included by default, and the old ?with_config=true / ?with_retry=true query parameters stopped working in the same release. Set include_config true or include_retry true in the monitor_agent <source> section to get them back.

How to diagnose it

flowchart TD
  A[SIGHUP or SIGUSR2 sent] --> B{Reload message in log?}
  B -- no --> C[Signal never reached supervisor: check container init, PID, signal routing]
  B -- yes --> D[Fetch /api/plugins.json]
  D --> E{Plugin set matches config?}
  E -- yes --> F[Reload applied; investigate elsewhere]
  E -- plugin missing --> G[Check log for plugin init failure at reload time]
  E -- old config still active --> H{SIGUSR2 used?}
  H -- yes --> I[Unreloadable plugin or ignored system section; restart required]
  H -- no --> J[Worker respawn failed partway; check for plugin load errors, worker_id collisions]
  G --> K[Fix plugin, then full restart, not another reload]
  I --> K
  J --> K
  1. Snapshot the expected plugin set from the config. Count the <source>, <filter>, <match>, and <label> plugin blocks in the effective config file. Include plugins that other plugins start implicitly only if you know they exist; the point is a baseline of the plugins you declared.

  2. Fetch the actual running set. Pull /api/plugins.json and list every plugin_id, type, and plugin_category. Do this per worker port in multi-worker mode.

  3. Diff expected versus actual. Any plugin in the config but not in the API output failed to load during the reload. Any plugin in the API output with stale behavior (old routing, old match rules) kept old state while others reloaded: the partial apply.

  4. Check the log at the reload timestamp. Look for plugin configure errors, LoadError, “Unreloadable plugin”, or “specified worker_id collisions is detected”. The reload log line and the failure line are usually close together but easy to miss if you only grepped for “reload”.

  5. Confirm which reload mechanism ran. SIGHUP respawns the worker; SIGUSR2 reloads in process. If you sent SIGUSR2 and one plugin is unreloadable, the entire reload fails atomically and everything stays on the old config. If you sent SIGHUP, the failure is per-plugin during respawn, which is what produces the true partial state.

  6. Check the process tree for anomalies. More workers than configured, a worker with a much older start time than its siblings, or leftover processes from before the reload all indicate the respawn did not complete cleanly.

  7. Decide: another reload will not fix it. Once the running state and the config diverge, the only reliable convergence is a full process restart. Repeated SIGHUPs against a wedged worker have produced SIGKILLed workers and leaked processes in reported issues.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Plugin inventory from /api/plugins.jsonThe ground truth of what is actually runningSet does not match config after a reload or restart; plugin count changed unexpectedly
emit_records per input pluginA new input that failed to load emits nothingNew source shows zero records after the config change that added it
write_count / emit_records per output pluginA missing output silently stops delivery to its destinationNew output flatlines at zero; old output you removed is still delivering
Monitor agent responsiveness (HTTP 200 within 5s)Confirms the event loop survived the reloadTimeout or non-200 right after SIGHUP means the respawn wedged the worker
Worker process count and start timesDetects leaked or mixed-generation workersMore processes than workers N, or divergent start times after a reload
Plugin load errors in Fluentd’s own logThe only place init failures are recordedAny LoadError or configure error at reload time

Fixes

Plugin failed to initialize during the reload

Find the init error in the log, fix the plugin configuration or the missing gem, then do a full restart of Fluentd rather than another SIGHUP. A restart guarantees every plugin initializes from a clean process. Tradeoff: a restart drops memory-backed buffer contents and resets in_tail to its pos_file positions, so expect a brief gap or duplicate window with file-backed buffers. That replay is normal.

Unreloadable plugin blocks SIGUSR2

The check is mechanical: Fluentd refuses to reload if any plugin class defines class variables. The known offender was fluent-plugin-prometheus before v1.8.2. Upgrade the plugin. If no upgrade exists, stop using graceful reload for that deployment and use SIGHUP or a full restart instead. Any plugin can be unreloadable; there is no maintained list, so treat the log error as the discovery mechanism.

A plugin breaks Fluentd core at load time

Plugins that require 'yajl/json_gem' replace Ruby’s JSON module globally, which breaks Fluentd’s config parser (it depends on JSON.parse raising on incomplete input; the yajl compatibility layer returns {} instead). This has broken reloads in the wild (splunkhec before v2.1). Remove or upgrade the offending plugin, then restart. This class of bug is why plugin upgrades deserve the same change control as config changes.

Container signal routing

In Docker with dumb-init, docker kill -s HUP <container> delivered SIGHUP to the entire process group, so the worker got the signal directly and died with SignalException instead of the supervisor handling a reload. Set DUMB_INIT_SETSID=0, or send the signal to the supervisor PID from inside the container, or use an orchestrator-native rolling restart instead of in-place reload.

Diverged state that will not converge

If the running plugin set and the config disagree and the cause is not obvious, do not keep sending signals. Capture the evidence (/api/plugins.json output, /api/config.json output, log excerpt, process list), then restart the process. Reload is a convenience; restart is the convergence mechanism.

Prevention

  • Make post-reload verification a step, not an afterthought. Any automation that sends SIGHUP should immediately fetch /api/plugins.json, compare the plugin set to the rendered config, and fail the pipeline on mismatch.
  • Baseline the plugin inventory. Record the expected plugin IDs per worker in your deployment tooling. Alert when the live set deviates, not just when the process dies.
  • Enable the monitor agent everywhere. Without in_monitor_agent, this failure class is undetectable except by noticing missing data downstream.
  • Prefer restart-friendly reload paths on modern versions. Since v1.18, zero-downtime restart (SIGUSR2 to the supervisor) is the documented recommendation over graceful reload, and it supports in_udp, in_tcp, and in_syslog via shared sockets.
  • Keep plugins current and few. Unreloadable plugins and core-breaking requires are third-party plugin behaviors. Fewer, newer plugins means fewer reload failure modes.
  • Test reloads in staging with the exact production plugin set. A config that reloads cleanly without your output plugins tells you nothing.

How Netdata helps

  • Netdata collects Fluentd’s monitor agent metrics continuously, so you have a per-second history of emit_records and write_count per plugin across the reload boundary. A plugin that vanished in a partial reload shows up as an abrupt stop in its series at the reload timestamp, even if nobody checked the API at the time.
  • Because counters are per plugin ID, you can correlate “new plugin never emitted” against “reload happened at 14:03” without parsing logs first.
  • Process-level metrics (process count, RSS, CPU per process) expose leaked or mixed-generation workers after a bad respawn, which pure Fluentd metrics miss.
  • Alerting on plugin throughput dropping to zero while its input source is still producing catches the silent half of a partial reload: the config that applied to some plugins but not the one that moves your data.
  • Per-worker history matters in multi-worker deployments, where one worker can fail a reload while its siblings apply it cleanly.