The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-unauthorized-forward ▌

Operations Guides

Fluentd unauthorized in_forward connections: log injection on the aggregator tier

You ran ss -tn on your Fluentd aggregator and saw connections to port 24224 from IP addresses you do not recognize. Or a security review turned up the fact that your aggregator-tier Fluentd accepts forwarded events from anything that can reach it. Either way: your in_forward input is unauthenticated, and any network-reachable host can inject events directly into your log pipeline.

An in_forward source without a <security> section performs zero authentication. Any host that can open a TCP connection to port 24224 can submit events with arbitrary tags and arbitrary record contents. The Fluentd project stated this plainly when authentication was introduced in v0.14.5: anyone who can connect to the TCP port of in_forward can inject events into the Fluentd process.

The aggregator tier is where this hurts most. A node agent tailing local files only trusts the host it runs on. An aggregator exists to accept events from the network, so its exposure surface is the network itself.

What this means

The forward protocol is how Fluentd instances ship events to each other. A node-level agent runs out_forward, the aggregator runs in_forward (default port 24224), and events flow over MessagePack. Without a <security> block, the aggregator does not verify who the sender is, what tags it uses, or what the records contain. The routing engine matches injected tags against your <match> directives and pushes forged events through filters, buffers, and outputs as legitimate traffic.

That single misconfiguration opens three attack modes:

flowchart LR
  A[Any reachable host] -->|TCP 24224, no auth| B[in_forward on aggregator]
  B --> C[Log injection: forged events, false audit trails]
  B --> D[Flooding: buffer fills, overflow fires]
  B --> E[Downstream poisoning: crafted payloads hit ES, SIEM]
  E --> F[tag-based path traversal in file outputs]
  1. Log injection. The attacker picks tags that match your routing rules and writes forged events. Audit trails, access logs, and security events can be fabricated wholesale. If you use logs for forensics or compliance, their evidentiary value is gone the moment an unauthenticated sender can write to them.

  2. Flooding and buffer-overflow DoS. A hostile sender pushes events faster than your outputs can drain. The buffer queue grows, buffer_available_buffer_space_ratios drops toward zero, and the overflow_action fires. With the default throw_exception, legitimate events are discarded. This is the backpressure cascade from buffer queue length growing, triggered deliberately.

  3. Downstream poisoning. Crafted record contents are delivered to systems that trust log data: Elasticsearch, SIEM correlation rules, alerting pipelines. A specific and severe variant: if any output plugin builds file paths with ${tag} placeholders, injected tags containing ../ sequences can write outside the intended directory. (CVE-2026-44024, GHSA-44hj-4m45-frj3 — arbitrary file write via ${tag} path validation, CVSS 9.8 critical, fixed in v1.19.3, official Fluentd security advisory). A related decompression issue lets a small gzip-compressed payload expand and exhaust memory, because older versions limited compressed payload size but not decompressed size. This is CVE-2026-44160, GHSA-j9cw-hwqf-85w7 — a gzip decompression bomb in in_http and in_forward (CVSS 7.5 high, fixed in v1.19.3; affected versions up to v1.19.2).

There is also a quieter exposure: in_http (default port 9880) has no built-in authentication mechanism. If it is enabled and reachable, it accepts events from anyone, with no <security> option available. Network-level controls are the only defense.

Common causes

CauseWhat it looks likeFirst thing to check
in_forward has no <security> sectionAny host can connect and emit events; unknown peer IPs in ss outputgrep -A5 '@type forward' /etc/fluent/fluentd.conf and look for a <security> block
<security> exists but allow_anonymous_source left at defaultShared key configured, yet unauthenticated senders still acceptedCheck for allow_anonymous_source false and <client> sections; the default is true even with <security> present
Port 24224 exposed beyond the expected networkForward port reachable from the internet, other VPCs, or the whole pod networkTest from a host that should NOT have access: nc -zv <aggregator> 24224
in_http enabled for convenience and forgottenEvents accepted on 9880 from anywhere; no auth option existsCheck config for @type http sources and firewall rules for the port
Shared key widely distributed or leakedEvery node agent has the same key; key appears in repos or imagesAudit where the key is stored; treat it as a credential
Fluentd version missing current security fixesOld version without the decompression limit or ${tag} traversal patchfluentd --version

Quick checks

All read-only and safe to run on a production aggregator.

# 1. Who is connected to the forward port right now? (peer address is column 5)
ss -tn src :24224 | awk 'NR>1 {print $5}' | sort | uniq -c | sort -rn

# 2. Every unique peer seen on the port
ss -tnp | grep ":24224" | awk '{print $5}' | sort -u

# 3. Does the forward source have a security section?
grep -A20 '@type forward' /etc/fluent/fluentd.conf | grep -iE 'security|shared_key|allow_anonymous|transport'
# (Use /etc/td-agent/td-agent.conf for td-agent installs.)

# 4. Is in_http also listening?
grep -B2 -A5 '@type http' /etc/fluent/fluentd.conf
ss -tlnp | grep -E '9880|24224'

# 5. Authentication and TLS errors in the Fluentd log
grep -iE "(tls|ssl|auth|unauthorized|forbidden|certificate)" \
  /var/log/fluent/fluentd.log | tail -20

# 6. Fluentd version (CVE exposure depends on it)
fluentd --version

# 7. Is the port reachable from outside the expected sender network?
# Run this FROM a host that should not be allowed:
nc -zv <aggregator-ip> 24224

Interpretation notes:

  • Checks 1 and 2 compare reality against your sender inventory. Every peer IP should map to a known node agent, forwarder, or load balancer. Anything else is an incident.
  • Check 3 returning nothing means no authentication at all. Returning shared_key without allow_anonymous_source false means partial protection only.
  • Check 7 is the one teams skip. Internal firewalls, security groups, and Kubernetes NetworkPolicies drift. Assume the config file lies and test the path.

How to diagnose it

  1. Inventory your legitimate senders. List the nodes, forwarders, and services that should be forwarding to this aggregator. Use IP ranges, not just hostnames, because that is what <client> sections and firewall rules match on.

  2. Diff live connections against the inventory. Use checks 1 and 2 above. Unknown source IPs mean either an unauthorized sender or a stale inventory. Resolve that ambiguity first; do not assume benign.

  3. Audit the effective config. @include directives mean the <security> block may live in a different file than the <source>. Verify against the running process: query http://localhost:24220/api/config.json on each worker’s monitor_agent port (24220, 24221, …) and inspect the forward source’s configuration as Fluentd actually loaded it.

  4. Prove the injection path, safely. From a host that is not an authorized sender, push a tagged test event:

# From an unauthorized host: this should FAIL on a locked-down aggregator
echo '{"msg":"unauthorized injection test"}' | \
  fluent-cat --host <aggregator-ip> --port 24224 security.test

Then check whether security.test events arrive at your outputs. If they do, the pipeline accepted forged events and you have confirmed the exposure end to end. If the connection is refused or the handshake fails, authentication is working.

  1. Check for signs it already happened. Look for tags in your outputs that no configured input produces, sudden unexplained input rate spikes (input emit_records rate far above baseline), or buffer pressure events with no corresponding application log volume. The emit records rate gap and unexplained queue growth are your historical breadcrumbs.

  2. Check version exposure. If the running version predates the decompression-limit and ${tag} traversal fixes, those mitigations are absent. On unauthenticated inputs they are directly exploitable. The fix landed in v1.19.3; decompression_size_limit (default 256MB) was introduced in that release, so prior versions have no decompressed-size limit at all.

Signals to monitor

SignalWhy it mattersWarning sign
Source IPs on port 24224The direct unauthorized-sender signalAny peer outside the sender allowlist
Input emit_records rate per sourceFlooding shows as a rate spike from one sender before buffers feel itSustained deviation >50% from rolling baseline
buffer_queue_length and buffer_available_buffer_space_ratiosA flood attack manifests here as it becomes a DoSQueue growing with available ratio under 20%
Tags observed at outputs vs tags produced by configured inputsInjected events arrive with attacker-chosen tagsTags present downstream that no input source emits
TLS/auth errors in Fluentd logFailed handshakes indicate probing or misconfigured sendersRepeated handshake failures from one IP
Fluentd version across the fleetDetermines exposure to the path-traversal and decompression CVEsAny aggregator running an unpatched version with a network-reachable forward port

How to lock it down

Enable shared-key authentication

Add a <security> section to the forward source on every aggregator. The handshake is a SHA512-based challenge-response (HELO/PING with a nonce); the key itself never crosses the wire. Built in since v0.14.5.

<source>
  @type forward
  port 24224
  <security>
    self_hostname aggregator.internal
    shared_key YOUR_LONG_RANDOM_KEY
    allow_anonymous_source false
    <client>
      network 10.1.0.0/16
      shared_key YOUR_LONG_RANDOM_KEY
    </client>
  </security>
</source>

Two details that bite operators:

  • allow_anonymous_source defaults to true even when <security> is present. Setting a shared_key alone does not lock the port down. You must explicitly set it to false and define <client> sections for your sender networks.
  • Senders must configure the matching key in out_forward (shared_key in their <server> section). Roll this out in coordination: senders first with the key, then the aggregator, or you will break legitimate traffic.

With authentication enabled, only MessagePack payloads are supported; JSON-format single events over forward are unavailable per the protocol spec. Standard out_forward uses MessagePack, so this mainly matters for custom senders.

Add TLS, preferably mutual TLS

Encryption without authentication only hides the injected traffic. Combine both. Built-in TLS transport exists since v0.14.12; before that, the third-party fluent-plugin-secure-forward gem was required and is now superseded. client_cert_auth true with a ca_path gives mutual TLS, where the aggregator verifies client certificates (mutual TLS support was added in v1.1.1).

<source>
  @type forward
  port 24224
  <transport tls>
    cert_path /etc/fluent/certs/aggregator.pem
    private_key_path /etc/fluent/certs/aggregator.key
    client_cert_auth true
    ca_path /etc/fluent/certs/ca.pem
  </transport>
  <security>
    self_hostname aggregator.internal
    shared_key YOUR_LONG_RANDOM_KEY
    allow_anonymous_source false
    <client>
      network 10.1.0.0/16
      shared_key YOUR_LONG_RANDOM_KEY
    </client>
  </security>
</source>

mTLS is the stronger control because it does not depend on a symmetric key copied to every node. The tradeoff is certificate lifecycle management: expiry causes cliff-edge simultaneous failures across all senders, so monitor certificate end dates.

Restrict the network path

Authentication is defense in depth, not a substitute for reachability control:

  • Bind in_forward to an internal interface only (bind 10.1.0.5), not 0.0.0.0, unless you genuinely accept events from multiple networks.
  • Firewall or security-group port 24224 to the sender CIDR ranges.
  • In Kubernetes, use NetworkPolicy to restrict which pods can reach the aggregator’s port.
  • For in_http, which has no auth option, network restriction is the only control. If you cannot restrict it, do not expose it.

Upgrade and harden against known CVEs

  • Upgrade to the current Fluentd release that carries the ${tag} path traversal and gzip decompression fixes. (A fix is in Fluentd v1.19.3, bundled for the fluent-package v6.0.4 LTS distribution; CVE-2026-44024 is CVSS 9.8 critical and CVE-2026-44160 is CVSS 7.5 high, and the new decompression_size_limit parameter defaults to 256MB — official Fluentd security advisory and v1.19.3 release notes).
  • Independently of version: do not use ${tag} in output file paths when any input accepts events from the network, and run Fluentd as a non-root user so a traversal write cannot touch system files.
  • In mixed fleets, check whether Fluent Bit instances face the same exposure; Fluent Bit is a separate codebase with its own in_forward authentication history. (CVE-2025-12969 is a real advisory: Fluent Bit’s in_forward does not enforce the security.users authentication under certain configuration conditions).

Prevention

  • Baseline config template. Every aggregator-tier forward source ships with <security>, allow_anonymous_source false, and <client> network sections from day one. Make “no security section” a code review failure.
  • Sender allowlist as data. Keep the expected sender CIDR list in one place and generate both the <client> sections and the firewall rules from it. Drift between the two is how gaps appear.
  • Continuous peer auditing. Alert on any source IP on 24224 outside the allowlist. It is a binary signal and cheap to evaluate.
  • Treat the shared key as a credential. Store it in your secrets system, rotate it on the same schedule as other credentials, and never commit it to a repo or bake it into an image.
  • Config file integrity monitoring. A tampered Fluentd config can disable the security section or add an unauthorized output. Alert on config file modification outside the deployment pipeline.
  • Version tracking. Track Fluentd versions fleet-wide. Security fixes only help if they are actually deployed.

How Netdata helps

  • Netdata collects Fluentd monitor_agent metrics per worker, so input emit_records rate spikes from a flooding sender are visible per aggregator instance rather than averaged away.
  • Correlating input rate with buffer_queue_length and buffer_available_buffer_space_ratios on one dashboard separates a flood-driven buffer fill from a destination-driven one: in a flood, input rate leads the queue growth.
  • Host-level network visibility shows which peers hold connections to the Fluentd process, making an unknown sender on 24224 visible without a manual ss session.
  • Process RSS trends catch the memory-exhaustion variant of an attack (decompression bombs, oversized events) before the OOM killer does.
  • Fluentd log collection lets you alarm on repeated TLS and auth handshake failures, which is what probing looks like before a successful injection.