The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-output-emit-lower-than-input ▌

Operations Guides

Fluentd output rate lower than input rate: the deficit that is quietly dropping logs

Your Fluentd process is up. The monitor agent responds. retry_count is zero. And yet, when you query the destination, whole time windows of logs are missing. The cause is usually the same: output emit_records has been running below input emit_records for hours or days, and nobody was comparing them.

Most teams chart input rate and output rate independently and never compute the ratio. A sustained 5% deficit on a host doing 200 events per second is over 6 million events lost per week. The two rates should converge over any reasonable window. When they do not, data is either being dropped or piling up in a buffer that will eventually overflow and drop it anyway.

What this means

Every event flows through Input, Parser, Filter chain, Buffer, Output. Input and output plugins each maintain a cumulative emit_records counter. In a healthy pipeline, the output rate tracks the input rate. There are only three ways output can stay below input:

  1. Events are accumulating in the buffer. The output cannot keep pace. buffer_queue_length grows and buffer_available_buffer_space_ratios declines. This is debt that must be repaid, either by draining or by overflow.
  2. Events are being dropped deliberately. overflow_action drop_oldest_chunk is discarding old chunks (drop_oldest_chunk_count increments), or retries have been exhausted and chunks were discarded or sent to a secondary (write_secondary_count increments).
  3. Events are being dropped silently. The buffer is full and overflow_action is the default throw_exception. New events are rejected at the buffer and lost. No counter reliably tracks these drops. The only evidence is the rate gap itself, plus buffer_available_buffer_space_ratios pinned near 0%.

There is also one legitimate cause: time-sliced outputs (daily S3 files, for example) hold chunks until the slice expires, so input and output rates diverge by design and converge only when the slice rolls over.

flowchart TD
  A[Output rate below input rate sustained] --> B{Buffer queue growing?}
  B -- Yes --> C[Output cannot keep pace]
  C --> C1{write_count incrementing?}
  C1 -- No --> C2[Destination down or retrying: check retry_count and logs]
  C1 -- Yes, slowly --> C3[Destination slow: check flush_time_count per write_count]
  B -- No, buffer full --> D{overflow_action?}
  D -- drop_oldest_chunk --> E[Confirmed loss: drop_oldest_chunk_count increments]
  D -- throw_exception --> F[Silent loss at buffer: no counter, gap is the signal]
  B -- No, buffer healthy --> G{Time-sliced output?}
  G -- Yes --> H[Expected: rates converge at slice boundary]
  G -- No --> I[Misrouting or filter dropping: verify end to end]

Common causes

CauseWhat it looks likeFirst thing to check
Destination slow or throttlingAverage flush time rising, slow_flush_count incrementing, queue growingflush_time_count / write_count trend
Destination down, retries cyclingwrite_count flat, retry_count and rollback_count incrementing, queue growingFluentd error logs for the specific failure
Buffer full, throw_exception (default)Input rate normal, output rate capped, buffer_available_buffer_space_ratios near 0%, no error counters movingoverflow_action in config
drop_oldest_chunk absorbing the deficitQueue stays below max, drop_oldest_chunk_count incrementing, gap persistsdrop_oldest_chunk_count delta
Retry exhaustion discarding chunkswrite_secondary_count non-zero, or chunks gone after retry_timeout (default 72 hours)write_secondary_count, retry object state
Time-sliced output holding chunksGap follows the slice schedule and closes at rolloverbuffer_oldest_timekey vs current time
Events misrouted or dropped by a filterRate gap but buffers healthy and outputs fineQuery the destination for a known tag

Quick checks

All of these are read-only. They assume the monitor agent is enabled on the default port 24220. In multi-worker mode, each worker has its own port (24220 + worker_id), so check each one.

# Sum input emit_records and output emit_records
curl -s http://localhost:24220/api/plugins.json | \
  jq '{input: [.plugins[] | select(.plugin_category=="input") | .emit_records // 0] | add,
       output: [.plugins[] | select(.plugin_category=="output") | .emit_records // 0] | add}'

# Run it twice, 60 seconds apart, and compute per-second rates from the deltas.
# These are cumulative counters; raw values are meaningless without a delta.

# Per-output breakdown: writes, retries, queue, available space
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") |
      {id: .plugin_id, writes: .write_count, retries: .retry_count,
       rollbacks: .rollback_count, queue: .buffer_queue_length,
       avail_pct: .buffer_available_buffer_space_ratios,
       dropped: .drop_oldest_chunk_count, secondary: .write_secondary_count}'

# Average flush time per output (run twice, divide deltas)
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") |
      {id: .plugin_id, flush_ms: .flush_time_count, writes: .write_count,
       slow: .slow_flush_count}'

# Buffer data age: how far behind is the oldest chunk
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") |
      {id: .plugin_id, oldest: .buffer_oldest_timekey, newest: .buffer_newest_timekey}'

# Recent output errors from the Fluentd log (path varies by package)
grep -E "failed to flush|retry|BufferOverflowError|could not connect" \
  /var/log/td-agent/td-agent.log | tail -20

One prerequisite that bites people constantly: on Fluentd versions before v1.19.0, input emit_records requires enable_input_metrics true in the <system> section. Without it, the input counter is always 0 and the comparison is impossible. If your input sum comes back as zero while logs are clearly flowing, check that setting first.

How to diagnose it

  1. Confirm the deficit is real, not sampling noise. Both counters are cumulative. Take two samples at least max(2 * flush_interval, 10 minutes) apart, compute per-second rates, and compare. Require input rate greater than 0 before computing any ratio, or you will divide by zero on idle hosts. Batching causes short-term spikes; only a sustained gap over that window matters.

  2. Check the buffer to classify the deficit. Look at buffer_queue_length and buffer_available_buffer_space_ratios per output. Growing queue with declining available space means accumulation: the output cannot keep pace. Available space pinned near 0% with a stable queue means the buffer is full and the overflow action is firing: active loss.

  3. If accumulating, find out why the output is slow. Flat write_count with rising retry_count and rollback_count means the destination is failing; read the Fluentd error log for the actual error (connection refused, 401/403, 429, TLS). Writes incrementing but too slowly, with rising flush_time_count / write_count and slow_flush_count increments, means the destination is degraded, not dead.

  4. If the buffer is full, check which loss path you are on. drop_oldest_chunk_count incrementing means confirmed, counted loss of the oldest chunks. If it is zero and overflow_action is unset, you are on the default throw_exception: new events are being rejected at the buffer with no counter. Grep the Fluentd log for BufferOverflowError to confirm.

  5. Check retry exhaustion. If retry_forever is false and retry_max_times is unset, retries are governed by retry_timeout (default 72 hours), after which the chunk is discarded. If a <secondary> is configured, write_secondary_count tells you chunks have already fallen through to the backup destination. Also check the retry object: if retry.next_time is far in the future, exponential backoff has pushed recovery out even though the process looks alive.

  6. Rule out the legitimate cause. For time-sliced outputs, compare buffer_oldest_timekey to the current time and the configured timekey and timekey_wait. A gap that closes at each slice boundary is expected behavior, not loss.

  7. If buffers and outputs look healthy, suspect routing. Events can flow through Fluentd and land in a null output or the wrong destination. Everything looks green, but the data never arrives where you query it. Verify end to end: inject a test event with fluent-cat and confirm it appears at the destination.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
output emit rate / input emit rateThe core health ratio; should approach 1.0 over the convergence windowSustained below 1.0 over max(2 * flush_interval, 10 min) with input > 0
buffer_queue_lengthShows whether the deficit is accumulating as backlogSustained upward trend
buffer_available_buffer_space_ratiosTime remaining before overflow firesBelow 20% and still declining; near 0% means loss is happening now
write_countProves chunks are actually being deliveredFlat while input is active
retry_count, rollback_countDestination is rejecting or unreachableAny sustained non-zero rate
flush_time_count / write_countAverage flush time; the earliest sign of destination degradationRising trend; approaching flush_interval means the output cannot keep up
slow_flush_countFlushes exceeding slow_flush_log_threshold (default 20s)Rising ratio of slow flushes to total writes
drop_oldest_chunk_countConfirmed, counted data lossAny increment
write_secondary_countPrimary output failed exhaustively; data went to fallbackAny non-zero value
buffer_oldest_timekeyAge of the oldest undelivered dataLag exceeding 2 * flush_interval (or timekey + timekey_wait for sliced outputs)

Fixes

Destination slow or failing

Fix the destination first. No Fluentd-side tuning compensates for an Elasticsearch cluster in red, an expired credential, or a throttled S3 bucket. Check the destination independently, then read the Fluentd error log for the specific failure. If retry backoff has run away (retry.next_time far in the future), a Fluentd restart resets retry state, but only do this after the destination is healthy. A restart also resets all cumulative counters, including the ones you were just comparing. With file-backed buffers, unflushed chunks survive the restart and will replay; with memory-backed buffers, everything buffered is lost on restart.

Buffer full with throw_exception

You are losing new events silently. Decide explicitly which overflow semantics you want per output: block stops the loss but exerts backpressure on inputs (upstream senders may drop instead), drop_oldest_chunk keeps the pipeline flowing but loses old data in a countable way. There is no zero-loss option once the buffer is full; the real fix is capacity. Raise total_limit_size if disk allows, and treat buffer_available_buffer_space_ratios as the early warning you should have had.

Output cannot keep pace with input

If average flush time is approaching flush_interval, the output has no headroom. Options: increase flush_thread_count (helps for I/O-bound outputs, since network writes release the GVL; it does not help if the bottleneck is CPU-bound serialization), increase chunk_limit_size so each flush carries more records, or add workers in multi-worker mode for true parallelism. Each option trades memory, destination load, or operational complexity. If the destination itself is undersized for the event volume, scale the destination or shed input volume at the source.

Retry exhaustion discarding chunks

Set explicit retry policy instead of inheriting the 72-hour retry_timeout default, and configure a <secondary> output so exhausted chunks land somewhere recoverable instead of being discarded with only a log line. Monitor write_secondary_count so you know when the fallback has engaged.

Prevention

  • Compare the rates, always. Alert when output rate stays below input rate over a window of at least max(2 * flush_interval, 10 minutes), with the input > 0 guard. This one check catches nearly every loss path in this article, including the silent ones no error counter exposes.
  • Enable input metrics. On Fluentd before v1.19.0, add enable_input_metrics true to <system>. Without it the comparison is impossible.
  • Choose overflow_action explicitly per output. Never inherit the default throw_exception without understanding that it means silent loss under a full buffer.
  • Use file-backed buffers in production. Memory buffers lose all unflushed data on restart or OOM kill. The deficit you detect today becomes unrecoverable loss the moment the process dies.
  • Alert on the confirmation signals too. drop_oldest_chunk_count any increment, write_secondary_count any increment, buffer_available_buffer_space_ratios below 20% and declining. The rate ratio tells you something is wrong; these tell you what.
  • Account for time-sliced outputs in the alert. Exclude or widen the window for outputs whose gap closes at slice boundaries, or you will train the team to ignore the alert.

How Netdata helps

  • Netdata collects per-plugin emit_records for both inputs and outputs and derives rates automatically, so the input/output comparison is a first-class chart rather than a manual jq exercise.
  • Buffer gauges (buffer_queue_length, buffer_available_buffer_space_ratios, stage vs queue byte sizes) are graphed alongside the throughput rates, so you can see in one view whether a deficit is accumulating, full, or draining.
  • Loss-confirmation counters (drop_oldest_chunk_count, write_secondary_count, retry_count, rollback_count) are tracked per output plugin, letting you move from “there is a gap” to “this output, this cause” without SSHing in.
  • Flush latency signals (flush_time_count, slow_flush_count) let you spot destination degradation as a rising average flush time before the rate deficit even appears.
  • In multi-worker deployments, per-worker visibility surfaces imbalances that aggregate metrics hide, since each worker has independent buffers and counters.