The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-retry-count-climbing ▌

Operations Guides

Fluentd retry_count climbing: the destination is rejecting or unreachable

You pulled /api/plugins.json from the monitor_agent and one of your output plugins shows a nonzero retry_count, and it keeps going up. At the same time, write_count has stopped incrementing and buffer_queue_length is growing. Fluentd itself is alive and inputs are still collecting.

This is the destination-unavailable failure pattern: the output plugin cannot deliver chunks, the retry engine has taken over, and Fluentd is now in exponential backoff against a destination that is rejecting connections, rejecting data, or simply gone. Every event that arrives from now on accumulates in the buffer. The clock you are racing is buffer capacity, not the retry count itself.

The dangerous part is that this state can persist for a long time before anything looks “down.” The process is up. Inputs are up. Retries are nominally a recovery mechanism. But with exponential backoff, each failure pushes the next attempt further into the future, and a pipeline that is “retrying” can be effectively stalled for tens of minutes at a time.

What this means

When a chunk flush fails, the output plugin does not drop the chunk. It rolls the chunk back into the queue and the retry engine schedules another attempt. With the default retry_type exponential_backoff, the wait between attempts grows with each failure. If retry_max_times is not set and retry_forever is false, retries are bounded by retry_timeout (default 72 hours), after which the chunk is discarded: data loss, with only a log line to mark it.

Two things in the monitor_agent response describe this state, and they are not interchangeable:

  • retry_count is a cumulative counter of retry error occurrences on that output plugin. It tells you failures have happened, and whether the rate of failures is ongoing. In current Fluentd it does not reset to zero on success; it clears on process restart. (Some older documentation and builds describe it as resetting on a successful flush, so verify against your version before you alert on the raw value.)
  • The retry object (retry.start, retry.steps, retry.next_time) describes the live retry cycle: when it started, how many attempts have been made in this cycle, and when the next attempt fires. This is the field that tells you how bad things are right now.

That distinction matters operationally. A retry_count of 40 with no retry object present means past failures that have since recovered. A retry_count of 6 with retry.next_time 25 minutes in the future means the pipeline is stalled right now and getting worse.

flowchart LR
  A[Destination rejects or unreachable] --> B[Chunk flush fails]
  B --> C[Chunk rolled back to queue]
  C --> D[Retry scheduled with growing backoff]
  D -->|fails again| B
  D -->|succeeds| E[Queue drains]
  C --> F[buffer_queue_length grows]
  F --> G[Buffer approaches total_limit_size]
  G --> H[overflow_action: drop, block, or exception]

Common causes

CauseWhat it looks likeFirst thing to check
Destination down or unreachableConnection refused, timeouts in Fluentd logs; write_count flatProbe the destination directly from the Fluentd host
Destination overloaded or throttlingHTTP 429, slow responses; slow_flush_count and flush_time_count rising before retries startedDestination-side health (Elasticsearch cluster state, S3/Kafka throttling)
Authentication or TLS failure401/403, handshake errors, expired certificate or rotated credentialsFluentd logs for auth/TLS error lines
Network partition or firewall changeSudden onset across all outputs to one destination; no application-level error, just timeoutsNetwork path: DNS resolution, security groups, firewall rules
Schema or index conflictDestination reachable but rejecting data (bulk requests rejected)Destination logs; Fluentd logs for rejection responses
Permanent misconfigurationRetries from the moment of a config change or deploy; never a single successRecent config diff, endpoint, credentials in config

Quick checks

All of these are read-only. Paths shown use the td-agent package layout; fluent-package uses /var/log/fluent/fluentd.log and /etc/fluent/fluentd.conf.

# 1. Per-output retry state and queue depth
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") | {id: .plugin_id, retries: .retry_count, queue: .buffer_queue_length, writes: .write_count, avail_pct: .buffer_available_buffer_space_ratios}'

# 2. The live retry object (start, steps, next_time)
# On v1.19.3+ add 'include_retry true' to the monitor_agent source config first:
# the ?with_retry=true query parameter stopped working in that release.
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") | {id: .plugin_id, retry: .retry}'

# 3. Recent output errors from Fluentd's own log
grep -E "failed to flush|retry|temporarily failed|could not connect|broken pipe" \
  /var/log/td-agent/td-agent.log | tail -20

# 4. Auth and TLS rejections
grep -iE "(tls|ssl|auth|401|403|unauthorized|forbidden|certificate)" \
  /var/log/td-agent/td-agent.log | tail -20

# 5. Is the destination reachable at all from this host?
curl -s -o /dev/null -w "%{http_code}\n" --max-time 5 https://your-destination-endpoint/

Two notes on the checks above. First, the retry object is not guaranteed to be in the response: from v1.19.3 onward, monitor_agent’s include_retry defaults to false, and the ?with_retry=true query parameter no longer overrides it (fluent/fluentd PR #5392). If the retry field is missing, set include_retry true in your monitor_agent source configuration rather than assuming there is no active retry. Second, in multi-worker mode each worker has its own monitor_agent port (24220, 24221, and so on): query every worker before concluding a retry is or is not happening.

How to diagnose it

  1. Confirm the pattern, not just the counter. Take two samples of retry_count, write_count, and buffer_queue_length a minute apart. The active-failure signature is: retry_count incrementing, write_count flat, buffer_queue_length growing. If write_count is incrementing and the queue is draining, you are looking at a recovered failure and a stale cumulative counter, not an incident.

  2. Read the retry object. retry.steps tells you how deep into the backoff curve this chunk is. retry.next_time tells you when Fluentd will try again. If next_time is more than a few minutes out, delivery latency is already severe and will keep growing even if you fix the destination this second, because the scheduled attempt is still far away.

  3. Get the actual error from the logs. The counters tell you that flushes fail; the log tells you why. Match the error to the cause table above: connection refused points at the destination or network, 401/403 at credentials, 429 at throttling, TLS errors at certificates. Auth errors often appear once at connection setup and are then masked by generic retry messages, so grep wider than the last few lines.

  4. Verify the destination independently. From the Fluentd host, probe the destination endpoint and check the destination’s own health signals (Elasticsearch cluster state, Kafka broker status, S3 error rates). Distinguish “destination is down” from “destination is up but rejecting this data,” because the fixes are completely different.

  5. Compute your runway. With the output stalled, the buffer fills at roughly the input rate. Check buffer_available_buffer_space_ratios and the growth of buffer_total_queued_size to estimate time to overflow. When the buffer hits total_limit_size, overflow_action fires: throw_exception (the default) raises BufferOverflowError, which for most inputs means the new event is dropped; drop_oldest_chunk discards old data; block stalls your inputs. None of these are good; all of them are worse than the retry state you are in now.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
retry_count (rate, not raw value)Confirms flush failures are ongoingDelta > 0 sustained across samples
retry.stepsDepth of the current backoff cycleLarge and growing; each step lengthens the wait
retry.next_timeWhen delivery might resumeMore than a few minutes in the future
write_countWhether any chunk is being deliveredFlat while input continues
rollback_countChunks returned to the queue after failed flushesSustained nonzero rate
buffer_queue_lengthBacklog depthGrowing trend, not just high absolute value
buffer_available_buffer_space_ratiosProximity to overflowBelow 20% and falling
flush_time_count / write_countAverage flush latencyRising before retries start; the earliest warning
slow_flush_countFlushes over slow_flush_log_threshold (default 20s)High ratio of slow to total flushes

Fixes

Destination down or unreachable

Fix the destination; there is no Fluentd-side fix for a dead endpoint. While it is down, your job is to buy buffer time: verify buffer_available_buffer_space_ratios is draining slowly enough to outlast the outage, and confirm your overflow_action is the failure mode you actually want when the buffer fills. If you have a <secondary> output configured, chunks that exhaust retries fall through to it instead of being discarded; watch write_secondary_count.

Authentication or TLS failure

Rotate the credential or renew the certificate in the Fluentd configuration and reload. Auth failures never self-heal: retries against a bad credential just generate backoff. After the fix, see the note below about resetting retry state, because the backoff schedule does not care that you fixed the cause.

Throttling or overload (HTTP 429, quota exceeded)

Reduce pressure on the destination or slow Fluentd’s delivery. Check slow_flush_count and average flush time: if flushes were degrading before retries started, the destination was already at capacity. Longer term, the destination needs more capacity or the pipeline needs less volume; retry tuning alone does not fix a throughput deficit.

Recovering from deep backoff

This is the counterintuitive part. Once you fix the destination, chunks sitting deep in exponential backoff may have retry.next_time far in the future, so recovery is slow even though the cause is gone. The standard reset is to restart Fluentd: retry counters and backoff state reset on restart, and file-backed buffers replay their chunks from disk. Expect a flush burst and a temporarily high queue while the backlog drains. That burst is normal; do not mistake it for a new problem. Do not restart before fixing the underlying cause, or you will re-enter backoff from step one.

If you run many Fluentd instances (a Kubernetes DaemonSet, for example), make sure retry_randomize true is in effect. Without jitter, every instance retries on the same backoff schedule, and the synchronized retry burst can knock over a destination that just recovered, starting the cycle again.

Retry tuning

Adjust retry_max_interval to cap how long the backoff can grow, and be deliberate about retry_timeout (default 72 hours) versus retry_forever. A shorter timeout bounds how long stale chunks occupy the buffer but guarantees discard when it expires. retry_forever never discards, but then the buffer is your only bound, and overflow becomes the data-loss path instead. Neither is universally right; pick based on whether losing old data or blocking new data is worse for your pipeline.

Prevention

  • Alert on the delta, not the value. retry_count is cumulative and does not reliably return to zero. Alert on delta(retry_count) > 0 sustained, or better, on the presence of a retry object with retry.next_time beyond a threshold.
  • Pair retry alerts with buffer state. A retry alert without buffer_queue_length and buffer_available_buffer_space_ratios context cannot tell you how urgent it is.
  • Watch flush latency as the leading indicator. Rising flush_time_count / write_count and slow_flush_count show destination degradation before the first retry fires.
  • Choose overflow_action explicitly per output. The default throw_exception drops new events at the buffer limit, and no counter reliably tracks those drops.
  • Use file-backed buffers in production. They survive restarts, which matters both for durability and for making the retry-state reset procedure safe.
  • Configure a <secondary> output for destinations where a multi-hour outage is plausible, so retry exhaustion routes to a fallback instead of the discard path.

How Netdata helps

  • Netdata collects the Fluentd monitor_agent fields as time series, so retry_count, write_count, rollback_count, and buffer_queue_length appear on one timeline instead of in manual curl samples. The stall signature (retries up, writes flat, queue growing) is visible at a glance.
  • Because Netdata computes rates from cumulative counters, you alert on delta(retry_count) rather than the raw value, which sidesteps the “counter never resets” trap.
  • Buffer saturation metrics (buffer_available_buffer_space_ratios, buffer_total_queued_size) next to retry metrics let you estimate time-to-overflow while the destination is down.
  • Flush latency derived from flush_time_count / write_count gives you the pre-retry degradation signal, so you see destination slowdowns before the first failure.
  • Per-plugin and per-worker breakdowns keep a single failing output, or a single struggling worker in multi-worker mode, from hiding inside aggregates.