The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-write-secondary ▌

Operations Guides

Fluentd write_secondary_count: the primary output has failed to its backup

write_secondary_count is a per-output counter in Fluentd’s monitor_agent API (exposed there since v1.19.0). A nonzero value means the primary output plugin exhausted its retries for at least one chunk, and Fluentd wrote that chunk to the configured <secondary> backup destination instead. The primary pipeline for that output is broken.

This is a ticket-level signal. Data is not lost yet (that is the point of the secondary), but it is no longer flowing where downstream systems expect it. Dashboards, SIEM rules, and alerts that read from the primary destination are now working from a gap.

Two facts make this counter easy to misread. First, it only exists when a <secondary> section is configured on the output. Without one, exhausted retries discard the chunk outright, with only a log line as evidence. Second, it is a cumulative in-memory counter: it resets on Fluentd restart, and a nonzero value from an incident last week looks identical to one from an incident happening right now. You need the delta, not the value.

What this means

Fluentd’s output path is: staged chunk, queued chunk, flush attempt, and on failure a retry cycle with backoff. Retries are bounded by retry_timeout (default 72h) or retry_max_times if set. When a chunk crosses the secondary threshold, Fluentd hands the chunk to the <secondary> plugin and increments write_secondary_count.

flowchart LR
  A[queued chunk] --> B[flush attempt]
  B -->|success| C[purged]
  B -->|failure| D[retry with backoff]
  D -->|recovers| B
  D -->|retry threshold crossed| E[secondary output]
  E --> F[write_secondary_count +1]
  D -->|no secondary configured| G[chunk discarded: data loss]

The trigger point matters. Fluentd switches a chunk to the secondary when the elapsed retry time exceeds retry_secondary_threshold (a ratio of retry_timeout, default 0.8). With the default 72h retry_timeout, a chunk only falls to the secondary after roughly 57 hours of continuous failure. If your write_secondary_count just moved, the primary has likely been down for a long time, or you have lowered retry_timeout or the threshold deliberately.

Common causes

CauseWhat it looks likeFirst thing to check
Primary destination permanently downretry_count climbing, write_count flat for hours, connection refused or timeout in Fluentd logsCurl the destination endpoint from the Fluentd host
Authentication or TLS failureRetries with 401/403 or handshake errors; starts suddenly after a credential or cert changegrep -iE "(tls|ssl|auth|401|403)" in the Fluentd log
Destination rejecting data (schema, quota)Flushes connect but fail; error logs show 4xx responses or per-document rejectionsRecent error lines mentioning the output plugin
Retry configuration too aggressiveSecondary engages quickly; retry_timeout or retry_secondary_threshold set lowOutput config block for retry parameters
Old incident, counter never resetCounter nonzero but retry_count flat, write_count incrementing normally, buffer healthyCompare counter against last Fluentd restart time

Quick checks

# Read the counter per output plugin
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.plugin_category=="output") | {id: .plugin_id, secondary: .write_secondary_count, retries: .retry_count, writes: .write_count, queue: .buffer_queue_length}'

A nonzero secondary with flat writes and growing queue is an active incident. A nonzero secondary with writes incrementing and an empty queue is residue from a past failure.

# Check current retry state, not just the counter
curl -s "http://localhost:24220/api/plugins.json?with_retry=true" | \
  jq '.plugins[] | select(.plugin_category=="output") | {id: .plugin_id, retry: .retry}'

If retry.steps is large and retry.next_time is far in the future, backoff has pushed the next attempt minutes or hours out. The pipeline is effectively stalled even though it looks like it is retrying.

# Find the failure reason in Fluentd's own log (adjust path for your package)
grep -E "failed to flush|retry|secondary|could not connect" \
  /var/log/td-agent/td-agent.log | tail -30

Look for the log line where chunks were handed to the secondary; it sits right after the last retry failure for each chunk and usually names the plugin and the underlying exception.

# Confirm whether a secondary is configured and where it writes
grep -A5 "<secondary>" /etc/td-agent/td-agent.conf
# Verify the destination independently
curl -sS -o /dev/null -w "%{http_code}\n" --max-time 5 https://your-destination-endpoint/

How to diagnose it

  1. Establish timing. Note the current write_secondary_count, then re-read it after 60 seconds. An incrementing counter means chunks are falling through now. A static counter means the failover already happened (or happened long ago) and the question is whether the primary has recovered.
  2. Check retry state. Pull the retry object for the affected output. Large retry.steps and a distant retry.next_time tell you recovery will be slow even after the destination is fixed.
  3. Read the error. The counter says the primary failed; the Fluentd log says why. Connection refused, auth errors, and data rejection each have different fixes.
  4. Test the destination directly. From the Fluentd host, verify DNS, TCP connectivity, TLS, and auth independently. This separates “destination is down” from “Fluentd’s view of the destination is broken” (expired cert in config, stale credentials).
  5. Locate the secondary data. If the secondary is secondary_file, chunks are being written to the configured directory. Confirm files are actually appearing and growing; a secondary that is itself failing (disk full, permission error) is the worst case, because the fallback is silently not a fallback. Note that write_secondary_count increments the moment a chunk is handed to the secondary, before the secondary write completes, so a failing secondary still increments it; watch the secondary’s own retries and error log for that.
  6. Check buffer headroom. While the primary is down, new events keep arriving. Watch buffer_available_buffer_space_ratios and buffer_queue_length so the ongoing outage does not turn into an overflow on top of the failover.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
write_secondary_count (delta)Chunks falling through to backup right nowAny increment
retry_count and retry objectPrimary is failing; backoff state predicts recovery speedSteps climbing, next_time far out
write_count (delta)Successful primary deliveriesFlat while input continues
rollback_countChunks put back in queue after failed flushesSustained increments
buffer_queue_lengthBacklog depth while the primary is downSustained growth
buffer_available_buffer_space_ratiosTime until overflow during the outageBelow 20% and shrinking
buffer_oldest_timekeyAge of the oldest undelivered dataHours or days behind now

Fixes

Destination down or unreachable

Fix the destination, then be patient or intervene. With exponential backoff, the next retry may be far away even after the destination recovers. Restarting Fluentd resets retry state (and the counter), at the cost of a brief pipeline pause and, with file buffers, a replay burst. That is usually acceptable to drain a long-stalled output quickly.

Authentication or TLS failure

Rotate the credential or certificate into the Fluentd config and reload. A single auth error line may be followed by generic retry messages, so grep broadly. After the fix, watch for write_count resuming and retry clearing.

Destination rejecting data

Schema conflicts and quota rejections do not heal themselves. Stop the flow of offending records (filter or re-route), fix the mapping or quota at the destination, then let retries drain. Do not just raise retry limits: the chunks will keep failing and eventually all land in the secondary.

Re-ingesting secondary data

Chunks in the secondary are not automatically replayed. For secondary_file, the standard recovery path is to read those files back with a separate input (for example in_tail with the matching parser) after the primary is healthy, then remove the files. Plan this before you need it: verify the secondary file format round-trips through your parser.

Retry tuning went wrong

If the secondary engaged too fast, revisit retry_timeout and retry_secondary_threshold. A low threshold trades “fast failover” for “failover during every transient blip,” which scatters data across two destinations constantly.

Prevention

  • Always configure a <secondary> for outputs where data loss is unacceptable. Without it, exhausted retries discard chunks with only a log line. secondary_file ships with Fluentd core; note it only works inside <secondary>, and only buffered outputs support secondary at all.
  • Size the secondary destination. A local directory on the same disk as the buffer is common, but it shares fate with buffer disk pressure. Know its capacity and monitor its growth.
  • Alert on the delta of write_secondary_count, not the value. The counter is cumulative and resets on restart. Any increment should page a human during business hours at minimum.
  • Set retry parameters deliberately. Decide how long the primary may fail before failover, and set retry_timeout and retry_secondary_threshold to match. The defaults imply about 57 hours before the secondary engages.
  • Monitor the leading signals. retry_count, flat write_count, and growing buffer_queue_length all fire long before the secondary threshold is crossed. See the monitoring checklist for the full set.

How Netdata helps

  • Netdata collects Fluentd monitor_agent metrics per output plugin, so write_secondary_count appears alongside retry_count, write_count, and buffer gauges on one timeline instead of separate curl snapshots.
  • The delta view matters here: a rate chart of write_secondary_count distinguishes an active failover (rising slope) from residue of an old incident (flat line at nonzero).
  • Correlating flat write_count, climbing retry_count, and growing buffer_queue_length in one dashboard shows the full arc: primary failing, backlog building, secondary engaging.
  • Anomaly detection on retry and buffer metrics surfaces the primary failure hours before the secondary threshold is crossed, which is when you want to know.
  • Per-plugin breakdowns keep multi-output configurations honest: one output can be failing over to its secondary while others stay healthy, and aggregated metrics would hide that.