The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-event-duplication ▌

Operations Guides

Logstash event duplication: in/out ratio drift, clones, and re-read sources

Logstash exposes three cumulative event counters at the pipeline level: events.in, events.out, and events.filtered. Their relationship reflects the pipeline’s intended transformation. A passthrough pipeline should see out approximately equal to in over any sustained window. A pipeline with drop {} filters should see out plus filtered approximately equal to in. Clone and split filters legitimately produce more output events than input events.

When this ratio drifts from what the configuration intends, the pipeline is duplicating events, losing events, or both. The worst case is silent duplication: the process is up, throughput looks healthy, and the only evidence is downstream indices containing more documents than the source produced. Clone, split, aggregate, and drop filters all change the ratio by design, so the diagnostic question is not whether the ratio differs from 1:1 but whether it matches what this specific pipeline’s configuration should produce.

What this means

Event duplication manifests as events.out exceeding what the pipeline’s filter logic should produce for a given events.in volume. The divergence may be constant (a filter always cloning or splitting) or sudden (a source re-read spike after a restart or sincedb corruption).

The counters work as follows. events.in counts events after codec processing, before filters. For persistent queues, this counts events written to the queue, not raw network receipt. events.filtered counts events removed from the pipeline, primarily by drop {} filters. events.out counts events successfully emitted by output plugins: sent to the destination, not necessarily acknowledged or persisted.

A pipeline without clone, split, drop, or conditional routing should see out approximately equal to in over any sustained window. Persistent divergence beyond normal in-flight buffering (queue events count plus batch_size times pipeline.workers) means events are being created, lost, or re-sent unexpectedly.

flowchart TD
    A["Unexplained in/out ratio drift"] --> B{"Sudden spike in events.in?"}
    B -->|Yes| C["Source re-ingestion"]
    B -->|No| D{"events.out consistently exceeds events.in?"}
    D -->|Yes| E["Filter fan-out or output resend"]
    D -->|No| F{"Output retries in logs?"}
    C --> G["Check sincedb state and restart history"]
    E --> H["Check clone/split per-plugin stats"]
    F --> I["Check output error and retry patterns"]
    G --> J["sincedb corruption, deletion, or inode change"]
    H --> K["Intentional clone/split or accidental?"]
    I --> L["PQ replay, destination rejects, or retry loop"]

Common causes

CauseWhat it looks likeFirst thing to check
Clone or split filterevents.out consistently exceeds events.in by a fixed multiplierPer-plugin events.out for clone and split filter instances
Sincedb corruption or deletionSudden spike in events.in after restart, downstream duplicatesSincedb file existence, timestamps, and inode references
PQ replay after crashBurst of events.out at startup exceeding recent events.inJVM uptime and recent unclean restart history
Output retry loopevents.out rising with retry/error patterns in logs; destination has duplicate documentsLog file for retry, error, reject, 429, 503 patterns
Multiple config files mergedDuplicate writes to same destination; events.out roughly 2x expectedCount of .conf files in the pipeline config directory
Multiline codec misassemblyInconsistent event counts; partial or merged events at destinationWhether multiline codec is used with Beats input

Quick checks

# Check pipeline-level event counters for all pipelines
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty \
  | python3 -c "
import sys,json
data = json.load(sys.stdin)
for pname, pdata in data.get('pipelines',{}).items():
    ev = pdata.get('events',{})
    print(f\"{pname}: in={ev.get('in',0)} out={ev.get('out',0)} filtered={ev.get('filtered',0)}\")
"
# Check per-plugin event counters for clone/split filters
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty \
  | python3 -c "
import sys,json
data = json.load(sys.stdin)
for pname, pdata in data.get('pipelines',{}).items():
    for f in pdata.get('plugins',{}).get('filters',[]):
        evts = f.get('events',{})
        name = f.get('name','?')
        if name in ('clone','split','aggregate'):
            print(f\"{pname}/{f.get('id','?')} ({name}): in={evts.get('in',0)} out={evts.get('out',0)}\")
"
# Check for output retry and error patterns in logs
grep -Ei '(retry|error|exception|failed|reject|unavailable|timeout|429|503)' \
  /var/log/logstash/logstash-plain.log | tail -n 200
# Check JVM uptime to correlate duplication with restarts
curl -sS http://127.0.0.1:9600/_node/stats/jvm?pretty | grep uptime_in_millis
# Check how many config files exist in the pipeline directory
find /etc/logstash/conf.d/ -name '*.conf' -type f
# Check sincedb file state for file inputs
ls -la $(grep -r 'sincedb_path' /etc/logstash/conf.d/ 2>/dev/null \
  | head -1 | awk -F'"' '{print $2}')/ 2>/dev/null || echo "sincedb_path not explicitly set"
# Check DLQ size to rule out DLQ-diverted events
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty \
  | grep -A 5 'dead_letter_queue'
# Sample two data points 10 seconds apart to compute live event rates
E1=$(curl -s http://127.0.0.1:9600/_node/stats/pipelines/main \
  | python3 -c "import sys,json; d=json.load(sys.stdin)['pipelines']['main']['events']; print(f\"{d['in']} {d['out']} {d['filtered']}\")")
sleep 10
E2=$(curl -s http://127.0.0.1:9600/_node/stats/pipelines/main \
  | python3 -c "import sys,json; d=json.load(sys.stdin)['pipelines']['main']['events']; print(f\"{d['in']} {d['out']} {d['filtered']}\")")
echo "T0: $E1"
echo "T1: $E2"
echo "Compute deltas to see live ratio"

How to diagnose it

  1. Establish the expected transformation ratio. Read the pipeline configuration. Count clone filters and the entries in each clones array. Identify split filters and what fields they operate on. Note any drop {} conditionals. This ratio is your baseline.

  2. Sample the live ratio. Poll events.in, events.out, and events.filtered twice with a known interval between samples. Compute the delta for each counter. The ratio of delta(out) to delta(in) over the window is the live transformation ratio. Compare it to the expected ratio from step 1.

  3. Classify the divergence. If delta(out) consistently exceeds delta(in) by a factor matching the clone or split configuration, the filter is working as configured and the duplication is intentional. If the factor does not match, or no clone/split filter exists, investigate further.

  4. Check for source re-ingestion if events.in spiked. A sudden increase in events.in that does not correspond to a known upstream volume change suggests the file input is re-reading files from the beginning. Check the sincedb file for the affected pipeline. Look for missing entries, stale inode references, or filenames with whitespace that break sincedb matching.

  5. Check restart history for PQ replay. If the divergence appeared immediately after a restart and the pipeline uses persistent queues, the burst may be PQ replay. On abnormal termination (OOM kill, SIGKILL, container force-kill), all persisted and unacknowledged in-flight events are replayed on restart. This is by-design at-least-once delivery. The burst should be temporary and self-correcting.

  6. Check output error and retry patterns. If the destination (typically Elasticsearch) rejects events and Logstash retries, the same events may be sent multiple times. Look for sustained retry patterns in the log file. Correlate with output plugin error stats from the stats API. An infinite retry loop on certain failure types, such as write blocks on Time Series Data Streams, can resend the same events repeatedly.

  7. Inspect per-plugin stats to localize fan-out. The per-plugin breakdown in /_node/stats/pipelines/<id> shows events.in and events.out for each filter and output plugin instance. A clone filter with three entries in its clones array should show out equal to in times four (original plus three clones). A split filter operating on a field with N elements should show proportional fan-out. If per-plugin numbers do not match expectations, the filter configuration needs review.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
events.in vs events.out ratioPrimary duplication and loss indicatorSustained deviation from the pipeline’s expected transformation ratio
Per-plugin events.out for clone/splitIsolates where fan-out occursPlugin out exceeds in by an unexpected factor
Sincedb file stateDetects file re-read risk for file inputsMissing entries, stale inode references, or unexpected file deletion
Output retry and error countIdentifies output-driven duplicationSustained nonzero retry pattern correlated with events.out growth
DLQ size and growthRules out DLQ diversion as cause of unexpected count changesGrowth when output rejections are active
JVM uptimeCorrelates duplication with restart events and PQ replayRecent unclean restart followed by output burst

Fixes

Clone or split filter fan-out

If the duplication comes from a clone or split filter, verify it is intentional. Check the clones array length in each clone filter. For split filters, identify the field being split and the typical number of elements.

Clone filter behavior changes with ECS compatibility. When ECS is disabled, clones receive a type field. When ECS is enabled (v1 or v8), clones receive a tags array entry instead. Downstream conditionals that check type may break after an ECS migration, causing clones to follow unexpected routing paths.

Note: the split filter pipeline-level counting difference tracks elastic/logstash issue #12746 and affects only the pipeline-level events.out counter; per-plugin stats report the real fan-out, which is why this guide reads per-plugin numbers.

If a split or clone filter was added accidentally (for example, by a config merge from multiple .conf files in the same pipeline directory), remove it or restructure the config directory so each pipeline has a single, explicit configuration source.

Sincedb corruption causing file re-reads

Sincedb corruption causes the file input to lose its read position and re-process files from the beginning. The result is a spike in events.in and downstream duplicates for every previously processed event.

Common triggers include deleting the sincedb file, inode changes from in-place file modification (the new inode is treated as a new file), and filenames containing whitespace that the file input does not escape correctly when writing sincedb entries.

To prevent recurrence, set an explicit sincedb_path to a persistent, backed-up location. Avoid modifying watched files in-place; rotate them instead. If sincedb corruption has already occurred, you cannot undo the duplicates already written to the destination. Deduplicate downstream or accept the duplicates and fix the sincedb path to prevent the next occurrence.

PQ replay after crash

PQ replay is by-design at-least-once delivery. When Logstash crashes or is force-killed, all unacknowledged in-flight events in the persistent queue are replayed on restart. This produces a burst of events.out that temporarily exceeds events.in.

To make replay idempotent at the destination, use the fingerprint filter to generate a consistent hash from one or more event fields, and set that hash as the document_id in the Elasticsearch output. On replay, the re-sent event overwrites the existing document with the same ID instead of creating a duplicate.

This approach requires a field or combination of fields that uniquely identifies each event. If events lack a natural unique key, the fingerprint approach cannot deduplicate them.

Output retry loops

When the Elasticsearch output retries failed bulk requests, events may be delivered multiple times. This is especially dangerous with write blocks on certain index types, where Logstash may enter an infinite retry loop resending the same events.

Check for sustained retry patterns in the log file. If the destination is rejecting events due to resource pressure (HTTP 429, bulk rejections, thread pool saturation), address the destination capacity first. See Logstash downstream backpressure cascade for the full diagnostic flow.

For idempotent writes, apply the same fingerprint-plus-document_id approach described for PQ replay. Retried events overwrite existing documents rather than creating duplicates.

Multiple config files merged

If multiple .conf files exist under one pipeline’s config directory, Logstash merges them. All filters and outputs from all files apply to all inputs. If two files each define an Elasticsearch output pointing to the same cluster, every event is written twice.

Consolidate the configuration into a single file per pipeline, or use pipelines.yml to assign separate config directories to separate pipelines. Verify the fix by checking that events.out returns to the expected ratio.

Multiline codec with Beats input

The multiline codec should not be used with the Beats input plugin or any input supporting multiple hosts. Doing so can mix streams from different hosts and corrupt event assembly. Configure multiline processing in Filebeat before it sends to Logstash.

If you see inconsistent event counts with partial or merged events at the destination, and the pipeline uses a multiline codec on a Beats input, move the multiline logic upstream to Filebeat.

Prevention

  • Document the intended transformation ratio for every pipeline. Store it alongside the configuration or in a runbook. Without this baseline, ratio drift is invisible until downstream users report duplicates.

  • Alert on ratio deviation, not absolute counts. Compare the live delta(out) / delta(in) ratio against the documented expected ratio. A pipeline that legitimately clones 1:4 should not alert at 4x, but it should alert if the ratio suddenly changes to 5x.

  • Protect sincedb files. Set an explicit sincedb_path on a persistent volume. Include sincedb health in your monitoring for file-input pipelines.

  • Use fingerprint plus document_id for idempotent Elasticsearch writes. This converts at-least-once delivery into effectively-once delivery at the destination, eliminating duplicates from PQ replay and output retries.

  • Audit config directories for accidental duplicate outputs. A single misplaced .conf file can silently double-index every event.

How Netdata helps

  • Per-second event counter collection makes ratio drift visible within seconds. Transient spikes such as PQ replay bursts or sincedb re-read floods show up immediately, not at the next polling interval.

  • Anomaly detection on event flow rates surfaces ratio changes without requiring you to pre-define the expected multiplier for every pipeline. The learned baseline flags deviations that static thresholds miss.

  • Correlation between JVM restarts and event counter resets helps distinguish PQ replay bursts (temporary, startup-correlated) from persistent duplication bugs (constant, config-correlated).

  • Per-plugin breakdown visibility narrows the search from “the pipeline duplicates” to “the clone filter with ID parse_syslog is fanning out 5x instead of 3x.”