The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-multi-pipeline-monitoring ▌

Operations Guides

Logstash multi-pipeline monitoring: why aggregate stats hide a failed pipeline

You migrated to pipelines.yml to isolate workloads: one pipeline per team, per source, or per destination. Each pipeline got its own queue, its own workers, its own failure domain. Right call for fault isolation. But if your monitoring still polls the node once and looks at global event counts, you have given back the visibility the isolation bought you.

The concrete scenario: five pipelines, roughly equal traffic, one of them wedges. An output stalls, a config reload fails, a Kafka consumer group gets stuck rebalancing. That pipeline’s throughput goes to zero; the other four keep flowing. Aggregate node throughput drops by about 20 percent. If your alert is “throughput drops more than 50 percent” or “events per second below N”, nothing fires. The failed pipeline can sit dead for hours while node-level graphs show a normal dip.

This is the most common multi-pipeline monitoring mistake: the node is green, one pipeline is dead, and nobody knows until a downstream consumer asks where their data went. Below: why the math works against you, how to check per-pipeline state right now, and how to restructure collection and alerting so a single failed pipeline pages.

What this means

Multiple pipelines in pipelines.yml run inside one JVM but keep their state separate. Each pipeline has its own queue (memory or persistent), its own worker pool, its own event counters, its own flow metrics, and its own reload state. Persistent queues and dead letter queues are namespaced on disk by pipeline ID. A stall in one pipeline does not directly block the others. It also does not show up in their stats.

The aggregate view, the one most dashboards and quick curl /_node/stats checks surface, sums or averages across all pipelines. That view is fine for JVM-level concerns (heap, GC, process CPU, file descriptors), because those are genuinely shared. It is misleading for pipeline-level health, because pipeline failure is localized by design.

The failure pattern:

flowchart TD
  A[Five pipelines in one JVM] --> B[Pipeline 3 output stalls]
  B --> C[Pipeline 3 queue fills, throughput 0]
  C --> D[Aggregate node throughput drops ~20%]
  D --> E{Alert threshold: -50%?}
  E -->|No| F[No alert fires]
  F --> G[Pipeline 3 PQ fills to max_bytes]
  G --> H[Pipeline 3 inputs block, upstream backs up]
  A --> I[Per-pipeline stats: /_node/stats/pipelines]
  I --> J[output_throughput = 0 while input > 0]
  J --> K[Alert fires per pipeline]

The fix is not a new tool. It is a monitoring topology change: query each pipeline’s stats individually and alert on each pipeline individually.

Common causes

The localized failures aggregate stats most often hide:

CauseWhat it looks likeFirst thing to check
Stalled output on one pipelineOne pipeline’s events.out flatlines, its queue grows, other pipelines healthyPer-pipeline flow.output_throughput and queue.events
Failed config reloadPipeline runs old config or stops; aggregate throughput barely movesPer-pipeline reloads.failures and reloads.last_error
Input failure on one pipeline (Kafka rebalance, Beats disconnect, JDBC connection lost)Partial drop in aggregate events.in, proportional to that pipeline’s sharePer-pipeline flow.input_throughput and per-input plugin stats
PQ full on one pipelineThat pipeline’s inputs blocked; node totals look like a modest slowdownPer-pipeline queue.capacity.queue_size_in_bytes / max_queue_size_in_bytes
Pipeline missing entirelyJVM healthy, expected pipeline ID absent from statsCompare /_node/stats/pipelines response keys to pipelines.yml
GC death spiralAll pipelines degrade together, API slow, high GC timeJVM stats: jvm.gc.collectors.old.collection_time_in_millis (this one is node-wide)

The last row matters: not everything needs per-pipeline scoping. Heap, GC, file descriptors, and process CPU are shared across the JVM and belong at the node level. The mistake is applying node-level scoping to per-pipeline things.

Quick checks

Safe, read-only. Assumes the monitoring API is on the default port 9600 on localhost.

# List every pipeline the node is actually running
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty | grep -E '^    "[^"]+": \{$'

# Compare against what you configured
grep -E 'pipeline.id' /etc/logstash/pipelines.yml

# Per-pipeline throughput and queue state (the core check)
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty | \
  python3 -c "
import sys, json
ps = json.load(sys.stdin)['pipelines']
for pid, p in ps.items():
    flow = p.get('flow', {})
    inp = flow.get('input_throughput', {}).get('current', 'n/a')
    out = flow.get('output_throughput', {}).get('current', 'n/a')
    q = p.get('queue', {})
    print(f'{pid}: in={inp} out={out} queue_type={q.get(\"type\")} queue_events={q.get(\"events\")}')
"

# Check for failed reloads per pipeline
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty | \
  python3 -c "
import sys, json
ps = json.load(sys.stdin)['pipelines']
for pid, p in ps.items():
    r = p.get('reloads', {})
    print(f'{pid}: successes={r.get(\"successes\")} failures={r.get(\"failures\")} last_error={r.get(\"last_error\")}')
"

# In Logstash 8.x, structured pipeline health
curl -sS http://127.0.0.1:9600/_health_report?pretty

Two caveats. First, the flow metrics (input_throughput, output_throughput, queue_backpressure) are per-pipeline flow rates available in Logstash 8.x; on 7.x you compute rates from the cumulative events.in / events.out counters by sampling twice. Second, do not poll faster than every 10 seconds; the stats API shares the JVM with the pipelines and adds load on a stressed node.

(/_node/stats/pipelines/<pipeline_id> is a supported per-pipeline path and returns 404 for a missing pipeline id; the commands above use the full pipelines endpoint and filter client-side, which always works.)

How to diagnose it

When you suspect a pipeline is silently dead, or when auditing whether your monitoring would catch one:

  1. Enumerate expected pipelines. Read pipelines.yml and list the pipeline IDs that should exist. This is your ground truth. Note that if Logstash was started with -e or -f, pipelines.yml is ignored entirely and a warning is logged; confirm how the service actually starts.
  2. Enumerate actual pipelines. Query /_node/stats/pipelines and compare the response keys to the expected list. A missing pipeline ID means the pipeline failed to start or was terminated, and aggregate counters will not tell you that.
  3. Check per-pipeline output throughput. For each pipeline, look at flow.output_throughput.current (or delta events.out over a sampling interval). Zero output while flow.input_throughput.current is positive means the pipeline is alive but not delivering: the “living dead” state, scoped to one pipeline.
  4. Check per-pipeline queue state. For the suspect pipeline, look at queue.events and, for persistent queues, queue.capacity.queue_size_in_bytes / queue.capacity.max_queue_size_in_bytes. A growing queue on one pipeline while others are flat confirms a localized output or worker problem, not a node-wide one.
  5. Check reload state. Non-zero reloads.failures on the suspect pipeline means the running config may not match the deployed config. The pipeline can be “running” and still be wrong.
  6. Check per-plugin stats within the pipeline. If the pipeline is flowing but slowly, plugins.filters[].events.duration_in_millis and plugins.outputs[].events.duration_in_millis tell you which stage is the bottleneck.
  7. Correlate with node-level signals only for shared resources. If multiple pipelines degrade simultaneously, pivot to /_node/stats/jvm (heap, GC) and /_node/stats/process (CPU, FDs). Simultaneous degradation across pipelines points at the shared JVM, not at individual pipeline configs.

Metrics and signals to monitor

Every row here should be collected and alerted per pipeline, with the pipeline ID as a label or dimension.

SignalWhy it mattersWarning sign
pipelines.<id>.flow.output_throughputThe real health signal per pipelineZero or >50% below baseline while input is non-zero
pipelines.<id>.flow.input_throughputDetects per-pipeline input failureDrop to zero on one pipeline while its sources are active
pipelines.<id>.queue.eventsLocalized backpressureMonotonic growth over 15 minutes on one pipeline
pipelines.<id>.queue.capacity.queue_size_in_bytes / max_queue_size_in_bytesPQ runway per pipeline>80% sustained, page at >90% with positive growth
pipelines.<id>.flow.queue_backpressureInput throttling per pipelineSustained rise above that pipeline’s own baseline
pipelines.<id>.reloads.failuresConfig drift per pipelineAny non-zero value not previously observed
Pipeline presence in stats responseCatches fully dead pipelinesExpected ID missing from /_node/stats/pipelines
pipelines.<id>.dead_letter_queue.queue_size_in_bytesSilent per-pipeline data lossAny unexpected growth

Keep these node-level, not per-pipeline: jvm.mem.heap_used_percent (post-GC floor), jvm.gc.collectors.old.collection_time_in_millis, process.cpu.percent, process.open_file_descriptors. Pipelines share the JVM, so these are genuinely global.

Fixes

Restructure collection to be per-pipeline

Whatever collects Logstash stats must query /_node/stats/pipelines and explode the response into per-pipeline series, preserving the pipeline ID as a dimension. If your collector emits a single “logstash events out” metric with no pipeline label, that is the gap. Some aggregation layers have historically flattened this: Metricbeat’s logstash module, for example, has been reported to aggregate queue stats across pipelines in default dashboards, and pipeline-level data has been slow to appear in Stack Monitoring. Verify what your stack actually stores, not what the API returns.

Alert per pipeline, not per node

Rewrite throughput alerts so evaluation is per pipeline ID. “Output throughput zero while input non-zero, for 5 minutes, for any pipeline” catches the single dead pipeline. “Node throughput below 1000 eps” never will. Use baseline-relative thresholds per pipeline, because pipelines legitimately have different traffic volumes and different transformation ratios (clone, split, drop filters all change the in/out relationship).

Add a pipeline-presence check

Alert if the set of pipeline IDs in the stats response differs from the set in pipelines.yml. This catches the hardest case: a pipeline that failed to start and emits nothing at all. In Logstash 8.x, the /_health_report endpoint gives structured pipeline health and is preferable to inferring state from stats presence.

Isolate runaway pipelines operationally

If one pipeline’s PQ is filling because its destination is down, you can reduce blast radius by shedding that pipeline’s non-critical input or temporarily stopping it while the others continue. This is only possible because queues are isolated per pipeline. The same isolation means a PQ max_bytes sized for the whole node’s disk is wrong: each pipeline’s queue competes for the same underlying volume, so sum the configured maxima and compare against actual disk.

Account for shared resource contention

Default settings are tuned for a single pipeline: each pipeline defaults to one worker per CPU core. Five pipelines on an 8-core host means 40 workers competing for 8 cores. Per-pipeline CPU attribution is not exposed, so when you see node CPU saturation, check per-pipeline flow.worker_utilization to find which pipeline is burning it.

Prevention

  • Treat pipeline ID as a mandatory label. No Logstash throughput, queue, or error metric should exist in your monitoring system without it. Make this a review item for any new pipeline.
  • Dashboard per pipeline, aggregate only for JVM. Node-level dashboards show heap, GC, CPU, FDs. Pipeline-level dashboards show throughput, queue, backpressure, reloads, DLQ. Do not mix the scopes.
  • Alert on baseline deviation per pipeline. Absolute thresholds fail because pipelines differ and workloads drift. Compare against each pipeline’s own rolling baseline, and gate zero-throughput alerts on jvm.uptime_in_millis > 300000 to avoid cold-start noise.
  • Watch for version-specific stats bugs. The pipelines stats endpoint has had regressions: 7.3.x returned stats for only one pipeline (fixed in 7.3.2, elastic/logstash#11021), and 8.17.3 returned an empty pipelines object with X-Pack monitoring enabled (fixed in 8.17.4 and 8.18.0, elastic/logstash#17185). If per-pipeline graphs suddenly go empty after an upgrade while aggregate counters still move, suspect the endpoint, not the pipelines. Pin this as an upgrade checklist item.
  • Remember stats reset on reload. A hot config reload resets pipeline counters, producing artificial dips to zero. Do not confuse a reload dip with a pipeline death; check reloads.successes timestamps.
  • Include pipeline isolation in capacity reviews. Queue isolation means one pipeline’s max_bytes is not the whole story. Sum per-pipeline PQ maxima against the shared disk, and sum per-pipeline worker counts against cores.

How Netdata helps

  • Per-pipeline series, not just node totals: Netdata collects Logstash node stats and keeps pipeline-level dimensions, so a single stalled pipeline shows as its own series rather than being averaged into a node-wide line.
  • Throughput correlation per pipeline: input throughput, output throughput, and queue depth side by side per pipeline makes the “input flowing, output zero, queue growing” signature visible in one view instead of three curl calls.
  • Queue backpressure and worker concurrency: per-pipeline flow.* metrics distinguish “this pipeline is throttled because its queue is full” from “this pipeline is starved because the JVM is busy”, which determines whether you fix the pipeline or the node.
  • Reload failure visibility: reload successes and failures per pipeline surface configuration drift that otherwise stays invisible until behavior changes.
  • Node-level shared resources in the same view: JVM heap, GC time, CPU, and file descriptors alongside per-pipeline data, so simultaneous multi-pipeline degradation is immediately attributable to the shared JVM.
  • Anomaly detection per series: per-pipeline ML anomaly scoring catches the “one pipeline deviates from its own baseline” case that static node-wide thresholds structurally miss.