The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-monitoring-maturity-model ▌

Operations Guides

Logstash monitoring maturity model: from survival to expert

Most Logstash deployments are monitored at the wrong level. Teams alert on process liveness and heap percentage, then get paged by users asking where Wednesday’s logs went. The process was up, the API returned 200, and throughput looked fine the whole time. The gap is not tooling. It is which signals the team decided to watch.

This article defines a four-level maturity model for Logstash monitoring: Survival, Operational, Mature, and Expert. Each level names the specific signals to collect, why they matter, and what class of failure becomes visible that was invisible at the level below. Use it as an audit checklist against your current setup, and as a roadmap for what to add next. The levels are cumulative: every level assumes everything below it is already in place.

Two caveats before the lists. First, pipeline flow metrics (flow.input_throughput, flow.output_throughput, flow.queue_backpressure, flow.worker_concurrency, flow.worker_utilization) exist only in Logstash 8.6 and later (pipeline-level worker_utilization was added in 8.13); on older versions you must derive rates from cumulative counters. Second, thresholds here are baseline-relative on purpose. Absolute thresholds like “alert below 1000 events/sec” fire all night on quiet pipelines and never fire after traffic doubles.

flowchart TD
  L1["Level 1: Survival
Is it alive and delivering?"] L2["Level 2: Operational
Per-pipeline and resource health"] L3["Level 3: Mature
Per-plugin breakdown and runway"] L4["Level 4: Expert
SLOs and silent-loss detection"] L1 --> L2 --> L3 --> L4

Level 1: Survival

The minimum to avoid complete blind spots. With only these signals you can detect total failure, but nothing about degradation, data quality, or impending problems.

  • Pipeline output rate. pipelines.<name>.events.out or flow.output_throughput from GET /_node/stats/pipelines. Output rate is the real health signal. Zero output while inputs are active is the “living dead” scenario that process checks miss entirely.
  • Queue growth. queue.events_count and, for persistent queues, queue_size_in_bytes / max_queue_size_in_bytes. A monotonically growing queue for more than 15 minutes means the pipeline is falling behind.
  • Output errors and retries. Log patterns for retry, error, reject, timeout, 429, and 503 in logstash-plain.log, plus per-output event stats. Sustained nonzero retries are abnormal and usually precede queue growth.
  • JVM heap usage. jvm.mem.heap_used_percent from /_node/stats/jvm. At this level it is a coarse memory-pressure check, nothing more.
  • API liveness. curl -sS --connect-timeout 5 http://127.0.0.1:9600/ plus systemctl is-active logstash. A 200 only means the JVM and HTTP server are alive; pair it with output rate or you will miss the living-dead case.

What Level 1 cannot see: partial pipeline failures in multi-pipeline setups, GC death spirals in their early stage, parse failures, and anything about data correctness.

Level 2: Operational

What a professional team running Logstash in production needs. The theme of this level is moving from global aggregates to per-pipeline and per-resource views, because aggregates average away localized failures. One dead pipeline out of five drops aggregate throughput by 20%, below most alert thresholds.

  • Per-pipeline stats. Query /_node/stats/pipelines/<pipeline_id> individually instead of relying on node-level aggregates. Also verify expected pipeline IDs are present in the response; a missing pipeline after a failed reload is a partial outage.
  • Worker utilization. flow.worker_utilization. Sustained values above 90% during normal peaks mean the pipeline cannot absorb bursts without queue growth. High utilization with low host CPU points at blocking I/O or lock contention, not compute.
  • GC overhead. delta(jvm.gc.collectors.*.collection_time_in_millis) / delta(wall time). Above 10% is concerning, above 20% is severe. Track old-gen separately: rising old.collection_count is far more dangerous than rising young-gen counts.
  • Disk on PQ, DLQ, and log volumes. queue.data.free_space_in_bytes from the API plus df -h /var/lib/logstash /var/log/logstash. PQ max_bytes limits queue size but not disk usage if the queue shares a partition with logs or the DLQ. A full filesystem crashes the process even when the queue is within limits.
  • Grok failure counter. plugins.filters[].failures for grok filters, direct from the API. This is the cheapest data-quality signal available and almost nobody monitors it.
  • DLQ growth. dead_letter_queue.queue_size_in_bytes. Any unexpected growth on production data is a correctness failure: events are being permanently diverted instead of delivered. Note that DLQ is disabled by default; without it, permanently failing events are logged and silently lost.
  • File descriptor ratio. process.open_file_descriptors / process.max_file_descriptors from /_node/stats/process. Alert above 80%. File inputs hold one FD per tailed file, and FD leaks grow for weeks before they bite.
  • Input vs output comparison. flow.input_throughput vs flow.output_throughput. A sustained input/output ratio above 1.1 for more than 15 minutes with no drain periods means a backlog is building. Account for pipelines that legitimately transform event counts with clone, split, or drop.

What Level 2 cannot see: which specific plugin is the bottleneck, how much runway remains before the queue fills, and composite failures that require correlating several signals at once.

Level 3: Mature

This is where monitoring starts answering “how long do we have” and “exactly which component is at fault” instead of “is something wrong.”

  • Per-plugin breakdown. plugins.filters[].flow.worker_utilization, worker_millis_per_event, and events.duration_in_millis. Logstash problems are usually plugin-local. One filter consuming more than 80% of pipeline processing time is your culprit. Assign explicit id values to plugins in config so the stats map back to readable names.
  • Queue runway. flow.queue_persisted_growth_bytes gives the fill rate directly; positive means growing, negative means draining. Runway is (max_queue_size_in_bytes - queue_size_in_bytes) / growth_rate. Page when occupancy exceeds 90%, smoothed growth is positive over a 5-15 minute window, output rate is below input rate, and runway is under 30 minutes.
  • Queue backpressure. flow.queue_backpressure, the fraction of time input threads spend blocked pushing into the queue. Treat it as baseline-relative: the magnitude depends heavily on pipeline shape and cannot be compared across pipelines.
  • Processing-duration trend. delta(events.duration_in_millis) / delta(events.filtered) for per-event average. A sustained 2x rise over baseline without a change in event complexity means filters or enrichment got more expensive.
  • Composite pattern detection. Correlate signals into the known archetypes: backpressure wedge (low CPU, growing queue, output errors), grok hell (high CPU, high worker utilization, growing queue), GC death spiral (rising post-GC floor, GC overhead above 20%, throughput collapse). Single signals page too often; the combinations are the reliable triggers.
  • Reload state. pipelines.<name>.reloads.successes, .failures, and .last_error. Any new failure means the running config has diverged from the deployed config, invisibly.
  • Cardinality drift. events.in vs events.filtered vs events.out against the intended transformation ratio. Unexplained drift catches accidental drops, clones, and sincedb corruption re-reading files.
  • Hot threads on demand. GET /_node/hot_threads during incidents. Repeated snapshots separate filter burn from output waits. Not an alerting signal; forensic evidence.

Level 4: Expert

Deep signals for catching subtle or chronic issues, and for turning “is it up” into “is it meeting its commitment.”

  • Per-pipeline latency and freshness SLOs. End-to-end freshness is not a built-in metric. Derive it by comparing source-set @timestamp against destination indexing time. Throughput can look perfect while events arrive minutes late, and for alerting and security use cases stale data is as bad as missing data.
  • Post-GC floor. The heap level after garbage collection, tracked via jvm.mem.pools.old trend. A rising floor is the early-warning leak signal; the sawtooth peak is noise. This replaces the useless “heap > 80%” alert that fires on every normal GC peak.
  • Sincedb health. For file inputs, sincedb corruption causes re-reads from the beginning: a spike in events.in with duplicates downstream. The Node Stats API has no sincedb health fields; this requires custom file/state checks on the sincedb files.
  • Elasticsearch per-document bulk failures. Bulk requests can return HTTP 200 while individual documents fail mapping or type checks. Logstash counts the batch as out. Detecting this requires ES-side bulk rejection metrics or document-level output stats, not Logstash output counters; the Node Stats API exposes no per-document bulk failure counter.
  • Config-drift detection. Running config versus source-controlled config. Reload counters tell you a reload failed; drift detection tells you the deployed and running configs differ even when nothing recently failed. This is a custom external check (e.g. config hash comparison) unless you use Central Pipeline Management, which sources pipeline config from Elasticsearch.

Choosing your target level

LevelCatchesStill misses
SurvivalTotal failure, living-dead processDegradation, partial pipeline failure, correctness
OperationalPer-pipeline failure, resource exhaustion, DLQ and parse failuresPlugin-level root cause, time-to-full, composite patterns
MatureBottleneck attribution, queue runway, reload drift, GC vs CPU vs backpressure splitLatency SLO breaches, silent ES bulk loss, config drift
ExpertSilent data loss, leak early warning, freshness SLOs, driftBusiness-logic correctness (out of scope for metrics)

A reasonable target for most production teams is Level 2 everywhere, Level 3 for pipelines that feed alerting or compliance reporting, and Level 4 only where a silent-loss incident has already cost you something.

How Netdata helps

  • Netdata polls the Logstash monitoring API directly, so Survival and Operational signals (output throughput, queue occupancy, heap, GC, FD ratio, reload counters) are collected per second without hand-rolled curl scripts.
  • Per-pipeline and per-plugin charts make Level 2 and Level 3 breakdowns visible by default, which is where aggregate-only monitoring hides failures.
  • Correlating worker_utilization against host CPU on one dashboard is the fastest way to split the two most common patterns: compute bottleneck (high both) versus output blocking (high utilization, low CPU).
  • Persistent queue occupancy alongside disk free space on the same volume exposes the PQ-masked-outage pattern before the queue hits max_bytes.
  • Baseline-relative anomaly detection on throughput handles the workload-variation problem that breaks absolute thresholds, without you hand-tuning rolling averages per pipeline.