The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-monitoring-checklist ▌

Operations Guides

Logstash monitoring checklist: the signals every production pipeline needs

Most Logstash monitoring setups answer the wrong question. They answer “is the process running?” when the question that matters is “are events leaving the pipeline?” A Logstash JVM can be alive, healthy by systemd’s standards, and returning 200 from its monitoring API while the queue is full, workers are blocked on a dead Elasticsearch, and zero events have been delivered for an hour. Process liveness is necessary. It is nowhere near sufficient.

This checklist organizes the signals a production Logstash deployment needs into four maturity levels, from the floor that avoids total blindness to the deep signals that catch silent correctness failures. Use it to audit an existing setup or to build one that will not betray you at 3 a.m.

Almost everything here comes from the Logstash monitoring API, which binds to 127.0.0.1:9600 by default (configurable via api.http.port). The two workhorse endpoints are /_node/stats/pipelines and /_node/stats/jvm. In containerized deployments the API still binds to loopback by default, so you must publish the port and may need to bind to a non-loopback address to scrape it. Do not poll faster than every 10 seconds; the stats API competes with the pipeline for JVM resources on loaded instances.

How to use this checklist

Work down the levels in order. Level 1 is not a recommendation; it is the floor. If you are missing a Level 1 signal, you have a blind spot that has already caused or will cause an undetected outage. Levels 2 and 3 are where most teams should aim. Level 4 signals pay off after you have been burned by the failure modes they detect.

One rule applies across every level: thresholds should be baseline-relative, not absolute. “Alert below 1000 events/sec” fires all night and never during a real incident after traffic doubles. Percent deviation from a rolling average for the same time window is more work and dramatically more useful. Also gate time-sensitive alerts on JVM uptime (jvm.uptime_in_millis > 300000); cold starts produce false positives in almost every signal for the first 30 to 120 seconds.

flowchart TD
  L1["Level 1: Survival - output rate, queue, output errors, heap, API liveness"]
  L2["Level 2: Operational - per-pipeline stats, workers, GC, DLQ, FDs, parse failures"]
  L3["Level 3: Mature - per-plugin breakdown, runway, backpressure, reloads, drift"]
  L4["Level 4: Expert - post-GC floor, freshness, recovery path, bulk failure detail"]
  L1 --> L2 --> L3 --> L4

Level 1: survival

The minimum to avoid complete blind spots.

  • Pipeline output rate. The rate of events successfully emitted to outputs: pipelines.<name>.events.out as a counter, or flow.output_throughput as a pre-computed rate. This is the primary functional health signal. Zero output with non-zero input (gate on flow.input_throughput > 0 to avoid idle-server false positives) means the pipeline is alive but delivering nothing. Alert when output drops more than 50% below the rolling baseline, or to zero sustained, while input is active.
  • Queue growth and occupancy. pipelines.<name>.queue.events_count and, for persistent queues, queue.queue_size_in_bytes against queue.max_queue_size_in_bytes. A steadily growing queue means events arrive faster than they leave. For memory queues, any monotonic increase over 15 minutes is concerning. For PQ, alert above 80% occupancy sustained.
  • Output error and retry activity. Grep the log for retry, error, reject, timeout, 429, and 503 patterns, and check per-output plugin stats in /_node/stats/pipelines. Sustained non-zero retry activity is abnormal; brief retries during downstream failover self-heal. Severity escalates when retries coincide with queue growth.
  • JVM heap usage. jvm.mem.heap_used_percent from /_node/stats/jvm. Heap pressure triggers GC, which pauses all processing. A steady 75% with an efficient sawtooth is normal; alert fatigue from naive “>80%” rules is why teams miss real crises. See Level 4 for the correct version of this signal.
  • Process and API reachability. curl -sS --connect-timeout 5 http://127.0.0.1:9600/ plus systemctl status logstash. The first check in any incident. A 200 only proves the JVM and HTTP server are alive; always pair it with output throughput. During severe GC pauses the API can time out while the process recovers, so treat standalone unreachability as a ticket and let the output-rate signal carry page weight.

Level 2: operational

Everything in Level 1, plus what a professional team needs to run Logstash without surprises.

  • Per-pipeline stats, not aggregates. In multi-pipeline deployments (pipelines.yml), one dead pipeline out of five drops aggregate throughput by 20%, below most thresholds. Query /_node/stats/pipelines/<pipeline_id> individually and verify each expected pipeline ID is present and running. On recent 8.x, the /_health_report endpoint gives structured pipeline health states; prefer it over inferring state from raw stats.
  • Input rate and input-vs-output ratio. events.in / flow.input_throughput versus the output equivalents. Input exceeding output by more than about 10% for over 15 minutes with no drain periods means a backlog is building. Know the intended transformation ratio per pipeline first: clone and split filters legitimately produce more output than input, drop filters fewer.
  • Worker utilization. flow.worker_utilization (flow metrics are available since Logstash 8.5, with per-plugin flow metrics since 8.6 and pipeline-level worker_utilization since 8.13). Sustained above 90% during normal peaks means no headroom for bursts. Interpret with CPU: high utilization plus high CPU is a compute bottleneck; high utilization plus low CPU is workers blocked on output I/O.
  • GC overhead. jvm.gc.collectors.old.collection_time_in_millis and collection_count, computed as a rate: delta(collection_time) / delta(wall time). Over 10% of wall clock is concerning, over 20% severe. Track old-gen separately; rising old-gen frequency is far more dangerous than young-gen activity.
  • Disk space on PQ, DLQ, and log volumes. queue.data.free_space_in_bytes reflects filesystem free space, not just PQ allocation. PQ max_bytes does not protect you if the queue shares a partition with logs or the DLQ. A full disk crashes the process even when the queue is within its configured limits.
  • Parse failure indicators. The grok filter’s plugins.filters[].failures counter is directly available in per-plugin stats and almost never monitored. Track its rate, not the absolute value. A rising failure rate with normal throughput is a silent correctness disaster: events reach the destination with wrong or missing fields. Cross-check downstream for _grokparsefailure and _jsonparsefailure tags, since aggregate rates can hide 100% failure on one small critical stream.
  • DLQ growth. pipelines.<name>.dead_letter_queue.queue_size_in_bytes. Any unexpected growth on production data is a ticket: events are being permanently rejected and diverted. Two caveats. First, the DLQ is disabled by default (dead_letter_queue.enable: false), and without it, permanently failed events are logged and silently lost. Second, teams that enable it often never monitor or replay it, which turns the safety net into silent data loss with extra disk usage.
  • File descriptor pressure. process.open_file_descriptors / process.max_file_descriptors from /_node/stats/process. Alert above 80% sustained. File inputs hold one FD per tailed file, so wildcards matching thousands of files can exhaust the limit; connection leaks show up as slow FD growth over weeks.

Level 3: mature

Everything above, plus the signals that surface degradation before it becomes an outage.

  • Per-plugin performance breakdown. plugins.filters[].flow.worker_utilization, plugins.filters[].events.duration_in_millis, and the output equivalents. Logstash problems are usually plugin-local; aggregate metrics hide the one bad grok pattern consuming 80% of pipeline time. Assign explicit id values to filters in config so stats map back to config without guesswork.
  • Queue fill rate and runway. flow.queue_persisted_growth_bytes gives you the PQ fill rate directly. Runway is (max_queue_size_in_bytes - queue_size_in_bytes) / current_fill_rate. This converts “PQ is growing” into “inputs block in 40 minutes.” The growth signal moves in chunks as pages are allocated and freed, not smoothly.
  • Queue backpressure. flow.queue_backpressure, the fraction of time input threads spend blocked pushing into the queue. This is the earliest backpressure signal, visible before the queue visibly fills. Alert on a sustained rise above the pipeline’s own baseline.
  • Event processing duration trend. Compute per-event average as delta(events.duration_in_millis) / delta(events.filtered). A sustained 2x rise without a workload change points at filter cost, enrichment latency, or pathological regex backtracking.
  • Configuration reload state. reloads.successes, reloads.failures, reloads.last_error. Any failure is a ticket: the old config keeps running, which is safe but creates invisible drift between deployed and running configuration. Teams routinely believe a fix shipped when it did not.
  • Composite failure pattern detection. The three big cascades have distinct signatures. Backpressure: output errors rise, queue grows, CPU stays moderate. Compute bottleneck (“grok hell”): CPU pegged, worker utilization pinned, no output errors. GC death spiral: post-GC floor rising, old-gen GC frequent, API flaky, throughput wobbling while the process looks alive. Alerting on the combination is far safer than any single leg.
  • Incident-time hot threads. GET /_node/hot_threads. Not an alert signal; take repeated snapshots during incidents to separate filter burn from output waits. Store snapshots from peak load for forensics.

Level 4: expert

Deep signals for catching subtle or chronic issues. Adopt these after the corresponding incident has cost you once.

  • Post-GC floor monitoring. The heap signal that actually works: the heap level after garbage collection, not the peak. A rising floor plus old-gen pool above 85% of max plus GC overhead above 20%, sustained with visible throughput impact, is the GC death spiral composite worth paging on. The raw heap_used_percent > 80% rule fires on every normal sawtooth peak and gets silenced before the real event.
  • End-to-end event freshness. Logstash exposes processing duration, not freshness. Derive it by comparing event @timestamp against destination indexing time. Throughput can stay flat while latency grows, and for alerting and security use cases stale data is as bad as missing data.
  • Elasticsearch per-document bulk failure detection. Logstash counts a bulk request as “out” when ES returns 200, even if individual documents inside it were rejected for mapping conflicts. The event counter looks healthy while data is lost. Catch this with documents.non_retryable_failures in the Elasticsearch output stats and ES-side rejection metrics.
  • Recovery-path monitoring. Watch queue drain rate after a downstream outage, not just failure detection. PQ drain can take far longer than fill, and high I/O plus reduced throughput during drain is expected, not a new incident.
  • PQ runway paging. The full page condition: PQ occupancy above 90%, smoothed positive growth over 5 to 15 minutes, output below input, runway under 30 minutes, active input, JVM uptime over 600 seconds. All legs together mean inputs block imminently with no self-recovery in sight.
  • Configuration drift detection. Compare running config against source-controlled config. File mtimes under /etc/logstash outside deploy windows plus reload log entries catch both accidents and unauthorized changes.
  • Sincedb health for file inputs. Sincedb corruption causes Logstash to re-read files from the beginning, producing a spike in events.in and duplicates downstream that look like a traffic surge.

What this checklist deliberately excludes

Three traps show up in almost every Logstash monitoring setup. Absolute throughput thresholds instead of baseline-relative ones. Heap percentage alerts instead of post-GC floor trends. And process liveness treated as health. If your current setup is built on those, fixing them is worth more than adding any new signal.

How Netdata helps

  • Netdata’s Logstash collector scrapes the monitoring API on port 9600 and charts pipeline event rates, queue depth, JVM heap, GC, and process stats per second, so output-rate drops and queue growth are visible without hand-rolled polling scripts.
  • Per-pipeline charts make multi-pipeline deployments readable: one dead pipeline stands out instead of being averaged away in a global aggregate.
  • JVM heap and GC charts side by side make the death-spiral signature (rising floor, rising old-gen time, falling throughput) a visual correlation rather than a log-diving exercise.
  • File descriptor usage against the configured limit is charted continuously, which catches slow FD leaks over weeks before the “too many open files” cliff.
  • Anomaly detection on throughput metrics handles the baseline-relative problem: deviations from the learned pattern for that time of day stand out without hand-tuned thresholds per pipeline.