The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-out-of-memory-java-heap-space ▌

Operations Guides

Logstash OutOfMemoryError: Java heap space and how to recover

java.lang.OutOfMemoryError: Java heap space in the Logstash log means the JVM could not satisfy an allocation request and GC could not reclaim enough heap to proceed. The process exits or gets kernel OOM-killed, and every pipeline it was running stops with it.

Before the hard crash, there is usually a warning period: the GC death spiral. Heap fills, garbage collection runs longer and more frequently, throughput collapses. This can last minutes or hours before the JVM finally fails to allocate. If you catch the spiral, you can intervene before data is lost. If you only catch the OOM, you are in recovery mode.

The default JVM heap in Logstash ships at 1GB (-Xms1g -Xmx1g in jvm.options), unchanged from 7.x through 9.x. This is too small for most production workloads.

What this means

OutOfMemoryError: Java heap space means the JVM’s garbage collector could not free enough heap to satisfy an allocation request. The allocation might be a batch of event objects, a large JSON document being parsed, a filter’s internal buffer, or any other heap consumer. When the JVM exhausts heap and GC cannot reclaim enough, it throws this error and terminates.

This is distinct from two related failure modes:

  • GC death spiral: The process stays alive but spends increasing time in garbage collection, leaving less time for event processing. Throughput approaches zero while the process appears healthy to a liveness check. This can persist for minutes or hours and often precedes a hard OOM.
  • Direct buffer memory OOM: The error reads java.lang.OutOfMemoryError: Cannot reserve N bytes of direct buffer memory instead of Java heap space. This happens when Netty-based inputs (Beats, TCP, HTTP) exhaust direct memory. The pipeline.buffer.type setting exists in logstash.yml (values direct/heap): on 8.x it defaulted to unset (Netty uses direct buffers), and since Logstash 9.0.0 the default is heap.

The critical operational question is whether you are seeing the sudden terminal crash, or the slow spiral that predicts it. Different symptoms, different response windows.

flowchart TD
    A[Heap fills with live objects] --> B[GC frequency and duration increase]
    B --> C[Throughput drops, events accumulate in queue]
    C --> D[Heap fills faster from backlog]
    D --> B
    B --> E[Post-GC floor rises, old-gen fills]
    E --> F{JVM can still allocate?}
    F -- No --> G["OutOfMemoryError: Java heap space"]
    F -- Barely --> H[Process alive, near-zero useful work]
    H --> G
    G --> I[Process exits or kernel OOM-kills]

Common causes

CauseWhat it looks likeFirst thing to check
Heap too small1GB default with production traffic; OOM under normal loadjvm.options for -Xms / -Xmx values
Oversized events or JSON field explosionSingle events consuming hundreds of MB during parse; stack trace in JSON parsing codeEvent source for unusually large payloads
In-flight events too largeOOM scales with pipeline.workers * pipeline.batch.sizeWorker count and batch size in config
In-memory queue at capacityMemory queue full, heap dominated by queued event objectsQueue type and events_count
Filter plugin memory leakPost-GC floor rises steadily over hours or days with no config changeOld-gen pool trend after GC

Quick checks

Run these read-only commands to assess the situation. If the process has already crashed, skip to checking logs and dmesg.

# Check if the process is still alive
pgrep -f org.logstash.Logstash

# Check API responsiveness (slow response indicates GC pressure)
time curl -sS --max-time 5 http://127.0.0.1:9600/_node/stats/jvm?pretty

# Check heap usage and memory pools
curl -sS http://127.0.0.1:9600/_node/stats/jvm | python3 -c "
import sys,json
j = json.load(sys.stdin)['jvm']
m = j['mem']
print(f\"Heap: {m['heap_used_in_bytes']//1048576}MB / {m['heap_max_in_bytes']//1048576}MB ({m['heap_used_percent']}%)\")
old = m['pools']['old']
print(f\"Old-gen: {old['used_in_bytes']//1048576}MB / {old['max_in_bytes']//1048576}MB\")
"

# Check GC time and frequency
curl -sS http://127.0.0.1:9600/_node/stats/jvm | python3 -c "
import sys,json
gc = json.load(sys.stdin)['jvm']['gc']['collectors']
for gen, s in gc.items():
    print(f\"{gen}: {s['collection_count']} collections, {s['collection_time_in_millis']}ms total\")
"

# Check queue depth (growing queue under GC pressure compounds the problem)
curl -sS http://127.0.0.1:9600/_node/stats/pipelines | python3 -c "
import sys,json
p = json.load(sys.stdin)['pipelines']
for name, d in p.items():
    q = d.get('queue', {})
    print(f\"{name}: {q.get('events_count',0)} events queued\")
"

# Check for kernel OOM kill in system logs (may require root)
dmesg -T | grep -i 'out of memory\|oom.*kill\|killed process'
journalctl -u logstash --since '1 hour ago' | grep -i 'OutOfMemoryError\|heap space'

# Check current JVM heap settings
grep -E '^-Xm' /etc/logstash/jvm.options

How to diagnose it

  1. Confirm whether the process is alive or dead. If dead, check dmesg or journalctl for OOM killer messages. If alive but unresponsive, suspect GC death spiral rather than a completed OOM.

  2. Distinguish hard OOM from GC death spiral. A hard OOM produces OutOfMemoryError in the log and the process exits. A GC death spiral shows high GC overhead (more than 20% of wall time), rising old-gen usage, and degraded throughput, but the process stays alive. The spiral may eventually trigger a hard OOM, or the kernel may OOM-kill the process when total RSS exceeds system limits.

  3. Compute GC overhead as a percentage of wall time. Take two samples of jvm.gc.collectors.old.collection_time_in_millis 60 seconds apart. Divide the delta by 60000 (wall time in ms). If old-gen GC alone consumes more than 10% of wall time, you are in significant memory pressure. Above 20%, throughput is severely impaired.

  4. Check the post-GC floor, not the peak. The meaningful heap signal is the level after garbage collection, not before. A rising post-GC floor (old-gen used_in_bytes after collections) indicates live objects accumulating. A sawtooth pattern where the floor stays stable is normal JVM behavior.

  5. Calculate the in-flight event count. Multiply pipeline.workers by pipeline.batch.size. On a 16-core machine with default settings (16 workers, batch size 125), that is 2000 events in flight. If events average 500KB after JSON expansion, the in-flight set alone needs approximately 1GB of heap. With the default 1GB heap, this is a guaranteed OOM.

  6. Look for oversized events. If the OOM stack trace shows jruby.RubyStringConverter or JSON parsing code, a single large event may be the trigger. The Jackson string length limit defaults to 200MB (property logstash.jackson.stream-read-constraints.max-string-length; the -D line sits commented in jvm.options because 200MB is the built-in default). A single event near that limit can consume most of the heap during parsing.

  7. Check for persistent queue overhead. Each pipeline with PQ enabled requires at least head and tail pages in native memory (default 64MB each). With 10 pipelines, PQ alone consumes approximately 1.28GB of native memory before any heap usage. This does not count toward heap, but it counts toward total process RSS and can trigger container OOM kills.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
jvm.mem.heap_used_percentOverall heap pressureSustained above 85% after GC
jvm.mem.pools.old.used_in_bytesLong-lived object accumulationRising post-GC floor over hours
jvm.gc.collectors.old.collection_time_in_millisGC overhead stealing processing timeMore than 10% of wall time (compute as rate)
flow.output_throughputWhether the pipeline is doing useful workDrops to zero while input continues
queue.events_countBacklog building under memory pressureGrowing while throughput drops
Process RSSTotal memory footprint vs system or container limitRSS approaching container limit

Fixes

Increase JVM heap size

The most common fix. The default 1GB heap is insufficient for production. Elastic’s guidance for typical ingestion workloads: no less than 4GB and no more than 8GB.

In /etc/logstash/jvm.options, set:

-Xms4g
-Xmx4g

Always set -Xms equal to -Xmx. If they differ, the JVM spends cycles resizing heap at runtime, which adds latency and can trigger unnecessary GC cycles. The equal setting also means heap_max_in_bytes is fixed from startup, making heap percentage metrics stable from the first second.

Do not set heap above 50% of available system memory. The JVM needs room for off-heap allocations (Netty buffers, PQ memory-mapped files, JVM internal structures). In containers, set heap relative to the container memory limit, not the host. The bundled JDK in Logstash 8.x and 9.x has container support enabled (-XX:+UseContainerSupport), so it detects cgroup limits.

After changing jvm.options, restart Logstash. There is no hot-reload for JVM settings.

Reduce in-flight event count

If increasing heap is not possible due to memory constraints, or the OOM persists after increasing, reduce the number of events held in memory simultaneously.

The in-flight count is pipeline.workers * pipeline.batch.size. Lower pipeline.batch.size in logstash.yml or pipelines.yml:

pipeline.batch.size: 64

This trades throughput for memory safety. Each worker holds fewer events, so the total in-flight set is smaller. Test the throughput impact, as smaller batches mean more overhead per event.

Do not reduce pipeline.workers below what your CPU can handle. Fewer workers means slower processing, which grows the queue, which can increase memory pressure rather than decrease it.

Address oversized events

If the OOM stack trace points at JSON parsing or string conversion, a single large event is the likely trigger. Sources of oversized events:

  • Log producers embedding full stack traces, JSON payloads, or binary blobs in a single log line
  • Multiline codec accumulating many lines into one event without a size limit
  • JSON with deeply nested or exploded field structures

Mitigations:

  • Add size guards at the source (log producer configuration)
  • Use the truncate filter to cap field lengths before they consume heap
  • Split large events earlier in the pipeline
  • Review the Jackson string length limit setting and lower it if your events should never approach 200MB

Enable heap dump on OOM

-XX:+HeapDumpOnOutOfMemoryError writes a heap dump file when the OOM occurs. This file is your best tool for post-mortem analysis. It is included in the default jvm.options file in recent Logstash versions. Verify it is present:

grep HeapDumpOnOutOfMemoryError /etc/logstash/jvm.options

If not present, add it. The dump is written to the JVM working directory by default. A heap dump for a 4GB heap produces a file of approximately 4GB. Ensure the working directory has enough disk space, and clean up old dumps.

Analyze the dump with Eclipse MAT, VisualVM, or jhat to identify which objects dominate the heap.

Distinguish from direct buffer OOM

If the error message is Cannot reserve N bytes of direct buffer memory rather than Java heap space, increasing -Xmx will not fix it. Direct memory has its own limit. By default, the JVM sets MaxDirectMemorySize equal to -Xmx, so a 1GB heap also implies 1GB of direct memory.

In Logstash 8.x, Beats/TCP/HTTP inputs allocate Netty buffers from direct memory (official jvm-settings docs name Beats, TCP, and HTTP inputs as direct-memory users). Options:

  • Set pipeline.buffer.type: heap in logstash.yml to move these allocations onto the Java heap (this is the default since Logstash 9.0.0)
  • Increase direct memory explicitly with -XX:MaxDirectMemorySize in jvm.options

Even with pipeline.buffer.type: heap, some plugins may still use direct memory. If direct buffer OOM persists, check plugin-specific buffer behavior.

Prevention

  • Size heap for production, not defaults. 1GB is a development default. Size to your workload, keep -Xms equal to -Xmx, and never exceed 50% of system or container memory.
  • Monitor the post-GC floor, not the peak. A heap alert that fires on heap_used_percent > 80% will trigger during every normal GC cycle and get silenced. Alert on the post-GC old-gen level rising over time instead.
  • Calculate in-flight against heap. Know your pipeline.workers * pipeline.batch.size product. Estimate average event size. Ensure the product fits comfortably in heap with room for GC overhead.
  • Track GC overhead as a rate. Compute delta(collection_time_in_millis) / delta(wall_time) for old-gen collections. Alert above 10% sustained. This catches the GC death spiral before it becomes a hard OOM.
  • Enable and verify heap dump. Confirm -XX:+HeapDumpOnOutOfMemoryError is in jvm.options. Ensure disk space for the dump file.
  • Watch total RSS in containers. Container OOM kills happen when RSS exceeds the limit, not when heap exceeds max. PQ pages, direct buffers, and JVM overhead all add to RSS.
  • Gate alerts on uptime. Cold start behavior (JIT compilation, class loading, PQ replay) can look like memory pressure. Require jvm.uptime_in_millis > 300000 before firing heap or GC alerts.

How Netdata helps

  • Per-second JVM heap metrics let you see the sawtooth pattern and, more importantly, the post-GC floor trend that predicts OOM before it happens. Standard polling intervals (15 to 60 seconds) often miss the GC floor between collections.
  • GC collection time and count are collected at per-second resolution, making it straightforward to compute GC overhead as a percentage of wall time and alert on the death spiral pattern.
  • Correlation between heap, GC, throughput, and queue depth in a single view lets you distinguish a GC death spiral (high GC, dropping throughput, growing queue) from an output bottleneck (low GC, dropping throughput, growing queue) or a CPU-bound filter (low GC, high CPU, growing queue).
  • ML-based anomaly detection on heap usage and old-gen trends can surface the slow post-GC floor climb that a fixed threshold would miss.
  • Container-aware memory metrics show RSS alongside heap, which is critical for diagnosing container OOM kills where heap is within limits but total process memory exceeds the cgroup cap.