The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-jvm-heap-sizing ▌

Operations Guides

Logstash JVM heap sizing: -Xms/-Xmx, the 1GB default, and why they should match

Logstash ships with a 1GB JVM heap (-Xms1g -Xmx1g in config/jvm.options), unchanged through 7.x, 8.x, and 9.x. That default works for trivial pipelines. It is catastrophically small for production and is the single most common cause of GC death spirals, throughput collapse, and OOM kills.

The fix: increase the heap, set minimum and maximum to the same value, and account for off-heap memory. Set the heap too large and you starve the OS and risk longer full-GC pauses. Set it without understanding off-heap allocation and you get OOM kills with a heap that looks comfortably under capacity. Ignore container cgroup limits and the JVM sizes itself against host RAM, not the container limit.

The 1GB default and what it costs

At 1GB, the heap holds in-flight events, filter plugin state, codec buffers, and the memory queue (if used). With default pipeline.batch.size of 125 and pipeline.workers set to CPU core count, a busy pipeline exhausts 1GB in seconds.

The failure pattern is the GC death spiral: heap fills, GC runs more frequently and for longer durations, throughput drops, events accumulate faster, heap fills faster, and full-GC pauses eventually freeze the pipeline. The process stays alive (process checks pass) while throughput goes to near zero. The monitoring API itself may become unresponsive during prolonged GC pauses.

Signals that distinguish this from normal GC cycling:

  • Post-GC heap floor above 90% sustained
  • Old-gen pool above 85% of its max
  • GC overhead above 20% of wall time
  • Output rate dropping while heap stays high

A 1GB heap reaches this state quickly under any real ingestion load. The first fix is increasing heap. The long-term fix is sizing it correctly.

Why -Xms and -Xmx must be equal

Elastic guidance is explicit: set Xms and Xmx to the same value to prevent the heap from resizing at runtime, which is costly.

When -Xms is less than -Xmx, the JVM starts with a smaller heap and grows it as needed. Each growth operation allocates and zeroes new memory pages, potentially triggers a full GC, and produces an application-visible pause during the resize.

For a streaming pipeline processor, heap resizing pauses are destructive. They happen under load (when the JVM decides it needs more heap), which is exactly when you can least afford pauses. Setting -Xms equal to -Xmx allocates the full heap at startup and eliminates runtime resizing.

The trade-off: the full heap is reserved immediately, even if the pipeline is idle. On a dedicated Logstash host this does not matter. On a shared host or memory-constrained container, you must size the heap to fit within the available memory budget from the start.

How to set the heap

The canonical location is jvm.options:

  • RPM/DEB installs: /etc/logstash/jvm.options
  • Tarball installs: <logstash_home>/config/jvm.options
  • Docker: set via LS_JAVA_OPTS environment variable or a custom jvm.options mounted into the container

Set both values:

-Xms4g
-Xmx4g

Remove or comment out any conflicting -Xms/-Xmx lines. The file is read top to bottom, and later flags on the JVM command line take precedence.

LS_JAVA_OPTS alternative: Set heap via the LS_JAVA_OPTS environment variable:

export LS_JAVA_OPTS="-Xms4g -Xmx4g"

This appends to the JVM arguments constructed from jvm.options. If jvm.options already contains -Xms1g -Xmx1g, the LS_JAVA_OPTS values appear later on the command line and take precedence.

LS_HEAP_SIZE is deprecated. Use jvm.options or LS_JAVA_OPTS instead.

Bug note (fixed in 7.17/8.0): In Logstash 7.12 through 7.16, LS_JAVA_OPTS was silently ignored if a readable jvm.options was absent (regression fixed in 7.17.0 and 8.0.0, elastic/logstash#13525). If LS_JAVA_OPTS seems to have no effect on those versions, verify that jvm.options exists (even if empty).

Sizing heap for production

Elastic recommends no less than 4GB and no more than 8GB for typical ingestion. Heap should not exceed 50-75% of total physical memory, leaving room for off-heap allocation, JVM overhead, and OS page cache.

What consumes heap

ConsumerDescription
In-flight eventspipeline.batch.size (default 125) multiplied by pipeline.workers (default CPU cores). Each event includes all fields.
Filter plugin stateGrok pattern caches, translate dictionaries, GeoIP database references, Ruby filter state.
Codec buffersMultiline codec accumulates partial events in memory until pattern completion.
Memory queueIf using the default memory queue (not PQ), queued events live in heap.
pipeline.buffer.type=heap allocationsSince Logstash 9.0, input plugin buffers (Beats, TCP, HTTP, Elastic Agent) default to heap instead of direct memory; the setting itself was introduced in 8.14 with a direct default.

Sizing checklist

  • Calculate peak in-flight events. batch_size * pipeline.workers * peak_event_size gives a floor. A pipeline with batch_size=125, 8 workers, and 10KB events holds approximately 10MB per batch cycle. Filter state and codec buffers add overhead.
  • Account for all pipelines. In multi-pipeline setups, all pipelines share one JVM heap. Sum the in-flight and state requirements across all pipelines.
  • Factor in pipeline.buffer.type (9.0+). If upgrading from 8.x where pipeline.buffer.type defaulted to direct, allocations that previously consumed direct memory now consume heap. Without a heap increase, this can cause OOM after upgrade.
  • Leave room for off-heap. Reserve at least the same amount as heap for off-heap allocation, JVM internal structures, and OS page cache.
  • Persistent queue overhead. Each PQ requires memory-mapped space for head and tail pages (64MB each, 128MB per pipeline). With 10 pipelines, that is 1.28GB before any heap or direct memory. This is off-heap but competes for physical memory.

When bigger is not better

An oversized heap lengthens full-GC pause times. With G1GC, a large heap can produce multi-second full-GC pauses when old-gen finally fills.

  • Larger heap: less frequent GC, but longer pauses when they happen
  • Smaller heap: more frequent GC, but shorter pauses and faster recovery

For a latency-sensitive streaming pipeline, a 4-8GB heap with frequent but short young-gen collections is preferable to a 16GB heap with rare but devastating full-GC pauses.

Off-heap memory: the hidden budget

JVM heap is not the total memory footprint. Off-heap memory includes:

ComponentWhat it is
Direct memory buffersNetty I/O buffers for network input/output plugins. Default MaxDirectMemorySize equals -Xmx.
JVM internalCompressed class space, code cache, thread stacks, GC data structures.
Memory-mapped filesPQ page files are memory-mapped. Each pipeline’s PQ uses at least 128MB of mapped space.
flowchart TD
    A["Total Physical or Container Memory"] --> B["JVM Heap
-Xms = -Xmx
4-8GB typical"] A --> C["Direct Memory
MaxDirectMemorySize"] A --> D["JVM Overhead
class space, code cache,
thread stacks, GC"] A --> E["PQ Memory-Mapped Files
128MB per pipeline"] A --> F["OS Page Cache + Other Processes"] C --> G{"pipeline.buffer.type"} G -->|"heap (9.0+ default)"| H["Input buffers on heap"] G -->|"direct (8.x default)"| I["Input buffers on direct memory"]

MaxDirectMemorySize

By default, the JVM sets MaxDirectMemorySize equal to -Xmx. A 4GB heap reserves up to 4GB of direct memory. On a host with 16GB RAM, a 4GB heap could consume up to 8GB (heap + direct) before JVM overhead or OS needs.

Elastic’s JVM-settings guide recommends setting -XX:MaxDirectMemorySize to half of heap size:

-XX:MaxDirectMemorySize=2g

for a 4GB heap. MaxDirectMemorySize is not included in the default jvm.options.

pipeline.buffer.type and the 9.0 migration

In Logstash 8.x, pipeline.buffer.type defaulted to direct for Beats, TCP, HTTP, and Elastic Agent inputs. In 9.0.0, the default changed to heap.

  • 8.x to 9.0 upgrade risk. If you upgrade to 9.0 without adjusting heap size, buffer allocations that previously lived in direct memory now consume heap. This can cause OOM. Logstash 8.16.0 already logged a deprecation warning when the setting was left unset, so check the startup logs before upgrading.
  • Migration options. Either set pipeline.buffer.type: direct to preserve old behavior, or set pipeline.buffer.type: heap and increase heap accordingly.

Setting pipeline.buffer.type: heap has a diagnostic advantage: buffer allocations are visible in heap dumps, making OOM debugging easier. Direct memory exhaustion does not appear in standard heap analysis tools.

Container and cgroup awareness

In containers (Kubernetes, Docker), the JVM must see the container memory limit, not host RAM. Without cgroup awareness, the JVM sizes itself against total host memory and can be OOM-killed when it exceeds the container limit.

JDK 10+ behavior: -XX:+UseContainerSupport is enabled by default on JDK 10+. The bundled JDK in modern Logstash supports this (Logstash 9.0 requires Java 17 as a minimum; Logstash 9.4.0 raised the minimum to Java 21). However, MaxRAMPercentage defaults to 25%, which may be too conservative for Logstash.

To set heap as a percentage of the container limit:

-XX:MaxRAMPercentage=50

Or set explicit -Xms/-Xmx values that fit within the container limit minus off-heap overhead.

Practical container sizing: For a container with an 8GB memory limit:

  • Heap: 4GB (-Xms4g -Xmx4g)
  • Direct memory: 2GB (-XX:MaxDirectMemorySize=2g)
  • JVM overhead + PQ mapped files + OS: approximately 2GB

This leaves no headroom for PQ growth or unexpected off-heap allocation. In practice, a container running Logstash with PQ needs more room. Either increase the container limit or reduce heap.

cgroup v1 vs v2. On older systems with cgroup v1, container memory detection may be less reliable. Verify by checking the actual heap the JVM selects after startup.

Verifying the configuration

After changing jvm.options and restarting Logstash, verify the JVM picked up the correct values:

# Check effective heap via the monitoring API
curl -sS http://127.0.0.1:9600/_node/stats/jvm | python3 -c "
import sys, json
m = json.load(sys.stdin)['jvm']['mem']
print(f\"Heap max: {m['heap_max_in_bytes'] / 1024 / 1024:.0f} MB\")
print(f\"Heap used: {m['heap_used_in_bytes'] / 1024 / 1024:.0f} MB ({m['heap_used_percent']}%)\")
"

If jcmd is available:

# jcmd requires the same user that runs the JVM
jcmd $(pgrep -f org.logstash.Logstash | head -1) VM.flags

If the reported heap max does not match your configured -Xmx, the configuration was not applied. Common causes:

  • Multiple jvm.options files (check installation vs config directory)
  • LS_JAVA_OPTS overriding the file on older versions
  • jvm.options file missing entirely (the LS_JAVA_OPTS bug on pre-7.17/8.1)

Common pitfalls

  • Heap too large for the host. An 8GB heap on a 12GB host leaves 4GB for off-heap, OS, and other processes. Leads to swapping or OOM kills even though heap metrics look fine.
  • Ignoring off-heap after upgrading to 9.0. If pipeline.buffer.type changed from direct to heap, the same heap size now carries more allocation. Increase heap or set pipeline.buffer.type: direct explicitly.
  • Setting -Xms less than -Xmx in production. Runtime heap resizing causes unpredictable pauses during traffic spikes.
  • Forgetting PQ memory-mapped space. Each PQ pipeline uses at least 128MB of memory-mapped space. Ten pipelines means 1.28GB of additional physical memory pressure invisible in heap metrics.
  • Container JVM seeing host RAM. Without cgroup support or explicit heap flags, the JVM may allocate based on host memory. Always verify effective heap after startup.
  • Not enabling heap dump on OOM. Without -XX:+HeapDumpOnOutOfMemoryError, an OOM kills the process and leaves no artifact for root cause analysis. Enable it and point the dump path to a volume with enough space.

Signals to monitor

SignalWhy it mattersWarning sign
jvm.mem.heap_used_percentOverall heap pressure.Post-GC floor above 85% sustained.
jvm.mem.pools.old.used_in_bytesOld-gen accumulation indicates leak or undersized heap.Rising post-GC floor in old-gen pool.
jvm.gc.collectors.old.collection_time_in_millisFull-GC time directly steals from processing.GC overhead above 10% of wall time (warning), above 20% (severe).
jvm.gc.collectors.old.collection_countFrequency of full-GC cycles.Rate increasing over time.
Process RSSTotal physical memory including off-heap.RSS approaching container or host limit.
Pipeline output throughputWhether useful work is happening.Throughput dropping while heap is high and GC time is rising.

Alert on the post-GC floor, not the peak. The heap sawtooth pattern means peak usage regularly crosses 80% during normal GC cycles. A rising floor (heap level after GC completes) indicates real pressure that will not self-resolve.

How Netdata helps

  • Per-second heap metrics. Netdata collects jvm.mem.heap_used_percent, heap_used_in_bytes, and heap_max_in_bytes at per-second resolution, enough to distinguish normal GC cycling from a rising post-GC floor.
  • Old-gen pool tracking. Per-pool breakdown shows whether pressure is in young-gen (normal churn) or old-gen (accumulation or leak), which determines whether to increase heap or investigate a plugin memory leak.
  • GC time correlation. Correlating gc.collectors.old.collection_time_in_millis with pipeline output throughput in the same view shows whether throughput drops coincide with GC pauses, distinguishing a GC death spiral from an output bottleneck.
  • Container memory awareness. Cgroup-aware collection shows container RSS alongside heap metrics, making off-heap pressure and OOM-kill risk visible without separate tooling.
  • Anomaly detection on heap floor. Anomaly detection on the post-GC floor trend catches slow memory leaks and progressive old-gen accumulation before they trigger a death spiral.