The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-persistent-queue-corruption-wont-start ▌

Operations Guides

Logstash won't start after a crash: persistent queue corruption and checkpoint errors

Logstash was killed uncleanly (OOM kill, kill -9, SIGKILL after a pod termination grace period, power loss) and now refuses to start. The log at /var/log/logstash/logstash-plain.log shows java.io.IOException and checkpoint-related errors during pipeline initialization, and the process exits before any events flow.

The persistent queue (PQ) is a page-based on-disk queue: events live in page files (default 64MB each, up to queue.max_bytes which defaults to 1GB), and checkpoint files record which events have been acknowledged as delivered. Both are updated continuously while the pipeline runs. An unclean shutdown can leave the checkpoint files and page files inconsistent with each other, and Logstash’s startup code refuses to open a queue it cannot reconcile.

Every recovery path involves a tradeoff between re-delivery (duplicates downstream) and loss (gaps downstream). There is no zero-impact repair, because the corrupted state is precisely the record of what has been delivered. Decide which side of that tradeoff your pipeline can tolerate before deleting or moving files.

If this deployment uses the default memory queue instead of PQ, this failure mode does not exist. A crash with a memory queue loses the in-flight events silently and Logstash starts cleanly. If you are seeing startup failure with these errors, PQ is enabled.

What this means

The PQ has two kinds of state on disk under the queue directory (typically /var/lib/logstash/queue/<pipeline_id>/):

  • Page files (page.N) hold the serialized events, written sequentially.
  • Checkpoint files (checkpoint.head, checkpoint.N) record which pages and events have been fully processed and acknowledged. Checkpoints are written periodically (by default after 1024 writes or 1024 acks, or every 1000ms), not on every event.

On a clean shutdown, the final checkpoint reflects the true state of the queue. On an unclean shutdown, the last checkpoint write may be torn, stale, or missing entirely. Logstash validates this on startup and fails closed rather than guessing.

The direction of the inconsistency determines the failure consequence:

  • Checkpoint behind actual state: events that were already delivered look unacked. Recovery replays them, so downstream sees duplicates. PQ is at-least-once by design, so this is the “safe” direction.
  • Corrupted or unreadable pages: events in those pages cannot be recovered. They are gone, and repair means discarding the damaged segments.
  • Corrupted or zero-byte head checkpoint: Logstash cannot determine queue state and aborts startup. This is the most common crash artifact.

For the broader model of how the queue sits between inputs and workers, see How Logstash actually works in production and Logstash memory queue vs persistent queue.

Common causes

CauseWhat it looks likeFirst thing to check
OOM kill during PQ writeLogstash OOMKilled (container) or killed by kernel OOM; on restart, checkpoint errors in logdmesg / journalctl -k for OOM messages; container restart events
kill -9 or SIGKILL after grace periodPod or service force-killed; PQ left mid-writeOrchestrator events; journalctl -u logstash for shutdown sequence
Power loss or host crashWhole host went down; PQ checkpoint tornHost uptime; filesystem journal state
Disk full on the PQ volumePQ could not complete writes; Logstash crashed or wedgeddf -h /var/lib/logstash and queue.data.free_space_in_bytes history
Filesystem corruption or NFS/slow storage under PQCheckpoints unreadable or stale independent of shutdown cleanlinessStorage backend under path.data; NFS amplifies PQ problems
Known version bugLogstash 9.2.0 refuses to start when queue.max_bytes is 2 GiB or larger, with no corruption involvedLogstash version and queue.max_bytes in logstash.yml

One non-corruption case worth ruling out first: if you recently upgraded to Logstash 9.2.0 and set queue.max_bytes to 2147483648 (2 GiB) or more, the process will not start even with a perfectly healthy queue. That is a documented known issue with no workaround other than downgrading. If the log shows checkpoint or IOException errors rather than a config validation failure, you are dealing with real corruption instead.

Quick checks

All read-only. Do not delete anything yet.

# 1. Confirm the failure reason from the log
grep -Ei '(IOException|checkpoint|queue|persistent)' /var/log/logstash/logstash-plain.log | tail -n 50

# 2. Confirm how Logstash died last time (root cause matters later)
journalctl -u logstash --since "2 hours ago" | tail -n 100
dmesg -T | grep -i -E '(oom|killed process)' | tail -n 20

# 3. Check queue type and PQ settings
grep -E 'queue\.(type|max_bytes|page_capacity|checkpoint)' /etc/logstash/logstash.yml

# 4. Inspect the queue directory contents (sizes, zero-byte files)
ls -la /var/lib/logstash/queue/*/

# 5. Check disk state on the PQ volume
df -h /var/lib/logstash
du -sh /var/lib/logstash/queue/*/

# 6. Verify Logstash is actually stopped before any repair
pgrep -f org.logstash.Logstash

If running in a container, replace journalctl -u logstash with your orchestrator’s pod event log or kubectl describe pod / docker inspect.

In step 4, look for a zero-byte checkpoint.head. That alone is enough to block startup. A checkpoint.head.tmp alongside checkpoint.head indicates a checkpoint write was interrupted mid-rename, and the .tmp file may be the newer, intact copy.

How to diagnose it

  1. Identify the exact error. The log line tells you which file is bad. A checkpoint checksum mismatch points at checkpoint.head. Errors referencing page numbers point at page files. The pipeline ID in the error tells you which queue subdirectory is affected in multi-pipeline setups.

  2. Run pqcheck. Logstash ships a diagnostic utility that reads the queue offline. Logstash must be stopped. For package installs (RPM/DEB), the tool is at /usr/share/logstash/bin/pqcheck; for tarball installs, it is under bin/ in the extraction directory.

# Point pqcheck at the affected pipeline's queue dir
/usr/share/logstash/bin/pqcheck /var/lib/logstash/queue/main

pqcheck walks the checkpoint files and reports each page’s state: page numbers, the first unacknowledged page and event, and whether pages are fully acknowledged.

A page whose size reports as “NOT FOUND” is corrupted. This tells you whether you have a checkpoint-only problem (cheap to fix, possible duplicates) or damaged pages (data loss is already real, repair just bounds it).

  1. Map damage to impact. If only the head checkpoint is bad and all pages check out, recovery risks re-delivery of recently acknowledged events but no loss. If pqcheck reports corrupted pages, the events in those pages are unrecoverable; the decision is how much surrounding state you discard to get a bootable queue.

  2. Check the root cause before recovering. If the OOM killer or a full disk caused the crash, recovering the queue without fixing that guarantees a repeat. Confirm heap sizing, container memory limits, and free space on the PQ volume now, not after the next crash.

The recovery decision tree:

flowchart TD
    A[Logstash fails to start with PQ errors] --> B[Back up the queue directory]
    B --> C[Run pqcheck on the queue dir]
    C --> D{What is damaged?}
    D -->|Head checkpoint only| E[Restore checkpoint.head from checkpoint.head.tmp, or delete checkpoint files]
    D -->|Corrupted pages| F[Run pqrepair to remove corrupt segments]
    D -->|pqrepair fails or damage too broad| G[Move queue dir aside, start fresh]
    E --> H[Start Logstash, expect some re-delivery]
    F --> H2[Start Logstash, expect gaps in corrupt pages]
    G --> H3[Start Logstash, queued events in old dir are lost]
    H --> I[Fix unclean-shutdown root cause]
    H2 --> I
    H3 --> I

Metrics and signals to monitor

You cannot watch these during the outage (the stats API on port 9600 is down with the process), but they are the signals that tell you the next crash is coming, and they confirm recovery afterwards.

SignalWhy it mattersWarning sign
JVM heap post-GC floor (jvm.mem.heap_used_percent, old-gen pool)Rising floor is the path to the OOM kill that corrupts the PQFloor trending up over hours; old-gen above 85% of max
GC overhead (jvm.gc.collectors.old.*)GC death spiral precedes OOM and can freeze the JVM mid-writeOld-gen GC time above 20% of wall time
Process RSS vs container/host limitThe OOM killer acts on RSS, not heapRSS approaching the cgroup or host limit
Disk free on PQ volume (queue.data.free_space_in_bytes)Full disk causes failed writes and crashes even within max_bytesFree space declining; max_bytes sized near partition size
PQ occupancy and growth (queue.queue_size_in_bytes, flow.queue_persisted_growth_bytes)A large, growing queue means more in-flight state at risk per crashOccupancy above 80% with positive growth
jvm.uptime_in_millisDetects restart loops and unexpected restartsUptime resetting between polls

Note the PQ sizing subtlety: max_bytes limits the queue’s logical size, but pages are allocated in fixed increments (default 64MB) and freed only when fully drained and checkpointed, so on-disk usage can exceed the configured limit. Size max_bytes against the partition, not against your comfort level.

Fixes

Every step below is destructive in different ways. Back up the queue directory before any of them, with Logstash stopped. Ensure the backup volume has enough free space for a copy of the full queue:

# Stop Logstash, then back up the queue for forensics or a second attempt
systemctl stop logstash
# In a container: docker stop / kubectl delete pod instead
cp -a /var/lib/logstash/queue /var/lib/logstash/queue.backup-$(date +%Y%m%d-%H%M)

Fix 1: repair with pqrepair (preferred when pages are corrupt)

# Removes corrupt queue segments. Run per pipeline queue dir. Use the logstash user
# so file ownership stays consistent. Path shown is for package installs.
sudo -u logstash /usr/share/logstash/bin/pqrepair /var/lib/logstash/queue/main

pqrepair produces no output on success. It discards corrupted segments, which means the events in those pages are lost; the rest of the queue survives and will drain normally. This is the least-bad option when pqcheck shows real page damage, because it preserves everything readable.

Fix 2: restore or remove the head checkpoint (checkpoint-only corruption)

If pqcheck shows healthy pages and the log points at the head checkpoint:

  • If a checkpoint.head.tmp exists, it may be the newer intact copy. Logstash writes each checkpoint through a temporary .tmp file that is then atomically renamed over the target, so an interrupted write can leave the old checkpoint.head beside a newer, intact .tmp. Restoring the .tmp over checkpoint.head is a reasonable recovery when pqcheck confirms the page files are healthy and only the head checkpoint is at fault.

  • If there is no usable .tmp file, or the head checkpoint is zero bytes, delete the checkpoint files. Logstash will rebuild checkpoint state from the page files on startup.

The tradeoff: rebuilt checkpoints treat already-delivered events as unacknowledged, so downstream will see some duplicates. Most Elasticsearch-destined pipelines tolerate this (duplicate documents overwrite by ID or appear as dupes); measure it against your destination’s dedup behavior.

Fix 3: move the queue aside (last resort)

# WARNING: queued events in the old directory will NOT be delivered.
mv /var/lib/logstash/queue/main /var/lib/logstash/queue/main.corrupt
systemctl start logstash

Logstash creates a fresh empty queue and starts. Everything that was queued is lost from Logstash’s perspective. Whether that data is truly gone depends on your sources: Kafka inputs can re-consume (consumer offsets are in Kafka, not the PQ), Beats sources resend from their own registry, but fire-and-forget sources (UDP syslog, some HTTP) are gone for good.

Fix 4: fix the root cause

Recovery without this step is a scheduled repeat incident:

  • OOM kill: raise container/host memory or lower JVM heap so RSS fits with headroom. Heap plus off-heap overhead should sit comfortably under the limit; heap alone is not the whole footprint.
  • Force kills: give Logstash a real shutdown window. For planned maintenance on PQ deployments, queue.drain: true drains the queue before shutdown so nothing is in flight. In Kubernetes, size the termination grace period around drain time, not the default 30s.
  • Disk full: separate the PQ volume from log and DLQ storage, or lower max_bytes relative to the partition.
  • Config trigger: if you set queue.max_bytes at or above 2 GiB on Logstash 9.2.0, downgrade; that version will not start regardless of queue health.

Prevention

  • Eliminate unclean shutdowns. They are the only common trigger for this failure. That means memory limits with headroom, graceful stops with adequate grace periods, and no kill -9 in runbooks.
  • Alert on the pre-crash signals from the table above: rising post-GC heap floor, GC overhead, RSS near limit, PQ volume free space. The crash is the last step of a visible trend.
  • Keep PQ on local, reliable storage. NFS and slow network storage under the PQ amplify both latency and corruption risk.
  • Size max_bytes against the partition with room for page-allocation overshoot, and keep the PQ partition dedicated.
  • Know your source replay capability before an incident. If sources cannot replay (UDP, transient HTTP), the blast radius of a queue wipe is permanent loss, which should factor into how aggressively you protect the queue volume.
  • Expect re-delivery after checkpoint recovery. Make sure downstream consumers tolerate duplicates, because at-least-once is the PQ’s design contract even in the best-case recovery.

How Netdata helps

  • Netdata’s Logstash collector polls the node stats API, so heap sawtooth, old-gen pool usage, and GC collection time are charted per second. A rising post-GC floor is visible hours before the OOM kill that corrupts the queue.
  • PQ occupancy, queue_size_in_bytes against max_queue_size_in_bytes, and persisted growth rate show how much in-flight state a crash would put at risk, and how full the queue was when it happened.
  • Correlating JVM metrics with host-level signals (RSS against cgroup limits, disk free space and I/O on the PQ volume, OOM kill events) connects the unclean shutdown to its cause in one view instead of three tools.
  • JVM uptime resets and API reachability gaps make restart loops and crash timing obvious, which helps line the crash up with orchestrator events.
  • After recovery, output throughput against input throughput confirms the queue is draining and the pipeline is actually processing, not just running.