The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-dead-letter-queue-growing ▌

Operations Guides

Logstash dead letter queue growing: DLQ diversion, replay, and disabled-by-default risk

A growing dead_letter_queue.queue_size_in_bytes means an output is permanently rejecting events and diverting them to the DLQ instead of delivering them downstream. This is a data-integrity failure, not just a metric ticking up. Every event in the DLQ is data your destination never received.

Two design properties compound the risk. The DLQ is disabled by default in all Logstash versions; many production deployments run without it, meaning permanently failed events are logged and silently lost with no on-disk record. And the DLQ only captures specific output failure classes. Filter and parse failures (grok, json, date) do not go to the DLQ. They pass through the pipeline with _grokparsefailure or similar tags and reach the output as malformed events.

When enabled, the DLQ defaults to a 1 GB maximum. Once full, storage_policy controls behavior: drop_newer (default) stops accepting new events; drop_older evicts the oldest entries. In both cases, a full DLQ means the safety net is itself discarding data.

What this means

The metric pipelines.<pipeline_id>.dead_letter_queue.queue_size_in_bytes reflects on-disk DLQ size. A rising value means events are accumulating because an output permanently rejected them after exhausting retries. The DLQ is not a retry queue. Events sit there until a human replays them, deletes them, or the queue fills and the storage policy kicks in.

What lands in the DLQ:

  • Events rejected by the Elasticsearch output with non-retriable HTTP response codes (400, 404). These are permanent failures: mapping conflicts, type mismatches, malformed documents the destination cannot accept.
  • Events that fail conditional statement evaluation in the pipeline configuration.

What does NOT land in the DLQ:

  • Filter and parse failures. A grok pattern that fails to match tags the event with _grokparsefailure and sends it through the normal output path. The DLQ never sees it.
  • Transient output errors. Connection timeouts, HTTP 429 rate limiting, and temporary unavailability are retried by the output plugin. Only permanent rejections after retries are exhausted go to the DLQ.

Official docs confirm DLQ support is limited to the Elasticsearch output and conditional-statement evaluation; other output types (Kafka, S3, HTTP) never write to the DLQ.

flowchart LR
    A[Input] --> B[Queue]
    B --> C[Workers + Filters]
    C -->|parse fail| D[Tagged event
bypasses DLQ] C -->|parsed OK| E[Output plugin] D --> E E -->|accepted| F[Destination] E -->|permanent reject
HTTP 400 or 404| G{DLQ enabled?} G -->|no| H[Logged and
silently lost] G -->|yes| I[DLQ on disk] I -->|full| J[drop_newer or
drop_older] I -->|manual replay| K[Replay pipeline] K --> C

The diagram shows the two paths that matter most. Parse failures bypass the DLQ entirely, reaching the destination as malformed data. If the DLQ is disabled, permanently rejected events vanish with only a log entry.

Common causes

CauseWhat it looks likeFirst thing to check
Elasticsearch mapping conflictBulk request returns 400 with per-document errors; DLQ grows steadilyES output documents.non_retryable_failures counter
Schema or type mismatch at sourceNew field appears with wrong type (string vs integer); specific events always failSample DLQ content and compare against current index mapping
Poison-pill eventsSmall number of specific events repeatedly fail; DLQ grows in small burstsInspect a few DLQ entries for common field patterns
Index write blockES index set to read-only (disk watermark exceeded); all writes failES cluster health and index settings
Version incompatibilityAfter ES upgrade, some documents fail due to breaking mapping changesES version and recent upgrade history

Quick checks

# Check DLQ size and max via the stats API
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty | grep -A 5 dead_letter_queue
# Check DLQ size on disk
du -sh /var/lib/logstash/dead_letter_queue/*/
# Check if DLQ is enabled and its settings
grep -E 'dead_letter_queue\.(enable|max_bytes|storage_policy)' /etc/logstash/logstash.yml
# Check ES output document-level failure counts
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty \
  | python3 -c "
import sys,json
outputs = json.load(sys.stdin)['pipelines']['main']['plugins']['outputs']
for o in outputs:
    name = o.get('name','?')
    docs = o.get('documents',{})
    print(f'{name}: non_retryable_failures={docs.get(\"non_retryable_failures\",\"N/A\")}')
"
# Check for rejection and mapping errors in logs
grep -Ei '(reject|mapping|error|exception|failed)' /var/log/logstash/logstash-plain.log | tail -n 100
# Check Elasticsearch cluster health
curl -s localhost:9200/_cluster/health?pretty
# Check the grok failures counter (parse failures bypass the DLQ)
curl -sS http://127.0.0.1:9600/_node/stats/pipelines?pretty \
  | python3 -c "
import sys,json
filters = json.load(sys.stdin)['pipelines']['main']['plugins']['filters']
for f in filters:
    if f.get('name') == 'grok':
        print(f\"grok ({f.get('id','?')}): failures={f['events'].get('failures', 0)}\")
"

How to diagnose it

  1. Confirm DLQ is enabled and growing. Query the stats API for dead_letter_queue.queue_size_in_bytes. Take two samples 60 seconds apart. If the value is rising, events are actively being diverted. If DLQ is not enabled, check logs for output rejection messages: those events were silently lost.

  2. Check the Elasticsearch output for non-retryable failures. In the stats API response, look at plugins.outputs[] for documents.non_retryable_failures or bulk_requests.failures. These tell you the destination is permanently rejecting events.

  3. Sample DLQ content. The DLQ stores events in a binary format at <path.data>/dead_letter_queue/<pipeline_id>/. Replay a few events through a temporary pipeline using the dead_letter_queue input plugin to inspect what is failing. Each DLQ entry includes the original event, the reason for failure, and the failure timestamp.

  4. Identify the pattern. Are all events failing (systemic issue like mapping conflict or index write block), or only some (poison-pill events or specific field values)?

  5. Check the destination mapping. Compare fields in failing events against the Elasticsearch index mapping. A field defined as long in the mapping receiving a string value is a classic mapping conflict.

  6. Check for recent source format changes. If the source application changed its log format, new fields or changed field types may conflict with the existing mapping.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
dead_letter_queue.queue_size_in_bytesPrimary DLQ growth indicatorAny sustained increase above zero
dead_letter_queue.max_queue_size_in_bytesConfigured DLQ capacityQueue size approaching this limit means data will be discarded
ES output documents.non_retryable_failuresEvents the destination permanently rejectedAny non-zero and growing value
Pipeline events.outWhether events are flowing at allNormal output rate with DLQ growth equals silent partial failure
Grok filter failures counterParse failures that bypass DLQ entirelyRising rate means data quality is degrading without DLQ involvement
Disk usage on DLQ partitionDLQ competes with PQ and logs for spaceShared partition filling up can crash Logstash

Fixes

Fix the root cause

The DLQ is a symptom, not the problem. The problem is that specific events cannot be accepted by the destination. Fix options depend on the cause:

Elasticsearch mapping conflict. Update the index mapping to accept the new field type, or transform the event in a Logstash filter to match the expected type. If the conflict comes from a new field the source should not be sending, drop or rename it before the output.

Index write block. If Elasticsearch set the index to read-only due to disk watermarks, free disk space on the data nodes. The write block lifts automatically once the disk watermark recovers.

Poison-pill events. If a small number of malformed events are causing rejections, add a filter condition to catch and route them before they reach the output. Quarantine them to a separate index for investigation.

Replay events from the DLQ

DLQ events do not automatically retry. To replay them, create a temporary pipeline that reads from the DLQ using the dead_letter_queue input plugin:

input {
  dead_letter_queue {
    path => "/var/lib/logstash/dead_letter_queue"
    pipeline_id => "main"
    commit_offsets => true
  }
}

The commit_offsets option (default true) records the read position so events are not re-read on restart. Since Logstash 8.4.0, the clean_consumed option automatically deletes consumed segments, keeping the DLQ from growing during replay.

The queue_size_in_bytes accounting fix for clean_consumed replays shipped in Logstash 8.15.0 (elastic/logstash#16195); earlier versions can report a stale DLQ size while consumed segments are deleted.

Replay risks. Replay re-sends events to the output. If the root cause is not fixed, the same events will fail again and re-enter the DLQ. Always confirm the root cause is resolved before replaying. Replay also adds load to the destination, which may already be under pressure.

Clear the DLQ

If the DLQ contains events you choose not to replay, you can clear it manually:

There is no online clearing path: the official docs require stopping the pipeline before deleting the DLQ directory, and no newer version has added an API to clear it while running.

  1. Stop the Logstash pipeline or the entire Logstash process.

  2. Delete the DLQ directory: rm -rf <path.data>/dead_letter_queue/<pipeline_id>

    WARNING: This is destructive and unrecoverable. Confirm the pipeline is fully stopped before deleting. Verify <path.data> resolves correctly before running the command.

  3. Restart Logstash.

The DLQ directory cannot be deleted while the pipeline is running. The path follows the pattern <path.data>/dead_letter_queue/<pipeline_id>/.

Enable the DLQ if it is disabled

If you discover DLQ is disabled on a production pipeline, permanently failed events have been silently lost. There is no way to recover them retroactively. Enable the DLQ going forward:

In logstash.yml:

dead_letter_queue.enable: true
dead_letter_queue.max_bytes: 1024mb
dead_letter_queue.storage_policy: drop_newer

The storage_policy option requires Logstash 8.3.0 or later (verified: the setting was added in 8.3.0). On earlier versions, the behavior is always drop_newer. Restart Logstash for the change to take effect.

Prevention

  • Monitor DLQ growth even when throughput looks healthy. A pipeline can show normal events.out while events are silently diverted. The output counter increments for successful documents in a partial bulk response, masking rejected ones.
  • Alert on any DLQ growth. Any non-zero growth rate on a production pipeline warrants investigation. The question is not “is the DLQ full?” but “why is even one event being diverted?”
  • Track DLQ capacity, not just size. Monitor queue_size_in_bytes relative to max_queue_size_in_bytes. A DLQ at 90% capacity with drop_newer is one event away from permanent data loss.
  • Do not confuse DLQ with parse failure monitoring. The DLQ catches output rejections. Parse failures are invisible to the DLQ. Monitor the grok filter failures counter and parse failure tags at the destination separately.
  • Prepare a replay pipeline before you need it. Most teams never replay DLQ events because they have no replay pipeline ready. Build and test the configuration in staging, and document the procedure.
  • Watch disk space. The DLQ competes with the persistent queue, log files, and the OS for the same partition. A growing DLQ can fill a disk and crash Logstash even if the queue is within its configured limit.

How Netdata helps

  • Per-second collection of dead_letter_queue.queue_size_in_bytes catches growth the moment it starts, before the DLQ fills and the storage policy begins discarding events.
  • Correlating DLQ growth with Elasticsearch output error rates (documents.non_retryable_failures, bulk request failures) pinpoints whether the destination is rejecting events.
  • ML anomaly detection on DLQ size flags unexpected growth patterns that static thresholds miss, especially on pipelines with intermittent DLQ activity.
  • Disk space monitoring on the DLQ partition warns when the queue competes with PQ and log volumes for the same filesystem.
  • Pipeline throughput metrics alongside DLQ size reveal the silent partial failure pattern where throughput looks normal but events are being diverted.
  • Grok filter failures counter monitoring catches parse failures that bypass the DLQ, covering the correctness gap the DLQ does not.