The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-jsonparsefailure ▌

Operations Guides

Logstash _jsonparsefailure: malformed JSON and codec mismatches

When events arrive at your Elasticsearch indices carrying the _jsonparsefailure tag, they contain raw, unparsed data instead of the structured fields your downstream consumers expect. Throughput metrics look healthy, events-out counts keep climbing, and no alerts fire. But the data is wrong.

The _jsonparsefailure tag is added by either the json codec or the json filter when it receives input it cannot parse as valid JSON. The event is not dropped. It flows through to the output carrying whatever raw data was received, plus the failure tag. This is a silent correctness failure, not an availability failure.

The critical diagnostic question: is the source data genuinely malformed, or is Logstash configured to parse data that was never JSON in the first place? Correcting a codec mismatch takes one line of config. Chasing a producer that occasionally truncates JSON objects may require upstream changes.

What this means

Both the json codec and the json filter can emit _jsonparsefailure. Their behavior on failure differs:

  • json codec: Falls back to plain text. The raw payload is stored in the message field, and the event continues through the pipeline with the _jsonparsefailure tag.
  • json filter: Leaves the event untouched (the source field retains its original value) and adds the _jsonparsefailure tag.

In both cases the event reaches the output. It is not diverted to the Dead Letter Queue. The DLQ captures output delivery failures, not filter or codec parse errors. An event tagged _jsonparsefailure still increments the events-out counter, so standard throughput monitoring stays green.

Because throughput metrics do not reflect this failure, it commonly goes unnoticed. Dashboards querying structured fields return fewer results, and the degradation can persist for weeks.

The json filter has a skip_on_invalid_json option (default false). When set to true, the filter returns without parsing or tagging the event on failure. This does not fix the data quality problem. It hides it.

flowchart TD
    A["_jsonparsefailure on events"] --> B["Sample failed events
from destination"] B --> C{"Raw message
is valid JSON?"} C -->|"No - plain text"| D["Codec mismatch:
fix input codec"] C -->|"No - truncated"| E["Source issue:
TCP drops or
multiline needed"] C -->|"Yes"| F{"Both json codec
and json filter
on same path?"} F -->|"Both"| G["Double-parsing:
remove filter
or change codec"] F -->|"Filter only"| H{"Keys have brackets
or metadata
non-object?"} H -->|"Yes"| I["Field-reference
or metadata bug"] H -->|"No"| J["Encoding issue:
check charset"]

Common causes

CauseWhat it looks likeFirst thing to check
Codec mismatchProducer sends plain text, syslog strings, or key=value pairs to an input with codec => json. Every event gets tagged.Inspect raw message field on failed events. Is it actually JSON?
Double-parsing (codec + filter)Input has codec => json and a filter does json { source => "message" }. The codec already parsed the JSON into fields. The filter then tries to parse the message field, which is now empty or contains a non-JSON string, and fails.Check pipeline config for both json codec and json filter on the same data path.
Truncated or partial JSONTCP connection drops mid-message, or a producer sends incomplete JSON objects. Failures are intermittent rather than constant.Check whether failures correlate with network events or source-side issues.
Concatenated JSON objectsProducer sends multiple JSON objects on one line without a delimiter (e.g., {"a":1}{"b":2}). The codec fails because trailing content after the first object makes the line invalid JSON.Inspect raw messages for multiple JSON objects on a single line.
Character encoding mismatchSource emits non-UTF-8 data (e.g., CP1252 from nxlog or Windows producers). The json codec default charset is UTF-8.Check the charset option and the source’s actual encoding.
@metadata set to non-object typeIncoming JSON contains {"@metadata": null}, {"@metadata": "string"}, or {"@metadata": 0}. Logstash throws ClassCastException and tags the event. Known bug (issue #13630). skip_on_invalid_json does not prevent this.Inspect failed events for @metadata field values.
Square brackets in field namesJSON keys containing [ or ] trigger Invalid FieldReference errors in the json filter path. The --field-reference-escape-style flag (added in Logstash 8.3.0) only fixes the codec path, not the filter path.Check if failed events have keys containing brackets.

Quick checks

# Check Logstash log for JSON parse errors and related exceptions
grep -iE 'JSON parse|_jsonparsefailure|Invalid FieldReference|ClassCastException' \
  /var/log/logstash/logstash-plain.log | tail -n 50

# Count _jsonparsefailure events in Elasticsearch.
# If `tags` is mapped as text, use `tags.keyword` instead.
curl -s 'http://localhost:9200/<index>-*/_count' -H 'Content-Type: application/json' -d '{
  "query": { "term": { "tags": "_jsonparsefailure" } }
}'

# Sample three failed events to inspect raw message content
curl -s 'http://localhost:9200/<index>-*/_search?size=3' -H 'Content-Type: application/json' -d '{
  "query": { "term": { "tags": "_jsonparsefailure" } }
}'

# Check pipeline config for double-parsing: json codec AND json filter on same path
grep -rn 'codec.*json' /etc/logstash/conf.d/
grep -rn 'json {' /etc/logstash/conf.d/

# Verify a sample raw message is valid JSON
echo '<paste raw message here>' | python3 -m json.tool

# Check json filter settings including skip_on_invalid_json
grep -B2 -A10 'json {' /etc/logstash/conf.d/*.conf

How to diagnose it

  1. Sample the failed events. Query your destination for events with the _jsonparsefailure tag. Look at the raw message field. This is your primary diagnostic signal. The raw content tells you what Logstash actually received.

  2. Determine if the raw message is valid JSON. If it is valid, the problem is on the Logstash side (double-parsing, field-reference issues, encoding). If it is not, the problem is at the source (producer sending non-JSON, truncated data, encoding mismatch).

  3. Check for double-parsing. If the input uses codec => json, the JSON is already deserialized before filters run. The message field no longer contains the original JSON string. A downstream json { source => "message" } filter then fails because message is empty or contains a non-JSON value. Fix: remove the json filter, or switch the input codec to plain and keep only the filter.

  4. Check for codec mismatch. If the raw message is not JSON (plain text, syslog, key=value), the input codec is wrong. Change the codec to plain or line and parse in a filter, or fix the producer to send JSON.

  5. Check for special key issues. If the raw message looks like valid JSON but events are still failing, inspect the keys. JSON keys containing [ or ] cause failures in the json filter path due to Logstash’s field-reference parser. A @metadata field set to a non-object type (null, string, number) causes a ClassCastException that also results in _jsonparsefailure. Neither is fixed by skip_on_invalid_json.

  6. Check for encoding issues. If the raw message contains non-ASCII characters and the source may not be UTF-8, set the charset option on the json codec to match the source encoding.

  7. Check for truncated or multiline JSON. If failures are intermittent and correlate with high traffic or network instability, the source may be sending partial JSON. Use a multiline codec on the input to reassemble fragmented messages, then parse with a json filter instead of a json codec. An input can only have one codec, so you cannot chain multiline and json codecs on the same input.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
_jsonparsefailure tag rate at destinationPrimary data quality indicator. Parse failures mean events arrive with raw data instead of structured fields.Sudden spike after a source-side change. Sustained non-zero rate on a stream that should be all JSON.
Filter failures counterThe Logstash stats API exposes per-filter metrics at _node/stats/pipelines. If the pipeline uses grok alongside json, correlated failures point to a common source format change.Counter rate increasing above baseline alongside jsonparsefailure spikes.
Output throughput (events.out)Stays green during parse failures. Events with failure tags still count as successfully delivered.Normal throughput with rising failure tags equals silent correctness degradation.
Config reload state (reloads.failures)A failed config reload can leave old config running that no longer matches the data format.Non-zero reloads.failures after a config deploy.
Log error rate for JSON parse messagesThe Logstash log records codec and filter parse errors. Since Logstash 7.15, these are logged at ERROR level instead of INFO, which can create feedback loops if logs are re-ingested.Spike in ERROR-level log lines mentioning JSON or parse, especially after an upgrade.

Fixes

Fix double-parsing

The most common cause of spurious _jsonparsefailure. If the input already uses codec => json, the JSON is deserialized at ingestion time. The original message field no longer contains the raw JSON string. Adding filter { json { source => "message" } } then tries to parse the now-empty or modified message field and fails.

Pick one path:

  • Option A: Use codec => json on the input. Remove the json filter entirely. The fields are already parsed into the event.
  • Option B: Use codec => plain (or line) on the input. Keep the json filter. The raw JSON arrives in message and the filter parses it.

Do not do both.

Fix codec mismatch

If the source sends non-JSON data to an input configured with codec => json, every event will be tagged. Check what the producer actually sends:

# Capture raw data hitting a TCP input without interfering with the pipeline.
# Run on the Logstash host. Adjust the port to match your input.
tcpdump -i any -A -s0 'tcp port <input_port>' -c 50

# If the source exposes its own TCP listener, you can connect directly.
# WARNING: this consumes data from the source. Do not run this while
# Logstash is actively reading from the same source.
nc <source_host> <source_port> | head -5

Either fix the producer to emit valid JSON, or change the codec to match the actual format and parse with appropriate filters.

Handle mixed formats conditionally

If the same input receives both JSON and non-JSON events (common with syslog inputs where some applications format as JSON and others as plain text), use conditional parsing:

filter {
  if [message] =~ /^\s*{/ {
    json {
      source => "message"
      target => "json_data"
    }
  } else {
    # Handle non-JSON format with grok, dissect, or kv
  }
}

This prevents non-JSON events from being tagged. The regex check avoids the parse attempt entirely for input that does not start with a JSON object delimiter.

Fix encoding issues

If the source emits data in an encoding other than UTF-8 (common with nxlog on Windows sending CP1252), set the charset option on the json codec:

input {
  tcp {
    port => 5000
    codec => json {
      charset => "CP1252"
    }
  }
}

The codec converts from the specified charset to UTF-8 before attempting JSON parsing.

Handle @metadata non-object type

This is a known Logstash bug (issue #13630). If the source JSON contains @metadata as null, a string, or a number, Logstash throws ClassCastException and tags the event _jsonparsefailure. Setting skip_on_invalid_json => true on the json filter does not prevent this.

The tracking issue (elastic/logstash#13630) remains open as of 2026 with no fix released; the sanitization workarounds below still apply in current Logstash.

Workaround: sanitize the @metadata field upstream at the producer, or use a ruby filter to remove or convert it before the json filter processes the event.

Handle square brackets in field names

JSON keys containing [ or ] conflict with Logstash’s field-reference syntax. In Logstash 8.3.0+, the --field-reference-escape-style flag was added, but it only works for the json codec path, not the json filter path. As of Logstash 8.5.3, the filter path still fails, and the tracking issue (elastic/logstash#14821) remains open as of 2026: the field-reference escape style covers event creation from codecs, not the Event API calls the json filter uses, so the workarounds below still apply.

Options:

  • Use codec => json instead of the json filter, combined with --field-reference-escape-style ampersand.
  • Use a ruby filter to sanitize keys with gsub before parsing. Note: this has severe performance impact on high-throughput pipelines (reported degradation from 4k events/sec to 150 events/sec).

Avoid re-ingesting Logstash’s own logs through a json codec

The json codec’s “Falling back to plain-text” message is logged at INFO level (it was logged at ERROR only in codec versions 2.1.0 through 2.1.3, released in 2016; there is no 7.15 log-level change). Even at INFO, if Logstash logs to syslog and syslog is re-ingested by a pipeline that uses the json codec, the fallback message can itself fail JSON parsing and reappear as a _jsonparsefailure event, feeding back into the same pipeline.

Check whether Logstash’s own logs are being re-ingested through the same pipeline that parses your data. Route Logstash’s own logs separately, or filter them out at the input level before they reach the json codec.

Prevention

  • Monitor the _jsonparsefailure tag rate. Do not wait to discover parse failures weeks later when a dashboard stops showing data. Query your destination regularly for failure tag counts, or add a metrics filter in the pipeline to count tagged events.
  • Route failed events to a separate index. This preserves raw data for debugging and keeps malformed events from polluting your primary data store. It also makes the failure rate trivially queryable.
  • Validate config changes for double-parsing. Before deploying a config that includes a json filter, verify the input codec is not also json. This is the single most common preventable cause.
  • Gate source format changes. Coordinate with application teams before they change log output format. A format change at the source that breaks JSON structure will silently degrade data quality until someone notices missing fields downstream.
  • Set skip_on_invalid_json deliberately. If you set it to true on the json filter, failed events pass through without the tag. You lose the ability to detect parse failures at all. Only use this when you have another mechanism to validate data quality.

How Netdata helps

  • Correlate parse failure spikes with deployment events. A sudden increase in _jsonparsefailure events that coincides with a config reload or application deploy pinpoints the change that introduced the mismatch.
  • Per-pipeline throughput visibility. In multi-pipeline setups, aggregate metrics mask parse failures in one pipeline. Per-pipeline event rates help identify which pipeline is producing tagged events.
  • Config reload state monitoring. Tracking reloads.failures alongside parse failure rates catches the scenario where a failed reload leaves old config running against a new data format.
  • Log-based anomaly detection. JSON parse error patterns in Logstash logs, especially after the 7.15 log-level change to ERROR, can be correlated with throughput and pipeline health metrics to distinguish data quality issues from availability issues.