The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / logstash / logstash-file-descriptor-pressure ▌

Operations Guides

Logstash file descriptor pressure: leaks, tailed files, and reconnection churn

Logstash is up, throughput looks normal, and then new connections start failing, file opens error out, and the process dies with “too many open files.” The failure looks sudden. It usually is not. File descriptor pressure builds over days or weeks as process.open_file_descriptors climbs toward process.max_file_descriptors, and exhaustion is a cliff edge: everything works until the limit is hit, then new connections, new file opens, and sometimes logging itself fail at once.

The three drivers that show up again and again are slow FD leaks (connections or files opened and never closed), file inputs tailing more files than expected, and output reconnection storms that churn sockets faster than the OS reclaims them. All three are visible in the same two numbers if you look before the cliff.

What this means

Every network connection (input listeners, output connections), every tailed file, every persistent queue page file, and every log file Logstash holds open consumes one file descriptor. The process has a ceiling: process.max_file_descriptors, the soft limit the service manager set at startup. As the open count approaches it, Logstash fails in partial, confusing ways: it is alive, the API may still respond, but it cannot accept new work.

The diagnostic split that matters:

  • A stable high baseline means you are legitimately using many FDs (tailing thousands of files, many Beats connections). This is a sizing problem.
  • A slow climb that never returns to baseline means a leak. Leaks can take weeks to manifest, which is why they are missed.
  • Sharp spikes correlated with downstream trouble mean reconnection churn: an output loses its connection, reconnects aggressively, and old sockets linger in states like CLOSE_WAIT while new ones pile up.
flowchart TD
  A[File input tails many files] --> D[open_file_descriptors rising]
  B[Connection leak: opened, never closed] --> D
  C[Output reconnection storm: CLOSE_WAIT buildup] --> D
  D --> E{open / max ratio}
  E -->|under 80%| F[Watch trend, compute runway]
  E -->|over 80% sustained| G[Act now: identify FD type]
  G --> H[Cliff: too many open files, partial failure]

If FDs are growing, estimate time to exhaustion:

runway = (max_file_descriptors - open_file_descriptors) / growth_rate

A leak growing 50 FDs per day with 8,000 FDs of headroom gives you roughly 160 days. That sounds comfortable until nobody notices for 160 days.

Common causes

CauseWhat it looks likeFirst thing to check
Connection/socket leak in an input or output pluginFD count climbs steadily over days or weeks, never returns to baseline after traffic dropsls -l /proc/<pid>/fd for sockets that never disappear; correlate FD growth with connection counts
File input with broad wildcardsHigh, stable FD count; matches the number of files discovered under paths like /var/log/**/*.logCount files matched by the input path; compare against open FD count of type regular file
Output reconnection stormFD spikes during downstream outages or flapping; many sockets in CLOSE_WAITss -tan state close-wait and output error/retry lines in logstash-plain.log
Persistent queue and DLQ page filesFD count grows during queue buildup, may lag behind queue drainls -l /proc/<pid>/fd for FDs pointing at the queue and dead_letter_queue directories
Limit set too low for the workloadFD ratio sits above 80% during normal peaks with no growth trendcat /proc/<pid>/limits and the systemd unit’s LimitNOFILE

Quick checks

All read-only and safe to run during an incident.

# FD pressure straight from the Logstash API
curl -sS http://127.0.0.1:9600/_node/stats/process?pretty
# Read process.open_file_descriptors, process.peak_open_file_descriptors,
# and process.max_file_descriptors. Compute the ratio.
# OS-level view of the same thing
PID=$(pgrep -f org.logstash.Logstash)
ls /proc/$PID/fd | wc -l
grep "Max open files" /proc/$PID/limits

The limits output shows both soft and hard limits; Logstash is bounded by the soft limit. Service managers override shell defaults, so the number here is the only one that matters. On systemd systems the effective limit comes from LimitNOFILE in the unit (or an override in /etc/systemd/system/logstash.service.d/), not from an interactive ulimit. On non-systemd (SysV-style) package installs, LS_OPEN_FILES in /etc/logstash/startup.options is still present (package default 16384) and is honored when the service is generated via bin/system-install, which passes it as --limit-open-files; on systemd installs the effective limit comes from LimitNOFILE=16384 in the shipped logstash.service unit.

# Classify what the FDs actually are
ls -l /proc/$PID/fd | awk '{print $NF}' | sed 's/[0-9]*$//' | sort | uniq -c | sort -rn | head -20

This buckets FDs by target: sockets, regular files under specific paths, pipes, and anon inodes. The bucket that is growing is your driver.

# Socket state breakdown: look for CLOSE_WAIT accumulation
ss -tanp | grep "pid=$PID" | awk '{print $1}' | sort | uniq -c

A large and growing CLOSE_WAIT count means the remote side closed the connection but Logstash never closed its end. That is a leak signature, and it typically points at an output plugin or a proxy in front of the destination.

# Count files matched by a broad file input path (example pattern)
find /var/log -name '*.log' -type f 2>/dev/null | wc -l
# Output retry/reconnect activity in the Logstash log
grep -Ei '(retry|reconnect|error|exception|timeout|refused)' /var/log/logstash/logstash-plain.log | tail -n 100
# Baseline trend: sample twice, 10 minutes apart
for i in 1 2; do
  curl -sS http://127.0.0.1:9600/_node/stats/process | grep open_file_descriptors
  sleep 600
done

Two samples say nothing about a multi-week leak, but during an active incident a rising delta over 10 minutes confirms active growth rather than a static high baseline.

How to diagnose it

  1. Confirm the ratio. Pull open_file_descriptors and max_file_descriptors from /_node/stats/process. Above 80% sustained, treat this as the incident, not a side note.

  2. Check the peak. peak_open_file_descriptors tells you whether you already brushed the ceiling. A peak near max with a current value well below it points at churn (spikes) rather than a leak (ratchet).

  3. Establish the trend shape. Look at FD count over the longest window you have. Three shapes: flat-and-high (sizing), slow monotonic climb (leak), sawtooth spikes tied to downstream events (reconnection churn).

  4. Classify the FDs. Use the /proc/<pid>/fd breakdown above. Sockets dominating means network paths: inputs (Beats, TCP, HTTP) or outputs. Regular files under log directories means the file input. Files under the queue directory means PQ pages.

  5. If sockets: check state and peer. Group ss -tanp output by state and by remote address. CLOSE_WAIT against your Elasticsearch or load balancer addresses, growing over time, is the classic reconnection-leak combination. Correlate the growth periods with output errors in the Logstash log.

  6. If regular files: count what the input sees. Compare the number of files your file input path patterns match against the number of open file FDs. The file input keeps one FD per actively tailed file, bounded by its max_open_files setting (default 4095) with a sliding window: it can track more files than that, but only keeps that many open at once. A “Reached open files limit” warning from filewatch in the log means the input itself is saturated and files are waiting to be opened. Idle files are closed after close_older (default 1 hour), which frees FDs for hotter files. If most matched files are constantly written, nothing goes idle and the window stays full.

  7. If churn: find the trigger. Reconnection storms follow downstream events: Elasticsearch restarts, load balancer idle timeouts cutting connections, TLS or auth failures. Line up the FD spike times with output error timestamps and with events on the destination side.

  8. Compute runway. (max - open) / growth_rate using a growth rate measured over days, not minutes. This tells you whether you have hours or weeks, and therefore whether the fix is urgent operational work or a planned change.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
process.open_file_descriptorsThe primary count. Trend shape distinguishes leak from churn from sizingAny climb that does not return to baseline after load drops
process.max_file_descriptorsThe ceiling the ratio is computed againstLower than you assumed (service manager overrides defaults)
process.peak_open_file_descriptorsReveals near-misses that current values hidePeak within 10% of max
Ratio open/maxThe alertable numberAbove 80% sustained
FD growth rateInput to the runway calculationAny sustained positive rate measured over days
Socket states (CLOSE_WAIT count)Early leak evidence before the FD total is alarmingGrowing count against one downstream peer
Output error/retry log linesReconnection churn triggerError bursts that coincide with FD spikes
File input saturationInput-side FD exhaustionfilewatch “Reached open files limit” warnings

Two traps worth naming. First, pipeline stats reset on config reload, but process FD stats do not; FD graphs survive reloads, which makes them reliable for long-window trend analysis. Second, do not poll the API faster than about every 10 seconds on a loaded instance; the stats endpoint shares the JVM with the pipeline.

Fixes

Fix the leak

A true leak (FDs opened and never closed) is almost always inside a plugin or its interaction with a proxy, not something you can tune away. Historical examples include Beats input connections leaked under backpressure, HTTP input connections leaked when clients outpaced the pipeline, and Elasticsearch output connections accumulating in CLOSE_WAIT when a proxy in the middle closed them first. The pattern to look for is FD growth that tracks connection events, not data volume.

  • Upgrade the affected plugin or Logstash version. Several FD leaks in the Beats input, HTTP input, and Elasticsearch output were fixed in specific releases. If your FD growth matches a known pattern (for example, growth only under backpressure), check the plugin’s CHANGELOG (and the versioned plugin documentation) to confirm the fix version for your deployment before upgrading.
  • Remove the middlebox trigger. If CLOSE_WAIT sockets pile up against a load balancer or proxy that idles connections out, align its idle timeout with the output plugin’s connection reuse behavior, or have the output side close connections first. The side that closes first avoids the lingering half-closed state.
  • Restart as mitigation, not as a fix. A restart releases every FD and resets the clock. With a persistent queue, in-flight events survive. With a memory queue, queued events are lost. If the leak took weeks to build, a restart buys you weeks, but schedule the real fix.

Tame the file input

  • Narrow the path patterns. Replace broad wildcards like /var/log/**/*.log with explicit directories per pipeline. Split unrelated log trees into separate pipelines so one hot tree does not starve another.
  • Size max_open_files deliberately. The default 4095 is independent of the OS limit. Raising the systemd limit without raising max_open_files does nothing for a saturated file input; raising max_open_files without OS headroom just moves the ceiling. Set both, and keep the input’s value comfortably below the process limit.
  • Tune close_older for your write pattern. If files go idle quickly, a shorter close_older frees FDs sooner. If everything is hot, closing and reopening just adds churn; the real fix is fewer matched files or more headroom.
  • Consider read mode for batch workloads. In read mode the input closes files at EOF instead of holding them open, which suits content-complete files. In the default tail mode, files stay open as long as they are active. See the file input and sincedb guide linked below for the re-read and duplication tradeoffs.

Stop reconnection churn

  • Fix the downstream first. Reconnection storms are a symptom. If Elasticsearch is flapping, rejecting, or being restarted, the FD churn stops when the destination stabilizes.
  • Watch retry behavior during incidents. Output retries preserve data but multiply connection attempts. During a prolonged downstream outage, FD spikes are expected; the danger is a spike that never decays after recovery, which converts churn into a leak.
  • Check authentication and TLS. Auth failures and handshake errors cause rapid connect-fail-retry loops that churn FDs without ever establishing a working connection. The log patterns in the quick checks section catch this.

Raise the limit (with eyes open)

Raising LimitNOFILE in a systemd override (/etc/systemd/system/logstash.service.d/override.conf, then systemctl daemon-reload and a restart) is legitimate when the workload genuinely needs more FDs, for example tailing thousands of files by design. Note that the restart is disruptive: with a memory queue, in-flight events are lost. A limit increase is not a fix for a leak; it only lengthens the runway. Always pair it with growth-rate monitoring so you know whether the new ceiling is being consumed.

Prevention

  • Alert on the ratio, not the absolute count. Threshold: open_file_descriptors / max_file_descriptors above 80% sustained. Absolute counts mean different things on different deployments; the ratio travels.
  • Alert on the trend, not just the level. A leak at 30% utilization is invisible to a level alert for weeks. A growth-rate alert (sustained positive slope over days, normalized for traffic) catches leaks while runway is still long.
  • Compute and track runway. (max - open) / growth_rate turns FD monitoring from a binary alarm into a capacity signal. A runway under 30 days deserves a ticket; under 7 days, treat it as urgent.
  • Baseline FDs per config change. File input path changes, new Beats sources, and new outputs all move the legitimate baseline. Record the expected FD footprint when configs change so drift is visible.
  • Watch socket states during downstream incidents. CLOSE_WAIT growth during an outage is normal in small amounts; a count that keeps climbing after the downstream recovers is your leak alarm.

How Netdata helps

  • Netdata charts open_file_descriptors against max_file_descriptors continuously, so the ratio and the long slow leak trend are visible on one graph instead of buried in periodic manual polls.
  • Long retention at per-second granularity is what makes week-scale leaks diagnosable: you can zoom from a month-long climb down to the hour the slope changed and line it up with a deploy or a downstream incident.
  • peak_open_file_descriptors alongside the current value surfaces near-misses from reconnection spikes that a current-value-only check would miss.
  • Correlating the FD curve with process CPU, JVM heap, and pipeline throughput on the same host timeline separates the three drivers: leak (FDs up, everything else flat), churn (FD spikes aligned with output errors), and sizing (FDs tracking input rate).
  • Alerting on the ratio with a sustained-duration condition gives you the >80% ticket signal the playbook recommends, without paging on brief reconnection spikes.