The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-pos-file-corruption ▌

Operations Guides

Fluentd pos_file corruption: duplicates and gaps after a restart

You restarted Fluentd (a deploy, an OOM kill, a pod reschedule) and now one of two things is wrong downstream: the same log lines appear twice, or a window of logs never arrived. Both symptoms point at the same component: the in_tail position file.

The pos_file is how in_tail remembers where it stopped reading. It is a plain text file with one line per tailed file, recording the file path, a byte offset in hexadecimal, and an inode number in hexadecimal. There are no checksums and no integrity metadata. On startup, Fluentd reads it and trusts it completely.

If the pos_file is missing, stale, corrupt, or was sitting on tmpfs, the restart degrades to one of two bad outcomes: re-reading from offset 0 (duplicates) or resuming from an offset past the current end of file (a permanent gap). This guide walks through confirming the symptom, isolating which failure you have, recovering with the least collateral damage, and preventing recurrence.

What this means

While running, in_tail tracks a byte offset per watched file and persists offsets to the pos_file. On startup, for each pos_file line, Fluentd checks whether the recorded inode still matches the file on disk. If it does, reading resumes at the recorded offset. Everything that can go wrong after a restart is a corruption of one of those three fields, or the absence of the line entirely:

  • Pos_file missing (tmpfs, deleted, fresh container filesystem): there is no entry for the file. With read_from_head true, the whole file is re-read: duplicates of every line still on disk, plus a CPU and memory spike while the backlog is parsed. With read_from_head false, reading starts at the current end of file: every line written before startup is skipped, which is a gap.
  • Corrupt or stale offset: an offset of 0 forces a full re-read (duplicates). An offset beyond the current end of file means the lines between the real end and the recorded offset are never read (gap).
  • Inode mismatch: the recorded inode no longer matches the current file, which means rotation handling failed. The current file is treated as untracked or the wrong file is followed.

One distinction before you start: a brief burst of duplicate deliveries right after a restart is normal with file-backed buffers, because unflushed chunks are replayed. Buffer replay duplicates are bounded to the chunks that had not been flushed. Pos_file-driven duplicates cover entire files and keep arriving until the re-read catches up. The volume and duration of the duplicate window tells you which mechanism you are looking at.

flowchart TD
  A[Fluentd restarts] --> B{pos_file intact?}
  B -- yes --> C{inode matches file on disk?}
  C -- yes --> D[Resume at recorded offset: clean]
  C -- no --> E[Rotation follow failed: skip or re-read]
  B -- no --> F{what is recorded or missing}
  F --> G[No entry or offset 0: full re-read, duplicates]
  F --> H[Offset past current EOF: lines skipped, gap]

Common causes

CauseWhat it looks likeFirst thing to check
pos_file on tmpfs or ephemeral container storageDuplicates across all tailed files after every reboot or pod reschedulefindmnt -T <pos_file path> and the volume type in the pod spec
pos_file deleted or not writableDuplicates if read_from_head true, gap from EOF if falseDoes the file exist, and is its mtime advancing while Fluentd runs
Corrupt pos_file (partial write, disk error)Per-file duplicates or gaps; truncated or merged lines in the filecat the pos_file and look for malformed lines
Inode mismatch after rotationOne file skipped or re-read; correlates with rotation timeCompare the inode in the pos_file with stat -c %i on the live file
copytruncate rotation raceA small duplicate or gap window at every rotationCheck the logrotate or application rotation method
Stale offsets after SIGKILL or OOM killDuplicates of lines written between the last pos_file write and the killCorrelate the duplicate window with the kill time
One pos_file shared between multiple in_tail sourcesInterleaved, corrupted lines and erratic resume positionsCount pos_file directives across the config; each source needs its own

The in_tail documentation explicitly warns against sharing one pos_file between in_tail configurations: concurrent updates from two plugins corrupt the content.

Quick checks

All of these are read-only.

Find where the pos_file lives. Paths vary by package: td-agent uses /etc/td-agent/, fluent-package uses /etc/fluent/, Kubernetes deployments usually mount config from a ConfigMap.

# Locate pos_file directives
grep -rn "pos_file" /etc/td-agent/ /etc/fluent/ 2>/dev/null

Check whether the pos_file is on durable storage. If the filesystem type is tmpfs, the file is lost on every reboot; in a container, the default writable layer is lost on every reschedule.

# Identify the filesystem backing the pos_file
findmnt -T /var/log/td-agent/td-agent.pos
df -T /var/log/td-agent/td-agent.pos

Inspect the file itself. A healthy pos_file has one clean line per tailed file and an mtime that advances while Fluentd runs. Truncated lines, merged hex fields, or lines referencing files that no longer exist are all signs of trouble.

# Check freshness and contents
ls -la /var/log/td-agent/td-agent.pos
cat /var/log/td-agent/td-agent.pos

Compare each pos_file entry against the file it references. The recorded inode should match the live inode, and the recorded offset should sit near (slightly behind) the current file size.

# Show live state for every file referenced in the pos_file
awk '{print $1}' /var/log/td-agent/td-agent.pos | while read -r f; do
  if [ -f "$f" ]; then
    stat -c 'inode=%i size=%s %n' "$f"
  else
    echo "MISSING $f"
  fi
done

The offset is stored in hexadecimal. Convert it to decimal before comparing with the size from stat:

# Convert a recorded hex offset to decimal bytes
printf '%d\n' 0x0000000000014bfa

Check the input-side counters on the monitor agent (default port 24220). A post-restart spike in input emit_records confirms a re-read; a drop to zero while source files keep growing confirms skipped files.

# Total input emit_records (cumulative counter, derive the rate)
curl -s http://localhost:24220/api/plugins.json | \
  jq '[.plugins[] | select(.plugin_category=="input") | .emit_records // 0] | add'

On Fluentd versions before v1.19.0, input emit_records is always 0 unless enable_input_metrics true is set in <system>. If you get zeros everywhere, check that first.

Check the in_tail-specific counters. rotated_file_count (v1.14.1+) tells you rotation detection is firing; tracked_file_count (v1.19.0+) tells you how many files are currently watched.

# in_tail tracking and rotation counters
curl -s http://localhost:24220/api/plugins.json | \
  jq '.plugins[] | select(.type=="tail") | {id: .plugin_id, tracked: .tracked_file_count, rotated: .rotated_file_count}'

If in_tail is pinned to a specific worker with <worker N> (it does not support multi-worker), query that worker’s monitor_agent port; worker 0 is 24220, worker 1 is 24221, and so on.

Establish when the process actually started, and whether it was killed rather than stopped cleanly. An OOM kill or SIGKILL means the pos_file was last flushed some time before death, so offsets are stale by that window.

# Process start time and recent OOM kills
ps -o lstart= -p $(pgrep -f fluentd | head -1)
dmesg | grep -i oom | tail -5

How to diagnose it

  1. Confirm the symptom at the destination. Query your log store for a known line from the affected host around the restart. Duplicates mean a re-read; a missing time window means skipped data. Do not trust Fluentd-side metrics alone for this.
  2. Fix the timeline. Get the restart time from ps or the pod’s restart count, and the kill cause from dmesg or the Fluentd log. An unclean kill implies stale offsets; a clean stop implies the pos_file should have been flushed on shutdown.
  3. Check whether the pos_file survived. Existence, mtime, and the filesystem it sits on. If it is on tmpfs, an emptyDir, or the container’s writable layer, the diagnosis ends here: positions were lost, and the restart behavior is governed by read_from_head.
  4. For each affected file, compare entry to reality. Inode in the pos_file versus stat -c %i, recorded offset versus current size. Offset at 0 with a large live file means full re-read. Offset beyond the live size means a gap. Inode mismatch means the rotation follow failed.
  5. Confirm with input emit_records. A sharp spike starting at the restart time is the re-read in progress. A flat line while stat shows the source file growing means the file is not being read at all.
  6. Correlate with rotated_file_count. If the symptom window lines up with a rotation event and the rotation counter did not increment as expected, rotation handling, not the restart itself, is the primary cause.
  7. Choose a recovery path based on whether you are containing duplicates (usually tolerable, dedup downstream) or recovering a gap (only possible if the skipped bytes still exist in a file Fluentd can be pointed at).

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Input emit_records rateStep changes are the primary pos_file symptom: spike means re-read, drop means skipped filesSharp spike beginning exactly at restart time; drop to zero while source files grow
rotated_file_count (v1.14.1+)Confirms rotation is detected and followedStops incrementing on the expected rotation schedule, or increments align with duplicate windows
tracked_file_count (v1.19.0+)Shows how many files are actually being watchedDrops below the expected file count for the host
pos_file mtime and size (host level)A pos_file that is not being updated means positions are not being savedStale mtime while Fluentd runs; missing file after restart
Inode match (host level)The exact mechanism behind rotation follow failuresRecorded inode differs from the live file’s inode
Input versus output emit_records balanceQuantifies the scale of duplication or loss over a windowInput rate far above the host’s baseline after a restart

No API metric directly exposes pos_file integrity; this is a known blind spot. The host-level checks above are the only way to cover it, which is why mature setups include pos_file freshness and inode validation as scripted checks.

Fixes

Recover a lost or corrupt pos_file

Warning: the blunt recovery, deleting the pos_file and restarting, re-reads or skips every tailed file on the host. With read_from_head true, every file is re-read in full: massive duplicates and a startup CPU and memory spike proportional to on-disk log volume. With read_from_head false, reading starts at the current end of each file: everything written before startup is skipped. Since v1.14.3 (PR #3542, also stated in the in_tail documentation), files discovered after startup are read from head regardless; read_from_head false only changes startup behavior, which is exactly the case you are in.

The surgical option is usually better. Stop Fluentd first, because a running in_tail periodically rewrites the pos_file and will clobber your edit. Then edit the hex offset on the affected file’s line: set it to 0 to force a full re-read of that one file (accepting duplicates of just that file), or to a known-good byte offset you computed from the file’s current content. Restart Fluentd and watch input emit_records for the expected step.

For gaps: if the skipped bytes are still in the live file, the same edit recovers them by moving the offset backwards. If the file was rotated away, the data only exists in the archived rotated files, and in_tail will not touch them unless your path pattern covers them. In that case, recovery means re-ingesting the archived files through a separate path or accepting the gap.

Move the pos_file to durable storage

This is the fix for the most common root cause. The pos_file must survive reboots, container restarts, and pod reschedules. On hosts, put it on a persistent local filesystem, not tmpfs. In Kubernetes DaemonSets, mount a hostPath volume or a PVC for the pos_file directory; an emptyDir is lost when the pod leaves the node, and a memory-backed emptyDir is tmpfs. Until this is fixed, every restart reproduces the incident.

Fix rotation follow failures

  • Prefer rename/create rotation over copytruncate. copytruncate has an inherent race: lines written between the copy and the truncate can be missed or read twice.
  • Set follow_inodes true when using wildcard paths. Available since v1.12.0, this tracks files by inode across rotation instead of by path, which is what prevents re-reads when the path pattern matches both the old and new file. Residual duplication risk in the rotate_wait window has been documented even after later fixes, so verify at the destination after rotations.
  • Tune rotate_wait (default 5s) and refresh_interval. Versions before v1.16.3 could silently stop tailing a file when rotate_wait exceeded refresh_interval; v1.16.2 and v1.16.3 fixed the watcher-stall bugs for both follow_inodes modes. If you are on an older release, upgrading is the real fix.
  • Kubernetes note: container log paths are symlinks, and the pos_file tracks the symlink target’s inode, which changes on rotation. An inode mismatch after rotation on a node is expected to resolve via follow_inodes; if it persists, the watcher is stuck.

Bound pos_file growth and avoid known bugs

When tailing many files with dynamic paths, the pos_file grows until restart because entries for unwatched files are only cleaned at startup. pos_file_compaction_interval (v1.9.2+) periodically removes unwatched, unparsable, and duplicated lines. All supported releases write the current line format (path, offset, inode as hex fields), but if you inherit an ancient td-agent pos_file, inspect it for malformed lines such as merged or truncated hex fields and compact or recreate the file. There is also an open issue (fluent/fluentd #4843, still open as of v1.19.x) where limit_recently_modified can cause an all-ones hex offset to be recorded for unwatched files, forcing a full re-read from head when the file is re-watched; if you use that parameter and see periodic duplicate bursts, this is a candidate cause.

Prevention

  • Durable pos_file storage. Positions must survive reboots and pod reschedules, or every restart is a duplicates-or-gaps incident.
  • One pos_file per in_tail source. Sharing corrupts the file and produces erratic resume positions.
  • Rotation method audit. Use rename/create where possible, follow_inodes true with wildcards, and rotate_wait long enough for your rotation tooling.
  • Version floor. Run v1.16.3 or later to carry the in_tail watcher fixes; add pos_file_compaction_interval if you tail dynamic paths.
  • Pos_file integrity checks. Scripted host-level checks for mtime freshness and inode match catch the failure mode no API metric exposes.
  • Post-restart input rate alerting. An alert on input emit_records deviation from baseline turns every future occurrence into a detection instead of a downstream surprise.

How Netdata helps

  • Netdata collects the monitor_agent counters (emit_records, rotated_file_count, tracked_file_count) at per-second granularity, so the restart-time step change is visible instead of averaged away.
  • Correlating the input emit_records spike with process restarts and OOM kills on the same host separates a pos_file re-read storm from a genuine application log storm.
  • Comparing input and output emit rates over a rolling window quantifies how much data was duplicated or lost, which decides whether downstream dedup or gap recovery is worth doing.
  • Tracking rotated_file_count against the expected rotation schedule catches follow failures before they compound across rotations.
  • Process uptime, OOM events, and file descriptor usage sit next to the Fluentd metrics in one view, so the kill-cause-to-symptom link takes minutes instead of a log archaeology session.