The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / fluentd / fluentd-process-not-running ▌

Operations Guides

Fluentd process not running: the log pipeline is dead and the host has gone dark

The Fluentd process is gone. No logs are being collected, parsed, buffered, or forwarded from this host. Your downstream systems (Elasticsearch, S3, a SIEM, an aggregator tier) are now receiving nothing from here, and most of them will not tell you that. Log pipelines fail silently at the consumer side: the absence of data looks identical to a quiet host.

This is a full observability blackout for the host and a page-worthy condition in most environments. The one nuance: brief absences are normal. Fluentd restarts during config reloads, rolling updates, package upgrades, and container rescheduling. Gate your page on sustained absence (more than about 2 minutes) so a routine restart at 3 a.m. does not wake anyone up.

This guide covers confirming the process is really dead (not just hidden behind a live supervisor), finding the crash reason, and recovering without losing more data than necessary.

What this means

A default Fluentd installation runs two Ruby processes: a supervisor and a worker. The supervisor owns the worker lifecycle and restarts it after a crash. In multi-worker mode (workers N in <system>), there are N worker processes under one supervisor. This creates a failure mode that fools naive checks:

  • The supervisor can be alive while every worker is dead. systemctl status td-agent reports “active” because systemd tracks the supervisor. Meanwhile, zero events are flowing.
  • A worker killed by the OOM killer may be restarted by the supervisor within seconds. The outage is real but invisible unless you track restarts or crash log lines.
  • In Kubernetes, CrashLoopBackOff produces rapid up/down oscillation. Sustained absence shows up as repeated brief absences rather than one long gap, so a naive “down for 2 minutes” check may never fire even though the pipeline is effectively dead.

The crash reason matters for data loss. File-backed buffers survive a restart and replay on recovery. Memory-backed buffers are gone the instant the process dies. If in_tail position files are intact, collection resumes where it left off; if the pos_file was lost or corrupted, you get duplicates or gaps.

flowchart TD
  A[Fluentd process check fails] --> B{Supervisor alive?}
  B -- No --> C[Check systemd unit and journal for boot failure]
  B -- Yes --> D{Workers alive?}
  D -- No --> E[Check supervisor log for worker exit signal]
  D -- Yes --> F[Process exists: check monitor_agent responsiveness]
  C --> G[Config error or plugin load failure at boot]
  E --> H{Exit signal}
  H -- SIGKILL --> I[OOM killer: check dmesg]
  H -- SIGSEGV --> J[Crash in Ruby or C extension: check core logs]
  H -- non-zero exit code --> G
  F -- Unresponsive --> K[Process hung, not dead: treat as separate incident]

Common causes

CauseWhat it looks likeFirst thing to check
OOM killProcess vanished or restarted; supervisor log shows worker finished with SIGKILL; possible crash loop under memory pressuredmesg | grep -i oom and process RSS history
Config syntax error at bootService fails to start or enters a restart loop right after a config change or package upgradejournalctl -u td-agent for parse errors since the last deploy
Fatal plugin exception at bootWorker exits immediately with a non-zero code; supervisor restarts it; repeatsStartup logs for LoadError or plugin load failures
Poison pill crash loopProcess oscillates up/down within seconds; CPU spike before each crash; crash recurs at the same pos_file positionFluentd error log for parse exceptions or regex backtracking
Supervisor alive, all workers deadsystemctl status shows active, but no worker processes exist and nothing flowsWorker process count via pgrep, per-worker monitor_agent ports

Quick checks

Run these read-only checks in order. They take under a minute and localize the failure.

# 1. Service-level status (adjust unit name for your package)
systemctl status td-agent   # td-agent package
systemctl status fluentd    # fluent-package

# 2. Actual process list: you should see a supervisor AND at least one worker
pgrep -af fluentd

# 3. For td-agent: two ruby processes expected by default (supervisor + worker)
ps w -C ruby -C td-agent --no-heading

# 4. Kernel OOM events
dmesg | grep -i oom | tail -20

# 5. Recent unit logs: crash reason, restart loop, config errors
journalctl -u td-agent --since "30 minutes ago" | tail -50

# 6. Fluentd's own log: worker exit lines and plugin errors
grep -iE "finished unexpectedly|LoadError|uncaught|error" /var/log/td-agent/td-agent.log | tail -30
# fluent-package path: /var/log/fluent/fluentd.log

# 7. Monitor agent responsiveness (only if configured; default port 24220)
curl -s -o /dev/null -w "%{http_code}\n" --max-time 5 http://localhost:24220/api/plugins.json

Interpretation notes:

  • Steps 2 and 3 are the checks that matter. A live supervisor with no workers means step 1 lies to you.
  • Step 6 looks for the supervisor’s “Worker N finished unexpectedly with signal …” lines. SIGKILL almost always means the OOM killer. SIGSEGV means a crash in Ruby or a C extension. A clean non-zero exit code points at config or plugin failure.
  • Step 7 distinguishes “dead” from “hung”. A process that exists but does not answer the monitor agent within 5 seconds is a different incident (deadlock, GC storm, blocked event loop), not this one.

How to diagnose it

  1. Confirm the blast radius. Is this one host or many? If a DaemonSet or a whole aggregator tier lost Fluentd at once, suspect a shared cause: a bad config push, a package upgrade, a destination outage that cascaded into OOM, or a certificate expiry. Check a second host before debugging the first in depth.

  2. Establish the process topology. Count supervisors and workers. Expected: one supervisor plus workers N workers (default N=1). If the supervisor is missing entirely, the failure is at the unit level (systemd gave up restarting, or the boot sequence never got there). If the supervisor is present but workers are missing, read the supervisor log for worker exit lines.

  3. Get the crash reason from the journal, not just Fluentd’s log. A Fluentd worker that dies from a Ruby segfault or an OOM kill may never write its final error to /var/log/td-agent/td-agent.log. The kernel and systemd records survive: journalctl -u td-agent for exit codes and restart counts, dmesg for OOM kills. Check whether systemd has been restarting the unit repeatedly; that masks the outage as flapping rather than a clean “down”.

  4. Correlate with change. Did a config deploy, package upgrade, or logrotate run happen just before the death? Config syntax errors and plugin load failures show up immediately at boot and produce a tight restart loop. Plugin load errors look like LoadError or “cannot load” lines in the journal at startup.

  5. Check the resource ceiling. If the evidence points at SIGKILL/OOM: was it the host OOM killer or a container memory limit? In Kubernetes, kubectl describe pod shows OOMKilled as the last state and the restart count. On a bare host, dmesg shows which process was chosen and the memory state at the time. Ruby RSS grows and plateaus; a plateau is normal, a monotonic rise is a leak or unbounded memory-backed buffers.

  6. Rule out the poison pill. If the process restarts and dies again within seconds, repeatedly, at roughly the same input position, suspect a malformed log line crashing a parser. The crash recurs because in_tail resumes from the saved pos_file position and re-reads the same bad line. Look for parse exceptions or regex errors in the Fluentd log immediately before each death.

  7. Verify recovery actually restored flow. After the process is back, “running” is not enough. Confirm the monitor agent returns 200, that input emit_records is incrementing, and that output write_count is incrementing. A live process with a full buffer and a dead destination is still a dark host.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Process liveness (supervisor + worker count)Dead process means zero collection and forwarding from the hostAny worker absent for more than 2 minutes sustained
systemd restart count / pod restart countAuto-restart hides outages; the restart count is the only recordRestart count climbing between checks
Kernel OOM eventsExplains SIGKILL deaths and predicts recurrenceNew OOM entries naming ruby/td-agent/fluentd
Process RSS vs limitRuby leaks and memory-backed buffer growth end in OOM killsMonotonic RSS growth, or RSS above 80% of the container limit
Monitor agent HTTP statusConfirms the event loop is functional, not just the PIDNon-200 or timeout over 5s when monitor_agent is configured
Input/output emit_records ratesA live process that emits nothing is still a blackoutBoth rates flat after a restart that should be replaying buffers
buffer_total_queued_size after recoveryFile-buffer replay should drain; a stuck queue means the destination is downReplay backlog not draining within minutes of startup

Two gating rules for paging:

  • Page on sustained absence, not instantaneous absence. More than 2 minutes covers restarts, rolling updates, and container rescheduling without delaying real pages materially.
  • In Kubernetes, also page on CrashLoopBackOff or a restart count climbing past a handful within a short window. Sustained-absence logic alone can miss a tight crash loop, because the process is technically “up” for a few seconds each cycle.

Fixes

OOM kill

  1. Confirm with dmesg or the pod’s last state, then decide whether the limit is wrong or the usage is wrong.
  2. If buffers are memory-backed, switch to file-backed (@type file in the buffer section). This removes the largest memory consumer and makes buffered data survive the next crash.
  3. Increase chunk_limit_size so fewer, larger chunks exist. Millions of small chunks create millions of Ruby objects and heavy GC pressure.
  4. Raise the container or systemd memory limit with headroom above the observed RSS plateau. Ruby fragmentation means the plateau sits higher than the configured buffer sizes suggest; leave at least 20-30% above normal RSS.
  5. If RSS grows monotonically even with file buffers and stable throughput, suspect a plugin leak. Check whether the growth correlates with a recent plugin addition or upgrade.

Tradeoff: file-backed buffers cost disk I/O and need disk headroom on the buffer path. That trade is almost always worth it for crash durability.

Config syntax error or plugin load failure at boot

  1. Pull the exact error from journalctl -u td-agent or the Fluentd log. Syntax errors name the file and line.
  2. Roll back the offending config change if it came from a deploy. Do not leave the unit crash-looping while you edit in place on the host.
  3. Validate before the next reload: fluentd --dry-run -c /etc/td-agent/td-agent.conf parses the config without starting the pipeline. Make this a CI or deploy-pipeline step so the next bad config never reaches a host.
  4. For LoadError-style failures, check for a missing gem or an incompatible plugin version after a package upgrade. Plugin load errors at startup appear in the journal even when the unit looks merely “flapping”.

Poison pill crash loop

  1. Confirm by matching crash times to the same pos_file position and to parse exceptions in the log.
  2. Stop the loop by moving Fluentd past the bad line: advance the position in the pos_file, or temporarily move the offending source log file aside. This is disruptive and skips data; do it deliberately. The bad line is lost either way, so the choice is between losing one line and losing everything behind it.
  3. Fix the parser: eliminate catastrophic regex backtracking, cap line sizes, or add a filter that excludes the malformed pattern.
  4. Set read_lines_limit on in_tail to bound per-cycle processing.

Supervisor alive, workers dead

  1. Treat this as a monitoring gap first: your process check was watching the parent PID. Fix the check to count workers.
  2. Restart the unit (systemctl restart td-agent) to get a clean supervisor/worker pair. This is disruptive to any in-flight in-memory state, but the workers are already dead, so there is little left to lose. If workers die again immediately, you have one of the causes above; the supervisor restart loop is a symptom, not the disease.
  3. If you run with --no-supervisor under an external supervisor (systemd, runit), verify that the external supervisor is actually restarting the worker and that you have not silently lost the auto-restart behavior the built-in supervisor would have provided.

After any recovery

  • Expect a brief throughput spike from file-buffer replay and a burst of duplicates. File-backed buffers replay unflushed chunks by design; exactly-once delivery is not guaranteed. Do not page on the replay burst.
  • If buffers were memory-backed, accept the gap and quantify it: compare input and output emit_records at the destination for the outage window.
  • Verify the pos_file survived. If it was on volatile storage (tmpfs, container ephemeral layer), in_tail will either re-read everything (duplicates, if read_from_head true) or skip history (gaps).

Prevention

  • Worker-aware liveness checks. Alert on the expected worker process count, not the unit status and not the parent PID. In multi-worker mode, check each worker’s monitor_agent port: each worker binds its own port by adding its worker id to the configured port (worker 0 is 24220, worker 1 is 24221, and so on).
  • Track restart counts. The single most reliable early signal of a crash loop. A unit that restarted 5 times in an hour is telling you the next OOM or poison pill is already scheduled.
  • File-backed buffers in production. Crash durability plus lower memory pressure. Verify with /api/plugins.json?with_config=true that each output’s buffer section actually says @type file.
  • Config validation in the deploy pipeline. Dry-run the config before it ships. Boot-time syntax errors are the cheapest class of Fluentd outage to eliminate entirely.
  • Memory headroom discipline. Set container limits at least 20-30% above the observed RSS plateau, and alert on the RSS trend, not the absolute value. Plateaus are normal for Ruby; slopes are not.
  • Persistent pos_file storage. Keep position files off tmpfs and off the container ephemeral filesystem so restarts resume instead of replaying or skipping.
  • Watch Fluentd’s own log. The log pipeline’s own error log is routinely the one log nobody collects. Ship it somewhere that survives the host going dark.

How Netdata helps

  • Process liveness and restart tracking. Netdata’s process and systemd-unit monitoring show the Fluentd process disappearing and, critically, the restart count climbing, which is what turns a masked crash loop into a visible signal.
  • Fluentd collector via monitor_agent. Netdata polls the monitor_agent endpoint and charts buffer queue length, retry counts, and emit rates per plugin, so you can confirm after recovery that replay is draining and flow is actually restored rather than just “process is up”.
  • Memory correlation. Per-process RSS alongside cgroup memory limits shows the monotonic growth trend that precedes an OOM kill, which is the difference between preventing the next crash and explaining the last one.
  • OOM and log context on one timeline. Kernel OOM events, RSS, Fluentd restarts, and buffer depth land on the same dashboard timeline, so the SIGKILL-to-OOM-to-restart chain takes minutes to confirm instead of a dmesg archaeology session.
  • Sustained-absence alerting. Netdata alert hysteresis lets you require the process to be absent for a sustained window before paging, matching the 2-minute gating rule and filtering out routine restarts.