The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-context-deadline-exceeded ▌

Operations Guides

NATS context deadline exceeded: JetStream publish and request timeouts

Your application logs fill up with nats: context deadline exceeded. It shows up on js.Publish(), on js.StreamInfo(), on consumer fetches, sometimes on nats CLI commands like nats stream view. The NATS server is running, /healthz returns ok, and core NATS pub/sub traffic is flowing fine. Yet every JetStream operation hangs until the client’s context fires.

This error is a client-side deadline expiring. The client sent a JetStream API request and no response came back before the deadline. The request is not being rejected; it is not being answered at all. Somewhere between the client and the JetStream subsystem, the request is stuck, and the server is usually still healthy enough to look innocent.

The fix is almost never “increase the client timeout.” The timeout is telling you the server is slow to ack publishes or slow to answer API calls, and there are a small number of server-side causes that produce exactly this shape: JetStream disk I/O stalls, Raft leader instability, meta-cluster elections, an overloaded server, or an API backlog the server cannot drain.

What this means

Every JetStream operation, publish-ack, consume, stream info, consumer create, is a request-reply exchange over $JS.API.> subjects. The client library wraps the call in a context with a deadline and waits for the reply. When the reply does not arrive in time, the context fires and you get context deadline exceeded. The client has no visibility into why; it only knows nobody answered.

Server-side, there are four ways a request gets stuck long enough to blow the deadline:

  1. The publish-ack is waiting on storage. A JetStream publish is not acked until the message is written to the stream storage or Raft write path. If storage write latency — including fsync when configured for durability — spikes, every ack can stretch out and publishes can time out.
  2. The write needs Raft consensus and there is no stable leader. In clustered JetStream, writes require the stream’s Raft group, and metadata operations require the meta group. During an election there is no leader, so proposals queue and clients wait.
  3. The API backlog is growing faster than the server can drain it. /jsz exposes api.inflight: the number of in-progress JetStream API requests. When inflight climbs and stays high, the server is slow to process API requests, usually because of the two causes above or plain CPU saturation.
  4. The server or the network path is shedding the request entirely. Stale connections and broken pipes on the server side present to the client as a deadline, not an error.

This is not the same failure as a core NATS request timeout. A core request that times out or returns no responders available means the responder application was slow or absent; the broker did its job. With JetStream timeouts, the broker itself is the slow responder. If you are seeing plain request-reply timeouts without JetStream involvement, see NATS no responders available for request instead.

flowchart TD
  A[context deadline exceeded
on JetStream call] --> B{api.inflight high
and sustained?} B -->|yes| C{meta_cluster leader
changing frequently?} B -->|no| D{Server overloaded?
CPU, stale connections} C -->|yes| E[Raft instability:
check route RTT, GC pauses,
disk latency on WAL path] C -->|no| F{Storage near limits?} F -->|no| G[Disk I/O stall:
iowait, fsync latency,
backup or noisy neighbor] F -->|yes| H[Storage exhaustion:
consumer stall or retention] D -->|yes| I[CPU saturation or
connection path issues] D -->|no| J[Per-stream Raft group issue:
check /raftz for that stream]

Common causes

CauseWhat it looks likeFirst thing to check
JetStream disk I/O stallPublish-acks stretch out, api.inflight high, storage usage NOT near limits, OS iowait elevatediostat on the JetStream storage path; look for backups, snapshots, noisy neighbors
Raft leader flapping (meta or per-stream)api.errors rising, leader name keeps changing, operations fail during elections/jsz meta_cluster leader stability and replica current/offline/lag
Single stream Raft group brokenOnly operations on one stream time out; rest of JetStream fine/raftz for that stream’s group; per-stream info
API backlog / overloaded serverapi.inflight sustained high, CPU saturated, many concurrent admin operations/jsz api stats, /varz CPU, what the clients are doing
Extreme consumer countTens of thousands of consumers, especially replicated ones; admin calls and CLI time out/jsz consumer count; replicated consumers each carry Raft state
Storage exhaustionapi.errors rising with publish rejections, storage near reserved limits, consumer stall upstream/jsz storage vs reserved_storage, consumer lag
Connection path problemsDeadlines plus “Stale Client Connection - Closing” and broken pipe entries in server logsServer logs, /varz stale_connections and stalled_clients
Cold start / recovery windowDeadlines right after a restart, bare /healthz failing during asset recovery/varz uptime, bare /healthz vs ?js-server-only=true

One cause worth singling out because it surprises operators: consumer-count scaling. Operators running on the order of 50-60K consumers on a stream have reported persistent context deadline exceeded, and the maintainer response on nats-server issue #5609 was that R3 consumers each create a full Raft group, which does not scale. The recommended alternatives are fewer consumers with subject filters, or republishing from the stream to core NATS subscribers. If your consumer count is in the tens of thousands, this is probably your root cause, not disk or network.

Quick checks

All read-only. Run them against the monitoring port (default 8222).

# Is JetStream up, and what does the API backlog look like?
curl -s http://localhost:8222/jsz | jq '{disabled, api: .api, streams, consumers}'

# Meta cluster health: leader name and replica state
curl -s http://localhost:8222/jsz | jq '.meta_cluster | {leader, replicas: [.replicas[]? | {name, current, offline, lag}]}'

# Server load and connection-path distress
curl -s http://localhost:8222/varz | jq '{cpu, mem, connections, slow_consumers, stale_connections, stalled_clients}'

# Is this a cold start or recent restart?
curl -s http://localhost:8222/varz | jq .uptime

# Health: full JetStream check vs basic server readiness
curl -s http://localhost:8222/healthz
curl -s http://localhost:8222/healthz?js-server-only=true

Then at the OS level on the JetStream nodes:

# Disk latency and saturation on the JetStream storage path
iostat -x 2 5

# Any backup, snapshot, or compaction process hammering the disk?
ps aux | grep -Ei 'backup|snapshot|rsync|tar'

In the server logs, look for Stale Client Connection - Closing, broken pipe, new leader, and Stepping down entries correlating with the timeout window.

Reading the results:

  • api.inflight near zero and errors flat: the server is not backlogged right now. Reproduce the failing call while watching, or suspect a single stream’s Raft group or the connection path.
  • api.inflight high and leader stable: point at disk I/O. Confirm with iostat (high await, high %util on the store device).
  • api.inflight high plus a leader name that changes between polls: Raft instability. This is the meta cluster if admin calls fail cluster-wide, or per-stream groups if only one stream is affected.
  • api.errors climbing with publishes rejected and storage near reserved_storage: exhaustion, a different incident. Deadlines here usually mean the consumer stall behind the exhaustion has also wedged the API.

How to diagnose it

  1. Confirm the scope. Does the deadline hit all JetStream calls or only one stream or one account? Compare nats stream info on a known-good stream versus the failing one. If everything times out, think meta cluster, disk, or server-wide overload. If one stream times out, think that stream’s Raft group.

  2. Sample api.inflight over 30-60 seconds. A single snapshot lies; brief spikes during stream creation are normal. Poll /jsz a few times. Sustained high inflight means the server is slow to respond, not that clients are misbehaving.

  3. Check Raft stability. Poll the meta_cluster leader field repeatedly during the incident window. More than one leader change in 5 minutes is concerning. Check replicas for offline=true or current=false. For a single affected stream, check /raftz and the stream’s cluster state: any replica lagging or offline reduces quorum headroom and can block writes. Standalone (single-node) JetStream has no meta cluster, so meta_cluster will be null; if that is your topology, skip to step 4.

  4. Check disk I/O on the storage path. JetStream file stores are fsync-latency sensitive. High iowait, high device await, or an active filesystem backup/snapshot during the incident window is enough to explain stretched publish-acks. Network-attached storage with variable latency is the classic offender and the number one cause of downstream Raft election storms, because slow WAL writes delay Raft heartbeats.

  5. Check server saturation. /varz CPU sustained above 90 percent, memory climbing toward the container limit, and non-zero stale_connections or stalled_clients all point to a server too busy to answer API calls promptly. Also count consumers: a very large consumer population, especially replicated, loads the server with Raft groups and is a known timeout cause.

  6. Rule out recovery and lifecycle artifacts. If uptime is low, JetStream may still be recovering assets; bare /healthz fails during this window on large stores while ?js-server-only=true stays ok. Also check whether JetStream was toggled by a config reload: if JetStream is disabled or was disabled during a reload, in-flight API calls hang until their contexts expire rather than erroring out.

  7. Correlate with the client. Note the client’s configured deadline and which operation timed out. Publish-ack timeouts point at the write path (storage, stream Raft). Admin call timeouts (stream info, consumer create) point at the meta group or API backlog. Consume/fetch timeouts can also be a stalled delivery path, so check num_ack_pending against MaxAckPending on the affected consumer before blaming the server.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
/jsz api.inflightDirect measure of JetStream API backlog; the closest server-side mirror of client deadlinesSustained elevation above baseline
/jsz api.errors vs api.totalDistinguishes “slow to answer” from “actively failing”Error ratio above 5% sustained, or rising during the timeout window
meta_cluster leaderLeader churn means elections, and elections pause writes and admin opsMore than 1 change per 5 minutes
meta_cluster replicas (current, offline, lag)Non-current or offline peers erode quorum and slow consensusAny peer offline > 60s or persistently not current
OS disk latency / iowait on store devicefsync latency directly stretches publish-acks and Raft WAL writesAwait climbing during incident windows; backup processes active
/varz cpu and memA saturated server answers API calls lateCPU > 90% for > 5 min; monotonic memory growth
/varz stale_connections, stalled_clientsConnection-path distress presents to clients as deadlinesAny non-zero value sustained > 5 minutes
/jsz storage vs reserved_storageRules exhaustion in or out as the causeAbove 90% of reserved
/jsz consumers countReplicated consumers carry Raft groups; extreme counts overload the meta layerTens of thousands of consumers, especially R>1
/raftz per-stream groupsCatches the single-broken-stream case that aggregate /jsz hidesA stream group with no leader or lagging replicas

Fixes

Disk I/O stall

Move JetStream storage to local SSDs if it is on network-attached storage. Schedule filesystem backups and snapshots away from peak publish windows, or use a mechanism that does not stall the store device. Reduce the number of streams sharing one disk, or reduce the publish rate while you remediate. Faster storage is the real fix; tuning retention and rates only buys headroom.

Raft instability

Fix the underlying trigger, not Raft itself: route RTT between nodes, GC pressure from memory headroom, disk latency on the WAL path. If one node is chronically slow, its stream leaders will flap; investigate that node’s CPU steal, disk, and network. Elections usually settle once the resource problem clears. Do not restart nodes as a first move; a rolling restart is a last resort to break a self-sustaining election storm, and it will itself trigger elections.

Single broken stream

If one stream’s Raft group has lost quorum (two of three replicas offline or partitioned), restore connectivity to the peers first. If the group cannot recover on its own, treat it as a data-path incident: verify replica state from /raftz and stream cluster state, then use the stream/peer administrative operations supported by your deployed version. Do not assume a single universal in-place recovery command is safe for every topology.

API backlog and overload

Reduce concurrent administrative operations: clients that poll stream info or recreate consumers in tight loops inflate both api.total and the backlog. Add jitter and backoff to admin tooling. If the backlog is from legitimate load, scale the JetStream nodes or redistribute stream leaders so one node is not answering for everything.

Consumer-count scaling

Consolidate to fewer consumers with subject filters, or republish from the stream to core NATS subscribers instead of giving every subscriber its own durable consumer. Avoid R3 consumers at high counts; each one is a full Raft group. This is an application-design fix, not a server tune.

Client-side mitigations

Use explicit contexts with deadlines sized to your real durability SLO, and treat deadline errors as retryable with backoff, since JetStream publish and admin operations are generally safe to retry when idempotent. Retry without backoff during a server-side stall just deepens the inflight pile. Do not paper over server stalls by raising timeouts to minutes; you will convert visible errors into invisible queueing.

Prevention

  • Alert on api.inflight as a leading indicator. It rises before clients start reporting deadlines. Baseline it and alert on sustained deviation.
  • Track meta leader stability as a metric, not a log grep. Leader-change rate per 5 minutes is one of the earliest JetStream distress signals.
  • Monitor disk latency on the JetStream device independently of capacity. The disk-stall failure mode has free space; it is purely a latency problem.
  • Keep storage below 80% of reserved and IOPS below roughly 70% of the device’s proven envelope, per the capacity guidance in the NATS monitoring checklist.
  • Set Kubernetes readiness probes with the /healthz semantics in mind. Use ?js-server-only=true for the page-level readiness check and treat bare /healthz failures during recovery as expected; overly aggressive probes restart pods mid-recovery and convert a slow start into a crash loop. See NATS /healthz explained.
  • Cap consumer fan-out in application design reviews. If a design calls for thousands of replicated consumers per stream, push back before it reaches production.
  • Load-test the write path with fsync-realistic storage before declaring capacity. Benchmarks on tmpfs or burst-credit volumes hide the fsync latency that causes this incident.

How Netdata helps

  • Netdata’s NATS collector polls the HTTP monitoring endpoints and charts api.inflight, api.errors, and api.total over time, so you can see whether the backlog built gradually (disk, overload) or step-changed (election, restart) before clients reported deadlines.
  • JetStream storage, memory, stream, and consumer counts are collected from /jsz, letting you rule storage exhaustion in or out at a glance during triage.
  • Server-side CPU, memory, slow consumers, and connection counts from /varz sit on the same dashboard as the JetStream signals, which is what makes the “overloaded server vs storage vs Raft” split fast.
  • Netdata also collects node-level disk I/O latency and iowait, so you can overlay the store device’s await curve on the api.inflight curve and confirm or eliminate the fsync-stall hypothesis in one view.
  • Per-second collection granularity catches the short election-and-backlog transients that 60-second scrape intervals miss.