The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-serve-stale-masking-failures ▌

Operations Guides

CoreDNS serve_stale: keeping resolution alive while masking upstream failure

The serve_stale option in the CoreDNS cache plugin is a resilience feature with a monitoring side effect that bites teams during real incidents. When it is enabled, CoreDNS answers queries from expired cache entries instead of failing them when the upstream resolver is unreachable. Clients keep resolving names through an upstream outage. Dashboards stay green. The upstream can be dead for an hour before anyone notices, because the one metric that would tell you is almost never graphed.

This article explains how serve_stale behaves, why it masks upstream failure from end-to-end success metrics, and how to instrument it so the resilience does not turn into blindness. It assumes you understand the CoreDNS plugin chain and the cache plugin’s role in it. If not, start with How CoreDNS actually works in production.

What serve_stale is and why it exists

Without serve_stale, the cache plugin has a hard rule: an entry past its TTL is useless. A query for an expired name goes upstream, and if the upstream is down or slow, the client gets SERVFAIL or a timeout. During a full upstream outage, every cached name expires within minutes to an hour, and resolution then fails completely. The playbook’s “Upstream Black Hole” pattern describes exactly this: cache temporarily masks the failure, then all queries fail once TTLs expire.

With serve_stale enabled, the cache plugin keeps expired entries for a configurable window (default 1 hour past expiry) and serves them to clients when a fresh answer cannot be fetched quickly. The design follows unbound’s serve-expired behavior, not RFC 8767: it deliberately favors fast responses over answer correctness. An expired entry younger than the configured duration is served immediately, and a background refresh attempts to fetch a fresh answer from the upstream.

The operational trade is explicit: you accept that some answers may be outdated in exchange for resolution surviving upstream blips, slow upstreams, and short network partitions. For most external dependencies, a stale A record from 20 minutes ago is far better than SERVFAIL. The problem is not the feature. The problem is that it changes what your success metrics mean.

How it works

The configuration lives inside the cache plugin block in the Corefile:

cache 30 {
    serve_stale 1h
}

The syntax is serve_stale [DURATION] [REFRESH_MODE [VERIFY_TIMEOUT]], with two refresh modes:

  • immediate (the default): the stale entry is sent to the client right away, and a refresh to the upstream is dispatched in the background. Client latency stays at cache-hit speed even while the upstream is degraded.
  • verify: CoreDNS tries the upstream first and falls back to the stale entry only if the upstream does not answer. In earlier releases this could block the client for the full upstream timeout; recent releases added an optional verify timeout, for example serve_stale 1h verify 100ms, after which the stale entry is served and verification continues in the background (the timeout argument was added in v1.14.4).

Two behavioral details matter operationally. First, stale responses are served with a TTL of 0, so downstream resolvers and clients do not re-cache the stale answer. Second, the served-stale path emits exactly one signal: the counter coredns_cache_served_stale_total (labels server, zones, view). That counter is the entire observability surface of the feature.

flowchart LR
  Q[Client query] --> C{Cache entry fresh?}
  C -- yes --> H[Serve from cache]
  C -- expired --> U{Upstream reachable?}
  U -- yes --> F[Fetch fresh answer and update cache]
  U -- no or slow --> S[Serve stale entry, TTL 0]
  S --> M[coredns_cache_served_stale_total increments]
  S --> B[Background refresh retries upstream]

The success metric the client sees (coredns_dns_responses_total{rcode="NOERROR"}) does not distinguish the fresh path from the stale path. That is the crux of the masking problem.

Where the masking bites in production

The playbook calls out the general version of this trap in its “cache masks upstream failure” anti-pattern: when the cache is warm, an upstream failure is invisible in end-to-end success rates, and teams see “everything is fine” until the cache drains. serve_stale extends the masking window from “until TTL expiry” to “until TTL expiry plus the stale duration.” With the default 1-hour stale window, an upstream can be completely unreachable for over an hour while your NOERROR ratio stays at 99.9%.

Three concrete consequences:

  1. Upstream health alerts tied to SERVFAIL stop firing. Without serve_stale, a dead upstream produces a SERVFAIL spike within minutes as entries expire. With serve_stale, SERVFAIL never comes, so any alert rule built on coredns_dns_responses_total{rcode="SERVFAIL"} stays silent for the duration of the stale window.
  2. Latency dashboards look healthy. In immediate mode, stale serves are cache-speed responses. The “Upstream Black Hole” pattern’s distinguishing feature (SERVFAIL with low latency) never appears, and the “Slow Upstream Drag” pattern’s high P99 never appears either.
  3. The failure surfaces as data staleness, not errors. Applications keep resolving, but to increasingly old answers. For records that change (failover IPs, traffic-shifted endpoints, recently rotated service addresses), this becomes a correctness incident that looks like an application bug.

There is also a deployment-specific caveat: CoreDNS guidance recommends enabling serve_stale for custom forwarded zones, not for the server block that contains the kubernetes plugin. Headless service pod IPs change frequently, and serving expired pod IPs during an API or upstream hiccup can send traffic to dead pods. For Kubernetes-internal staleness failure modes, see the API disconnect pattern in CoreDNS query rate dropped to zero while the process looks healthy and the monitoring checklist.

Two historical bugs are worth knowing if you run older versions: a race between serve_stale and prefetch caused redundant upstream fetches, and stale negative-cache entries could mask a name that had started resolving. The prefetch/stale race was fixed in CoreDNS v1.8.1 (PR #4367) and the stale negative-cache masking in v1.6.9 (PR #3744); both are years old at this point. On current releases these are not concerns, but they explain confusing behavior reports from older clusters.

Tradeoffs and when to use it

  • Use it for external forwarded zones. SaaS endpoints, package registries, identity providers: for these, a 30-minute-old answer during an upstream blip is almost always better than an error.
  • Prefer immediate mode when client latency matters and you accept staleness. Prefer verify with an explicit timeout when correctness matters more but you still want a bounded fallback.
  • Do not use it on the Kubernetes server block for the reasons above. Stale pod IPs are worse than a fast failure.
  • Size the stale window deliberately. The 1-hour default is generous. Ask how stale an answer can be before it causes harm for the records you actually serve, and set the duration accordingly.
  • Treat it as a warning system, not a comfort blanket. Every stale serve means an upstream interaction failed. If the counter climbs for hours, you do not have resilience, you have an outage you have not paged on.

Signals to watch in production

SignalWhy it mattersWarning sign
coredns_cache_served_stale_total (rate)The only direct measure of stale serving; each increment is a query that could not get a fresh answerAny sustained nonzero rate; rising trend
coredns_proxy_healthcheck_failures_total{to=...}Per-upstream health, independent of what clients seeSustained increments for any upstream while stale serves rise
coredns_forward_healthcheck_broken_totalAll upstreams failing health checks simultaneouslyAny increment, especially alongside rising stale serves
coredns_proxy_request_duration_seconds{to=...}Shows whether the upstream is slow (dragging) rather than deadP99 climbing while stale serves rise
Cache hit ratio (coredns_cache_hits_total / coredns_cache_requests_total)Stale serving inflates apparent cache effectivenessHit ratio looks stable or better while upstream health degrades
coredns_dns_responses_total{rcode="SERVFAIL"}With serve_stale, absence of SERVFAIL proves nothing by itselfFlat zero during a known upstream event means the stale path absorbed it

The alert rule that matters most is simple: a rate of coredns_cache_served_stale_total greater than zero sustained over a few minutes is a ticket, and a steeply rising rate combined with coredns_forward_healthcheck_broken_total incrementing is a page. You are then in the “all upstreams down” incident, just with the client-facing symptoms delayed. The triage path in CoreDNS all upstreams down: the forwarding black hole and healthcheck_broken applies; the only difference is that clients are not yet feeling it.

One limitation: the stale counter does not tell you whether the upstream was slow or fully down. There is no label for that. Correlation with the per-upstream health and latency metrics above is how you tell the difference, which is why monitoring upstream health independently of end-to-end success is non-negotiable when this feature is on.

How Netdata helps

  • Netdata collects the CoreDNS Prometheus endpoint and charts coredns_cache_served_stale_total alongside cache hits, request rates, and response codes, so a rising stale-serve rate is visible in the same view as the success metrics it is masking.
  • Per-second granularity catches short upstream blips that minute-resolution scraping averages away, which is exactly the timescale at which serve_stale activates.
  • Upstream health metrics (coredns_proxy_healthcheck_failures_total per upstream, coredns_forward_healthcheck_broken_total) sit next to the stale counter, making the slow-versus-dead distinction a two-chart correlation instead of a log dive.
  • Latency percentiles for forwarded queries are charted per upstream, so you can see a degrading upstream before the stale rate climbs.
  • ML-based anomaly detection on the stale counter flags the first deviations from the normal zero baseline rather than waiting for a static threshold.