The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-secondary-zone-expired ▌

Operations Guides

BIND secondary zone expired: the SOA expire timer runs out and the zone returns SERVFAIL

A single zone on your BIND secondary starts returning SERVFAIL. Other zones on the same server answer normally. The named process is up, CPU and memory look fine, and port 53 responds to health checks. Your monitoring shows green because it checks process liveness or queries a different zone. Only clients asking for that one zone are failing.

This is the end state of the zone staleness cascade. Transfers from the primary have been failing silently for days or weeks while the secondary kept serving the last-known-good data. The SOA expire timer counted down to zero, and BIND removed the zone from its database. The transition is binary: at one moment the zone works (stale but functional), at the next it is gone. Recovery requires a full AXFR, and that transfer must succeed or the cycle repeats.

What this means

The SOA record’s expire field defines how long a secondary may serve zone data after the last successful transfer from the primary. BIND counts down from that value every time a refresh or retry attempt fails. When the countdown reaches zero, BIND removes the zone from its authoritative database and stops answering queries for it. Clients receive SERVFAIL.

Key characteristics of this failure:

  • Zone-scoped, not server-scoped. Only the expired zone breaks. All other zones, recursive resolution (if enabled), and the named process itself continue normally.
  • Long silent period before failure. A zone with a 7-day or 41-day expire timer can serve stale data for that entire window with no errors and no log complaints beyond quiet transfer retry messages.
  • Binary cliff. There is no degraded mode between “serving stale data” and “zone gone.” The transition is instantaneous.
  • Recovery needs a full AXFR. Once the zone is expired and removed, the secondary cannot use IXFR to catch up. It must transfer the entire zone from scratch.

BIND enforces constraints on the expire value regardless of what the primary advertises in its SOA. The expire is capped at a hard-coded maximum of 14,515,200 seconds (24 weeks). It is also floored at the sum of refresh and retry values, each of which has a 5-minute minimum. There are no min-expire-time or max-expire-time configuration knobs; only min-refresh-time, max-refresh-time, min-retry-time, and max-retry-time exist.

flowchart TD
    A["Primary reachable, transfers succeed"] --> B["Transfer starts failing"]
    B --> C["Secondary serves stale data"]
    C --> D["SOA expire countdown ticking"]
    D --> E{"Expire reached?"}
    E -->|No, transfer recovers| A
    E -->|Yes| F["Zone removed, SERVFAIL"]
    F --> G["Full AXFR required"]
    G --> A

Common causes

CauseWhat it looks likeFirst thing to check
Primary decommissioned or IP changedTransfer logs show connection refused or timeout; primary not at old addressdig @primary-ip <zone> SOA +short from the secondary
Firewall blocks TCP/53 between primary and secondaryUDP queries to primary work; AXFR/IXFR times outdig @primary-ip <zone> AXFR from the secondary host
TSIG key mismatchTransfer denied; xfer-in logs show “bad signature” or TSIG errorsCompare key statements in named.conf on both sides
Primary overloadedIntermittent transfer failures; transfers succeed during low-traffic periodsCheck primary’s CPU, FD usage, and TCP connection count
Disk full on secondaryTransfer completes but write fails; xfer-in logs show I/O errordf -h on the partition holding zone and journal files

Quick checks

Safe, read-only commands. Run them on the affected secondary.

# Check the zone's current state and expire countdown
rndc zonestatus example.com

# Compare SOA serials between primary and secondary
dig @primary-ip example.com SOA +short
dig @127.0.0.1 example.com SOA +short

# Query the zone directly to confirm SERVFAIL
dig @127.0.0.1 example.com SOA +time=2 +tries=1

# Check for transfer failure and expiry log messages
journalctl -u named --since "7 days ago" | grep -i "example.com.*transfer\|example.com.*expired\|xfer-in"

# Verify TCP connectivity to the primary on port 53
dig @primary-ip example.com SOA +tcp +time=3 +tries=1

# Check disk space on the partition holding zone files
df -h /var/named

# Confirm total loaded zones (process is fine, just one zone missing)
rndc status | grep -i "zones"

How to diagnose it

  1. Confirm the zone is expired. Run rndc zonestatus <zone>. If the zone is expired or not loaded, the output shows an error or the “expires” field is in the past. The log line zone <zone>/IN: expired confirms it.

  2. Verify only that zone is broken. Query a different zone hosted on the same secondary. If other zones answer normally, the problem is zone-scoped transfer failure, not a server-wide issue.

  3. Check primary reachability from the secondary. Run dig @primary-ip <zone> SOA from the secondary’s host. If this times out, the problem is network reachability. If it returns REFUSED or SERVFAIL, the primary itself may be broken.

  4. Test the transfer path directly. Run dig @primary-ip <zone> AXFR from the secondary. If this fails, the transfer ACL, TSIG key, or TCP path is broken. If it succeeds, the zone can be recovered with rndc retransfer <zone>.

  5. Review transfer failure history. Search xfer-in logs for the timeline of when transfers started failing. This tells you how long the zone has been stale and helps identify the triggering event (firewall change, key rotation, primary migration).

  6. Check whether a restart already masked the problem. Restarting named resets the expire timer if the zone file still exists on disk. The secondary reloads the old data and resumes serving it, buying time but not fixing the underlying transfer failure. If someone restarted named recently, the zone may appear healthy now but will expire again after the full SOA expire duration unless the transfer path is fixed.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
SOA expire runwayThe actual countdown to outage, not just serial lagRunway below 25% of expire value or below 24 hours
SOA serial mismatch (primary vs secondary)Earliest indicator that transfers are failingMismatch persisting beyond the refresh interval
Transfer success/failure rate (xfer-in logs)Shows the mechanism that feeds the zone is brokenRepeated transfer timeout or denial entries
Zone-specific SERVFAILThe user-visible symptom once expiry hitsSERVFAIL limited to one zone, zero for others
rndc zonestatus expires fieldAuthoritative local source for when the zone will die“expires” timestamp within 24 hours or in the past
Primary reachability from secondaryThe upstream dependency that feeds the timerConnection refused, timeout, or wrong serial from primary

Fixes

Primary unreachable or decommissioned

If the primary has been decommissioned or its address changed without updating the secondary’s masters (or primaries in BIND 9.18+ terminology) statement, no transfer will ever succeed.

  • Fix: Update the masters / primaries address in the secondary’s zone configuration. Then run rndc retransfer <zone> to force an immediate full transfer.
  • Tradeoff: If the primary is permanently gone and no other source exists, rebuild the zone data from a backup or from another secondary that still has a valid copy.

Firewall blocks TCP/53

Zone transfers use TCP. A firewall rule that permits UDP/53 but blocks TCP/53 between the secondary and primary will break transfers while allowing SOA queries to succeed.

  • Fix: Open TCP/53 between the secondary and primary in both directions. Verify with dig @primary-ip <zone> AXFR from the secondary host.
  • Tradeoff: If your security policy restricts TCP/53 broadly, consider a dedicated transfer network, a VPN tunnel, or an intermediate hidden primary that the secondaries can reach.

TSIG key mismatch

TSIG keys used for transfer authentication can drift if one side is updated but not the other. The secondary’s transfer request fails authentication, and BIND logs a signature or TSIG error in the xfer-in category.

  • Fix: Copy the current key from the primary to the secondary (or vice versa), ensure the allow-transfer and masters / primaries statements reference the same key, reload the configuration, and run rndc retransfer <zone>.
  • Tradeoff: Key rotation procedures should update all secondaries atomically. Running both old and new keys temporarily during the transition window avoids a gap.

Disk full on the secondary

If the secondary cannot write the received zone data or its journal file, the transfer completes over the network but the write fails. BIND logs an I/O error.

  • Fix: Free disk space on the affected partition. Then run rndc retransfer <zone>.
  • Tradeoff: If the disk filled due to unbounded named_stats.txt growth from repeated rndc stats calls, add log rotation or truncate the file.

Forcing recovery after expiry

Once the zone is expired and removed, the only path to recovery is a successful full AXFR. The secondary requests IXFR by default, but with the zone missing from memory it falls back to AXFR.

# Force a full transfer of the expired zone
rndc retransfer example.com

# Verify the zone loaded successfully
rndc zonestatus example.com

# Confirm queries now succeed
dig @127.0.0.1 example.com SOA +short

If the transfer fails again, the zone remains unserved. Fix the transfer path first.

Prevention

  • Monitor SOA expire runway, not just serial mismatch. Serial mismatch tells you transfers are failing; expire runway tells you when the zone will actually break. Alert when runway drops below 25% of the expire value or below 24 hours, whichever is shorter. The rndc zonestatus <zone> output includes the “expires” timestamp for this purpose.

  • Track serial consistency between primary and all secondaries. A SOA serial comparison catches transfer failures within minutes rather than days. Query the SOA record against each server and compare the serial field.

  • Use dig +expire for remote expire monitoring. BIND 9.10+ supports the EDNS EXPIRE option, which lets you query a secondary’s remaining expire time remotely without rndc access. Useful for monitoring secondaries you do not control directly.

  • Validate the transfer path after any firewall or network change. Run dig @primary-ip <zone> AXFR from each secondary to confirm TCP/53 works end to end. Do not assume that successful UDP SOA queries mean transfers will work.

  • Alert on xfer-in failures. BIND logs every transfer attempt in the xfer-in category. A sustained stream of failures is the earliest indicator that the expire countdown has started. Forward these logs to your monitoring system.

  • Remember that restarting named resets the timer. If the zone file is still on disk, restarting named reloads the old data and resets the expire countdown. This can mask an ongoing transfer failure for another full expire duration. Never treat a restart as a fix.

How Netdata helps

  • SOA expire runway as a first-class signal. Netdata collects zone-level statistics from BIND’s statistics channel and can surface the expire countdown as a trendable metric, turning a cliff-edge failure into something you can see coming days in advance.

  • Serial mismatch correlation. By comparing SOA serials between primary and secondary at regular intervals, Netdata can alert on divergence before the expire countdown becomes critical.

  • Zone-scoped SERVFAIL detection. Netdata’s per-second query metrics can show SERVFAIL rising for one specific zone while other zones remain clean, narrowing the problem from “DNS is broken” to “this zone is expired.”

  • Transfer failure signal correlation. BIND’s zone transfer counters and xfer-in log streams, when correlated with expire runway trends, expose the full causal chain: transfer failures start, serial diverges, runway shrinks, zone expires.

  • Process health in context. Netdata shows named process liveness, memory, CPU, and file descriptor usage alongside zone health. When only one zone is expired, the process metrics stay green, which is itself diagnostic: the server is fine, the transfer path is not.