The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-soa-serial-mismatch ▌

Operations Guides

BIND SOA serial mismatch: a secondary serving stale data behind the primary

The secondary’s SOA serial is behind the primary. Clients querying the secondary get stale A values, missing new entries, deleted hosts still resolving. The server responds NOERROR because its zone data is internally consistent. It is not current.

This is the precursor to zone expiry, and it is silent. The secondary serves stale but functional answers for the entire SOA expire period, typically 1 to 4 weeks. Monitoring that checks “can I resolve this zone?” passes. Health checks on port 53 pass. Everything looks fine except the data is wrong. The only signal is the serial number gap between primary and secondary, and BIND’s statistics channel does not expose it. You must probe externally.

The mismatch becomes actionable when it persists beyond the zone’s SOA refresh interval. At that point, transfers are failing, not delayed. The expire countdown is running. This article covers how to detect the gap, diagnose why transfers are failing, and fix it before the secondary stops serving the zone entirely.

What this means

A BIND secondary keeps zone data current through polling and NOTIFY:

  1. The primary sends NOTIFY over UDP when the zone changes.
  2. The secondary schedules a refresh. It queries its configured primaries list, not the NOTIFY sender directly.
  3. For each primary, the secondary queries the SOA and compares serials using RFC 1982 arithmetic.
  4. If the primary’s serial is greater, the secondary initiates AXFR or IXFR over TCP/53.
  5. If the serials match or the primary is unreachable, the secondary moves to the next primary.
  6. If no primary offers a newer serial, the secondary keeps its current data and waits for the next refresh interval.

Persistent failure here means the secondary falls behind. The serial mismatch is the symptom. The cause is almost always one of: primary unreachable, transfer denied (ACL or TSIG), network path broken, or serial arithmetic confusion.

RFC 1982 defines serial comparison on a 32-bit sequence space. Serials wrap at 2^32. A serial of 4294967290 followed by 5 is valid: the primary incremented past the boundary. The maximum defined increment is 2^31 - 1 (2147483647). Pairs exactly 2^31 apart are undefined. In practice, serial wraparound is rare and usually signals an operational mistake.

The expire timer runs independently of the refresh mechanism. If the secondary has not completed a successful transfer within the SOA expire interval (the sixth field in the SOA record), it stops serving the zone entirely. Serial mismatch is the warning. Expiry is the cliff.

flowchart TD
    A["Serial mismatch detected"] --> B{"Primary SOA query works?"}
    B -- No --> C["Primary down or unreachable"]
    B -- Yes --> D{"AXFR from primary succeeds?"}
    D -- No --> E["allow-transfer, TSIG, or TCP/53"]
    D -- Yes --> F["Check RFC 1982 arithmetic"]
    E --> G["Fix root cause, then rndc retransfer"]
    F --> G
    G --> H{"Serial converges?"}
    H -- No --> I["Check BIND version, min-transfer-rate-in, journal"]
    H -- Yes --> J["Fixed. Track expire runway."]

Common causes

CauseWhat it looks likeFirst thing to check
allow-transfer ACL denies the secondaryTransfer denied in primary security log. XfrFail increments in zonestats. BIND 9.20+ defaults allow-transfer to none.named-checkconf -p on primary, grep for allow-transfer
Firewall blocks TCP/53 between serversSOA query over UDP succeeds but AXFR or IXFR over TCP fails or times out.dig @primary-ip example.com AXFR +time=5 +tcp from secondary host
TSIG key mismatchTransfer starts but fails with NOTAUTH or BADKEY. xfer-in logs show TSIG errors.Compare key name, algorithm, and secret on both servers
Primary unreachable from secondarySOA query to primary fails entirely. No path to primary.dig @primary-ip example.com SOA +time=2 +tries=1 from secondary
NOTIFY not reaching secondarySerial converges eventually but only on next scheduled refresh, not promptly after a zone change.Firewall rules for UDP/53; verify notify-source is routable
Serial arithmetic edge caseSerials appear to mismatch but RFC 1982 comparison shows they are equal or the lower number is actually ahead.Manual RFC 1982 comparison

Quick checks

# Compare SOA serials: third field is the serial number
dig @primary-ip example.com SOA +short
dig @secondary-ip example.com SOA +short

# Check expire runway and transfer state on the secondary
rndc zonestatus example.com

# Check transfer counters from statistics channel
# Note: replace 8653 with your configured statistics-channels port
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{k}: {v}') for k,v in d.get('zonestats',{}).items() if 'Xfr' in k]"

# Look for transfer failures in logs
journalctl -u named --since "1 hour ago" | grep -iE "transfer|xfr|notify|denied"

# Check allow-transfer configuration on the primary
named-checkconf -p | grep -i "allow-transfer"

# Test AXFR from the primary (run from the secondary host)
dig @primary-ip example.com AXFR +time=5

# Check BIND version
rndc status | head -1

Zero-valued counters are omitted from statistics channel output by default. If XfrFail does not appear in the output, the counter is zero.

How to diagnose it

  1. Confirm the mismatch is real and persistent. Compare serials from an external probe. If the gap appeared moments ago, it may be a normal in-flight transfer. If it persists beyond the SOA refresh interval, transfers are failing.

  2. Check the expire runway. Run rndc zonestatus example.com on the secondary. The expire field tells you how long until the zone is removed. If the runway is below 25% of the total expire value or below 24 hours, this is urgent.

  3. Verify primary reachability from the secondary host. Run a SOA query directly against the primary’s IP from the secondary. If this fails, the problem is network path or primary health, not transfer configuration.

  4. Test whether transfers are authorized. Attempt an AXFR from the secondary host against the primary. If it fails with REFUSED, the allow-transfer ACL is denying the secondary. If it fails with NOTAUTH, TSIG is the problem.

  5. Check transfer logs on both servers. On the primary, look in the security and xfer-out log categories for denied or failed transfers. On the secondary, look in xfer-in for transfer failures, timeouts, or TSIG errors.

  6. Check for BIND version-specific default changes. BIND 9.20.0 changed allow-transfer from the old any default to none; outgoing transfers now require an explicit zone, view, or options ACL. Also check for removed options: alt-transfer-source, alt-transfer-source-v6, and use-alt-transfer-source were removed in BIND 9.20.0 and will cause named-checkconf errors if still present.

  7. Consider serial arithmetic and refresh internals. BIND caches an unreachable primary for 10 minutes, or until it receives NOTIFY from that address. NOTIFY can therefore accelerate recovery even if the regular refresh cycle has not fired yet. BIND allows one active refresh and one queued refresh per zone; additional NOTIFY messages arriving while a refresh is already queued are discarded.

  8. Force a retransfer and observe. Run rndc retransfer example.com on the secondary. This forces a full AXFR regardless of serial comparison. If it succeeds, the serial converges and the transfer path works. If it fails, the failure mode is visible in the logs.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
SOA serial consistency (primary vs secondary)Direct measure of zone data freshness. Requires external probing, not in stats channel.Any difference persisting beyond the SOA refresh interval.
SOA expire runwayTime remaining before the secondary stops serving the zone. This is the real danger metric.Below 50% of total expire value. Critical below 25% or 24 hours.
XfrFail counter (zonestats)Transfer failure rate from the statistics channel. Sustained increase means transfers are repeatedly failing.Any sustained increase from baseline.
XfrSuccess counter (zonestats)Successful transfer count. Should increment on each zone change on the primary.Flatlined while primary serial is advancing.
SERVFAIL rate for specific zoneTerminal symptom: the zone has expired on the secondary and it stopped serving it.Any SERVFAIL for a zone you control served by a secondary.

Fixes

Fix the allow-transfer ACL

On the primary, explicitly authorize the secondary:

zone "example.com" {
    type primary;
    allow-transfer { <secondary-ip>; };
};

BIND 9.20.0 changed the default for allow-transfer to none. Any zone that relied on the old implicit permissive default now silently refuses all transfers. After upgrading, audit every zone that should allow transfers and add an explicit allow-transfer statement.

Open TCP/53 between primary and secondary

Zone transfers use TCP port 53. If the primary responds to SOA queries over UDP but transfers fail, the firewall is likely blocking TCP/53 in one direction. This is common in environments where UDP/53 was opened but TCP/53 was assumed unnecessary. NOTIFY messages use UDP/53, so NOTIFY may succeed while transfers fail.

Resync TSIG keys

If transfers fail with NOTAUTH or BADKEY, the TSIG key on the primary and secondary has drifted. Compare the key name, algorithm, and secret on both servers. Regenerate and redistribute if the source of the drift is unclear. Verify that the server statement on the secondary references the correct key for the primary, and that allow-transfer on the primary requires the matching key.

Recover from serial number errors

If the primary’s serial was accidentally set lower than the secondary’s (for example, a zone file restored from an old backup), BIND will not accept a downgrade. The two-step recovery:

  1. Set the primary serial to the secondary’s current serial plus 2147483647 (2^31 - 1). This value is guaranteed to be greater by RFC 1982 arithmetic. Reload the primary and let the secondary transfer.
  2. After the transfer succeeds, set the primary serial to the desired correct value. Reload again. The secondary transfers again and converges.

Avoid setting the serial to zero. Some DNS implementations treat zero as special and BIND may exhibit unexpected behavior.

Handle slow-transfer termination

BIND 9.20.7 introduced min-transfer-rate-in, with default 10240 5: terminate an inbound transfer if it delivers less than 10240 bytes in 5 minutes. On large zones transferred over slow or congested links, this can cause repeated transfer failures with no obvious ACL or TSIG error. If transfers start but never complete, check for transfer-rate termination in the xfer-in logs.

Handle journal corruption

If IXFR transfers repeatedly fail and fall back to AXFR (visible in xfer-in logs), the journal file on the primary may be corrupt.

WARNING: rndc sync -clean writes the journal to the zone file and deletes the journal. Run this only on the primary, and ensure no dynamic DNS updates are in flight.

# On the primary: write journal to zone file, remove journal
rndc sync -clean example.com
rndc reload example.com

This commits current journal contents to the zone file and removes the journal. Dynamic updates since the last sync are preserved in the written zone file.

Force an immediate transfer

After fixing the root cause, force convergence on the secondary:

# Forces a full AXFR regardless of serial comparison
rndc retransfer example.com

Use this only after resolving the underlying issue. If the root cause is not fixed, the retransfer will fail the same way. For zones in views, specify the view: rndc retransfer example.com IN external.

Prevention

  • Monitor SOA serial consistency externally. Probe both primary and secondary with dig SOA at regular intervals. Alert on any mismatch persisting beyond the SOA refresh interval. This signal is not available in the statistics channel.
  • Track the expire runway. Use rndc zonestatus output to compute remaining time before expiry. Alert when runway drops below 50% of the total expire value.
  • Audit allow-transfer after BIND upgrades. If the default changed to none, any upgrade requires an explicit audit of all transfer ACLs before the upgrade is considered complete.
  • Verify transfer health after every config change. After rndc reload, confirm that the secondary’s serial converges. Do not assume transfers work because the reload command succeeded.
  • Monitor XfrFail and XfrSuccess counters. A trending XfrFail or flatlined XfrSuccess while the primary serial advances indicates a transfer problem before the serial gap is large enough to matter.
  • Keep NOTIFY paths open. NOTIFY uses UDP/53. If NOTIFY is blocked, the secondary only learns about changes on its next scheduled refresh, adding unnecessary latency to convergence.

How Netdata helps

  • Netdata collects XfrFail and XfrSuccess from the BIND statistics channel. A rising XfrFail trend or flatlined XfrSuccess while the primary is known to be updating is the earliest statistics-channel signal that transfers are broken.
  • QrySERVFAIL collected per second lets you see the exact moment a zone starts failing on the secondary. If SERVFAIL for a specific zone appears after a sustained period of transfer failures, the causal chain from broken transfer to zone expiry is confirmed.
  • Per-second collection means you can correlate a transfer failure timestamp with other events on the same host: BIND restarts, config reloads, network interface flaps, or CPU spikes on the primary that made it too slow to respond.
  • The serial mismatch itself requires external SOA probing, which is outside the statistics channel. Pair the transfer counter trends from Netdata with an external serial check for complete coverage of the failure-to-expiry cascade.
  • Netdata collects transfer counters alongside cache, recursive client, and query rate metrics on the same timeline, so you can distinguish a transfer-specific problem from general BIND degradation.