The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-any-query-amplification ▌

Operations Guides

CoreDNS ANY query flood: DNS amplification and reflection abuse

You are looking at coredns_dns_requests_total and the type="ANY" series has jumped from a flat zero to a sustained rate, or ANY queries have crept past a few percent of total traffic. That pattern is the classic fingerprint of DNS amplification and reflection abuse: an attacker sends small ANY queries with a spoofed source address, and your CoreDNS replies with large responses delivered to the victim’s IP. Your server is the reflector; someone else absorbs the blast.

This is not always an attack. Some debug tooling and monitoring scripts use ANY queries legitimately at low volume, and RFC 8482 acknowledges debugging as a valid use. The operational question is not “is ANY nonzero” but “did ANY change shape”: a spike from a zero baseline, or ANY crossing roughly 5% of total traffic, is unusual and worth investigating.

What this means

An ANY query asks the server for all records it holds for a name. A normal A or AAAA answer is small. An ANY answer can contain every record type for the name at once, which makes the response far larger than the query. That asymmetry is the whole attack:

  1. The attacker crafts an ANY query with the source IP set to the victim’s address (UDP makes source spoofing trivial on networks without egress filtering).
  2. CoreDNS receives a small query and sends a large response to the spoofed source.
  3. Thousands of open resolvers doing this at once produce a volumetric DDoS against the victim, and your egress bandwidth, conntrack table, and UDP buffers pay part of the cost.
flowchart LR
    A[Attacker] -->|"small ANY query, spoofed source = victim IP"| B[CoreDNS]
    B -->|"large ANY response"| V[Victim IP]
    B -.->|"also burns your egress, UDP buffers, conntrack"| N[Your node]

RFC 8482 exists precisely because of this. It deprecates the expectation that ANY returns everything and blesses minimal responses. CoreDNS behavior depends on which plugins you have configured: without anything specific in place, ANY answers come from whatever plugin serves the zone, at full size.

Two measurement realities matter for diagnosis. First, CoreDNS exposes no per-source-IP metrics, so identifying who is sending the flood requires log analysis or dnstap, not Prometheus. Second, UDP packets dropped by the kernel before reaching CoreDNS are invisible in CoreDNS metrics, so a severe flood can make your query rate look lower than reality while the node drowns. See the hub’s failure pattern catalogue for the UDP buffer cliff and conntrack exhaustion patterns.

Common causes

CauseWhat it looks likeFirst thing to check
Active reflection attackANY rate spiking from zero baseline, response size distribution shifting upward, egress saturatedcoredns_dns_requests_total{type="ANY"} rate and coredns_dns_response_size_bytes histogram
Reconnaissance or scanningModerate ANY rate against many names, possibly mixed with AXFR attemptscoredns_dns_requests_total{type="AXFR"} alongside ANY
Legitimate debug toolingLow, steady ANY baseline from one or two internal sourcesEnable the log plugin temporarily and aggregate by source IP
Misbehaving internal clientANY queries from a single pod, often a monitoring script or resolver testQuery logs grouped by source IP
CoreDNS reachable from untrusted networksANY flood combined with total QPS far above expected internal demandWhere the service is exposed: ClusterIP only, or node port / LoadBalancer / host port

Quick checks

All of these are read-only.

# ANY query rate and share of total traffic
curl -s http://localhost:9153/metrics | grep 'coredns_dns_requests_total' | grep 'type="ANY"'

# Response size distribution: is the histogram shifting upward?
curl -s http://localhost:9153/metrics | grep 'coredns_dns_response_size_bytes'

# AXFR attempts often travel with ANY reconnaissance
curl -s http://localhost:9153/metrics | grep 'type="AXFR"'

# Total query rate for context (is overall QPS also abnormal?)
curl -s http://localhost:9153/metrics | grep '^coredns_dns_requests_total'

Then check what the node sees, because CoreDNS metrics cannot:

# UDP buffer errors: packets dropped before CoreDNS ever counted them
netstat -su | grep -i "buffer errors"
cat /proc/net/snmp | grep Udp

# Conntrack pressure on the node (Kubernetes with iptables DNAT)
cat /proc/sys/net/netfilter/nf_conntrack_count
cat /proc/sys/net/netfilter/nf_conntrack_max

Finally, see what an ANY response from your server actually looks like. This tells you the amplification factor you are offering:

# Compare ANY response size to a plain A response
dig @<coredns_ip> <a_name_you_serve> ANY +noall +answer
dig @<coredns_ip> <a_name_you_serve> A +noall +answer

If the ANY answer comes back with a full record set, your server is a good reflector. If it comes back with a single short HINFO record, someone already deployed the fix described below.

How to diagnose it

  1. Confirm the ANY share. Compute ANY as a percentage of total coredns_dns_requests_total. Under 5% and steady is plausibly baseline tooling. A spike from zero, or a sustained rate well above 5%, moves you to step 2.

  2. Correlate with response size. Check coredns_dns_response_size_bytes. A flood of ANY queries producing large answers shifts the histogram’s upper buckets. If ANY is up but response sizes are flat and small, your ANY responses are already minimal (or the queries are failing), and the reflection risk is low even though the query pattern is suspicious.

  3. Identify the source. Prometheus cannot help here. Temporarily enable the log plugin and aggregate by client IP. The log plugin’s combined format puts the client address in the second field, as IP:port:

    # Group queries by source IP (requires the log plugin)
    kubectl logs -n kube-system <coredns-pod> | awk '{print $2}' | cut -d: -f1 | sort | uniq -c | sort -rn | head -20
    

    Many spoofed sources, or internet-routable sources you do not recognize, point to abuse. One internal IP points to a misbehaving or compromised workload. Note the tradeoff: the log plugin adds real overhead at high QPS, so enable it for the investigation and remove it after.

  4. Check exposure. In a default Kubernetes deployment, CoreDNS sits behind a ClusterIP and is only reachable inside the cluster. If your instance answers queries from untrusted networks (host port, node port, LoadBalancer, or a standalone deployment on a public interface), the attack surface is much larger and source spoofing is much easier.

  5. Check collateral damage. During a real flood, the first thing to break may not be CoreDNS itself. Look at netstat -su receive buffer errors and conntrack utilization from the quick checks. If those are climbing, the node is shedding packets for every workload on it, not just DNS.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
coredns_dns_requests_total{type="ANY"}The primary abuse indicatorANY > 5% of total traffic, or a spike from a zero baseline
coredns_dns_response_size_bytesConfirms whether ANY queries produce large (amplifying) responsesSustained shift in the distribution, or responses above 4096 bytes
coredns_dns_requests_total{type="AXFR"}Zone transfer attempts often accompany ANY reconnaissanceAny nonzero AXFR where transfers are not expected
Total coredns_dns_requests_totalSeparates an ANY-specific event from a general QPS surgeRate above 2-3x the rolling baseline with no known cause
Node UDP buffer errors (/proc/net/snmp, RcvbufErrors)Reveals floods the CoreDNS process never seesAny nonzero, incrementing count
Node conntrack usageUDP DNS flows consume conntrack entries; floods exhaust the tableUsage above 80% of nf_conntrack_max

Fixes

Reduce the amplification factor with the any plugin

Since CoreDNS 1.5.1, the built-in any plugin answers ANY queries with a minimal HINFO response per RFC 8482 instead of the full record set. Adding it to the Corefile is a one-word change in the server block:

. {
    any
    # ... rest of your plugin chain
}

The response becomes a short synthesized record ("ANY obsoleted" "See RFC 8482"), which collapses the amplification factor to roughly 1:1. Your server stops being an attractive reflector.

Tradeoffs to understand:

  • This is not a rate limiter. The attacker can still send the queries; you just stop multiplying their effect. Query volume still costs you CPU, sockets, and conntrack entries.
  • Legitimate dig ANY debugging against your server now returns the HINFO stub, not real data. That is the RFC-sanctioned behavior, but expect confused colleagues.
  • The upstream Kubernetes default Corefile ships without any: the standard manifest defines only errors, health, ready, kubernetes, prometheus, forward, cache, loop, reload, and loadbalance. Check your Corefile before assuming the protection is in place.

Rate limit with the rrl plugin

For volumetric mitigation, the rrl plugin provides BIND-style response rate limiting: it tracks response rates per category and can drop or truncate responses that exceed limits.

Tradeoffs:

  • rrl is an external plugin. It is not in standard CoreDNS builds, so you must compile a custom CoreDNS image with it added to plugin.cfg. For teams running the stock Kubernetes CoreDNS image, this is a real operational burden.
  • The rrl README documents a known wildcard-flooding weakness: an attacker spreading queries across unlimited unique names synthesized by a wildcard record keeps per-name rates under the limits. Mitigating this requires the metadata plugin so responses synthesized by known wildcards are accounted under the wildcard’s parent domain; the README itself leaves the minimum CoreDNS version as TBD.
  • Rate limiting that is too aggressive will clip legitimate clients during bursts. Start permissive and tighten against observed baselines.

Remove the exposure

If CoreDNS is reachable from untrusted networks and does not need to be, fixing that beats every in-process mitigation:

  • Keep cluster DNS behind the ClusterIP; do not publish it via node port, host port, or LoadBalancer.
  • For standalone deployments, bind to internal interfaces only and filter port 53 at the network edge for untrusted sources.
  • Egress filtering (BCP 38) on your own network prevents your hosts from being the source of spoofed packets, which is the neighborly half of the same problem.

Survive the active flood

While a flood is ongoing, the risk is that the node fails before CoreDNS does. Watch UDP receive buffer errors and conntrack utilization, and be prepared to raise net.core.rmem_max and nf_conntrack_max as stopgaps. Both are covered in depth in the related guides below. Do not restart CoreDNS as a first response: it does nothing to stop spoofed traffic and adds a cold-cache thundering herd on top of the attack.

Prevention

  • Baseline the ANY share. Record your normal ANY percentage so “spike from zero” and “>5%” are detectable, not vibes.
  • Ship any in the Corefile by default. It is cheap, built in, and removes the reflection value of your server.
  • Alert on the pair, not the point. ANY rate up plus response size distribution shifting is a much stronger signal than either alone, and it filters out low-volume legitimate debugging.
  • Audit exposure after infrastructure changes. New load balancers, node ports, and firewall rules are how internal resolvers accidentally become public reflectors.
  • Monitor node-level signals. UDP buffer errors and conntrack utilization are where a flood actually hurts first, and neither appears in CoreDNS metrics.

How Netdata helps

  • Netdata charts coredns_dns_requests_total broken out by query type, so a rising type="ANY" series is visible next to A and AAAA traffic without writing PromQL during an incident.
  • The coredns_dns_response_size_bytes distribution is charted alongside query rates, making the “ANY up plus responses getting bigger” correlation a single-screen check.
  • Per-second collection catches short flood bursts that minute-resolution scraping averages into invisibility.
  • Node-level metrics (UDP errors, conntrack usage) are collected from the same hosts, so you can correlate the DNS-layer symptom with the kernel-layer damage in one view.
  • Baseline deviation on the ANY share is the kind of pattern ML anomaly detection flags well, because the normal value is a flat near-zero line.