The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-goroutine-leak ▌

Operations Guides

CoreDNS goroutine count climbing: blocked upstream calls and leaks

Your go_goroutines graph for CoreDNS is trending up. Maybe it spiked during an incident and never came back down. Maybe it has been climbing slowly for days. Either way, the question is the same: are queries piling up behind a slow upstream, or is something leaking goroutines that will never exit?

The distinction matters because the fixes are completely different. A blocked-upstream spike resolves when the upstream recovers. A leak grows until the OOM killer ends the debate for you.

What this means

CoreDNS handles each in-flight DNS query on its own goroutine, plus a small set of background goroutines for cache maintenance, Kubernetes API watches, health checks, and metrics. Go’s runtime spawns and reaps these cheaply, so the goroutine count is a live proxy for concurrent work.

A healthy instance sits at a baseline of roughly 20-50 goroutines. Under load the count tracks approximately with QPS multiplied by average query duration: fast queries mean goroutines are created and destroyed so quickly the count stays low. Two failure shapes break this pattern:

  • Blocked upstream calls. A query that needs forwarding holds its goroutine until the upstream responds or the forward timeout fires. If an upstream gets slow but not dead, in-flight queries accumulate. Goroutines climb, latency climbs with them, and each goroutine holds memory. This pattern is self-limiting: when the upstream recovers, the backlog drains and the count returns to baseline.
  • A goroutine leak. A goroutine blocked on something that will never complete: an unclosed connection, a deadlocked channel, a plugin bug. The count ratchets upward regardless of load and never returns to baseline. At roughly 4KB minimum stack per goroutine plus per-goroutine allocations, enough leaked goroutines push the process toward its memory limit and an OOM kill.

The operational dividing line: a sustained count above 2x your stable baseline for more than 10 minutes warrants investigation. Above 5000 goroutines combined with rising memory is a strong leak signal. A startup spike that settles within about 30 seconds is normal Go runtime behavior, not a leak.

Common causes

CauseWhat it looks likeFirst thing to check
Slow upstream (blocked forward calls)Goroutines climb alongside P99 latency; count drops when upstream recoversPer-upstream latency: coredns_proxy_request_duration_seconds{to=...}
Upstream connection churnConnection cache misses rising; FDs climbing with goroutinescoredns_proxy_conn_cache_misses_total and process_open_fds
Traffic burst / retry stormGoroutines scale proportionally with QPS; no leak signatureQuery rate vs goroutine count ratio
Plugin goroutine leakCount ratchets up independent of QPS; never returns to baselinepprof goroutine profile; CoreDNS version vs known leak fixes
DoQ stream exhaustion (CVE-2025-47950, CVE-2026-32934)Explosive goroutine growth on versions before 1.12.2 / 1.14.3 with DoQ enabledCoreDNS version and whether the quic Corefile block is in use
Memory pressure feedback loopGoroutines, heap, and GC pauses all rising togethergo_memstats_heap_inuse_bytes and go_gc_duration_seconds

The version-specific leaks are worth naming because the fix is an upgrade, not a config change: the transfer plugin leaked a goroutine on AXFR errors before 1.12.4, and the secondary plugin leaked one on reload before 1.13.2. The DoQ server spawned an unbounded goroutine per QUIC stream before 1.12.2, and a regression (CVE-2026-32934) kept spawning unbounded waiter goroutines until 1.14.3. If you are on an affected version and the leak signature matches, upgrading is the fix.

Quick checks

These are all read-only and safe to run against a production instance.

# Current goroutine count
curl -s http://localhost:9153/metrics | grep '^go_goroutines'

# Per-upstream forward latency (which upstream is slow?)
curl -s http://localhost:9153/metrics | grep 'coredns_proxy_request_duration_seconds'

# Per-upstream health check failures
curl -s http://localhost:9153/metrics | grep 'coredns_proxy_healthcheck_failures_total'

# Forward concurrency backpressure (if max_concurrent is set)
curl -s http://localhost:9153/metrics | grep 'coredns_forward_max_concurrent_rejects_total'

# Heap and GC state
curl -s http://localhost:9153/metrics | grep -E '^(go_memstats_heap_inuse_bytes|go_gc_duration_seconds|process_resident_memory_bytes)'

# Open file descriptors (connection leak signature)
curl -s http://localhost:9153/metrics | grep '^process_open_fds'

# Query rate for proportionality check
curl -s http://localhost:9153/metrics | grep '^coredns_dns_requests_total'

Note on metric names: from CoreDNS 1.5.0 through 1.11.x the per-upstream metrics appear as coredns_forward_request_duration_seconds, coredns_forward_healthcheck_failures_total, and coredns_forward_conn_cache_misses_total; since CoreDNS 1.12.0 they are exported as coredns_proxy_request_duration_seconds{proxy_name="forward"}, coredns_proxy_healthcheck_failures_total{proxy_name="forward"}, and coredns_proxy_conn_cache_misses_total{proxy_name="forward"}. If a grep returns nothing, try the other name before concluding the signal is absent.

Also check the CoreDNS version against the leak fixes listed above, and look at the pod restart count in Kubernetes: repeated OOMKilled restarts with a climbing goroutine history before each kill is the leak archetype reaching its conclusion.

How to diagnose it

The core diagnostic is shape analysis: does the curve track load, or does it ratchet?

flowchart TD
  A[go_goroutines climbing] --> B{Tracks QPS and returns to baseline?}
  B -- Yes --> C[Blocked upstream calls]
  B -- No, ratchets up --> D[Goroutine leak]
  C --> E[Find slow upstream via per-upstream latency]
  E --> F[Remove or replace slow upstream]
  D --> G[Capture goroutine profile with pprof]
  G --> H{Stack trace points to a plugin?}
  H -- Yes --> I[Check version vs known leak fixes]
  H -- No --> J[Check DoQ enabled + version vs DoQ CVEs]
  I --> K[Upgrade CoreDNS]
  J --> K
  1. Establish the baseline. Look at go_goroutines over the last 24-72 hours, not just the incident window. You need to know what “normal” is for this instance before judging the spike. Baselines of 20-50 are typical; yours may differ with heavy watches or DoQ enabled.

  2. Test proportionality. Overlay the goroutine count against query rate (coredns_dns_requests_total) and P99 latency (coredns_dns_request_duration_seconds). If goroutines grew 10x while QPS grew 2x, queries are blocking, not just arriving. If goroutines grow with no QPS change at all, suspect a leak immediately.

  3. Check the recovery behavior. This is the single most diagnostic observation: after the load or latency spike ends, does the count return to baseline? Blocked-call accumulation drains within minutes of the upstream recovering. A leak does not drain. Watch the 10-30 minutes after the event.

  4. Identify the slow upstream. If it is blocked calls, break down coredns_proxy_request_duration_seconds by the to label. One upstream at 500ms+ while others are at 20ms tells you exactly where the backlog is forming. Cross-check coredns_proxy_healthcheck_failures_total{to=...} for the same upstream.

  5. Check memory correlation. Pull go_memstats_heap_inuse_bytes and process_resident_memory_bytes over the same window. Goroutines climbing with flat heap leans toward transient blocking. Goroutines climbing with monotonically rising post-GC heap minimums leans hard toward a leak, and tells you the runway to the OOM kill.

  6. Profile if it is a leak. If the count never drains, capture a goroutine profile from the Go pprof endpoint (served by the pprof plugin at http://localhost:6053/debug/pprof/goroutine?debug=1 by default). A leak shows thousands of goroutines parked in the same stack frame, usually a channel receive or network read inside one plugin. That stack tells you which plugin and whether it matches a known fix.

  7. Rule out the DoQ cases. If you serve DNS-over-QUIC and run a version before 1.12.2 (CVE-2025-47950) or before 1.14.3 (CVE-2026-32934), unbounded stream goroutines are a documented remote-trigger condition. Treat this as a security issue, not just a capacity issue.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
go_goroutinesThe symptom itself; proxy for concurrent in-flight work>2x baseline sustained 10 min; >5000 with rising memory
coredns_proxy_request_duration_seconds{to=...}Per-upstream latency; finds the slow upstream causing blockingOne upstream 3-5x slower than the others
coredns_dns_request_duration_secondsOverall P99; confirms user impact of blocked queriesP99 climbing in step with goroutines
go_memstats_heap_inuse_bytesLeak confirmation; goroutine stacks and allocations consume heapPost-GC minimum rising and never reclaiming
process_resident_memory_bytesWhat the OOM killer actually measuresTrending toward container memory limit
coredns_forward_max_concurrent_rejects_totalBackpressure signal if max_concurrent is configuredAny sustained nonzero rate
process_open_fdsConnection leak signature alongside goroutine leakClimbing without returning to baseline
go_gc_duration_secondsHeap pressure feedback into tail latencyPauses increasing as goroutines climb

Fixes

Slow upstream causing blocked calls

The direct fix is restoring upstream health: remove or replace the slow upstream in the Corefile forward block. The per-upstream to label tells you which one. If the slowness is transient (network path, upstream overload), the backlog drains on its own once latency normalizes.

If you need guardrails against recurrence, the forward plugin’s max_concurrent option caps in-flight forwarded queries and rejects excess with REFUSED. That is deliberate backpressure: bounded failures instead of unbounded goroutine growth. Size it at roughly 3x your peak QPS multiplied by average upstream latency. The tradeoff is that excess queries fail fast instead of waiting, which is usually what you want from a shared resolver.

A restart will clear the accumulated goroutines, but treat it as a last resort: it also flushes the cache, which creates a cold-cache thundering herd on your upstreams at the worst possible moment. If the upstream has recovered, the backlog drains without a restart.

Confirmed goroutine leak

There is no config fix for a plugin leak. The path is: identify the leaking stack from a pprof profile, match it to a known issue, and upgrade. Known fixed cases include the transfer plugin AXFR leak (1.12.4), the secondary plugin reload leak (1.13.2), and the DoQ stream goroutine exhaustion fixes (1.12.2 and 1.14.3). If your stack does not match a known fix, capture the profile and report it upstream with the version.

As an interim measure while you schedule the upgrade, you can raise the container memory limit to extend the runway, but that only delays the OOM kill. A leak does not self-heal. If the leak rate is fast, scheduled rolling restarts, staggered to avoid cold-cache stampedes, are a defensible temporary mitigation.

Memory pressure compounding the problem

If heap is climbing with the goroutines, you are on the OOM clock regardless of cause. Estimate runway as (container limit minus current RSS) divided by the post-GC minimum growth rate. If the runway is shorter than your time to fix, raise the limit temporarily as a defensive move while you work the actual cause.

Prevention

  • Alert on the shape, not just the value. Two alerts: goroutines >2x baseline sustained 10 minutes (catches both patterns early), and goroutines >5000 with rising RSS (catches the leak-before-OOM case). A single static threshold misses slow leaks under the line.
  • Monitor per-upstream latency as a leading indicator. Blocked-call accumulation always starts as upstream latency. Alerting on one upstream drifting 3x slower than its peers catches the problem before goroutines pile up.
  • Keep CoreDNS current. The documented goroutine leaks are all version-specific and fixed. Running an old version with the transfer, secondary, or quic plugins active is running known-leaky code.
  • Set max_concurrent deliberately if your upstreams are unreliable. Bounded REFUSED responses are a better failure mode than unbounded goroutine accumulation when an upstream drags.
  • Baseline your normal. Record the idle goroutine count, QPS, and heap for each deployment variant you run. “2x baseline” is only actionable if you know the baseline.
  • Track restarts against goroutine history. If a pod was OOMKilled, check whether goroutines were climbing before the kill. That retro-check often reveals a slow leak that restarts were masking.

How Netdata helps

  • Netdata charts go_goroutines at per-second resolution, which makes the distinguishing shapes visible: a spike that tracks a traffic burst and drains, versus a ratchet that never returns to baseline.
  • Correlating goroutine count with go_memstats_heap_inuse_bytes and process_resident_memory_bytes on one dashboard separates “blocked but transient” from “leaking toward OOM” without switching tools.
  • Per-upstream latency from coredns_proxy_request_duration_seconds{to=...} sits next to the goroutine curve, so you can see which upstream started dragging before the accumulation began.
  • GC pause duration (go_gc_duration_seconds) alongside heap shows the memory-pressure feedback loop that turns a goroutine problem into a latency problem.
  • Anomaly detection on go_goroutines flags deviations from the learned baseline, which catches slow leaks that sit under static thresholds for days.