The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-dashboard-api-exposed ▌

Operations Guides

Traefik dashboard and API exposed: your routing table on the public internet

If Traefik’s dashboard, API endpoints (/api/rawdata, /api/http/services, /api/http/routers), or /debug/pprof answer requests from the public internet, you are publishing a map of your infrastructure. The API returns every router rule, backend server URL, middleware configuration, and upstream health status. /debug/pprof adds Go runtime profiling data. Automated scanners fingerprint Traefik continuously, and the dashboard API has a history of information-disclosure issues. The fix is usually a ten-minute configuration change, but only if you know the exposure exists. Many teams never test from outside their own network.

This guide covers how to probe for the exposure from an untrusted network, the common misconfiguration patterns, and how to lock the management surface down with an IP allowlist, authentication, network policy, or a non-public entrypoint.

What this means

Traefik has two planes. The data plane serves your entrypoints (ports 80/443) and proxies traffic to backends. The control plane serves the dashboard, the REST API, and optionally the Go debug endpoints. The control plane is where the sensitive material lives:

  • /api/rawdata returns the complete dynamic configuration: routers, services, middlewares, their status, and dependency relations. In effect, your entire routing table.
  • /api/http/services reveals every backend server URL, so an attacker learns internal hostnames, ports, and IP addresses directly.
  • /api/http/routers shows every routing rule, including rules for internal-only services that should not be publicly known.
  • /dashboard/ is the web UI backed by the same API.
  • /debug/pprof/ (when api.debug is enabled) exposes Go profiling data: goroutine dumps, heap profiles, CPU profiles. Beyond reconnaissance value, profiling endpoints can be abused to burn CPU on the proxy itself.

An attacker with this map knows which backends exist, which are unhealthy, which middleware protects which route, and where the soft targets are. In some configurations the API also exposes certificate details. Combined with known Traefik header- and routing-trust vulnerabilities in unpatched versions, an exposed API is both an information leak and an attack surface. CVE-2024-45410 lets an HTTP client remove or manipulate Traefik-added forwarded headers; GitHub Advisory GHSA-62c8-mh53-4cqv lists v2.11.9 and v3.1.3 as first patched versions.

The rule is simple: the dashboard entrypoint must never be published to the public internet. Not “protected later”, not “obscured on a weird port”. Never published.

flowchart LR
  Internet[Public internet] -->|probe /api, /dashboard, /debug| Entry{Entrypoint}
  Entry -->|published, no auth| API[Traefik API + dashboard]
  Entry -->|allowlist + auth, or not published| Deny[403 / connection refused]
  API --> Raw[/api/rawdata: full routing table/]
  API --> Svcs[/api/http/services: backend URLs + health/]
  API --> Pprof[/debug/pprof: Go profiling/]

Common causes

CauseWhat it looks likeFirst thing to check
api.insecure=true left onAPI and dashboard served directly on the traefik entrypoint (default port 8080), no authentication, middleware does not applyStatic config: api.insecure or --api.insecure flag
Dashboard entrypoint bound to a public interfacePort 8080 (or the custom dashboard port) reachable from the internet via the load balancer or host firewallProbe the public IP on the dashboard port from outside
Kubernetes Service exposing the dashboard portA Service of type LoadBalancer or NodePort includes the dashboard/API port alongside 80/443kubectl get svc and check the published ports
Router for the API without auth middlewareapi@internal routed on a public entrypoint with no BasicAuth, ForwardAuth, or IPAllowList attachedDynamic config for the router targeting api@internal
api.debug=true in production/debug/pprof/ and /debug/vars answering on the API surfaceStatic config: api.debug or --api.debug flag
Firewall/security group wider than intendedThe management port is firewalled on some paths but the cloud security group allows 0.0.0.0/0Probe from a genuinely external network, not from a VPN

A subtle trap with api.insecure=true: it bypasses middleware entirely. Operators often enable insecure mode “temporarily”, then try to bolt on BasicAuth via labels or dynamic config, and it silently does nothing because the insecure API handler never passes through the middleware chain. If you want middleware-based protection, you must disable insecure mode and route api@internal explicitly.

Quick checks

Run these from a network that is genuinely outside your perimeter: a personal hotspot, a VPS in another cloud, or a scanner host. Probing from the office LAN or VPN will give you false reassurance.

# Probe the management surface from an external host.
# Any 200 here from the public internet is a finding.
for path in /api/rawdata /api/http/services /api/http/routers /dashboard/ /debug/pprof/; do
  code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 "http://traefik-public-ip:8080${path}")
  echo "${code}  ${path}"
done
# Check whether the dashboard port is even open on the public address
nc -zv traefik-public-ip 8080
# On the host or in the pod: inspect how the API is configured
# Static config flags (process args)
ps aux | grep -o '\-\-api[^ ]*' 

# Or in the static config file
grep -A5 '^api:' /etc/traefik/traefik.yml
# Kubernetes: check which ports the Traefik Service publishes
kubectl -n traefik get svc -o wide
kubectl -n traefik get svc traefik -o jsonpath='{.spec.ports[*].port}'
# See what the API actually discloses (run against the internal address)
curl -s http://localhost:8080/api/http/services | head -c 2000

Interpretation: 401 or 403 from outside means some protection exists (verify it is real and not bypassable). 404 or connection refused means the surface is not published. 200 means you have an exposure to fix now.

How to diagnose it

  1. Establish the external truth first. Run the probe loop above from an untrusted network. Record which paths return 200, which return 401/403, and which are unreachable. This is your exposure inventory.

  2. Determine how the API is being served. Check the static configuration for api.insecure. If it is true, the API is bound to the traefik entrypoint directly and no middleware can protect it. If it is false (the default), look for a router in the dynamic configuration that targets the api@internal service and note which entrypoints it is attached to.

  3. Map the network path. If the API answers on the public IP, find out why: is the dashboard entrypoint bound to 0.0.0.0 and published by a cloud load balancer? Is there a Kubernetes Service or Ingress that includes the management port? Is a security group allowing the port from 0.0.0.0/0?

  4. Check whether debug mode is on. Look for api.debug=true in the static config or --api.debug in the process arguments. If /debug/pprof/ answered in step 1, treat it as part of the exposure even if the dashboard itself is protected.

  5. Check the Traefik version. An exposed management surface on an unpatched version compounds the risk. Recent examples include path-normalization bypass CVE-2025-66490 (fixed in v2.11.32 and v3.6.4), forwarded-header alias spoofing and StripPrefixRegex bypass CVE-2026-39858/CVE-2026-40912 (fixed in v2.11.43, v3.6.14, and v3.7.0-rc.2), and the CVE-2024-45410 lowercase-token bypass CVE-2026-29054 (fixed in v2.11.38 and v3.6.9). If you must keep the API reachable from anywhere beyond localhost, being current on patches is not optional.

  6. Review the audit trail. If the API was exposed, pull access logs (Traefik’s and any upstream LB logs) and look for requests to /api/*, /dashboard/*, and /debug/* from source IPs you do not recognize. Assume anything returned by /api/rawdata is now known.

Metrics and signals to monitor

This exposure is primarily detected by external probing and network audit, not by metrics. But Traefik’s Prometheus metrics give you supporting evidence of probing and exploitation attempts:

SignalWhy it mattersWarning sign
Entrypoint 404 rate (traefik_entrypoint_requests_total{code="404"})Scanners enumerating paths and hosts show up as unmatched requests at the entrypointSudden increase over 3x baseline, especially concentrated on unknown Host headers
401/403 rate (traefik_entrypoint_requests_total{code=~"40[13]"})Tells you whether auth middleware on the API is being exercised or attackedSustained rate above 5x baseline from concentrated sources
Requests hitting the management entrypointA management entrypoint that should see near-zero traffic should see near-zero trafficAny sustained request rate on the dashboard/API entrypoint from outside your admin ranges
Access log review of /api/*, /debug/* pathsConfirms actual reads of the routing table, not just probes200 responses to these paths from unrecognized source IPs

Source IP caveat: if Traefik sits behind a CDN or cloud load balancer, access log source IPs are the intermediary’s. Parse X-Forwarded-For for the real client address.

Fixes

Disable insecure mode and route the API deliberately

Turn off api.insecure. It exists for local development, and it disables any possibility of middleware-based protection. In production, the API and dashboard are disabled by default (api: false); if you need them, enable the API and expose api@internal through an explicit router in the dynamic configuration, with protection attached. That router is where your defenses live.

Attach an IPAllowList middleware

Restrict the management router to known admin networks using the IPAllowList middleware with sourceRange in CIDR notation. The default rejection status is 403; rejectStatusCode can be changed if you prefer to return 404 and not confirm the endpoint exists. If Traefik is behind a proxy or load balancer, configure ipStrategy (depth or excludedIPs) so the allowlist evaluates the real client IP from X-Forwarded-For rather than the LB address, otherwise you either allow nothing or allow everything.

Add authentication as a second layer

IP allowlisting plus BasicAuth or ForwardAuth on the management router gives you defense in depth: a network change alone does not open the door. Put auth middleware before any path-manipulation middleware in the chain; several published auth-bypass CVEs exploit the strip-before-auth ordering. Never ship default or placeholder credentials.

Bind the dashboard to a non-public entrypoint

The cleanest fix is architectural: serve the dashboard and API on a dedicated entrypoint bound to localhost or an internal interface only, and reach it through a VPN, bastion, or kubectl port-forward. If the port is never published by a load balancer, Service, or security group, there is nothing to misconfigure later. In Kubernetes, double-check that no Service of type LoadBalancer or NodePort includes the management port.

Turn off debug mode

Set api.debug=false (the default) in production. If you need pprof for an investigation, enable it temporarily on an internal-only entrypoint and turn it back off when done. Go profiling endpoints are both a disclosure risk and a cheap way to burn proxy CPU.

Rotate what the exposure leaked

If the API was publicly reachable for an unknown period, treat the leaked data as known: internal hostnames, ports, backend topology, and any secrets that were visible in the dynamic configuration. Review access logs for reads of /api/*, and patch Traefik to a current release before re-exposing anything, even internally.

Prevention

  • Probe continuously, not once. Add an external check that periodically requests /dashboard/ and /api/rawdata against your public addresses and alerts if any return 200. Exposure often reappears after a Helm upgrade, a new Service, or a security group change.
  • Keep the API disabled unless you use it. api: false is the default. If nobody on the team queries the API or opens the dashboard, leave it off.
  • Deny by default in network policy. In Kubernetes, use NetworkPolicy so only intended namespaces can reach the management port. At the cloud layer, scope the security group for the management port to admin CIDRs, or do not create a rule at all.
  • Review middleware ordering in CI. Lint dynamic configuration for routers serving api@internal and verify auth and allowlist middleware are present and ordered before any path rewriting.
  • Patch Traefik on a schedule. Auth-middleware bypass fixes land in patch releases regularly. An allowlist you trust today is only as strong as the proxy version enforcing it.

How Netdata helps

  • Netdata collects Traefik’s Prometheus metrics per second, so a sudden burst of entrypoint 404s or 401/403s from scanner activity is visible as it happens, not five minutes later.
  • Per-entrypoint request charts make it obvious when a management entrypoint that should be idle starts serving real traffic, which is exactly what an accidental exposure looks like from the inside.
  • Code-level breakdowns of traefik_entrypoint_requests_total let you separate scanner noise (404s on unknown hosts) from auth pressure (401/403 on the API router) without touching raw logs.
  • ML-based anomaly detection on request rates and error ratios catches the low-and-slow probing pattern that static thresholds miss, like a scanner spreading requests across hours.
  • Alerting on the 4xx ratio at the entrypoint level gives you an early tripwire when new paths, including management paths, start receiving external attention after a config or network change.