The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-scanning-probing ▌

Operations Guides

Traefik under scanning: exploit-path probing and request smuggling signals

Any Traefik instance with a public entrypoint is being scanned right now. Requests for /.env, /.git/config, /wp-admin, /phpMyAdmin, and /actuator arrive continuously from botnets and research scanners, and on a correctly configured instance they all get the same answer: a 404 generated by Traefik itself, because no router matches. That is normal background radiation, and paging on it will burn out your on-call in a week.

The operator problem is not “are we being scanned” but “did anything change that makes this scan dangerous.” The three states that matter are: a probe for a sensitive path returns 200 with content, scanning volume or targeting shifts from generic to focused, or request patterns show smuggling or SSRF indicators that Go’s strict parser does not fully neutralize because of what sits behind Traefik. This guide covers how to separate those states from noise, what to check first, and which signals are worth alerting on.

What the baseline looks like

When a request matches no router, Traefik returns 404 at the entrypoint level, before any service is selected. This shows up in traefik_entrypoint_requests_total{code="404"} and in the access log. It does not appear in service-level metrics, because no service was ever chosen. This is the exact opposite of a backend returning 404 for a missing resource, which shows up in traefik_service_requests_total{code="404"} and means the request was routed. Confusing these two layers sends investigations to the wrong team, and the distinction is your first scanning filter: entrypoint 404s on paths like /.env are uninteresting, service 404s or 200s on those paths are not.

The access-log discriminator is RouterName/ServiceName plus OriginStatus. For an unmatched request, RouterName and ServiceName are empty and there is no origin status. For a backend-generated 404, those router/service fields are populated and OriginStatus is 404.

In JSON access logs, use those fields together: an unmatched request has neither a router nor an origin response. If they are not empty for a request to /.git/config, that request was routed somewhere, and you need to find out where.

Noise versus a targeted attack

The useful triage axis is not the path list, which every scanner on the internet shares, but volume, diversity, and success.

SignalBackground noiseTargeted or escalated
Source distributionMany IPs, low request count eachOne or few IPs, high request count
Path diversitySmall set of well-known pathsMany distinct paths, or paths matching your real route structure
Response codesOverwhelmingly entrypoint 404Probes that return 200, 301, or 403 instead of 404
TimingConstant low rate, day and nightBursts, or scans that start right after a new route or DNS record is published
User agentsKnown scanner agents, empty agentsBrowser-like agents, or agents mimicking your own clients

The single highest-severity condition: a request for a path that should not exist returns 200 with content. A /.env that returns environment variables, a /.git/config that returns repository configuration, or an /actuator endpoint that returns Spring internals is not scanning noise anymore. It is an exposure, and it usually means a backend was deployed with a permissive route (for example a broad PathPrefix("/") router) that forwards everything to an application serving those files.

Also watch for requests that get past a middleware chain you expected to block them. Encoded path variants (such as requests containing %2e%2e or encoded slashes) have historically bypassed routing rules and middleware on specific Traefik versions. If a request path in the access log contains encoded traversal or encoded restricted characters and the response is not a 404, check your Traefik version against the security advisories for path-traversal and path-normalization bypasses. Advisories for these patterns appeared in 2025 and continued in 2026 across the v2.11 and v3.x lines, so an unpatched instance may be routing requests its configuration says it should not.

Quick checks

These are read-only and safe to run during an incident.

# Probe paths that must always return 404 from the edge
for p in /.env /.git/config /wp-admin /actuator /phpMyAdmin; do
  printf '%s -> ' "$p"
  curl -s -o /dev/null -w '%{http_code}\n' "https://your-traefik-host$p"
done

# Entrypoint-level 404 rate (scanning pressure and route correctness)
curl -s http://localhost:8080/metrics | grep 'traefik_entrypoint_requests_total' | grep 'code="404"'

# Service-level 404s and 4xx (backend-generated; different meaning)
curl -s http://localhost:8080/metrics | grep 'traefik_service_requests_total' | grep -E 'code="4..'

# Scan for exploit-path hits in a JSON access log
jq -r 'select(.RequestPath | test("(\\.env|\\.git|wp-admin|phpMyAdmin|actuator)")) |
  [.ClientHost, .RequestPath, .DownstreamStatus] | @tsv' /var/log/traefik/access.log | \
  sort | uniq -c | sort -rn | head -30

# Find sensitive-path requests that did NOT return 404 (the ones that matter)
jq -r 'select(.RequestPath | test("(\\.env|\\.git|actuator)")) |
  select(.DownstreamStatus != 404) |
  [.ClientHost, .RequestPath, .DownstreamStatus, .RouterName] | @tsv' /var/log/traefik/access.log

# Check for odd Host headers (SSRF probing, host-header attacks)
jq -r '.RequestHost' /var/log/traefik/access.log | sort | uniq -c | sort -rn | head -20

Two notes on these checks. First, the jq field names depend on your access log configuration; Traefik’s JSON access log field names are version- and configuration-dependent, so adjust to what your log actually emits. Second, if Traefik sits behind a CDN or cloud load balancer, ClientHost will be the CDN’s address. You need the X-Forwarded-For or X-Real-IP header captured in the access log to see the real source.

Triage flow

flowchart TD
  A[Spike in exploit-path requests] --> B{Response codes?}
  B -->|All entrypoint 404| C{Volume or targeting changed?}
  B -->|200 / 301 / 403 on sensitive path| D[PAGE: exposure or routing bypass]
  C -->|No, steady background rate| E[Log and ignore; this is internet noise]
  C -->|Yes, single source or many distinct paths| F[TICKET: targeted reconnaissance]
  D --> G[Identify matched router in access log]
  D --> H[Check Traefik version against path-traversal advisories]
  F --> I[Correlate source IP across 404 and 403 counts]
  F --> J[Check for smuggling indicators in request headers]
  G --> K[Fix over-broad route or remove exposed file from backend]

Request smuggling indicators

Request smuggling is a different class of problem from path scanning. It exploits disagreement between how the front-end proxy and the back-end server parse a single HTTP request, so the backend sees a different request boundary than Traefik did. A successful smuggle can bypass authentication middleware, poison caches, or reach routes that Traefik’s router chain would have rejected.

Traefik is built on Go’s net/http, which has strict HTTP parsing. Go prioritizes Transfer-Encoding over Content-Length per RFC 9112 and rejects conflicting multiple Content-Length values. This blocks most classic CL.TE and TE.CL smuggling variants at the entrypoint, before a router is even evaluated. If your backends also parse strictly, the front door is largely closed.

The residual risk lives in two places:

  • Lenient backends. Node.js and some Java servlet containers accept request shapes that Go rejects or normalizes. If Traefik normalizes a request one way and the backend interprets it another way, the middleware chain (authentication, IP allow lists, rate limits) can be bypassed for the smuggled request.
  • HTTP/2 to HTTP/1.1 downgrade. Clients commonly speak HTTP/2 to Traefik while Traefik speaks HTTP/1.1 to backends. HTTP/2 frame boundaries are unambiguous, but after downgrade the request is re-serialized into HTTP/1.1 headers, and header ambiguity (H2.CL and H2.TE style disagreements) re-enters the picture. Go’s strict parser does not protect you here because the disagreement is between the downgraded byte stream and the backend’s parser.

What to look for in access logs and request data:

  • Requests with conflicting Content-Length and Transfer-Encoding headers. Standard access logs do not capture request headers by default; you need access log fields configured to include them, or a middleware that inspects them.
  • Host headers that do not match any configured router, especially Host values containing private IP addresses, localhost, or internal service names. This is SSRF and host-header probing.
  • Header names using underscore aliases of trusted forwarding headers (for example X_Forwarded_Host instead of X-Forwarded-Host). Backends that normalize underscores and dashes equivalently can be tricked into trusting injected forwarding context. Check your Traefik version against the header-sanitization advisories.
  • Unusually long header values or methods your service never legitimately receives (TRACE, DELETE on a read-only API).

Keep patching current. In 2025, a critical Go net/http smuggling CVE (CVE-2025-22871, bare LF in chunked coding) was disclosed; Traefik path-normalization advisories appeared in 2025 and additional ones followed in 2026 across supported release lines. An edge proxy that is six months behind on patches has known, published bypasses that scanners actively probe for. Traefik’s default Server response header also discloses version information, which makes version-targeted scanning cheaper for an attacker.

Signals to monitor

SignalWhy it mattersWarning sign
traefik_entrypoint_requests_total{code="404"}Baseline scanning pressure; also catches route lossSustained rate more than 5% of total entrypoint requests, or sudden 3x jump over baseline
traefik_service_requests_total{code=~"4.."}401 spikes suggest credential stuffing; 403 spikes suggest authorization probingSustained 4xx rate more than 5x baseline
Sensitive-path requests returning non-404 in access logsThe difference between noise and exposureAny 200 with content on /.env, /.git/*, /actuator
Per-source 404/403 aggregation in access logsSeparates distributed noise from one focused scannerA single source generating hundreds of 404/403 across many distinct paths in a short window
traefik_entrypoint_requests_tls_total by version and cipherSudden appearance of unusual ciphers or old TLS versions can indicate downgrade probingSustained TLS 1.0/1.1 traffic, or abrupt cipher distribution shift
Host header distribution in access logsSSRF and host-header attack surfaceHost values that are IPs, localhost, or unknown domains at non-trivial volume

A practical detection rule for focused scanning, adapted to whatever log pipeline you run: alert when a single source address produces more than roughly a hundred 404 or 403 responses across more than roughly fifty distinct paths within a polling window. Distributed botnets each send a handful of requests and never trip a per-source threshold; a human-driven or single-origin scanner trips it quickly. Tune both numbers to your traffic.

Response and hardening

Fix the exposure, not the scanner. If a sensitive path returns 200, the fix is on your side: remove the file from the backend image, tighten the router so only intended paths match, or add an explicit deny rule. Blocking the scanning IP is cosmetic; the next botnet arrives in minutes.

Verify middleware actually applies to the probed path. Encoded-path bypasses mean a request can reach a backend without passing through the middleware chain attached to the “obvious” router. Test with encoded variants of your protected paths, not just the plain form, and patch Traefik if the behavior differs.

Reduce fingerprintability. Suppress or genericize the Server header so version-targeted scanning gets less signal. Confirm the dashboard, /api/*, and /debug/pprof/* are not reachable from the public network; an exposed API hands an attacker your complete route table and backend addresses.

Constrain what backends accept. Smuggling succeeds in the gap between parsers. Where possible, terminate strict parsing at the backend too, reject ambiguous requests at the application layer, and avoid forwarding hop-by-hop header ambiguity through the downgrade path.

Do not alert on raw 404 volume. Entrypoint 404s from scanning are constant. Alert on state changes: non-404 responses to sensitive paths, per-source threshold breaches, and 404-rate shifts that correlate with recent deployments (which suggest route misconfiguration rather than scanning; see the related guide on unmatched requests).

How Netdata helps

  • Netdata collects Traefik’s Prometheus metrics per second, so the code label on entrypoint and service request counters lets you split Traefik-generated 404s (scanning, route loss) from backend-generated 404s (application behavior) without log spelunking first.
  • Per-code request rate dashboards make the scanning baseline visible, so a deviation, whether a focused scanner or a sudden 401/403 burst, stands out against established normal rather than a static threshold.
  • ML anomaly detection on the entrypoint 404 and 4xx series flags rate shifts that correlate with deployments, which is how you catch “this 404 spike is a broken route, not a botnet.”
  • Correlating entrypoint 404 rate with traefik_config_last_reload_success on one dashboard separates scanning noise from provider desync, which produces the same 404 symptom for a completely different reason.
  • TLS version and cipher distribution charts surface downgrade-probing patterns alongside the request-rate data, so security signals and traffic signals are compared in the same time window.