The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-default-certificate-served ▌

Operations Guides

Traefik serving the default certificate: SNI and certificate selection

A client hits your site and gets a certificate warning. The certificate presented is a self-signed cert with CN “TRAEFIK DEFAULT CERT” instead of the certificate you configured. Everything else looks fine: Traefik is up, /ping returns 200, backends are healthy, routes work. The proxy is completely healthy and the only broken thing is which certificate it picked during the TLS handshake.

This failure is invisible to most Traefik monitoring. There is no error counter, no 5xx, no failed health check. The TLS handshake succeeds; it just succeeds with the wrong certificate. Unless you probe TLS externally, you find out from a user’s screenshot.

This article covers how Traefik selects which certificate to present, why it falls back to the default, and how to find the specific mismatch in your configuration.

What this means

Traefik terminates TLS at the entrypoint and chooses which certificate to present during the TLS handshake, before any HTTP routing happens. The selection input is the Server Name Indication (SNI) extension the client sends in the TLS ClientHello. Router rules like Host() are evaluated after the handshake completes, so they play no role in certificate selection. This is the most common misconception behind this incident.

The selection logic, in order:

  1. If the client sends an SNI server name and Traefik has a certificate whose CN or SANs match it, Traefik presents that certificate.
  2. If there is no SNI, or no configured certificate matches the SNI, Traefik falls back to the default certificate from the TLS store, if you configured one (via defaultCertificate or defaultGeneratedCert).
  3. If no default certificate is configured either, Traefik generates a self-signed certificate on the fly, with CN “TRAEFIK DEFAULT CERT”, and serves that.

So “serving the default certificate” is not an error state inside Traefik. It is the defined fallback for “I could not match this handshake to any certificate you gave me.”

flowchart TD
    A[Client TLS ClientHello with SNI] --> B{SNI matches a configured cert CN or SAN?}
    B -- yes --> C[Serve matched certificate]
    B -- no --> D{Default certificate configured in TLS store?}
    D -- yes --> E[Serve configured default certificate]
    D -- no --> F[Serve generated self-signed cert - TRAEFIK DEFAULT CERT]
    C --> G[Handshake completes - routing happens after]
    E --> G
    F --> G

One important consequence: the cert presented and the router that eventually handles the request are chosen independently. A request can get the correct certificate and still 404, or get the default certificate and then route perfectly once the client accepts it. Diagnose the certificate layer separately from the routing layer. See Traefik 404 not found for the routing side.

There is also one benign case: during first-boot ACME acquisition, Traefik serves the self-signed default certificate until issuance completes. If you see the default cert for the first minutes after a fresh deploy and then the real cert appears, that is expected behavior, not a fault.

Common causes

CauseWhat it looks likeFirst thing to check
No certificate matches the SNIClients for one specific hostname get the default cert; other hostnames on the same Traefik are fineopenssl s_client -servername for the failing hostname, then compare against certs Traefik has loaded
No default certificate configuredEvery unmatched or SNI-less connection gets the self-signed “TRAEFIK DEFAULT CERT”TLS store configuration: is defaultCertificate or defaultGeneratedCert set?
First-boot ACME acquisition in progressDefault cert served briefly after a fresh deploy, then replaced by the real certTraefik logs for ACME issuance; check again after a few minutes
Router without TLS enabled, or cert not associated with the routeHTTP works, HTTPS serves default cert for that route’s hostnameRouter’s TLS configuration and which certificate resolver or store it references
TLS store misconfigurationA configured defaultCertificate is silently ignored; generated cert still servedWhether the store definition is in the dynamic configuration and named default
Certificate loaded but for a different nameCert exists in Traefik but its CN/SANs do not cover the requested hostnameInspect the certificate’s SAN list against the failing hostname

Quick checks

All of these are safe and read-only.

# See which certificate Traefik presents for a specific hostname.
# This is the single most diagnostic command for this issue.
echo | openssl s_client -servername app.example.com -connect traefik-host:443 2>/dev/null \
  | openssl x509 -noout -subject -issuer -dates -ext subjectAltName

If the subject shows CN = TRAEFIK DEFAULT CERT, Traefik had no matching certificate for app.example.com and no configured default. If the subject is a real cert but for a different domain than you expected, your default certificate is configured and the SNI matched nothing specific.

# Check what an SNI-less client gets (approximates scanners and very old clients).
echo | openssl s_client -connect traefik-host:443 2>/dev/null \
  | openssl x509 -noout -subject -issuer
# List certificates Traefik has loaded, with expiry.
# Labels include cn and sans, so you can check coverage for the failing hostname.
curl -s http://localhost:8080/metrics | grep traefik_tls_certs_not_after
# Inspect loaded routers and confirm TLS is enabled on the route you expect.
curl -s http://localhost:8080/api/http/routers | jq '.[] | {name, rule, tls}'
# Check TLS versions and ciphers actually being negotiated, per entrypoint.
curl -s http://localhost:8080/metrics | grep traefik_entrypoint_requests_tls_total
# Search logs for ACME issuance activity if this is a fresh deployment.
# Exact wording varies by version.
grep -iE "acme|certificate" /var/log/traefik/traefik.log | tail -50

How to diagnose it

Work from the client inward. Each step narrows which layer dropped the match.

  1. Reproduce from a controlled client. Run the openssl s_client check above against the failing hostname, both through your normal path and directly against Traefik. If the direct connection gets the right cert but the load balancer path gets the default, something in front (another TLS terminator, a CDN) is intercepting, and Traefik is not your problem.

  2. Record exactly what was served. Subject, issuer, SAN list, expiry. Three outcomes: (a) self-signed “TRAEFIK DEFAULT CERT”, meaning no configured default and no SNI match; (b) a real certificate for a different domain, meaning a configured default is being served because the SNI matched nothing specific; (c) the correct certificate, meaning the problem is client-side (trust store, old CA bundle) rather than selection.

  3. Check what certificates Traefik actually has. Use traefik_tls_certs_not_after and compare the cn/sans labels against the failing hostname. If no loaded certificate covers it, the selection behavior is correct and the problem is upstream: the certificate was never obtained or never loaded. If ACME manages that cert, this becomes a renewal or issuance failure. See Traefik certificate expired and Traefik ACME rate limit.

  4. Confirm the route references TLS correctly. Via /api/http/routers, check that the router for the hostname has TLS enabled and references the resolver or store you expect. A route without TLS configuration will still terminate TLS at the entrypoint if the entrypoint itself does TLS, but with no per-route certificate association, selection falls to SNI matching against the store.

  5. Check the TLS store default. If you intended a specific certificate as the default (for SNI-less clients and unmatched names), verify defaultCertificate or defaultGeneratedCert is defined in the dynamic configuration and that the store is the one Traefik actually uses. If both are defined, defaultCertificate wins. If the store is defined in the wrong place (for example, in static config or a file the provider is not watching), it is silently absent.

  6. Rule out the benign case. If the affected hostname was deployed minutes ago and ACME is still acquiring the certificate, wait for issuance and re-run step 1. Check logs for challenge errors if it does not resolve. See Traefik ACME rate limit if issuance keeps failing.

  7. Check for version-specific provider and selection bugs. Traefik v3.7.0 requires Gateway API v1.5.1 CRDs; and, beginning with v3.7.10, mismatched Standard-channel v1.6.1 CRDs with TCPRoute support enabled can prevent the Gateway provider startup with no logged error, so no Gateway API resource is served and HTTPS falls back to the default self-signed cert. Verify the installed Gateway API CRD version and the Traefik migration notes. Upstream also fixed nondeterministic certificate selection when multiple loaded certificates share SANs; if your certificates overlap, test after upgrading to the patched line.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_tls_certs_not_after (labels: cn, sans, serial)Inventory of loaded certificates and their expiry; confirms whether the cert for the failing hostname exists at allHostname absent from the label set, or a cert covering it approaching expiry
traefik_entrypoint_requests_tls_total (labels: tls_version, tls_cipher)Confirms TLS handshakes are completing, and on which entrypoint; a mismatch between handshake success and client-reported errors points at selection, not transportHandshakes succeeding while clients report cert warnings
traefik_entrypoint_requests_total{code="404"}Rising entrypoint 404s alongside default-cert reports suggest routes are missing too, which points at provider desync rather than a cert problem404 rate rising on hostnames that previously worked
traefik_config_last_reload_successA frozen timestamp means stale configuration; a newly deployed cert or TLS store change may simply not be loadedTimestamp not advancing while config changes are being made
External synthetic TLS probes per production hostnameThe only signal that actually detects “wrong cert served”. Traefik has no metric for certificate selection outcomeProbe sees unexpected subject/issuer or a self-signed cert

Nothing in Traefik’s Prometheus metrics tells you which certificate was presented for a handshake. The selection failure is only observable from the client side. Without external TLS probes, you are blind to this entire failure class.

Fixes

Add the missing certificate for the hostname

If the SNI simply has no matching cert, obtain or load one that covers the hostname. For ACME-managed domains, fix whatever is blocking issuance (challenge reachability, rate limits, corrupted acme.json). For manually managed certs, make sure the certificate and key are defined where your provider can load them, and that the SAN list covers every hostname you serve, including bare-domain and www variants.

Tradeoff: none. This is the correct fix when the inventory check showed the cert missing.

Configure a deliberate default certificate

If you serve many hostnames and want unmatched or SNI-less connections to get a real, valid certificate instead of the self-signed fallback, set defaultCertificate (or defaultGeneratedCert for a generated cert with controlled parameters) in the TLS store. The default certificate is what non-SNI clients and scanners will see.

Tradeoff: whatever you choose as the default is publicly visible for any unmatched name, so pick something you do not mind being associated with arbitrary traffic. It also masks misconfigurations: a hostname with a missing cert now serves the default instead of failing loudly, which delays detection of the missing cert.

Enable strict SNI if you want mismatches to fail closed

If your policy is “no valid cert, no handshake”, strict SNI checking makes Traefik refuse the handshake when no certificate matches, instead of serving the default. This converts silent wrong-cert serving into an immediate, loud connection failure, which is easier to alert on and does not leak a certificate for the wrong name.

Tradeoff: clients without SNI (rare, but they exist) and any hostname with a missing cert get a hard failure instead of a working-but-wrong-cert connection. During a cert outage, the blast radius is total for the affected name rather than degraded. Set sniStrict: true under tls.options.<name>; use the default TLS option for the generic fallback and a named option only for routers that reference that option.

Fix the TLS store or route association

If a configured default is being ignored, verify the store definition lives in the dynamic configuration (file provider, or the appropriate CRD on Kubernetes), that it is the store Traefik actually consults, and that the route in question has TLS enabled with the expected reference. Remember the selection order: per-SNI match first, store default second, generated cert last.

Wait out first-boot ACME, then investigate if it persists

For the benign first-boot case, no action is needed if the real cert appears within a few minutes. If the default cert persists, treat it as an ACME issuance failure and debug the challenge path.

Prevention

  • Probe TLS externally, per production hostname. This is the only reliable detector for wrong-cert serving. Check subject, issuer, expiry, and SAN coverage, not just “handshake succeeded”. A probe that only checks TCP 443 open will never catch this.
  • Alert on traefik_tls_certs_not_after well before expiry. A missing or expiring cert today is a default-cert incident next week. Traefik starts renewal attempts 30 days out for Let’s Encrypt, so treat renewal as broken long before expiry; ticket at under 7 days remaining.
  • Decide your default-cert policy explicitly. Either configure a deliberate default certificate, or enable strict SNI so mismatches fail closed. The worst option is the implicit one: the self-signed fallback served to real users while nobody watches.
  • Baseline which hostnames should serve which certificates and alert when the served subject changes. Cert changes outside a deployment window are a signal worth paging on.
  • After Traefik upgrades, verify certificate serving before declaring success. Version-specific selection bugs and provider sync failures (the v3.7 Gateway API case above) present as default-cert serving with green health checks. A one-line openssl s_client check per hostname in your post-upgrade checklist catches this class entirely.

How Netdata helps

Netdata surfaces the signals that bracket this failure, even though the selection decision itself is only client-observable:

  • Certificate inventory and expiry via traefik_tls_certs_not_after, with per-certificate CN and SAN labels, so you can confirm whether a cert for the failing hostname is loaded and how long it has left.
  • TLS negotiation breakdown via traefik_entrypoint_requests_tls_total, showing handshake volume by TLS version and cipher per entrypoint, which confirms handshakes are completing while clients complain.
  • Config freshness via traefik_config_last_reload_success, so you can tell immediately whether a newly added certificate or TLS store change was ever loaded, or whether Traefik is running stale config.
  • Entrypoint 404 rate to distinguish a pure certificate problem from a wider routing or provider desync problem, since the two often travel together after a bad deploy or provider outage.
  • Correlation across these signals in one view: default-cert reports plus a frozen config timestamp plus missing cert inventory points at provider desync; default-cert reports plus a present-but-expiring cert points at ACME renewal failure. Seeing them together is what turns a user screenshot into a root cause in minutes.