The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-certificate-expired ▌

Operations Guides

Traefik certificate expired: when ACME renewal has been failing silently

Browsers are rejecting your site with a certificate error, and Traefik is the TLS terminator. The process is up, /ping returns 200, traffic flows on port 80, and every HTTPS client reports an expired or expiring certificate. This is the signature of silent ACME renewal failure.

The nasty part is the timeline. Traefik attempts renewal 30 days before expiry. For a 90-day Let’s Encrypt certificate, that means renewal has been attempted since day 60. If a certificate is 7 days from expiry, renewal has already been failing for roughly 53 days. There is no Prometheus metric for ACME failure. The only metric you get is the consequence: traefik_tls_certs_not_after sliding toward the current time while nothing alerts.

This guide covers confirming the failure, finding the cause in the logs, forcing renewal, and closing the monitoring gap so you catch the next failure at day 61 instead of day 89.

What this means

Traefik’s ACME resolver runs as a background process. It renews certificates, writes them to its store (acme.json or a KV backend), and the TLS handshake layer keeps serving whatever certificate it has. Serving and renewing are decoupled: Traefik will serve an expiring, and eventually expired, certificate while the renewal loop errors in the background. If Let’s Encrypt is unreachable, Traefik falls back to previously generated certificates, then expired ones, then any provided certificates. There is no circuit breaker that stops serving a dying cert.

The failure is silent by design: renewal errors go to the log, not to any metric. The system “works” right up until client trust stores reject it.

flowchart LR
  A["Day 0: cert issued, 90-day validity"] --> B["Day 60: renewal window opens"]
  B --> C["Renewal failing, retries log errors only"]
  C --> D["Day 83: less than 7 days left"]
  D --> E["Day 90: expiry, browsers reject"]
  B -.->|healthy path| F["Renewal succeeds, not_after jumps forward"]

The moment you are debugging an expired certificate, you are at the end of a multi-week failure. Your job is twofold: fix the renewal now, and move detection to the start of the window.

Common causes

CauseWhat it looks likeFirst thing to check
HTTP-01 challenge blockedFirewall, security group, or another service took port 80; challenge never completesIs port 80 reachable from the internet and answered by Traefik?
DNS-01 challenge failurePropagation timeouts; errors like propagation: time limit exceeded in logsIs the DNS provider API credential still valid and the env var named correctly?
Rotated DNS API tokenRenewal worked for months, broke after a credential rotation, no config change on your sideTraefik uses Lego env var names; verify them against the Lego provider docs
Let’s Encrypt rate limiturn:ietf:params:acme:error:rateLimited in logs; often after frequent restarts or many domainsCount recent issuances; limits include 50 certs per registered domain per week and 5 duplicate certs per week
acme.json permissions not 600Traefik skips the resolver or logs “permissions … are too open”; renewal silently stopsls -l on acme.json, must be exactly 600
acme.json corruptedHard failure at startup; mid-operation corruption leaves in-memory certs serving while renewal silently failsParse the file as JSON; check startup logs
Staging CA left configuredcaServer points at Let’s Encrypt staging; renewals “succeed” but browsers reject the Fake LE chainCheck the caServer value and the cert issuer
“No ACME certificate generation required”DEBUG-level log line; Traefik found a matching (possibly stale or wildcard) cert in the store and skippedInspect acme.json for stale entries covering the domain
HA ACME lock stuckMulti-instance with Consul/etcd lock; one instance crashed holding the lock; logs show lock acquisition failureInspect and clear the stale lock in the KV store

Quick checks

All read-only and safe to run during the incident.

# What cert is actually being served, and when does it expire?
echo | openssl s_client -servername your.domain.com -connect your.domain.com:443 2>/dev/null \
  | openssl x509 -noout -dates -issuer -serial

# What does Traefik think it has? Expiry timestamps per CN/serial/SANs.
curl -s http://localhost:8080/metrics | grep traefik_tls_certs_not_after

# Seconds remaining per certificate (run in Prometheus):
# traefik_tls_certs_not_after - time()

# acme.json permissions: must be exactly 600
ls -l /path/to/acme.json

# acme.json parses as valid JSON
python3 -c "import json; json.load(open('/path/to/acme.json'))" && echo OK

# Recent ACME errors in the logs (pick your variant)
grep -iE 'acme|renew|challenge' /var/log/traefik/traefik.log | tail -50
# docker logs traefik --since 168h 2>&1 | grep -iE 'acme|renew'
# kubectl logs deploy/traefik -n traefik --since=168h | grep -iE 'acme|renew'

To inspect every certificate stored in acme.json and its real expiry (requires the Python cryptography package):

python3 -c "
import json, base64
from cryptography import x509
with open('/path/to/acme.json') as f:
    data = json.load(f)
for resolver in data:
    for cert in data[resolver].get('Certificates', []):
        c = x509.load_pem_x509_certificate(base64.b64decode(cert['certificate']))
        print(cert['domain']['main'], 'expires', c.not_valid_after_utc)
"

How to diagnose it

  1. Confirm which certificate is being served. The openssl s_client check above shows the live cert’s expiry, issuer, and serial. If the issuer is a staging CA (“Fake LE Intermediate” or similar), the resolver is pointed at the staging server and every “successful” renewal has been producing unusable certs.

  2. Check what the store holds. Compare the served serial against the serials in traefik_tls_certs_not_after and in acme.json. If the store contains a newer, valid cert that is not being served, the problem is cert selection, not renewal. If the store only has the expiring cert, renewal itself is broken.

  3. Read the renewal errors. Search the logs for “Error renewing certificate” and ACME challenge errors. The error string routes you:

    • rateLimited means Let’s Encrypt quota; no amount of restarting helps. Wait out the window or deploy a stop-gap cert.
    • propagation: time limit exceeded or DNS errors mean DNS-01 is broken: credentials, resolver reachability, or the DNS provider API.
    • Connection or timeout errors on the HTTP-01 path mean port 80 is not reaching Traefik.
    • “permissions … are too open” means acme.json mode is wrong.
    • “unable to acquire lock” in HA setups means a stale distributed lock.
  4. Verify the challenge path manually. For HTTP-01, confirm port 80 is reachable from outside and terminates on this Traefik instance. For DNS-01, confirm the DNS provider credentials work and that Traefik’s environment uses the variable names Lego expects, which are not Traefik-specific names. The Lego provider documentation is the authoritative list.

  5. Rule out the store. If the logs show nothing useful, check acme.json permissions and integrity. Also look for stale entries: Traefik renews certificates that are no longer referenced, which burns rate limit quota and can mask the real problem.

  6. Check the DEBUG-level skip. If “No ACME certificate generation required for domains” appears only at DEBUG level, Traefik found an existing certificate in the store it considers a match (including a wildcard or stale entry covering the hostname) and skipped issuance entirely. Cleaning the stale entry or correcting the router’s certificate resolver is the fix.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_tls_certs_not_afterThe only metric in the whole failure chain; expiry timestamp per cn, serial, sansWithin 30 days: renewal window open. Within 14 days: renewal has failed. Within 7 days: it has failed for ~53 days
Certificate serial change over timeA successful renewal produces a new serial and a jump forward in not_afterSame serial for weeks inside the renewal window means renewal is not happening
ACME error lines in logsThe only place the failure cause existsAny “Error renewing certificate”, challenge error, or rateLimited line
External TLS probes per production hostnametraefik_tls_certs_not_after covers all certs in the store, including dormant ones; it cannot tell you what is actually servedProbe shows an actively served cert nearing expiry: page-worthy
traefik_entrypoint_requests_total on the HTTPS entrypointCollapse after expiry confirms user impactSudden drop on 443 while port 80 is normal

One caveat on traefik_tls_certs_not_after: after a successful renewal, the old certificate’s series may not be cleaned up, so Prometheus can show multiple serials for the same CN and produce false-positive expiry alerts. A common workaround is alerting on topk(1, traefik_tls_certs_not_after) - time() so only the newest serial counts. As Traefik v3.7.12, issue #8606 is still open and PR #11659 is not merged.

Fixes

Challenge path broken (HTTP-01 or DNS-01)

Restore the challenge path and let the retry loop recover on its own; Traefik keeps retrying. For HTTP-01, fix the firewall or the process squatting on port 80. For DNS-01, fix the credentials using the exact env var names from the Lego provider docs, then restart Traefik if env vars changed. If propagation checks are flaky against your DNS, look at the DNS challenge propagation options for your Traefik version. In v3.3 the older delayBeforeCheck and disablePropagationCheck keys were deprecated in favor of propagation.delayBeforeChecks and propagation.disableChecks.

Rate limited by Let’s Encrypt

You cannot force your way through a rate limit. If the certificate has days left, wait out the window and stop whatever is burning quota: frequent restarts triggering re-issuance, staging or test domains on the production resolver, or stale certs in acme.json that Traefik keeps renewing. If the cert expires before the window resets, deploy a stop-gap certificate from another CA or a manually obtained cert as a provided certificate, then fix the quota consumption.

acme.json permissions or corruption

chmod 600 /path/to/acme.json

If the file is corrupted, back it up, remove it, and restart Traefik to trigger fresh issuance. This is disruptive: Traefik will re-issue every certificate, and fresh issuance for many domains at once can itself hit rate limits, so prune dead domains from your configuration first. Note the behavioral asymmetry: a corrupted acme.json is a hard failure at startup (visible), but corruption mid-operation fails silently while in-memory certs keep serving.

Stale store entries and the silent skip

If Traefik logs “No ACME certificate generation required” while the cert expires, remove the stale or overly broad entry (for example an old wildcard) from acme.json so the resolver no longer considers the domain covered. Back up the file before editing, then restart Traefik: the local ACME store is read once into memory, so in-process resolver passes do not reread manually edited acme.json.

HA lock contention

Traefik Proxy 2/3 has no built-in Consul or etcd ACME lock, so a stale built-in KV lock is not a cause here. Run one ACME-capable instance or use a product designed for ACME HA; if a custom or legacy v1 lock backend is involved, clear the stale lock in that backend.

Emergency stop-gap

If the cert expires within 24 hours and the cause is not yet fixed, deploy a manually obtained certificate as a provided certificate. This restores service immediately and decouples recovery from the ACME investigation.

Prevention

  • Alert on the renewal window, not the expiry. Page at <7 days only with a synthetic probe confirming the cert is actively served; ticket at <14 days; track at <30 days, because 30 days is when renewal should have succeeded. A serial that has not changed by day 75 of a 90-day cert is a failed renewal.
  • Log-based alerting on ACME errors. Since no metric exists for renewal failure, ship Traefik logs somewhere you can alert on “Error renewing certificate”, challenge errors, and rateLimited.
  • Synthetic TLS probes per production hostname. This closes the dormant-cert false-positive gap and catches cert selection problems that store metrics cannot see.
  • Protect the store. Keep acme.json at 600, back it up, and monitor it for parse failures. Prune certificates for decommissioned domains so Traefik stops renewing them and burning quota.
  • Track rate limit headroom. In dynamic environments where CI/CD creates subdomains, certificate issuance count is a capacity resource. Stay well under 50 per registered domain per week.
  • Test renewal after upgrades. The v2 to v3 migration changed rule syntax (for example Host() accepting a single domain), which broke ACME detection for some operators. After any Traefik upgrade, verify renewal still works before the old cert approaches its window.

How Netdata helps

  • Certificate countdown as a first-class chart. Netdata collects traefik_tls_certs_not_after and lets you alert on time-to-expiry, so the renewal window itself becomes the alerting threshold instead of browser complaints.
  • Per-second correlation with entrypoint traffic. When a cert expires, correlating the expiry timestamp with the HTTPS entrypoint request rate shows the exact moment client impact started and confirms recovery after the fix.
  • Log and metric correlation on one host. Netdata’s per-host view puts Traefik metrics next to system state (restarts, disk, FD usage), which matters for causes like a volume remount resetting acme.json permissions or frequent restarts driving rate limits.
  • Restart visibility. process_start_time_seconds alongside cert metrics reveals the “Traefik restarts, re-requests certs, hits rate limits” loop that quietly produces this failure.
  • Anomaly detection on the expiry series. A not_after value that keeps approaching now without ever jumping forward is the pattern ML-based anomaly detection flags, without hand-tuning a per-domain threshold.