The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-tls-certs-not-after ▌

Operations Guides

Traefik TLS certificate expiry monitoring: reading traefik_tls_certs_not_after

traefik_tls_certs_not_after is the only signal Traefik gives you about certificate expiry. There is no metric for ACME renewal failures, no counter for rate limit rejections, no gauge for “the challenge did not complete.” You get a Unix timestamp per certificate, and everything else has to be inferred from it or pulled from logs.

Most teams wire it to a pager with a threshold like “alert at 7 days.” Then the first ACME renewal happens, the old certificate series stays in the output, and the pager fires for a certificate that is no longer serving anything. Or the alert fires for a staging certificate, or for a dormant cert left in the store from a decommissioned hostname, and the on-call learns to ignore it. This guide covers reading the metric correctly and building an alert chain that only wakes someone up when a certificate that is actually serving production traffic is about to expire.

What the metric actually reports

traefik_tls_certs_not_after is a gauge. Its value is the Unix epoch timestamp of the certificate’s notAfter field, and it carries three labels:

LabelContent
cnThe certificate’s Common Name
sansSubject Alternative Names
serialThe certificate serial number

A typical exposition looks like:

traefik_tls_certs_not_after{cn="example.com",sans="example.com,www.example.com",serial="03ab..."} 1.7935e+09

To get days remaining, subtract time() and divide by 86400. The important property is what the metric counts: every certificate in Traefik’s certificate store. That includes certificates currently terminating traffic, certificates that were renewed and replaced, certificates for routers that no longer exist, and certificates issued against the Let’s Encrypt staging CA. The store is a cache, not a live view of what is on the wire.

Two prerequisites before you see anything at all:

  • Prometheus metrics must be enabled (--metrics.prometheus=true, typically scraped from a dedicated metrics entrypoint).
  • The metric only appears once Traefik has at least one TLS-terminating entrypoint and certificates loaded. If your entrypoints are plain HTTP, the series is absent entirely. Absence of the metric is not “no certs expiring soon”; it is “nothing to report” or “not configured,” and those are very different states.

The three traps

1. The store outlives the certificate’s usefulness

A cert for a hostname you migrated off Traefik six weeks ago still emits a series, still ticks toward expiry, and still crosses your alert threshold. There is no label that says “this cert is currently serving traffic,” and no metric that maps a certificate to an active router. This is the fundamental reason paging on any single cert’s expiry false-fires.

2. Stale serials after renewal

There is a long-standing open bug (traefik/traefik#8606): when ACME renews a certificate, the old certificate’s series is not removed from the Prometheus output. You end up with multiple series for the same cn/sans pair, differing only by serial. The old series keeps counting down toward its original expiry date. With a 90-day Let’s Encrypt certificate renewed at 30 days remaining, a 7-day alert fires about 23 days after renewal (a 30-day plan threshold can fire immediately), and the stale series remains until the process restarts. Users have confirmed this persists in recent v3 releases. A proposed fix (setting stale series to 0) stalled because maintainers considered the sentinel value approach unsatisfactory, so plan on working around it in your queries, not waiting for it to disappear.

3. Staging certificates look valid to the metric

If your caServer points at the Let’s Encrypt staging endpoint (acme-staging-v02.api.letsencrypt.org), Traefik issues certificates signed by “Fake LE Intermediate X1” and reports them in traefik_tls_certs_not_after like any other valid cert. Traefik does not distinguish staging from production ACME in its metrics. Browsers reject them. So the metric can show a healthy 80 days remaining while every client that matters is throwing TLS errors. This comes up constantly when a staging config leaks into production, or when someone flips the CA server to debug rate limits and forgets to flip it back.

Alert design that does not false-fire

A three-tier design maps cleanly onto the failure modes above. The tiers differ in who gets woken up and why:

  • PLAN at 30 days. Let’s Encrypt certificates live 90 days and Traefik attempts renewal 30 days before expiry. A cert inside 30 days with no renewal yet means the renewal window just opened. No alert fatigue risk here; this is a ticket or a dashboard annotation telling you to verify the automation is alive.
  • TICKET at 7 to 14 days. A cert 14 days from expiry means renewal has been failing for roughly 16 days. At 7 days, it has been failing for about 53. This is a genuine signal that something in the ACME pipeline (challenge reachability, DNS provider credentials, rate limits, acme.json permissions) is broken, but you still cannot tell from the metric alone whether the expiring cert serves real traffic.
  • PAGE only with an external synthetic probe. The metric cannot tell you which certificate is on the wire for your production hostname. Pair the internal signal with an external TLS probe (blackbox-style) that connects to the real production hostname, completes a handshake, and checks the served chain’s validity and expiry. When the probe confirms the actively-serving certificate for a known production name is about to expire, that is page-worthy. Until then, it is a ticket.
flowchart TD
  A["traefik_tls_certs_not_after series"] --> B{Days remaining?}
  B -->|"> 30d"| C["No action"]
  B -->|"14-30d"| D["PLAN: verify renewal automation"]
  B -->|"7-14d"| E["TICKET: investigate ACME pipeline"]
  B -->|"< 7d"| F{"External probe confirms production hostname serving this cert?"}
  F -->|"Yes"| G["PAGE"]
  F -->|"No / dormant cert"| H["TICKET: clean up store, fix renewal"]

Collapsing the stale serials

Because of the stale-serial bug, group away the serial label before thresholding. The community-standard pattern from Awesome Prometheus Alerts uses min by (instance, sans):

# Critical: any cert under 7 days, collapsing duplicate serials per SAN set
min by (instance, sans) (
  last_over_time(traefik_tls_certs_not_after[5m]) - time()
) / 86400 < 7
# Warning: same expression with a 14-day threshold
min by (instance, sans) (
  last_over_time(traefik_tls_certs_not_after[5m]) - time()
) / 86400 < 14

A caveat worth knowing: min by (sans) takes the earliest expiry among duplicate serials, which means a stale series from a just-renewed cert can still drag the group below threshold for a while. An alternative is max by (instance, sans) (...) on the expiry expression, which selects the largest expiry timestamp—the freshest renewed certificate—in each group. It masks stale serials in the query; it does not delete them. A Traefik restart clears its process-local stale gauge state, but the series can reappear after subsequent renewals if upstream does not change collector cleanup.

Neither workaround fixes staging certs. If you run a shared Traefik for staging and production, split staging onto its own instance so the stores never mix. The certificate metric exposes only cn, sans, and serial; it has no issuer or CA label, so there is no reliable label pattern for filtering “Fake LE Intermediate” certificates.

When the metric lies about the filesystem

Two adjacent failure modes matter for operators who provide certs by file rather than ACME:

  • On some network filesystems (the reported cases involve GlusterFS and NFS), Traefik does not detect certificate file changes and keeps serving the old cert. The metric reflects the old cert’s expiry even though the new file is on disk. A config reload or restart picks it up. If your certs live on a network mount, verify a rotation actually changes the metric value before relying on it.
  • acme.json must be readable and writable with mode 600. If permissions change (volume remount, restore from backup), Traefik may silently stop persisting renewals. The metric keeps counting down with no error surfaced anywhere except logs. Traefik hard-fails at startup on a corrupted acme.json, but mid-operation corruption fails silently while in-memory certs keep serving.

Version notes

  • The Prometheus metric name traefik_tls_certs_not_after is stable across v2 and v3. It first appeared in v2.5.
  • In v3.5.4, the OpenTelemetry variant was renamed from traefik_tls_certs_not_after_milliseconds to traefik_tls_certs_not_after_seconds to match its actual unit. The migration note documents a rename, not a dual-emission grace period; verify the exact name emitted by your exporter after upgrading.
  • In HA deployments, each instance reports its own store. If ACME storage is not shared (or the distributed lock in Consul/etcd is stuck because an instance died holding it), instances can show divergent expiry values for the same hostname. Compare the metric per instance, and watch logs for lock acquisition failures.

Signals to correlate

Certificate expiry never exists in isolation. These signals tell you whether “cert expiring” is “renewal automation hiccup” or “production outage in 6 days”:

SignalWhy it mattersWarning sign
traefik_tls_certs_not_afterThe countdown itself; the only cert metric that existsAny series < 14 days, or a production-name series < 7 days
Traefik logs (ACME errors)The only place renewal failure causes appear; there is no failure metric“Error renewing certificate”, challenge failures, rate limit responses
External synthetic TLS probeConfirms which cert is actually served on the production hostnameProbe sees the near-expiry or staging-signed chain
traefik_config_last_reload_successA frozen config can mean cert updates are not being applied eitherTimestamp not advancing while renewals should be occurring
traefik_entrypoint_requests_tls_totalTLS version/cipher distribution; drops can indicate clients rejecting handshakesSudden change in handshake mix near expiry
Per-instance cert values (HA)Detects store divergence and stuck ACME locksSame hostname showing different not_after across replicas

For Let’s Encrypt specifically, keep the rate limits in mind when diagnosing: 50 certificates per registered domain per week, 5 duplicate certificates per week. Hitting the duplicate limit is a common cause of “renewal silently stopped” in crash-looping or HA-without-shared-storage deployments, and the metric alone will not tell you that is what happened.

How Netdata helps

  • Netdata charts traefik_tls_certs_not_after per certificate with its cn/sans/serial labels, so you can see the full store at a glance and spot dormant or duplicate-serial series before they page you.
  • Per-second collection catches the moment a renewal lands: a new series appears with a fresh expiry while the stale one keeps counting down, which is exactly the pattern the alert queries need to handle.
  • Correlating the cert countdown with traefik_config_last_reload_success on the same dashboard separates “renewal failing” from “config frozen, nothing is being applied,” two causes with very different fixes.
  • Comparing the metric across HA instances side by side surfaces store divergence and stuck ACME locks without per-instance manual checks.
  • Pairing Netdata’s internal view with an external synthetic probe of the production hostname closes the gap the metric cannot: knowing whether the expiring cert is the one on the wire.