The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-acme-challenge-failure ▌

Operations Guides

Traefik ACME challenge failed: HTTP-01, DNS-01, and TLS-ALPN-01 renewal errors

Your certificates are approaching expiry and Traefik’s logs contain lines like “Unable to obtain ACME certificate for domains” or “Error renewing ACME certificate”. HTTPS still works for now because Traefik keeps serving the old certificate, but the clock is running: Let’s Encrypt certificates live 90 days, Traefik attempts renewal 30 days before expiry, and if you are seeing renewal errors today, they have likely been failing silently for weeks.

The uncomfortable part: Traefik exposes no Prometheus metric for ACME renewal failures. The only metric-adjacent signal is traefik_tls_certs_not_after counting down toward expiry. The actual failure details exist only in the logs. That means this failure mode is invisible until someone either reads the logs, alerts on certificate expiry, or gets a browser warning.

The fix depends entirely on which challenge type you are using. HTTP-01, TLS-ALPN-01, and DNS-01 fail for completely different reasons, and the log error for one looks nothing like the log error for another.

What this means

Traefik’s ACME resolver is a background process that asks Let’s Encrypt (or another CA) to prove domain control before issuing a certificate. The CA picks the proof mechanism based on what you configured:

  • HTTP-01 (httpChallenge): Let’s Encrypt makes an HTTP request to http://<domain>/.well-known/acme-challenge/<token>. Traefik must be reachable on port 80 and must answer that request itself.
  • TLS-ALPN-01 (tlsChallenge): Let’s Encrypt opens a TLS connection on port 443 with the ALPN protocol acme-tls/1. Traefik must terminate that handshake itself and present a temporary challenge certificate.
  • DNS-01 (dnsChallenge): Traefik uses your DNS provider’s API to publish a TXT record at _acme-challenge.<domain>. Let’s Encrypt queries DNS for it. This is the only challenge type that can issue wildcard certificates.

Each mechanism has exactly one hard dependency: port 80 reachability, port 443 TLS reachability, or a working DNS provider API plus propagation. Break that dependency, and renewal fails with no signal other than a log line and a shrinking expiry window.

flowchart TD
  A[Renewal due: 30 days before expiry] --> B{Configured challenge}
  B -->|httpChallenge| C[LE requests token over port 80]
  B -->|tlsChallenge| D[LE handshake on port 443, ALPN acme-tls/1]
  B -->|dnsChallenge| E[Traefik creates TXT record via DNS API]
  C --> F{Port 80 reaches Traefik?}
  D --> G{TLS handshake reaches Traefik?}
  E --> H{TXT record visible in public DNS?}
  F -->|No| I[Challenge error in logs]
  G -->|No| I
  H -->|No| I
  F -->|Yes| J[Certificate stored in acme.json]
  G -->|Yes| J
  H -->|Yes| J
  I --> K[traefik_tls_certs_not_after keeps counting down]

Common causes

CauseWhat it looks likeFirst thing to check
Firewall or security group blocks port 80HTTP-01 challenge timeouts in logs; everything else healthycurl -v http://<domain>/.well-known/acme-challenge/test from an external host
Port 443 terminated by something that is not TraefikTLS-ALPN-01 fails; a TLS-terminating load balancer in front of Traefik cannot forward the ALPN challengeWhich component actually terminates TLS on 443
Rotated or expired DNS provider API tokenDNS-01 fails with provider-specific API errors; HTTP-01 domains keep renewing fineTest the token against the provider API directly
DNS propagation slower than the check timeoutTXT record exists but Let’s Encrypt’s resolvers do not see it yetdig TXT _acme-challenge.<domain> against public resolvers
Let’s Encrypt rate limits exhaustedRenewal was working, then all orders rejected; often after a crash-looping Traefik re-requested certs repeatedlyCount recent issuance for your registered domain
Corrupted or unwritable acme.jsonRenewal errors about storage, or Traefik hard-fails at startup after restartls -l on acme.json (must be 600); validate JSON
caServer pointed at stagingCertificates “renew” but are issued by a fake staging CA that browsers rejectCheck the caServer value in your certificate resolver config
Stale distributed ACME lock (HA)All instances healthy but no renewal happens; logs mention lock acquisitionCheck the lock key in Consul/etcd

Quick checks

All of these are read-only and safe to run during an incident.

# 1. What is actually being served, and when does it expire?
echo | openssl s_client -servername <domain> -connect <traefik-host>:443 2>/dev/null \
  | openssl x509 -noout -dates -issuer

# 2. Check the expiry metric Traefik exposes
curl -s http://localhost:8080/metrics | grep traefik_tls_certs_not_after

# 3. Find the ACME errors in the logs (adjust for your logging setup)
journalctl -u traefik --since "24 hours ago" | grep -iE 'acme|renew'
# or: docker logs traefik --since 24h 2>&1 | grep -iE 'acme|renew'

# 4. HTTP-01: is port 80 reachable at Traefik from the outside?
# Run this from a host OUTSIDE your network:
curl -v http://<domain>/.well-known/acme-challenge/test
# You want a Traefik response (usually 404 or 503 from Traefik itself),
# not a connection timeout and not an answer from a CDN or other proxy.

# 5. TLS-ALPN-01: does the TLS handshake on 443 land on Traefik?
echo | openssl s_client -connect <domain>:443 -servername <domain> 2>/dev/null \
  | openssl x509 -noout -subject -issuer
# The presented cert should be one Traefik serves, not your LB's or CDN's.

# 6. DNS-01: can public resolvers see the challenge record?
dig TXT _acme-challenge.<domain> @1.1.1.1 +short
dig TXT _acme-challenge.<domain> @8.8.8.8 +short

# 7. Check acme.json permissions and validity
ls -l /path/to/acme.json        # must be 600
python3 -m json.tool /path/to/acme.json > /dev/null && echo "valid JSON"

How to diagnose it

  1. Confirm the symptom is renewal, not serving. Check traefik_tls_certs_not_after for the affected CN/SANs. If the cert is 30-60 days from expiry, renewal is failing but you have runway. If it is under 14 days, renewal has been broken for weeks and this is now urgent.

  2. Identify the configured challenge type. Look at your certificate resolver configuration: httpChallenge.entryPoint, tlsChallenge, or dnsChallenge.provider. The challenge type determines the entire diagnostic path. If you are requesting a wildcard (*.example.com), you must be on DNS-01; the other two cannot issue wildcards.

  3. Read the actual error. Grep the logs for acme around the renewal attempts. The error string tells you which of the three dependencies broke: a timeout or connection error points at HTTP-01/TLS-ALPN-01 reachability; a provider API error points at DNS-01 credentials; a propagation or TXT mismatch error points at DNS.

  4. Test the dependency directly. For HTTP-01, replay an external request to the challenge path (check 4 above) and confirm the response comes from Traefik, not a CDN, WAF, or load balancer in front. For TLS-ALPN-01, confirm the handshake on 443 terminates at Traefik; a TLS-terminating load balancer in front of Traefik makes this challenge type unworkable because the ALPN challenge never reaches it. For DNS-01, create a test TXT record via the same API credentials Traefik uses, then verify it resolves publicly. If the API call fails, it is credentials. If the API call succeeds but resolvers do not see it, it is propagation.

  5. Rule out rate limits. Let’s Encrypt enforces 50 certificates per registered domain per week, 5 duplicate certificates per week, and 300 new orders per account per 3 hours. A Traefik instance crash-looping with a broken resolver config can burn through the duplicate limit in an afternoon. If every other check passes, this is the likely cause, and the log error will say so.

  6. Check the storage and the CA endpoint. Verify acme.json is valid JSON with 600 permissions, and confirm caServer is not pointing at the Let’s Encrypt staging directory. Staging renewals “succeed” from Traefik’s perspective but produce certificates no browser trusts.

  7. In HA deployments, check the lock. Traefik Proxy stores ACME state in a local file (not distributed), so instances do not share acme.json through Consul or etcd in v2/v3. If your deployment still uses a v1-era KV-backed lock, clear the orphaned lock only after confirming the actual storage design.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_tls_certs_not_afterThe only metric that reflects ACME health, and only as a consequenceAny cert under 30 days means the renewal window opened and nothing happened
Traefik logs: “Unable to obtain ACME certificate for domains” / “Error renewing ACME certificate”The only place the actual failure cause appearsAny occurrence; one failure is a warning, repeated failures over days are the pattern that kills you
Certificate serial/CN changes over timeA renewing cert gets a new serial; a stuck one does notSame serial for 60+ days on a 90-day cert
caServer configurationStaging vs production determines whether issued certs are trustedAny resolver pointing at the staging directory in production
acme.json file stateCorruption causes silent failure mid-operation and hard failure at startupPermission drift away from 600, invalid JSON

Fixes

HTTP-01: restore port 80 reachability

Open the firewall or security group rule for port 80 to the Traefik instance and make sure the traffic lands on Traefik’s HTTP entrypoint, not on a CDN, WAF, or another proxy. If something sits in front of Traefik, the challenge request must be forwarded to Traefik unmodified. Let’s Encrypt does follow redirects, so an HTTP-to-HTTPS redirect on the entrypoint does not by itself break HTTP-01, but the redirect target must still be Traefik. After fixing, renewal retries on Traefik’s own schedule; you do not need to restart.

TLS-ALPN-01: give Traefik the handshake

The TLS connection on port 443 must be terminated by Traefik itself, because only Traefik can present the temporary acme-tls/1 challenge certificate. If a load balancer terminates TLS in front of Traefik, either reconfigure it for TLS passthrough or switch to HTTP-01 or DNS-01. In most cloud load balancer setups, switching challenge type is the pragmatic choice.

DNS-01: fix credentials or propagation

  • Rotated token: update the API credentials in Traefik’s environment or config and reload. Verify the new token works by creating a test TXT record through the provider API before waiting for Traefik’s next attempt.
  • Propagation delay: Traefik verifies the TXT record before telling Let’s Encrypt to check it. If your DNS provider is slow to propagate, increase the verification delay (dnsChallenge.propagation.delayBeforeChecks in current Traefik v3; the legacy dnsChallenge.delayBeforeCheck option remains available but is deprecated, so confirm your installed version before using it).
  • Wildcard plus apex: remember DNS-01 is mandatory for wildcards. If only your wildcard domains fail while plain domains renew, that asymmetry confirms a DNS-01-specific problem.

Rate limits: stop the bleeding, then wait

Fix the underlying cause first so Traefik stops generating failing orders. If you are testing configuration, point caServer at the Let’s Encrypt staging endpoint, which has much higher limits. If you have already hit the production weekly limit, there is no override: you wait for the window to slide. This is why rate-limit failures must never be the first time you notice renewal is broken.

Corrupted acme.json or expired-certificate deadlock

Back up acme.json, remove it, and restart Traefik. This forces fresh issuance for all configured domains. Two cautions: restarting Traefik drops in-flight connections, so do it deliberately; and fresh issuance for many domains at once counts against rate limits, so do not do this repeatedly. Note also the asymmetry from the playbook: mid-operation corruption fails silently (in-memory certs keep being served), while corrupted storage at startup makes Traefik hard-fail. If Traefik refuses to start after a restart, check acme.json first.

Stale HA lock

For a legacy KV-backed design only, clear the orphaned lock key manually and consider shorter lock TTLs. For current Traefik Proxy, diagnose each instance’s own acme.json and resolver configuration instead.

Prevention

  • Alert on expiry with real runway. Ticket at under 30 days (the renewal window opened and failed) and page at under 7 days. Waiting for 7 days as your only alert means renewal has been broken for almost two months before anyone looks.
  • Alert on the log lines. Since there is no ACME failure metric, scrape or alert on “Unable to obtain ACME certificate for domains” and “Error renewing ACME certificate” in Traefik logs. Two consecutive days of renewal errors should page before the expiry metric moves at all.
  • Do all resolver testing against staging. Every config change, new domain pattern, or DNS provider migration should be validated against the staging caServer first. Production rate limits do not forgive experimentation.
  • Protect acme.json. Enforce 600 permissions, include it in backups, and monitor it for unexpected modification. It is a single point of failure for every certificate you serve.
  • Track DNS provider credential rotation. Put DNS API token expiry/rotation on the same calendar as certificate management. A rotated token breaks DNS-01 silently, and wildcards have no fallback.
  • Keep certificate inventory visible. Track CN, SANs, and serial numbers so “the cert never changed” is an observable fact, not something you reconstruct during an incident.

How Netdata helps

Netdata surfaces the exact signals this failure mode hides behind:

  • traefik_tls_certs_not_after per CN/SAN, charted as time-to-expiry, so a certificate whose renewal has silently failed shows up as a shrinking countdown weeks before browsers complain.
  • Anomaly detection on the expiry curve flags certificates that are not being replaced on the expected 30-day-before-expiry cadence, which is the earliest metric-based hint of ACME trouble.
  • Log-based correlation puts the ACME error lines from Traefik’s logs next to the expiry metric and any restart events, so you can see “renewal started failing right after this deploy” instead of discovering it from a browser warning.
  • Process restart tracking (process_start_time_seconds) catches crash-looping Traefik instances before they burn through Let’s Encrypt’s duplicate-certificate rate limit.
  • Config reload freshness (traefik_config_last_reload_success) helps rule in or out a provider-side config problem when the resolver configuration itself may be stale.

The correlation that matters: expiry countdown dropping, ACME errors in logs, and no certificate serial change. Any two of those three means act now.