The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-acme-lock-contention ▌

Operations Guides

Traefik ACME lock contention: stuck distributed locks blocking renewal

Every Traefik instance in your HA deployment is healthy. /ping returns 200, traffic flows, backends are up. Yet traefik_tls_certs_not_after keeps creeping toward the current time, and the logs on each instance repeat some variant of “unable to acquire lock”. Certificates are drifting toward expiry and nobody is renewing them.

This is the multi-instance ACME failure mode: Traefik uses a distributed lock in a KV store (Consul or etcd) so that exactly one instance performs ACME issuance and renewal at a time. When the lock holder crashes mid-renewal, is force-killed, or loses its session to a network partition, the lock can be left behind. Every remaining instance refuses to renew because the lock appears held, and the cluster slides toward certificate expiry while looking healthy on every conventional health signal.

This article covers detection, safe recovery, and prevention for deployments that coordinate ACME through a KV store. If you run Traefik v3 Community Edition, read the version note in “What this means” first, because the clustered ACME model changed.

What this means

In a single-instance deployment, Traefik’s ACME resolver stores certificates in acme.json and renews them roughly 30 days before expiry. In a multi-instance deployment, letting every replica renew independently would produce duplicate orders, conflicting writes to shared storage, and Let’s Encrypt rate limit consumption. To prevent that, Traefik elects a single ACME leader through a lock stored in the KV backend. Only the lock holder runs the renewal loop.

The failure is a split-brain by accident. The lock holder dies before releasing the lock. In the v1 KV design the datastore lock has a fixed 20-second TTL, so a holder that disappears should block acquisition only until that TTL expires; if lock state or leadership behaves otherwise, inspect the KV/session directly. Since certificate renewal is the only thing the lock gates, data-plane traffic is unaffected.

flowchart TD
  A[Instance A acquires ACME lock in KV store] --> B[A begins renewal]
  B --> C{A crashes or is partitioned mid-renewal}
  C --> D[Lock key remains in Consul/etcd]
  D --> E[Instances B and C attempt renewal]
  E --> F[Lock appears held - acquisition fails]
  F --> G[Log: unable to acquire lock]
  G --> H[traefik_tls_certs_not_after drifts toward now]
  H --> I[Clients start rejecting expired certs]

Version note. Clustered ACME with a KV-store lock is a Traefik v1-era design; official v3 documentation says the KV approach was dropped in 2.0. Traefik v3 Community Edition does not support shared or clustered ACME at all: each instance manages its own acme.json and its own ACME account, so the stale-lock failure mode does not apply, but neither does coordinated renewal. The officially supported multi-instance ACME path in v3-era deployments is Traefik Enterprise’s distributed ACME agent. If you are on v3 CE with multiple replicas, your equivalent risks are duplicate orders and rate limits, not lock contention. See Traefik ACME rate limit for that failure mode.

Common causes

CauseWhat it looks likeFirst thing to check
Lock holder crashed mid-renewalLock key persists in KV store; all instances log lock acquisition failures; one instance restarted recentlyprocess_start_time_seconds across replicas; inspect the lock key in Consul/etcd
Force kill or OOM during renewalSame as above, with an OOM event or SIGKILL in the instance’s historyContainer/pod restart reason, kernel OOM log on the node
Network partition between holder and KV storeHolder lost its KV session but kept running; another instance may have taken the lock, or the lock is orphanedKV connectivity logs on each instance; session/key TTL state in the KV store
Lock TTL longer than cert runwayLock eventually expires, but not before the renewal window closesCompare lock age against traefik_tls_certs_not_after
Uncoordinated rolling restartNew instance comes up while the old one’s lock is still valid; renewal deferred repeatedly during churnCorrelate lock errors with deployment/rollout timestamps

Quick checks

All read-only. Run these before touching the KV store.

# 1. Confirm the symptom: which certs are closest to expiry
curl -s http://localhost:8080/metrics | grep traefik_tls_certs_not_after

# 2. Confirm all instances are otherwise healthy (the trap: they will be)
curl -s http://localhost:8080/ping

# 3. Look for lock acquisition failures in the logs
#    (patterns seen in the field include "unable to acquire lock"
#    and "Existing key does not match lock use")
grep -iE 'lock|acme' /var/log/traefik/traefik.log | tail -50

# 4. Check for recent restarts that could have orphaned the lock
curl -s http://localhost:8080/metrics | grep process_start_time_seconds

# 5. Inspect the ACME area of the KV store (read-only)
#    The exact lock key path depends on your configured storage prefix.
#    For the v1 datastore, the lock key is exactly the configured ACME KV storage key plus "/lock"; for storage "/traefik/acme/account", inspect "/traefik/acme/account/lock".
consul kv get -recurse /traefik/acme/        # Consul example
etcdctl get --prefix /traefik/acme/          # etcd example

Note on check 5: the key path under your storage prefix varies with how you configured the KV provider and ACME storage. Browse the tree rather than assuming a fixed path. A lock key will typically stand out by name.

How to diagnose it

  1. Establish the timeline. Pull traefik_tls_certs_not_after for the affected CNs and SANs. Let’s Encrypt certificates have a 90-day lifetime and renewal is attempted about 30 days before expiry. If a cert is inside 30 days and has not renewed, renewal is blocked. If it is inside 7 days, renewal has been failing for roughly three weeks and you are in the escalation window.

  2. Confirm the lock story in logs. On each instance, search for ACME and lock-related errors. The distinguishing signature of lock contention, versus challenge failure or rate limiting, is that renewal never starts: instances complain about acquiring the lock, not about challenge responses or ACME server errors. If you instead see challenge errors, go to Traefik ACME challenge failed. If you see rate limit responses, go to Traefik ACME rate limit.

  3. Identify the lock holder. Read the lock key in the KV store. Lock implementations typically record the holder’s identity in the key value. Compare it against your running instances. If the recorded holder no longer exists (terminated pod, decommissioned node), the lock is stale.

  4. Rule out an active holder. Before declaring the lock stale, verify the holder is genuinely gone and not merely partitioned. If the holder is alive and mid-renewal, deleting the lock underneath it lets a second instance start issuing concurrently, which is exactly what the lock exists to prevent. Check whether the holder’s instance ID, pod name, or address appears among live replicas.

  5. Check for a post-deletion race history. If someone already deleted the lock once and instances then logged errors like “Datastore sync error: object lock value: expected X, got Y”, multiple instances raced to re-acquire and the KV state may be inconsistent. The safe recovery from that state is a coordinated restart of all Traefik instances after the lock is cleared, so exactly one instance wins the acquisition cleanly.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
traefik_tls_certs_not_after (per cn, sans, serial)The only metric that surfaces this failure; there is no ACME failure metricAny cert under 30 days not renewing; under 7 days is urgent
process_start_time_seconds per replicaA recent restart of one replica is the usual way a lock gets orphanedOne instance restarted while others log lock failures
Traefik logs: lock acquisition errorsThe direct evidence; metrics alone only show the consequenceRepeated “unable to acquire lock” across all instances
KV store session/lock key ageTells you whether the lock is live (session-backed) or orphanedLock key present with no corresponding live holder
traefik_config_last_reload_success per instanceRules out config drift as a contributing factor in HADivergent timestamps across replicas

Fixes

Clear the stale lock manually

This is the direct fix, and it is disruptive if done carelessly. Deleting a lock that is actually held by a live, mid-renewal instance permits concurrent issuance: duplicate orders, wasted rate limit budget, and potentially conflicting writes to shared certificate storage.

Procedure:

  1. Confirm via the diagnosis steps that the recorded lock holder is gone, not just quiet.
  2. Delete the lock key only, not the whole ACME tree. The account data and issued certificates in the KV store must be preserved.
# DESTRUCTIVE if the holder is still alive. Verify first.
# Replace the key path with the one you confirmed in quick check 5.
consul kv delete /traefik/acme/account/lock     # Consul example
etcdctl del /traefik/acme/account/lock          # etcd example
  1. Watch the logs. One instance should acquire the lock and begin renewal. If several instances log sync errors while racing, restart the Traefik replicas together so acquisition happens cleanly once.

There is no non-disruptive way to revoke a lock. The mitigation for the risk is verification discipline, not a safer command.

Shorten the lock TTL

Check the fixed datastore-lock TTL

In the v1 KV-backed datastore, the transaction lock TTL is fixed in source at 20 seconds and is not exposed as a Traefik configuration option. A crashed holder is bounded by that TTL; do not tune it in Traefik. If a lock remains after the TTL, inspect the KV store and session/lease implementation rather than deleting it blindly.

Tradeoff: a TTL shorter than a worst-case renewal (multiple domains, slow DNS-01 propagation) can expire the lock under a live holder and cause the concurrent-issuance problem from the other direction. Size the TTL to comfortably exceed your slowest observed renewal, not to minimize blocking time.

Designate a single ACME instance

Remove the coordination problem by removing the coordination. Run ACME on exactly one instance (or a small dedicated pair behind health-checked failover) and distribute issued certificates to the other replicas through your own mechanism, or terminate TLS only on the ACME-owning tier.

Tradeoff: you take on certificate distribution yourself, and the ACME instance becomes a renewal single point of failure that needs its own alerting. In exchange, lock contention, duplicate orders, and post-crash races all disappear.

Move off KV-coordinated ACME (v3 CE)

If you are on Traefik v3 Community Edition, clustered ACME is not available, so the fix is architectural: one instance owns ACME per domain set, or use DNS-01 with separate ACME accounts per instance so replicas never collide on orders, or adopt Traefik Enterprise’s distributed ACME agent if coordinated issuance is a hard requirement.

Prevention

  • Alert on renewal lag, not expiry. A 90-day certificate that has not renewed by day 70 has been failing for ten days. Alert when traefik_tls_certs_not_after crosses 30 days without the serial changing, instead of waiting for the 7-day cliff. Monitor the renewal process, not just the expiry timestamp.
  • Track restarts against lock errors. A replica restart followed by lock acquisition errors on peers is the canonical precursor. Correlating process_start_time_seconds across replicas with ACME log errors catches this in hours, not weeks.
  • Make shutdowns clean. OOM kills and force kills during renewal are how locks get orphaned. Give Traefik enough memory headroom and enough shutdown grace to finish or abort an in-flight renewal. On context cancellation the leadership candidate resigns; an in-flight datastore lock remains bounded by its fixed 20-second TTL.
  • Document the lock key location. During an incident is the wrong time to discover where your storage prefix puts the lock. Record the path and the read/delete commands in your runbook.
  • Prefer a single ACME owner at small scale. Below the replica count where distributed issuance genuinely pays for itself, one ACME-owning instance plus certificate distribution is operationally cheaper than a distributed lock.

How Netdata helps

  • Netdata charts traefik_tls_certs_not_after per certificate, so renewal lag is visible as a countdown long before clients start rejecting certificates.
  • Process uptime per replica makes the “one instance restarted, lock orphaned” correlation visible on a single dashboard instead of across three log tails.
  • Per-second metric collection catches the restart-and-recover pattern that hourly scrapes miss, which matters when the lock holder flaps.
  • Because Traefik exposes no ACME failure metric, Netdata’s value here is pairing the expiry gauge with instance restarts and traffic health so you can confirm renewal is the only thing broken.
  • Alerts on certificates crossing the 30-day renewal window turn a silent three-week drift into a ticket while there is still runway.