The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix / postfix-tls-certificate-expiry ▌

Operations Guides

Postfix TLS certificate expiry: expired certs and handshake failures

An expired TLS certificate on a Postfix server breaks mail delivery in two distinct ways, and teams frequently detect only one. The inbound symptom is loud: clients that require STARTTLS fail handshakes, connections drop, and complaints arrive quickly. The outbound symptom is quiet: mail to destinations enforcing mutual TLS or DANE silently defers into the deferred queue, where it sits under Postfix’s increasing backoff schedule until someone notices the queue growing.

Postfix has no built-in certificate expiry monitoring. There is no metric, no warning at startup, and no periodic check. The daemon loads certificate files into memory at process spawn time and serves them until reloaded. If certbot renews a certificate but the deploy hook is missing or broken, Postfix continues serving the old cert from memory. The file on disk is valid. The cert in flight is not.

The diagnostic discipline must cover both sides: the smtpd certificate presented to inbound clients, and the smtp client certificate used for outbound connections to destinations that require mutual TLS or DANE. Certificate expiry is not a Postfix metric. It requires an external openssl x509 check against the actual cert files, and the check should also verify what Postfix serves over the wire, not just what is on disk.

What this means

flowchart TD
    A["TLS failures or silent deferrals"] --> B{Inbound or outbound?}
    B -->|Inbound smtpd| C["Check smtpd cert expiry"]
    B -->|Outbound smtp| D["Check smtp client cert expiry"]
    C --> E{Expired?}
    E -->|Yes| F["Renew and reload Postfix"]
    E -->|No| G["Check cipher or protocol mismatch"]
    D --> H{DANE enabled?}
    H -->|Yes| I["TLSA hash matches new key?}
    H -->|No| E
    I -->|No| J["Update TLSA or reuse key"]
    I -->|Yes| E

An expired smtpd certificate breaks inbound TLS for every client that requires it. Clients with mandatory-TLS policies (smtp_tls_security_level = encrypt) refuse to deliver. Clients with opportunistic TLS (may) silently fall back to plaintext, which may violate security policy without triggering any alert.

An expired smtp client certificate breaks outbound mutual TLS. Destinations that require authenticated TLS reject the connection, typically with a 4xx deferral. Postfix retries on its backoff schedule, so the deferred queue grows steadily. Because opportunistic outbound TLS falls back silently, the failure may only become visible when a destination enforces it.

DANE-EE (TLSA usage 3) does not check certificate expiration. An expired cert can still pass DANE-EE validation if the TLSA hash matches. But the same expired cert will fail standard PKIX validation at non-DANE destinations. This asymmetry means an expired cert can cause selective failures: DANE-validating destinations continue working while PKI-validating destinations reject.

Common causes

CauseWhat it looks likeFirst thing to check
Inbound smtpd cert expiredClients fail STARTTLS or fall back to plaintextopenssl x509 -enddate on the smtpd cert file
Outbound client cert expiredMail silently defers to mutual-TLS destinationsopenssl x509 -enddate on the smtp client cert file
Postfix not reloaded after renewalCert on disk is valid but Postfix serves old cert from memoryCompare served cert via openssl s_client vs file on disk
DANE TLSA mismatch after renewalOutbound DANE mail defers after certbot runCompare TLSA record hash with renewed cert public key
Expired intermediate CA in chainRemote servers reject cert with “certificate expired” despite valid leafCheck chain for expired intermediates, verify --preferred-chain

Quick checks

All commands below are safe and read-only.

# Show configured cert paths (inbound and outbound)
postconf -h smtpd_tls_cert_file smtpd_tls_key_file
postconf -h smtp_tls_cert_file smtp_tls_key_file
# Postfix >= 3.4: chain files combine key and cert in one file
postconf -h smtpd_tls_chain_files smtp_tls_chain_files

# Check expiry of configured inbound smtpd cert
openssl x509 -in "$(postconf -h smtpd_tls_cert_file)" -noout -enddate -subject -issuer 2>&1

# Check expiry of configured outbound smtp client cert
openssl x509 -in "$(postconf -h smtp_tls_cert_file)" -noout -enddate -subject -issuer 2>&1

# Check if cert expires within 30 days (exit 0 = still valid, exit 1 = expires soon)
openssl x509 -in "$(postconf -h smtpd_tls_cert_file)" -checkend 2592000 -noout 2>&1; echo "exit: $?"

# Verify what Postfix is actually serving on port 25 (inbound cert in memory)
echo QUIT | openssl s_client -connect localhost:25 -starttls smtp 2>/dev/null | openssl x509 -noout -enddate

# Scan for TLS-related failures from the previous hour (GNU date)
grep "$(date -d '1 hour ago' '+%b %e %H')" /var/log/mail.log | grep -iE 'TLS.*fail|SSL.*error|certificate' | tail -20

# Check deferred queue for TLS-related deferral reasons
grep 'status=deferred' /var/log/mail.log | grep -iE 'TLS|certificate|handshake|SSL' | tail -20

# Verify certbot deploy hook exists for Postfix
ls -la /etc/letsencrypt/renewal-hooks/deploy/ 2>/dev/null

How to diagnose it

  1. Identify which cert paths are configured. Run postconf -h smtpd_tls_cert_file for inbound and postconf -h smtp_tls_cert_file for outbound. If you are running Postfix >= 3.4, also check smtpd_tls_chain_files and smtp_tls_chain_files, which combine key and chain in a single file and are the preferred interface.

  2. Check expiry of each cert file. Run openssl x509 -in <path> -noout -enddate against each configured cert. Pay attention to both the end-entity cert and any intermediate CA certs in the chain. An expired intermediate causes the same “certificate expired” rejection from remote validators even when the leaf is valid.

  3. Compare what Postfix serves versus what is on disk. After any renewal, there is a window where the file has been updated but Postfix has not been reloaded. Connect with openssl s_client -connect localhost:25 -starttls smtp and pipe the output through openssl x509 -noout -enddate. If the served end date differs from the file, Postfix is holding a stale cert in memory.

  4. Scan logs for TLS failure patterns. Default Postfix logging (level 0) hides failed opportunistic TLS negotiations. Temporarily elevate with postconf -e 'smtp_tls_loglevel = 1' and postconf -e 'smtpd_tls_loglevel = 1' followed by postfix reload to capture negotiation details. This modifies main.cf and triggers a reload; revert to level 0 after debugging. Look for SSL_accept error, SSL_connect error, certificate expired, and TLS handshake failed.

  5. Check DANE/TLSA state if outbound mail to DANE-adopting destinations is deferring. After certbot renewal, the new certificate typically has a new public key unless --reuse-key was specified. The published TLSA record, which hashes the public key, no longer matches. Query the TLSA record and compare against the new cert.

  6. Verify the reload mechanism. Check /etc/letsencrypt/renewal-hooks/deploy/ for a script that runs postfix reload or systemctl reload postfix. If the hook is missing, certbot renews the cert on disk but Postfix never picks it up. Note that certbot renew --dry-run does not execute deploy hooks (certbot issue #9916), so test the hook by running it manually rather than trusting a dry run.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Certificate expiry date (external check)Not a Postfix metric; must be checked externally with opensslLess than 30 days: alert. Less than 7 days: page.
TLS negotiation failure rateIndicates cert problems, cipher mismatch, or interceptionMandatory TLS: any failure. Opportunistic: more than 5% failure rate.
Deferred queue growth rateOutbound TLS failures cause silent 4xx deferralsSustained growth correlated with TLS errors in logs
Delivery rate vs injection rateTLS deferrals reduce delivery while injection continuesDelivery rate diverging downward from injection rate
DANE/TLSA validation failuresTLSA record mismatch after renewal breaks outbound DANE mailMail to DANE destinations deferring with TLS errors
Opportunistic TLS fallback rateSilent downgrade to plaintext is a security concernIncrease in plaintext connections where TLS was previously negotiated

Fixes

Expired inbound smtpd certificate

Renew the certificate through your standard process (certbot or equivalent). Then reload Postfix so it reads the new cert into memory:

# Reload Postfix to pick up the renewed cert
postfix reload

Verify the served cert matches the renewed file using openssl s_client as described in the diagnostic steps. If Postfix does not pick up the new cert after reload, a full restart may be needed.

Postfix serving stale cert after renewal

The cert file on disk is valid but Postfix is serving the old one from memory. This means the deploy hook is missing or failed. Create a deploy hook so Postfix reloads automatically after every renewal:

# Create the deploy hook
cat > /etc/letsencrypt/renewal-hooks/deploy/reload-postfix.sh << 'EOF'
#!/bin/sh
/usr/sbin/postfix reload
EOF
chmod +x /etc/letsencrypt/renewal-hooks/deploy/reload-postfix.sh

certbot renew --dry-run does not trigger deploy hooks (certbot issue #9916), so verify the hook by running it manually after a real renewal.

DANE TLSA mismatch after renewal

When certbot generates a new key pair on renewal, the published TLSA record no longer matches. Three options, in order of operational simplicity:

  1. Use --reuse-key in certbot. The public key stays the same across renewals, so the TLSA hash remains valid. Add --reuse-key to your certbot renewal configuration. This is the simplest fix but means the same key persists across renewal cycles.

  2. Switch to DANE-TA (TLSA usage 2) with the CA certificate. Instead of hashing the end-entity cert, hash the issuing CA. The TLSA record survives leaf cert rotation as long as the CA remains the same. This requires republishing the TLSA record once.

  3. Update the TLSA record on every renewal. Add a deploy hook that recomputes the TLSA hash from the new cert and updates DNS via your provider API. Publish the new TLSA record before the cert rotation and account for DNS TTL propagation to avoid a validation gap.

Expired intermediate CA in chain

If remote servers reject your cert with “certificate expired” but your leaf cert is valid, the intermediate CA in your chain may be expired. This occurred with the Let’s Encrypt DST Root CA X3 cross-signature expiration in September 2021: remote servers running outdated OpenSSL rejected certs chained to the expired cross-sign, even though the leaf was valid. Request certs with the current preferred chain:

# Request with ISRG Root X1 chain
certbot certonly --preferred-chain "ISRG Root X1" -d mail.example.com

Verify the chain served by Postfix includes only valid intermediates by inspecting the full chain with openssl s_client -showcerts.

Prevention

  • Alert on certificate expiry externally. Postfix provides no internal metric. Run openssl x509 -checkend 2592000 (30 days) in a cron job or monitoring check against both smtpd and smtp cert files. Alert at 30 days out, page inside 7 days.
  • Monitor both sides. Most teams watch only the smtpd (inbound) certificate. The smtp (outbound) client certificate is equally critical for mutual-TLS and DANE destinations. Outbound failures are silent deferrals, not obvious errors.
  • Verify deploy hooks after any certbot configuration change. A certbot reconfigure or version update can silently drop or break the deploy hook. Verify the hook after a real renewal - certbot renew --dry-run does not execute deploy hooks.
  • Elevate TLS log levels. Default Postfix logging (level 0) hides failed opportunistic TLS negotiations, creating blind spots for downgrade detection and interoperability issues. Set smtp_tls_loglevel = 1 and smtpd_tls_loglevel = 1 to capture negotiation details. Use level 2 only for active debugging, as it adds significant verbosity.
  • Use smtpd_tls_chain_files and smtp_tls_chain_files (Postfix >= 3.4). These combine key and chain in a single file, avoiding the race condition present when key and cert are specified in separate files during rollover.
  • Prefer --reuse-key with certbot for DANE deployments. This keeps the TLSA record valid across renewals without requiring DNS updates on every cycle.
  • If you use deprecated TLS parameters, migrate. Postfix 3.9 deprecated smtpd_use_tls, smtpd_enforce_tls, and smtp_tls_per_site (postconf logs a warning). Replace smtpd_use_tls/smtpd_enforce_tls with smtpd_tls_security_level (may or encrypt), smtp_use_tls/smtp_enforce_tls with smtp_tls_security_level (may, encrypt, or dane), and smtp_tls_per_site with smtp_tls_policy_maps to avoid postconf warnings and future removal.

How Netdata helps

Netdata does not replace the external openssl x509 expiry check, because Postfix does not expose certificate expiry as a metric. What Netdata provides is early detection of the downstream effects of an expired cert:

  • Deferred queue growth rate. An expired outbound cert causes silent deferrals to TLS-enforcing destinations. Per-second queue metrics catch the growth before the queue reaches crisis depth.
  • Mail flow velocity (injected vs delivered). When delivery rate drops while injection continues, the divergence appears within minutes. Correlating this with TLS error patterns in logs pinpoints the cause.
  • TLS negotiation failure rate. Elevated TLS failures in mail logs surface as a rate anomaly. Mandatory TLS failures should be zero; any nonzero rate warrants investigation.
  • Bounce rate monitoring. Some destinations reject expired certs with a 5xx permanent failure rather than a 4xx deferral. A bounce rate spike to specific destinations can indicate cert rejection.
  • Process count anomalies. If smtpd processes are consumed by clients retrying failed TLS handshakes, the process count anomaly appears alongside the TLS errors.

Correlating deferred queue growth with TLS failure patterns and delivery rate drops shortens the diagnosis from “mail is slow” to “outbound TLS is broken to specific destinations, check cert expiry.”