The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix
POSTFIX · OPERATIONS PLAYBOOK

Postfix rarely crashes — it stalls: mail backs up in the deferred queue, the inodes run out, and delivery stops before the process ever does

A modular mail transfer agent where a single-threaded queue manager schedules delivery through stateless agents, every queued message is a small file on disk, and almost every outbound step waits on DNS. We trace how that design behaves under load, where 'slow mail' turns into a halted queue, and what to do when it does.

"

Postfix's defaults get mail flowing in minutes, then hand you a set of quiet failure modes that only surface once the queue is already deep.

The defaults work. Until a destination slows down, its mail hogs the fair-queue slots in the single-threaded queue manager, and healthy destinations starve behind it. Until the resolver hiccups and every MX lookup fails with Host or domain name not found. Name service error — so mail defers everywhere, and the system looks slow rather than broken. Until the deferred queue fills the partition with tiny files and Postfix dies with No space left on device while df -h still shows free gigabytes, because it ran out of inodes, not bytes. Until a content filter stalls and incoming grows while active stays empty. Until the submission port takes a brute-force run and SASL LOGIN authentication failed floods the log. Until a server accepts mail for unknown recipients, bounces it to forged senders, and lands on a blocklist before anyone files a ticket.

These guides are written for engineers who already run Postfix, not for people setting up their first mail server. The goal is the mental model of how the MTA actually behaves under load, the failure patterns that keep recurring, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last delivery stall.

How Postfix actually runs in production

Postfix is not one program. It is a master supervisor spawning stateless daemons, a chain of disk-backed queues, and a single-threaded scheduler that hands mail to delivery agents which almost always wait on DNS. Most production failures live between these layers, not inside any one of them.

01
master supervisor
The parent process reads <code>master.cf</code> and spawns every other daemon on demand, then reaps them. It never touches mail itself — it owns the process table and the per-service <code>maxproc</code> limits. If the master is gone, the whole MTA is down; if it is up, individual daemons can still be starved.
MASTER
02
smtpd / postscreen (ingress)
Accepts SMTP on port 25 and submission on 587, applies HELO / sender / recipient restrictions and SASL auth. Optional <code>postscreen</code> sheds obvious zombies before smtpd. Each connection is one smtpd process and one file descriptor — slow clients or a connection storm exhaust the pool.
INGRESS
03
cleanup + milters
Sanitises headers, rewrites addresses, runs <code>header_checks</code> / <code>body_checks</code> and any milters, then writes the queue file. CPU-bound on large messages and complex regex; a slow or hung milter blocks the DATA phase and ties up smtpd behind it.
CLEANUP
04
incoming queue
Newly accepted mail waiting for the queue manager to pick it up. When <code>incoming</code> grows while <code>active</code> stays small, the bottleneck is cleanup or a <code>content_filter</code> — backpressure, not a delivery problem.
INCOMING
05
queue manager (qmgr)
The single-threaded brain: it moves mail from <code>incoming</code> to <code>active</code>, schedules retries, and enforces fair per-destination concurrency. Fair, not priority — one slow destination can consume the schedule and starve everyone else. If qmgr hangs, nothing moves.
QMGR
06
active queue
Mail currently eligible for delivery, capped by <code>qmgr_message_active_limit</code> (default 20000). At the cap, qmgr cannot schedule more concurrent deliveries no matter how healthy the destinations are — a cliff-edge, not a slowdown.
ACTIVE
07
delivery agents + DNS
<code>smtp</code>, <code>local</code>, <code>virtual</code> and <code>pipe</code> agents are spawned per delivery, resolve the destination MX via DNS, and hand the message off. They are stateless — all state lives in the queue file. DNS capacity and destination reachability are usually the real constraint, not Postfix.
DELIVER
08
deferred queue + retry
Temporary failures — a 4xx, a timeout, a DNS error — land the message in <code>deferred</code>, which retries with exponential backoff up to <code>maximal_backoff_time</code>. It grows whenever failures outpace retry successes, and the growth compounds because retried mail piles onto new arrivals.
DEFER

Why this matters: 'mail is slow' or 'mail is stuck' can be smtpd exhaustion at ingress, cleanup or milter backpressure filling the incoming queue, the active queue at its cap, one slow destination monopolising fair-queue slots, a DNS resolver failing every MX lookup, or the deferred queue backing up behind a blocklist. The symptom rhymes but each layer has a different signal — and a different fix.

The failures you'll actually see

Most Postfix incidents fall into a small set of recurring patterns. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The deferred queue backup

Mail keeps arriving but stops leaving. A destination becomes unreachable and deliveries defer with Connection timed out; the deferred queue climbs, retries pile onto new arrivals, and exponential backoff keeps messages queued for hours even after the cause clears. Injection greater than delivery is the leading read — long before queue depth confirms it, and before the partition fills.

  • status=deferred (connect to mx…: Connection timed out) repeating in the log
  • Injection rate sustained above delivery rate; delivery near zero for minutes
  • deferred/ file count growing, oldest message age climbing
  • Active queue low — the mail is not moving, not gridlocked
Investigate
CRITICAL

The invisible DNS failure

Delivery collapses to near zero while injection stays normal, and every destination is affected equally. Postfix depends on DNS for every MX lookup; when the resolver fails or times out, mail defers with Host or domain name not found. Name service error — correct behaviour, but the cause is the resolver, not the network. The system reads as 'slow', which is why teams waste the first hour on the wrong metrics.

  • Name service error for name=… type=MX: Host not found, try again
  • All destinations deferring at once, not one domain
  • dig / host failing from the mail server itself
  • No TLS errors and no Connection refused — rules out the network
Investigate
CRITICAL

The inode and disk cliff

Postfix logs fatal: … No space left on device and stops creating queue files, PID files and locks — while df -h still shows free gigabytes. The queue filesystem ran out of inodes, not bytes, because Postfix writes one small file per message and a runaway queue exhausts them first. A cliff-edge with no graceful degradation: acceptance and delivery fail unpredictably at once.

  • fatal: … No space left on device with df -h showing free space
  • df -i on /var/spool/postfix at or near 100%
  • Queue subdirectory with tens of thousands of files
  • Deferred or bounce queue that has grown for hours unnoticed
Investigate
ACTIVE

File descriptor exhaustion

New connections are refused and daemons cannot fork. Every connection and every in-flight queue file holds a descriptor, and the default soft limit of 1024 is far too low for production. At the ceiling the node stops accepting mail and cannot open queue files — a symptom often misdiagnosed as a disk problem. The fix is raising both the OS ulimit and the systemd LimitNOFILE, not restarting into the same wall.

  • too many open files / unable to fork in the log
  • Open descriptors above 80% of the soft ulimit
  • New SMTP connections refused or immediately closed
  • Defunct / zombie Postfix processes accumulating
Investigate
ACTIVE

The rejection wall

Legitimate mail suddenly refused after a deploy, a network change, or a credential rotation. A client that is neither in mynetworks nor SASL-authenticated is told 554 5.7.1 Relay access denied, or a recipient missing from the maps gets 550 … User unknown. The denial usually means your restrictions are working — the fault is a misconfigured sender or a stale map, not the server. Restriction order is where this hides.

  • 554 5.7.1 <addr>: Relay access denied for an app that used to work
  • 550 5.1.1 Recipient address rejected: User unknown spiking
  • Rejections concentrated on one client IP or one deploy window
  • sasl_method empty on the rejected attempts
Investigate
IMMINENT

Brute force on the submission port

A sustained flood of SASL LOGIN authentication failed on port 587 is a credential brute-force in progress. Each failure closes the connection and drives reconnect churn, burning smtpd processes and descriptors; a single guessed password becomes a compromised account that sends spam and gets you blocklisted. A backend auth outage looks similar but fails everyone at once — tell them apart before you act.

  • SASL LOGIN / PLAIN authentication failed flooding from one or many IPs
  • Auth failure rate many times baseline, or >100/min from a single source
  • Connection churn and smtpd process pressure on 587
  • A sudden burst of outbound mail if an account is already compromised
Investigate
Choosing a tool

Best Postfix Monitoring Tools: 10 Compared (2026)

A ranked review of the tools teams actually shortlist here, what each one is genuinely good at, and how the pricing behaves as you scale.

Postfix monitoring maturity levels

Postfix observability works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much your mail flow matters. Most production servers should land at the second level.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: is Postfix alive and is mail moving at all? You will not learn what broke, but you will learn that something broke before users do. Survival is enough for a dev relay or a null client.

  • Master process running PID file present and pointing at a live master — not a stale PID after a crash.
  • SMTP ports listening Ports 25 and 587 accept a connection and return a 220 greeting.
  • Total queue not growing unbounded The sum of all queue directories is stable, not climbing.
  • Queue filesystem disk free /var/spool/postfix has bytes left; a full partition halts all mail I/O.
  • Basic send / receive works A test message actually enters and leaves the system.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production mail servers should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose an incident on its own: queue backups, delivery stalls, resource pressure, reputation risk.

  • Injection / delivery / deferred rates Parsed from the log — the imbalance predicts a backup before depth does.
  • Deferred queue size AND growth rate A shrinking 5000 is fine; a growing 5000 is a crisis onset. Watch the derivative.
  • Active queue vs qmgr_message_active_limit At the cap, qmgr cannot schedule more deliveries — a cliff-edge.
  • Inodes AND disk on the queue filesystem df -i, not just df -h; inodes run out before bytes.
  • Open file descriptors vs ulimit At the limit, connections are refused and queue files cannot open.
  • Bounce rate 5xx failures read as 'success' in naive metrics; watch them explicitly.
  • TLS certificate expiry Not a Postfix metric — an expired cert is a silent, total outage.
  • Auth-failure and relay-denied rates Brute force and open-relay probes show here first.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. A single slow destination consuming the schedule, a content filter creeping past two seconds, DNS latency rising, the maildrop queue quietly stalling. None of these page you on day one — they become the incident on day thirty.

  • Per-destination delivery and defer rates Which domain is consuming the fair-queue slots?
  • Content filter / milter response time 'Is it responding in under 2s', not just 'is it running'.
  • Queue subsystem breakdown maildrop / incoming / active / deferred / hold sized separately.
  • Average and oldest message age Latency the depth number alone cannot show.
  • DNS resolver latency and failure rate Monitored independently of Postfix — the top misdiagnosed cause.
  • Maildrop queue age Local submissions failing are invisible in the SMTP logs.
  • Process spawn / 'to limit' warnings smtpd or cleanup pinned at maxproc, holding connections.
  • Relay recipient map latency A slow or stale map rejects valid mail or hangs smtpd.

Level 4: Expert

Reactive instrumentation after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. Reputation and blocklist checks, per-transport saturation, TLS downgrade detection, message-age distribution, queue-manager latency. Most teams never need every signal here. Add the ones your incident history says you do.

  • Sender reputation / blocklist status Programmatic RBL and feedback-loop checks before deliverability craters.
  • Per-transport throughput and saturation smtp / local / virtual / pipe measured separately.
  • postqueue latency under load A slow postqueue -p is qmgr under internal pressure.
  • Message-age distribution, not just count The long tail of the queue, not the average.
  • DNS lookup latency distribution Slow resolution degrades delivery long before it fails.
  • TLS downgrade / policy-failure detection Mandatory TLS silently falling back to plaintext.
  • Milter DATA-phase timing Where acceptance latency actually goes.
  • Multi-instance resource accounting One Postfix instance starving another on shared hardware.

Operating mistakes worth avoiding

The traps Postfix teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after an incident.

Watching disk space but not inodes

Almost every major queue-growth incident eventually hits inode exhaustion, yet teams monitor only <code>df -h</code>. Postfix writes one small file per message, so a runaway queue exhausts inodes long before bytes — and the error is <code>No space left on device</code> with the disk half empty, which wastes precious minutes. Alert on <code>df -i</code> for <code>/var/spool/postfix</code> explicitly, and keep at least 20% free inodes.

Alerting on queue size, not growth rate

A static threshold like 'deferred > 10000 pages' misses the actual signal. A queue of 5000 that is draining is healthy; the same 5000 growing fast is a crisis beginning. Track the derivative — messages per hour in versus out — because queue buildup compounds exponentially once delivery is saturated. Absolute size tells you where you are; the rate tells you where you're going.

Treating DNS as a reliable dependency

Postfix depends entirely on DNS for MX, PTR and blocklist lookups, but teams monitor Postfix and assume the resolver is fine. A failing resolver looks like 'mail is slow' or 'everything defers', and the deferral text is vague, so the resolver is the single most common misdiagnosed root cause. Monitor resolver latency and failure rate independently — and give Postfix more than one resolver.

Judging the content filter by liveness, not latency

Teams alert on 'is Amavis running', not 'is Amavis answering in under two seconds'. A filter that slows down backs mail up in the <code>incoming</code> queue long before the process dies — incoming growing while <code>active</code> stays small is the tell, and it is rarely monitored. Measure filter response time and treat the incoming queue as the backpressure gauge it is.

Ignoring the maildrop queue

Locally submitted mail — cron jobs, monitoring scripts, <code>sendmail</code> callers — flows through <code>maildrop</code> and the single-threaded pickup daemon, and none of it appears in the SMTP logs. When pickup fails or a <code>postdrop</code> setgid permission breaks after a package update, local mail vanishes silently. A maildrop age check catches it; nothing else will.

Bounce-rate and backscatter blindness

Bounces read as 'successful' in simple delivery metrics, so teams never alert on them — but a spike in 5xx bounces, or a backscatter storm sending NDRs to forged senders, gets you blocklisted fast, and the discovery usually arrives as an outside complaint. Monitor bounce rate and outbound MAILER-DAEMON generation, and reject invalid recipients at SMTP time instead of accepting-then-bouncing.

Leaving TLS failure logging off

Default Postfix logs successful TLS but not failed opportunistic TLS, so a silent fallback to plaintext — a downgrade attack, a broken cipher policy, an expired peer cert — is invisible. Raise <code>smtpd_tls_loglevel</code> and <code>smtp_tls_loglevel</code> to at least 1, and remember outbound TLS failures surface as deferrals, not obvious errors. And certificate expiry is not a Postfix metric at all — it needs an external check.

Accepting default limits in production

The defaults are tuned for a near-idle dev box: 1024 file descriptors, modest <code>maxproc</code> values, conservative concurrency. Production load quietly exceeds them and the symptoms — occasional refused connections, slow delivery, processes stuck at the limit — look like everything except a limit. Raise the fd limit (both <code>ulimit</code> and systemd <code>LimitNOFILE</code>), size <code>maxproc</code>, and tune per-destination concurrency before load forces the lesson.

Postfix runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up Postfix monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing.