The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix / postfix-monitoring-maturity-model ▌

Operations Guides

Postfix monitoring maturity model: from survival to expert

Postfix is a modular, queue-based MTA where mail flow depends on a chain of cooperating daemons, filesystem-backed queues, DNS resolution, and external filters. Monitoring maturity is not about collecting more metrics for their own sake. It is about closing the gap between what you can detect and what actually causes incidents: inode exhaustion hiding behind “No space left on device,” active queue saturation from one slow destination, silent filter backpressure, or deferred queues that grow for hours before anyone notices.

This four-level progression maps survival checks to expert-grade observability. Identify your current level, then target the signals at the next level that close your biggest blind spots. The model reflects incident patterns where teams had monitoring in place but still missed the signal that mattered. The gap was rarely “no monitoring” – it was monitoring the wrong layer, or collecting a metric without alerting on its rate of change.

flowchart TD
    L4["Level 4: Expert
Latency distributions, per-transport
throughput, reputation checks"] L3["Level 3: Mature
Per-destination metrics, filter health,
queue breakdown, pattern detection"] L2["Level 2: Operational
Flow rates, bounce rate, TLS,
security signals, inode monitoring"] L1["Level 1: Survival
Master alive, ports listening,
queue not growing, disk not full"] L1 --> L2 --> L3 --> L4

Level 1: Survival

Survival monitoring answers one question: is Postfix running and can it accept mail? If any of these checks fail, you have a complete outage. These are the signals your paging system must cover before anything else.

SignalWhy it mattersWhat to check
Master process livenessNo master means no mail reception or delivery. Complete outage.PID file at /var/spool/postfix/pid/master.pid exists and references a running process named master.
SMTP listener responsiveSocket present but no greeting means master issue or all smtpd processes busy.TCP connect to port 25, expect a 220 greeting with your hostname within 2 seconds.
Total queue not growingUnbounded queue growth is the earliest indicator of delivery failure.Count files across all queue subdirectories. Compare against a rolling baseline.
Queue partition has spaceDisk full halts all mail I/O immediately.df -h /var/spool/postfix below your alert threshold (typically 90%).

Master PID file check. The PID file must exist and point to a running process. A stale PID file after an unclean shutdown is a common false positive. Compare the PID file mtime against the process start time to detect this. In containers, PID namespace confusion can cause the PID file to reference PID 1 or the wrong namespace entry.

Greeting test. A TCP socket that accepts connections but never sends a 220 banner means the smtpd pool is exhausted or the master process is stuck. Test the full path, not just the port bind. If postscreen is enabled, it answers on port 25 and passes to smtpd internally, so test both the postscreen path and the backend smtpd path.

Queue size as a survival signal. At this level you do not need per-queue breakdowns. You need to know whether the total file count is trending up. A static threshold (for example, “page if total queue exceeds 50,000”) is sufficient for survival. Velocity-based alerting comes at Level 2.

Disk space only. Level 1 tracks disk space, not inodes. This is a known blind spot. Most teams at this level have been surprised by “No space left on device” when df -h shows 40% free. Inode monitoring arrives at Level 2, but if you take one thing from this article before finishing, add df -i to your survival checks today.

Level 2: Operational

Operational monitoring adds early warning. You can detect degradation before it becomes an outage. The focus shifts from “is it running?” to “is it draining?”.

SignalWhy it mattersWhat to check
Injection vs delivery ratePersistent injection exceeding delivery causes queue growth.Parse logs for client= (injection) and status=sent (delivery) over 5-minute windows. Ratio should be near 1:1.
Deferred rateHigh deferred ratio indicates destination problems or policy rejections.Count status=deferred lines. Healthy systems run below 5% of injection. Above 20% warrants investigation.
Active queue vs limitActive queue at qmgr_message_active_limit means the queue manager cannot schedule new deliveries.Compare active queue file count against postconf -h qmgr_message_active_limit (default 20,000).
Deferred queue growth rateVelocity matters more than absolute size. A static 5,000 that is shrinking is fine. The same 5,000 growing at 1,000/hour is not.Track the derivative of deferred queue file count over time. Any sustained positive growth over 4 hours is abnormal.
Bounce rateHigh bounce rates damage IP reputation and indicate data quality or routing problems.Count status=bounced lines. Above 1% sustained for transactional mail warrants a ticket. Above 5% is critical.
Inode utilizationPostfix creates one file per queued message plus metadata. Inodes exhaust before disk space.df -i /var/spool/postfix. Alert below 10,000 free inodes or above 90% usage.
File descriptor countEach connection and queued message consumes an fd. Exhaustion halts new connections.Sum open fds across Postfix processes. Compare against ulimit -n. Alert above 80% of soft limit.
SMTPD process utilizationServices at maxproc cause “to limit” log messages and connection delays.Check for to limit in logs. Monitor smtpd process count against maxproc in master.cf.
TLS certificate expirationExpired certificates cause mandatory TLS destinations to defer silently.Monitor both inbound (smtpd) and outbound (smtp) certificate expiry dates. Do not forget client certificates for mutual TLS.
Authentication failure rateBrute force against submission ports (587, 465).Count SASL authentication failed lines. Baseline is very low. Above 100/minute from a single IP is active attack.
Relay denial rateConfirms smtpd_recipient_restrictions are blocking unauthorized relay attempts.Count relay access denied lines. Any successful unauthorized relay is a security incident.

The deferred growth trap. Teams at this level often set a static threshold (“page if deferred exceeds 10,000”) and miss the transition from linear to exponential growth. Once throughput saturates, retried messages compound new arrivals. Track the rate of change, not just the absolute count. A deferred queue growing at 1,000 messages/hour with no plateau is a page, regardless of total size.

Inodes, not just bytes. This is the single most common gap at every level. Postfix stores each queued message as an individual file with associated metadata. A deferred queue with hundreds of thousands of messages can exhaust inodes while df -h shows ample free space. The error message “No space left on device” is technically correct but deeply misleading. XFS dynamically allocates inodes and is less susceptible, but ext4 with default inode density is vulnerable on /var partitions.

File descriptors in production. The default soft limit (often 1024) is inadequate for production mail servers. Each smtpd process, each queued message being processed, and each network connection consumes file descriptors. Monitor the sum across all Postfix processes and compare against the configured limit. Production servers commonly need 65,536 or higher.

Level 3: Mature

Mature monitoring moves from individual metrics to correlation. You can identify which destination is causing queue gridlock, which filter is creating backpressure, and which composite failure pattern is unfolding.

SignalWhy it mattersWhat to check
Per-destination delivery metricsOne slow destination can monopolize active queue slots via fair queueing. Identifying the destination is the first step to resolution.Parse deferred logs grouped by recipient domain. Look for one or few domains dominating.
Content filter / Milter healthFilter slowdown causes queue backup long before process failure.Monitor filter response time. Baseline is typically under 1 second. Alert at p99 above 5 seconds.
Queue subsystem breakdownKnowing which queue is growing (maildrop vs incoming vs active vs deferred) narrows the problem immediately.Track file counts per queue subdirectory independently.
Anvil connection stateRapid connection table growth indicates dictionary attack or connection pool abuse.Parse connect from lines, group by client IP. NAT/proxy environments appear as a single client.
Relay recipient map latencySlow recipient verification causes smtpd to hang before queueing.time postmap -q test@example.com hash:/etc/postfix/relay_recipients. Network-backed maps should be under 500ms.
Composite pattern detectionIndividual metrics within thresholds can still form a known failure pattern.Correlate active queue saturation with per-destination deferred rates and filter response times.

Queue gridlock pattern. Active queue near qmgr_message_active_limit with deferred growing steadily, CPU and network low, and one destination dominating deferred entries with “connection timed out” or rate-limit 4xx responses. The queue manager is working as designed, but fair queueing without priority means one slow destination stalls everything. Detecting this pattern requires correlating active queue depth, deferred growth rate, and per-destination breakdown. No single metric tells the whole story.

Filter backpressure pattern. Incoming queue growing while active queue remains small. This is the reverse of normal gridlock and is easy to miss if you only monitor deferred. The filter is accepting mail (incoming grows) but not completing processing (active stays empty because cleanup cannot finish). SMTP clients may time out and retry, amplifying load. The tell is incoming queue growth with healthy active queue depth.

Postscreen and anvil. If postscreen is enabled, its statistics reveal how effectively zombies are being blocked before reaching smtpd. Anvil tracks per-client connection counts and rates in memory. Both are cleared on restart, so there is a brief window after restart where rate limits reset. In NAT or proxy environments, all clients appear as the same IP address, which can trigger false connection count limits.

Postfix does not expose lookup-table timing natively; time postmap -q is a manual probe, not a continuous signal, so continuous monitoring of map latency requires log parsing or external probing of the map backend.

Level 4: Expert

Expert monitoring adds distributions, per-transport granularity, and external reputation signals. The focus shifts from detecting problems to predicting them.

SignalWhy it mattersWhat to check
Message age distributionAverage queue age hides long tails. A distribution reveals whether most mail flows quickly with a stuck minority, or whether the entire queue is aging.Analyze mtime of files in deferred and active directories. Bucket by age ranges.
Per-transport throughputsmtp, local, virtual, and pipe transports have different bottlenecks. Aggregate throughput hides per-transport saturation.Parse delivery logs grouped by transport. Compare throughput against configured concurrency limits per transport.
DNS latency distributionDNS is Postfix’s most critical external dependency. Latency distribution, not just failure rate, reveals resolver degradation before it causes deferrals.Postfix does not export resolver timing directly. Instrument the system resolver (unbound, systemd-resolved) or parse delay= fields from postfix/smtp log entries.
Postqueue responsivenessSlow postqueue -p response indicates queue manager under pressure.Time postqueue -p execution. Normal is under 5 seconds. Above 10 seconds indicates qmgr stress.
Predictive queue growthTime-to-full estimation based on current growth rate and remaining capacity.Calculate runway: free inodes divided by growth rate (files/hour). Alert when runway drops below a threshold.
Reputation / blocklist checksBlocklistings appear within hours of sustained high bounce rates. Proactive checks catch listings before users report them.Query major DNSBLs programmatically for your sending IP(s). Monitor sender score externally.

Postfix’s delay=/delays= fields cover queue residence and connection time, not resolver latency, so a DNS latency distribution requires instrumenting the resolver independently.

SORBS was decommissioned in June 2024; use currently active blocklist sources for reputation checks.

Message age vs message count. A deferred queue with 5,000 messages where the oldest is 10 minutes old is a transient blip. The same 5,000 messages where the oldest is 6 hours old is a delivery crisis. Count alone is insufficient. The distribution of message ages reveals whether retries are succeeding for recent messages while old messages are stuck, or whether the entire queue is aging uniformly.

Reputation as a monitoring signal. Most teams discover blocklistings when external parties report them. By then, deliverability is already damaged. Querying DNSBLs programmatically (Spamhaus, Barracuda) for your sending IPs provides early detection. Correlate with bounce rate trends: sustained bounce rates above 2% precede most blocklist appearances.

Multi-instance contention. If you run multiple Postfix instances (inbound vs outbound, per-customer, submission vs receiving), aggregate metrics hide resource starvation. One instance can exhaust file descriptors or disk I/O while others appear healthy. Per-instance resource accounting is essential at this level.

How to use this model

Assess your current state honestly. Most teams operate at Level 1 with some Level 2 signals. The gap between “we collect it” and “we alert on it” is where most incidents live. A metric you collect but never page on might as well not exist.

Pick one blind spot to close. Do not attempt to jump from Level 1 to Level 4. Pick the single most impactful signal at the next level and implement alerting for it. For most teams, that is deferred queue growth rate (velocity, not absolute size) or inode utilization.

Calibrate thresholds to your workload. The thresholds in this article are starting points. A backup MX legitimately maintains large deferred queues. A marketing sender has higher bounce baselines than a transactional sender. Establish baselines during stable operation, then set thresholds relative to those baselines.

Test recovery procedures, not just detection. Monitoring tells you something is wrong. Knowing which lever to pull (reduce destination concurrency, bypass a failed filter, hold a problem destination’s mail) is a separate skill. Document playbook steps before you need them at 3 a.m.

What teams consistently get wrong

  • Inode blind spot. Almost every major Postfix queue-growth incident eventually hits inode exhaustion. Teams monitor disk space, not inodes. Add df -i today.
  • Static thresholds on dynamic queues. A deferred queue of 5,000 that is growing is more urgent than 20,000 that is draining. Track velocity.
  • DNS as unmonitored dependency. Postfix depends entirely on DNS. Resolver failure looks like “slow mail to everyone.” Monitor resolver latency and failure rate independently.
  • Filter health measured as process liveness. “Is Amavis running” is the wrong question. “Is Amavis responding in under 2 seconds” is the right one. Slow filters cause queue backup before they crash.
  • Bounce rate blindness. Bounces count as “sent” in simple delivery metrics. Explicit bounce rate monitoring is required separately.
  • Ignoring maildrop queue. Local submission failures (cron, monitoring scripts) are invisible in SMTP logs. Pickup daemon failure manifests as missing local mail.

How Netdata helps

Netdata’s per-second collection and anomaly detection shorten the path from symptom to root cause for Postfix incidents.

  • Disk and inode metrics at per-second resolution catch the inode exhaustion cliff before it halts mail delivery. Correlating inode usage with queue file count confirms whether growth is Postfix-driven.
  • Process and file descriptor metrics reveal smtpd pool exhaustion, fd pressure, and process spawn anomalies before they cause connection failures.
  • System-level DNS latency and resolver health provide the independent DNS visibility that Postfix itself cannot surface.
  • ML-based anomaly detection on queue-related filesystem activity (file creation rate, directory growth) flags unusual patterns without requiring static thresholds that miss workload-specific baselines.
  • Correlation across system, disk, network, and process metrics in a single timeline makes it faster to distinguish between a Postfix-internal problem (active queue saturation) and an external dependency failure (DNS, filter, destination reachability).