The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix / postfix-destination-concurrency-limit ▌

Operations Guides

Postfix destination concurrency limit: tuning per-destination delivery

Postfix does not open unlimited outbound connections to a single destination. The queue manager (qmgr) enforces a per-destination concurrency cap controlling how many simultaneous delivery attempts may target the same recipient domain at once. This parameter, smtp_destination_concurrency_limit, is behind two common operational failures: queue gridlock from one slow destination monopolizing active queue slots, and reputation penalties from hammering a fragile destination that then throttles or blocklists your IP.

The default works for general-purpose MTAs sending moderate volume to healthy destinations. But with rate-limited providers, corporate servers with strict connection caps, or a backlog that needs draining fast, the default can be either too aggressive or too conservative. Since Postfix 2.5, the concurrency limit is a ceiling that an adaptive feedback algorithm approaches and retreats from based on delivery outcomes, similar to TCP congestion control. Understanding both the limit and the feedback algorithm is necessary before making changes.

What it controls

smtp_destination_concurrency_limit (default: $default_destination_concurrency_limit, which defaults to 20) controls the maximum number of simultaneous outbound SMTP deliveries to a single destination. “Destination” here means a resolved next-hop address: typically the MX host of the recipient domain, or a relayhost if one is configured.

When mail for a destination enters the active queue, qmgr opens delivery connections up to the concurrency limit. If 500 messages are queued for example.com and the limit is 20, Postfix opens at most 20 simultaneous connections to example.com’s MX hosts. The remaining messages wait in the active queue until a slot frees.

Key facts:

  • This is a per-destination limit, not a global limit. Postfix will open up to 20 connections to gmail.com and simultaneously up to 20 to outlook.com, and so on for every destination with queued mail.
  • There is no built-in global concurrency cap across all destinations. The global process pool is governed by default_process_limit (default: 100), which caps total spawned smtp delivery agents across all destinations. If you need a hard cap on total outbound connections, tune default_process_limit or use a dedicated relay transport.
  • The limit applies per transport. smtp_destination_concurrency_limit governs the smtp transport (outbound SMTP delivery). The local, virtual, and pipe transports each have their own *_destination_concurrency_limit parameters.

If the limit is too high for a destination’s tolerance, the remote MTA may respond with 4xx deferrals, greylisting, or outright connection throttling. Sustained aggressive concurrency can trigger reputation penalties or blocklistings. If the limit is too low, you cannot drain a backlog fast enough. Mail piles up in the active and deferred queues even though destinations are healthy and willing to accept more.

Adaptive concurrency feedback

Since Postfix 2.5, the per-destination concurrency limit is not a static number. It is a ceiling that an adaptive feedback algorithm approaches and retreats from based on delivery outcomes.

The algorithm operates on pseudo-cohorts: groups of messages sent to the same destination in a burst. The lifecycle:

  1. Initial concurrency: Postfix starts with initial_destination_concurrency (default: 5). The first pseudo-cohort to a destination opens at most 5 simultaneous connections.
  2. Positive feedback: When deliveries in a cohort succeed without connection or handshake failures, Postfix increments the effective concurrency toward the configured ceiling. The rate is controlled by default_destination_concurrency_positive_feedback (default: 1).
  3. Negative feedback: When a delivery fails with a connection or handshake failure (a network-level failure, not a 4xx SMTP response), Postfix decrements concurrency. The amount is controlled by default_destination_concurrency_negative_feedback (default: 1).
  4. Failed cohort limit: If an entire pseudo-cohort fails, Postfix counts it. After default_destination_concurrency_failed_cohort_limit (default: 1) failed pseudo-cohorts, Postfix declares the destination “dead” and suspends delivery to it until a retry is scheduled after $minimal_backoff_time (default 300s).

Officially, feedback values are fractions from 0 to 1: with the default constant value 1, the concurrency window grows at the end of each sequence of length 1/feedback (one successful delivery), so it doubles after a fully successful pseudo-cohort, matching pre-2.5 behavior; the number/concurrency form grows the window by number after each successful pseudo-cohort.

flowchart TD
    A[Mail enters active queue for destination] --> B[Start at initial_destination_concurrency]
    B --> C{Delivery outcome}
    C -->|Success| D[Positive feedback: grow concurrency]
    C -->|Connection or handshake failure| E[Negative feedback: reduce concurrency]
    D --> F{Concurrency < smtp_destination_concurrency_limit?}
    F -->|Yes| C
    F -->|No| G[Hold at ceiling]
    E --> H{Cohort fully failed?}
    H -->|No| C
    H -->|Yes| I[Increment failed cohort counter]
    I --> J{Counter >= failed_cohort_limit?}
    J -->|Yes| K[Declare destination dead, suspend delivery]
    J -->|No| C
    G --> C
    K --> L[Retry after minimal_backoff_time]

The defaults for positive_feedback (1), negative_feedback (1), and failed_cohort_limit (1) are documented as reproducing pre-2.5 Postfix behavior. You do not need to change these unless you have a specific reason.

Critical distinction: 4xx SMTP deferrals are not treated as connection or handshake failures by the concurrency algorithm. A 4xx deferral moves the message to the deferred queue with exponential backoff, but it does not trigger negative feedback on the concurrency window. Only network-level failures (connection refused, connection timed out, TLS handshake failure) affect the adaptive concurrency.

Where it shows up in production

The most common scenario is queue gridlock: a single slow or failing destination consumes a disproportionate share of active queue slots.

The active queue is capped at qmgr_message_active_limit (default: 20,000). When mail for one destination fills the active queue and that destination is slow to accept deliveries, all other destinations are starved. The queue manager uses fair scheduling across destinations but does not deprioritize a destination consuming slots without delivering.

Diagnostic pattern:

  • Active queue is near qmgr_message_active_limit
  • Deferred queue is growing steadily
  • One destination dominates deferred entries with “connection timed out” or rate-limit deferrals
  • Overall delivery rate is flat or declining despite queue depth
  • System CPU and network utilization are low (the constraint is the destination, not your hardware)

In this state, reducing the concurrency limit for the problem destination frees active queue slots for healthy destinations. The TUNING_README warns: “Knee-jerk changes to these parameters in the face of congestion can actually make problems worse.” The right approach is a targeted per-transport override, not a global limit reduction.

The second scenario is the opposite: a backlog of valid mail to a healthy destination where the default concurrency of 20 is too low to drain it. A dedicated transport with a higher concurrency limit can help, but verify the destination can tolerate the load before raising it.

Per-transport overrides

Dedicated transports for fragile destinations

The recommended pattern for destinations with strict rate limits or fragile infrastructure is a dedicated transport in master.cf with an overridden concurrency limit in main.cf.

Step 1: Define the transport in master.cf. The chroot field (5th column) must match your existing smtp service definition:

slow    unix  -       -       n       -       -       smtp
    -o smtp_connect_timeout=5

Step 2: Set the per-transport parameters in main.cf using the transport name as a prefix:

postconf -e 'slow_destination_concurrency_limit=2'
postconf -e 'slow_destination_rate_delay=1s'
postconf -e 'slow_destination_recipient_limit=10'

Step 3: Route specific domains to this transport. Verify your transport map path first; the transport_maps parameter in main.cf must point to this file:

# Confirm the transport map is configured
postconf transport_maps

# Check existing entries before appending
grep fragile-provider.com /etc/postfix/transport

# Add the routing entry
echo 'fragile-provider.com    slow:' >> /etc/postfix/transport
postmap /etc/postfix/transport
postfix reload

The common mistake: operators add the transport in master.cf but forget the corresponding *_destination_concurrency_limit in main.cf. Without it, the transport inherits the default concurrency of 20, defeating the purpose.

smtp_connect_timeout per-transport override

smtp_connect_timeout does not have a per-transport name parameter in main.cf. It must be overridden with -o directly in the master.cf service definition. The timeout applies to TCP connection establishment before any SMTP protocol exchange begins.

destination_rate_delay interaction

When you set destination_rate_delay (which pauses between deliveries to the same destination) alongside a reduced concurrency limit, the failed_cohort_limit parameter becomes critical. With rate delay, a pseudo-cohort takes longer to complete, and the default of 1 may cause destinations to be declared dead too aggressively. The TUNING_README notes that failed_cohort_limit is critical when destination_rate_delay is used. Consider raising it to 2 or 3.

Raising the global default

The TUNING_README notes that the default of 20 “seems enough to noticeably load a system without bringing it to its knees.” For most deployments this is correct. If you need higher throughput to healthy destinations, prefer dedicated transports with elevated concurrency rather than raising the global default. A global increase affects all destinations including fragile ones that may already be on the edge of throttling you.

Signals to watch

SignalWhy it mattersWarning sign
Active queue size vs qmgr_message_active_limitApproaching the limit blocks new deliveries from being scheduled regardless of destination healthRatio above 80% sustained for over 10 minutes
Deferred queue growth rateGrowing deferred queue with a single dominant destination indicates the feedback loop is not keeping upSustained positive growth over 4 hours
Per-destination deferral reason codes“4.7.1 rate limited” or “connection timed out” for one destination while others succeed means concurrency is too high for that destinationOne destination accounting for over 50% of deferrals
Per-destination delivery rateA destination whose delivery rate drops while others stay flat is being throttled or is failingSudden drop for one destination with flat or rising queue depth
SMTP connection latency to specific destinationsElevated connect time indicates network path issues or destination-side throttling before it shows in deferralsConnection establishment consistently over 10 seconds
smtp process count vs default_process_limitThe global process pool can starve high-concurrency destinations if total smtp agents hit the capProcess count sustained above 80% of limit

Monitoring with Netdata

Netdata’s Postfix collector surfaces queue depth and mail flow signals that reveal concurrency-related problems before they become incidents.

  • Per-second active and deferred queue depth: the ratio between them distinguishes a destination problem (deferred growing, active near limit) from a global problem (both growing uniformly).
  • Mail flow velocity correlation: the injected-vs-delivered rate alongside queue depth shows whether the system is draining or accumulating. A widening gap between injection and delivery rates with a growing active queue is the classic gridlock precursor.
  • Deferred queue growth rate: Netdata computes rates per second, so you see the derivative of queue depth, not just the absolute value. A positive slope sustained over minutes is actionable long before any static threshold is crossed.
  • System-level signals: CPU utilization, network connections, and process counts alongside Postfix metrics distinguish “destination is slow” (low CPU, low network) from “we are the bottleneck” (high CPU, process exhaustion).
  • Anomaly detection: ML-based anomaly flags on queue depth and delivery rate catch early divergence from baseline that precedes gridlock, even when absolute values are still within nominal ranges.

For deeper context on how Postfix’s queue architecture creates the failure patterns this parameter controls, see How Postfix actually works in production: a mental model for operators.