The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix / postfix-connection-refused ▌

Operations Guides

Postfix Connection refused: blocked port 25 and rejected outbound delivery

You see status=deferred (connect to mx.example.com[192.0.2.1]:25: Connection refused) filling your mail logs. The deferred queue is growing. Postfix cannot establish outbound TCP connections to destination mail servers.

“Connection refused” is not the same as “Connection timed out.” Connection refused means the TCP handshake was actively rejected with a RST packet: the destination IP is reachable, but nothing is listening on that port, or something in the path is rejecting the connection. Connection timed out means the SYN packet was silently dropped: a firewall rule, a network partition, or an ISP dropping egress traffic.

A sub-second connect delay in the log’s delays= field confirms Connection refused. A connect delay near your smtp_connect_timeout (default 30 seconds) confirms Connection timed out.

What this means

When Postfix logs Connection refused, the smtp delivery agent attempted a TCP connect to the destination MX, relayhost, or transport_maps target on port 25 and received an immediate RST. Postfix classifies this as a temporary failure (DSN 4.4.1), defers the message, and schedules a retry with exponential backoff. The deferred queue grows as long as the condition persists.

The canonical log pattern:

postfix/smtp[PID]: QUEUEID: to=<user@domain>, relay=none, delay=0.8,
  delays=0.1/0.1/0.5/0.1, dsn=4.4.1, status=deferred
  (connect to mx.example.com[192.0.2.1]:25: Connection refused)

Three fields tell you what happened:

  • relay=none: Postfix never reached the destination. No SMTP session was established.
  • delays= third field (connect time): Sub-second means actively refused. Near 30 seconds means timed out. This is your primary triage signal.
  • connect to <host>[<ip>]:<port>: The specific host and port Postfix tried to reach. This tells you whether the problem is a content filter on localhost, a relayhost, a transport_maps target, or a destination MX.

Common causes

CauseWhat it looks likeFirst thing to check
Cloud or ISP egress blocking port 25All outbound deliveries to any MX refused; relay=none on every messageTest direct TCP to a well-known MX on port 25 from the server
relayhost set but not listeningEvery outbound message defers to the same relayhost address and portpostconf -h relayhost, then test that host on port 25
Content filter port downLogs show connect to 127.0.0.1[127.0.0.1]:10024: Connection refusedCheck whether amavisd, rspamd, or other filter process is running
transport_maps overriding relayhostSome domains deliver normally; others refuse despite relayhost being setpostconf -h transport_maps and inspect entries for affected domains
Destination MX not accepting connectionsOnly one destination affected; all others deliver normallydig MX domain, then test TCP to each returned MX on port 25
Postfix delivering to itselfRefused to own IP on port 25; smtpd only listening on localhostCheck mydestination, virtual_alias_domains, and inet_interfaces

Quick checks

# Confirm the error pattern and count recent refusals
grep 'Connection refused' /var/log/mail.log | tail -20

# Check connect delay timing (refused is sub-second; timed out is near 30s)
grep 'Connection refused' /var/log/mail.log | grep -oE 'delays=[^ ]+' | tail -10

# Check if relayhost is configured
postconf -h relayhost

# Check transport_maps for domain-specific overrides
postconf -h transport_maps

# Check smtp_connect_timeout (confirms timeout boundary)
postconf -h smtp_connect_timeout

# Test direct TCP to a known MX on port 25
nc -w 5 -zv gmail-smtp-in.l.google.com 25

# Check for content filter refusals to localhost
grep 'connect to 127.0.0.1' /var/log/mail.log | tail -10

# See which destinations are being refused and how often
grep 'Connection refused' /var/log/mail.log | grep -oE 'connect to [^ ]+' | sort | uniq -c | sort -rn | head

# Current deferred queue size
find /var/spool/postfix/deferred -type f 2>/dev/null | wc -l

On RHEL and CentOS systems, substitute /var/log/maillog for /var/log/mail.log.

How to diagnose it

Work through the decision tree in order. The log line tells you which path to take based on the connect to target.

flowchart TD
    A["Log shows Connection refused"] --> B{"Refused target?"}
    B -->|"127.0.0.1 port"| C["Content filter down"]
    B -->|"relayhost"| D["Test relayhost port 25"]
    B -->|"destination MX"| E{"All destinations
or one domain?"} E -->|"All destinations"| F["Port 25 egress blocked"] E -->|"One domain"| G["MX down or transport override"] C --> H{"Filter process running?"} H -->|No| I["Restart filter"] H -->|Yes| J["Check filter port config"] D --> K{"Reachable?"} K -->|No| L["Fix relayhost or remove"] K -->|Yes| M["Check transport_maps"] F --> N["Route via relay on 587 or 465"] G --> O["dig MX and inspect transport_maps"]

Step 1: Identify the refused target. Parse the connect to <host>[<ip>]:<port> field from the log line. If it is 127.0.0.1 or localhost, this is a content filter problem, not an outbound delivery problem. If it is a relayhost address, skip to step 4. If it is a destination MX, continue.

Step 2: Determine scope. Are all destinations refused, or just one domain? Extract and group the refused targets:

grep 'Connection refused' /var/log/mail.log | grep -oE 'connect to [^ ]+' | sort | uniq -c | sort -rn | head

If every destination shows Connection refused, the problem is on your side: port 25 egress is blocked, or Postfix cannot route traffic. If only one or a few domains are affected, the destination itself is the issue, or a transport_maps entry is routing those domains somewhere unexpected.

Step 3: Test port 25 egress directly. Pick a well-known MX and test TCP connectivity:

# Test TCP connectivity to a known MX on port 25
nc -w 5 -zv gmail-smtp-in.l.google.com 25

If this fails with Connection refused or times out, your server cannot reach the internet on port 25. This is almost certainly ISP or cloud provider egress filtering. AWS EC2 blocks outbound port 25 by default and lifts the block only on request; Azure blocks outbound SMTP on port 25 for VMs (exemption requestable); GCP blocks outbound port 25 to destinations outside your VPC; OVH blocks it on VPS (requestable); DigitalOcean blocks it (unblock via support ticket). Hetzner does not block outbound port 25 by default. Many residential ISPs (Comcast, Verizon, AT&T) do the same. Check your provider’s documentation.

Step 4: Check relayhost configuration. If postconf -h relayhost returns a value, every outbound message should go through that host. Show the relayhost, then test it manually:

# Show relayhost configuration
postconf -h relayhost
# Test connectivity manually based on output, e.g.:
# nc -w 5 -zv smtp.relay.example 587

If the relayhost refuses connections, it is down or not listening on the configured port. Fix the relayhost, or temporarily remove it to allow direct delivery (if your provider does not block port 25).

Step 5: Check transport_maps overrides. If both relayhost and transport_maps are set, transport_maps takes precedence for matching domains. A transport map entry for a specific domain bypasses the relayhost entirely, causing direct connections that may be refused if port 25 egress is blocked:

# Show transport_maps configuration
postconf -h transport_maps

# Test lookups for affected domains
postmap -q example.com hash:/etc/postfix/transport

Step 6: Check content filter health. If logs show connect to 127.0.0.1[127.0.0.1]:10024: Connection refused, your content filter (typically Amavis, Rspamd, or a commercial filter) is not running or not listening on the configured port. This is not an outbound delivery problem, but it looks identical in the logs:

# Check if content_filter is configured
postconf -h content_filter

# Test the filter port directly
nc -zv 127.0.0.1 10024

# Check filter process status
systemctl status amavisd 2>/dev/null || systemctl status rspamd 2>/dev/null

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Deferred queue growth rateDeferrals accumulate with exponential backoffSustained positive growth over multiple 5-minute windows
Delivery rate vs injection rateDivergence means outbound delivery is failing while mail keeps arrivingDelivery rate dropping below 80% of injection rate
Deferral reason distributionDistinguishes Connection refused from timed out, 4xx, or DNS failuresSpike in “Connection refused” as a percentage of all deferrals
Connect delay from delays= fieldSub-second means refused; approximately 30s means timed outShift from sub-second to 30s delays indicates change from port closed to port filtered
Content filter responseFilter port down appears as Connection refused to localhostAny Connection refused to 127.0.0.1
Relayhost reachabilityIf relayhost is your single outbound path, its failure blocks all mailIntermittent or sustained refusal from relayhost address

Fixes

Port 25 egress blocked by ISP or cloud provider

This is the most common cause for servers in cloud environments. AWS EC2 blocks outbound port 25 by default. Many residential ISPs do the same. The fix is to route outbound mail through a relay service on port 587 or 465.

These commands modify live Postfix configuration. Run postfix reload after applying changes. Test with a single message before flushing the queue.

# Set relayhost to a submission relay (adjust host and port)
postconf -e 'relayhost = [smtp.relay.example]:587'

# Enable SASL authentication for the relay
postconf -e 'smtp_sasl_auth_enable = yes'
postconf -e 'smtp_sasl_password_maps = hash:/etc/postfix/sasl_passwd'

# Create the password map
echo '[smtp.relay.example]:587 username:password' > /etc/postfix/sasl_passwd
postmap /etc/postfix/sasl_passwd
chmod 600 /etc/postfix/sasl_passwd /etc/postfix/sasl_passwd.db

# Reload configuration
postfix reload

This change takes effect immediately for new delivery attempts. Existing deferred messages will be retried against the new relayhost on the next scheduled attempt.

relayhost configured but not listening

If the relayhost itself is down or not accepting connections, all outbound mail stalls. Fix the relayhost service. As a temporary measure, you can remove the relayhost to allow direct delivery, but this only helps if your provider does not block port 25:

# Temporarily remove relayhost for direct delivery
postconf -e 'relayhost ='
postfix reload

Warning: if port 25 egress is also blocked (common in cloud environments), removing relayhost makes things worse, not better. Verify port 25 egress works before relying on direct delivery.

Content filter port down

If the filter process crashed or was never started after a reboot, mail queues up with Connection refused to the filter port. Restart the filter:

# Restart the content filter
systemctl restart amavisd  # or rspamd, or your filter service

In an emergency, you can bypass the filter entirely. This accepts unfiltered mail rather than queueing indefinitely:

# Emergency bypass: mail will be unscanned
postconf -e 'content_filter ='
postfix reload

Re-enable filtering as soon as the filter process is healthy.

transport_maps overriding relayhost

If transport_maps routes specific domains directly (bypassing relayhost), those domains will fail when port 25 egress is blocked. Inspect the transport map and either remove the override or point it through the relay:

# View the transport map source file
postconf -h transport_maps

# Look up routing for a specific domain
postmap -q problem.domain.com hash:/etc/postfix/transport

Fix the transport map entry, rebuild it with postmap, and reload Postfix.

Destination MX not accepting connections

If only one destination is affected, the problem is on their end. Postfix will retry per its backoff schedule. No action is needed unless the destination is business-critical, in which case consider using a backup relay or contacting the destination postmaster.

Prevention

  • Know your provider’s port 25 policy before deploying. AWS and many ISPs block outbound port 25. Plan for relay-based delivery from the start.
  • Monitor deferred queue growth rate, not just absolute size. A static queue of 5,000 messages that is shrinking is fine. The same size growing rapidly is a crisis.
  • Monitor content filter process health and response time. A filter that is running but unresponsive causes the same symptoms as one that is down.
  • Document relayhost and transport_maps dependencies. A transport_maps entry that silently overrides relayhost is easy to forget during incident triage.
  • Include port 25 connectivity tests in deployment checks. A simple nc -zv to a known MX during provisioning catches egress blocks before production traffic hits them.
  • Alert on deferral reason distribution. A spike in “Connection refused” as a percentage of deferrals distinguishes port 25 blocking from other delivery problems.

How Netdata helps

  • Per-second deferred queue depth and growth rate. The derivative (growth rate) catches queue buildup from Connection refused before it reaches crisis, and is more actionable than absolute count.
  • Mail flow velocity (injection vs delivery rate). Divergence between these two rates is the earliest indicator that outbound delivery is failing, often minutes before the deferred queue becomes visibly large.
  • Postfix log parsing for deferral reason codes. Surfaces whether the dominant error is Connection refused (port closed), Connection timed out (port filtered), 4xx responses (destination throttling), or DNS failures, determining which diagnostic path to follow.
  • Content filter process and port health. Independent checks on the filter process and its listening port catch filter outages before they manifest as Postfix delivery failures.
  • TCP connectivity probes. Tests relayhost and content filter ports independently of Postfix, confirming whether the problem is configuration or network reachability.
  • DNS resolver latency and failure rate. Rules out DNS as a contributing factor; DNS failure causes a different deferral pattern that requires a different fix.