The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / postfix / postfix-file-descriptor-limits ▌

Operations Guides

Postfix file descriptor limits: raising ulimit and systemd LimitNOFILE

Postfix logs fatal: socket: Too many open files or begins silently dropping connections when per-process file descriptor limits are too low. The default soft limit on many distributions is 1024, which is adequate for a development mail server but insufficient for any production MTA handling concurrent SMTP sessions and active queue processing simultaneously.

The most common operator mistake is editing /etc/security/limits.conf, restarting Postfix, and finding the limit unchanged. On systemd-managed distributions, services do not read PAM limits. The effective file descriptor limit for a systemd service comes from the unit file’s LimitNOFILE directive, the global DefaultLimitNOFILE in /etc/systemd/system.conf, or the kernel’s compiled-in default. Editing only one layer leaves the limit silently unchanged.

Where the limit comes from

Three layers can set the per-process file descriptor limit for Postfix. Which one wins depends on how the process was started.

LayerMechanismApplies when
Kernelfs.file-max (system-wide total, not per-process)Always; caps total open FDs across all processes
systemdLimitNOFILE in unit file or DefaultLimitNOFILE in system.confPostfix started by systemd (the default on modern distros)
PAM/etc/security/limits.conf with nofile entriesPostfix started from a login shell or legacy init script

systemd overrides PAM for services it manages. When systemd starts the Postfix master process, it applies the LimitNOFILE value from the unit file (or the DefaultLimitNOFILE from /etc/systemd/system.conf if the unit file does not specify one). It does not consult /etc/security/limits.conf at all. The master process then passes its limits to all child daemons it spawns: smtpd, smtp, cleanup, qmgr, bounce, and so on.

The kernel’s fs.file-max is a separate constraint. It caps the total number of open file descriptors across all processes system-wide. Even if you set LimitNOFILE=1000000 on the Postfix unit, the system can still run out of file descriptors globally if fs.file-max is too low.

flowchart TD
    A["Postfix master starts"] --> B{"Launched by systemd?"}
    B -->|"Yes (modern distros)"| C["Unit LimitNOFILE or
DefaultLimitNOFILE from system.conf"] B -->|"No (init script / shell)"| D["PAM limits.conf or
inherited shell ulimit"] C --> E["Per-process RLIMIT_NOFILE
set at exec time"] D --> E F["Kernel fs.file-max"] --> G["System-wide total FD cap
across all processes"] E --> H["Effective limit for
master + all child daemons"] G -.->|"bounded by kernel total"| H

A compile-time constraint also exists but is irrelevant on modern systems. Postfix versions before 2.4 required recompilation with a larger FD_SETSIZE to support more than 1024 file descriptors per process. Postfix 2.4 and later use scalable I/O multiplexing (epoll on Linux 2.6+, kqueue on BSD) and are not constrained by FD_SETSIZE. If you are running Postfix 3.x on any modern Linux kernel, the runtime limit is the only constraint.

Prerequisites

  • Postfix running on a systemd-managed Linux distribution (CentOS/RHEL 7+, Debian 8+, Ubuntu 16.04+)
  • root or sudo access
  • Knowledge of your current default_process_limit (check with postconf -h default_process_limit)

Procedure

Step 1: Check the current limit

Confirm what limit the Postfix master process actually has before changing anything.

# Check the master process file descriptor limit
MASTER_PID=$(cat /var/spool/postfix/pid/master.pid)
cat /proc/$MASTER_PID/limits | grep "Max open files"

# Alternative using prlimit
prlimit -p $MASTER_PID --nofile

The output shows both a soft limit and a hard limit. The soft limit is what the kernel enforces at runtime. Postfix does not raise its own limit programmatically, so the soft limit is the operative constraint. A soft limit of 1024 is the default on many distributions and is almost certainly too low for production.

Step 2: Raise the limit via systemd drop-in

Use systemctl edit to create a drop-in override rather than editing the packaged unit file directly. Package updates (RPM or DEB) can overwrite /usr/lib/systemd/system/postfix.service, silently reverting your changes. Drop-in overrides in /etc/systemd/system/postfix.service.d/ survive package updates.

# Create or edit the drop-in override
systemctl edit postfix.service

In the editor, add the following:

[Service]
LimitNOFILE=65536

This sets both the soft and hard limits to 65536. If you want different soft and hard values, use the colon syntax: LimitNOFILE=1024:65536.

Avoid LimitNOFILE=infinity. On systemd releases before 240, infinity resolves to a hard-coded 65535; from systemd 240 onward it resolves to the kernel’s /proc/sys/fs/nr_open value (default 1048576). Either way the result is finite and depends on kernel state, so set an explicit numeric limit.

After saving the drop-in:

# Reload systemd to pick up the override
systemctl daemon-reload

# Restart Postfix (reload is not enough; LimitNOFILE is applied at process start)
# WARNING: this drops all active SMTP connections
systemctl restart postfix

A postfix reload is not sufficient here. LimitNOFILE is set by systemd at exec time when the master process starts. Only a full restart causes systemd to re-exec the process with the new limit.

Step 3: Raise the kernel-wide cap if needed

Check whether the system-wide limit is adequate for your total FD budget across all processes, not just Postfix.

# Check current system-wide limit
sysctl fs.file-max

# Check current system-wide usage
cat /proc/sys/fs/file-nr
# Output columns: allocated  unused-allocated  system-max

If fs.file-max is too low relative to your expected total usage, raise it persistently:

echo "fs.file-max = 2097152" > /etc/sysctl.d/99-postfix-fd.conf
sysctl -p /etc/sysctl.d/99-postfix-fd.conf

This change takes effect immediately and survives reboots. No Postfix restart is needed for kernel parameter changes, but the per-process limit still needs to come from step 2.

Step 4: Handle non-systemd Postfix installations

If Postfix is not managed by systemd (legacy init script, manual start from shell, or container without systemd), use PAM limits instead.

Edit /etc/security/limits.conf:

postfix  soft  nofile  65536
postfix  hard  nofile  65536

Then restart Postfix from a session that has loaded the new PAM limits. This typically means logging out and back in, or starting a new login shell before running postfix start. The root user starting Postfix must also have the limit applied; add entries for both root and postfix if root starts the master process.

In containers, the limit is usually inherited from the container runtime or the host’s default cgroup settings. Check with cat /proc/1/limits | grep "Max open files" inside the container. You may need to set the limit in the container runtime (Docker’s --ulimit nofile=65536:65536 or the equivalent in your orchestrator).

Verifying it works

After restarting Postfix, confirm the limit propagated to both the master and child processes.

# Verify master process limit
MASTER_PID=$(cat /var/spool/postfix/pid/master.pid)
cat /proc/$MASTER_PID/limits | grep "Max open files"

# Verify a child smtpd process (child inherits master's limits)
# Note: smtpd processes only exist when there are active or recent connections
SMTPD_PID=$(pgrep -x smtpd | head -1)
if [ -n "$SMTPD_PID" ]; then
  cat /proc/$SMTPD_PID/limits | grep "Max open files"
else
  echo "No smtpd process currently running"
fi

Check current FD usage to confirm you have headroom:

# Count open FDs per Postfix process, sorted by usage
for pid in $(pgrep -f postfix); do
  count=$(ls /proc/$pid/fd 2>/dev/null | wc -l)
  name=$(cat /proc/$pid/comm 2>/dev/null)
  echo "$count $name (pid $pid)"
done | sort -rn

# Total FDs across all Postfix processes
for pid in $(pgrep -f postfix); do
  ls /proc/$pid/fd 2>/dev/null | wc -l
done | awk '{s+=$1} END {print s}'

Sizing file descriptor limits

File descriptor consumption in Postfix comes from multiple sources per daemon process. Each smtpd or smtp process typically consumes:

  • 1 FD per accepted client connection
  • 1+ FDs for queue file access during message processing
  • Variable FDs for lookup table connections (LDAP, MySQL, PostgreSQL, hash maps)

The master process also holds 1 FD per listening socket (port 25, 587, etc.) and 1 FD per child process for IPC.

The sizing heuristic is:

LimitNOFILE >= (default_process_limit * FDs_per_process) + headroom_for_lookups + queue_files

With default_process_limit at its default of 100, and assuming 3-5 FDs per process plus lookup table connections, a minimum of 4096 is reasonable for low-to-moderate volume sites. For high-volume relays with 500+ concurrent processes or heavy use of SQL/LDAP lookup tables, 65536 is a common production setting.

# Check your process limit
postconf -h default_process_limit

# Check per-service maxproc settings in master.cf
postconf -M | awk '{print $1, $2, $7}'  # service, type, maxproc

Factor in the active queue depth. Messages in the active queue are being actively processed, and each may hold file descriptors open during delivery. If your active queue routinely runs at thousands of messages, ensure your FD limit accounts for that concurrent access, not just the process count.

Common pitfalls

Editing the packaged unit file directly. Files under /usr/lib/systemd/system/ are owned by the package manager and will be overwritten on the next Postfix package update. Always use systemctl edit postfix.service to create a drop-in under /etc/systemd/system/postfix.service.d/.

Using LimitNOFILE=infinity. The effective limit from infinity varies by systemd version and may resolve to 65535 on older releases. Set an explicit numeric limit instead.

Only raising the soft limit. If the hard limit remains at 1024 and a subprocess tries to raise its soft limit, it cannot exceed the hard limit. systemd’s LimitNOFILE=65536 sets both soft and hard to the same value, which is the simplest approach.

Confusing fork failures with FD exhaustion. On systems using cgroup v2 (the default on most modern distributions), the cgroup’s pids.max limit can cause fork failures that produce similar symptoms: “unable to fork” in logs, new connections refused, processes not spawning. If you raise LimitNOFILE and the problem persists, check the cgroup PID limit:

# Check the current cgroup PID limit for Postfix (cgroup v2)
cat /sys/fs/cgroup/system.slice/postfix.service/pids.max 2>/dev/null
# Check process/thread counts
cat /proc/$(cat /var/spool/postfix/pid/master.pid)/status | grep Pid

If pids.max is set low, raise it in the systemd unit:

[Service]
LimitNOFILE=65536
TasksMax=infinity

Forgetting systemctl daemon-reload. After creating or modifying a drop-in, systemd must reload its configuration. Without daemon-reload, the restart uses the old limit.

Restarting only the master, not the children. When you restart Postfix via systemd, the master process and all children restart. If you manually kill and restart only the master, old child processes may retain their old limits. Always use systemctl restart postfix for a clean restart.

Monitoring FD utilization

If you run Netdata alongside Postfix, the relevant correlations are:

  • Per-process FD counts. Netdata collects open FD counts per process. Watch smtpd and smtp trends relative to the soft limit to catch growth before it becomes an outage.
  • System-wide FD utilization. Netdata tracks fs.file-nr against fs.file-max, which tells you whether the kernel-wide cap is the binding constraint rather than the per-process limit.
  • Queue depth correlation. Cross-referencing FD usage spikes against active and deferred queue sizes distinguishes backlog-driven growth from connection-volume-driven growth.
  • Process counts vs. maxproc. If smtpd process counts approach the per-service limit, FD pressure follows predictably since each process holds multiple descriptors.