The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-file-descriptor-limits ▌

Operations Guides

uWSGI file descriptor limits: raising ulimit -n and systemd LimitNOFILE

The default per-process file descriptor limit on many Linux distributions is 1024. For a uWSGI instance running multiple workers, each holding connections, sockets, log files, and database handles, that ceiling is too low for production.

File descriptor exhaustion in uWSGI is silent. When the limit is hit, accept() and open() calls fail with EMFILE. New connections are rejected with no uWSGI-level error, no log entry, and no stats counter reflecting the problem. Clients see connection resets or timeouts. The master process stays alive. Worker status looks normal. The only evidence is at the OS level.

This guide covers raising the limit durably across the service manager, the OS, and uWSGI itself, plus leak detection, headroom sizing, and verifying the running limit via /proc.

The limit hierarchy

File descriptor limits cascade through several layers. Each layer can constrain the effective limit, and the lowest value wins.

flowchart TD
    A["Kernel fs.file-max\nhost-wide ceiling"] --> B["RLIMIT_NOFILE\nper-process limit"]
    B --> C["systemd LimitNOFILE\nOR shell ulimit -n"]
    C --> D["uWSGI --max-fd\noptional internal cap"]
    D --> E["Master process\ninherits effective limit"]
    E --> F["Worker processes\ninherit at fork"]

Setting ulimit -n in your shell has no effect on a systemd-managed service. systemd applies its own limits from the unit file at service start, overriding whatever the shell environment provided. Always verify the running limit via /proc/<pid>/limits, not ulimit -n.

What proper fd limits give you

Beyond preventing EMFILE, a correctly sized limit provides:

  • Headroom for traffic bursts. A burst that doubles normal connection count should not approach the ceiling.
  • Leak detection window. A steadily climbing per-worker fd count over hours or days indicates a leak. A generous limit gives you time to detect and respond before exhaustion. A tight limit means you hit the wall first.
  • Coverage for all fd consumers. File descriptors are not just connections. They include the listening socket, log files, the stats server socket, spooler files, pipes, and anything the application opens directly (database connections, cache connections, temp files).

Prerequisites

Before choosing a target limit, gather these values:

InputWhere to find itWhy it matters
Worker countuWSGI config (--processes) or stats workers[]Each worker has its own fd table
Expected connections per workerTraffic patterns, upstream proxy configOne fd per active connection
Non-connection fd consumersApplication code, database pool size, loggingOften overlooked, can be significant
Current effective limit/proc/<pid>/limits for the running processBaseline for sizing the increase

Checking the current limit

The shell ulimit -n shows the limit for your current shell session, not the limit of a running uWSGI process managed by systemd. The authoritative source is /proc/<pid>/limits.

# Check the limit of a running uWSGI master process
awk '/^Max open files/ {print $4}' /proc/$(cat /tmp/uwsgi.pid)/limits

# Check per-worker fd usage and effective limit
for pid in $(pgrep -P $(cat /tmp/uwsgi.pid)); do
    count=$(ls /proc/$pid/fd 2>/dev/null | wc -l)
    limit=$(awk '/^Max open files/ {print $4}' /proc/$pid/limits)
    echo "pid=$pid fds=$count limit=$limit"
done

If you do not have a pidfile, find the master PID with pgrep:

# Find the uWSGI master process
pgrep -f 'uwsgi.*master'

Setting the limit under systemd

For uWSGI running under systemd, the service unit file controls the fd limit. The LimitNOFILE directive in the [Service] section sets RLIMIT_NOFILE for the process.

The cleanest approach is a systemd override, which persists across package updates:

# Create or edit an override without modifying the original unit file
systemctl edit uwsgi

In the editor, add:

[Service]
LimitNOFILE=65536

Apply the change:

# Reload systemd unit files
systemctl daemon-reload

# Restart the service to pick up the new limit
# WARNING: this interrupts in-flight requests unless using chain reload
systemctl restart uwsgi

Setting the limit from a shell

If uWSGI is started from a shell or init script rather than systemd, set the limit before launching:

# Set the soft and hard limit for the current shell session
ulimit -n 65536

# Then start uWSGI
uwsgi --ini /etc/uwsgi/app.ini

This does not persist across reboots. For persistence without systemd, use /etc/security/limits.conf:

# /etc/security/limits.conf
uwsgi  soft  nofile  65536
uwsgi  hard  nofile  65536

/etc/security/limits.conf applies to PAM-based logins and may not affect system services started by systemd. Verify the effective limit after restart.

uWSGI’s –max-fd option

uWSGI accepts --max-fd to cap the maximum number of file descriptors it will use internally. This matters when the OS reports a very high RLIMIT_NOFILE and you want to prevent uWSGI from sizing internal data structures based on that large number.

# uWSGI configuration
max-fd = 65536

In Emperor mode, --max-fd set on the Emperor process is expected to propagate to vassals, avoiding per-vassal configuration.

--max-fd works by calling setrlimit() to set both the soft and the hard RLIMIT_NOFILE to the requested value. The uWSGI documentation states that it requires root privileges; in practice setting a value within the current hard limit works without root, while raising the hard limit above its inherited value needs root (or CAP_SYS_RESOURCE).

Emperor propagation is via rlimit inheritance: the Emperor forks vassals, so a vassal that does not set max-fd of its own inherits the Emperor’s raised RLIMIT_NOFILE. A vassal whose own configuration sets max-fd applies that value instead.

Verifying the running limit

After restarting or reloading uWSGI, verify that the new limit is in effect:

# Verify the master process limit (soft and hard)
awk '/^Max open files/ {print "soft="$4, "hard="$5}' /proc/$(cat /tmp/uwsgi.pid)/limits

# Verify a worker process limit
WORKER_PID=$(pgrep -P $(cat /tmp/uwsgi.pid) | head -1)
awk '/^Max open files/ {print "soft="$4, "hard="$5}' /proc/$WORKER_PID/limits

Workers inherit the master’s limits at fork time. If the master was started with the correct limit, workers should match.

If the limit did not change, check:

  • systemd: confirm the override is active. Run systemctl cat uwsgi to see the merged configuration including overrides.
  • shell: confirm ulimit -n was run in the same session or script that starts uWSGI.
  • limits.conf: confirm the PAM user matches the uWSGI service user.

The headroom rule

Keep fd usage below 50-80% of the soft limit under normal operating conditions. This provides margin for:

  • Traffic bursts that temporarily increase connection counts
  • Application code that opens additional fds during specific operations (file processing, bulk database transactions)
  • Slowly developing fd leaks that need time to detect before they cause incidents

The degradation curve for fd exhaustion is a cliff. Below the limit, everything works. At the limit, new accept() and open() calls fail with EMFILE. There is no gradual degradation and no uWSGI-level signal.

Sizing a target limit

Estimate peak fd usage per worker:

peak_fds_per_worker = max_concurrent_connections + database_pool_size + logging_fds + application_fds + overhead

Then apply headroom:

target_limit = peak_fds_per_worker * 1.5

The 1.5x multiplier provides 33% headroom above peak. For conservative deployments, use 2x.

Deployment profileSuggested limitRationale
Small (2-4 workers, behind nginx)65536Comfortable headroom, low overhead
Medium (8-16 workers, direct or proxied)65536Sufficient for most workloads
High-connection (async/gevent, many connections per worker)131072 or higherAsync workers multiplex many connections

Detecting fd leaks

A steadily climbing per-worker fd count that never stabilizes or decreases indicates a leak. Common causes:

  • mmap’d files not closed. The fd is opened, the file is mapped, but the fd is never closed after mmap returns. VSZ grows alongside fd count in this pattern.
  • Database connections not returned to the pool. Each leaked connection holds an fd.
  • Temp files opened but not closed. Common in file upload processing or report generation.
  • Socket connections without timeouts. Sockets that never close accumulate indefinitely.

Track the correlation between fd count and virtual memory size:

# Compare fd count and VSZ per worker over time
for pid in $(pgrep -P $(cat /tmp/uwsgi.pid)); do
    fds=$(ls /proc/$pid/fd 2>/dev/null | wc -l)
    vsz=$(awk '/^VmSize/ {print $2}' /proc/$pid/status)
    echo "$(date +%s) pid=$pid fds=$fds vsz_kb=$vsz"
done

VSZ growing alongside fd count is a strong indicator of mmap-related leaks. VSZ stable with growing fd count points to socket or pipe leaks. Either pattern warrants investigation before the limit is reached.

Common pitfalls

systemd LimitNOFILE overrides shell ulimit

The most common mistake. You run ulimit -n 65536 in your shell, restart the service with systemctl restart uwsgi, and assume the new limit is in effect. It is not. systemd applies LimitNOFILE from the unit file or override, ignoring the shell environment entirely. Always verify via /proc/<pid>/limits.

setrlimit after privilege drop

If uWSGI is configured with --uid and --gid to drop privileges after startup, and the process attempts to raise its fd limit after the drop, the setrlimit call fails with Operation not permitted. Raising the hard limit requires elevated privileges.

The fix is to ensure the limit is set before privilege drop occurs. uWSGI itself applies --max-fd during initialization, before --uid/--gid take effect, so an instance started as root can raise the limit and then drop privileges. Once the process is running as an unprivileged user, raising the hard limit fails with Operation not permitted. systemd handles this correctly because LimitNOFILE is applied before exec. For Emperor mode, set the limit on the Emperor process (which runs as root) so vassals inherit it.

High OS limits causing internal allocation overhead

Setting an extremely high RLIMIT_NOFILE can cause uWSGI to allocate oversized internal data structures, because several subsystems size tables from uwsgi.max_fd (derived from the soft limit): the poll/event queue allocates a struct pollfd array and a pointer array of max_fd entries, and the async subsystem allocates two struct wsgi_request * arrays of max_fd entries each. On a container with a limit in the billions, that alone reaches multiple gigabytes. Use --max-fd to cap uWSGI’s internal allocation independently of the OS limit.

Not accounting for all fd consumers

Database connection pools, Redis connections, log files, the stats server socket, and application-opened files all consume fds. Counting only incoming connections underestimates actual usage. Audit all fd consumers before sizing the limit. A worker with 50 concurrent connections but a 20-connection database pool and 5 log files needs at least 76 fds, not 50.

Signals to monitor

SignalWhy it mattersWarning sign
Per-worker fd count (/proc/<pid>/fd)Primary utilization metricSteady upward trend indicates a leak
fd count vs. soft limit ratioHeadroom indicatorAbove 50-80% under normal load
Worker VSZCorrelates with mmap’d fd leaksVSZ growing in lockstep with fd count
Worker RSSMemory pressure from leaked resourcesSteady growth without recycling
System-wide fd usage (/proc/sys/fs/file-nr)Host-wide fd consumptionApproaching fs.file-max

How Netdata helps

  • Per-process fd counts. Netdata collects open fd counts per process group, replacing manual /proc/<pid>/fd polling with continuous collection.
  • VSZ/RSS correlation. Memory metrics appear alongside fd counts. VSZ rising in lockstep with fd count signals an mmap-related leak.
  • Anomaly detection on fd trends. Netdata’s anomaly detection flags unusual fd growth patterns, catching acceleration before exhaustion.
  • Per-second resolution. fd leaks can accelerate nonlinearly. Per-second collection catches the inflection point where growth rate changes.
  • System-wide fd context. Netdata tracks /proc/sys/fs/file-nr alongside per-process metrics, so you can distinguish a uWSGI-specific problem from host-wide fd pressure.