The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-file-descriptor-exhaustion ▌

Operations Guides

Tomcat file descriptor usage: OpenFileDescriptorCount vs the ulimit

Tomcat’s JVM holds file descriptors for every open socket, every log file, every JAR on the classpath, the NIO selector itself, and various internal pipes. The pool is finite and process-scoped. When it runs out, Tomcat stops accepting connections and stops writing to logs in the same instant. The error is java.net.SocketException: Too many open files, and by the time you see it the failure is already total.

Operators have two views into this resource. JMX exposes OpenFileDescriptorCount and MaxFileDescriptorCount from the java.lang:type=OperatingSystem MBean. The OS exposes the same count through /proc/<pid>/fd and the limit through ulimit -n and /proc/<pid>/limits. The two views should agree within one descriptor. When they do not, you are usually looking at the wrong process, the wrong container namespace, or at lsof output (which counts more than file descriptors).

What it is and why it matters

File descriptors are the kernel’s handle for open files, sockets, pipes, and on Linux a long list of other resources. For a Tomcat process the dominant consumers are:

ConsumerTypical countNotes
HTTP/HTTPS/AJP socketsone per accepted connectionincludes idle keepalive connections
NIO selector and registered socketsa handful, plus one per registered socketthe poller’s epoll FD
Access and application logsone or two per active log filegrows if rotation leaves handles open
Classpath JARshundreds on a large Spring classpathbaseline cost, not a leak
Internal pipesa fewbetween JVM subsystems

The limit is hit suddenly. There is no graceful degradation, no queueing, no backpressure. The next accept(), open(), or socket() call returns EMFILE, and Java surfaces it as SocketException: Too many open files. Accepting new HTTP connections fails. Writing to the access log fails. Opening a freshly rotated log file fails. All of these happen in the same instant.

The operational signal is the ratio OpenFileDescriptorCount / MaxFileDescriptorCount, and the trajectory of OpenFileDescriptorCount relative to connection count. The playbook’s headroom rule: keep peak FD usage under 50% of the limit, and size production Tomcat with an ulimit of at least 65535. Default distro ulimits of 1024 or 4096 are routinely too low and fail under modest keepalive load.

How it works

flowchart TD
  subgraph C["FD consumers in the Tomcat JVM"]
    Sockets["Sockets"]
    NIO["NIO selector"]
    Logs["Logs"]
    JARs["Classpath JARs"]
    Pipes["Pipes"]
  end

  C --> Pool["Open file descriptors"]

  Pool -->|JMX view| JMXView["OpenFileDescriptorCount
java.lang:type=OperatingSystem"] Pool -->|OS view| ProcView["/proc/pid/fd"] JMXView -.->|agree within 1| ProcView Pool -->|bounded by| Limits["Limit chain:
soft ulimit,
fs.nr_open,
fs.file-max"]

The JMX view

The java.lang:type=OperatingSystem platform MBean exposes two long attributes:

  • OpenFileDescriptorCount - the current count of file descriptors held by the JVM process.
  • MaxFileDescriptorCount - the process’s current soft RLIMIT_NOFILE ceiling, not the hard ceiling.

On Linux, OpenFileDescriptorCount is computed by listing /proc/self/fd: the OpenJDK implementation opens the directory with opendir(), counts the entries, and subtracts the descriptor it opened for the count itself. The behavior is the same from JDK 8 through JDK 21, so JMX and ls /proc/<pid>/fd | wc -l agree.

MaxFileDescriptorCount reflects the soft limit at the time of the query. If the JVM or a wrapper script raised the soft limit below the hard limit, JMX reports the raised value. That is the limit the process actually operates under. Do not confuse it with the hard ceiling.

The OS view

ls /proc/<pid>/fd | wc -l is the authoritative count. cat /proc/<pid>/limits | grep "Max open files" shows the soft and hard RLIMIT_NOFILE for that process.

Avoid lsof for FD budget accounting. lsof lists more than file descriptors: it includes memory-mapped files, the current working directory, the root directory, executable text segments, and other entries that do not consume FD slots. lsof | wc -l systematically overcounts relative to /proc/<pid>/fd. For pure FD budget work, /proc/<pid>/fd is the truth and JMX mirrors it.

To identify what is consuming descriptors during a suspected leak:

# Count FDs grouped by target
ls -l /proc/<pid>/fd | awk '{print $NF}' | sed 's/\[.*\]//' | sort | uniq -c | sort -rn | head -20

# Show socket states for FDs that are sockets
ls -l /proc/<pid>/fd | grep socket | wc -l
ss -tanp | grep <pid> | awk '{print $1}' | sort | uniq -c

The first command groups FDs by what they point to. A large count against a single log file or a growing count of socket:[nnnn] entries narrows the leak source. The second cross-references sockets against TCP state; a rising CLOSE_WAIT count is the dominant leak pattern.

The limit chain on Linux

Three limits interact, and the smallest one wins:

  1. Per-process soft limit (ulimit -n, RLIMIT_NOFILE cur) - the value the process operates under and can raise up to the hard limit. This is what MaxFileDescriptorCount reports.
  2. Per-process hard limit (RLIMIT_NOFILE max, bounded by fs.nr_open) - the ceiling the process can raise to; raising the hard limit requires CAP_SYS_RESOURCE, and going above fs.nr_open first requires raising that sysctl.
  3. System-wide limit (/proc/sys/fs/file-max) - the total FDs the kernel will allocate across all processes.

For Tomcat the per-process soft limit is almost always the binding constraint. fs.file-max defaults to a large value on modern kernels and is rarely the bottleneck for a single JVM. The binding constraint shifts to fs.file-max only when many JVMs share a host and each has a generous per-process limit.

In containers the limit comes from the runtime

In containers the FD limit comes from the container runtime, not the host. The host’s ulimit -n is not what the container process sees.

For Docker and containerd, the runtime’s own systemd unit sets a very high LimitNOFILE, and containers inherit it unless overridden - dockerd ships LimitNOFILE=1048576; containerd’s unit varies by distro (some use infinity). This is generous, sometimes problematically so: tools that iterate over all possible FDs on fork/exec (Python’s subprocess, rpm) can slow down noticeably when the limit is six figures. Python’s subprocess closes descriptors up to the soft RLIMIT_NOFILE when close_fds applies; modern CPython batches that with the close_range() syscall, but the range still scales with the limit. To narrow the limit per container, set default-ulimits in /etc/docker/daemon.json or pass --ulimit nofile=65536:65536 at run time.

For Kubernetes there is no native pod spec field for FD limits. The container runtime’s defaults apply. To change them, CRI-O exposes default_ulimits in crio.conf; containerd’s CRI has no config.toml ulimit setting, so raise LimitNOFILE on the containerd service (inherited by pods) or set per-container limits via the runtime spec. cgroup v2 has no file descriptor controller: pids.max caps processes and threads, not FDs, so you cannot constrain FDs through the cgroup layer.

Where it shows up in production

Default ulimit too low

The most common failure is shipping Tomcat with the distro default of 1024 or 4096. With NIO, every keepalive connection consumes a descriptor, the selector consumes a few, the classpath JARs consume hundreds, and the access log consumes a couple. Under moderate keepalive load a 1024 limit is exhausted within minutes. The fix is to set the ulimit at the layer that owns the process: a systemd unit LimitNOFILE=65535, an init script ulimit -n, or the container runtime’s defaults.

For systemd-managed Tomcat, verify after restart:

cat /proc/<pid>/limits | grep "Max open files"

FD growth without connection growth

The playbook’s leak heuristic: FD count growing without proportional connection count growth is an FD leak. Sockets or files are being opened and never closed.

The dominant Tomcat pattern is CLOSE_WAIT sockets. The remote peer sends FIN, the kernel moves the socket to CLOSE_WAIT, and the application never reads EOF and never calls close(). The socket stays in CLOSE_WAIT indefinitely, consuming an FD. Other recurring patterns: log file rotation that leaves the old handle open, JDBC connections checked out but never returned, and HTTP client responses whose bodies are not fully consumed before the connection is reused or closed.

Distinguishing signal: OpenFileDescriptorCount rises monotonically while connectionCount (the JMX metric on Catalina:type=ThreadPool,name="http-nio-8080") stays flat or oscillates with traffic. If both rise together, the FD growth is explained by real connections; the fix is capacity, not a leak hunt.

Container FD surprises

Two container-specific surprises recur:

  • The container sees a much higher ulimit -n (1048576) than the host. Operators comparing the host’s 1024 to the container’s reading misdiagnose a configuration change. The container reading is correct and comes from the runtime.
  • The container hits the host’s fs.file-max. Because the runtime’s per-container limit is so generous, many containers on one host can collectively exhaust the system-wide pool. Check with cat /proc/sys/fs/file-nr; the first column is allocated, the third is the max. This is rare but worth checking when MaxFileDescriptorCount is high but open() still fails with EMFILE.

NIO and classpath descriptors are non-obvious consumers

NIO uses descriptors for the selector itself, plus one per registered socket. A Tomcat instance with maxConnections=8192 is structurally capable of holding 8192 connection descriptors plus the selector FDs plus the classpath. The JVM also holds descriptors for JARs on the classpath; on a large Spring deployment this can be several hundred. These are baseline cost, not leaks. They compress headroom under a low ulimit and explain why a Tomcat that “isn’t doing much” still reports a non-trivial OpenFileDescriptorCount at idle.

Common misuses

  • Alerting on absolute FD count. A count of 5000 is fine on a host with ulimit -n 65535 and fatal on a host with ulimit -n 4096. Alert on the ratio OpenFileDescriptorCount / MaxFileDescriptorCount, not the raw count.
  • Using lsof | wc -l as the FD count. It overcounts. Use /proc/<pid>/fd or JMX.
  • Reading the host’s ulimit -n and assuming the container sees the same value. It does not. Read MaxFileDescriptorCount from inside the JVM, or cat /proc/<pid>/limits for the container’s actual limit.
  • Treating MaxFileDescriptorCount as the hard ceiling. It is the soft limit. If you raised the soft limit at JVM startup, JMX reports the raised value, which is correct, but the hard ceiling is RLIMIT_NOFILE max (bounded by fs.nr_open).
  • Assuming FD growth is always a leak. Rule out legitimate growth from connection count, log file count, or classpath changes first. Correlate with connectionCount.

Signals to watch in production

SignalWhy it mattersWarning sign
OpenFileDescriptorCount / MaxFileDescriptorCount ratioPrimary saturation signalSustained above 70%, or any value above 90%
OpenFileDescriptorCount growth rateLeak indicatorMonotonic growth not matched by connectionCount
connectionCount on Catalina:type=ThreadPoolExplains legitimate FD consumptionFDs rising without this rising means suspect a leak
java.net.SocketException: Too many open files in logsThe failure signatureEven one occurrence means you hit the cliff
Access log write failuresFD exhaustion corollaryErrors writing to the access log coincide with socket errors
CLOSE_WAIT socket countDominant leak patternss -tan state close-wait count rising

How Netdata helps

Netdata’s Java and Tomcat collectors surface OpenFileDescriptorCount and MaxFileDescriptorCount at per-second resolution alongside connectionCount, thread pool, and GC metrics. The value is correlation across these signals, not any single number:

  • Plot OpenFileDescriptorCount against connectionCount from the same connector. A leak shows up as FD count climbing while connection count oscillates with traffic.
  • Plot the FD ratio against currentThreadsBusy and accept queue depth. A simultaneous rise across all three points to a connection-level event; an isolated FD rise points to a leak.
  • ML anomaly detection flags FD growth that deviates from the instance’s normal pattern, which catches slow leaks before the 70% threshold trips.
  • Per-second resolution matters because FD exhaustion is a cliff. The warning window between 70% and 100% can be minutes long, and minute-level polling can step over it entirely.
  • OS-level /proc collectors provide independent corroboration of the JMX count, so you can verify the two views agree rather than trusting one instrument.
  • Container-aware collection means the FD ratio reflects the runtime’s actual limit, not the host’s, which removes the most common false positive in containerized Tomcat.