The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-maxthreads-tuning ▌

Operations Guides

Tomcat maxThreads and minSpareThreads: sizing the executor correctly

The <Executor> thread pool, or the connector-level pool when no explicit executor is referenced, is the most critical bounded resource in a Tomcat instance. maxThreads sets the ceiling on concurrent request processing. minSpareThreads sets the floor of pre-warmed threads. Every active HTTP request occupies one worker thread for its entire processing duration, unless you use async servlets. When the pool is full, connections keep arriving but requests wait.

In a typical servlet workload, threads spend most of their lifetime blocked on I/O: database queries, downstream HTTP calls, cache lookups. A pool sized to CPU count is almost always too small. A pool sized to peak concurrent requests times worst-case processing time is closer to correct, but the queueing behavior between connections, the poller, and worker threads makes the accounting subtle.

What it is and why it matters

maxThreads and minSpareThreads are attributes of the <Executor> element in server.xml, or of the <Connector> when no shared executor is referenced. They define the worker thread pool that processes HTTP requests after the connector accepts the TCP connection and the NIO poller detects incoming data.

Default values per the Tomcat reference configuration: maxThreads=200 with minSpareThreads=10 on the connector-level pool, and maxThreads=200 with minSpareThreads=25 on a shared <Executor> element. The two defaults differ, so check which one applies to your setup.

The defaults work for many small deployments and are catastrophically wrong for others. Below roughly 80% utilization, throughput scales linearly with load. Above that, queue time grows non-linearly. At 100%, throughput is pinned to the rate at which threads free up, and the accept queue begins to fill.

The pool interacts with two other bounded resources:

  1. Connections (maxConnections, default 8192 for NIO): the poller multiplexes up to this many TCP connections. Idle keepalive connections hold a socket and a buffer, not a thread.
  2. Accept queue (acceptCount, default 100): the OS-level TCP backlog. When both the poller and the accept queue are full, the kernel rejects connections with RST.

Which layer is saturated determines whether you need more threads, more connections, or a faster backend.

How it works

Thread creation is lazy

Tomcat does not pre-allocate maxThreads threads at startup. It starts with minSpareThreads and creates more on demand as requests arrive, up to maxThreads. This is why currentThreadCount, the number of live threads in the pool, is almost always less than maxThreads in steady state. The first burst of traffic after startup may temporarily saturate the minSpareThreads pool before additional threads are created. Gate any thread-pool alerting on uptime greater than 120 seconds to avoid false positives during cold start.

One request, one thread (unless async)

In the default synchronous model, a worker thread is bound to a request from the moment the poller dispatches it until the response is committed. If your application blocks on a database query for 500ms, that thread is occupied for 500ms. Multiply by concurrent requests and you get the pool pressure.

Async servlets (AsyncContext) release the thread during async processing. The request stays in flight but currentThreadsBusy undercounts the actual number of concurrent in-flight requests. If your application makes heavy use of async, thread-pool utilization becomes a misleading capacity signal. Connection-based metrics become more meaningful.

After GC, a brief busy spike

When the JVM pauses for garbage collection, all worker threads are frozen. When the pause ends, they resume and report as busy simultaneously. This produces a transient spike in currentThreadsBusy that does not reflect real load. Correlate with GC pause time before interpreting a brief saturation event as capacity exhaustion.

The three JMX signals

The MBean Catalina:type=ThreadPool,name="http-nio-8080" exposes three attributes that define pool health:

AttributeMeaning
maxThreadsConfigured ceiling.
currentThreadCountLive threads in the pool. Grows on demand from minSpareThreads toward maxThreads.
currentThreadsBusyThreads currently processing a request.

The ratio that matters for capacity decisions is currentThreadsBusy / maxThreads. currentThreadCount tells you how much of the pool has been provisioned, not how much is under pressure.

flowchart LR
  Client["Client"] -->|TCP SYN| AcceptQ["Accept queue
acceptCount = 100"] AcceptQ --> Poller["NIO poller
maxConnections = 8192"] Poller -->|socket readable| Workers["Worker pool
maxThreads = 200"] Workers -->|busy reaches maxThreads| Wait["Request waits
for a free thread"] Wait --> Workers

Connections are cheap in NIO; threads are expensive. Saturation propagates backward through the layers: worker threads fill first, then the poller reaches maxConnections, then the accept queue fills, then the kernel sends RST.

Where it shows up in production

Sizing against processing time, not CPU

The most common mistake is sizing the pool by CPU count. In an I/O-bound servlet workload, threads spend most of their time waiting. A database call that takes 200ms occupies a thread for 200ms regardless of CPU. The correct sizing inputs are peak concurrent request rate and worst-case processing time under normal backend conditions.

Little’s Law gives the steady-state approximation: concurrent threads needed = request rate times average processing time.

# Little's Law demand estimate
# ConcurrentThreadsNeeded = RequestRate (req/s) * AvgProcessingTime (s)
# Example: 1000 * 0.05 = 50 threads at steady state

At 1000 req/s with 50ms average processing time, you need roughly 50 threads. That covers steady state, not bursts or tail latency. If the Little’s Law product exceeds maxThreads under sustained load, the pool will saturate. The time to saturation depends on how fast demand grows, which requires trend observation, not a static formula.

The operational headroom rule: keep peak currentThreadsBusy / maxThreads below 0.60 during the busiest hour. This leaves room for traffic spikes, slow requests, and the non-linear queueing that begins above 80%.

The cliff above 80%

Thread pools degrade gracefully below 80% utilization. Between 80% and 100%, queue time increases non-linearly. At 100%, throughput drops to the rate at which threads free up, and every additional arriving request extends the queue. From the client perspective, latency explodes even though the JVM is healthy and CPU may be low. This is the classic signature of I/O-bound exhaustion: threads busy, CPU low, throughput dropping.

Client retries add load, accelerating the cliff. Under sustained overload, each retried request consumes a thread that a legitimate request needed.

Shared executors across connectors

When multiple connectors (HTTP, HTTPS, AJP) reference the same <Executor>, the maxThreads value is shared across all of them. A single maxThreads=200 pool serving both HTTP and AJP means 200 total, not 200 each. If one connector dominates traffic, the other can starve. Monitor per-connector currentThreadsBusy and sum them against the shared maxThreads.

Spring Boot embedded Tomcat

In Spring Boot, the executor is internal. The relevant properties are server.tomcat.threads.max and server.tomcat.threads.min-spare (since Spring Boot 2.3, renamed from server.tomcat.max-threads and server.tomcat.min-spare-threads).

Two version-specific gotchas: Spring Boot 3.3.0 had a regression where the tomcat.threads.config.max metric always returned -1 (spring-projects/spring-boot#40957), fixed in 3.3.1. Spring Boot 3.3.0 also failed to start when server.tomcat.threads.max was set below the default server.tomcat.threads.min-spare of 10, because that release stopped letting Tomcat create its own executor; the fix landed in 3.3.1 and the workaround is to set min-spare below max. If you set a small pool for a low-traffic service and the application fails to start, check the version and the min-spare/max relationship.

Virtual threads change the model entirely

With JDK 21+ and Tomcat’s StandardVirtualThreadExecutor (available since Tomcat 10.1.10), one virtual thread is created per task. maxThreads and minSpareThreads are not meaningful. currentThreadsBusy reports -1. The bounded resource shifts from threads to connections. In this mode, connectionCount - keepAliveCount is the proxy for active concurrent work. Do not apply the traditional 0.6 headroom rule to a virtual-thread executor.

Tradeoffs and when to use it

Large maxThreads: more capacity, more memory

Every thread reserves a stack. The default -Xss is typically 512KB to 1MB. A pool of 500 threads with 1MB stacks reserves 500MB of address space; actual committed memory grows as stacks are used. Verify the stack size before extrapolating:

# Check JVM default thread stack size
java -XX:+PrintFlagsFinal -version 2>&1 | grep ThreadStackSize

Large minSpareThreads: faster burst response, permanent overhead

minSpareThreads is the floor of threads always kept alive. Raising it pre-warms the pool so the first burst does not pay thread-creation latency. Those threads and their stacks exist permanently, whether or not traffic arrives. For a service with spiky traffic where the first 50ms of a burst matters, raising minSpareThreads is a reasonable trade. For a steady-state service, the default is fine.

maxQueueSize: the silent trap

The executor’s maxQueueSize defaults to Integer.MAX_VALUE. This means the executor queues tasks indefinitely before rejecting them. If you raise maxThreads but leave maxQueueSize at default, sustained overload does not produce fast failures. It produces unbounded queue growth, rising memory pressure, and eventually OOM. For a fail-fast posture, set maxQueueSize to a finite value so overload produces a rejection quickly rather than a slow death.

When not to increase maxThreads

If currentThreadsBusy sits at maxThreads because threads are stuck on a slow backend, adding threads treats the symptom and accelerates the underlying problem. More threads means more concurrent backend calls, which can overwhelm the backend further. The correct response to backend-driven exhaustion is to fix the backend, add timeouts on outbound calls, or shed load at the load balancer. Take a thread dump to confirm whether threads are genuinely processing or blocked waiting:

# Take a thread dump to see what worker threads are doing
# For Spring Boot, replace the process selector with your jar name or main class
jstack $(pgrep -f 'catalina.startup.Bootstrap' | head -1) | grep -A 5 "http-nio-8080-exec"

If the stack traces show dozens of threads parked on a socket read or a database connection acquire, the problem is downstream. Adding threads will not help.

Signals to watch in production

SignalWhy it mattersWarning sign
currentThreadsBusy / maxThreadsPrimary capacity ratio. The single most important Tomcat signal.Sustained above 0.60 at peak. Above 0.80 for 5+ minutes. At 1.0 for 2+ minutes.
currentThreadCountHow much of the pool has been provisioned.Persistently equals maxThreads means the pool is at its ceiling and cannot grow further.
requestCount rate (throughput)Confirms whether busy threads are processing or stuck.Threads at max but throughput dropping means threads are blocked, not working.
Average request processing timeThreads held longer fill the pool faster.Upward trend in processing time at constant load.
HTTP 503 rateWith a bounded maxQueueSize, a full executor rejects the task and Tomcat closes the connection without an HTTP response; it does not send 503 for capacity. A 503 in the access log means application code or a proxy produced it.Any sustained 503 burst from the application.
GC pause timeGC freezes threads, causing a false busy spike on resume.Busy spike that coincides with a GC event.
JVM CPU utilizationDistinguishes I/O-bound exhaustion from CPU-bound.Threads busy plus CPU low means blocked on I/O. Threads busy plus CPU high means compute-bound or GC.
Accept queue depthThe last buffer before connection refusal. Invisible to JMX.Non-zero Recv-Q on the listening socket, sustained.
# Check accept queue depth (Recv-Q column on the listening socket)
ss -tnl | grep 8080

How Netdata helps

Netdata turns executor sizing from guesswork into measurement.

  • Per-second thread pool metrics: currentThreadsBusy, currentThreadCount, and maxThreads from the Catalina:type=ThreadPool MBean, collected every second so transient spikes are visible rather than averaged away.
  • Capacity ratio at per-second resolution: the currentThreadsBusy / maxThreads ratio reveals whether the 0.60 headroom rule holds during the busiest minute, not just the busiest hour.
  • GC correlation: when a busy-thread spike coincides with a GC pause, the correlated timeline shows the spike is a GC artifact, not real load. This prevents over-sizing the pool in response to false positives.
  • Throughput alongside busy threads: plotting requestCount rate next to currentThreadsBusy distinguishes a pool full of working threads from a pool full of stuck threads. If throughput drops while busy stays at max, the problem is downstream, not the pool size.
  • CPU context: I/O-bound exhaustion (threads busy, CPU low) looks completely different from CPU-bound saturation (threads busy, CPU high). Seeing both on one timeline prevents adding threads to a CPU-bound workload.
  • Uptime-gated alerting: thread-pool alerts can be gated on uptime greater than 120 seconds, eliminating cold-start false positives where the first burst saturates minSpareThreads before the pool grows.