The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-http-503-service-unavailable ▌

Operations Guides

Tomcat HTTP Status 503 Service Unavailable: the connector is out of threads

A 503 in a Tomcat topology means something between the user and the servlet gave up. The instinct is to blame the connector thread pool, and that is often the right place to look, but the mechanism is more layered than it appears. Treating “503” and “out of threads” as the same condition leads to misdiagnosis.

With the default NIO connector, Tomcat decouples worker threads from TCP connections. A request thread pool can be saturated while the connector keeps accepting connections, and the connection pool can be saturated while the JVM still passes a TCP health check. The path from a full thread pool to a visible 503 passes through two more buffers before any client sees a failure.

What this means

Key mechanics from the Tomcat connector model (covered in depth in the mental model hub):

  • NIO is the default since Tomcat 8.5. A small set of poller threads multiplexes connections; request processing is dispatched to a bounded worker pool. Default sizing: maxThreads=200, minSpareThreads=10, maxConnections=8192, acceptCount=100, connectionTimeout=60000ms.
  • Worker threads and connections are separate limits. Every active request occupies one worker thread for its full duration. Idle keepalive connections occupy a poller slot and a file descriptor, but not a thread.
  • Saturation cascades through two queues. When all worker threads are busy, new connections are still accepted up to maxConnections. When maxConnections is reached, additional connections land in the OS accept queue (bounded by acceptCount). When the accept queue fills, the kernel rejects new connections with RST.

The crucial implication: when the worker thread pool is exhausted, the NIO connector does not synthesize an HTTP 503 response. The failure surfaces at the TCP layer as connection refused or, upstream of that, as a timeout. The 503 that users see is generated by an upstream load balancer or reverse proxy translating the failed backend connection into a 5xx response, or by the application itself returning 503 from servlet code. Tomcat itself only synthesizes a 503 while the connector is paused (for example during a graceful stop); thread-pool exhaustion surfaces as queuing and, at the limit, TCP-level refusal rather than an HTTP 503. The dominant pattern in production is proxy-synthesized 503 triggered by Tomcat refusing or timing out on the backend connection.

This distinction matters because the fix differs. If the 503 is proxy-synthesized from a saturated thread pool, raising maxThreads or fixing the slow backend resolves it. If the 503 is application-emitted, the thread pool may be fine and the problem is in servlet code. If the 503 is from connection exhaustion (poller full, not thread pool full), adding threads does nothing.

Common causes

CauseWhat it looks likeFirst thing to check
Worker thread exhaustioncurrentThreadsBusy == maxThreads, CPU low, throughput collapsed, proxy returns 502/503/504jstack to see what threads are waiting on
Connection (poller) saturationconnectionCount near maxConnections, accept queue depth non-zero, thread pool may still have headroomss -tnl 'sport = :8080' Recv-Q
Application-emitted 503Access log shows 503 from a specific servlet path, thread pool and connections both have headroomGrep application code for sendError(503) or SC_SERVICE_UNAVAILABLE
Slow backend holding threadsThread dump shows many threads in WAITING on socket read or DB connection acquire, processing time climbingCorrelate thread dump stack traces with backend latency
Stuck threads draining the poolcurrentThreadsBusy sticks at maxThreads and does not drop, StuckThreadDetectionValve (if configured) reports stuck countjstack, look for identical BLOCKED/WAITING stack traces
Connection leak or slowloris-style clientconnectionCount high with low request rate, low bytes receivedss per-source-IP connection counts

Quick checks

Safe read-only commands. Run these on the Tomcat host. JMX commands assume a JMX remote endpoint (for example, -Dcom.sun.management.jmxremote.port=9090) is configured.

# Worker thread pool state via JMX
java -jar jmxterm.jar -l localhost:9090 -n -v silent -e \
  "get -b Catalina:type=ThreadPool,name=\"http-nio-8080\" currentThreadsBusy currentThreadCount maxThreads connectionCount maxConnections"

# Accept queue depth (Recv-Q is current backlog, Send-Q is the configured max)
ss -tnl 'sport = :8080'

# Error count from JMX (lumps 4xx + 5xx, cannot isolate 503)
java -jar jmxterm.jar -l localhost:9090 -n -v silent -e \
  "get -b Catalina:type=GlobalRequestProcessor,name=\"http-nio-8080\" errorCount requestCount"

# 5xx rate from the access log (assumes default field positions, $9 is status)
awk '$9 ~ /^5/' /var/log/tomcat/localhost_access_log.$(date +%Y-%m-%d).txt | wc -l

# 503 specifically
awk '$9 == 503' /var/log/tomcat/localhost_access_log.$(date +%Y-%m-%d).txt | wc -l

# Thread dump: count http-nio worker threads and inspect their state
# If multiple Tomcat JVMs run on the host, narrow the pgrep pattern or use the specific PID
jstack "$(pgrep -f 'catalina.startup.Bootstrap' | head -1)" | grep -c "http-nio-8080-exec"

# Manager status XML (if Manager app is enabled and reachable)
curl -s -u "$USER:$PASS" 'http://localhost:8080/manager/status?XML=true' | \
  grep -oP '(maxThreads|currentThreadsBusy|currentThreadCount|connectionCount|maxConnections)="[0-9]+"'

How to diagnose it

  1. Confirm where the 503 originates. Parse the access log on the Tomcat host. If Tomcat’s access log records the request with status 503, Tomcat (or the application) generated it. If the access log shows no matching request at all, the 503 was synthesized upstream and the Tomcat-side symptom is a refused or timed-out connection.

  2. Check thread pool saturation. Read currentThreadsBusy and maxThreads. If busy is at or near max, the pool is the constraint. With virtual threads (JDK 21+), currentThreadsBusy reports -1 and maxThreads is not meaningful; the connector’s ThreadPool MBean still exposes connectionCount and keepAliveCount, so monitor connectionCount (minus keepAliveCount) as a proxy instead.

  3. Check connection saturation. Read connectionCount against maxConnections. If connections are at the limit while threads still have headroom, the poller is the bottleneck, not the worker pool. Adding maxThreads will not help.

  4. Check the accept queue. Run ss -tnl 'sport = :8080'. A non-zero Recv-Q means connections are backing up in the OS backlog. If Recv-Q equals Send-Q (the acceptCount), the kernel is refusing connections. This is the layer that produces the TCP RST an upstream proxy turns into a 503.

  5. Take a thread dump. jstack shows exactly what each worker thread is doing. Many threads in the same WAITING or BLOCKED state, on the same socket read or connection acquire, points at the slow backend.

  6. Correlate CPU. If currentThreadsBusy == maxThreads and CPU is low, threads are blocked on I/O (database, downstream HTTP, DNS). If CPU is high, the application is compute-bound or GC is thrashing. Low CPU with full threads is the classic thread-exhaustion signature.

The cascade below shows where each layer fails and what surfaces to the client:

flowchart TD
  A[Inbound request] --> B{Free worker thread?}
  B -- yes --> C[Dispatch and process]
  B -- no --> D{connectionCount less than maxConnections?}
  D -- yes --> E[Queued, awaiting worker thread]
  D -- no --> F{accept queue less than acceptCount?}
  F -- yes --> G[Queue in OS backlog]
  F -- no --> H[Kernel sends RST]
  H --> I[Proxy synthesizes 502/503/504]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
currentThreadsBusy / maxThreadsPrimary capacity ratio. At 1.0, requests queue.Sustained >0.80, or == 1.0 for >120s
connectionCount / maxConnectionsPoller saturation. Independent of thread pool.Sustained >0.80 with threads available
Accept queue Recv-Q (ss -tnl)OS backlog depth. Invisible to JMX.Any sustained non-zero value; == Send-Q means RST
Request throughput (requestCount delta)Throughput collapsing while traffic arrives means requests are not being processed.Drop >50% from baseline during expected traffic
Request processing time (processingTime / requestCount)Threads held longer fill the pool faster.Average trending up >2x baseline
Access log 5xx rateThe only way to isolate 503 from 4xx. JMX errorCount cannot.5xx count rising, especially 503 specifically
CPU utilizationLow CPU with full threads means blocked on I/O. High CPU means compute or GC bound.Either extreme combined with thread saturation

Fixes

Thread pool exhaustion from a slow backend

The durable fix is on the backend, not in Tomcat. Use the thread dump to identify the dependency, then add or tighten timeouts on the outbound call (JDBC query timeout, HTTP client socket timeout, DNS lookup timeout). Raising maxThreads buys headroom but does not solve the underlying problem; if the backend is slow, more threads just means more threads waiting. If you do raise maxThreads, watch file descriptor usage and connectionCount, since each thread and connection consumes an FD.

Thread pool exhaustion from insufficient sizing

If the thread dump shows threads processing quickly (not blocked) and currentThreadsBusy is still pegged at maxThreads under peak load, the pool is undersized for the workload. Increase maxThreads and verify the backend (especially the database connection pool) can absorb the additional concurrent load. A common mistake is raising Tomcat maxThreads past the database pool maxActive, which just moves the bottleneck downstream.

Connection (poller) saturation

If connectionCount is at maxConnections but threads have headroom, the issue is too many open connections, typically from upstream keepalive pools or a slowloris-style pattern. Tune keepAliveTimeout down from the 60s default (it defaults to connectionTimeout if unset), coordinate keepalive timers with the reverse proxy so the proxy does not reuse connections Tomcat has already closed (for example, nginx holding connections 75s while Tomcat closes at 60s causes resets), and investigate per-source connection counts with ss.

Accept queue overflow

If Recv-Q is at Send-Q, the OS is refusing connections. Increase acceptCount in server.xml, but verify net.core.somaxconn on the host, since the kernel caps the actual backlog at somaxconn regardless of what Tomcat requests. Also investigate why the acceptor thread is not keeping up; a stop-the-world GC pause blocks the acceptor along with every other thread.

Stuck threads

If jstack shows identical BLOCKED or WAITING stack traces across many worker threads, those threads are not coming back without intervention. Configure StuckThreadDetectionValve (not enabled by default; default threshold is 600s, consider 60s for user-facing services) for visibility. The fix is the blocking call: add a timeout or break a deadlock. Restarting Tomcat clears the threads but they will re-accumulate if the root cause remains.

Application-emitted 503

If Tomcat’s access log shows the 503 and the thread pool and connection metrics are healthy, the 503 is coming from servlet code. Grep the codebase for sendError(503), SC_SERVICE_UNAVAILABLE, or framework-level “service unavailable” responses (for example, a circuit breaker or a rate limiter tripping). The fix is in the application.

Prevention

  • Monitor currentThreadsBusy / maxThreads as a first-class signal. Sustained >0.80 is a ticket. == 1.0 for >120s with maxThreads > 50 and uptime >120s is a page. Gate on uptime to avoid cold-start false positives.
  • Set explicit timeouts on every outbound call. Default socket timeouts in JDBC drivers and HTTP clients are infinite. A hung downstream service will consume threads forever without them.
  • Configure StuckThreadDetectionValve with a realistic threshold. The default 600s is too long for most user-facing services. Page when stuckThreadCount > maxThreads * 0.5.
  • Coordinate keepalive timers across the proxy and Tomcat. Mismatched keepAliveTimeout and proxy keepalive produces connection resets that look like intermittent 503s.
  • Monitor the accept queue with ss. It is invisible to JMX. A non-zero Recv-Q is the leading indicator of connection refusal.
  • Parse the access log for 5xx specifically. JMX errorCount lumps 4xx and 5xx; it cannot tell you whether a 503 is happening.
  • Confirm acceptCount against net.core.somaxconn. Setting acceptCount higher than somaxconn has no effect.

How Netdata helps

  • Correlating currentThreadsBusy, maxThreads, connectionCount, and maxConnections on one per-second timeline makes it obvious whether the constraint is the worker pool or the poller, which determines the fix.
  • ML anomaly detection on request throughput and processing time surfaces a slow backend draining the thread pool before currentThreadsBusy reaches maxThreads.
  • The OS-level accept queue (ss Recv-Q) and file descriptor counts sit alongside the JVM metrics, so connection-layer saturation is not hidden behind a healthy JVM dashboard.
  • Per-second resolution means the thread-exhaustion cascade (busy threads rising, throughput collapsing, errors following) is visible as a sequence rather than a single aggregated point.
  • CPU and GC metrics next to thread pool state distinguish I/O-blocked exhaustion (low CPU) from compute-bound or GC-thrashing saturation (high CPU), which changes the response.
  • Virtual-thread deployments, where currentThreadsBusy reports -1, still surface connection-count signals so the exhaustion profile remains observable.