The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-monitoring-maturity-model ▌

Operations Guides

Tomcat monitoring maturity model: from survival to expert

Most Tomcat monitoring setups start with a process check, a health endpoint probe, and maybe a 5xx alert. That catches crashes. It does not catch the failures that actually page teams: thread pool exhaustion, GC death spirals, accept queue overflow, classloader leaks after hot redeploys. This article maps the signals that matter across four maturity levels so you can audit what you have, identify what you are missing, and prioritize what to add next.

The model is cumulative. Each level includes everything below it. Signals are drawn from the Tomcat-specific failure archetypes operators encounter in production, not from generic JVM advice.

flowchart TD
  L1["Level 1: Survival
process, connector, context, OOM, 5xx"] L2["Level 2: Operational
threads, heap, GC, throughput, latency"] L3["Level 3: Mature
stuck threads, Metaspace, RSS, accept queue"] L4["Level 4: Expert
NMT, allocation rate, G1 regions, cgroup"] L1 -->|adds| L2 L2 -->|adds| L3 L3 -->|adds| L4

How to read these levels

Each level assumes the previous level’s signals are in place. The tables list the signal, the failure mode it catches, and the source. The note at the end of each level explains the gap that motivates the next level.

Deployment variants matter. Standalone Tomcat and embedded Tomcat (Spring Boot) expose the same JMX MBeans, but the Manager application is absent in embedded deployments, and application properties replace server.xml. Clustered Tomcat with session replication adds its own monitoring surface (DeltaManager and BackupManager MBeans, cluster heartbeat) that is outside this model.

Virtual threads on JDK 21+ with Tomcat change the thread pool model fundamentally. When virtual threads are enabled, currentThreadsBusy reports -1 and maxThreads is not meaningful. Connection-based monitoring (connectionCount minus keepalive estimate) becomes the proxy for active work, and the standard thread pool utilization thresholds at Level 2 no longer apply.

Level 1: survival

The absolute minimum to know Tomcat is alive and serving.

SignalWhat it catchesSource
JVM process aliveProcess crash, OS OOM kill, init failureOS process table: pgrep -f org.apache.catalina.startup.Bootstrap
HTTP connector respondsConnector failure, total thread exhaustion, all apps failed to startHTTP request to configured port (not just TCP connect)
Context state STARTEDFailed deployment, startup exceptionJMX Catalina:type=Context state attribute, or Manager /manager/text/list
OutOfMemoryError in logsHeap exhaustion eventcatalina.out or application log grep
HTTP 5xx error rateServer-side application failuresAccess log status filtering (JMX errorCount mixes 4xx and 5xx)

A TCP connect success does not mean Tomcat is healthy. The OS accept queue holds connections even when Tomcat has no threads to process them. Your health check must complete an HTTP request through the servlet pipeline, not just open a socket.

A context in FAILED state at startup may not produce HTTP errors. Requests to that context return 404, which can be confused with a missing route. Tomcat does not restart failed applications by default. They remain in FAILED state until manually redeployed or the JVM restarts.

What this misses: thread pool saturation (process up, port open, requests hang or 503), GC death spirals (process alive, latency climbing to seconds), Metaspace exhaustion from classloader leaks, and accept queue overflow (clients get connection refused, Tomcat logs nothing).

Level 2: operational

What a competent production team monitors. These signals catch the top Tomcat failure archetype: thread pool exhaustion.

SignalWhat it catchesSource
Thread pool utilizationThread exhaustion, slow backends, stuck threadsCatalina:type=ThreadPool currentThreadsBusy / maxThreads
Post-GC heap utilizationMemory leaks (not instantaneous usage)java.lang:type=Memory HeapMemoryUsage, correlated with GC events
GC pause time and frequencyGC death spiraljava.lang:type=GarbageCollector CollectionTime, CollectionCount
Request throughputProcessing constraints, backend failuresCatalina:type=GlobalRequestProcessor requestCount delta
Request processing timeLatency degradationprocessingTime / requestCount delta, or access log %D for percentiles
4xx/5xx error splitClient errors vs server errorsAccess log status code parsing (JMX cannot separate them)
File descriptor countFD exhaustion, socket leaksjava.lang:type=OperatingSystem OpenFileDescriptorCount / MaxFileDescriptorCount
Active session countSession leaks, bot session spamCatalina:type=Manager activeSessions
JDBC pool utilizationConnection pool exhaustion, connection leaksCatalina:type=DataSource numActive / maxActive, or pool-specific MBean
JVM CPU utilizationGC CPU consumption, compute saturationjava.lang:type=OperatingSystem ProcessCpuLoad

The thread pool signal deserves emphasis. currentThreadsBusy / maxThreads is the single most important Tomcat-specific metric. When it reaches 1.0, requests are queuing. With NIO, connections are still accepted up to maxConnections (default 8192 for all connector types), but no thread is available to process them. Users experience timeouts while the JVM sits at low CPU, looking healthy.

Alerting on post-GC heap means tracking the valley of the sawtooth, not the peak. High instantaneous heap usage before GC is normal. A rising post-GC baseline over hours or days indicates a memory leak. Alerting on raw “heap above 80 percent” produces constant false positives because the value spends most of its time there by design.

The default access log pattern (%h %l %u %t "%r" %s %b) omits processing time. Without adding %D (milliseconds in Tomcat 9.0.x, microseconds from Tomcat 10.0 on), you cannot compute per-request latency percentiles. JMX processingTime is a cumulative total, not an average. Dividing by requestCount yields an arithmetic mean that hides tail latency: a few 30-second requests averaged with a thousand 10-millisecond requests looks fine.

What this misses: threads stuck indefinitely (no StuckThreadDetectionValve configured), Metaspace growth from classloader leaks, the OS accept queue depth, per-endpoint latency breakdowns, and the full process RSS (heap metrics exclude non-heap memory).

Level 3: mature

Full coverage for a production-grade deployment. These signals close the gaps that cause the most confusing incidents: the ones where Tomcat looks healthy but users are failing.

SignalWhat it catchesSource
Stuck thread countThreads blocked indefinitely on backends, deadlocksStuckThreadDetectionValve MBean stuckThreadCount
Metaspace utilizationClassloader leaks on hot redeployjava.lang:type=MemoryPool,name=Metaspace Usage
Process RSSNon-heap memory growth, OS OOM kill riskOS /proc/<pid>/status VmRSS or cgroup memory
GC overhead ratioGC death spiral leading indicatorComputed: CollectionTime delta / wall clock time
Connection count vs maxConnectionsNIO poller saturationCatalina:type=ThreadPool connectionCount, maxConnections
Accept queue depthKernel-level connection refusalss -tnl Recv-Q on the listen socket
Per-endpoint latency and errorsEndpoint-specific slow pathsAccess log with %D, grouped by URI
Class loading countClassloader leak confirmationjava.lang:type=ClassLoading LoadedClassCount

Stuck thread detection requires explicit configuration. The StuckThreadDetectionValve is not enabled by default. Add it to server.xml or context.xml with a threshold. The default of 600 seconds is too high for most production applications. Without it, stuck threads are indistinguishable from normal busy threads until the pool drains.

Metaspace is the silent killer. If MaxMetaspaceSize is not set, Metaspace grows toward the JVM’s computed ceiling. In containers with tight memory limits, the cgroup OOM killer can terminate the process before the JVM reaches its own limit, leaving no OutOfMemoryError in the log. When the JVM does hit its Metaspace limit, it throws OutOfMemoryError: Metaspace. Teams that watch heap but ignore Metaspace are blindsided by this after N hot redeploys. In containerized environments with immutable deploys, this is rarely an issue. In environments that hot-deploy, it is inevitable.

The accept queue is invisible to JMX. When maxConnections is reached and acceptCount (default 100) fills, the kernel rejects TCP connections with RST. Tomcat logs nothing. The only detection is OS-level ss -tnl monitoring showing non-zero Recv-Q, or client-side connection error rates.

Two security signals belong at this level. Manager application access from unexpected sources warrants alerting, since the Manager allows WAR upload and remote code execution. AJP connector posture matters too: the AJP connector should bind to localhost with secretRequired=true. An exposed AJP port on 0.0.0.0 without authentication is the Ghostcat (CVE-2020-1938) attack surface, which allows unauthenticated file read and potential RCE.

What this misses: the native memory breakdown behind RSS growth, allocation and promotion rates that predict future heap pressure, and container-level CPU throttling that extends GC pauses.

Level 4: expert

Signals that teams add after their third or fourth major Tomcat incident. These require deeper JVM expertise and higher instrumentation overhead, but they catch failures hours or days before they become outages.

SignalWhat it catchesSource
Native memory trackingOff-heap leaks, thread stack bloat, JNI allocationsjcmd <pid> VM.native_memory summary with -XX:NativeMemoryTracking=summary
Allocation and promotion rateFuture heap pressure before it manifestsGC log analysis or JDK Flight Recorder
G1GC region statisticsHumongous allocations, evacuation failuresjcmd <pid> GC.heap_info
Cgroup CPU throttlingContainer-induced latency spikes and extended GC pauses/sys/fs/cgroup/cpu.stat nr_throttled (v2) or /sys/fs/cgroup/cpu/cpu.stat (v1)
Scheduled thread dumpsLock contention patterns before they cause outagesjcmd <pid> Thread.print on a recurring schedule
JIT deoptimization eventsSudden latency regression from deoptimized hot methods-XX:+PrintCompilation analysis

Native memory tracking explains the gap between heap usage and RSS. When heap looks stable but RSS keeps climbing, the growth is in thread stacks, native buffers, direct memory, or JNI allocations. NMT breaks this down by category. The tradeoff: enabling NMT adds overhead. Reserve it for diagnosis or run it permanently only on instances where you have CPU headroom.

Allocation rate (bytes per second into young generation) is a better predictor of GC pressure than absolute heap usage. A sudden increase means the heap fills faster, GC runs more frequently, and pauses lengthen. Promotion rate (objects surviving to old gen) predicts old gen pressure specifically. Neither is exposed via standard JMX. They require GC log analysis or JDK Flight Recorder.

Cgroup CPU throttling is the hidden latency source in containerized Tomcat. CFS quota throttling can deschedule GC threads mid-collection, extending a 50-millisecond pause to seconds. The nr_throttled counter in the cgroup CPU controller reveals this. Java 17+ has proper cgroup v2 support; older Java versions may misread container limits.

Scheduled thread dumps, not just incident-driven ones, reveal contention patterns that build slowly. A thread dump taken every few minutes during peak load, compared across days, shows whether threads are converging on the same lock or backend call before the pool drains.

What most teams get wrong regardless of level

A few gaps persist across all maturity levels because they are instrumentation or configuration failures, not missing metrics:

  • Monitoring the PID, not the thread pool. JVM alive does not mean Tomcat is serving. Thread pool exhaustion leaves the process up, the port open, and requests hanging or returning 503.
  • Alerting on instantaneous heap. The sawtooth is normal GC behavior. Track the post-GC baseline. Alerting on raw utilization above 80 percent produces constant false positives.
  • Average latency instead of percentiles. JMX processingTime / requestCount is an arithmetic mean. Configure %D in the access log and compute p95 and p99 to see actual user experience.
  • No timeouts on outbound connections. Default socket timeouts in java.net.HttpURLConnection, Apache HttpClient, and most JDBC drivers are infinite. One hanging backend consumes threads forever. This is a code problem, but without StuckThreadDetectionValve, monitoring cannot surface it.
  • Confusing thread exhaustion with connection exhaustion. Thread pool full does not mean connections are refused. With NIO, connections continue to be accepted up to maxConnections. Connections are only refused when both maxConnections and acceptCount are exceeded.

How Netdata helps

Netdata collects JVM and Tomcat metrics at per-second resolution and correlates them with OS-level signals, which is where the gap between symptom and root cause usually closes:

  • Thread pool utilization (currentThreadsBusy / maxThreads) collected per-second and correlated with request processing time and throughput distinguishes backend-driven saturation from load-driven saturation within seconds.
  • JVM heap, GC pause time, and GC frequency tracked continuously expose the post-GC baseline trend without manual jstat polling or GC log parsing.
  • Metaspace and class loading counters captured across redeploy events surface classloader leaks before they become silent OOM kills.
  • OS-level file descriptor count, CPU, and RSS alongside JVM metrics give the full memory picture (heap, non-heap, native) in a single correlated view.
  • Anomaly detection on connector error rates, session counts, and connection counts flags deviations from baseline before static thresholds trip.