The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-monitoring-checklist ▌

Operations Guides

Tomcat monitoring checklist: the signals every production instance needs

Tomcat’s capacity model rests on a small number of bounded resources: a worker thread pool, a connection poller, the JVM heap, Metaspace, and the OS file descriptor table. Most production outages are one of these hitting its limit while the JVM process keeps running. A monitoring setup that only tracks CPU and memory will miss the dominant failure mode: thread pool exhaustion with the JVM at low CPU and a healthy process.

This checklist is organized into four cumulative maturity levels: survival, operational, mature, and expert. Each level assumes the previous one is in place. The single most important Tomcat-specific signal, thread pool utilization, sits at survival level because without it you cannot distinguish “Tomcat is up” from “Tomcat is accepting connections but cannot process them.”

The signals come from three sources: the Tomcat Manager status XML at /manager/status?XML=true, JMX MBeans under the Catalina: and java.lang: domains, and OS-level tools such as ss and /proc. Several critical signals, including active session count, GC activity, and file descriptor usage, are not exposed through Manager XML and require JMX. Plan for JMX from the start.

Read your current dashboards against each level. Anything missing is a gap. The “Common blind spots” section lists the gaps that most reliably produce 3 a.m. pages.

How the levels are organized

The four levels build upward. Survival tells you the service is alive. Operational tells you it is fast enough. Mature tells you it is not about to saturate. Expert gives you leading indicators for leaks and queues before users notice.

flowchart TD
    L4[Level 4 expert
accept queue, post-GC trend, metaspace leak] L3[Level 3 mature
sessions, GC ratio, fd used/limit, stuck threads] L2[Level 2 operational
throughput, processing time, error split, memory pools] L1[Level 1 survival
process alive, connector responds, context STARTED, thread pool] L1 --> L2 --> L3 --> L4

Level 1: survival

The absolute minimum to know Tomcat is alive and serving. Without all four of these, you are flying blind.

  • JVM process alive. The process exists. Source: OS process table, pgrep -f 'org.apache.catalina.startup.Bootstrap' for standalone, or the application JAR for embedded. Why it matters: catches JVM crashes and OS OOM kills. Limitation: a live process does not mean requests are being served. The OS OOM killer leaves no Tomcat-level log, only the kernel log.
  • Connector responds to a real request. A TCP connect is not enough. You must issue an HTTP request that exercises the servlet pipeline. Why it matters: a TCP connect can succeed while the accept queue holds the connection and no worker thread picks it up. A 503 returned by Tomcat itself signals thread pool or connector saturation, not an application error.
  • Application context is STARTED. Each deployed Context must be in the STARTED LifecycleState. Source: Catalina:type=Context,host=localhost,context=/<app> state attribute, or manager/text/list. Why it matters: a FAILED context returns 404 to every request, which looks like a missing route rather than a down app. Tomcat does not restart failed contexts automatically.
  • Thread pool utilization. currentThreadsBusy / maxThreads per connector. Source: Catalina:type=ThreadPool,name="http-nio-8080". Why this is survival-level: when busy reaches maxThreads, Tomcat stops processing new requests even though the JVM is healthy and the port is open. Default maxThreads is 200 and default minSpareThreads is 10.

Threshold guidance for the thread pool: sustained busy at 100% of maxThreads (with maxThreads > 50 and uptime greater than 120 seconds) is a page. Sustained ratio above 0.80 for more than five minutes is a ticket. Gate alerts on uptime to avoid cold-start spikes, because the pool starts at minSpareThreads and the first burst can briefly saturate before more threads are created.

Level 2: operational

What a competent team monitors in addition to survival. These turn “Tomcat is slow” into a diagnosable signal.

  • Request throughput. Rate derived from cumulative requestCount on Catalina:type=GlobalRequestProcessor,name="http-nio-8080". Warning sign: throughput drops below 50% of same-daypart baseline while upstream reports traffic arriving. The counter resets on restart, so compute deltas with reset awareness.
  • Request processing time. Cumulative processingTime in milliseconds divided by the delta of requestCount gives an average. Warning sign: average trending upward more than 2x over baseline. Important limitation: JMX gives averages only. For percentiles you need the access log with %D (request duration in milliseconds) configured, which is not present in the default or combined log pattern.
  • Error count, split by class. errorCount on the GlobalRequestProcessor MBean lumps 4xx and 5xx together and cannot safely be paged on, because crawler 404s would fire the alert. For 5xx-only alerting, parse the access log by status code. Warning sign: any sustained 503 from Tomcat itself means thread or connector exhaustion, not an application bug.
  • JVM memory pools. Per-pool breakdown from java.lang:type=MemoryPool,name=.... Old Gen post-GC rising indicates a memory leak. Metaspace growing monotonically across redeploys indicates a classloader leak. Pool names vary by GC algorithm: “G1 Old Gen”, “PS Old Gen”, “Tenured Gen”.

Level 3: mature

Full coverage for a production-grade deployment. These signals catch saturation and internal state before users do.

  • Post-GC heap baseline. Not instantaneous usage. The sawtooth fills then drops on GC, so high instantaneous usage is normal. What matters is the valley after GC. Warning sign: post-GC old gen consistently above 85% of max and trending up. Alerting on raw heap above 80% produces constant false positives.
  • GC time ratio. Cumulative CollectionTime divided by wall clock time, from java.lang:type=GarbageCollector,name=.... Healthy is below 5%. Above 10% is concerning. Above 20% is a GC death spiral. With G1GC (default since JDK 9), any Full GC warrants investigation.
  • Open file descriptors vs limit. OpenFileDescriptorCount / MaxFileDescriptorCount from java.lang:type=OperatingSystem, or ls /proc/<pid>/fd | wc -l. Warning sign: above 80% of limit, or monotonic growth unrelated to connection count. Production needs ulimit at 65535 or higher; defaults of 1024 or 4096 are routinely too low.
  • Active session count per context. activeSessions on Catalina:type=Manager,host=localhost,context=/<app>. Not available via Manager XML, requires JMX. Warning sign: monotonic growth without plateau. Sessions persist until timeout (default 30 minutes), so a traffic drop does not immediately reduce the count. Bots that do not send cookies can create one session per request.
  • Stuck thread detection. Requires StuckThreadDetectionValve configured explicitly. It is not enabled by default and its default threshold is 600 seconds. Without it, stuck threads are indistinguishable from a normal backend slowdown in JMX metrics. The valve exposes stuckThreadCount via JMX.
  • JVM CPU utilization. ProcessCpuLoad from java.lang:type=OperatingSystem. If more than 30% of CPU is GC, you have a memory problem, not a CPU problem. Exclude cold start from alerting, since JIT compilation consumes CPU for the first few minutes after startup.

Level 4: expert

Leading indicators and leak detection. These are the signals that prevent the next major incident.

  • Accept queue depth. ss -tnl 'sport = :8080' Recv-Q. No JMX counter exists for this. When Recv-Q is non-zero and approaching acceptCount (default 100), connections are about to be refused with RST. This is invisible to Tomcat logs.
  • Connection count vs maxConnections. connectionCount / maxConnections on the ThreadPool MBean. Default maxConnections is 8192 for all connectors since Tomcat 9.0.30 (in 8.5.x and 9.0.0-9.0.29: 10000 for NIO/NIO2, 8192 for APR/native). Idle keepalive connections consume poller slots and file descriptors but not threads, so connection count can legitimately far exceed busy threads.
  • Post-GC heap trend over days. Extrapolate the rising valley to project when old gen hits 90% of max. GC overhead feedback often makes the cliff arrive sooner than a linear projection suggests.
  • Metaspace growth per redeploy. Measure the step increase after each undeploy and redeploy cycle. If Metaspace does not return to within 10% of its pre-deploy value, you have a classloader leak. If -XX:MaxMetaspaceSize is not set, Metaspace grows until the OS kills the process with no JVM-level OOM.
  • Per-endpoint latency percentiles. Access log %D parsed for p95 and p99. A p50 of 100 ms with a p99 of 15 seconds averages to “fine” while 1% of users time out. Bimodal distributions indicate two distinct failure modes.
  • Total JVM thread count. ThreadCount on java.lang:type=Threading. Expected baseline is roughly maxThreads plus 50 overhead threads. Each thread consumes stack memory (default 512 KB to 1 MB via -Xss), so 500 threads silently commits around 500 MB off-heap.
  • JDBC connection pool utilization. numActive / maxActive and wait count on the pool MBean (Tomcat JDBC pool: Active / MaxActive / WaitCount; Commons DBCP2: NumActive / MaxTotal). Pool exhaustion cascades directly into thread pool exhaustion, because a thread holding an HTTP request and waiting on getConnection() is doubly expensive.
  • 5xx-only error rate from access logs. The JMX errorCount mixes 4xx and 5xx and cannot be paged on safely. Log parsing by status code is the only reliable source for server-error alerting.

Common blind spots

These are the gaps that most reliably produce incidents, drawn from the failure patterns that recur across Tomcat deployments.

  • Monitoring the PID, not the thread pool. JVM alive does not mean Tomcat is serving. A process check alone misses thread pool exhaustion, the most common Tomcat outage.
  • Alerting on instantaneous heap. The sawtooth is supposed to fill before GC runs. Only the post-GC valley indicates a leak.
  • Relying on average latency. JMX processingTime / requestCount is an average. Stuck threads averaged with fast requests look healthy while users time out. Percentiles require access log %D.
  • Missing the accept queue. When threads and connections are both exhausted, requests queue in the OS TCP backlog and are then refused with RST. No JMX counter covers this.
  • Default access log omits timing. The default and combined patterns lack %D. Without it, per-request latency analysis is impossible retroactively.
  • Ignoring Metaspace. Teams watch heap and are blindsided by OutOfMemoryError: Metaspace after hot redeploys. If MaxMetaspaceSize is unset, the process dies from OS OOM kill with no JVM error.
  • Confusing thread exhaustion with connection exhaustion. With NIO, the poller keeps accepting connections up to maxConnections (10000 default) even when the thread pool is full. Connections are only refused when both maxConnections and acceptCount are exceeded.
  • Leaving the Manager app exposed. Default installations include the Manager and Host Manager apps, which allow WAR upload and remote code execution. Remove them or restrict to localhost in production.

How Netdata helps

Netdata’s per-second collection is useful for Tomcat specifically because the dominant failure modes are saturation cascades that develop over tens of seconds, not minutes. The value is in correlating signals on a shared timeline.

  • Thread pool saturation correlation. Netdata surfaces currentThreadsBusy and maxThreads alongside request throughput and processing time on the same timeline, so you can see whether a busy pool is caused by slow requests (throughput dropping, processing time rising) or a pure load spike (throughput rising).
  • Post-GC heap versus instantaneous heap. Correlating heap utilization with GC collection time makes the sawtooth valley visible without manual chart inspection, and the valley is what actually indicates a leak.
  • GC overhead ratio. Collection time as a fraction of wall clock time, plotted next to request latency, shows whether latency spikes align with GC pauses rather than application code.
  • File descriptor pressure. FD count plotted against connection count separates a connection surge from an FD leak, which require different responses.
  • Counter-reset-aware rate computation. requestCount, errorCount, and processingTime are cumulative counters that reset on restart. Netdata handles the reset so throughput and error-rate charts do not flatline or spike after a deploy.