The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-heap-dump-before-restart ▌

Operations Guides

Tomcat heap dump before restart: capturing evidence with jmap and jstack

When Tomcat is failing, the instinct is to restart. That instinct is right for recovery and wrong for diagnosis. The JVM process holds the only copy of the evidence: the object graph that explains the heap exhaustion, the thread stack traces that explain the pool stall, the GC state that explains the death spiral. Once the process exits, that evidence is gone. You are left restarting blind with nothing but an OutOfMemoryError line in catalina.out.

This guide covers the capture procedure you run in the narrow window between “Tomcat is broken” and “Tomcat is restarted.” The goal is two artifacts: a heap dump (.hprof) and one or more thread dumps (jstack output) for offline analysis with Eclipse MAT or VisualVM. The procedure is the same whether the trigger is an OutOfMemoryError: Java heap space, a GC death spiral, or unexplained thread pool exhaustion. See the GC overhead pattern and thread pool exhaustion for broader diagnostic context.

The procedure assumes the JVM process is still alive. If the process has already been OOM-killed or has crashed, the heap is gone. Your remaining evidence is the GC log, the kernel OOM log, and any automatic heap dump produced by -XX:+HeapDumpOnOutOfMemoryError if it was configured.

What this captures

Heap dump (.hprof). A full snapshot of the JVM heap at capture time: every live object, every reference, every class instance count and retained size. This tells you what is consuming memory and who is retaining it. Load it in Eclipse MAT and run the “Leak Suspects” report to find the dominant retainer. Without this, you are guessing at the leak source.

Thread dump (jstack output). A snapshot of every JVM thread and its current stack trace, including thread state (RUNNABLE, BLOCKED, WAITING, TIMED_WAITING). For thread pool exhaustion, it reveals whether threads are stuck on a socket read, a database connection acquire, a lock, or a GC pause. Multiple thread dumps taken a few seconds apart are more useful than one, because they distinguish a thread that is permanently stuck from one that is just slow.

Capture both. A heap dump without a thread dump tells you what is retained but not what is actively blocking. A thread dump without a heap dump tells you what threads are doing but not whether the heap is the constraint.

Prerequisites

  • JDK, not JRE. jmap and jstack ship with the full JDK. They are absent from JRE-only distributions. This bites teams running minimal container images. Verify before you need them.
  • Disk space. A heap dump can be as large as the JVM’s maximum heap (-Xmx). A 4 GB heap produces a dump approaching 4 GB. On a full disk, the dump fails partway through and is useless. Check available space on the target volume before capturing.
  • Process access. You must run jmap and jstack as the same user that owns the JVM process, or as root. The tools attach via the JVM Attach API, which requires matching credentials.
  • The process must still be alive. If the JVM has exited, the Attach API has nothing to attach to. Check with pgrep first.
  • Local shell access to the host. In Kubernetes, you need kubectl exec into the pod, or a sidecar that can run the tools. Remote JMX is not sufficient for jmap and jstack, which are local attach tools.

Procedure

Run these steps in order. The order matters: thread dumps are cheap and fast, heap dumps are slow and disruptive.

1. Confirm the process is alive and find the PID

# Find the Tomcat JVM process
pgrep -f 'catalina.startup.Bootstrap'

For embedded Tomcat (Spring Boot), match on your application JAR or main class instead. If pgrep returns nothing, the process is already gone and this procedure cannot proceed. Check dmesg | grep -i oom for an OOM kill, or the container restart count.

Record the PID. Every subsequent command uses it.

2. Capture thread state first

jstack is fast and does not trigger a Full GC. Capture it before anything that might pause or destabilize the JVM.

# Capture three thread dumps, 5 seconds apart, for comparison
for i in 1 2 3; do
  jstack <pid> > /tmp/tomcat-threads-$i.txt
  sleep 5
done

The three-dump variant is worth the extra 10 seconds. Comparing stack traces across dumps tells you which threads are permanently stuck (same frame in all three) versus which are just slow (frames moving).

If jstack fails to attach, try the jcmd equivalent:

# jcmd alternative for thread dump
jcmd <pid> Thread.print > /tmp/tomcat-threads-1.txt

jcmd is the supported successor to jstack. Both work on current JDKs, but jmap and jstack carry an “experimental and unsupported” label in the Oracle documentation and may not be present in future JDK releases.

3. Check disk space before the heap dump

# Check available space on the target volume
df -h /tmp

# Check the JVM's max heap to estimate dump size
jcmd <pid> VM.flags | grep -i MaxHeapSize

The dump will approach the size of the live heap, up to -Xmx. If /tmp does not have enough space, write to a mounted volume that does. In containers, /tmp is often an overlay filesystem backed by a small ephemeral layer. Mount a persistent volume or an emptyDir with adequate capacity.

4. Capture the heap dump

Use one of the two commands below. The first is the traditional jmap form. The second is the jcmd equivalent, which is the recommended path on modern JDKs.

# Option A: jmap (traditional, still works on JDK 17/21)
jmap -dump:format=b,file=/tmp/tomcat-heap.hprof <pid>

# Option B: jcmd (recommended)
jcmd <pid> GC.heap_dump /tmp/tomcat-heap.hprof

Warning: the live suboption triggers a Full GC. jmap -dump:live,format=b,file=<path> <pid> forces a stop-the-world Full GC before dumping, which on a large heap can freeze the JVM for tens of seconds. On a JVM already in a GC death spiral, this pause can trip health checks and cause a load balancer to drain the node. Use live only if you specifically want the post-GC object set; omit it if you want the full heap as it currently sits.

# Full heap, no forced GC (preferred for forensics)
jmap -dump:format=b,file=/tmp/tomcat-heap.hprof <pid>

# Post-GC heap only (triggers Full GC, disruptive)
jmap -dump:live,format=b,file=/tmp/tomcat-heap.hprof <pid>

The output file must not already exist. Both jmap and jcmd fail if the target file is present. Use a unique filename including the PID and timestamp:

# Unique filename to avoid the "file exists" failure
jmap -dump:format=b,file=/tmp/tomcat-heap-$(date +%Y%m%d-%H%M%S)-<pid>.hprof <pid>

Wait for the command to return. A heap dump of a multi-gigabyte heap can take minutes. Do not interrupt it. A partial dump cannot be opened by MAT.

5. Verify the dumps are complete

# Check file sizes (heap dump should be large, thread dumps small)
ls -lh /tmp/tomcat-heap-*.hprof /tmp/tomcat-threads-*.txt

# Check the heap dump magic bytes (should start with "JAVA PROFILE")
head -c 20 /tmp/tomcat-heap-*.hprof | xxd | head -2
# Expected: 4a 41 56 41 20 50 52 4f 46 49 4c 45 ...

A heap dump smaller than a few hundred megabytes for a JVM with a multi-gigabyte -Xmx is suspicious. It may indicate the dump failed silently or the file was truncated by a full disk.

Confirm the thread dumps contain real application frames. Open the jstack output and look for org.apache.catalina.connector.CoyoteAdapter.service or your application code. A thread dump full of JVM-internal frames with no application code suggests you captured during a GC pause or startup, not during the failure.

If you captured three dumps and all three are byte-identical, either the JVM was completely frozen (likely a GC pause or deadlock) or jstack returned an error. Check for “Unable to attach” messages at the top of the file.

6. Copy the artifacts off the host

# From a Kubernetes pod (specify the exact filename; kubectl cp does not glob)
kubectl cp <namespace>/<pod>:/tmp/tomcat-heap-20240101-120000-12345.hprof ./tomcat-heap.hprof

Do not analyze the dump on the production host. Eclipse MAT needs heap at least as large as the dump file to parse it, and running MAT on the same host as the failing JVM adds memory pressure to an already pressured system.

7. Now you can restart

Once the heap dump and thread dumps are safely copied off the host, restart the JVM to recover service. The evidence is preserved.

flowchart TD
    A[Tomcat failing: OOM or GC spiral] --> B{Process alive?}
    B -- No --> C[Evidence gone. Check GC log, OOM killer, auto-dump]
    B -- Yes --> D[Find PID: pgrep catalina.startup.Bootstrap]
    D --> E[jstack: capture 3 thread dumps, 5s apart]
    E --> F[Check disk space vs -Xmx]
    F --> G[jmap or jcmd: capture heap dump]
    G --> H[Verify file size and hprof magic]
    H --> I[Copy artifacts off host]
    I --> J[Restart JVM to recover]
    J --> K[Analyze offline: MAT or VisualVM]

Common pitfalls

jmap or jstack not found in containers. Most production container images ship only the JRE. Running jmap inside the pod fails with executable file not found in $PATH. You have three options: build the image from a JDK base (for example eclipse-temurin:17-jdk instead of -jre), use jcmd if it happens to be present, or capture the dump from outside the container by exec’ing into the host’s JDK installation against the container’s PID (which requires host PID namespace sharing and is fragile). The clean fix is to build diagnostic-capable images or run a sidecar with the JDK.

The eclipse-temurin JRE images for JDK 17 and 21 ship neither jcmd, jmap, nor jstack (their bin contains only java, jfr, jrunscript, keytool, and - on 21 - jwebserver), so a JDK-based image or a sidecar is required for heap capture; kill -3 still works for a thread dump.

Disk fills before the dump completes. A heap dump can be as large as -Xmx. In a container with a small writable layer, the dump fills the overlay filesystem, the pod gets evicted, and you lose both the dump and the process. Always write to a mounted volume with capacity greater than -Xmx, and check df before capturing.

HeapDumpOnOutOfMemoryError does not fire before OOM kill. The flag -XX:+HeapDumpOnOutOfMemoryError is the right production default, but it only triggers when the JVM throws OutOfMemoryError. In a container with a memory cgroup limit lower than -Xmx, the kernel OOM killer can terminate the process before the JVM ever sees the OOM condition. No OutOfMemoryError, no automatic dump. This is why manual capture matters: the kernel does not wait for the JVM to be polite.

jmap -dump:live makes a bad situation worse. The live option forces a Full GC before dumping. On a heap that is already at 95% and barely collecting, this Full GC can run for tens of seconds or minutes, during which the JVM is unresponsive. If health checks fail, the orchestrator may kill the pod before the dump completes. Prefer the non-live form for forensics unless you specifically need the post-GC snapshot.

SIGQUIT (kill -3) is not a substitute for jstack. kill -3 <pid> writes a thread dump to the JVM’s stdout, which for Tomcat usually means catalina.out. It works when jstack cannot attach, but the output is less structured and interleaves with application logging. Use it as a fallback, not a first choice.

OpenJ9 and other non-HotSpot JVMs use different tooling. jmap and jstack are HotSpot tools. On OpenJ9 (used in some IBM and Semeru images), the heap dump format is .phd, not .hprof, and the capture mechanism differs. If your Tomcat runs on OpenJ9, the commands in this guide do not apply directly. Check the OpenJ9 documentation for the equivalent -Xdump options.

OpenJ9 writes heap dumps as .phd (capture with jcmd <pid> GC.heap_dump or the -Xdump:heap trigger). Eclipse MAT cannot open .phd by itself: it needs the IBM DTFJ feature installed (IBM publishes a MAT build with it pre-installed).

Signals to monitor

These signals tell you when to initiate the capture procedure. If you wait for a hard crash, the evidence is gone.

SignalWhy it mattersWarning sign
Post-GC heap utilizationRising valley means live data growing toward OOMPost-GC heap above 85% of -Xmx and trending up
Full GC frequencyA Full GC on G1 means the concurrent cycle is not keeping upMore than one Full GC per minute
GC overhead ratioGC consuming CPU that should serve requestsGC time above 20% of wall clock
Thread pool utilizationThreads at max with low CPU means blocked, not busycurrentThreadsBusy at maxThreads sustained over 60s
OutOfMemoryError in catalina.outJVM has already failed to allocateAny occurrence; capture immediately if process survives

See the monitoring checklist for the full signal catalog.

How Netdata helps

The capture procedure is reactive. Continuous monitoring tells you when to trigger it.

  • Per-second JVM heap metrics show the sawtooth pattern and, more importantly, the post-GC baseline trend. A rising valley over hours or days is the leading indicator that a heap dump will soon be necessary.
  • GC collection count and time, broken down by young and old generation, surface the transition from normal collection to a death spiral. Sustained Full GCs on G1 are a page-worthy trigger for capture.
  • Thread pool utilization (currentThreadsBusy relative to maxThreads) distinguishes a heap problem from a thread problem. If threads are at max but heap is fine, the issue is blocking, and jstack is the priority capture, not jmap.
  • Anomaly detection on heap and GC metrics flags deviations from the baseline before thresholds are crossed.
  • Cgroup memory usage in containerized deployments shows the gap between JVM heap and the container limit, indicating how close the kernel OOM killer is to firing before the JVM throws OutOfMemoryError.