The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-outofmemoryerror-gc-overhead-limit-exceeded ▌

Operations Guides

Tomcat OutOfMemoryError: GC overhead limit exceeded: GC running but freeing nothing

java.lang.OutOfMemoryError: GC overhead limit exceeded means the heap is functionally full. GC has been running almost continuously and recovering almost nothing. This is the terminal stage of a GC death spiral, thrown just before a hard Java heap space OOM.

The exact rule the JVM applies: more than 98% of CPU time spent in GC and less than 2% of heap recovered, sustained across five consecutive collections. When both conditions hold, the JVM aborts rather than burn CPU indefinitely. Instantaneous heap usage can still read under -Xmx when this fires. By that point, the live set is effectively pinned to the ceiling.

On Tomcat, the root cause is almost always one of two things: an application or session leak that grows the live set, or a heap sized too small for the deployed apps’ working set. The fix is different for each. Suppressing the check with -XX:-UseGCOverheadLimit is not a fix. It removes the early warning, makes the JVM unresponsive on CPU for longer, and still ends in a hard OOM.

What this means

The GC overhead limit is a self-protection mechanism. When the JVM detects that GC is consuming the process and not reclaiming memory, it aborts. The error is thrown by the JVM, not by Tomcat. You will usually see it in catalina.out or in your application’s log framework, often as an unhandled exception bubbling up through a request thread.

The behaviour depends on the collector in use:

  • Parallel GC (default through JDK 8): throws GC overhead limit exceeded when the 98%/2%/five-collection conditions are met.
  • G1 GC (default since JDK 9): historically did not honour -XX:+UseGCOverheadLimit. OpenJDK bug JDK-8212084 implemented the check for G1 (integrated in JDK 26) and has been backported to OpenJDK 21.0.12 and 25.0.3; no 11u or 17u backport has been released. On older G1 builds, the same heap state produces a hard Java heap space OOM instead.
  • ZGC: does not implement the overhead limit at all. The same underlying leak will eventually produce Java heap space.
  • CMS: removed in JDK 14, but threw the overhead limit while it was still shipped.

Key diagnostic point: if you are running an older G1 build and you only ever see Java heap space, the same root cause applies. The error string is just different.

flowchart TD
    A[Heap live set grows] --> B[GC runs more often]
    B --> C[Post-GC baseline climbs]
    C --> D[GC dominates wall clock]
    D --> E{Recovers 2% heap?}
    E -- No, 5 cycles --> F[GC overhead limit exceeded]
    E -- Yes --> B
    F --> G[No heap dump configured means no evidence]

Common causes

CauseWhat it looks likeFirst thing to check
Application memory leakPost-GC old gen climbs monotonically over hours/days; heap dump shows one dominator (cache, collection, queue)Dominator tree in MAT
HTTP session accumulationactiveSessions grows without plateau; sessions promoted to old gen; bot or crawler trafficJMX Catalina:type=Manager activeSessions
Undersized -XmxDeath spiral triggers under peak load but heap returns to baseline after restart; no obvious dominator-Xmx vs peak live set after GC
Hot-redeploy classloader leakOutOfMemoryError: Metaspace typically, but heap can fill with classloader-retained objects; occurs after redeploy cyclesMetaspace pool; WebappClassLoader count in heap dump
HTTP/2 priority header leak (CVE-2025-31650)Heap grows under HTTP/2 traffic; affects Tomcat 9.0.76 to 9.0.102, 10.1.10 to 10.1.39, 11.0.0-M2 to 11.0.5Tomcat version; upgrade to 9.0.104+, 10.1.40+, 11.0.6+

Quick checks

These are read-only. None will worsen the situation.

# Confirm the error is in the logs and capture the timestamp window
grep -n "GC overhead limit exceeded" $CATALINA_BASE/logs/catalina.out

# Check OS OOM killer involvement (orthogonal but worth ruling out)
dmesg -T | grep -iE "oom|killed process"

# Look at the JVM flags actually in effect, including GC algorithm
jcmd $(pgrep -f 'catalina.startup.Bootstrap') VM.flags

# Heap occupancy snapshot via jstat (1000ms refresh)
jstat -gcutil $(pgrep -f 'catalina.startup.Bootstrap') 1000

# Per-pool breakdown; old gen column is the one that matters
jstat -gc $(pgrep -f 'catalina.startup.Bootstrap')

# Confirm HeapDumpOnOutOfMemoryError is set (so future OOMs leave evidence)
jcmd $(pgrep -f 'catalina.startup.Bootstrap') VM.flags | grep -i heapdump

# Active sessions per context (frequent culprit)
# jmxterm has no -e execute flag (-e is exit-on-failure); pipe the command on stdin.
echo "get -b Catalina:type=Manager,host=localhost,context=/ activeSessions" | \
  java -jar jmxterm.jar -l localhost:9090 -n -v silent

# Tomcat version (relevant to CVE-2025-31650 check)
$CATALINA_HOME/bin/version.sh

If multiple Tomcat instances run on the host, replace the $(pgrep ...) subshell with the specific PID.

If the JVM is still alive but in the spiral, capture a heap dump now, before you restart. The dump is the only durable evidence of what was on the heap.

How to diagnose it

  1. Capture a heap dump from the live process. Use jcmd <pid> GC.heap_dump /tmp/tomcat-$(date +%s).hprof or jmap -dump:format=b,file=/tmp/tomcat.hprof <pid>. Both pause the JVM while the dump is written. jcmd GC.heap_dump without -all and jmap -dump:live,... additionally trigger a Full GC. If the process is already thrashing, the dump will take longer. If the process has already died, you can only rely on whatever -XX:+HeapDumpOnOutOfMemoryError produced.

  2. Restart the JVM. Once the dump is captured, restart to restore service. Do not restart before capturing the dump unless the process is dead.

  3. Confirm HeapDumpOnOutOfMemoryError is configured going forward. If it was not set on the dying instance, set it now. Add to CATALINA_OPTS in setenv.sh:

    -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/tomcat/dumps/
    

    The dump path must be on a volume with enough free space. A full heap dump for a 4GB heap is roughly 4GB on disk.

  4. Load the dump in Eclipse MAT. Run the Leak Suspects report. The dominator tree and “incoming references” views identify the single object graph retaining most of the heap.

  5. Map the dominator to a cause. Common patterns:

    • A ConcurrentHashMap or HashMap field growing without bound: application cache without eviction.
    • A large Object[] under a session manager or context attribute: session accumulation.
    • Multiple WebappClassLoader instances with started=false: hot-redeploy leak. Look for extra classloaders where started is false, then trace their GC roots.
  6. If no clear dominator, suspect undersized heap. Compare peak post-GC old gen against -Xmx. If the working set genuinely needs more headroom than the configured heap provides, sizing is the fix.

  7. Cross-check session count. Use the JMX Catalina:type=Manager MBean. If activeSessions is climbing without plateau, sessions are the leak.

  8. Cross-check Metaspace. If the error string was Java heap space but the live set is largely classloader metadata, the real failure may be classloader-driven. Look at the Metaspace memory pool. Set -XX:MaxMetaspaceSize if it is unset.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Post-GC old gen utilizationTracks the live set after collection. The only heap metric that predicts OOM.Valleys rising over hours or days
GC overhead ratio (GC time / wall clock)Approaches the 98% threshold before the error firesSustained >20%
Full GC frequencyOn G1, any Full GC is abnormal; concurrent collection should keep upMultiple per minute
GC pause durationDirectly injects latency into requests in flightPauses >1s
Active session countSessions are the largest heap consumer in most web appsMonotonic growth without plateau
Process RSSCaptures native memory (thread stacks, metaspace, direct buffers) the heap metric missesRSS growing while heap is stable
Metaspace poolDetects classloader leaks on hot-redeployStep up after each redeploy that does not fall
catalina.out for OutOfMemoryErrorBinary canary that the spiral has already firedAny occurrence

Fixes

Application memory leak

The durable fix is in code: bound the cache, evict on size or TTL, close resources in finally or try-with-resources. There is no JVM flag that substitutes. After the code fix, watch the post-GC old gen baseline for several days to confirm it has stopped climbing.

If you cannot ship a code fix immediately, scheduling periodic rolling restarts is a stopgap. It does not address the leak. It bounds the blast radius.

HTTP session accumulation

If sessions are the leak, the application is creating sessions faster than they expire. Common patterns:

  • Bots and crawlers that ignore cookies force a new session per request when the app calls getSession(true).
  • JSP pages create a session by default. Add <%@ page session="false" %> where sessions are not needed.
  • Default session timeout is 30 minutes. Long timeouts with sustained traffic fill old gen.
  • maxActiveSessions=-1 by default, so there is no backstop.

Set maxActiveSessions on the Manager element to enforce a ceiling, and reduce session timeout where the application allows it. If the leak is bot-driven, avoid calling getSession() for bot user-agents, or front the app with a bot filter that prevents session creation.

Undersized -Xmx

If the heap dump shows no abnormal dominator and the post-GC old gen baseline genuinely exceeds about 75% of -Xmx at peak, increase -Xmx. Before doing so, confirm the host or container has the headroom. If actual heap usage plus native memory (thread stacks, metaspace, direct buffers) exceeds the cgroup memory limit, the kernel OOM-kills the process before the JVM reports heap pressure.

When increasing heap, also revisit GC tuning. Larger heaps with Parallel GC mean longer pauses. G1 handles multi-GB heaps better. ZGC avoids long pauses but does not throw GC overhead limit exceeded at all, so the same leak eventually surfaces as Java heap space.

Hot-redeploy classloader leak

The reliable fix is to stop hot redeploying. Use full JVM restarts for deployment. Containerised deployments that replace pods rather than redeploy into a running JVM avoid this entirely.

If hot redeploy is required, the JreMemoryLeakPreventionListener (enabled by default in server.xml) mitigates some JRE singleton leaks but cannot fix application code that retains references through ThreadLocals, JDBC driver registration, or unclosed log appenders. Tomcat’s “Find Leaks” feature in the Manager app identifies leaked contexts, but it calls System.gc() (forcing a full garbage collection), so avoid running it during production traffic.

CVE-2025-31650 (HTTP/2 priority header leak)

If your Tomcat version falls in the affected ranges (9.0.76 to 9.0.102, 10.1.10 to 10.1.39, 11.0.0-M2 to 11.0.5) and heap pressure correlates with HTTP/2 traffic, upgrade. The fix shipped in 9.0.104, 10.1.40, and 11.0.6. The vulnerability is a memory leak from improper cleanup of failed requests with invalid HTTP priority headers.

Prevention

  • Set -XX:+HeapDumpOnOutOfMemoryError and HeapDumpPath on every Tomcat JVM. Without it, an OOM leaves zero evidence. The dump path must be on a volume with enough free space.
  • Alert on post-GC old gen, not instantaneous heap. Instantaneous heap follows a sawtooth and is supposed to fill before GC. The valley is the signal. Track the post-GC baseline as a trend. Alert when it crosses about 75% of -Xmx.
  • Alert on the GC overhead ratio. Trend GC time as a fraction of wall clock. Sustained >10% is concerning. Sustained >20% means the death spiral is starting.
  • Set -XX:MaxMetaspaceSize. Without it, metaspace grows until the OS kills the process with no JVM-level error. Pick a value that leaves headroom for normal classloading.
  • Monitor activeSessions per context. Growing sessions without a plateau is the leading indicator of session-driven OOM.
  • Disable hot redeploy in production. Use full restarts or pod replacement.
  • Patch to a fixed Tomcat version if you are in the CVE-2025-31650 range.
  • Do not suppress the check. -XX:-UseGCOverheadLimit is occasionally suggested on forums. It removes the early warning and still ends in a hard OOM. Do not use it as a fix.

How Netdata helps

Netdata surfaces the signals that precede GC overhead limit exceeded at per-second resolution, so you can catch the spiral before the error fires.

  • JVM memory pools per second (java.lang:type=MemoryPool): the G1 Old Gen or PS Old Gen trend is the early warning. Rising post-GC valleys over hours are the leak signature.
  • GC collection count and time (java.lang:type=GarbageCollector): shows Full GC frequency and the GC time / wall clock ratio. A spike here correlates directly with request latency spikes.
  • ML-based anomaly detection on heap and GC signals flags subtle baseline drift that static thresholds miss. Useful for slow leaks that take days to manifest.
  • Active HTTP session count (Catalina:type=Manager): the leading indicator for session-driven OOM.
  • Correlation between GC pause time and request processing time (Catalina:type=GlobalRequestProcessor processingTime): confirms latency spikes are GC-induced rather than backend-induced.
  • CPU utilization split by process: if GC threads are dominating CPU, the memory problem is masquerading as a CPU problem.