The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-hot-redeploy-vs-restart ▌

Operations Guides

Tomcat hot redeploy vs clean restart: when autoDeploy leaks and when it doesn't

When you deploy a new WAR, Tomcat gives you two paths. Let autoDeploy watch webapps/ and swap the application in place, or stop the JVM and start a clean one. Both deliver the new code. They differ in what they leave behind, measured in Metaspace.

The root cause is the WebappClassLoader. Each Context gets its own. On a hot redeploy the old one is supposed to be garbage collected. Often it is not. A lingering ThreadLocal, a DriverManager registration, a timer thread, or a static field can pin the old classloader and every class it loaded. On a clean restart the JVM dies, and every classloader dies with it. There is nothing to pin.

For the failure pattern in depth, see the companion guide on the WebappClassLoader that never dies.

What it is and why it matters

Tomcat isolates applications by giving each Context its own classloader. When the application loads a class, the WebappClassLoader reads it from the WAR’s WEB-INF/classes and WEB-INF/lib. All of that class metadata, method bytecode, constant pools, and field layout lives in Metaspace on JDK 8 and later (PermGen on JDK 7 and earlier). The classloader object itself sits on the heap, but the classes it defined live in Metaspace, and they are reachable only through the classloader that defined them.

The garbage collection rule is strict and asymmetric. A WebappClassLoader can be collected only when nothing references it and nothing references any class it loaded. A single live reference to a single application class pins the classloader, and through it, every class it loaded. This is why a classloader leak is expensive: not one class leaks, but the whole application’s class set on every redeploy.

The operational concern is that Metaspace is not heap. Most teams watch heap and set -Xmx. Far fewer set -XX:MaxMetaspaceSize. With no cap, Metaspace grows until the operating system kills the process. The JVM raises no OutOfMemoryError, writes no clean shutdown, and leaves only a kernel OOM log line. Heap looks fine, but RSS climbs and the JVM dies.

That is the asymmetry. Heap leaks surface slowly as a rising post-GC baseline. Classloader leaks surface as a staircase in Metaspace that only a clean restart can flatten.

How it works

Two deployment paths. Both end with the new application running. Only one inherits the old application’s ghosts.

flowchart TD
  subgraph Hot[Hot redeploy - autoDeploy true]
    HA[Drop WAR into webapps] --> HB[Tomcat stops old Context]
    HB --> HC[Old WebappClassLoader queued for GC]
    HC --> HD{Reference still held?}
    HD -- yes --> HE[Classloader retained]
    HE --> HF[Metaspace grows one step]
    HD -- no --> HG[Classloader collected]
  end
  subgraph Clean[Clean JVM restart]
    CA[Stop the JVM] --> CB[Process exits]
    CB --> CC[All classloaders die]
    CC --> CD[Start fresh JVM]
    CD --> CE[Metaspace at baseline]
  end

The hot redeploy path

autoDeploy defaults to true on the Host element in current Tomcat releases. On a periodic check of appBase (the webapps/ directory), Tomcat notices a new or updated WAR. It stops the old Context, creates a new WebappClassLoader, loads the new classes, and starts the new Context. The old classloader is queued for collection. If nothing pins it, it is collected and Metaspace returns to roughly its previous level. If anything pins it, the old classes stay, and the next redeploy adds another copy on top.

Tomcat has accumulated defenses against common pins, and they do real work. The JreMemoryLeakPreventionListener forces certain JRE threads to start at container startup so they inherit the system classloader rather than the first webapp classloader. On stop, Tomcat deregisters JDBC drivers the webapp registered, flushes the java.beans.Introspector cache, and renews the thread pool so ThreadLocal values held by pooled threads do not pin the old classloader. It also logs warnings when it detects threads the application started but did not stop.

These defenses reduce the leak rate. They do not eliminate it. The remaining pins live in application or library code: a ThreadLocal whose value indirectly references a webapp class, a logging appender that captured the classloader, a Timer or ScheduledExecutorService the app never cancelled, a JMX MBean registered but not unregistered, or a static field in a shared library that references an application class. Tomcat’s stop-time detection does not catch indirect references. The Manager app’s “Find Leaks” analysis does, but that button invokes System.gc(), which is disruptive on a loaded production JVM.

The result is a staircase. Each hot redeploy adds a step in Metaspace. The degradation curve is “staircase to cliff.” If MaxMetaspaceSize is set, you eventually hit it and get OutOfMemoryError: Metaspace. If it is not set, you eventually hit the container or host memory limit and the kernel kills the process.

The clean restart path

A clean restart is the trivial case. The JVM shuts down. The process exits. Every classloader, pinned or not, is destroyed because the address space is gone. When the JVM starts again, it builds a fresh classloader hierarchy from zero. Metaspace starts at baseline. There is nothing to pin because there is nothing left.

This is why a full restart is the reliable fix and why containerized environments that replace the whole JVM on deploy rarely see classloader leaks. The leak is a property of the hot redeploy path. A clean restart has no old classloader to pin.

There is a third knob worth naming so you do not confuse it with autoDeploy. The Context attribute reloadable defaults to false. When true, Tomcat watches WEB-INF/classes and WEB-INF/lib for changes and reloads the Context automatically. The Tomcat documentation explicitly discourages it for deployed production applications. autoDeploy on the Host and reloadable on the Context are different mechanisms, but both take the hot redeploy path and both can leak.

Where it shows up in production

The failure pattern has a signature shape, and it almost always involves an environment mismatch.

  • The classic trap. A team hot-deploys all day in CI and staging but restarts cleanly in production. The leak never surfaces in those environments because they are short-lived or get restarted nightly. It appears the first time someone hot-deploys in production “just this once,” and Metaspace starts climbing in steps that never come back down.
  • Spring Boot WAR on standalone Tomcat. A Spring Boot application packaged as a WAR and deployed to a standalone Tomcat has a well-documented Metaspace cost per redeploy. Each redeploy builds a new WebappClassLoader and reinitializes the Spring context, and the previous one is frequently retained.
  • Long-lived shared dev or staging boxes. A Tomcat that absorbs dozens of redeployed builds a day is the fastest way to hit the staircase. These boxes look fine until they do not.
  • The first hot redeploy after a long clean-run stretch. Production Tomcats that have been restarted on every deploy for years can absorb one accidental hot deploy without incident, which lulls teams into thinking hot deploy is safe. It is not the first one that kills you, it is the cumulative step count.

The common thread: hot redeploy only looks safe when you rarely do it. The cost is paid per redeploy, and it accumulates.

When to use each

The decision is simpler than teams make it.

Prefer a clean JVM restart in production. Treat each production deploy as an immutable replacement of the JVM, not an in-place swap. In containerized environments this is the natural model. Outside containers, a deploy script that stops Tomcat, replaces the WAR, and starts Tomcat avoids the leak entirely. You lose a few seconds of cold start and JIT warmup, which is almost always cheaper than a Metaspace incident.

If you must hot-deploy, budget Metaspace per redeploy. Measure how much Metaspace climbs on a single undeploy and redeploy cycle. Divide remaining headroom by that number to estimate how many hot redeploys you have before trouble. Then schedule a clean restart before you get there, or set -XX:MaxMetaspaceSize so the JVM fails loudly with OutOfMemoryError: Metaspace instead of dying silently to the kernel OOM killer.

Set -XX:MaxMetaspaceSize regardless of path. Without it, Metaspace is unbounded and the only termination is an OS kill with no JVM-level error. A cap converts a silent catastrophe into a loud, diagnosable one.

Keep reloadable=false in production. It defaults to false and should stay that way. Use autoDeploy only where you have explicitly budgeted for the leak, and prefer it never in production.

Use the Manager “Find Leaks” function carefully. It is a diagnostic, not a routine. It invokes System.gc() and reports webapps whose classloader failed to collect. Run it on a staging box that mirrors production, not under live production load.

Signals to watch in production

SignalWhy it mattersWarning sign
Metaspace usageClass metadata lives here. Hot redeploy leaks show up as steps that do not drop.Step increase after each redeploy that never returns to baseline
LoadedClassCountCorroborates Metaspace growth with class loading activity.Monotonic growth after an undeploy and redeploy cycle
catalina.out leak warningsTomcat logs detected pins at stop time.“appears to have started a thread … but has failed to stop it”
Process RSSIncludes non-heap memory the heap graphs hide.RSS climbing while heap baseline is flat
Deploy events vs MetaspaceConfirms causation, not just correlation.Each deploy event lines up with a Metaspace step

The diagnostic rule is concrete: after a full undeploy and redeploy cycle, Metaspace should return to within roughly 10 percent of its pre-undeploy value. Any persistent growth indicates a leak. If you see growth, the fix is not to tune around it but to switch that environment to clean restarts.

How Netdata helps

  • Per-second Metaspace tracking from the JVM memory pools exposes the staircase pattern that daily or hourly polling smooths over. A single hot redeploy that adds a permanent step is visible immediately, not after the process dies.
  • Correlating deploy events with Metaspace steps turns “memory keeps growing” into “every deploy adds 80 MB that never comes back,” which is the difference between a leak you act on and one you live with until it pages you.
  • LoadedClassCount alongside Metaspace confirms whether the growth is class metadata (classloader leak) rather than something else pushing native memory.
  • Process RSS next to heap utilization catches the case where heap looks healthy but the process is walking toward the OOM killer through non-heap growth.
  • Heap and GC context keeps the diagnosis honest. If you also see rising post-GC heap baseline, you may have a heap leak in parallel, which needs a different fix than a classloader leak.