The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / tomcat / tomcat-session-memory-leak ▌

Operations Guides

Tomcat sessions eating the heap: bots, getSession(true), and unbounded growth

Post-GC heap baseline climbing steadily, Full GC frequency increasing, and eventually an OutOfMemoryError or a GC death spiral that leaves the JVM effectively unresponsive. Thread dumps look normal, CPU is dominated by GC threads, and a restart clears the problem temporarily before the cycle repeats within hours or days.

A common and overlooked root cause is HTTP session accumulation. Under Tomcat’s default StandardManager, every active session lives in JVM heap. When sessions are created faster than they expire, they fill the heap, promote to Old Gen, and drive major GCs. Sessions are frequently the largest heap consumer in a Tomcat app, and unbounded session count is a classic memory leak vector.

The specific pattern covered here is crawler traffic combined with application code that calls getSession(true). Bots that do not send cookies never join a session. Each request creates a new session, the Set-Cookie header is ignored, and the next request starts the cycle again. With the default session timeout of 30 minutes and maxActiveSessions at -1 (unlimited), a modest crawler storm can create hundreds of thousands of live sessions, each consuming heap.

What this means

Tomcat’s default session manager, StandardManager, stores all session state in memory. There is no eviction beyond timeout-based expiry. Sessions persist until they time out (default 30 minutes in the global web.xml) or are explicitly invalidated. The maxActiveSessions attribute defaults to -1, meaning the manager imposes no upper bound on live sessions.

The Servlet API contract for getSession(true) is the key mechanic. If no session exists for the request, a new one is created and a Set-Cookie header is sent back. If the client never returns that cookie, the session is never rejoined. The next request from that client creates yet another new session. The Servlet API documents this explicitly: if the client chooses not to join the session, getSession returns a different session on each request.

This is documented behavior, not a bug. But when the clients are bots, scrapers, monitoring agents, or misconfigured HTTP clients that drop cookies, and the application calls getSession(true) (or bare getSession(), which defaults to create=true) on every request path, the result is unbounded session creation.

flowchart TD
    A[Bot request, no cookie] --> B[App calls getSession true]
    B --> C[New session created in heap]
    C --> D[Set-Cookie sent, bot ignores it]
    D --> E[Next request: no cookie]
    E --> B
    C --> F[Sessions accumulate]
    F --> G[Promote to Old Gen]
    G --> H[Full GC frequency rises]
    H --> I[Post-GC baseline climbs]
    I --> J[OOM or GC death spiral]

JSPs make this worse. Unless a JSP page declares <%@page session="false"%>, the JSP engine implicitly calls getSession() on every request. A site serving JSP pages to crawlers without that directive will create a session for every bot hit, even for a page that renders static HTML and never touches session data.

Common causes

CauseWhat it looks likeFirst thing to check
Bots without cookies plus getSession(true)activeSessions spikes correlate with crawler User-Agent traffic in the access logAccess log User-Agent distribution against activeSessions trend
JSP pages without session="false"Every JSP hit creates a session, even static contentgrep for session="false" in JSP files
Session timeout too longactiveSessions plateaus at a high steady state under normal trafficweb.xml session-config session-timeout value
maxActiveSessions at default -1No protection, unbounded growth until heap exhaustsManager configuration in context.xml or server.xml
Large per-session objectsFew sessions but disproportionate heap impactHeap dump, dominant retainer analysis

Quick checks

These are read-only and safe on a production instance. The JMX and jstack calls attach to the JVM; they do not mutate state.

# Active session count via JMX for the ROOT context.
# Requires JMX enabled on the JVM (com.sun.management.jmxremote.*).
# Port 9090 is an example; use whatever your CATALINA_OPTS sets.
java -jar jmxterm.jar -l localhost:9090 -n -v silent -e \
  "get -b Catalina:type=Manager,host=localhost,context=/ activeSessions sessionCounter expiredSessions rejectedSessions"

# Old Gen usage and GC count over 5 seconds.
# If multiple Tomcat JVMs run on the host, pgrep returns several PIDs;
# pick the right one explicitly instead of command substitution.
jstat -gcutil <pid> 1000 5

# Count bot/crawler requests in today's access log.
# Path and extension vary by distribution; adjust to your server.xml AccessLogValve config.
grep -iE 'bot|slurp|crawler|spider' /var/log/tomcat/localhost_access_log.$(date +%Y-%m-%d).txt | wc -l

# Top User-Agents by request count.
# The awk field index assumes the combined log format; verify against your pattern.
awk -F'"' '{print $6}' /var/log/tomcat/localhost_access_log.$(date +%Y-%m-%d).txt | \
  sort | uniq -c | sort -rn | head -20

# Configured session timeout (global default).
grep -A2 'session-config' $CATALINA_BASE/conf/web.xml

# maxActiveSessions setting. Absent means default -1, unlimited.
grep -rn 'maxActiveSessions' $CATALINA_BASE/conf/ 2>/dev/null

The ratio that matters is sessionCounter (cumulative sessions ever created) versus expiredSessions (cumulative sessions expired). If sessionCounter is climbing much faster than expiredSessions, sessions are being created faster than the 30-minute timeout can reap them.

activeSessions is not reliably exposed via Manager Status XML across all Tomcat versions. JMX is the authoritative source. The Manager HTML status page shows session counts, but for scripted monitoring use the JMX bean Catalina:type=Manager,host=localhost,context=/<app>.

How to diagnose it

  1. Confirm activeSessions is growing monotonically. Sample the JMX attribute every minute. If it never plateaus during a traffic window, sessions are accumulating. A drop that lags falling traffic by roughly the session-timeout interval is normal, because sessions persist until timeout.

  2. Correlate the growth with access log bot traffic. Pull the top User-Agents from the access log for the same window where activeSessions is climbing. A surge of requests from Googlebot, Bingbot, semrush, Ahrefs, or generic scrapers that tracks the session growth curve is the signature.

  3. Estimate the session memory footprint. Multiply activeSessions by an estimate of per-session size. Per-session size varies enormously by application, from roughly 1KB for minimal sessions to 1MB or more for sessions holding large object graphs. A reasonable operating budget is that session count times per-session size should not exceed about 30% of max heap. If you do not know your per-session size, capture a heap dump and measure it directly.

  4. Confirm Old Gen pressure tracks session growth. Watch the Old Gen memory pool post-GC. If it rises in step with activeSessions, sessions are the dominant heap consumer and are promoting into Old Gen where they drive Full GCs.

  5. Audit the application for getSession calls. Search the codebase for getSession(true), getSession(false), and bare getSession(). Any call path reachable from a stateless request (a page that does not need session state) is a candidate for removal. Check JSPs for the absence of session="false".

Metrics and signals to monitor

SignalWhy it mattersWarning sign
activeSessions (JMX)Direct count of live sessions consuming heapMonotonic growth without plateau
sessionCounter vs expiredSessionsCreation rate vs reaping ratesessionCounter delta greatly exceeds expiredSessions delta
rejectedSessionsIndicates maxActiveSessions is set and being hitAny non-zero value means requests are failing session creation
Old Gen post-GC baselineShows whether live data is growingValleys rising in step with activeSessions
Full GC countA G1 Full GC means concurrent collection failed to keep upSustained Full GC frequency warrants investigation
Bot request rate (access log)Source of session creation pressureSpike in crawler User-Agents correlates with session growth

Fixes

Install CrawlerSessionManagerValve

Tomcat ships a valve for this problem: CrawlerSessionManagerValve. It detects requests from known crawler User-Agents and associates all of them with a single session per client identifier, regardless of whether they send a cookie. This collapses the per-request session explosion into one session per bot.

The valve is added to a Host or Context in server.xml or context.xml. It must be explicitly configured; it is not in the default pipeline. If you also use RemoteIpValve, place CrawlerSessionManagerValve after it so the valve sees the resolved client IP.

<Valve className="org.apache.catalina.valves.CrawlerSessionManagerValve"
       crawlerUserAgents=".*[bB]ot.*|.*Yahoo! Slurp.*|.*Feedfetcher-Google.*"
       sessionInactiveInterval="60"/>

The default crawlerUserAgents regex catches many common bots but will miss crawlers that do not match. Tune the regex against your actual bot traffic pulled from the access log. The sessionInactiveInterval governs how long the crawler-to-session mapping is cached.

Tradeoff: this only helps if the bot matches your regex. Sophisticated scrapers that forge browser User-Agents will not be caught, and you may need to add crawlerIps for known datacenter ranges.

Reduce session timeout

The global default is 30 minutes. If your application does not need long-lived sessions, reducing this to 5 or 10 minutes directly limits how many sessions can accumulate before the reaper clears them.

<session-config>
  <session-timeout>10</session-timeout>
</session-config>

Tradeoff: shorter timeouts log out idle users sooner. Measure the impact on your real session duration distribution before cutting aggressively.

Set maxActiveSessions

Setting maxActiveSessions to a finite value caps the damage. The failure mode is abrupt: when the limit is reached, any attempt to create a new session throws TooManyActiveSessionsException (an IllegalStateException subclass), which surfaces as a 500 error to the user. This is a circuit breaker, not a graceful throttle. Use it as a backstop, not as the primary control.

Tradeoff: legitimate users may see 500s during a bot storm. Pair this with CrawlerSessionManagerValve so bots rarely consume slots.

Disable session creation in JSPs

Add <%@page session="false"%> to JSP pages that do not need session access. This is especially important for landing pages, error pages, and any content frequently crawled. Without this directive, the JSP engine calls getSession() implicitly on every request, creating a session even for a page that renders static HTML.

Audit and remove unnecessary getSession(true) calls

Search the codebase for getSession(true) and bare getSession() calls on request paths that do not need session state. Replace with getSession(false) where the code only needs to read existing session data, or remove the call entirely. This is the root fix for applications that create sessions as a side effect of touching the request object.

Consider PersistentManager

PersistentManager can swap idle sessions to disk, reducing heap pressure. This is a heavier change and introduces disk I/O and serialization overhead. It is appropriate when sessions are legitimately large and numerous, not as a workaround for a bot-driven leak. It does not solve the underlying creation-rate problem, only the residency problem.

Prevention

  • Monitor activeSessions per context as a first-class metric. Alert on monotonic growth that does not plateau. The signal requires JMX; it is not reliably exposed via Manager Status XML.

  • Track sessionCounter and expiredSessions deltas. If creation consistently outpaces expiry, investigate before it becomes an OOM at 3 a.m.

  • Watch the correlation between bot traffic and session creation. A sudden bot spike on a site without CrawlerSessionManagerValve is a leading indicator.

  • Budget session memory explicitly. Estimate activeSessions peak times per-session size against max heap. If that number exceeds roughly 30% of heap, you are operating without margin.

  • Capture a heap dump before restarting during an incident. Once the JVM restarts, the evidence is gone. See the heap dump capture guide in Related guides.

How Netdata helps

  • Per-second activeSessions collection via JMX shows session growth as it happens, not on a minute lag. Correlating the activeSessions curve with the Old Gen post-GC baseline on the same dashboard confirms whether sessions are the heap consumer rather than a cache or application leak.

  • Anomaly detection on session count flags monotonic growth that a static threshold would miss. A session count that is normal at peak traffic may be a leak at off-peak hours; the detection adapts to the diurnal pattern rather than firing false positives.

  • Correlation of access log bot traffic with session creation shortens diagnosis from “heap is full” to “this specific crawler spike created 80,000 sessions.” Surfacing request rate by User-Agent alongside JVM metrics pinpoints the source.

  • GC frequency and pause duration alongside heap pools lets you watch the spiral form in real time and confirm the fix (CrawlerSessionManagerValve, shorter timeout) actually flattened the curve.

  • Post-GC heap baseline trending distinguishes a real leak from normal sawtooth. The valleys, not the peaks, indicate retained live data.