The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-memory-heap-fragmentation ▌

Operations Guides

Envoy memory_heap_size vs memory_allocated: tcmalloc fragmentation that fools dashboards

Envoy exposes three memory gauges that operators naturally try to map onto process RSS: server.memory_allocated, server.memory_heap_size, and server.memory_physical_size. On a busy proxy the first two diverge by 2x-4x routinely, and dashboards that graph heap_size look like a slow leak even when nothing is wrong.

The gap is not a bug. Envoy ships with tcmalloc as its default allocator, and tcmalloc is built for allocation latency, not for prompt return of freed memory to the OS. It parks freed pages in per-thread caches, central caches, and the page heap free lists so the next allocation is fast. Those pages stay mapped into the process, counted in memory_heap_size, and very often still resident in RSS.

What it is and why it matters

server.memory_allocated and server.memory_heap_size both come from tcmalloc’s MallocExtension. memory_allocated reflects bytes currently handed out to Envoy’s own data structures: connection buffers, stats storage, config caches, filter state, the deferred-delete list. memory_heap_size reflects the total bytes tcmalloc has reserved from the OS, which includes everything in memory_allocated plus everything tcmalloc has freed at the application layer but is holding onto internally.

The ratio memory_allocated / memory_heap_size is the fragmentation signal. Low ratios are normal. A heap that is several times the size of allocated is not pathology, it is tcmalloc holding the working set hot so it does not have to call back into the kernel on the next burst. The operational signal you want is growth of memory_allocated itself without a corresponding traffic change. That, not the ratio, is what indicates a leak or unbounded accumulation.

server.memory_physical_size exists to give operators a number that is closer to actual RSS. It maps to tcmalloc’s generic.total_physical_bytes and includes allocator overhead the kernel still counts as resident. Even memory_physical_size is not a perfect RSS proxy because some allocation paths in Envoy bypass tcmalloc. For process RSS itself, read /proc/<pid>/status or cgroup memory.current.

How tcmalloc holds freed memory

When Envoy frees an object, the bytes do not go back to the kernel. They walk through several layers of tcmalloc internal caches first, and each layer has different implications for RSS.

flowchart TD
    APP[Envoy frees object] --> TC[Per-thread cache]
    TC --> CC[Central cache]
    CC --> PHF[Page heap free list: mapped, in RSS]
    PHF -. madvise DONTNEED .-> PHU[Page heap unmapped: virtual only]
    PHU -. slow release .-> OS[Returned to OS]
  • Per-thread caches: each worker thread has its own fast bins, and most frees return here. Bytes in thread caches are still mapped and still in RSS.
  • Central cache: when a thread cache fills, it spills into a process-wide central cache that other threads can pull from. Still mapped, still resident.
  • Page heap free list: when spans of pages are fully free, they sit in the page heap free lists, classified by size class. Mapped, resident, ready for reuse.
  • Page heap unmapped: eventually tcmalloc may madvise(MADV_DONTNEED) a span. The virtual range stays reserved but the kernel is free to reclaim the physical pages.
  • Returned to OS: the final step, which only the unmapped path leads to.

Only the last step reduces RSS, and tcmalloc does it on its own schedule. The practical effect is that a workload which allocates a burst of memory (loading many routes, a large config push, a body-buffering event) and then frees it will see memory_allocated drop, memory_heap_size stay flat, and RSS drift down slowly for minutes afterward.

The /memory admin endpoint shows the breakdown directly:

# Inspect tcmalloc internal breakdown
curl -s http://localhost:9901/memory | jq .

The fields returned are:

FieldWhat it represents
allocatedbytes currently handed to Envoy (matches server.memory_allocated)
heap_sizetotal bytes reserved from OS (matches server.memory_heap_size)
pageheap_freebytes in free, mapped pages (still in RSS)
pageheap_unmappedbytes in free, unmapped pages (virtual only, kernel may reclaim)
total_thread_cachebytes held in per-thread caches
total_physical_bytesphysical memory estimate (matches server.memory_physical_size)

pageheap_free + total_thread_cache + allocated is roughly the gap between heap_size and what tcmalloc could plausibly return. Watching those three fields while you load and unload clusters tells you whether the gap is mostly thread caches (will drift down over time) or pageheap_free (will linger).

Where it shows up in production

The fragmentation gap is most visible in deployment shapes that produce allocation spikes followed by quiet periods.

Sidecar meshes with frequent xDS churn. Each cluster add or remove drives a burst of small allocations: cluster metadata, host state, per-worker connection pool state, route table entries. Removing the cluster frees the objects, but the spans they lived in are not necessarily returned. Operators running Istio at scale with services deploying and removing continuously see heap_size march upward over hours or days even though memory_allocated tracks traffic. This is the workload shape most likely to fool a dashboard.

Edge proxies handling large buffered bodies. A buffer filter, or an ext_authz configuration that buffers the full request body, allocates a large contiguous region for the duration of the request. Under burst traffic many such regions exist simultaneously. When the requests complete the memory is freed, but heap_size and RSS will reflect the high-water mark for a while.

Config reloads and hot restart. Loading a new config, even one that ends up smaller, allocates the new graph before the old one is released. Hot restart runs two Envoy processes side by side. During the handover window RSS roughly doubles as both processes hold their heaps. This is expected, but if your alert is a fixed threshold on RSS it will fire every deploy.

Lua and Wasm filters. Lua extensions in particular may manage some allocations through their own runtime rather than tcmalloc , so memory used by Lua may appear in process RSS without being fully reflected in memory_allocated or memory_heap_size. If RSS is climbing and Envoy’s memory stats are flat, look at filter-level allocations before assuming the stat is wrong.

Common misuses

Alerting on memory_heap_size directly. This is the mistake the title is about. A threshold like “page if heap_size > 1GB” will fire on a perfectly healthy proxy that has seen a normal burst in the last hour. The signal you want is memory_allocated growth without traffic growth.

Graphing heap_size and allocated on the same panel without context. Two lines diverging by 3x looks alarming to anyone who has not seen tcmalloc behavior before. Either graph the ratio (allocated / heap_size) alongside the raw bytes, or graph allocated and memory_physical_size, and document that heap_size is allocator working set, not memory pressure.

Treating shrink_heap as a fix. The overload manager exposes a shrink_heap action that calls into tcmalloc’s release path . It is worth enabling as a soft pressure-relief mechanism, but it is not guaranteed to reduce RSS . Treat it as best-effort, not as a way to lower your alert thresholds.

Assuming memory_physical_size is RSS. It is closer than heap_size, but it still reflects tcmalloc-tracked allocations. Filter runtimes, segment allocations from other allocators, and any third-party allocator linked into the process are not in it. For cgroup-bounded workloads, watch memory.current on the cgroup, not Envoy’s internal stat.

Blaming Envoy for sidecar OOMs without per-process breakdown. In a sidecar pod the kernel kills the whole pod when memory.current hits the limit. The OOM is often blamed on Envoy because heap_size looks high, while the real consumer was the application. Always look at per-process RSS inside the pod before assigning blame.

Signals to watch in production

SignalWhy it mattersWarning sign
server.memory_allocatedActual bytes Envoy is using for live data structuresMonotonic growth without traffic growth is a real leak or cardinality explosion
server.memory_allocated / server.memory_heap_sizeFragmentation ratio; baseline for the deploymentSustained drop well below the historical baseline for the same workload
server.memory_physical_sizeCloser to RSS than heap_size, includes tcmalloc overheadApproaching cgroup memory limit
cgroup memory.currentActual RSS the kernel will OOM onSustained growth toward memory.max
Total stat count (curl -s localhost:9901/stats | wc -l)Cardinality explosions inflate the stats regionCount growing without infrastructure growth
Overload manager action gaugesConfirm whether the overload manager has triggeredAny active memory-related action
server.live and drain gaugesCorrelate memory events with hot restart or drainDrain in progress during a memory spike is expected

The two signals worth alerting on are server.memory_allocated sustained growth against traffic, and memory_physical_size (or cgroup memory.current) approaching the limit. The ratio is a dashboard annotation, not an alert.

Working with the allocator, not against it

If the fragmentation gap is genuinely causing operational pain (RSS so high you cannot pack pods the way you want), there are a few real levers.

Configure the overload manager deliberately. The overload manager uses the envoy.resource_monitors.fixed_heap resource monitor, which depends on max_heap_size_bytes. If it is unset or set too high, the overload manager never triggers and Envoy goes straight from “fine” to OOM with no graceful degradation. Setting max_heap_size_bytes deliberately, with shrink_heap at the lower threshold and stop_accepting_requests at the upper threshold, gives Envoy a self-protection path. This is a capacity decision, not a fragmentation fix.

Investigate jemalloc as an alternative allocator. Operators running Envoy builds that link jemalloc instead of tcmalloc have reported lower fragmentation and lower RSS for the same workload, particularly on workloads with heavy allocation churn. This is a build-time decision; you cannot swap allocators at runtime.

Reduce churn at the source. The biggest lever is usually reducing how often you allocate and free large graphs of small objects. For xDS-driven deployments this means less frequent, larger config pushes instead of many small pushes. For body-buffering workloads it means streaming instead of buffering where possible.

Force kernel reclaim under cgroups. On cgroup v2 you can write a byte count to memory.reclaim to nudge the kernel into reclaiming pages. This is a workaround, not a fix, and it does not help if the pages are still mapped in tcmalloc’s pageheap_free.

How Netdata helps

  • Netdata’s per-second collection of server.memory_allocated, server.memory_heap_size, and server.memory_physical_size makes the fragmentation ratio visible as its own chart. Plotting the ratio next to the raw bytes makes “low ratio is normal” obvious to anyone reading the dashboard.
  • ML anomaly detection on memory_allocated separates genuine growth (the leak signal) from the normal sawtooth of allocator behavior. Alerts fire on the in-use signal, not on heap_size.
  • Correlating memory_allocated with downstream_cx_active, upstream_cx_active, and total stat count tells you whether growth is connection-driven, traffic-driven, or cardinality-driven. Three different root causes, three different fixes.
  • The overload manager action gauges appear alongside the memory stats. When shrink_heap or stop_accepting_requests flips active, the memory chart is one click away.
  • Cgroup-level memory.current and memory.max from the same Netdata node let you compare Envoy’s self-reported memory to what the kernel will actually OOM on. In sidecar deployments this is the comparison that matters.
  • Hot restart epoch and server.state transitions are collected per second, so memory spikes during drain are easy to identify and exclude from baseline calculations.