The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / traefik / traefik-out-of-memory-oom ▌

Operations Guides

Traefik out of memory: OOM kills and the crash loop that follows

Traefik was proxying traffic normally. Then the process vanished. In Kubernetes you see OOMKilled in the pod’s last terminated state and a restart count that keeps climbing. On bare metal or Docker you see exit code 137 and a container that came back up, ran for a while, and died again.

An OOM kill has no graceful degradation phase. Go’s garbage collector absorbs growing memory pressure by running more often, right up until the moment it cannot. Then the kernel OOM killer terminates the process instantly. Every in-flight connection drops at once: client connections, backend connections, WebSockets, gRPC streams. From the outside it looks like a total, simultaneous outage of everything behind the proxy.

The crash loop that follows makes diagnosis harder, not easier. Each restart clears the heap, so by the time you look at the process it is healthy and climbing toward the same wall. The evidence you need is in the growth curve before the kill, not in the process after it.

What this means

The kernel (or the container runtime enforcing the cgroup memory limit) killed the Traefik process because it ran out of memory. Two things follow:

  1. The limit that matters is the cgroup limit, not host RAM. A Traefik container with a 1 GiB limit on a 64 GiB host dies at 1 GiB. Comparing process_resident_memory_bytes against total system memory tells you everything is fine right up to the kill.
  2. RSS is not live data. With the default GOGC=100, the Go runtime triggers GC when the heap doubles, so steady-state RSS is roughly 2x the live heap. A process whose live data is 400 MiB will sit around 800 MiB of RSS. This is normal, and it is why headroom calculations go wrong when teams size the container limit to observed RSS without understanding the multiplier.
flowchart TD
  A[Memory driver: leak, buffering, routing table, log buffer] --> B[RSS climbs toward cgroup limit]
  B --> C{GC keeping up?}
  C -->|yes| D[Longer GC pauses, latency jitter]
  D --> B
  C -->|no| E[Kernel OOM kill: instant, all connections dropped]
  E --> F[Restart: heap cleared, process healthy]
  F --> G[Traffic returns, memory climbs again]
  G --> B

Each trip around that loop drops all connections and produces a burst of client-visible errors. The loop period is your diagnostic window: headroom divided by growth rate tells you roughly how long each cycle takes.

Common causes

CauseWhat it looks likeFirst thing to check
Goroutine leak (hung backend connections)go_goroutines climbs monotonically, disconnected from traffic; RSS climbs in lockstepGoroutine profile from pprof; look for goroutines blocked on backend reads
Request/response body-buffering middlewareMemory tracks concurrent request count and payload size; spikes under large uploads/downloadsWhich routers have buffering or compress middlewares attached
Access log buffer growthEntrypoint latency rises while service latency stays normal; memory grows when log output is slow or blockedWhether the log destination (disk, pipe) is keeping up
Routing table sizeHigh baseline memory at rest that scales with router/service count; grows after onboarding wavesRouter and service counts in the provider; memory per config reload
Config-object accumulation across reloadsStepwise memory increases correlated with traefik_config_reloads_total incrementsOverlay reload count onto the RSS curve
Metric cardinality (per-router labels)Memory growth after enabling addRoutersLabels; series count explodesCardinality of the metrics endpoint

Two version-specific notes. Traefik v3.0 through v3.3 selected zstd before gzip by default; maintainers reproduced high-memory behavior and measured Brotli as especially memory-intensive. The fix reverted compression priority to gzip first and shipped in v3.4.0. Operators on affected v3 versions reported memory returning to normal after configuring gzip only or removing compression. Separately, one operator reported that the file provider with watch: true correlated with their leak and that watch: false resolved it; this was never confirmed as a core bug.

Quick checks

All read-only. Run these against the host and the metrics endpoint (adjust the port and pid for your deployment).

# 1. Confirm the kill actually was OOM (host kernel log)
dmesg -T | grep -i -E 'oom|killed process' | tail -20

# 2. In Kubernetes: confirm OOMKilled and see the restart count
kubectl get pod -n <ns> <traefik-pod> -o jsonpath='{.status.containerStatuses[*].lastState.terminated}{"\n"}'
kubectl get pod -n <ns> <traefik-pod> -o jsonpath='{.status.containerStatuses[*].restartCount}{"\n"}'

# 3. In Docker: check the last exit state (137 = 128 + SIGKILL)
docker inspect <container> --format '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Restarting}}'

# 4. Current RSS of the live process
grep VmRSS /proc/$(pgrep -x traefik)/status

# 5. The limit that actually applies (cgroup, not host RAM)
cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes

# 6. Goroutine count right now
curl -s http://localhost:8080/metrics | grep '^go_goroutines'

# 7. Heap in use (live data) vs RSS
curl -s http://localhost:8080/metrics | grep -E '^go_memstats_heap_inuse_bytes|^process_resident_memory_bytes'

# 8. GC pause behavior
curl -s http://localhost:8080/metrics | grep '^go_gc_duration_seconds'

# 9. Config reload rate (is memory growth tracking reloads?)
curl -s http://localhost:8080/metrics | grep '^traefik_config_reloads_total'

Note on the cgroup path: v2 uses memory.max, v1 uses memory/memory.limit_in_bytes. A value of max or a very large number means no limit is set, which shifts the question to host-level memory pressure.

How to diagnose it

The goal is to separate three cases: a leak (unbounded growth, needs a code or config fix), normal scaling (growth proportional to load, needs more headroom), and GC arithmetic (the process is fine but the limit is too tight for Go’s 2x RSS behavior).

  1. Confirm the OOM and find the limit. Steps 1-3 above confirm the kill; step 5 gives you the number that matters. Everything else is measured against that number.

  2. Reconstruct the growth curve. You need the RSS trend before the kill, which means a metrics system with history, not the live process. Plot process_resident_memory_bytes over the hours before the last few restarts. Three shapes:

    • Monotonic climb over hours/days, disconnected from traffic: leak. Continue to step 3.
    • Climb that tracks request rate and connection count, flattening off-peak: normal scaling or buffering proportional to load. Jump to step 5.
    • Sawtooth that resets at each restart and climbs again at the same rate: either, but the restart period gives you the growth rate: (limit minus baseline) divided by cycle time.
  3. Correlate with goroutines. Overlay go_goroutines on the same window. If goroutines and RSS climb together while request rate is flat, you have the classic goroutine leak: something (usually a backend that accepts connections and never responds, with no effective timeout) is pinning goroutines, and each leaked goroutine holds its closure’s memory. This is the most common Traefik OOM mechanism.

  4. Capture a goroutine profile before the next kill. With the debug API enabled, pull the profile while memory is high:

    # Requires the debug/pprof endpoint to be enabled (--api.debug=true)
    curl -s 'http://localhost:8080/debug/pprof/goroutine?debug=1' > goroutines.txt
    

    Group the stacks. Thousands of goroutines parked in the same backend-read or transport stack point at the offending backend and at missing timeouts (dialTimeout, responseHeaderTimeout, idleConnTimeout in serversTransport). Do not enable the debug API on a publicly reachable entrypoint; it exposes profiling data about your infrastructure.

  5. Check the middleware chain for buffering. Buffering and compression middlewares hold memory proportional to request/response body size times concurrency. If growth tracks large request or response bodies (check traefik_service_requests_bytes_total and traefik_service_responses_bytes_total alongside RSS), audit which routers carry buffering or compress middlewares, and on Traefik v3.0-3.2 treat the compress middleware as a prime suspect.

  6. Check reload correlation and routing-table size. Overlay traefik_config_reloads_total increments onto RSS. Stepwise increases after reloads in a high-churn environment suggest config-object accumulation. A high memory floor at rest that scales with router count is routing-table cost; very large routing tables (thousands of routers) can consume hundreds of MB before a single request arrives.

  7. Rule out log-buffer growth. If entrypoint latency rose while service latency stayed flat in the same window, suspect the access log writer blocking on a slow or full destination; blocked log goroutines accumulate memory the same way hung backend goroutines do. Check the filesystem the access log writes to.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
process_resident_memory_bytes vs cgroup limitThe actual OOM distanceSustained above ~80% of the limit, or rising trend
go_memstats_heap_inuse_bytesLive heap; separates real growth from GC headroomGrowing without corresponding traffic growth
go_goroutinesEarliest leak indicator; each goroutine pins memorySustained growth without matching connection/request growth
go_gc_duration_secondsGC struggling is the last warning before the wallp99 pauses climbing as RSS approaches the limit
traefik_config_reloads_totalLinks memory steps to config churnReload rate high and RSS stepping up with it
process_start_time_seconds / restart countDetects the crash loop itselfMultiple restarts within 30 minutes
traefik_open_connectionsConnections rising without request-rate growth feeds both FD and memory pressureSteady climb with flat request rate

Alert on the ratio of RSS to the container limit, not the absolute value. Memory above 80% of the limit is a ticket; the reason it cannot be a page from metrics alone is that the monitoring system often does not know the cgroup limit. Fix that gap explicitly: record the limit (or read it from the cgroup) and alert on the ratio, because by the time the OOM kill fires you have already lost every connection.

Fixes

Goroutine leak from hung backends

Set explicit transport timeouts in serversTransport so no request can pin a goroutine indefinitely: dialTimeout, responseHeaderTimeout, and idleConnTimeout. The specific failure this addresses is a backend that accepts the TCP connection but never responds; without responseHeaderTimeout, Traefik waits forever and each such request is a permanent goroutine. Identify the offending backend from the goroutine profile and remove it from rotation while you fix it. Tradeoff: timeouts that are too aggressive produce 504s for legitimately slow backends, so size them from observed p99 service latency, not guesses. See Traefik 504 Gateway Timeout.

Buffering and compression middleware

Remove buffering middlewares from routers that do not need them, and avoid buffering in front of large-payload routes (file upload/download, media). On Traefik v3.0-3.3, upgrade to v3.4.0 or later for the compression-priority fix. If you cannot upgrade immediately, configure gzip only or remove the compress middleware as the reported workaround. Tradeoff: dropping compression increases egress bandwidth and response times for compressible content; dropping buffering removes its protection for slow backends.

Access log pressure

Point access logs at stdout or an async destination, reduce verbosity (drop fields you do not use), and make sure the destination filesystem has space and write throughput. Tradeoff: less log detail for forensics; mitigate by keeping full detail on error responses only if your log format supports it.

Routing table and config churn

Consolidate routers where rules can be merged, and split very large configurations across multiple Traefik instances by responsibility. In high-churn environments, raise providersThrottleDuration so bursts of provider events batch into fewer rebuilds; this reduces both rebuild CPU and the allocation churn that feeds GC pressure. Tradeoff: longer throttle means slower convergence when you legitimately deploy.

Headroom and runtime limits

Size the container memory limit at roughly 2x the observed stable peak heap, per the GC arithmetic above: live data plus one doubling. On Go 1.19 and later, set GOMEMLIMIT below the cgroup limit so the runtime GCs harder as it approaches the wall instead of coasting into it; this converts some OOM kills into elevated GC CPU, which is survivable and visible. Tradeoff: aggressive GC costs CPU and adds latency jitter, which beats an instant kill but is not free.

Prevention

  • Alert on the ratio, not the absolute. RSS against the cgroup limit, with a ticket at ~80% sustained. A dead proxy is a page; a climbing ratio is the ticket that prevents it.
  • Trend goroutines against traffic. go_goroutines growing faster than connections or requests is the cheapest early-warning signal you have. Baseline it once.
  • Set transport timeouts before you need them. Every Traefik in production should have explicit dialTimeout, responseHeaderTimeout, and idleConnTimeout. Their absence is what turns one wedged backend into a leaked-goroutine farm.
  • Restart-count alerting. Multiple restarts in 30 minutes is the crash-loop signature. Catch the second kill, not the tenth.
  • Audit middleware chains after changes. Buffering and compression middlewares are memory multipliers; review which routers carry them during change review.
  • Load-test at production concurrency before raising traffic. Buffering and goroutine counts scale with concurrency, not request rate, so memory behavior at 10x your test concurrency is not linear.

How Netdata helps

  • Per-second RSS and Go runtime metrics on the Traefik process show the exact growth curve before each kill, which is the evidence the post-restart process no longer has.
  • Goroutine count alongside connection and request rates makes the leak-versus-scaling distinction visible on one screen instead of three terminals.
  • GC pause duration correlated with request latency shows when GC pressure is degrading traffic before the process dies, giving you an earlier tripwire than the OOM itself.
  • Container memory usage versus the cgroup limit is collected together, so the ratio that matters is computed from the right denominator rather than host RAM.
  • Restart events overlaid on memory and traffic turn the crash loop into a readable cycle: climb, kill, restart, climb, with the growth rate measurable from the chart.