The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-workspace-client-overflow ▌

Operations Guides

Varnish workspace_client overflow: 500 errors from oversized headers and cookies

HTTP 500 responses from Varnish when the backend is healthy, the object is in cache, the thread pool is not saturated, and the 503 backend-fetch path is not involved. The 500s come from Varnish itself, and they cluster on requests carrying large Cookie headers, long URLs, deep Via or X-Forwarded-For chains, or responses where VCL has added many synthetic headers.

The root cause is per-request workspace exhaustion. Varnish allocates a fixed block of memory (workspace_client) for each client request to hold the parsed HTTP object, headers, and VCL string operations. When request headers plus any data VCL pushes into that workspace exceed the allocation, the request fails with a 500.

The ws_client_overflow counter (Varnish 6.2+) is the definitive signal. On older releases, rely on client_resp_500 and losthdr. The fix is either to raise workspace_client or to reduce the size of the headers and cookies arriving at Varnish. The first option works but multiplies across every worker thread, so the second is usually the better long-term fix.

What this means

workspace_client is a per-request memory budget. The default is 64 kB on Varnish 6.x (64-bit) and 96 kB on Varnish 7.0+ (64-bit).

That budget holds:

  • the parsed HTTP request, including every header line
  • the Cookie header, often the largest single contributor
  • strings VCL builds during vcl_recv, vcl_hash, and vcl_deliver (for example regsuball results, synthetic headers, cookie manipulation)
  • ESI scratch space in some configurations

When the sum exceeds the allocation, the workspace allocator refuses the request and Varnish records the overflow. On Varnish 6.2+ the counter MAIN.ws_client_overflow increments directly. MAIN.client_resp_500 is a broader counter: it tracks all 500 responses delivered to clients, not just workspace-caused ones. MAIN.losthdr tracks a related but distinct condition: headers dropped because they exceeded http_max_hdr (default 64 headers), which also consumes workspace slots.

flowchart TD
  A[Client request with large Cookie/headers] --> B[Request parsed into workspace_client]
  B --> C{Workspace budget exceeded?}
  C -- no --> D[Normal request processing]
  C -- yes --> E[ws_client_overflow increments]
  E --> F[client_resp_500 increments]
  F --> G[500 returned to client]
  B --> H[VCL adds regsuball/synthetic headers]
  H --> C

The failure is per-request and stateless. A single abusive bot, a broken mobile SDK that appends tracking cookies without bound, or an application deploy that adds a new large cookie can trigger sustained 500s on a narrow slice of traffic while everything else looks fine.

Common causes

CauseWhat it looks likeFirst thing to check
Large Cookie header500s concentrated on authenticated or personalized paths; ws_client_overflow increments with request ratevarnishlog -g request -q 'RespStatus == 500' -i ReqHeader and inspect Cookie: length
VCL adding too much data500s appear after a VCL reload; regsuball or synthetic header logic in vcl_recv or vcl_delivervarnishadm vcl.list for recent reload; review VCL for regsuball, set req.http.*, synthetic
Deep proxy chainX-Forwarded-For or Via grows at each hop; 500s correlate with traffic from upstream proxiesvarnishlog -i ReqHeader:X-Forwarded-For and measure header length
http_max_hdr exceededlosthdr increments alongside or instead of ws_client_overflow; headers silently droppedvarnishstat -1 -f MAIN.losthdr and varnishadm param.show http_max_hdr
Long URLs500s on specific endpoints with long query strings; req.url dominates workspacevarnishlog -i ReqURL for affected transactions
Undersized workspace_client after tuning500s after increasing thread_pool_max without adjusting workspacevarnishadm param.show workspace_client and thread_pool_max

Quick checks

These commands are read-only and safe to run during an incident.

# Check workspace overflow counters (Varnish 6.2+)
varnishstat -1 -f 'MAIN.ws_*_overflow' -f MAIN.client_resp_500 -f MAIN.losthdr

# Current workspace_client and thread pool configuration
varnishadm param.show workspace_client
varnishadm param.show workspace_backend
varnishadm param.show thread_pool_max
varnishadm param.show http_max_hdr

# Find the failing transactions
varnishlog -g request -q 'RespStatus == 500' -i ReqURL -i ReqHeader -i RespStatus -i Debug

# Search logs for the workspace overflow message directly
varnishlog -q 'Debug ~ "workspace overflow"' -g request

# Confirm backend health is not the cause (500 here is Varnish-side, not backend)
varnishadm backend.list

# Identify the largest Cookie headers in live traffic
varnishlog -i ReqHeader:Cookie -g request | awk '{print length($0), $0}' | sort -rn | head -20

# Check loaded VCL versions and recent reload timestamps
varnishadm vcl.list

The combination that confirms workspace overflow: ws_client_overflow (or client_resp_500 on older releases) is incrementing, varnishlog shows workspace_client overflow in the Debug tag followed by RespStatus 500, and the affected requests carry visibly large Cookie or other headers.

How to diagnose it

  1. Confirm the counter is moving. Run varnishstat -1 -f 'MAIN.ws_*_overflow' -f MAIN.client_resp_500 twice, a few seconds apart, and compute the delta. A nonzero rate means active client-facing failures.

  2. Capture the failing transactions. Run varnishlog -g request -q 'RespStatus == 500' during the incident. Look for the Debug tag containing workspace_client overflow, then inspect ReqHeader:Cookie, ReqURL, and any ReqHeader lines to find the oversized contributor.

  3. Measure the actual header footprint. For the affected requests, sum the byte length of all request headers plus the URL. Compare against the configured workspace_client value. If the headers alone consume most of the budget, the overflow is data-driven, not VCL-driven.

  4. Rule out VCL as the amplifier. If the failing requests do not have unusually large headers, inspect VCL for regsuball, set req.http.* operations that append data, and any VMOD calls that allocate workspace (cookie manipulation VMODs are common offenders). A VCL reload that correlates with the start of the 500s is a strong signal.

  5. Check whether losthdr is also incrementing. If losthdr is nonzero alongside ws_client_overflow, the request has too many distinct headers (over http_max_hdr, default 64) in addition to or instead of raw size. The fix path differs: raise http_max_hdr rather than workspace_client.

  6. Verify the version-specific defaults. On Varnish 6.x the default workspace_client is 64 kB. On 7.0+ it is 96 kB. If you recently upgraded Varnish but not the explicit parameter, or if you migrated to a host with more threads, the effective memory pressure changed even if the workspace value did not.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
MAIN.ws_client_overflowDirect count of client workspace allocation failures (V6.2+)Any sustained nonzero rate
MAIN.client_resp_500All 500 responses delivered to clients; broader than workspace-onlyNonzero rate with healthy backends and cache hits
MAIN.losthdrHeaders dropped because http_max_hdr was exceededAny nonzero value; indicates too many distinct headers, not just raw size
MAIN.ws_backend_overflowSame failure on the backend response side; response headers too large for workspace_backendNonzero rate; fix is workspace_backend, not workspace_client
Process RSS vs configured storageConfirms workspace increases did not push Varnish toward OOMRSS growing after raising workspace_client or thread_pool_max
Cookie header size distributionLeading indicator before overflow triggersP99 cookie length approaching workspace_client budget

Fixes

Raise workspace_client

The immediate remediation is to increase workspace_client:

# Runtime change (persists until restart; add to startup params for permanence)
varnishadm param.set workspace_client 128k

Note: workspace_client has the DELAYED_EFFECT flag — it can be changed at runtime but the new value takes effect for newly created connections, not existing ones.

This works, but it multiplies. workspace_client is allocated per worker thread. The real memory cost is workspace_client multiplied by thread_pool_max multiplied by thread_pools. With thread_pool_max at its default of 5000 per pool and 2 pools, raising workspace_client from 96 kB to 128 kB adds roughly (128 - 96) * 1024 * 5000 * 2, approximately 320 MB, to steady-state memory consumption. On a host where Varnish storage is already sized close to available RAM, this can push the process toward OOM.

Before applying a larger workspace value in production, check the arithmetic against actual memory headroom:

varnishadm param.show workspace_client
varnishadm param.show thread_pool_max
varnishadm param.show thread_pools

Verify the change resolved the overflow by watching ws_client_overflow and client_resp_500 drop to zero rate.

The more durable fix is to stop the oversized data from reaching Varnish’s workspace:

  • Strip cookies that Varnish does not need for caching or forwarding. In vcl_recv, unset analytics, A/B testing, and feature-flag cookies before the request enters the hash and deliver phases. This directly shrinks the workspace footprint.
  • Cap or normalize X-Forwarded-For. In deep proxy chains, each hop appends an IP. If Varnish is several hops in, the header can dominate workspace. Normalize it to the client IP plus one trusted proxy.
  • Review regsuball and synthetic header logic. Each set req.http.X = regsuball(...) allocates workspace for the result. If VCL builds large strings (for example, reconstructing a cookie jar or assembling a forwarding header), that allocation counts against the budget. Prefer unset over rewrite where possible.
  • Shorten URLs. The URL is already in workspace by the time vcl_recv runs, so early return(pass) does not reduce its footprint. If a specific endpoint accepts arbitrarily long query strings, work with the application to move long parameters into the request body.

Adjust http_max_hdr separately

If losthdr is the incrementing counter, the problem is header count, not header size. Raise http_max_hdr:

varnishadm param.show http_max_hdr
varnishadm param.set http_max_hdr 96

Each additional header slot consumes a small amount of workspace, so this interacts with workspace_client. If you raise both, account for the combined memory cost.

Consider http_req_overflow_status (Varnish 7.x)

Varnish 7.x exposes http_req_overflow_status, which controls the HTTP status returned when http_req_size is exceeded. The default is 0, meaning Varnish closes the connection silently. Setting it to 400 or 414 makes the failure explicit to the client and to your logs, which helps distinguish oversized-request rejections from genuine workspace exhaustion:

varnishadm param.show http_req_overflow_status

Note: http_req_overflow_status applies only to the http_req_size limit (oversized request bodies), not to workspace_client overflow. These are distinct failure modes.

This does not fix workspace overflow, but it surfaces a related class of oversized-request failures that would otherwise look like silent connection drops.

Prevention

  • Monitor ws_client_overflow, client_resp_500, and losthdr continuously. These counters are frequently unmonitored. Alert on any sustained nonzero rate.
  • Track cookie header size as a leading indicator. If P99 cookie length is creeping toward the workspace_client budget, you will hit overflow before the counter fires. Capture this via varnishlog sampling or a log pipeline.
  • Size workspace_client deliberately, not by accident. The default changed between 6.x and 7.0. If you rely on defaults, know which default you are on. If you set it explicitly, document the memory arithmetic (workspace_client times thread_pool_max times thread_pools) alongside the value.
  • Review VCL changes for workspace pressure. A VCL reload that adds regsuball or synthetic header logic can push borderline requests over the edge. Treat the first 500 after a reload as a signal to check ws_client_overflow.
  • Distinguish client-side from backend-side overflow. ws_backend_overflow increments when backend response headers exceed workspace_backend. The fix is workspace_backend, not workspace_client, and the failure mode is different. On some versions, backend workspace overflow silently drops headers instead of failing the request, which is worse for correctness because it produces a 200 with missing headers rather than a visible 500.

How Netdata helps

  • Per-second collection of ws_client_overflow, client_resp_500, and losthdr lets you see the exact second the overflow rate began, narrowing the correlation window with deploys, traffic spikes, or bot activity.
  • Correlating workspace overflow counters with cache_hit and cache_miss rates confirms whether the 500s are hitting cacheable traffic (where header size is the issue) or pass traffic (where VCL logic is the amplifier).
  • Tracking threads and thread_pool_max alongside workspace_client gives the memory arithmetic context needed to decide whether raising the workspace is safe or will push the process toward OOM.
  • Process RSS monitoring catches the memory consequence of a workspace increase before the OOM killer does.
  • Anomaly detection on cookie and header volume patterns surfaces a creeping increase in header size before it crosses the workspace threshold.