The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-thread-pool-tuning ▌

Operations Guides

Varnish thread pool tuning: thread_pool_min, thread_pool_max, and thread_pools

Varnish uses a thread-per-request concurrency model with a bounded pool. Every client request occupies a worker thread for its entire lifecycle, from accept through response delivery. When the pool is full and the overflow queue overflows, sessions are dropped.

The three parameters that control this behavior, thread_pools, thread_pool_min, and thread_pool_max, determine how Varnish responds to load spikes, slow backends, and idle periods. The dominant factor in thread pool sizing is backend response time, not request rate. A cache miss that holds a thread for 2 seconds consumes the same pool capacity as thousands of sub-millisecond cache hits. Slow backends are the single most common cause of thread pool exhaustion, and they make thread_pool_max the parameter that matters most under stress.

This guide covers how the three parameters interact, how to recognize misconfiguration from runtime metrics, and how to size pools against your actual backend latency profile rather than your peak request rate.

How the thread pool works

Varnish’s child process maintains one or more thread pools. Each pool is an independent set of worker threads with its own lock. When a client connection arrives, the acceptor hands it to an idle worker from one of the pools. If no worker is available, Varnish either creates a new thread (if below the per-pool maximum) or queues the request. If the queue overflows, the session is dropped.

CPU can be idle, memory can be plentiful, and Varnish will still refuse connections if all worker threads are blocked waiting on slow backend responses. This is the most common Varnish outage pattern: thread pool exhaustion with idle resources.

The pool grows by creating new threads, but the rate of creation is governed by thread_pool_add_delay. In Varnish 6.0+, the default is 0 seconds, meaning threads are created as fast as needed. In older versions, the default was 2 milliseconds between thread creations, which meant a sudden traffic spike could cause temporary session drops while the pool ramped up.

Threads above thread_pool_min that sit idle for longer than thread_pool_timeout (default 300 seconds) are destroyed. If idle traffic oscillates around thread_pool_min, you will see constant create/destroy cycling in the thread counters, which wastes CPU and indicates the minimum is set too low for your traffic floor.

flowchart TD
    A[Connection arrives] --> B{Idle worker?}
    B -->|yes| C[Process request]
    B -->|no| D{Below pool max?}
    D -->|yes| E[Create thread]
    E -->|gated by add_delay| C
    D -->|no| F{Queue space?}
    F -->|yes| G[Queue request]
    G --> C
    F -->|no| H[Drop session]
    C --> I{Idle above min?}
    I -->|timeout| J[Destroy thread]
    I -->|active| K[Keep in pool]

The three core parameters

ParameterDefaultWhat it controlsFlags
thread_pools2Number of independent thread poolsDelayed; decreases need restart
thread_pool_min100Minimum idle threads kept alive per poolDelayed
thread_pool_max5000Maximum threads per pool (hard ceiling)Delayed

thread_pools: pool count and lock contention

Each thread pool has its own mutex. More pools means less lock contention on multi-core systems. The default of 2 is adequate for most deployments. Do not exceed one pool per CPU core; beyond that, lock overhead outweighs the parallelism benefit.

Increasing thread_pools can be done at runtime via varnishadm param.set. Decreasing it requires a restart to take effect.

Total worker threads across the entire Varnish process equals thread_pools multiplied by thread_pool_max. With defaults (2 pools, 5000 max each), the theoretical ceiling is 10,000 threads.

thread_pool_min: the idle floor

This is the minimum number of threads each pool always maintains. When traffic drops, threads above this count are destroyed after thread_pool_timeout seconds of idleness. When traffic picks up again, threads are created back up to demand, rate-limited by thread_pool_add_delay.

If your idle traffic baseline requires 300 concurrent threads, setting thread_pool_min to 100 means Varnish will destroy 200 threads during quiet periods, then recreate them when traffic returns. The result is visible as churn in MAIN.threads_created and MAIN.threads_destroyed. Set thread_pool_min high enough that idle traffic does not cause the pool to shrink below what the next traffic pulse will immediately need.

thread_pool_max: the ceiling

This is the hard limit on threads per pool. When the pool hits this ceiling and all threads are busy, new requests queue. When the queue overflows, sessions are dropped.

The default of 5000 per pool is adequate for most workloads with fast backends and good cache hit rates. With slow backends, each thread is held longer, reducing effective concurrency. A backend with 500ms P99 TTFB that serves 10,000 cache misses per second needs at least 5,000 threads just for backend fetches, assuming misses are evenly distributed across time.

Each thread consumes stack memory. The thread_pool_stack default is 64KB on 6.0 LTS, 56KB on 6.4–6.6, and 80KB on 7.0 and later (64-bit; smaller defaults on 32-bit). At the default ceiling of 2 pools times 5000 threads, that is approximately 800MB for thread stacks on 7.x (640MB with the 64KB default of 6.0 LTS). This is real memory that cannot be used for cache storage. Setting thread_pool_max higher than necessary wastes RAM without improving throughput.

When thread_pool_max is the active bottleneck, MAIN.threads_limited increments. This counter is the definitive signal that the ceiling is too low for the current workload.

Several additional parameters interact with the three core knobs:

  • thread_pool_add_delay: Controls how fast threads ramp up under load. Default is 0 seconds in Varnish 6.0+. Setting this too high (even 10ms) causes slow ramp-up under spikes, leading to temporary session drops during warmup after a restart or sudden load increase.
  • thread_pool_timeout: How long excess idle threads survive before destruction. Default is 300 seconds. Lower values make the pool shrink faster, increasing create/destroy churn. Higher values keep threads alive longer, reducing ramp-up latency but holding memory longer.
  • thread_pool_stack: Per-thread stack size. Default is 64KB on 6.0 LTS, 56KB on 6.4–6.6, 80KB on 7.0+. Total stack memory equals thread_pools * thread_pool_max * thread_pool_stack. This interacts directly with thread_pool_max for memory budgeting.
  • thread_queue_limit: Maximum queued requests per pool before drops. Default is 20. This is the last buffer before MAIN.sess_dropped or MAIN.req_dropped starts incrementing.
  • thread_pool_reserve: Available in every supported release, including 6.0 LTS. Reserves threads for vital internal tasks to prevent lower-priority work from starving critical operations. Default is 0, which auto-tunes to 5% of thread_pool_min; minimum is 1 otherwise, maximum 95% of thread_pool_min.

Sizing against backend TTFB, not request rate

The most common tuning mistake is sizing thread_pool_max against peak request rate without accounting for how long each request holds a thread. The number that matters is concurrent in-flight requests, not requests per second.

For cache hits, a thread is held for microseconds. For cache misses, a thread is held for the full backend fetch duration: time to first byte plus body transfer time. If your backend P99 TTFB is 1 second and your cache miss rate produces 2,000 concurrent misses at peak, you need at least 2,000 threads just for miss traffic. Cache hit traffic adds negligible thread demand by comparison.

This is why slow backends cause thread pool exhaustion while CPU sits idle. The threads are not doing work; they are blocked waiting on backend responses. Increasing thread_pool_max without addressing backend performance just delays the cliff.

To measure backend TTFB:

# Backend time-to-first-byte for recent fetches
varnishlog -g request -i Timestamp -q 'BerespStatus gt 0'
# Look for Bereq (request sent) and Beresp (response header received) timestamps

Varnish does not expose latency as varnishstat counters. You must use varnishlog or varnishncsa for timing data.

Diagnosing pool exhaustion from metrics

SymptomLikely causeCounter to check
MAIN.threads at thread_pools * thread_pool_maxPool at capacity, backend latency likelyMAIN.threads, MAIN.threads_limited
Constant threads_created / threads_destroyed churnthread_pool_min too low for idle trafficMAIN.threads_created, MAIN.threads_destroyed
Slow ramp after restart or traffic spikethread_pool_add_delay too conservativeMAIN.threads during warmup window
thread_queue_len sustained above zerothread_pool_max too low or backends too slowMAIN.thread_queue_len, MAIN.threads_limited
sess_dropped or req_dropped incrementingPool exhausted, queue overflowMAIN.sess_dropped, MAIN.req_dropped
threads_failed above zeroOS refusing thread creation (ulimit, memory)MAIN.threads_failed
Threads at max with idle CPUBackend latency holding threadsBackend TTFB via varnishlog

The key diagnostic command for pool state:

# Check thread pool state at a glance
varnishstat -1 -f MAIN.threads -f MAIN.thread_queue_len \
  -f MAIN.threads_limited -f MAIN.threads_failed -f MAIN.pools

If MAIN.threads equals thread_pools times thread_pool_max and MAIN.threads_limited is incrementing, the pool is at capacity. If MAIN.thread_queue_len is also nonzero, drops are imminent or already happening. Check MAIN.threads_failed separately: a nonzero value means the OS is blocking thread creation (ulimits, cgroup memory limits), which is a system-level problem rather than a Varnish tuning problem.

For alerting, monitor these counters at per-second resolution:

CounterAlert condition
MAIN.thread_queue_lenSustained nonzero for >10 seconds
MAIN.threads_limitedAny nonzero rate
MAIN.threads_failedAny nonzero value
MAIN.sess_dropped / MAIN.req_droppedAny sustained nonzero rate
MAIN.threads_created / MAIN.threads_destroyedConstant nonzero rate when traffic is stable

Runtime tuning: what you can change live

Most thread pool parameters can be changed at runtime without restarting the child process:

# Show current values
varnishadm param.show thread_pool_max
varnishadm param.show thread_pool_min
varnishadm param.show thread_pools

# Increase the per-pool ceiling live
varnishadm param.set thread_pool_max 8000

# Raise the idle floor to reduce create/destroy churn
varnishadm param.set thread_pool_min 200

The exception is thread_pools: increasing it works at runtime, but decreasing requires a restart. All changes made via varnishadm param.set are ephemeral and will be lost on restart. Persist them in your startup configuration using -p flags or your systemd unit.

Changes to thread_pool_max take effect immediately for new thread creation. Existing threads are not affected.

How Netdata helps

Netdata collects MAIN.threads, MAIN.thread_queue_len, MAIN.threads_limited, and MAIN.threads_failed per second. This resolution matters for thread pool saturation, which can spike and recover inside a 15-second scrape interval and produce no visible trace in coarser monitoring.

Correlating thread count with MAIN.sess_dropped and MAIN.req_dropped on the same timeline shows exactly when pool exhaustion translates to user-visible drops. Tracking threads_created and threads_destroyed rates reveals create/destroy cycling caused by an undersized thread_pool_min, which is hard to spot from cumulative counters alone.

Anomaly detection on thread_queue_len flags the transition from brief queueing to sustained saturation before drops begin. Per-second collection of backend health (VBE.*.happy) alongside thread metrics lets you distinguish “backends are slow” from “pool is too small” in a single view, because the two produce identical symptoms (threads at max) but need different fixes.