The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-sess-dropped-req-dropped ▌

Operations Guides

Varnish sess_dropped vs req_dropped: HTTP/1 connection drops and HTTP/2 stream drops

Two Varnish counters track the worst outcome a cache can produce: a client gets nothing. MAIN.sess_dropped counts HTTP/1 sessions dropped because the worker thread queue was full. MAIN.req_dropped counts HTTP/2 streams and other request types dropped for the same reason. Both mean Varnish refused to serve a client because no thread was available and the bounded queue was at its limit.

The distinction matters because HTTP/1 and HTTP/2 multiplex differently. An HTTP/1 client occupies a connection. An HTTP/2 client multiplexes many concurrent streams over a single connection. When Varnish runs out of threads, the rejection mechanism differs by protocol, and so does the counter that increments. Teams that monitor only sess_dropped can see zero drops while HTTP/2 streams are silently refused.

What these counters mean

CounterWhat it countsProtocolVersion scope
MAIN.sess_droppedHTTP/1 sessions (connections) dropped because the queue was too longHTTP/1Standard in modern Varnish
MAIN.req_droppedHTTP/2 streams and other request types dropped because the queue was too longHTTP/2 and otherAll supported releases (counter present since the 5.x series)

Both are cumulative counters. They only increase until the child process restarts, at which point all MAIN.* counters reset to zero. Monitor the rate of increase (delta per second), not the absolute value. A counter reading 50,000 tells you nothing without knowing whether those drops accumulated over a year or over the last 30 seconds.

Both counters increment when the session queue reaches thread_queue_limit. New arrivals are dropped instead of queued. The queue depth is visible as MAIN.thread_queue_len, a gauge that updates once per second.

Check both counters in one command:

# Check both drop counters simultaneously
varnishstat -1 -f MAIN.sess_dropped -f MAIN.req_dropped

HTTP/1 vs HTTP/2: why there are two counters

Varnish’s concurrency model is thread-per-request with a bounded pool. When a client request arrives and no worker thread is available, the request enters a bounded queue. When that queue hits thread_queue_limit, the request is dropped.

The protocol determines what “dropped” means:

  • HTTP/1: Each connection carries one request at a time. Dropping a session means closing the TCP connection. The client sees a connection reset or refused. This increments MAIN.sess_dropped.
  • HTTP/2: A single TCP connection multiplexes multiple concurrent streams. Varnish can refuse an individual stream without closing the connection. The client sees a stream reset (RST_STREAM) while the connection itself remains open. This increments MAIN.req_dropped.

A drop at the connection level (HTTP/1) and a drop at the stream level (HTTP/2) are fundamentally different events, even though the root cause (no thread available, queue full) is the same.

flowchart TD
    A[Client request arrives] --> B{Worker thread available?}
    B -- Yes --> C[Request processed]
    B -- No --> D[Enters session queue
thread_queue_len] D --> E{Queue at thread_queue_limit?} E -- No --> F[Waits for worker] F --> C E -- Yes --> G{Protocol} G -- HTTP/1 --> H["sess_dropped increments
connection reset"] G -- HTTP/2 --> I["req_dropped increments
stream reset"]

The legacy sess_drop naming trap

If you have been running Varnish for years, you may know the counter as MAIN.sess_drop (no trailing “ped”). This is the older spelling. It was removed from the codebase in Varnish Cache 6.4.0 (the 6.4 changelog notes “The MAIN.sess_drop counter is gone”). The current counter is MAIN.sess_dropped.

The trap works in two directions:

  1. Monitoring a dead counter. If your dashboards or alerts reference MAIN.sess_drop, they show zero forever. The counter either does not exist (removed) or is never incremented (deprecated but present). Meanwhile, HTTP/1 connections are being dropped and counted under sess_dropped.

  2. Assuming the name change was cosmetic. req_dropped was not introduced with the 6.4 removal of sess_drop — it has existed alongside it since the 5.x series, so the drop of the old spelling did not change what is counted. The pair (sess_dropped, req_dropped) covers both protocols. Monitoring sess_dropped alone still misses HTTP/2 drops.

If you are unsure which counter name your Varnish version exposes:

# List available session/request drop counters
varnishstat -1 | grep -E 'sess_dropp|req_dropp|sess_drop'

What a drop means for the client

Each drop is a hard failure. The client receives nothing: no cached response, no 503 error page, no synthetic response from VCL. The connection or stream is reset.

This is different from a 503 response. A 503 means Varnish accepted the request, processed it through VCL, attempted a backend fetch (or decided to synthesize an error), and returned a response. A drop means Varnish never got that far. The request never entered VCL processing because no thread was available to handle it.

Dropped requests do not appear in varnishlog or varnishncsa output. There is no transaction log entry because the request was never assigned to a worker thread. The counters are the only detection mechanism. If you rely solely on log-based monitoring, you have a blind spot for exactly the failure mode that represents the most severe user impact.

The monitoring mistake: watching only one

With HTTP/2, overload manifests as req_dropped (stream drops), not sess_dropped (connection drops). Teams monitoring only sess_dropped miss HTTP/2 traffic loss entirely. This is common because environments that originally deployed monitoring for HTTP/1-only traffic still have those dashboards. When TLS termination was added in front of Varnish (via Hitch, HAProxy, or similar) with ALPN negotiation for HTTP/2, the traffic profile changed but the monitoring did not.

The fix is to monitor both counters as a combined drop rate:

# Combined drop rate (take two readings, compute delta)
varnishstat -1 -f MAIN.sess_dropped -f MAIN.req_dropped

Alert on the combined rate (sess_dropped rate + req_dropped rate) sustained above zero. Any nonzero sustained drop rate with active traffic means Varnish is refusing clients.

For the full set of common Varnish monitoring gaps, see the Varnish monitoring checklist.

Reading the rate, not the cumulative counter

Both sess_dropped and req_dropped are monotonically increasing counters that reset only on child restart. A single snapshot of the raw value is useless for alerting. You need the rate of change.

To compute the rate manually, take two readings and divide by the interval:

# Manual rate check: two readings, 10 seconds apart
T1=$(varnishstat -1 -f MAIN.sess_dropped -f MAIN.req_dropped)
sleep 10
T2=$(varnishstat -1 -f MAIN.sess_dropped -f MAIN.req_dropped)
echo "$T1"; echo "$T2"
# Subtract per-counter values, divide each delta by 10 for drops/sec

For production alerting, use a monitoring system that computes per-second rates automatically. The key conditions for a page-worthy alert:

  • Combined drop rate (sess_dropped + req_dropped) sustained greater than 0 for more than 120 seconds.
  • MAIN.uptime greater than 300 seconds (excludes the warmup window after restart, where brief drops can occur as the thread pool ramps from thread_pool_min).
  • MAIN.client_req greater than 0 (confirms live traffic is present, preventing false fires on idle nodes).
  • MAIN.thread_queue_len greater than 0 (confirms queue saturation as the cause).

Zero is the only acceptable sustained value. Any nonzero combined rate, sustained past the warmup window, indicates a fault that is actively refusing client traffic.

Correlating with thread pool saturation

Drops are the terminal symptom. Thread pool saturation is the cause. When you see drops incrementing, the following chain has already occurred:

  1. All worker threads are busy (typically blocked on slow backend responses).
  2. New requests enter the queue (thread_queue_len rises above zero).
  3. The queue reaches thread_queue_limit.
  4. New arrivals are dropped.

The diagnostic signals to correlate:

SignalWhat it tells youWhy it matters
MAIN.thread_queue_lenCurrent queue depth (gauge, updates once per second)Leading indicator. When this rises, drops follow.
MAIN.threadsCurrent total worker threadsIf equal to thread_pool_max * thread_pools, the pool is at capacity.
MAIN.threads_limitedCount of times thread creation hit thread_pool_maxIncrementing means Varnish wanted more threads but the config prevented it.
MAIN.threads_failedCount of times the OS refused thread creationIncrementing means a system-level limit (ulimit, memory, cgroup) is blocking thread creation. This is worse than threads_limited.
MAIN.client_reqClient request rateConfirms traffic is present. Drops on an idle server are suspicious.
# Full thread pool saturation snapshot
varnishstat -1 -f MAIN.threads -f MAIN.thread_queue_len \
  -f MAIN.threads_limited -f MAIN.threads_failed -f MAIN.pools \
  -f MAIN.sess_dropped -f MAIN.req_dropped

If threads_limited is incrementing alongside drops, the pool is too small for the traffic volume. If threads_failed is incrementing, the OS is refusing to create threads (check ulimit -u, available memory, and cgroup limits). Both require different remediation: the first needs a thread_pool_max increase, the second needs OS-level configuration changes.

For a deeper treatment of thread pool exhaustion diagnosis and remediation, see the thread pool exhaustion guide.

Distinguishing capacity exhaustion from attack

Not all drops mean your capacity is too low. A sudden spike in drops without a corresponding increase in legitimate traffic can indicate an attack.

The CVE-2024-30156 “Broke Window Attack” (VSV00014) targets HTTP/2 flow control in Varnish Cache. All releases with HTTP/2 support are affected except the fixed ones: 7.3.2, 7.4.3, 7.5.x, and 6.0 LTS 6.0.13 onward. An attacker exhausts HTTP/2 connection flow control credits, causing denial of service. Detection includes spikes in sess_dropped or req_dropped without matching legitimate traffic increase.

To distinguish capacity exhaustion from attack:

  • Capacity exhaustion: client_req rate is elevated, thread_queue_len is high, backend response times are elevated (slow backends holding threads). threads_limited increments steadily. The drop rate correlates with traffic volume.
  • Attack pattern: Drop rate spikes without proportional increase in client_req. Backend response times may be normal. HTTP/2-specific abuse counters (MAIN.sc_rapid_reset, MAIN.sc_bankrupt) may also increment. The drop pattern may be bursty or sustained regardless of backend health.

Additional HTTP/2 security signals to check:

# HTTP/2 protocol abuse counters
varnishstat -1 -f MAIN.sc_rapid_reset -f MAIN.sc_bankrupt -f MAIN.req_reset

All three counters exist under MAIN.. sc_rapid_reset ships with the Rapid Reset mitigation (7.3.1, 7.4.2, 6.0 LTS 6.0.12) and detects the CVE-2023-44487 pattern. sc_bankrupt ships with the Broke Window mitigation (7.3.2, 7.4.3, 6.0 LTS 6.0.13) and counts HTTP/2 sessions failed because all streams were waiting for flow-control window credits when h2_window_timeout triggered. req_reset counts requests whose client went away before VCL processing completed — for HTTP/2, a stream reset by the client’s RST_STREAM frame or a stream/connection error.

Cold start and warmup considerations

During the first few minutes after a child process restart, brief session drops can occur. The thread pool ramps from thread_pool_min, and if a traffic spike hits during warmup with a conservative thread_pool_add_delay, temporary drops are possible. These are transient and self-resolve.

This is why the alerting threshold includes MAIN.uptime > 300. The five-minute window excludes the warmup period from alerting. The MGT.* counters (management process) do not reset on child restart, but all MAIN.* counters do. Rate calculations that span a child restart will produce incorrect values for that interval.

After restart, expect the cache hit rate to start at zero and climb as the cache warms. Drops during warmup with an empty cache are particularly likely because every request is a cache miss, every miss requires a backend fetch, and each fetch holds a thread longer than a cache hit would.

Correlating drops in Netdata

Netdata collects these counters at per-second resolution. For drop diagnosis, the useful correlations are:

  • Protocol-specific drop rates. sess_dropped and req_dropped rate charts show whether losses are concentrated in HTTP/1 or HTTP/2. This is the signal that catches the “monitoring only sess_dropped” blind spot.
  • Queue depth timeline. thread_queue_len charts alongside drop rates, making the causal chain visible. Per-second collection catches sub-second spikes that the Varnish gauge itself (which updates once per second) can smooth over.
  • Thread pool limits vs OS refusals. threads_limited and threads_failed alongside drops distinguish a config ceiling from an OS-level limit without running separate varnishstat invocations.
  • HTTP/2 abuse counters. sc_rapid_reset and sc_bankrupt appear alongside drop rates when attack patterns are involved.