The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / varnish / varnish-session-close-reasons ▌

Operations Guides

Varnish session close reasons: reading the sc_* counters

The MAIN.sc_* counters in Varnish record why every client session ended. Each close reason is a DIAG-level counter visible through varnishstat. The distribution across these counters is one of the fastest triage signals available: it tells you whether sessions are ending normally or whether the client, the network, or Varnish itself is causing problems.

The sc_* family has grown across versions. Some counters shifted accounting between releases, and several error counters are benign under certain traffic patterns. Treating every nonzero error counter as an incident leads to alert fatigue; treating them all as noise means missing real problems.

The sc_* family

Every Varnish client session ends with a close reason. Varnish records these as counters prefixed MAIN.sc_, each incremented when a session terminates for that cause. The counters are classified at DIAG level:

varnishstat -1 -f 'MAIN.sc_*'

On a healthy system, the vast majority of closes are normal client-initiated disconnects. A shift toward error counters is the signal to investigate.

Normal close reasons

These counters reflect expected session lifecycle and should dominate the distribution.

CounterClose reasonWhat it means
sc_rem_closeClient closedClient closed the TCP connection. The most common close reason on any healthy system.
sc_req_closeClient requested closeClient sent Connection: close, ending the session after the request. Normal for HTTP/1.0 clients or clients that disable keepalive.
sc_tx_pipePiped transactionSession ended after a pipe transaction. Normal for traffic handled via return(pipe) in VCL.
sc_tx_eofEOF transmissionSession closed after Varnish sent an EOF to the client. Normal for streaming or chunked responses that complete naturally.
sc_resp_closeBackend/VCL requested closeVarnish closed the session because the backend or VCL indicated Connection: close.

A high count of sc_rem_close is not a problem. Clients close connections when done, especially mobile clients and browser connection pools. What matters is the ratio of normal closes to error closes, and whether previously zero error counters start incrementing.

Close reasons that warrant investigation

These counters indicate something went wrong during the session. Each points to a different layer of the stack.

sc_overload: Varnish ran out of a resource

The official description is “Out of some resource.” This is a Varnish-side problem. The most common trigger is thread pool exhaustion: no worker thread was available to handle the session, and the queue was full.

When sc_overload is nonzero and sustained, correlate with:

  • MAIN.threads at or near thread_pool_max * thread_pools
  • MAIN.thread_queue_len greater than zero
  • MAIN.threads_limited incrementing

If these are elevated, either increase thread_pool_max or address the upstream cause (slow backends holding threads, traffic spike exceeding pool capacity).

sc_rx_bad: malformed request from the client

The client sent a request that Varnish could not parse as valid HTTP. Causes include buggy clients, protocol mismatches, or deliberate attack traffic such as request smuggling, fuzzing, or buffer overflow attempts.

varnishlog -q 'BogoHeader or HttpGarbage'

A low, steady rate is normal internet background noise from scanners and bots. A sudden spike above baseline warrants source IP analysis. Varnish is strict in HTTP parsing compared to some other proxies, so some legitimate but non-compliant clients may trigger this counter.

sc_rx_timeout and sc_rx_close_idle: idle or slow clients

These counters were split in Varnish 6.4. Before 6.4, all idle timeouts were counted as sc_rx_timeout. After 6.4:

  • sc_rx_close_idle counts sessions where timeout_idle was exceeded while waiting for a client request. This is the idle-keepalive case: the client opened a connection, sent a request, then sat idle past the configured timeout (default 5 seconds).
  • sc_rx_timeout now counts timeouts during active request receiving. The client started sending a request but stalled partway through.

sc_rx_close_idle is usually the dominant error counter on a healthy system. Clients open keepalive connections and abandon them. This is benign. Increasing timeout_idle reduces sc_rx_close_idle but keeps idle connections alive longer, consuming file descriptors.

sc_rx_timeout in the post-6.4 meaning is more interesting. A client started but did not finish a request. This can indicate slow uploads, broken clients, or slowloris-style attacks. Correlate with timeout_req, the parameter governing how long Varnish waits for a complete request.

sc_tx_error: error transmitting to the client

Varnish encountered an error while writing the response to the client. The most common cause is a broken pipe: the client disconnected while Varnish was still sending the response body.

This is often normal on mobile networks and unreliable connections. A moderate sc_tx_error rate is expected for traffic with a significant mobile client base.

Important version change: starting in Varnish 6.4, send_timeout events are reported as sc_tx_error instead of sc_rem_close. If you upgraded from pre-6.4 and see sc_tx_error increase, check whether send_timeout is appropriate for your workload. The Varnish 6.4 upgrading guide warns about this for HTTP/1 clients with long-running backend fetches: the client connection may time out while Varnish waits for the backend, and the close is now accounted as sc_tx_error.

sc_rapid_reset and sc_bankrupt: HTTP/2 abuse

CounterWhat it detects
sc_rapid_resetHTTP/2 Rapid Reset attack pattern (CVE-2023-44487): a client rapidly opens and resets streams to overwhelm the server.
sc_bankruptHTTP/2 credit bankruptcy: all streams were waiting for flow-control window credits when h2_window_timeout triggered. CVE-2024-30156 is the related Broke Window credit-exhaustion attack.

On a healthy system, both should be zero. A sustained nonzero sc_rapid_reset rate indicates a possible DDoS attack against your HTTP/2 endpoint. Version availability differs: VSV00013 added rapid-reset protection in 7.3.1/7.4.2/6.0.12, while the credit-exhaustion mitigation appeared in 7.3.2/7.4.3/6.0.13 and h2_window_timeout was generalized in 7.5.

Other error close counters

CounterWhat it meansWhen to worry
sc_rx_junkReceived junk data that could not be parsedSpike indicates scanning or protocol abuse
sc_rx_overflowReceived buffer overflowSpike indicates oversized requests or attack
sc_rx_bodyFailure receiving request bodyClient disconnected during POST/PUT or network issue
sc_pipe_overflowSession pipe buffer overflowPiped request exceeded pipe buffer limits
sc_range_shortInsufficient data for requested byte rangeClient requested a range beyond object size
sc_req_http10Session used HTTP/1.0 protocolInformational; normal for legacy clients
sc_req_http20HTTP/2 not acceptedVarnish rejected an HTTP/2 preface because HTTP/2 was disabled or not accepted on that listener
sc_vcl_failureVCL failure during sessionCorrelate with MAIN.vcl_fail; indicates VCL runtime error

Version changes that shift counter accounting

Three version transitions materially change how the sc_* distribution looks. If you compare data across these boundaries, the numbers are not directly comparable.

Varnish 6.4:

  • MAIN.sess_drop was removed entirely. Any monitoring referencing it must be updated.
  • Idle timeouts split from sc_rx_timeout into the new sc_rx_close_idle. Pre-6.4 idle timeout counts will appear as sc_rx_timeout.
  • send_timeout events reclassified from sc_rem_close to sc_tx_error. This is the most impactful change for operators who treated sc_rem_close as “all normal.” The sc_tx_error counter will increase after upgrade even if traffic behavior is unchanged.

Varnish 6.6:

  • Fixed the close reason to properly report sc_resp_close where previously only sc_req_close was reported. After upgrading from pre-6.6, sc_req_close will appear to drop and sc_resp_close will appear to rise. This is the fix taking effect, not a behavior change in your traffic.

Varnish 7.3:

  • VXID format changed to 64-bit. In-memory and on-disk VSL format is not compatible with previous versions. Log dumps from prior releases are unreadable. This does not affect counter semantics but affects log analysis tooling.

Varnish 7.5:

  • Added sc_rapid_reset and sc_bankrupt for HTTP/2 attack mitigations.

Reading the distribution as a triage tool

The value of the sc_* family is not in any single counter but in the distribution. When diagnosing client-reported problems or investigating an anomaly, the distribution tells you where to look first.

flowchart TD
    A["Session close distribution"] --> B{"Dominant counters?"}
    B -->|"sc_rem_close, sc_req_close,
sc_tx_pipe, sc_tx_eof"| C["Normal:
client-initiated closes"] B -->|"Error counters elevated"| D{"Which error counter?"} D -->|"sc_overload"| E["Varnish-side:
check thread pool, queue"] D -->|"sc_rx_bad, sc_rx_junk,
sc_rx_overflow"| F["Client-side:
malformed HTTP or attack"] D -->|"sc_rx_timeout,
sc_rx_close_idle"| G["Idle/slow clients:
check timeout_idle, timeout_req"] D -->|"sc_tx_error"| H["Transmit error:
check send_timeout, mobile baseline"] D -->|"sc_rapid_reset,
sc_bankrupt"| I["HTTP/2 abuse:
DDoS or credit exhaustion"] D -->|"sc_vcl_failure"| J["VCL error:
check MAIN.vcl_fail"]

Triage logic:

  1. Normal counters dominate, error counters near zero. System is healthy.
  2. sc_overload incrementing. Problem is on the Varnish side. Check thread pool saturation, queue length, and resource limits before looking at client behavior.
  3. sc_rx_bad or sc_rx_junk spikes. Problem is on the client side. Determine whether it is a broken client or deliberate attack by examining source IPs and request patterns in varnishlog.
  4. sc_rx_close_idle dominates error counters. Usually benign idle keepalive behavior. Investigate only if the rate is consuming file descriptors or if timeout_idle is set so high that idle connections accumulate.
  5. sc_tx_error spikes after a pre-6.4 upgrade. Verify that the increase is due to send_timeout reclassification rather than a real network problem. Compare with the pre-upgrade sc_rem_close rate.
  6. sc_rapid_reset nonzero. Treat as a potential security incident and investigate HTTP/2 traffic patterns.
  7. sc_vcl_failure incrementing. Correlate with MAIN.vcl_fail. The compiled VCL hit a runtime error during session processing. Check for a recent VCL reload.

Monitoring sc_* counters with Netdata

Netdata collects all sc_* counters with per-second resolution. This matters during traffic spikes or attacks where the distribution shifts in seconds.

When sc_overload spikes, the Varnish collector shows MAIN.threads, MAIN.thread_queue_len, and MAIN.threads_limited in the same dashboard, confirming whether the overload is thread-pool-driven. Per-second rates mean a brief sc_rapid_reset burst is visible before it becomes sustained. Historical retention lets you compare the distribution before and after a version upgrade, so a reclassification artifact does not look like a phantom regression.

Anomaly detection flags deviations from the learned baseline for each counter, which helps catch gradual increases in sc_rx_bad or sudden spikes in sc_tx_error without manual threshold tuning.