The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / clickhouse / clickhouse-authentication-failures

Operations Guides

ClickHouse authentication failures: system.session_log, brute force, and credential drift

You notice a spike in failed connection attempts to ClickHouse: a security scanner flags repeated TCP 9000 probes, or an application logs connection timeouts after a secrets rotation. ClickHouse exposes authentication events through system.session_log, but only if the feature is enabled. Without it, fallback to server error logs and system.query_log exceptions.

The failures split into two patterns. Malicious: brute-force or credential-scanning campaigns against exposed TCP 9000 or HTTP 8123. Operational drift: a rotated password not updated in a client config, or a deployment shipping an old connection string. Distinguish them fast. Block an external attacker at the network layer; fix credential drift on the client.

This guide shows how to read system.session_log, correlate failures with connection counts and network exposure, and fix the root cause.

What this means

system.session_log records session lifecycle events, including LoginFailure entries with the username, source IP (client_address), authentication type, and failure reason. More than ten LoginFailure events per minute from one IP is a strong brute-force signal. Any failure from a new source IP needs tracing to a known application or user.

If session_log is disabled, the same events may appear in ClickHouse server logs or as AUTHENTICATION_FAILED exceptions in system.query_log. These are harder to aggregate and may lack source-address detail. Regardless of source, failures correlate with network exposure. If listen_host is bound to 0.0.0.0 or a public interface, any host that can reach ports 9000 and 8123 is in the attack surface.

flowchart TD
    A[Auth failures detected] --> B{session_log enabled?}
    B -->|No| C[Enable session_log or use query_log fallback]
    B -->|Yes| D[Aggregate by client_address and user]
    D --> E{Single IP > 10/min?}
    E -->|Yes| F[Brute force or scanning]
    E -->|No| G{Service account failing?}
    G -->|Yes| H[Credential drift]
    G -->|No| I[Check network exposure and client configs]

Common causes

CauseWhat it looks likeFirst thing to check
Brute force or credential scanning> 10 LoginFailure events per minute from one IP; many distinct usernamessystem.session_log aggregated by client_address
Credential rotation driftSteady failures from a known app server or service accountuser and client_address in system.session_log; secrets manager sync status
Application misconfigurationFailures begin after a deployment; usually one userDeployment timeline and user in system.query_log
Overly permissive network bindingExternal IPs reaching TCP 9000 or HTTP 8123 at allss -tlnp output for listen_host
Misconfigured monitoring probesRegular, low-rate failures from internal infra hostsSource IP of monitoring checkers against known probe config

Quick checks

Run these read-only checks to characterize the failure pattern without changing any state.

-- Recent authentication failures from session_log
SELECT event_time, user, client_address, auth_type, failure_reason
FROM system.session_log
WHERE type = 'LoginFailure'
  AND event_time > now() - INTERVAL 1 HOUR
ORDER BY event_time DESC;
-- Aggregate failures by source IP to detect brute force
SELECT
    client_address,
    user,
    count(*) AS failures,
    max(event_time) AS last_failure
FROM system.session_log
WHERE type = 'LoginFailure'
  AND event_time > now() - INTERVAL 10 MINUTE
GROUP BY client_address, user
HAVING failures > 10
ORDER BY failures DESC;
-- Fallback: authentication errors from query_log
SELECT event_time, user, exception, query_id
FROM system.query_log
WHERE exception LIKE '%AUTHENTICATION_FAILED%'
  AND event_time > now() - INTERVAL 1 HOUR
ORDER BY event_time DESC;
# Check network exposure: what interfaces is ClickHouse bound to
ss -tlnp | grep clickhouse
-- Check connection volume for context
SELECT metric, value
FROM system.metrics
WHERE metric IN ('TCPConnection', 'HTTPConnection');

If system.session_log does not exist or returns no rows, the feature is not enabled. Enable it in the server configuration to capture these events.

How to diagnose it

  1. Confirm the event source. Query system.session_log for type = 'LoginFailure'. If the table is empty, use the system.query_log fallback with exception LIKE '%AUTHENTICATION_FAILED%'. Note that query_log lacks the precise source-address detail of session_log.

  2. Identify the failure pattern. Aggregate by client_address and user. A single IP with more than ten failures per minute suggests brute force or automated scanning. A single service account failing from a known application host suggests credential drift.

  3. Correlate with changes. Check whether the onset of failures aligns with a recent deployment, secrets rotation, or infrastructure change. Credential drift almost always starts within minutes of a password or key rotation.

  4. Audit network exposure. Run ss -tlnp | grep clickhouse and inspect the bound addresses. If ClickHouse is listening on 0.0.0.0 or a public interface and you see brute-force attempts from external IPs, the immediate priority is reducing that exposure.

  5. Review server error logs. Check the ClickHouse server log file for connection failure details. On standard Linux installations this is /var/log/clickhouse-server/clickhouse-server.log. Look for unknown user, wrong password, or protocol mismatch messages.

  6. Map internal failures to consumers. For operational drift, filter session_log by the failing user and map client_address to known applications or hosts. Verify the connection strings and credentials in the corresponding secrets manager or configuration store.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
system.session_log LoginFailure rateCaptures every failed authentication event with source IP and reason> 10 failures per minute from one IP
system.query_log AUTHENTICATION_FAILEDFallback when session_log is disabled; tracks exceptions across all queriesSustained failures from service accounts
Client connection countDistinguishes brute force from connection leaks or retry stormsSpike in TCPConnection matching auth failure times
Network interface bindingUnnecessary exposure invites scanning and widens the blast radiusListening on 0.0.0.0 or public interfaces

Fixes

Brute force or credential scanning

Block the source IP at your firewall, cloud security group, or reverse proxy. Do not rely on ClickHouse for rate limiting. If ClickHouse is directly exposed because of a permissive listen_host, restrict it to internal interfaces or specific addresses. If the source is an internal misconfigured health check, fix the checker instead of blocking the IP.

Credential drift after rotation

Identify the failing user from system.session_log.user. Update the password or key in the application’s connection string, environment variable, or secrets manager. Restart or reload the client to clear cached credentials. Verify by watching LoginFailure entries for that user stop. If old and new credentials overlap during rotation, revoke the old credential to prevent silent fallback.

Application misconfiguration

Correlate the start of failures with a deployment timestamp. Roll back if ongoing, or patch the configuration. Use client_address from system.session_log or user and event_time in system.query_log to identify the emitting host if session detail is insufficient.

Missing session_log coverage

If system.session_log is disabled, failed authentication events are invisible to native SQL audit. Enable it in the ClickHouse server configuration. Until then, use system.query_log and server error logs as fallbacks.

Prevention

  • Enable and retain system.session_log.
  • Bind ClickHouse to specific internal interfaces via listen_host; audit with ss -tlnp after any configuration change.
  • Store credentials in a secrets manager and automate rotation with application restarts or hot-reload.
  • Monitor for LoginFailure spikes as an infrastructure security signal, not just a database issue.
  • Run periodic audits of active users and their expected source IP ranges.

How Netdata helps

Netdata collects ClickHouse TCPConnection and HTTPConnection metrics and query error rates. Correlate connection spikes with error-rate jumps to distinguish brute-force scans from client misconfiguration. Set alerts on unusual connection counts or error rates to catch authentication issues without manual polling.

The Netdata solution

ClickHouse monitoring with Netdata

Netdata monitors ClickHouse with per-second metrics and ML anomaly detection. Track merge debt, memory usage, replication lag, Keeper/ZooKeeper saturation, and disk headroom against the host signals that drive them.