The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kubernetes / kubernetes-api-server-rate-limited

Operations Guides

Kubernetes API Server Rate Limited: How To Fix It

Your API server is running. /healthz returns 200. /readyz passes. Yet nodes drop to NotReady, the scheduler stops placing pods, and controller logs fill with context deadline exceeded. The cluster is not down, but it is frozen. This pattern often points to API Priority and Fairness (APF) starvation: low-priority traffic consumes the API server’s concurrency budget, and critical control plane requests queue or get rejected.

APF is enabled by default in Kubernetes 1.20+. It classifies every API request into a priority level via FlowSchema rules, then schedules requests against a per-level concurrency limit. When a priority level exhausts its seats, requests queue. If the queue fills, the server returns HTTP 429. When the queue grows in system or leader-election, kubelets cannot renew leases, controllers cannot write status, and the cluster degrades from the inside out. This guide shows how to confirm APF starvation, identify the culprit, and fix the allocation without turning the API server into a free-for-all.

What This Means

APF replaces the older global --max-requests-inflight and --max-mutating-requests-inflight flat limits with a fair-queuing system. Two CRDs control APF:

  • FlowSchema assigns incoming requests to a priority level based on user, verb, resource, namespace, or source.
  • PriorityLevelConfiguration defines the concurrency share and queue size for each level.

The default levels include exempt (no limits, including for system:masters), system (for nodes), leader-election, workload-high (built-in controllers), workload-low (general authenticated traffic), and catch-all (unauthenticated or unmatched traffic).

Each non-exempt level receives an effective concurrency limit proportional to its nominalConcurrencyShares relative to the total shares across all levels. For example, in a cluster where the server concurrency limit is 600 and total shares are 100, a level with 10 shares gets an effective limit of 60 concurrent requests.

Starvation happens when lower-priority traffic, such as a runaway operator or CI pipeline, consumes its own seats plus any available headroom, leaving system or leader-election without capacity. Those critical requests then sit in queue until they time out or are rejected. The symptoms look like a slow control plane, but the root cause is distribution, not total volume.

Common Causes

CauseWhat it looks likeFirst thing to check
Runaway controller or operatorworkload-low executing at 100% of its limit; system queue depth risingapiserver_flowcontrol_current_executing_requests by priority_level
Insufficient concurrency for critical levelsleader-election or system queues grow during normal loadprioritylevelconfigurations shares and total cluster shares
Misconfigured FlowSchemaKubelet or controller traffic classified into catch-allflowschemas matching rules for critical users
Thundering herd after recoveryAll priority levels show queue spikes simultaneouslyRequest rate by flow schema
Global API server saturationapiserver_current_inflight_requests near the hard limit; 429s across all levelsGlobal inflight vs --max-requests-inflight

Quick Checks

# Check APF queue depth by priority level
kubectl get --raw /metrics | grep apiserver_flowcontrol_current_inqueue_requests

# Check concurrency utilization per priority level
kubectl get --raw /metrics | grep apiserver_flowcontrol_current_executing_requests

# Check APF rejected requests by priority level and flow schema
kubectl get --raw /metrics | grep apiserver_flowcontrol_rejected_requests_total

# Check 429 rate from the API server
kubectl get --raw /metrics | grep 'apiserver_request_total.*code="429"'

# Check global inflight requests
kubectl get --raw /metrics | grep apiserver_current_inflight_requests

# View current APF configuration
kubectl get prioritylevelconfigurations -o custom-columns=NAME:.metadata.name,CONCURRENCY:.spec.limited.nominalConcurrencyShares
kubectl get flowschemas

A healthy cluster shows zero sustained queue depth in system and leader-election, a 429 rate near zero, and inflight requests well below the hard limit.

How To Diagnose It

  1. Confirm APF is actively throttling. Check apiserver_flowcontrol_rejected_requests_total and apiserver_request_total{code="429"}. If 429s are present, APF is the bottleneck. If absent, look at etcd latency or admission webhooks instead.

  2. Identify which priority levels are queuing. Check apiserver_flowcontrol_current_inqueue_requests by priority_level. Any sustained queue depth in system or leader-election is critical. Queueing in workload-low or catch-all is expected under load and is APF working as designed.

  3. Find the level consuming all concurrency. Compare apiserver_flowcontrol_current_executing_requests against apiserver_flowcontrol_request_concurrency_limit for each priority level. If workload-low is at 100% while system is queuing, a noisy neighbor is starving critical traffic.

  4. Pinpoint the specific client or flow. Use apiserver_flowcontrol_rejected_requests_total broken down by flow_schema, or inspect audit logs for the user-agent and username generating the flood. A single flow schema dominating the request count indicates a runaway controller, aggressive CI job, or misconfigured operator.

  5. Distinguish local saturation from global overload. Check apiserver_current_inflight_requests. If global inflight is well below --max-requests-inflight and --max-mutating-requests-inflight but APF is rejecting traffic, the issue is share misallocation. If inflight is at the global limit, the server is universally overloaded.

  6. Correlate with downstream impact. Check node Ready conditions and controller logs. If kubelets miss heartbeats or the scheduler times out on leader election, APF starvation is already causing cluster-wide degradation. This confirms urgency.

flowchart TD
    A[Runaway controller floods workload-low] --> B[APF concurrency exhausted]
    B --> C[System and leader-election requests queue]
    C --> D[Kubelet heartbeats delayed]
    D --> E[Nodes marked NotReady]
    C --> F[Controller updates timeout]
    F --> G[Scheduling and reconciliation lag]

Metrics & Signals To Monitor

SignalWhy it mattersWarning sign
APF queue depth (system / leader-election)Critical control plane traffic is waiting instead of executingQueue depth > 0 sustained for more than 30 seconds
APF rejected requests (system / leader-election)Critical traffic is being dropped with 429Any non-zero rate in these levels
APF concurrency utilization per levelHow close each level is to its effective limit> 80% of limit sustained
429 response rateActive throttling by APF> 5% of total API requests
Inflight requests (mutating / read-only)Global API server saturation> 80% of --max-requests-inflight or --max-mutating-requests-inflight
Controller timeout errorsDownstream impact of queue delayscontext deadline exceeded in controller or kubelet logs

Fixes

If The Cause Is A Runaway Controller Or Operator

Identify the offending client from audit logs or the flow_schema label on rejected requests. Throttle the client at the source: add client-side rate limits, reduce polling frequency, or fix the reconciliation loop. Do not raise APF limits to absorb bad behavior; the client will keep growing until it hits the next ceiling.

If The Cause Is Insufficient Concurrency Shares

Edit the PriorityLevelConfiguration for system and leader-election to increase nominalConcurrencyShares. Remember that shares are relative to the total across all levels; increasing shares for one level reduces the effective limit of others unless you also raise the server concurrency limit via --max-requests-inflight and --max-mutating-requests-inflight. Ensure the API server’s CPU, memory, and etcd backing can handle the additional load before raising global limits.

If The Cause Is Misconfigured Flow Schemas

Ensure that kubelet, controller-manager, and scheduler traffic match dedicated high-priority flow schemas. The default schemas cover built-in components, but custom controllers or infrastructure agents often fall into catch-all. Create specific FlowSchema resources for these components, matching on their service account or user group, and assign them to workload-high or a custom high-priority level.

If The Cause Is A Thundering Herd

If the traffic is legitimate but bursty, add jitter to client retry logic and ensure exponential backoff respects 429 responses. Temporarily increasing concurrency shares can provide relief, but the permanent fix is client behavior.

If Cluster Stability Is At Risk

As a last resort, you can temporarily move a critical service account to the exempt priority level. This bypasses all queuing and can destabilize the API server if the client floods requests. Revert immediately after recovery. Long-term exemptions defeat the purpose of APF.

Prevention

  • Review APF configuration quarterly and after adding major operators. New controllers change the request mix.
  • Monitor system and leader-election queue depth as a leading indicator, not a lagging one.
  • Ensure every critical controller has a dedicated FlowSchema resource. Do not let important traffic fall into catch-all.
  • Document which service accounts and user groups each FlowSchema matches. Stale selectors silently reclassify traffic after deployments change.
  • Set client-side rate limits and backoff on all custom controllers and automation.
  • Test APF behavior under load. A deployment of 1,000 replicas should not push workload-low into a state that starves leader-election.

How Netdata Helps

  • Correlate APF queue depth with API server request latency to distinguish queuing delays from etcd latency.
  • Alert on sustained queue depth in system or leader-election before nodes transition to NotReady.
  • Track 429 spikes alongside etcd disk latency and webhook latency to isolate the true bottleneck.
  • Visualize per-priority-level concurrency utilization to spot noisy neighbors before they cause cluster-wide impact.
  • Monitor controller workqueue depth as a downstream signal that APF throttling delays reconciliation.
The Netdata solution

Kubernetes monitoring with Netdata

Netdata monitors Kubernetes with per-second metrics across the control plane, nodes, and every pod, with ML anomaly detection and zero per-pod configuration. Correlate API-server and etcd latency, kubelet PLEG stalls, scheduling pressure, and OOMKills in one place.