The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kubernetes / kubernetes-iptables-sync-stall

Operations Guides

Kubernetes kube-proxy iptables sync stall: causes and recovery

Pods fail to start. Services intermittently route traffic to dead endpoints. The kube-proxy health endpoint still returns HTTP 200, so the DaemonSet looks healthy, yet rules drift further behind with every sync cycle.

An iptables sync stall is not a crash. It is a slowdown or blockage in the control loop that translates Service and EndpointSlice state into kernel NAT rules. When kube-proxy cannot acquire the global xtables lock, when iptables-restore hangs, or when the rule set grows too large to reconcile within the sync period, the node forwards packets using stale rules. New endpoints are invisible. Terminated pods still receive connections. CNI plugins that also need the xtables lock time out, and pod sandbox creation fails.

This guide covers failure mechanics, distinguishing sync stalls from other network problems, and recovery steps.

What this means

kube-proxy in iptables mode reconciles desired state by building a complete set of iptables NAT rules and applying them atomically via iptables-restore. This requires the global xtables lock (/run/xtables.lock). While kube-proxy holds that lock, no other process on the node can modify iptables rules.

The default syncPeriod is 30 seconds. Kubernetes considers kube-proxy unhealthy on the :10256/healthz endpoint when network programming exceeds twice that value, or 60 seconds by default. A stall begins when a single sync takes long enough to miss the next cycle, when lock contention causes repeated restore failures, or when the process hangs entirely inside iptables-restore.

Since Kubernetes v1.28, the iptables proxier uses incremental updates, rewriting only rules for changed Services and EndpointSlices. On older versions, or at extreme scale, every sync may still rewrite the full table. Either way, the symptoms are identical: the gap between API server state and kernel state grows, and the node silently drops or misroutes connections.

flowchart TD
    A[EndpointSlice change or syncPeriod tick] --> B{Can kube-proxy acquire xtables lock?}
    B -->|Yes| C[Run iptables-restore]
    B -->|No| D[Sync aborts or retries]
    C --> E{Sync duration < syncPeriod?}
    E -->|Yes| F[Rules updated]
    E -->|No| G[Sync backlog grows]
    D --> H[Stale endpoints persist]
    G --> H
    H --> I[CNI plugins timeout on lock]
    H --> J[Traffic routed to dead pods]
    I --> K[Pod startup fails or slows]
    J --> L[Connection resets or drops]

Common causes

CauseWhat it looks likeFirst thing to check
xtables lock contentionsync_proxy_rules_iptables_restore_failures_total increasing; kube-proxy logs show lock acquisition failures; pods stuck in ContainerCreating with CNI timeoutslsof /run/xtables.lock and kube-proxy logs for xtables lock
Rule bloat (large clusters)kubeproxy_sync_proxy_rules_duration_seconds p99 approaching or exceeding 30s; thousands of iptables rulesiptables -t nat -S | wc -l
iptables-restore hangSync duration flatlines; iptables-restore process visible in ps for minutes; kube-proxy stops advancing its last sync timestampps aux | grep iptables-restore and sync timestamp metric
Endpoint churn exceeding sync capacitykubeproxy_sync_proxy_rules_endpoint_changes_pending growing during rolling updatesPending endpoint changes vs processed total
API server watch death (silent)Last sync timestamp frozen; no error logs; new Services unreachable from the nodess -tnp | grep kube-proxy | grep 6443 and sync timestamp age

Quick checks

# Check last successful sync age (should be well under 60s)
curl -s http://localhost:10249/metrics | grep kubeproxy_sync_proxy_rules_last_timestamp_seconds

# Check p99 sync duration against the 30s syncPeriod
curl -s http://localhost:10249/metrics | grep kubeproxy_sync_proxy_rules_duration_seconds

# Count kube-proxy managed iptables rules
sudo iptables -t nat -S | grep -c "KUBE-"

# Check for lock contention messages
kubectl logs -n kube-system -l k8s-app=kube-proxy | grep -i "xtables\|lock"

# Check pending endpoint changes
curl -s http://localhost:10249/metrics | grep endpoint_changes_pending

# Check iptables restore failure rate
curl -s http://localhost:10249/metrics | grep sync_proxy_rules_iptables_restore_failures_total

# Identify what holds the xtables lock
sudo lsof /run/xtables.lock

# Check conntrack utilization (adjacent shared resource)
echo "scale=2; $(cat /proc/sys/net/netfilter/nf_conntrack_count) * 100 / $(cat /proc/sys/net/netfilter/nf_conntrack_max)" | bc

# Check kube-proxy binary health (200 does not mean fresh rules)
curl -s -o /dev/null -w "%{http_code}" http://localhost:10256/healthz

# Check CNI timeouts that correlate with lock contention
journalctl -u kubelet --since "5 minutes ago" | grep -i "CNI.*timeout\|sandbox"

How to diagnose it

  1. Confirm kube-proxy is running but stale. Query http://localhost:10256/healthz. A 200 means the initial sync completed, not that rules are current. If traffic fails while healthz returns 200, you have a silent stall.

    • Why it matters: Restarting a crashed kube-proxy differs from unsticking a live one.
    • Next: check the last sync timestamp.
  2. Measure sync freshness. Inspect kubeproxy_sync_proxy_rules_last_timestamp_seconds and subtract it from the current epoch. A gap greater than twice the configured syncPeriod means the node is programming rules too slowly or not at all.

    • Why it matters: This distinguishes a transient spike from a sustained stall.
    • Next: inspect sync duration percentiles.
  3. Compare sync duration to syncPeriod. Look at the p99 of kubeproxy_sync_proxy_rules_duration_seconds. In iptables mode, values above 10 seconds are concerning; values approaching the 30-second syncPeriod indicate the loop is about to backlog.

    • Why it matters: Once sync duration exceeds syncPeriod, kube-proxy can never catch up on a busy node.
    • Next: determine if the cause is lock contention or rule bloat.
  4. Check for xtables lock contention. Search kube-proxy logs for messages about the xtables lock. Run lsof /run/xtables.lock to see which process holds it. Common culprits include CNI portmap plugins, Flannel, and VPC CNIs that refresh rules periodically.

    • Why it matters: The lock is global. If a CNI plugin holds it, kube-proxy blocks, and vice versa.
    • Next: check CNI plugin logs for iptables timeouts.
  5. Assess rule set scale. Count rules with iptables -t nat -S | wc -l. In iptables mode, performance degrades linearly as rule count grows. Above 10,000 rules, sync times become a scaling bottleneck.

    • Why it matters: This tells you whether the stall is architectural (iptables mode limits) or environmental (lock contention).
    • Next: if rule count is high, evaluate proxy mode migration.
  6. Look for hung iptables-restore processes. Run ps aux | grep iptables-restore. If a process has been running for minutes or hours, the nf_tables backend may be stuck. The documented case is iptables v1.8.4 with the nf_tables backend, where nf_tables CHAIN_USER_DEL failures (“Device or resource busy”) cause indefinite hangs. This is an iptables-userspace bug dating to 2019, not a kube-proxy defect.

    • Why it matters: A hung restore blocks the entire sync loop until the process is killed.
    • Next: kill the hung process and restart kube-proxy.
  7. Check endpoint churn backlog. Query kubeproxy_sync_proxy_rules_endpoint_changes_pending. If the number is non-zero and growing, the API server is generating changes faster than the node can apply them.

    • Why it matters: This happens during large rolling updates or HPA storms.
    • Next: slow the churn or increase sync capacity.
  8. Verify API server watch connectivity. Check rest_client_requests_total for 5xx or 429 errors, and verify with ss that kube-proxy has an active TCP connection to the API server on port 6443.

    • Why it matters: A dead watch causes silent staleness with no iptables errors.
    • Next: restart kube-proxy to re-establish the watch.
  9. Rule out conntrack exhaustion. Check nf_conntrack_count against nf_conntrack_max. Sync stalls often occur alongside connection churn that fills the conntrack table, producing identical timeout symptoms.

    • Why it matters: Fixing kube-proxy does not help if the kernel is dropping packets because the conntrack table is full.
    • Next: if utilization is above 90%, increase the limit or reduce connection churn.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
kubeproxy_sync_proxy_rules_duration_seconds p99Measures time to reconcile all iptables rulesp99 > 10s, or > 80% of syncPeriod
Age of kubeproxy_sync_proxy_rules_last_timestamp_secondsIndicates how stale the programmed rules areAge > 2 x syncPeriod
kubeproxy_sync_proxy_rules_endpoint_changes_pendingBacklog of unprocessed endpoint updatesNon-zero and growing
sync_proxy_rules_iptables_restore_failures_total rateDirect signal of lock contention or restore failureAny sustained increase
iptables rule countScaling indicator in iptables mode> 10,000 rules
nf_conntrack_count / nf_conntrack_maxShared kernel resource consumed by NAT> 80% utilization
kube-proxy healthz / readyzBinary readiness stateNon-200, or readyz returning 503
rest_client_requests_total errorsAPI server connectivity from kube-proxy5xx or 429 responses
CNI sandbox creation latencyDownstream symptom of xtables lock contentionTimeouts during pod creation
Pod startup duration on nodeEnd-to-end impact of stalled rulesp99 > 60s

Fixes

If the cause is xtables lock contention

Identify the competing process with lsof /run/xtables.lock. If it is a CNI plugin or daemon, restart it and reduce its iptables refresh frequency. As a longer-term fix, evaluate migrating the cluster to nftables mode or IPVS. Note that IPVS is deprecated as of Kubernetes v1.35; nftables is positioned as its replacement. IPVS uses hash-based lookups and does not hold the global xtables lock during updates. nftables mode uses per-table locking, which reduces contention between kube-proxy and CNI plugins.

If the cause is rule bloat or high sync duration

Audit the cluster for unnecessary Services and large EndpointSlices. If the cluster has grown beyond the comfortable limit for iptables mode, increase the syncPeriod temporarily to allow full syncs to complete, then plan a migration to nftables. Do not increase sync frequency; that worsens the problem.

If the cause is a hung iptables-restore

Kill the hung iptables-restore process, then delete the kube-proxy pod to force a restart and full re-sync. If the node runs iptables v1.8.4 with the nf_tables backend, upgrade iptables or the host image.

If the cause is endpoint churn exceeding capacity

Reduce deployment rollout surge or HPA scale-out rate. Spread large deployments across time windows. If the cluster is legitimately high-churn, move to nftables mode, which supports incremental updates without rewriting the entire table.

If conntrack is exhausted

Immediately increase nf_conntrack_max to buy headroom. Then identify whether the root cause is a connection leak, excessive UDP traffic, or overly long TIME_WAIT timeouts. Tune nf_conntrack_udp_timeout_stream for UDP-heavy workloads.

Prevention

  • Alert on kubeproxy_sync_proxy_rules_duration_seconds p99 crossing 5 seconds (or 25% of your configured syncPeriod), not just on kube-proxy restarts.
  • Collect and alert on kube-proxy log messages containing xtables lock to catch contention before it causes sync failures.
  • Monitor iptables rule count per node and establish a runway projection. Plan a migration to nftables before reaching 10,000 rules.
  • Size nf_conntrack_max for peak traffic plus headroom. Account for TIME_WAIT and UDP entries, not just established TCP connections.
  • If you run Kubernetes v1.28 or later, revisit legacy minSyncPeriod tunings that were previously used to mitigate full-table rewrite overhead. Incremental updates make large values unnecessary and can delay convergence.
  • Exercise failure modes in staging: terminate kube-proxy watches, simulate high endpoint churn, and measure sync latency under load.

How Netdata helps

  • Correlate sync duration with node CPU softirq time, CNI timeouts, and conntrack utilization.
  • Alert on sync timestamp age without manual metric scraping across nodes.
  • Flag xtables lock contention from kube-proxy logs.
  • Track iptables rule count trends to forecast scaling limits.
The Netdata solution

Kubernetes monitoring with Netdata

Netdata monitors Kubernetes with per-second metrics across the control plane, nodes, and every pod, with ML anomaly detection and zero per-pod configuration. Correlate API-server and etcd latency, kubelet PLEG stalls, scheduling pressure, and OOMKills in one place.