The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere
VMWARE VSPHERE · OPERATIONS PLAYBOOK

vSphere's invisible taxes: ready time the guest can't see, a memory cliff, and a vCenter that goes dark all at once

A two-plane stack where the ESXi scheduler deschedules vCPUs behind the guest's back, memory reclamation escalates from balloon to swap in minutes, one datastore stalls every VM on it, and the vCenter appliance — a constellation of 30-odd services on an embedded database — fails hard the moment a certificate expires or a /storage partition fills. We trace how that design behaves under load, where degradation turns into an outage, and what to do when it does.

"

vSphere's defaults get you virtualized quickly, then hand you a set of cliff-edges that most teams only meet during an incident — and half of them live in vCenter, not on the hosts.

The defaults work. Until a VM is sized with more vCPUs than the scheduler can co-schedule, ready time climbs, and the guest reports low CPU while everything runs slow. Until a host crosses its memory line and reclamation cascades from ballooning to compression to .vswp swap, whose disk I/O then drags down every VM on the datastore. Until a forgotten snapshot's delta disk fills the datastore and power-ons fail with No space left on device. Until a host PSODs and HA races to restart its VMs. And until the vCenter side breaks: an STS signing certificate expires and every login fails at once, or /storage/db fills and vPostgres stops, taking the whole management plane with it.

These guides are written for engineers who already run vSphere, not for people learning what a hypervisor is. The goal is the mental model of how ESXi and vCenter actually behave under load, the failure patterns that keep recurring on both planes, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident.

How vSphere actually runs in production

vSphere is two planes, not one. The ESXi data plane schedules vCPUs onto pCPUs, reclaims memory in tiers, and moves every I/O through the VMkernel; the vCenter management plane orchestrates HA, DRS, and vMotion on top of an appliance database. Most production failures live between these layers — a symptom in the guest can originate at the scheduler, the datastore, or a certificate three layers away.

01
VMs + VMware Tools
Each VM is a sandbox with emulated hardware. VMware Tools (or open-vm-tools) provides the balloon driver, the guest heartbeat, and time sync. Without Tools, the hypervisor is partly blind to the guest — and reclamation skips ballooning straight to swap.
GUEST
02
CPU scheduler + NUMA
The VMkernel maps vCPUs to pCPUs with a proportional-share scheduler. A vCPU that is runnable but waiting shows as <code>ready time</code>; multi-vCPU VMs pay <code>co-stop</code> to stay co-scheduled. The NUMA scheduler tries to keep each VM's memory local to its vCPUs.
SCHED
03
memory reclamation hierarchy
Four tiers, escalating in desperation: page sharing, then <code>ballooning</code> (guest pages internally), then <code>compression</code> into a cache, then host-level <code>swap</code> to a <code>.vswp</code> file. Each tier is an order of magnitude worse than the last. There is no graceful slope near the edge.
MEM
04
storage I/O path + datastores
Guest to virtual SCSI to the VMkernel queue to HBA/NFS client to the array. Latency decomposes into <code>KAVG</code> (VMkernel/queue) and <code>DAVG</code> (device/array). Datastores hold VMDKs, snapshot deltas, and swap files — and fill silently as thin disks and snapshots grow.
STORAGE
05
network I/O path
Guest to vNIC to port group on a vSwitch/dvSwitch to a physical uplink. A NIC team distributes flows but never aggregates one flow's bandwidth. VM, vMotion, storage, and management traffic often share the same wire unless separated.
NET
06
ESXi host (VMkernel)
The microkernel arbitrating all hardware. Its own health — hardware sensors, thermal state, driver stability — is the floor everything stands on. A PSOD kills every VM on the host instantly and triggers HA restarts elsewhere.
ESXi
07
cluster services: HA + DRS + vMotion
HA restarts VMs from failed hosts using network and datastore heartbeats; DRS rebalances placement every few minutes; vMotion live-migrates memory with a brief switchover stun. Misconfigured isolation response, admission control, or affinity rules turn these safety nets into incidents.
CLUSTER
08
vCenter management plane (vCSA)
The vCenter Server Appliance: <code>vpxd</code> holding an in-memory inventory cache, an embedded <code>vPostgres</code> database on dedicated <code>/storage/*</code> partitions, and <code>STS</code>/SSO issuing the tokens every login and integration depends on. Running VMs survive its outage, but DRS, HA reconfiguration, vMotion, and all management stop.
VCENTER

Why this matters: 'the VM is slow' can come from ready time the guest can't see, a memory reclamation cascade, one saturated datastore, a wide VM straddling NUMA nodes, a snapshot chain, or a vCenter that can't authenticate. The symptom rhymes but each layer has a different signal — and a different fix. And when vCenter itself is the problem, the VMs are usually fine; you've lost the ability to manage them.

The failures you'll actually see

Most vSphere incidents fall into a small set of recurring patterns, split across the ESXi data plane and the vCenter management plane. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The storage latency cliff

The array slows — a controller failover, a RAID rebuild, a noisy neighbour — and DAVG jumps from a few milliseconds to hundreds. The VMkernel queue backs up, KAVG follows, and every VM on the datastore stalls at once. Databases time out, guests report I/O errors, and heartbeats can drop. This is the number-one 'everything is slow' incident in vSphere, and it hits the whole datastore, not one VM.

  • DAVG spiking from under 5ms to 50-500ms on a datastore
  • KAVG rising and QUED sustained above zero (VMkernel queuing)
  • Every VM on the same datastore degraded simultaneously
  • Guest I/O wait high, application timeouts, possible heartbeat loss
Investigate
CRITICAL

The memory reclamation cascade

A host runs short of physical memory and reclamation escalates: ballooning forces guests to page internally, compression fills its cache, then the VMkernel swaps VM memory to .vswp files on the datastore. That swap I/O competes with real VM I/O, so latency spikes feed more memory pressure — the death spiral. It can go from fine to catastrophic in minutes, and a single node's pressure degrades every VM on it.

  • Balloon (MCTLSZ) inflating, then compression, then swap-out
  • Any sustained swap-IN rate (SWR/s) — active disk-speed memory access
  • Balloon at zero with swap active — VMware Tools not reclaiming
  • Datastore latency rising in step with swap file I/O
Investigate
IMMINENT

The snapshot time bomb

A snapshot is created — manually before a change, or by a backup job — and never removed. Its delta VMDK grows with every write, day after day, until the datastore fills. Then thin disks can't grow, swap files can't be created, power-ons fail, and consolidating the accumulated chain becomes a multi-hour, VM-stunning operation. The guest is completely unaware; performance just slowly gets worse.

  • Snapshot older than 24-72 hours and still growing
  • Datastore free space declining in step with one VM's write rate
  • Large *-delta.vmdk files in the VM folder
  • No space left on device / consolidation-needed warnings appearing
Investigate
ACTIVE

CPU starvation the guest can't see

A host is overcommitted, or a VM is oversized with more vCPUs than the scheduler can place at once. The vCPUs wait for physical CPUs — ready time and co-stop climb — but the guest OS has no idea it was descheduled, so in-guest CPU looks moderate. The result is the classic mystery: the VM is slow, guest CPU looks fine, and nobody can explain it. Adding vCPUs makes it worse.

  • CPU ready (%RDY) above 5-10% while guest CPU looks moderate
  • Co-stop (%CSTP) elevated on multi-vCPU VMs
  • Host CPU only moderate (60-80%), yet specific VMs are starved
  • Larger-vCPU VMs performing worse than smaller ones on the same host
Investigate
CRITICAL

The vCenter certificate expiry cascade

An internal certificate expires — most devastatingly the STS signing certificate, which is separate from the machine SSL cert and invisible in a browser. STS can no longer issue valid tokens, so every authentication fails at once: the vSphere Client won't log in, PowerCLI and every integration drop, and ESXi hosts disconnect in bulk. There is no gradual warning — everything works perfectly right up to expiry, then breaks together. VMCA certs default to a two-year life.

  • All logins failing simultaneously (not one account)
  • ESXi hosts disconnecting in bulk rather than one by one
  • SSL/TLS handshake errors across multiple services in the logs
  • checksts.py / vecs-cli showing a certificate at or past expiry
Investigate
CRITICAL

The vCenter disk death spiral

A dedicated vCSA partition fills — commonly /storage/db or /storage/log — while root looks healthy. When /storage/db fills, vPostgres can't write its WAL and crashes; vpxd has no database and the whole management plane goes down. When /storage/log fills, services that can't log crash, vmon restarts them, they log the failure — a self-reinforcing loop. The trigger is often a different issue (an SSO error storm, a failed purge) that flooded a partition.

  • df -h shows one /storage/* partition at 100% while / is fine
  • vPostgres or vpxd crash-looping (check vmon status)
  • Log generation rate paradoxically rising as services fail
  • SDK/API returning 503 or refusing connections
Investigate
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.

vSphere monitoring maturity levels

vSphere observability works in four practical levels, and each spans both the ESXi hosts and the vCenter appliance. Each level is a complete operation, not a stepping stone. Pick the one that matches how much your environment matters. Most production clusters should land at the second level.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: are the hosts and VMs alive, is there space to run, and is vCenter reachable? You won't learn what broke, but you'll learn that something broke before users do. Survival is enough for lab and non-critical clusters.

  • ESXi host up / VM power state Are the hosts reachable and are the VMs powered on?
  • Datastore free space (% and absolute) A full datastore halts every VM writing to it.
  • Host CPU and memory (aggregate) Blunt, but catches a host that is clearly saturated.
  • vCenter SDK / API reachable Can consumers actually talk to vCenter right now?
  • vCSA /storage/db and /storage/log The two partitions whose filling takes vCenter down.
  • Certificate expiry (machine SSL + STS) Expiry is a silent, total, cliff-edge outage.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production environments should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose an incident alone: CPU starvation, memory pressure, storage latency, snapshots, HA/DRS state, and the vCenter side.

  • CPU ready time per VM The starvation the guest can't see; the top CPU signal.
  • Memory balloon and swap (host + VM) Read the reclamation tiers, not one blended percentage.
  • Datastore latency (GAVG/DAVG/KAVG) Decompose it: array-slow vs VMkernel-queuing.
  • Snapshot age and count per VM Check daily — the #1 preventable vSphere incident.
  • HA cluster health + host connection Isolation, admission control, hosts Not Responding.
  • Per-partition vCSA disk + NTP offset All /storage/* mounts; clock skew breaks auth.
  • vpxd / vPostgres service status The two services whose loss is a full outage.
  • Certificate expiry for ALL cert types Not just the browser cert — the STS cert too.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. Co-stop creeping on an oversized VM, NUMA locality drifting, a datastore queue filling, the vPostgres database bloating, stats rollup falling behind. None of these page you on day one. They become page-out incidents on day thirty.

  • CPU co-stop and max-limited per VM Oversizing and forgotten CPU limits, invisible in %RDY.
  • NUMA locality per VM Wide VMs paying the remote-memory tax.
  • Memory compression + queue depth The tier before swap; QUED/ACTV before latency bites.
  • Per-vdisk latency and datastore IOPS Localise the noisy VM and the saturated LUN.
  • Dropped packets + vMotion stun time Ring/uplink pressure and connection-dropping pauses.
  • vPostgres database + SEAT table size Retention and statistics-level bloat, weeks early.
  • Task queue depth + stats rollup lag vpxd overload and rollup falling behind its interval.
  • vCSA VM seen from the hypervisor Ready time / balloon on vCenter itself — half the picture.

Level 4: Expert

Reactive instrumentation after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. SCSI sense codes, VMkernel world CPU, STS heap and GC, vmon restart counts, vPostgres autovacuum and WAL, VCHA replication. Most teams never need every signal here. Add the ones your incident history says you do.

  • SCSI sense codes + path state changes Classify storage errors; catch a flapping fabric path.
  • VMkernel world / system-time CPU NSX firewall and storage-driver overhead, not VM CPU.
  • STS Java heap + GC frequency The intermittent-SSO-failure root cause.
  • vmon service restart counts Silent crash loops that brief probes never catch.
  • vPostgres autovacuum + pg_wal size Dead-tuple bloat and unbounded WAL, 30-60 min early.
  • SDK session count by client IP The misbehaving backup job or script overloading vpxd.
  • VCHA / vmdir replication health Protection that silently isn't; ELM sites diverging.
  • Guest time drift + host power state NTP offset under contention; thermal throttling.

Operating mistakes worth avoiding

The traps vSphere teams keep falling into, on both the host and the vCenter side. Each has a clear, well-known fix. Most teams only learn it after an incident.

Monitoring guest CPU instead of hypervisor ready time

The single most common vSphere monitoring mistake. The in-guest agent reports utilisation of the time the VM <em>got</em>, not the time it <em>wanted</em> — so a VM at 30% guest CPU with 15% <code>ready time</code> looks healthy while running at 85% speed. The guest cannot see that it was descheduled. You must monitor CPU ready from the hypervisor level, per VM.

Treating memory as a single utilisation percentage

vSphere memory is a multi-tier reclamation system, not a number. A host at 85% 'memory utilisation' might be fine (large gap between consumed and active, no balloon) or actively swapping into a death spiral. The percentage tells you nothing on its own. Monitor the reclamation indicators independently: <code>balloon</code>, <code>compression</code>, and <code>swap</code>.

Not monitoring snapshots

Most vSphere setups have zero snapshot monitoring, yet forgotten snapshots are the number-one cause of preventable incidents. Teams discover two-week-old, 300GB snapshots when the datastore fills and VMs crash. Snapshot count and age must be checked <em>daily</em> — and a failed backup job is the classic source of an orphaned, ever-growing delta.

Reading datastore latency at the wrong layer

Teams that do watch storage latency often see 'high latency' and blame the array. But if <code>KAVG</code> is high and <code>DAVG</code> is normal, the problem is in the VMkernel (queue depth, VAAI locks, driver), not the array — and the diagnosis goes down the wrong path. Decompose latency into <code>DAVG</code> and <code>KAVG</code> before pointing anywhere.

Not monitoring vCenter as infrastructure

Many teams monitor thousands of VMs but never watch the vCSA's own disk partitions, database size, or service health. vCenter is the control plane: if it degrades, DRS stops, HA reconfiguration stops, provisioning stops, and during a host failure you may find HA can't be managed. Monitor vCenter like the critical infrastructure it is — including from the hypervisor layer.

Watching only the machine SSL certificate

The STS signing certificate is the most critical certificate in the vCSA ecosystem and has caused more outages than any other. It is separate from the machine SSL cert, invisible in a browser, and needs specific tooling (<code>checksts.py</code>, <code>vecs-cli</code>) to inspect. When it expires, all authentication fails at once. Track every certificate type, and alert 90 and 30 days out — not 7.

Monitoring disk on / but not the /storage/* partitions

The vCSA partition layout means root <code>/</code> can be perfectly healthy while <code>/storage/db</code>, <code>/storage/log</code>, or <code>/storage/seat</code> sits at 100%. Standard Linux monitoring that only checks <code>/</code> misses the most common vCenter disk failures entirely. Watch every <code>/storage/*</code> mount separately, and check inodes (<code>df -i</code>) on the log partition too.

Not monitoring CPU limits (max-limited)

A VM (or resource pool) with a forgotten MHz limit is throttled with <em>low</em> <code>ready time</code> but high <code>max-limited</code> time — so teams that only watch ready miss it. Limits get set during testing or inherited from a template and then silently cap production VMs for years. Monitor <code>cpu.maxlimited.summation</code> alongside <code>cpu.ready.summation</code>.

vSphere and vCenter runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up vSphere monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing — on the hosts or in vCenter.