The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph
CEPH · OPERATIONS PLAYBOOK

Ceph's cliff-edges: a cluster that throttles at 85%, stops writing at 95%, and heals only if you left it room

A distributed store where CRUSH places every object across OSDs, monitors hold the maps by Paxos majority, and the same disks and network carry client I/O, replication, recovery, and scrub all at once. We trace how that design behaves under load, where a single full OSD or a flapping daemon turns local trouble into a cluster-wide stop, and what to do when it does.

"

Ceph is self-healing until it isn't, and the line between the two is capacity you forgot to leave and flags you forgot to clear.

The defaults keep you alive. Until one OSD crosses the full ratio and Ceph refuses every write cluster-wide — a hard stop, not a slowdown, and reads keep working so it hides in plain sight. Until a cluster running hot loses a disk and recovery has nowhere to put the data, so backfill_toofull appears and degraded PGs stop healing. Until an OSD with a slow disk misses a heartbeat, gets marked down, triggers a peering storm across hundreds of PGs, and the storm slows its neighbours into missing heartbeats too. Until a monitor's clock drifts past 50ms and the quorum churns through elections while the whole cluster freezes. Until deep scrub finds a checksum mismatch on a weekend and the good replica's disk dies before anyone repairs it.

These guides are written for engineers who already run Ceph, not for people deciding whether to. The goal is the mental model of how RADOS, CRUSH, the OSDs, and the monitors actually behave under load, the failure patterns that keep recurring, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident.

How Ceph actually runs in production

Ceph is not one thing. It is a placement algorithm, a fleet of per-disk daemons, a Paxos control plane, and a local storage engine — with client I/O, replication, recovery, and scrub all competing for the same disks and network. Most production failures live between these layers, not inside any one of them.

01
clients + RADOS interfaces
Block (RBD), file (CephFS), and object (RGW / S3) interfaces all sit on top of RADOS. Every client independently computes where its data lives — there is no central lookup — and talks straight to the primary OSD for each object.
CLIENT
02
CRUSH placement
CRUSH deterministically maps objects into placement groups and PGs onto OSDs using a failure-domain hierarchy (host, rack, row). No lookup table, so any topology change recomputes placement and moves data. A wrong CRUSH rule can leave PGs unable to reach <code>active+clean</code> even with disks to spare.
CRUSH
03
placement groups (PGs)
The sharding and coordination unit. Objects hash into thousands of PGs; each PG has an acting set, a primary, and replicas, and lives in exactly one state. <code>active+clean</code> is health — every other state (<code>degraded</code>, <code>undersized</code>, <code>peering</code>, <code>incomplete</code>, <code>down</code>, <code>inconsistent</code>) is work or trouble.
PG
04
OSD daemons
One OSD per physical device. Each one simultaneously serves client I/O as primary, replicates to secondaries (or computes EC chunks), peers after topology changes, scrubs for integrity, runs recovery, and heartbeats with peers and monitors. Miss a heartbeat past <code>osd_heartbeat_grace</code> and peers mark it down.
OSD
05
BlueStore (device + DB/WAL)
The OSD's local engine, writing raw block devices with a RocksDB metadata store. Three devices may be in play: bulk data, a fast DB partition, and a WAL. When the DB device fills, metadata spills to the slow device and latency jumps by 10-100x — a cliff, invisible in capacity metrics.
STORE
06
monitors + cluster maps
Monitors hold the OSD, MON, CRUSH, PG, and MDS maps and agree on every update by Paxos majority. They are off the data path but critical during any topology change. Lose quorum and no map updates commit; clients coast on cached maps until those go stale, then I/O stalls.
MON
07
managers (MGR)
The active/standby pair running the Prometheus exporter, dashboard, balancer, and PG autoscaler. Almost every metric in these guides comes from the MGR module. If MGR fails, monitoring and some PG operations halt while client I/O keeps flowing — so a blind spot, not an outage.
MGR
08
capacity thresholds
The cluster-wide circuit breakers. At the nearfull ratio (0.85) Ceph warns and throttles; at backfillfull (0.90) it blocks recovery onto full OSDs; at full (0.95) it stops <em>all</em> writes. One OSD hitting full is the whole cluster's problem, and a full cluster cannot heal itself.
CAPACITY

Why this matters: 'Ceph is slow' or 'writes are failing' can come from a single full OSD, a flapping daemon triggering a peering storm, BlueStore metadata spilling to a slow disk, a recovery storm starving client I/O, a monitor quorum that keeps re-electing, or an OMAP storm on one RGW bucket index. The symptom rhymes but each layer has a different signal — and a different fix.

The failures you'll actually see

Most Ceph incidents fall into a small set of recurring patterns. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The write freeze

One OSD — or the cluster average — crosses the full ratio (default 0.95) and Ceph refuses all writes cluster-wide. It is a hard stop that cannot self-resolve, and because reads keep working it hides until applications start failing on write. A cluster left running hot gets there fast when losing an OSD subtracts its capacity from the total.

  • ceph_health_detail{name="OSD_FULL"} active, writes returning ENOSPC
  • Cluster or any OSD utilization at or above ceph_osd_full_ratio
  • One OSD near full while the cluster average looks moderate
  • backfill_toofull appearing as recovery runs out of room
Investigate
CRITICAL

Stuck and unavailable placement groups

After OSD losses, PGs cannot find enough authoritative data to serve I/O. incomplete means the write history is missing; down means no replica is available at all. Both fail reads and writes for the affected objects and will not self-resolve once peering has had time to finish.

  • sum(ceph_pg_incomplete) > 0 sustained past normal peering
  • sum(ceph_pg_down) > 0 — reads and writes to those PGs fail
  • PGs stuck in peering for more than ~10 minutes
  • Unfound objects blocking recovery of the affected PGs
Investigate
CRITICAL

Monitor quorum loss

Fewer than a majority of monitors agree, so no cluster map update can commit. Clients keep running on cached maps for a while, then new connections fail and I/O stalls as maps go stale. Clock skew, a MON-to-MON partition, a full MON disk, or one monitor on slow storage can all trigger the election churn behind it.

  • sum(ceph_mon_quorum_status) below floor(count/2)+1
  • ceph -s itself taking seconds to respond
  • MON election epoch incrementing, leader changing repeatedly
  • MON_CLOCK_SKEW active alongside the instability
Investigate
ACTIVE

The OSD flapping cascade

An OSD with a slow disk, network jitter, or memory pressure misses heartbeats and is marked down; its PGs peer and recover elsewhere; it returns and peering reverses; the underlying fault makes it miss again. Each flap mints an OSD map epoch every OSD must process, and the peering overhead can slow healthy OSDs into flapping too — the feedback loop that has killed more clusters than any other.

  • One or more OSDs alternating up/down in ceph osd tree
  • OSD map epoch incrementing several times per minute
  • SLOW_OPS appearing across many OSDs, not just one
  • PG states cycling through peering / recovering / active+clean
Investigate
ACTIVE

Slow requests and blocked I/O

Operations exceed osd_op_complaint_time (default 30s) and register as slow ops — stuck, not merely slow. Because a write waits for every replica to acknowledge, one slow OSD can block client I/O far beyond itself. Disk stalls, op-queue saturation, a blocked BlueStore kv_sync thread, or RocksDB compaction are the usual causes; deep scrub and recovery cause transient versions that must be ruled out.

  • ceph_healthcheck_slow_ops > 0 sustained past a couple of minutes
  • Commit or apply latency spiking on specific OSDs
  • Client latency elevated cluster-wide, not localised
  • Blocked ops concentrated behind a single primary OSD
Investigate
IMMINENT

Silent data corruption

Deep scrub finds a replica whose data does not match the others — bit rot, a firmware bug, non-ECC RAM, or a torn write. Ceph detected it, which is the system working as designed, but reads still succeed from a good copy so it is easy to ignore. If that copy's OSD fails before the inconsistency is repaired, the data is gone. A blind ceph pg repair can copy the corrupt primary over the good replica.

  • sum(ceph_pg_inconsistent) > 0 / OSD_SCRUB_ERRORS active
  • PG state active+clean+inconsistent, or PG_DAMAGED
  • ceph_pg_failed_repair > 0 after an automatic repair attempt
  • SMART degradation on an OSD in the affected acting set
Investigate

Ceph monitoring maturity levels

Ceph observability works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much your cluster matters. Most production clusters should land at the second level.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: is the cluster healthy, and can it still accept data? You will not learn what broke, but you will learn that something broke before users do. Survival is enough for lab clusters and non-critical storage.

  • Cluster health status ceph_health_status — OK / WARN / ERR, the umbrella over every check.
  • Monitor quorum A majority of MONs in quorum, or no map update can commit.
  • OSD up / in counts How many disks are alive and participating in placement.
  • Cluster capacity utilisation Raw used vs total against the nearfull and full ratios.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production clusters should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose an incident on its own: stuck PGs, slow OSDs, stalled recovery, capacity pressure, corruption.

  • Individual PG states Not just aggregates — degraded, undersized, stale, down, incomplete, inconsistent.
  • Per-OSD commit and apply latency Outliers localise a failing disk or a saturated DB/WAL device.
  • Slow ops count ceph_healthcheck_slow_ops — operations stuck past 30s.
  • Recovery / backfill rate Degraded PGs with zero recovery means healing has stalled.
  • Unfound / degraded / misplaced objects Redundancy status and the potential-data-loss signal.
  • Scrub inconsistency detection ceph_pg_inconsistent and OSD_SCRUB_ERRORS — corruption found.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. A flag left set after maintenance, one OSD filling faster than the rest, clock drift creeping toward election churn, scrubs falling behind. None of these page you on day one. They become the incident on day thirty.

  • OSD cluster flags The noout trap: noout / norecover / nobackfill left set.
  • Per-OSD capacity outliers One OSD near full stops writes even at a moderate average.
  • Per-pool capacity and quotas ceph_pool_percent_used and pool quota headroom.
  • Clock skew MON_CLOCK_SKEW before it turns into quorum churn.
  • Recovery-flag correlation Degraded PGs plus norecover/nobackfill means healing is off.
  • Scrub recency PG_NOT_DEEP_SCRUBBED — verification debt piling up.
  • OSD flapping detection OSD_FLAPPING and rapid map-epoch churn.
  • CRUSH placement safety TOO_FEW_OSDS / POOL_NO_REDUNDANCY from bad rules.

Level 4: Expert

Subsystem-specific observability after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. BlueStore internals, RGW OMAP health, MDS cap pressure, per-device-class capacity. Most teams never need every signal here. Add the ones your incident history says you do.

  • BlueStore DB/WAL and compaction BLUEFS_SPILLOVER and compaction stalls behind periodic latency.
  • RGW latency and OMAP health GET/PUT latency, queue length, LARGE_OMAP_OBJECTS.
  • MDS cap pressure and recall Cache oversized, recall throttle, client eviction (CephFS).
  • Per-device-class capacity HDD vs SSD vs NVMe fill tracked separately.
  • Erasure-coding shard recovery EC pools recover differently and cost more CPU.
  • PG autoscaler activity Splits and merges cause brief I/O stalls.
  • OSD memory target vs RSS Eviction pressure and the OOM kills that read as flapping.

Operating mistakes worth avoiding

The traps Ceph teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after an incident.

Dismissing nearfull warnings

The 85% warning gets waved off as premature. But recovery needs spare capacity: at 85%, losing OSDs pushes the survivors past backfillfull (90%) and blocks the very recovery meant to heal the failure. The cluster loses the ability to self-heal exactly when it needs to. Treat nearfull as a capacity deadline, not a nag, and keep headroom below backfillfull for a full OSD's worth of data.

Monitoring only cluster-wide metrics

Cluster-average capacity looks fine while one OSD sits at 96% — and a single full OSD stops writes cluster-wide. Cluster-average latency looks fine while one OSD stalls every request routed through it. Ceph's failures are lumpy; per-OSD capacity and per-OSD latency are where they show up first. Aggregate dashboards hide the outlier that pages you.

The noout trap

Setting <code>noout</code> for maintenance and forgetting to unset it is the single most common preventable Ceph outage. DOWN OSDs never get marked OUT, so recovery never starts, degraded PGs pile up silently, and the next failure lands on already-reduced redundancy. Alert whenever noout is set for more than a maintenance window, and urgently when noout is set while an OSD is down.

Treating HEALTH_WARN as background noise

WARN fires constantly during normal recovery, so teams mute it — and then miss the WARN that means nearfull, scrub errors, a set noout flag, or clock skew. Never silence the umbrella status. Split transient recovery warnings from structural ones using the name labels in ceph_health_detail, and alert on the specific checks that matter.

Ignoring scrub inconsistencies

Deep scrub finds a corrupt replica on a weekend; by Monday the cluster reads HEALTH_OK again and the ticket is closed. Those inconsistent replicas are a ticking bomb — if the good copy's OSD fails before repair, the data is lost. Investigate every ceph_pg_inconsistent promptly, and never run a blind <code>ceph pg repair</code>, which may overwrite the good replica with the corrupt primary.

Not watching recovery rate

Teams watch whether PGs are degraded but not whether recovery is actually progressing. A cluster with degraded PGs and a healthy recovery rate is healing normally; the same degraded count with zero recovery is a stalled, dangerous state one failure away from loss. Alert on degraded-and-not-recovering, not on degraded alone, or you drown in false tickets during normal healing.

Underestimating BlueStore DB spillover

When the dedicated DB device fills, RocksDB metadata spills onto the slow data device and performance collapses by 10-100x — but the OSD does not crash and capacity metrics look fine, so it is easily misread as a random slowdown. Only commit latency and BlueFS slow-device usage reveal it. Size the DB partition for the object count you actually store, and monitor for BLUEFS_SPILLOVER.

Not separating public and cluster networks

When client traffic and replication/recovery share the same NICs, a recovery storm saturates the link that clients depend on, and a routine OSD failure becomes a latency incident. A separate cluster network keeps rebuild traffic off the client path; if you cannot separate physically, at least throttle recovery so it cannot consume the whole link.

Ceph runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up Ceph monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing.