Ceph is self-healing until it isn't, and the line between the two is capacity you forgot to leave and flags you forgot to clear.
The defaults keep you alive. Until one OSD crosses the full ratio and Ceph refuses every write cluster-wide — a hard stop, not a slowdown, and reads keep working so it hides in plain sight. Until a cluster running hot loses a disk and recovery has nowhere to put the data, so backfill_toofull appears and degraded PGs stop healing. Until an OSD with a slow disk misses a heartbeat, gets marked down, triggers a peering storm across hundreds of PGs, and the storm slows its neighbours into missing heartbeats too. Until a monitor's clock drifts past 50ms and the quorum churns through elections while the whole cluster freezes. Until deep scrub finds a checksum mismatch on a weekend and the good replica's disk dies before anyone repairs it.
These guides are written for engineers who already run Ceph, not for people deciding whether to. The goal is the mental model of how RADOS, CRUSH, the OSDs, and the monitors actually behave under load, the failure patterns that keep recurring, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident.