<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>SLA-SLO on Netdata</title><link>https://www.netdata.cloud/tags/sla-slo/</link><description>Recent content in SLA-SLO on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 22 Aug 2026 05:09:03 +0300</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/sla-slo/index.xml" rel="self" type="application/rss+xml"/><item><title>Designing Error Budget Policies For SLOs At Scale</title><link>https://www.netdata.cloud/academy/designing-error-budget-policies/</link><pubDate>Thu, 04 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/designing-error-budget-policies/</guid><description>&lt;p&gt;In every engineering organization, there&amp;rsquo;s a constant, fundamental tension: the push to ship new features versus the need to maintain a stable, reliable service. Move too fast, and you risk outages that erode user trust. Move too slowly, and you risk being outpaced by the competition. For years, this balancing act was managed by intuition, late-night heroics, and tense priority meetings. Site Reliability Engineering (SRE) offers a better way: the error budget.&lt;/p&gt;</description></item><item><title>Understanding Error Budgets And Their Importance In SRE</title><link>https://www.netdata.cloud/academy/error-budget/</link><pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/error-budget/</guid><description>&lt;p&gt;In the quest for flawless digital experiences, the reality is that 100% uptime is an elusive, if not impossible, goal. Systems inevitably encounter issues, and services can experience disruptions. This is where the concept of an &lt;strong&gt;error budget&lt;/strong&gt; becomes a cornerstone for modern &lt;a href="https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/"&gt;Site Reliability Engineering (SRE) and DevOps&lt;/a&gt; practices. Understanding and effectively managing your &lt;strong&gt;error budget&lt;/strong&gt; can mean the difference between fostering innovation and constantly firefighting.&lt;/p&gt;&#10;&lt;p&gt;So, &lt;strong&gt;what is an error budget&lt;/strong&gt;? Simply put, it&amp;rsquo;s the quantifiable amount of unreliability or downtime that a service can tolerate over a specific period without breaching its Service Level Objectives (SLOs) or upsetting users. It&amp;rsquo;s the acknowledged margin for error, a critical component in balancing the drive for new features with the imperative of maintaining a stable and reliable service. For SRE teams, the &lt;strong&gt;sre error budget&lt;/strong&gt; is not just a metric- it&amp;rsquo;s a vital tool for decision-making.&lt;/p&gt;</description></item><item><title>What Is Uptime Monitoring? All SREs &amp; DevOps Teams Must Know</title><link>https://www.netdata.cloud/academy/what-is-uptime-monitoring/</link><pubDate>Thu, 19 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-uptime-monitoring/</guid><description>&lt;h2 id="uptime-monitoring-explained"&gt;Uptime Monitoring Explained&lt;/h2&gt;&#10;&lt;p&gt;Keeping an eye on how well and how often servers, apps, services, and all the parts of your system are up and running is what &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/"&gt;uptime monitoring&lt;/a&gt; is all about. For people in &lt;a href="https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/"&gt;Site Reliability Engineering (SRE) and DevOps teams&lt;/a&gt;, making sure everything works almost all the time is super important. Keeping your services up and running means users run into less trouble and enjoy a more seamless connection without outages. This cuts down on the chance of expensive interruptions in business.&lt;/p&gt;</description></item><item><title>CockroachDB Backup Job Failures: How To Fix Them</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-backup-job-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-backup-job-failures/</guid><description>&lt;p&gt;CockroachDB scheduled backups run through the internal jobs system, visible via &lt;code&gt;crdb_internal.jobs&lt;/code&gt; (backed by &lt;code&gt;system.jobs&lt;/code&gt;). When a scheduled backup fails silently, stalls indefinitely, or grows so slowly that it cannot complete within its interval, your recovery point objective (RPO) is at risk. The failure is insidious: backups often appear healthy until you realize the last successful completion was 36 hours ago.&lt;/p&gt;&#10;&lt;p&gt;The severity distinction is sharp. A single backup failure with a recent prior success is a TICKET. The retry will probably succeed. But when no successful backup exists within your RPO window (for example, last success over 24 hours ago with a 24-hour RPO), you are in PAGE territory.&lt;/p&gt;</description></item></channel></rss>