<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Incident-Response on Netdata</title><link>https://www.netdata.cloud/tags/incident-response/</link><description>Recent content in Incident-Response on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 10 Apr 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/incident-response/index.xml" rel="self" type="application/rss+xml"/><item><title>Alert Acknowledgement: Mark It as Seen, Keep Working</title><link>https://www.netdata.cloud/blog/alert-acknowledge/</link><pubDate>Fri, 10 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/alert-acknowledge/</guid><description>&lt;p>If you&amp;rsquo;ve ever opened the alerts tab during a busy period, you know the problem. There are alerts you&amp;rsquo;ve already looked at, alerts someone on your team is handling, and alerts that fired on a known issue that&amp;rsquo;s being worked on. They all sit together in the same list alongside the new ones you haven&amp;rsquo;t seen yet. There&amp;rsquo;s no way to say &amp;ldquo;I&amp;rsquo;ve seen this, move on&amp;rdquo; without silencing or disabling the alert entirely, which is a much heavier action than the situation calls for.&lt;/p></description></item><item><title>Introducing Real-Time Conversations with Netdata AI</title><link>https://www.netdata.cloud/blog/ai-conversations/</link><pubDate>Tue, 23 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ai-conversations/</guid><description>&lt;p>Over the past few months, we&amp;rsquo;ve seen incredible adoption of our AI Investigations and Insights reports. Teams are using them to automate the deep, thoughtful analysis required for complex post-mortems, capacity planning, and performance optimization. These comprehensive reports are fantastic when you need a well-researched, shareable document.&lt;/p>
&lt;p>But what about the moments &lt;em>during&lt;/em> an investigation? What about the rapid-fire &amp;ldquo;what if&amp;rdquo; questions and the quick exploration of hypotheses that happen in the heat of the moment? For that, you need speed and interactivity. You need a partner you can have a real-time dialogue with.&lt;/p></description></item><item><title>Blast Radius Detection For Faster Incident Response</title><link>https://www.netdata.cloud/features/aiml/blast-radius-detection/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/blast-radius-detection/</guid><description>Netdata reveals blast radius dynamically through real-time anomaly correlation and ML-powered pattern recognition, showing the complete story from first failure to full impact in seconds.</description></item><item><title>Operations Center Monitoring &amp; 24/7 NOC Teams</title><link>https://www.netdata.cloud/solutions/built-for/ops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/ops/</guid><description>Real-time observability platform designed for 24/7 operations centers. Per-second monitoring, zero-configuration deployment, and AI-powered troubleshooting reduce MTTR by 80%.</description></item><item><title>Real-Time Observability For 24/7 Operations</title><link>https://www.netdata.cloud/solutions/use-cases/continuous-operations/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/continuous-operations/</guid><description>Real-time observability built for continuous operations. Per-second metrics, ML anomaly detection, and AI root cause analysis keep your infrastructure running around the clock.</description></item><item><title>AI Troubleshooting GA With On-Demand Credits</title><link>https://www.netdata.cloud/blog/ai-credits/</link><pubDate>Tue, 02 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ai-credits/</guid><description>&lt;p>Since launching our AI investigations and insights in a research preview, one thing has become clear: &lt;strong>automated root cause analysis delivers a significant return on investment.&lt;/strong> Teams have confirmed that instant insights don&amp;rsquo;t just save a few minutes; they fundamentally shorten incident response cycles, free up valuable engineering hours, and reduce the business impact of downtime.&lt;/p>
&lt;p>The preview successfully demonstrated this value, with 10 free AI sessions per month allowing teams to integrate AI into their workflows. Now, based on the success and maturity of the capabilities, we are proud to announce that &lt;strong>Netdata&amp;rsquo;s AI investigations and insights are graduating from research preview to General Availability.&lt;/strong>&lt;/p></description></item><item><title>Save Hours on Troubleshooting with Automated Investigations</title><link>https://www.netdata.cloud/blog/automated-investigations/</link><pubDate>Mon, 04 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/automated-investigations/</guid><description>&lt;p>How many times has your team stared at a dashboard, pointed to a spike, and asked a question that charts alone can&amp;rsquo;t answer? &amp;ldquo;What was the real impact of that deployment?&amp;rdquo; &amp;ldquo;Why are our Kubernetes pods in the us-east-1 cluster suddenly crashing?&amp;rdquo; &amp;ldquo;Are we wasting money on overprovisioned servers?&amp;rdquo;&lt;/p>
&lt;!--truncate-->
&lt;p>Answering these questions is the real work of operations and SRE. It often kicks off a time-consuming scramble, sending engineers down rabbit holes for hours, days, or even weeks. You dig through logs, correlate metrics across services, and piece together clues from Slack conversations and Jira tickets.&lt;/p></description></item><item><title>Netdata Now Troubleshoots Your Alerts for You</title><link>https://www.netdata.cloud/blog/automated-alert-troubleshooting/</link><pubDate>Sun, 03 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/automated-alert-troubleshooting/</guid><description>&lt;p>The 2 AM pager alert. For anyone in Ops, SRE, or IT administration, those words trigger a familiar sense of dread. An alert has fired. Is it a real fire, or another false alarm waking you from a dead sleep? The pressure is on. Every minute of downtime costs money and reputation, but troubleshooting a complex system when you&amp;rsquo;re sleep-deprived is a Herculean task.&lt;/p>
&lt;!--truncate-->
&lt;p>This cycle is a massive drain on engineering resources. The daily grind of sifting through alerts, trying to distinguish signal from noise, and manually correlating metrics to find a root cause consumes countless hours. This constant firefighting leads to alert fatigue, where even critical notifications start to get ignored. The core questions are always the same: Is this a real problem? What is the potential impact? Why did this trigger? What do I do next? Answering them is a slow, manual, and often stressful process.&lt;/p></description></item><item><title>ilert Integration: Streamline Monitoring &amp; Response</title><link>https://www.netdata.cloud/blog/netdata-integration-with-ilert-streamlining-monitoring-and-incident-response/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-integration-with-ilert-streamlining-monitoring-and-incident-response/</guid><description>&lt;p>Netdata now integrates with ilert, a leading incident response platform. With this integration, the incident management features and alerting capabilities of ilert and the real-time systems monitoring provided by Netdata can be leveraged. By combining both systems, users can not only monitor their infrastructure with fine detail as never before, but also assure the responsiveness of critical alerts to the correct teams swiftly.&lt;/p>
&lt;h2 id="what-is-ilert">What is ilert?&lt;/h2>
&lt;p>&lt;a href="https://www.ilert.com/?utm_campaign=Netdata&amp;utm_source=integration&amp;utm_medium=organic" target="_blank">ilert&lt;/a> is an end-to-end platform for alerting, on-call management, and status pages, built for the new-age DevOps and SRE teams. It streamlines the entire incident management process by automating key aspects and providing powerful tools to improve efficiency, such as automated on-call duty, multi-level escalations, alert grouping, postmortem document creation, and much more. ilert integrates with monitoring systems, like Netdata, and sends alerts through multiple channels (SMS, phone, push, Slack, Microsoft Teams, etc.), ensuring that teams can promptly respond to issues before they impact service.&lt;/p></description></item></channel></rss>