<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>GPU on Netdata</title><link>https://www.netdata.cloud/tags/gpu/</link><description>Recent content in GPU on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 23 Aug 2026 02:03:02 +0300</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/gpu/index.xml" rel="self" type="application/rss+xml"/><item><title>10 Best GPU Monitoring Tools for AI Workloads (2026)</title><link>https://www.netdata.cloud/resources/best-gpu-monitoring-tools/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-gpu-monitoring-tools/</guid><description/></item><item><title>Native macOS Monitoring: Logs, Sensors, GPU &amp; Hardware Health</title><link>https://www.netdata.cloud/blog/macos-monitoring/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/macos-monitoring/</guid><description>&lt;p&gt;&lt;img src="../images/macos-monitoring.svg" alt="Native macOS monitoring with Netdata: unified logs, power, sensors, GPU, per-app metrics, storage, and network"&gt;&lt;/p&gt;&#10;&lt;p&gt;We&amp;rsquo;ve overhauled macOS monitoring in the latest Netdata release. Netdata already collects system metrics on Macs at per-second resolution; this release completes the picture with logs and hardware telemetry, areas that previously required users to run CLI tools like &lt;code&gt;log show&lt;/code&gt; and &lt;code&gt;powermetrics&lt;/code&gt;. The new collectors read this data through Apple&amp;rsquo;s own frameworks, allowing users to trace application and OS errors and catch hardware issues early.&lt;/p&gt;</description></item><item><title>NVIDIA DCGM Collector: Deep GPU Monitoring For AI</title><link>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</link><pubDate>Mon, 04 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</guid><description>&lt;p&gt;&lt;img src="../images/dcgm-collector.svg" alt="NVIDIA DCGM Collector: Deep GPU Monitoring for Data Center and AI Infrastructure"&gt;&lt;/p&gt;&#10;&lt;p&gt;GPU infrastructure is expensive and increasingly central to production workloads. Whether you&amp;rsquo;re running ML training jobs, inference serving, video transcoding, or HPC workloads, understanding what your GPUs are actually doing, and what&amp;rsquo;s going wrong when performance degrades, is not optional. The problem is that NVIDIA&amp;rsquo;s Data Center GPU Manager (DCGM) exposes an enormous amount of telemetry, but getting that data into a monitoring system in a useful, organized way has traditionally required significant setup and custom dashboarding work.&lt;/p&gt;</description></item><item><title>HPC Monitoring Software With Per-Second Metrics</title><link>https://www.netdata.cloud/solutions/use-cases/hpc/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/hpc/</guid><description>Transform HPC operations with distributed edge-native monitoring that delivers sub-2-second insights, 90% cost reduction, and linear scalability to 100,000+ nodes - without the complexity.</description></item><item><title>LLM Monitoring Platform With Per-Second Visibility</title><link>https://www.netdata.cloud/solutions/use-cases/llm-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/llm-monitoring/</guid><description>Real-time infrastructure monitoring for LLM deployments. Track GPU utilization, container resources, database performance, and system metrics with per-second granularity and ML-powered anomaly detection.</description></item><item><title>Real-Time GPU Monitoring For AI Infrastructure</title><link>https://www.netdata.cloud/solutions/industries/ai/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/ai/</guid><description>Monitor AI training clusters, inference APIs, and ML workloads with Netdata&amp;rsquo;s edge-native platform. Per-second visibility, ML-powered insights, predictable pricing.</description></item><item><title>TTIC Case Study: Efficient System Monitoring</title><link>https://www.netdata.cloud/case-studies/education/ttic/</link><pubDate>Thu, 28 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/ttic/</guid><description>&lt;h2 id="enhancing-system-management-with-scarce-time-resources"&gt;Enhancing System Management with Scarce Time Resources&lt;/h2&gt;&#10;&lt;p&gt;At the Toyota Technological Institute at Chicago, the challenge of system administration is uniquely compounded by the scarcity of time. As a one-man IT department, Adam Bohlander juggles various responsibilities, making efficient time management crucial. The institute&amp;rsquo;s evolving needs demand a monitoring solution that aligns with its mission of leading in computer science and information technology research and education.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&amp;ldquo;I find the install quick and easy via ansible (my own playbooks) and defaults to be quite sane. This allows me to use it without investing that much time.&amp;rdquo;&lt;/p&gt;</description></item><item><title>University of Calgary: Netdata in Academia</title><link>https://www.netdata.cloud/case-studies/education/calgary-university/</link><pubDate>Tue, 26 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/calgary-university/</guid><description>&lt;h2 id="empowering-academic-research-with-real-time-monitoring"&gt;Empowering Academic Research with Real-Time Monitoring&lt;/h2&gt;&#10;&lt;p&gt;At the &lt;strong&gt;Machine Learning Lab&lt;/strong&gt; at the &lt;strong&gt;University of Calgary&lt;/strong&gt;, managing a robust infrastructure of servers and workstations is critical for advancing their research in machine learning (ML). Assistant Professor &lt;strong&gt;Yani Ioannou&lt;/strong&gt; oversees this infrastructure, ensuring that these essential resources remain operational and are used to their fullest potential. The lab&amp;rsquo;s challenge lies not just in the maintenance of these resources but in minimizing downtime to keep the research moving forward.&lt;/p&gt;</description></item><item><title>Enhancing Research Through Advanced Monitoring</title><link>https://www.netdata.cloud/case-studies/education/landcare/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/landcare/</guid><description>&lt;h2 id="helping-scientists-focus-on-the-science"&gt;Helping Scientists focus on the Science&lt;/h2&gt;&#10;&lt;p&gt;At &lt;a href="https://www.landcareresearch.co.nz/"&gt;Manaaki Whenua – Landcare Research&lt;/a&gt;, managing multi-user Linux workstations for scientific computing presented a unique challenge. Tasked with overseeing system utilization without a dedicated role in the organization, the need for a solution that was intuitive and efficient became paramount. Netdata&amp;rsquo;s simplicity in setup and the ability to monitor systems remotely via the cloud addressed these challenges head-on, offering a seamless solution for workstation management.&lt;/p&gt;</description></item><item><title>AMD CPU &amp; GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/amd_smi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/amd_smi-monitoring/</guid><description>&lt;h2 id="amd-cpu--gpu-monitoring"&gt;AMD CPU &amp;amp; GPU Monitoring&lt;/h2&gt;&#10;&lt;h3 id="what-is-amd-cpu--gpu"&gt;What Is AMD CPU &amp;amp; GPU?&lt;/h3&gt;&#10;&lt;p&gt;AMD CPUs and GPUs form the core of many computational systems, driving performance in everything from gaming computers to data centers that power cloud applications. The AMD System Management Interface (SMI) is crucial for monitoring the performance and health of these processors, helping users optimize hardware performance through precise metrics.&lt;/p&gt;&#10;&lt;h3 id="monitoring-amd-cpu--gpu-with-netdata"&gt;Monitoring AMD CPU &amp;amp; GPU With Netdata&lt;/h3&gt;&#10;&lt;p&gt;Monitoring AMD CPU &amp;amp; GPU is essential for ensuring that your systems operate efficiently and without interruption. Netdata leverages the &lt;a href="https://github.com/amd/amd_smi_exporter"&gt;AMD SMI Exporter&lt;/a&gt; to monitor these devices. By using an openmetrics (Prometheus) exporter, Netdata can seamlessly ingest metrics without the need for a Prometheus server or Grafana. Once connected, users can enjoy automated dashboards and alerts, making it a comprehensive AMD CPU &amp;amp; GPU monitoring tool.&lt;/p&gt;</description></item><item><title>Intel GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/intelgpu-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/intelgpu-monitoring/</guid><description>&lt;h2 id="intel-gpu-monitoring"&gt;Intel GPU Monitoring&lt;/h2&gt;&#10;&lt;h3 id="what-is-intel-gpu"&gt;What Is Intel GPU?&lt;/h3&gt;&#10;&lt;p&gt;Intel Graphics Processing Units (GPUs) are integrated into Intel CPUs, offering powerful graphics capabilities for a variety of applications. These GPUs are particularly important for managing graphics-intensive workloads and improving the overall performance of systems using Intel-based hardware.&lt;/p&gt;&#10;&lt;h3 id="monitoring-intel-gpu-with-netdata"&gt;Monitoring Intel GPU With Netdata&lt;/h3&gt;&#10;&lt;p&gt;Netdata&amp;rsquo;s Intel GPU monitoring tool provides real-time, insightful data that allows you to track the performance and health of your Intel GPU seamlessly. With Netdata, you can monitor Intel GPU metrics like frequency, power consumption, and engine utilization without compromising system security or performance.&lt;/p&gt;</description></item><item><title>Nvidia GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nvidia-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nvidia-monitoring/</guid><description>&lt;h2 id="what-is-nvidia-gpu"&gt;What is Nvidia GPU?&lt;/h2&gt;&#10;&lt;p&gt;Nvidia GPU (Graphic Processing Unit) is a specialized electronic circuit designed to rapidly process and manipulate graphics data. Nvidia GPUs are typically found in high-end gaming computers, workstations, and servers. They are used to power complex video games and other graphics-intensive tasks.&lt;/p&gt;&#10;&lt;h2 id="monitoring-nvidia-gpu-with-netdata"&gt;Monitoring Nvidia GPU with Netdata&lt;/h2&gt;&#10;&lt;p&gt;The prerequisites for monitoring Nvidia GPU with Netdata are to have a system with an Nvidia GPU and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/"&gt;Netdata installed&lt;/a&gt; on your system.&lt;/p&gt;</description></item><item><title>Nvidia GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nvidia_smi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nvidia_smi-monitoring/</guid><description>&lt;h2 id="nvidia-gpu-monitoring"&gt;Nvidia GPU Monitoring&lt;/h2&gt;&#10;&lt;h3 id="what-is-nvidia-gpu"&gt;What Is Nvidia GPU?&lt;/h3&gt;&#10;&lt;p&gt;Nvidia GPUs are specialized processing units designed by &lt;a href="https://www.nvidia.com/en-us/"&gt;Nvidia&lt;/a&gt; primarily for graphics rendering, though they are widely used in computational tasks such as deep learning, scientific simulations, and cryptocurrency mining. Nvidia&amp;rsquo;s advanced GPU technology empowers applications to perform complex tasks efficiently.&lt;/p&gt;&#10;&lt;h3 id="monitoring-nvidia-gpu-with-netdata"&gt;Monitoring Nvidia GPU With Netdata&lt;/h3&gt;&#10;&lt;p&gt;Netdata provides real-time monitoring for Nvidia GPUs by leveraging the &lt;code&gt;nvidia-smi&lt;/code&gt; CLI tool. This setup allows you to keep an eye on various performance metrics, ensuring optimal operation and helping diagnose potential issues as they occur.&lt;/p&gt;</description></item></channel></rss>