The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

DevOps

GitLab Runner Executor Failures: Docker & Kubernetes

A deep dive into diagnosing and resolving common problems with Docker and Kubernetes executors- including cache- DIND- and RBAC configuration
by Netdata Team · July 24, 2025

You’ve been there. You push your code, a new CI/CD pipeline kicks off in GitLab, and you wait for that satisfying green checkmark. Instead, you get a dreaded red ‘X’. The pipeline failed. Digging into the job logs, you find a cryptic message: “ERROR: Job failed (system failure): prepare environment: exit code 1”. The culprit is often the GitLab Runner executor—the very engine responsible for running your jobs—failing in its environment.

These failures can be notoriously difficult to debug. They often stem not from your code or tests, but from the complex interaction between the runner and its underlying infrastructure, whether it’s a Docker daemon or a Kubernetes cluster. In this guide, we’ll unravel the most common GitLab Runner executor failures, from Docker socket permission errors to Kubernetes pod scheduling problems and inefficient cache configurations. We’ll show you how to fix them and, more importantly, how to proactively monitor your runners to prevent these failures from ever happening again.

The Docker Executor: Common Pitfalls and Solutions

The Docker executor is one of the most popular choices for GitLab Runner. It provides a clean, isolated environment for each CI/CD job. However, this isolation comes with its own set of challenges, primarily centered around permissions, networking, and state management.

The Dreaded “Docker Socket Permission Denied”

One of the most frequent errors you’ll encounter when setting up a Docker executor is a permissions issue with the Docker socket. The job log might show something like: “Got permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock”.

This happens because the GitLab Runner process, which runs as the gitlab-runner user by default, is trying to communicate with the Docker daemon by writing to its socket file. However, this socket is typically owned by the root user and the docker group. If the gitlab-runner user isn’t part of the docker group, the Docker daemon will deny access.

The most direct solution is to add the gitlab-runner user to the docker group and then restart the runner service for the changes to take effect.

Security Consideration: Giving direct access to the Docker socket is equivalent to granting root access on the host machine. A process with socket access can start, stop, and manage containers, including a docker socket bind mount, leading to potential docker privilege escalation. For security-conscious environments, you should explore rootless Docker or an alternative from the runner executor comparison, like the Kubernetes executor.

The Docker-in-Docker (DinD) Dilemma

A common gitlab ci docker task is building a Docker image. To do this within a Docker executor job, you need access to a Docker daemon. This leads to the “Docker-in-Docker GitLab” (DinD) pattern, where a dedicated Docker daemon container is spun up alongside your job’s container.

While it works, DinD has significant drawbacks:

  • Security Risk: It requires running the executor in --privileged mode, which disables nearly all container security mechanisms. A compromised job could potentially take over the entire runner host.
  • Performance Overhead: You’re running a full Docker daemon inside another container, which consumes extra resources and can slow down your pipelines, especially with many gitlab runner concurrent jobs.
  • Caching Complexity: Layer caching with DinD is notoriously tricky and often inefficient, leading to slow image builds as layers are rebuilt on every job.

Alternatives to DinD: For building container images, consider more modern, secure, and efficient tools that don’t require a Docker daemon:

  • Kaniko: A tool from Google that builds container images from a Dockerfile inside a container or Kubernetes cluster, without needing privileged access.
  • img: A standalone, unprivileged, and daemon-less container image builder.

Inefficient GitLab Runner Cache

The gitlab runner cache is designed to speed up jobs by persisting files between runs, such as node_modules or maven dependencies. With the gitlab runner docker executor, this cache is often stored on the host machine or in a distributed object store like S3.

Cache issues manifest as slow jobs or pipelines that seem to re-download dependencies every time. This can be caused by:

  • Slow Disk I/O: If the runner host is using slow network-attached storage (NAS/NFS) for the cache directory, read/write operations can become a major bottleneck.
  • Misconfigured Distributed Cache: When using S3, incorrect credentials, bucket policies, or high network latency to the S3 endpoint can make caching slower than not using it at all.
  • Incorrect Cache Keys: If your cache key is too dynamic (e.g., based on commit SHA), you may never get a cache hit, defeating its purpose.

The gitlab runner kubernetes executor is a powerful and scalable option that runs each CI job as a separate Pod in your cluster. This provides excellent isolation and leverages Kubernetes’ scheduling and resource management capabilities. However, its complexity introduces new potential points of failure, requiring specific gitlab runner troubleshooting.

Pod Creation Errors and Service Account Woes

When a job starts, the GitLab Runner manager Pod communicates with the Kubernetes API server to create a new Pod for the job. If this fails, you’ll see errors in the gitlab runner logs pointing to a problem with Pod creation.

This is almost always an RBAC issue. The kubernetes service account gitlab runner uses needs specific permissions in the target namespace to manage the lifecycle of job pods. If the associated Role or ClusterRole is missing permissions, the API server will reject the requests with a Forbidden error. This is a key part of kubernetes rbac gitlab runner configuration.

The ServiceAccount for the runner typically needs permissions to manage Pods, Services, Secrets, and ConfigMaps. Carefully review the official GitLab documentation for the exact RBAC manifest required for your version.

Authentication Failures with the Container Registry

A common runner executor failed scenario in the gitlab kubernetes executor is an ImagePullBackOff error. This means Kubernetes tried to pull the Docker image for your CI job but failed due to a container registry authentication issue.

To resolve this, you need to create a Kubernetes Secret of type docker-registry containing your registry credentials and then specify this secret in your runner’s config.toml. This ensures that every job Pod created by the runner has the necessary credentials to pull images. The initial runner registration token does not handle this.

Hitting Resource Quotas and Limits

Kubernetes administrators often use ResourceQuota and LimitRange objects to control resource consumption within a namespace. If your CI job Pod requests more CPU or memory than is allowed by the namespace’s quota, the Kubernetes scheduler will refuse to schedule it.

The job will get stuck in a Pending state. Inspecting the Pod’s events will reveal a message like FailedScheduling with a reason of exceeded quota. This requires you to either increase the quotas or adjust the resource requests in your gitlab runner configuration, which is a core part of gitlab runner scaling.

From Reactive to Proactive: Monitoring Your Runners with Netdata

Troubleshooting executor failures by digging through logs is a reactive process. You’re fixing something that’s already broken and has already delayed a deployment. The modern SRE approach is to use comprehensive, real-time monitoring to detect the signs of impending failure and act before the pipeline turns red. This is where Netdata excels.

Netdata provides unparalleled, high-granularity visibility into the entire stack supporting your GitLab Runners, allowing you to correlate executor behavior with system health.

Beyond Runner Logs: Correlating Failures with System Health

A runner log might tell you a job failed, but it rarely tells you why the environment was unhealthy.

  • Did a Docker job fail because the host’s CPU was pegged at 100%, causing the build process to time out?
  • Did a Kubernetes Pod get OOMKilled because the node ran out of memory?
  • Was slow GitLab Runner cache access caused by a disk I/O bottleneck on the underlying storage?

Netdata answers these questions by automatically collecting thousands of metrics from your systems. It allows you to see a spike in job failures on your GitLab dashboard and immediately correlate it with a CPU saturation alert, a memory pressure chart, or a disk latency heatmap in your Netdata dashboard—all on the same timeline.

Monitoring the Docker Host and Kubernetes Cluster

For the Docker executor, Netdata automatically monitors the health of the runner host, including CPU, memory, disk I/O, and network statistics. It even uses eBPF to provide deep insights into individual Docker containers, so you can see exactly how much resource each CI job is consuming.

For the Kubernetes executor, Netdata provides a holistic view of your cluster’s health. It monitors:

  • Node Health: CPU/memory/disk/network usage for every node.
  • Pod Status: Tracks the number of Pending, Failed, and Evicted pods, which can be early indicators of scheduling or resource problems.
  • Kubelet & API Server Performance: Monitors the health and latency of critical Kubernetes control plane components.

Smart Alerts for Failure Prevention

With Netdata’s extensive metric collection and health monitoring, you can set up intelligent alerts that warn you of conditions likely to cause executor failures:

  • Docker Host: Alert when CPU utilization is high, available memory is low, or disk latency exceeds a threshold.
  • Kubernetes: Alert when a node enters a NotReady state, when available Pod capacity in a namespace is low, or when the Kubernetes API server error rate spikes.

These proactive alerts give you time to scale your runners, adjust resource limits, or fix the underlying infrastructure issue before your CI/CD pipelines start failing.

Building Resilient CI/CD Pipelines

GitLab Runner executor failures are a frustrating but solvable problem. By understanding the common pitfalls of the Docker and Kubernetes executors—from permissions and caching to RBAC and resource quotas—you can effectively debug and fix your pipelines.

However, the key to truly resilient CI/CD is to shift from a reactive troubleshooting mindset to a proactive, observability-driven one. By implementing a powerful monitoring solution like Netdata, you gain the deep visibility needed to understand the health of your runner fleet, correlate failures with their root causes, and prevent outages before they impact your developers.

Ready to stop chasing red pipelines? Explore how Netdata can bring real-time observability to your GitLab Runners.