The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Databases

Using FOR UPDATE SKIP LOCKED For Queue Workflows

How a single SQL clause can transform your database into a high-throughput- parallel-processing task queue
by Netdata Team · August 29, 2025

One of the most common and powerful patterns in modern application development is the job queue. Whether you’re sending emails, processing images, or running complex calculations, offloading tasks to background workers is essential for building responsive and scalable systems. Many developers reach for dedicated queueing software like RabbitMQ or Redis, but for many use cases, your primary PostgreSQL database already has all the tools you need to build a robust, transactional, and incredibly performant job_queue_postgres.

The challenge, however, lies in coordination. How do you allow multiple concurrent_consumers (workers) to safely grab jobs from a table without them all trying to work on the same task, or worse, getting stuck in a traffic jam of database locks? The naive approach often leads to race conditions, while traditional locking (FOR UPDATE) can serialize your workers, defeating the purpose of parallelism.

This is where a powerful but often overlooked PostgreSQL feature comes into play: FOR UPDATE SKIP LOCKED. This single clause provides an elegant and highly efficient mechanism to create a deadlock_free_queue, allowing your workers to operate in parallel with minimal contention.

TL;DR

  • FOR UPDATE SKIP LOCKED lets a query skip rows already locked by other transactions instead of waiting, which turns a PostgreSQL table into a high-throughput job queue.
  • It fixes the contention that breaks naive queues: plain logic causes duplicate processing and plain FOR UPDATE forces workers into a single-file convoy, while SKIP LOCKED keeps them running in parallel.
  • Because the lock is tied to the transaction, a worker crash automatically rolls back and returns the job to pending, making it safer and simpler than manually managed advisory locks.
  • Performance depends on a composite index like (status, created_at), an ORDER BY for processing order, partitioning for huge tables, and monitoring queue depth, throughput, and lock contention.

The Classic Job Queue Problem: A Recipe For Contention

Let’s imagine a simple jobs table with a status column (pending, processing, completed). The most basic worker logic looks something like this:

  1. Select a Job: Find a job where the status is ‘pending’.
  2. Update the Status: Change the job’s status to ‘processing’ to mark it as taken.
  3. Process the Job: Perform the actual work.

The problem occurs between steps 1 and 2. If two workers run the SELECT query at nearly the same time, they will both see the same ‘pending’ job. They will then both try to UPDATE it, leading to either duplicate processing (a disaster for non-idempotent tasks) or one worker’s update failing.

A more robust approach uses pessimistic locking with FOR UPDATE. The worker’s query would look for a pending job and lock the row immediately. When a second worker tries to select that same row, the database forces it to wait until the first worker’s transaction is complete. This prevents duplicate work but introduces a new bottleneck. Your workers, which are supposed to run in parallel, now form a “convoy,” waiting in a single-file line for the first worker to finish. This leads to high blocking_time and poor queue_performance.

The Elegant Solution: FOR UPDATE SKIP LOCKED

Introduced in PostgreSQL 9.5, SKIP LOCKED is a modifier for FOR UPDATE that completely changes its behavior. When a query with SKIP LOCKED tries to acquire a row_level_lock on a row that is already locked by another transaction, it doesn’t wait. Instead, it simply skips that row as if it didn’t exist and moves on to the next one that meets the WHERE clause criteria.

This is the perfect primitive for a postgres_worker_queue. Each worker can now execute a query to find and lock the first available pending job.

A highly efficient dequeue_script can be built into a single, atomic query. This query finds a pending job, locks it (skipping any that are already locked by other workers), updates its status to ‘processing’, and returns the job’s data to the worker application all in one go. Because this happens in a single statement, it’s incredibly fast and eliminates any chance of a race condition between selecting and updating.

Building A Robust Worker

The full workflow for a worker using this transactional_queue model would be:

  1. Begin Transaction: The worker starts a new database transaction.
  2. Dequeue a Job: The worker executes the atomic SKIP LOCKED query.
  3. Check for Work: If the query returns a job, the worker proceeds. If it returns nothing (meaning all pending jobs are currently locked by other workers), the worker can sleep for a short period before trying again.
  4. Process the Job: The worker performs the business logic using the data from the dequeued job.
  5. Finalize the Job:
    • On success, the worker updates the job’s status to completed.
    • On failure, it updates the status to failed and potentially records the error message. This is a good place to implement retry_logic_sql, perhaps by incrementing a retry_count and setting the status back to pending if the count is below a threshold.
  6. Commit Transaction: The worker commits the transaction, releasing the lock and making the final status change permanent.

If the worker process crashes for any reason, the database automatically rolls back the transaction. The lock is released, and the job’s status reverts to ‘pending’, making it available for another worker to pick up.

SKIP LOCKED vs Advisory Locks

Before SKIP LOCKED was available, a common pattern for building a postgres_queue involved using advisory locks. An advisory_lock is a cooperative lock that developers manage manually in their application code. A worker would try to acquire a lock using a function call. If successful, it would process the job; if not, it would move on.

So, how does advisory_lock_vs_skip_locked stack up?

  • Complexity: Advisory locks are more complex. The application is responsible for managing the lock lifecycle, including ensuring locks are always released. A bug could lead to a “stuck” lock that requires manual intervention. SKIP LOCKED ties the lock directly to the row and the transaction, so it’s managed automatically by the database.
  • Safety: SKIP LOCKED is safer. Because the lock is tied to the transaction, a worker crash guarantees the lock is released. With session-level advisory locks, a crashed worker could hold a lock until the database connection times out, which could be a very long time.
  • Performance: For this specific use case, SKIP LOCKED is generally more performant as it’s a native, engine-level feature designed for this purpose.

For building a task_queue_db, FOR UPDATE SKIP LOCKED is the idiomatic, simpler, and more robust solution in modern PostgreSQL.

Performance Considerations & Best Practices

To ensure your optimistic_queue runs at peak performance, follow these best practices:

  • Create the Right Index: This is the most critical factor for queue_performance. Your jobs table must have a composite index on the columns used for finding work. For example, an index on (status, created_at) would allow PostgreSQL to very quickly find the oldest pending jobs.
  • Order Your Queue: Use an ORDER BY clause in your query to define the order in which jobs are processed (e.g., ORDER BY created_at for a FIFO queue, or ORDER BY priority, created_at for a priority queue).
  • Partition Large Tables: If your jobs table grows to hundreds of millions of rows, consider using PostgreSQL’s native table partitioning. You could partition by date or status, which can significantly improve query performance and make maintenance easier.
  • Monitor Your Queue: You can’t optimize what you can’t see. Continuously monitor key queue metrics:
    • Queue Depth: The number of jobs with status = 'pending'. A constantly growing number indicates your workers can’t keep up.
    • Worker Throughput: The rate of message_processing (jobs moved to completed or failed per second).
    • Lock Contention: Even with SKIP LOCKED, monitoring overall database lock contention is a good health check.

By combining the power of FOR UPDATE SKIP LOCKED with smart indexing and proactive monitoring, you can build a highly concurrent, reliable, and scalable job queue directly within the database you already trust, without adding another piece of infrastructure to your stack.

Netdata’s auto-discovery and detailed PostgreSQL dashboards make it easy to monitor these critical queue metrics in real-time, helping you fine-tune your worker pool and ensure your background job processing runs smoothly. Try Netdata for free today.

FOR UPDATE SKIP LOCKED FAQs

What Does FOR UPDATE SKIP LOCKED Do In PostgreSQL?

FOR UPDATE SKIP LOCKED is a modifier for the FOR UPDATE row-locking clause. When a query tries to lock a row that another transaction has already locked, the SKIP LOCKED part tells PostgreSQL not to wait. Instead, it skips that row as if it didn’t exist and moves on to the next row matching the WHERE clause. That behavior makes it ideal for job queues: multiple workers can each grab the first available pending job without blocking each other or processing the same row twice.

What Problem Does SKIP LOCKED Solve In A Job Queue?

It solves the contention problem that plain queue logic creates. A naive worker selects a pending job, then updates its status, but if two workers run the select at nearly the same moment they both grab the same job, causing duplicate processing or a failed update. Plain FOR UPDATE fixes the duplication but forces workers into a single-file “convoy,” each waiting for the one ahead. SKIP LOCKED lets every worker skip locked rows and pick up a different job, restoring true parallelism with minimal contention.

When Was SKIP LOCKED Introduced In PostgreSQL?

SKIP LOCKED arrived in PostgreSQL 9.5 as a modifier for FOR UPDATE. Before it existed, developers building a Postgres-backed queue typically reached for advisory locks, which the application had to manage manually. SKIP LOCKED made that pattern far simpler by tying the lock to the row and the transaction, so the database handles the lock lifecycle natively rather than leaving it to application code.

Does FOR UPDATE SKIP LOCKED Prevent Deadlocks In A Worker Queue?

Yes, it avoids the lock contention that stalls concurrent workers. Because a query with SKIP LOCKED never waits on a row another transaction holds, workers don’t pile up behind each other in a blocking convoy. Each one moves straight to an available job. The lock is also bound to the transaction, so if a worker crashes mid-job, the database rolls the transaction back, releases the lock, and the job’s status reverts to pending for another worker to claim. That combination keeps a multi-worker queue moving without stuck locks.

How Is SKIP LOCKED Different From Advisory Locks?

Both can build a Postgres queue, but they differ in three ways. Advisory locks are more complex: the application manages the full lock lifecycle and a bug can leave a “stuck” lock needing manual cleanup. SKIP LOCKED is safer because the lock is tied to the transaction, so a worker crash guarantees the lock is released, whereas a session-level advisory lock can linger until the connection times out. And for this use case SKIP LOCKED is generally more performant, since it’s a native engine feature built for exactly this purpose. For most modern Postgres queues it’s the simpler, more robust choice.

What Happens To A Job If A Worker Crashes Mid-Processing?

The job is automatically returned to the queue. Each worker wraps its work in a transaction: it begins the transaction, dequeues and locks a job with the SKIP LOCKED query, processes it, then commits to finalize the status. If the worker process crashes before committing, PostgreSQL rolls the transaction back. The row lock is released and the job’s status reverts to pending, making it available for another worker to pick up. You get crash recovery without writing any extra cleanup logic.

How Do I Make A SKIP LOCKED Job Queue Performant?

Indexing is the most important factor. Add a composite index on the columns your dequeue query filters and orders by, such as (status, created_at), so PostgreSQL can find the oldest pending jobs instantly. Use an ORDER BY clause to control processing order, like created_at for FIFO or priority, created_at for a priority queue. If the jobs table grows into hundreds of millions of rows, consider native table partitioning by date or status. Together these keep the queue fast as it scales.

What Metrics Should I Monitor For A Postgres Job Queue?

Watch three things. Queue depth is the count of jobs with status = ‘pending’; a number that keeps climbing means your workers can’t keep up. Worker throughput is the rate of jobs moving to completed or failed per second, which tells you how fast the pool is draining work. And lock contention is worth tracking as a database health check even with SKIP LOCKED in place. Tools like Netdata auto-discover PostgreSQL and chart these metrics in real time, so you can size your worker pool against live data.