The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Metrics, Logs & Traces
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Predictable Pricing.

Predictable per-node pricing for infrastructure monitoring. No per-metric or per-GB ingest charges for metrics and logs.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

S3 Compatible Object Storage icon

S3 Compatible Object Storage

S3 Compatible Object Storage documentation

S3 Compatible Object Storage

Plugin: go.d.plugin Module: s3check

Overview

Monitor S3-compatible object storage with active checks that write, read, list, and delete small probe objects and time their replication between sites.

The collector runs on the Netdata Agent as an ordinary S3 client, so every result includes what a real client experiences on that path: DNS, network, proxies, TLS, authentication, and the service itself. It reports the outcome and duration of each S3 operation, whether the content read back matches what was written, how long replication took compared with your objectives, how many probe objects still await cleanup, and whether new probes are paused because cleanup is backlogged.

One job runs one of three modes:

ModeWhat it checksBucket requirements
lifecycleOne endpoint. Writes a probe object, reads it back and verifies its content, confirms it is listed, deletes it, and confirms it is gone. Works with AWS S3, Ceph RGW, and other S3-compatible services.Versioning never enabled
ceph_multisiteOne replication direction between two Ceph RGW zones. Writes at the source, waits until the same object is readable at the destination, deletes it at the source, and waits until the destination no longer serves it.Versioning never enabled on either bucket
aws_replicationOne replication direction between two AWS S3 buckets, following AWS replication semantics: the object version created by the write and the delete marker created by the delete are each expected to replicate.Versioning enabled on both buckets and a replication rule that replicates delete markers

The lifecycle mode checks the life cycle of a probe object; it does not inspect S3 Lifecycle Management rules. Replication modes check one direction only: to validate a bidirectional setup, configure one job per direction.

Every job owns a private key space inside the configured prefix, <prefix><owner>/probe-..., where the owner segment is derived from this Agent and job. The collector writes, reads, lists, and deletes only keys it has recorded in that space. It never scans or touches other objects in the bucket, and Agents or jobs that share a bucket and prefix never interfere with each other.

Before a probe object is written, its key is recorded in a local journal under the Agent’s state directory. If a probe is interrupted by a restart, a network failure, or an ambiguous response, the recorded objects are deleted during later collections, a few per collection. Cleanup is best effort: it requires the same job name, mode, endpoints, buckets, prefix, and working credentials, and it cannot remove objects when the service or the permissions are gone. The number of objects awaiting cleanup is bounded. When that bound is reached, the collector keeps observing and cleaning up but stops creating new probe objects until a slot is free, and the mutation_backpressure chart shows the pause.

Replication modes poll the destination once per collection until the configured timeout. Each replication job has objectives, the propagation time you consider healthy, and timeouts, the point at which the probe is declared failed and its object is scheduled for cleanup. Delete propagation is measured as the moment the destination stops serving the object; how the storage system reclaims space behind that is not observed. In aws_replication mode cleanup deletes only the exact object versions and delete markers the probe created.

This collector is supported on all platforms.

This collector supports collecting metrics from multiple instances of this integration, including remote instances.

Grant each endpoint only the operations its role needs, scoped to the configured bucket and, for object operations, to the configured prefix:

ModeSource endpointDestination endpoint
lifecycles3:GetBucketVersioning, s3:PutObject, s3:GetObject, s3:ListBucket, s3:DeleteObjectnot used
ceph_multisites3:GetBucketVersioning, s3:PutObject, s3:GetObject, s3:DeleteObjects3:GetBucketVersioning, s3:GetObject, s3:DeleteObject
aws_replications3:GetBucketVersioning, s3:GetReplicationConfiguration, s3:PutObject, s3:GetObject, s3:GetObjectVersion, s3:ListBucketVersions, s3:DeleteObject, s3:DeleteObjectVersions3:GetBucketVersioning, s3:GetObject, s3:ListBucketVersions, s3:DeleteObjectVersion

The collector never writes to a destination; its delete permission is used only to remove the replicated copy of a probe object during cleanup. When assume_role is configured, the base credentials also need sts:AssumeRole on that role.

Default Behavior

Auto-Detection

There is no auto-detection. A job needs a mode, a source region and bucket, credentials or an AWS SDK credential source on the Agent host, and a destination for the replication modes.

Limits

Each job runs one probe at a time and cleans up a few pending objects per collection. The number of objects tracked for cleanup is a fixed internal bound; when it is reached, new probes pause until cleanup frees a slot. Probe objects are 4 KiB. Replication objectives must be at least one collection interval, because the destination is polled once per collection, and objectives and timeouts are capped at 24 hours.

Performance Impact

The impact on the Agent is negligible. On the storage side each collection issues a handful of small requests under the configured prefix: at most one new probe object with its reads, listing, and delete, plus a few cleanup deletes. Use a dedicated bucket or reserved prefix, credentials restricted to it, and a collection interval that matches how often you want to exercise the service.

Setup

You can configure the s3check collector in two ways:

MethodBest forHow to
UIFast setup without editing filesGo to Nodes → Configure this node → Collectors → Jobs, search for s3check, then click + to add a job.
FileIf you prefer configuring via file, or need to automate deployments (e.g., with Ansible)Edit go.d/s3check.conf and add a job.

Important

UI configuration requires paid Netdata Cloud plan.

Prerequisites

Choose a mode and prepare the buckets

The collector verifies the bucket contract of its mode before every collection and refuses to run when it is not met:

ModeBucketsVersioningReplication policy
lifecycleone bucketnever enabled; a suspended state is also rejectednone
ceph_multisiteone bucket per zone, names may differnever enabled on either bucketthe zones must sync the bucket so that a key written at the source appears unchanged at the destination
aws_replicationone bucket per side, names may differenabled on both buckets, MFA Delete offan enabled replication rule on the source bucket that covers the prefix without tag filters, targets the destination bucket, and has delete marker replication enabled; no other enabled rule may replicate the prefix to a different bucket

The destination key is always identical to the source key. Policies that rename or re-prefix objects are not supported. A dedicated bucket is the simplest way to meet these requirements, but any bucket that satisfies them works.

Every probe object is written with a conditional If-None-Match: * request so that a probe can never overwrite an existing key; the service must support conditional writes.

Reserve an object prefix

Pick a prefix, by default netdata-s3check/, and do not store your own data under it. The collector creates and deletes objects only under <prefix><owner>/, where the owner segment is unique to this Agent and job, so several Agents and jobs can safely share a prefix and even a bucket. Keys outside the collector’s own namespace are never listed, read, or deleted.

Create credentials with the minimum permissions

Each endpoint authenticates independently. Omit credentials to use the AWS SDK default credential chain on the Agent host (environment variables, shared config and profiles, instance or task roles), or provide static keys, preferably as go.d secret references such as ${env:NAME} rather than plaintext. Optionally add assume_role to switch to a role with those base credentials.

Grant only the operations listed in the permissions table of the Overview, scoped to the bucket and, for object operations, to the prefix. When assume_role is configured, the base credentials must also be allowed to call sts:AssumeRole on that role.

Run the job where your clients run

Results reflect the network path of the Agent that runs the job: DNS, routing, proxies, TLS, and authentication are all part of the measurement. Run each job on the Agent that best represents the clients you care about. Replication checks cover one direction; add a second job with source and destination swapped to check the reverse direction.

Configuration

Options

Set mode and configure only the matching mode_* object; the other two must be absent. Durations accept human-readable values such as 10s, 5m, 1h, or 1d.

Endpoint settings. Every source and destination object takes the same connection settings:

  • endpoint, region, bucket, path_style: where the bucket lives. For AWS S3 leave endpoint empty and set path_style: no; for Ceph RGW and other S3-compatible services set the full URL and keep path-style addressing.
  • credentials: static access keys. Omit the object to use the AWS SDK default credential chain of the Agent host. Prefer go.d secret references such as ${env:NAME} over plaintext values.
  • assume_role: optional. The base credentials call STS AssumeRole and the returned temporary credentials are used for this endpoint only.
  • timeout, proxy_url, and the tls_* options apply to every S3 request and to the STS AssumeRole request of that endpoint. An empty proxy_url honors the HTTP_PROXY, HTTPS_PROXY, and NO_PROXY environment of the Agent.

Prefix. The collector only writes, reads, lists, and deletes keys under <prefix><owner>/, where the owner segment is derived from this Agent and job. Reserve the prefix for Netdata; several Agents and jobs may share it.

Objectives and timeouts (replication modes). An objective is the propagation time you consider healthy: exceeding it marks the write_visibility_objective or delete_visibility_objective chart as breached while the probe keeps waiting. A timeout is the hard limit: exceeding it fails the probe with reason visibility_timeout or delete_timeout and schedules the object for cleanup. Objectives must be at least update_every, because the destination is polled once per collection, timeouts must be at least their objective, and all four are capped at 24h.

GroupOptionDescriptionDefaultRequired
ModemodeWhich check this job runs: lifecycle, ceph_multisite, or aws_replication.lifecycleno
Mode / Lifecyclemode_lifecycleSettings of the lifecycle mode. Required when mode is lifecycle.no
mode_lifecycle.prefixKey prefix reserved for probe objects. Must end with /.netdata-s3check/no
Mode / Lifecycle connectionmode_lifecycle.sourceThe S3 endpoint to check.no
mode_lifecycle.source.nameLabel for this endpoint on charts.sourceno
mode_lifecycle.source.endpointService URL with scheme and host only. Leave empty for the AWS regional endpoint.no
mode_lifecycle.source.regionSigning region, for example us-east-1. Required, also for non-AWS services.no
mode_lifecycle.source.bucketBucket that holds the probe objects.no
mode_lifecycle.source.path_stylePut the bucket in the request path (https://host/bucket/key) instead of the hostname. Turn off for AWS S3.yesno
mode_lifecycle.source.credentialsStatic access keys for this endpoint. Leave empty to use the AWS SDK default credential chain.no
mode_lifecycle.source.credentials.access_key_idAccess key ID of the static credentials.no
mode_lifecycle.source.credentials.secret_access_keySecret access key of the static credentials.no
mode_lifecycle.source.credentials.session_tokenSession token for temporary credentials. Leave empty for long-lived keys.no
mode_lifecycle.source.assume_roleIAM role to assume with this endpoint’s base credentials.no
mode_lifecycle.source.assume_role.role_arnARN of the IAM role to assume.no
mode_lifecycle.source.assume_role.external_idExternal ID expected by the role’s trust policy, if any.no
mode_lifecycle.source.timeoutMaximum duration of one S3 or STS request, for example 10s. Up to 1m.10sno
mode_lifecycle.source.proxy_urlHTTP proxy for this endpoint’s requests. Leave empty to use the proxy environment variables.no
mode_lifecycle.source.tls_skip_verifySkip TLS certificate and hostname verification. Insecure.nono
mode_lifecycle.source.tls_caAbsolute path to a CA bundle for a private certificate authority.no
mode_lifecycle.source.tls_certAbsolute path to the client certificate. Requires the client key.no
mode_lifecycle.source.tls_keyAbsolute path to the client certificate key. Requires the client certificate.no
Mode / Ceph multisitemode_ceph_multisiteSettings of the ceph_multisite mode. Required when mode is ceph_multisite.no
mode_ceph_multisite.prefixKey prefix reserved for probe objects. Must end with /.netdata-s3check/no
Mode / Ceph multisite sourcemode_ceph_multisite.sourceWhere probe objects are written and deleted.no
mode_ceph_multisite.source.nameLabel for this endpoint on charts.sourceno
mode_ceph_multisite.source.endpointService URL with scheme and host only. Leave empty for the AWS regional endpoint.no
mode_ceph_multisite.source.regionSigning region, for example us-east-1. Required, also for non-AWS services.no
mode_ceph_multisite.source.bucketBucket that holds the probe objects.no
mode_ceph_multisite.source.path_stylePut the bucket in the request path (https://host/bucket/key) instead of the hostname. Turn off for AWS S3.yesno
mode_ceph_multisite.source.credentialsStatic access keys for this endpoint. Leave empty to use the AWS SDK default credential chain.no
mode_ceph_multisite.source.credentials.access_key_idAccess key ID of the static credentials.no
mode_ceph_multisite.source.credentials.secret_access_keySecret access key of the static credentials.no
mode_ceph_multisite.source.credentials.session_tokenSession token for temporary credentials. Leave empty for long-lived keys.no
mode_ceph_multisite.source.assume_roleIAM role to assume with this endpoint’s base credentials.no
mode_ceph_multisite.source.assume_role.role_arnARN of the IAM role to assume.no
mode_ceph_multisite.source.assume_role.external_idExternal ID expected by the role’s trust policy, if any.no
mode_ceph_multisite.source.timeoutMaximum duration of one S3 or STS request, for example 10s. Up to 1m.10sno
mode_ceph_multisite.source.proxy_urlHTTP proxy for this endpoint’s requests. Leave empty to use the proxy environment variables.no
mode_ceph_multisite.source.tls_skip_verifySkip TLS certificate and hostname verification. Insecure.nono
mode_ceph_multisite.source.tls_caAbsolute path to a CA bundle for a private certificate authority.no
mode_ceph_multisite.source.tls_certAbsolute path to the client certificate. Requires the client key.no
mode_ceph_multisite.source.tls_keyAbsolute path to the client certificate key. Requires the client certificate.no
Mode / Ceph multisite destinationmode_ceph_multisite.destinationWhere replicated copies must appear. The collector only reads there and cleans up its own copies; it never creates objects.no
mode_ceph_multisite.destination.nameLabel for this endpoint on charts.destinationno
mode_ceph_multisite.destination.endpointService URL with scheme and host only. Leave empty for the AWS regional endpoint.no
mode_ceph_multisite.destination.regionSigning region, for example us-east-1. Required, also for non-AWS services.no
mode_ceph_multisite.destination.bucketBucket that holds the probe objects.no
mode_ceph_multisite.destination.path_stylePut the bucket in the request path (https://host/bucket/key) instead of the hostname. Turn off for AWS S3.yesno
mode_ceph_multisite.destination.credentialsStatic access keys for this endpoint. Leave empty to use the AWS SDK default credential chain.no
mode_ceph_multisite.destination.credentials.access_key_idAccess key ID of the static credentials.no
mode_ceph_multisite.destination.credentials.secret_access_keySecret access key of the static credentials.no
mode_ceph_multisite.destination.credentials.session_tokenSession token for temporary credentials. Leave empty for long-lived keys.no
mode_ceph_multisite.destination.assume_roleIAM role to assume with this endpoint’s base credentials.no
mode_ceph_multisite.destination.assume_role.role_arnARN of the IAM role to assume.no
mode_ceph_multisite.destination.assume_role.external_idExternal ID expected by the role’s trust policy, if any.no
mode_ceph_multisite.destination.timeoutMaximum duration of one S3 or STS request, for example 10s. Up to 1m.10sno
mode_ceph_multisite.destination.proxy_urlHTTP proxy for this endpoint’s requests. Leave empty to use the proxy environment variables.no
mode_ceph_multisite.destination.tls_skip_verifySkip TLS certificate and hostname verification. Insecure.nono
mode_ceph_multisite.destination.tls_caAbsolute path to a CA bundle for a private certificate authority.no
mode_ceph_multisite.destination.tls_certAbsolute path to the client certificate. Requires the client key.no
mode_ceph_multisite.destination.tls_keyAbsolute path to the client certificate key. Requires the client certificate.no
Mode / Ceph multisite objectivesmode_ceph_multisite.write_objectiveHealthy time for a new object to appear at the destination, for example 15m. Exceeding it flags the objective chart only.15mno
mode_ceph_multisite.write_timeoutHard limit for a new object to appear at the destination, for example 30m. Exceeding it fails the probe.30mno
mode_ceph_multisite.delete_objectiveHealthy time for a source delete to hide the destination copy, for example 5m.5mno
mode_ceph_multisite.delete_timeoutHard limit for the destination copy to disappear, for example 15m. Exceeding it fails the probe.15mno
Mode / AWS replicationmode_aws_replicationSettings of the aws_replication mode. Required when mode is aws_replication.no
mode_aws_replication.prefixKey prefix reserved for probe objects. Must end with /.netdata-s3check/no
Mode / AWS replication sourcemode_aws_replication.sourceWhere probe objects are written and deleted.no
mode_aws_replication.source.nameLabel for this endpoint on charts.sourceno
mode_aws_replication.source.endpointService URL with scheme and host only. Leave empty for the AWS regional endpoint.no
mode_aws_replication.source.regionSigning region, for example us-east-1. Required, also for non-AWS services.no
mode_aws_replication.source.bucketBucket that holds the probe objects.no
mode_aws_replication.source.path_stylePut the bucket in the request path (https://host/bucket/key) instead of the hostname. Turn off for AWS S3.yesno
mode_aws_replication.source.credentialsStatic access keys for this endpoint. Leave empty to use the AWS SDK default credential chain.no
mode_aws_replication.source.credentials.access_key_idAccess key ID of the static credentials.no
mode_aws_replication.source.credentials.secret_access_keySecret access key of the static credentials.no
mode_aws_replication.source.credentials.session_tokenSession token for temporary credentials. Leave empty for long-lived keys.no
mode_aws_replication.source.assume_roleIAM role to assume with this endpoint’s base credentials.no
mode_aws_replication.source.assume_role.role_arnARN of the IAM role to assume.no
mode_aws_replication.source.assume_role.external_idExternal ID expected by the role’s trust policy, if any.no
mode_aws_replication.source.timeoutMaximum duration of one S3 or STS request, for example 10s. Up to 1m.10sno
mode_aws_replication.source.proxy_urlHTTP proxy for this endpoint’s requests. Leave empty to use the proxy environment variables.no
mode_aws_replication.source.tls_skip_verifySkip TLS certificate and hostname verification. Insecure.nono
mode_aws_replication.source.tls_caAbsolute path to a CA bundle for a private certificate authority.no
mode_aws_replication.source.tls_certAbsolute path to the client certificate. Requires the client key.no
mode_aws_replication.source.tls_keyAbsolute path to the client certificate key. Requires the client certificate.no
Mode / AWS replication destinationmode_aws_replication.destinationWhere replicated copies must appear. The collector only reads there and cleans up its own copies; it never creates objects.no
mode_aws_replication.destination.nameLabel for this endpoint on charts.destinationno
mode_aws_replication.destination.endpointService URL with scheme and host only. Leave empty for the AWS regional endpoint.no
mode_aws_replication.destination.regionSigning region, for example us-east-1. Required, also for non-AWS services.no
mode_aws_replication.destination.bucketBucket that holds the probe objects.no
mode_aws_replication.destination.path_stylePut the bucket in the request path (https://host/bucket/key) instead of the hostname. Turn off for AWS S3.yesno
mode_aws_replication.destination.credentialsStatic access keys for this endpoint. Leave empty to use the AWS SDK default credential chain.no
mode_aws_replication.destination.credentials.access_key_idAccess key ID of the static credentials.no
mode_aws_replication.destination.credentials.secret_access_keySecret access key of the static credentials.no
mode_aws_replication.destination.credentials.session_tokenSession token for temporary credentials. Leave empty for long-lived keys.no
mode_aws_replication.destination.assume_roleIAM role to assume with this endpoint’s base credentials.no
mode_aws_replication.destination.assume_role.role_arnARN of the IAM role to assume.no
mode_aws_replication.destination.assume_role.external_idExternal ID expected by the role’s trust policy, if any.no
mode_aws_replication.destination.timeoutMaximum duration of one S3 or STS request, for example 10s. Up to 1m.10sno
mode_aws_replication.destination.proxy_urlHTTP proxy for this endpoint’s requests. Leave empty to use the proxy environment variables.no
mode_aws_replication.destination.tls_skip_verifySkip TLS certificate and hostname verification. Insecure.nono
mode_aws_replication.destination.tls_caAbsolute path to a CA bundle for a private certificate authority.no
mode_aws_replication.destination.tls_certAbsolute path to the client certificate. Requires the client key.no
mode_aws_replication.destination.tls_keyAbsolute path to the client certificate key. Requires the client certificate.no
Mode / AWS replication objectivesmode_aws_replication.write_objectiveHealthy time for a new object to appear at the destination, for example 15m. Exceeding it flags the objective chart only.15mno
mode_aws_replication.write_timeoutHard limit for a new object to appear at the destination, for example 30m. Exceeding it fails the probe.30mno
mode_aws_replication.delete_objectiveHealthy time for a source delete to hide the destination copy, for example 5m.5mno
mode_aws_replication.delete_timeoutHard limit for the destination copy to disappear, for example 15m. Exceeding it fails the probe.15mno
Collectionupdate_everyData collection interval, in seconds. Replication modes poll the destination once per interval.120no
autodetection_retryHow often to retry the initial bucket check when the job fails to start, in seconds. Zero disables retries.0no
Virtual NodevnodeVirtual Node that owns the charts of this job.no

mode

Select the mode first and configure only its mode_* object. A job whose mode object is missing, or that also sets another mode’s object, is rejected. The Overview compares the three modes.

mode_lifecycle

Each collection writes one 4 KiB probe object under the prefix, reads it back and verifies its content, confirms it appears in a listing, deletes it, and confirms it is gone. A successful collection leaves nothing behind.

The bucket must never have had versioning enabled; a suspended state is also rejected, because a plain delete must remove the object completely. Required permissions on the bucket: s3:GetBucketVersioning, s3:PutObject, s3:GetObject, s3:ListBucket, s3:DeleteObject.

mode_ceph_multisite

Writes a probe object at the source zone, polls the destination zone once per collection until the same key is readable there with the same content, deletes it at the source, and polls until the destination no longer serves it. Write lag is measured from the source write, delete lag from the source delete.

Both buckets must be unversioned. The source needs s3:GetBucketVersioning, s3:PutObject, s3:GetObject, s3:DeleteObject; the destination needs s3:GetBucketVersioning, s3:GetObject, s3:DeleteObject. The collector never writes to the destination and deletes there only its own replicated copies during cleanup. Use one job per direction.

mode_aws_replication

Follows AWS S3 replication semantics: the source write creates an object version, the source delete creates a delete marker, and both are expected to replicate. The collector records the exact version and marker IDs on both sides and deletes only those during cleanup.

Both buckets need versioning enabled with MFA Delete off. The source bucket needs an enabled replication rule that covers the prefix without tag filters, targets the destination bucket, and has delete marker replication enabled; no other enabled rule may replicate the prefix to a different bucket. Source permissions: s3:GetBucketVersioning, s3:GetReplicationConfiguration, s3:PutObject, s3:GetObject, s3:GetObjectVersion, s3:ListBucketVersions, s3:DeleteObject, s3:DeleteObjectVersion. Destination permissions: s3:GetBucketVersioning, s3:GetObject, s3:ListBucketVersions, s3:DeleteObjectVersion. Use one job per direction.

via UI

Configure the s3check collector from the Netdata web interface:

  1. Go to Nodes.
  2. Select the node where you want the s3check data-collection job to run and click the :gear: (Configure this node). That node will run the data collection.
  3. The Collectors → Jobs view opens by default.
  4. In the Search box, type s3check (or scroll the list) to locate the s3check collector.
  5. Click the + next to the s3check collector to add a new job.
  6. Fill in the job fields, then click Test to verify the configuration and Submit to save.
    • Test runs the job with the provided settings and shows whether data can be collected.
    • If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.

via File

The configuration file name for this integration is go.d/s3check.conf.

The file format is YAML. Generally, the structure is:

update_every: 1
autodetection_retry: 0
jobs:
  - name: some_name1
  - name: some_name2

You can edit the configuration file using the edit-config script from the Netdata config directory.

cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
sudo ./edit-config go.d/s3check.conf
Examples
Lifecycle check of an S3-compatible endpoint

Checks one Ceph RGW or other S3-compatible endpoint with static access keys read from environment variables. The bucket must never have had versioning enabled.

jobs:
  - name: s3_lifecycle
    mode: lifecycle
    mode_lifecycle:
      source:
        endpoint: https://s3.example.net
        region: us-east-1
        bucket: netdata-s3check
        credentials:
          access_key_id: ${env:NETDATA_S3CHECK_ACCESS_KEY_ID}
          secret_access_key: ${env:NETDATA_S3CHECK_SECRET_ACCESS_KEY}
        path_style: yes
Lifecycle check through a proxy with a private CA

The proxy and the CA bundle apply to this endpoint’s S3 requests (and to STS if assume_role were set). Credentials are omitted, so the AWS SDK default credential chain on the Agent host is used.

jobs:
  - name: s3_lifecycle_proxied
    mode: lifecycle
    mode_lifecycle:
      prefix: monitoring/netdata/
      source:
        endpoint: https://s3.internal.example.net
        region: us-east-1
        bucket: netdata-s3check
        proxy_url: http://proxy.example.net:3128
        tls_ca: /etc/ssl/private-ca.pem
        timeout: 15s
Ceph RGW multisite, one direction

Measures replication from zone A to zone B. Add a second job with source and destination swapped to check the reverse direction. Both buckets must be unversioned, and each zone uses its own keys.

jobs:
  - name: ceph_site_a_to_site_b
    mode: ceph_multisite
    mode_ceph_multisite:
      source:
        name: site-a
        endpoint: https://rgw-site-a.example.net
        region: us-east-1
        bucket: netdata-s3check
        credentials:
          access_key_id: ${env:NETDATA_S3CHECK_SITE_A_ACCESS_KEY_ID}
          secret_access_key: ${env:NETDATA_S3CHECK_SITE_A_SECRET_ACCESS_KEY}
      destination:
        name: site-b
        endpoint: https://rgw-site-b.example.net
        region: us-east-1
        bucket: netdata-s3check
        credentials:
          access_key_id: ${env:NETDATA_S3CHECK_SITE_B_ACCESS_KEY_ID}
          secret_access_key: ${env:NETDATA_S3CHECK_SITE_B_SECRET_ACCESS_KEY}
      write_objective: 15m
      write_timeout: 30m
      delete_objective: 5m
      delete_timeout: 15m
AWS S3 replication with a role per account

Checks replication from a bucket in one region to a bucket in another. Base credentials come from the AWS SDK default credential chain on the Agent host, for example an instance profile, and each endpoint then assumes its own role. Both buckets must be versioned and the source rule must replicate delete markers. path_style is disabled because AWS S3 uses virtual-hosted-style addressing.

jobs:
  - name: aws_replication
    mode: aws_replication
    mode_aws_replication:
      source:
        name: us-east-1
        region: us-east-1
        bucket: source-bucket
        path_style: no
        assume_role:
          role_arn: arn:aws:iam::[ACCOUNT]:role/netdata-s3-source
      destination:
        name: us-west-2
        region: us-west-2
        bucket: destination-bucket
        path_style: no
        assume_role:
          role_arn: arn:aws:iam::[ACCOUNT]:role/netdata-s3-destination

Metrics

Metrics grouped by scope.

The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.

All charts carry the mode, source, and destination labels (destination is empty in lifecycle mode). Work that did not happen during a collection leaves a gap rather than a zero, so a missing point means “not measured”. Call counters count logical S3 operations as the collector issued them; retries performed inside the AWS SDK are not counted separately. The reason label uses a fixed vocabulary: none, request, payload_mismatch, visibility_timeout, delete_timeout, cleanup, ownership, internal. Raw provider errors appear only in the collector log.

Per job status

Outcome of the collection and of the current or last completed probe.

Labels:

LabelDescription
modeConfigured mode.
sourceDisplay name of the source endpoint.
destinationDisplay name of the destination endpoint; empty in lifecycle mode.
reasonWhy the outcome failed; none otherwise.

Metrics:

MetricDescriptionDimensionsUnit
s3check.runtime_statusS3 check runtime statussuccess, failedstatus
s3check.probe_statusS3 check probe statussuccess, waiting, failedstatus
s3check.payload_mismatchS3 payload verification mismatchmismatchstatus

Per replication objectives

How long a source change took to become visible at the destination and whether that exceeded the configured objective. Replication modes only.

Labels:

LabelDescription
modeConfigured replication mode.
sourceDisplay name of the source endpoint.
destinationDisplay name of the destination endpoint.
reasonWhy the probe failed; none otherwise.

Metrics:

MetricDescriptionDimensionsUnit
s3check.write_visibility_lagS3 write visibility laglagseconds
s3check.write_visibility_objectiveS3 write visibility objective statusbreachedstatus
s3check.delete_visibility_lagS3 delete visibility laglagseconds
s3check.delete_visibility_objectiveS3 delete visibility objective statusbreachedstatus

Per ownership

Probe objects still awaiting cleanup and whether new probes are paused because the cleanup bound is reached.

Labels:

LabelDescription
modeConfigured mode.
sourceDisplay name of the source endpoint.
destinationDisplay name of the destination endpoint; empty in lifecycle mode.

Metrics:

MetricDescriptionDimensionsUnit
s3check.cleanup_pending_objectsS3 owned objects pending cleanuppendingobjects
s3check.mutation_backpressureS3 mutation backpressure statusactivestatus

Per operation

One logical S3 operation class on one endpoint: setup (precondition checks), put, read, list, delete, write_visibility, delete_visibility, reconcile, or cleanup.

Labels:

LabelDescription
modeConfigured mode.
endpointsource or destination.
operationLogical operation name.
sourceDisplay name of the source endpoint.
destinationDisplay name of the destination endpoint; empty in lifecycle mode.
reasonWhy the operation failed; none otherwise. Not present on s3check.operation_calls.

Metrics:

MetricDescriptionDimensionsUnit
s3check.operation_statusS3 logical operation statussuccess, failedstatus
s3check.operation_durationS3 logical operation durationdurationseconds
s3check.operation_callsS3 logical callscalls, failurescalls/s

Alerts

The following alerts are available:

Alert nameOn metricDescription
s3check_runtime_faileds3check.runtime_statusThe collector could not validate provider safety or persist ownership state. Inspect the collector log for the raw provider or local-state error.
s3check_probe_faileds3check.probe_statusThe active or last terminal probe failed with a bounded reason. Raw provider errors are available only in the collector log.
s3check_payload_mismatchs3check.payload_mismatchThe verification read returned different bytes for the exact key written by the probe.
s3check_write_visibility_objectives3check.write_visibility_objectiveThe exact source key has not become readable with the expected payload at ${label:destination} within write_objective.
s3check_delete_visibility_objectives3check.delete_visibility_objectiveThe destination current object remains readable after the source delete for longer than delete_objective.
s3check_mutation_backpressures3check.mutation_backpressureThe bounded ownership queue is full. Cleanup and observation continue, but no new probe object is created until a slot is released.

Troubleshooting

Diagnostics

Debug Mode

Important: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.

To troubleshoot issues with the s3check collector, run the go.d.plugin with the debug option enabled. The output should give you clues as to why the collector isn’t working.

  • Navigate to the plugins.d directory, usually at /usr/libexec/netdata/plugins.d/. If that’s not the case on your system, open netdata.conf and look for the plugins setting under [directories].

    cd /usr/libexec/netdata/plugins.d/
    
  • Switch to the netdata user.

    sudo -u netdata -s
    
  • Run the go.d.plugin to debug the collector:

    ./go.d.plugin -d -m s3check
    

    To debug a specific job:

    ./go.d.plugin -d -m s3check -j jobName
    

Getting Logs

If you’re encountering problems with the s3check collector, follow these steps to retrieve logs and identify potential issues:

  • Run the command specific to your system (systemd, non-systemd, or Docker container).
  • Examine the output for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
System with systemd

Use the following command to view logs generated since the last Netdata service restart:

journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep s3check
System without systemd

Locate the collector log file, typically at /var/log/netdata/collector.log, and use grep to filter for collector’s name:

grep s3check /var/log/netdata/collector.log

Note: This method shows logs from all restarts. Focus on the latest entries for troubleshooting current issues.

Docker Container

If your Netdata runs in a Docker container named “netdata” (replace if different), use this command:

docker logs netdata 2>&1 | grep s3check

Other Problems

The job fails its initial check

Before writing anything, the collector reads the bucket versioning state of every endpoint and, in aws_replication mode, the replication rules of the source bucket. A failure here means one of:

  • the endpoint is not reachable from the Agent (DNS, routing, proxy, TLS, or region);
  • the credentials are rejected or lack s3:GetBucketVersioning (or s3:GetReplicationConfiguration);
  • the bucket does not meet the mode’s versioning contract: never versioned for lifecycle and ceph_multisite, versioning enabled with MFA Delete off for aws_replication;
  • in aws_replication mode, no enabled rule covers the prefix and targets the destination bucket with delete marker replication, or another enabled rule replicates the prefix elsewhere.

Chart labels carry only a bounded reason; the exact provider error is in the collector log.

Objects remain pending cleanup or new probes are paused

After an interrupted probe the collector deletes the objects it recorded a few per collection. cleanup_pending_objects shows the backlog and mutation_backpressure turns active when the backlog reaches its bound, which pauses new probes but not observation or cleanup.

Cleanup needs the same job name, mode, prefix, endpoints, buckets, and working credentials that created the objects. If you changed the mode, prefix, endpoint, or bucket while objects were pending, the job refuses to start with an ownership mismatch error; if you renamed the job, the old objects are left behind. In both cases restore the previous configuration under the previous job name until the backlog drains, then change it. A cleanup reason means a probe object could not be confirmed gone, usually a permission problem; an ownership reason means the local state could not be locked or saved. The collector log has the details.

Replication probes stay in waiting or time out

A replication probe reports waiting until the destination serves the object (write) or stops serving it (delete), and fails with visibility_timeout or delete_timeout when the configured timeout passes. Check that the replication policy or zone sync covers the prefix, that the destination credentials can read the bucket, and that the policy delivers the object under exactly the same key. Persistent waits usually mean replication is slower than your objectives and timeouts assume; align them with what the replication setup promises.

The observability platform companies need to succeed

Sign up for free

Want a personalised demo of Netdata for your use case?

Contact Sales