The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Amazon CloudWatch icon

Amazon CloudWatch

Amazon CloudWatch

Plugin: go.d.plugin Module: cloudwatch

Overview

Monitor AWS infrastructure through Amazon CloudWatch. This collector discovers CloudWatch metrics for a curated set of AWS services and renders them as Netdata charts, with minimal configuration.

Monitored services:

  • Amazon EC2 (compute)
  • Amazon RDS (relational databases)
  • Elastic Load Balancing – Classic (ELB), Application (ALB), and Network (NLB) load balancers
  • Amazon S3 (object storage)
  • AWS Lambda (serverless functions)
  • Amazon SQS (message queues)
  • Amazon DynamoDB (NoSQL databases)
  • Amazon API Gateway (REST APIs)
  • AWS Step Functions (workflow orchestration)
  • NAT Gateway (VPC networking)
  • AWS PrivateLink endpoints (endpoint connections, traffic, and packet problems)
  • Amazon Kinesis Data Streams (streaming ingestion)
  • Amazon Data Firehose (delivery streams)
  • Amazon SNS (pub/sub messaging)
  • Amazon EBS (block storage volumes)
  • Amazon EFS (elastic file systems)
  • Amazon ECS (container services)
  • Amazon ElastiCache (in-memory cache)
  • Amazon OpenSearch Service (search and analytics)
  • Amazon DocumentDB (document database)
  • Amazon Redshift (data warehouse)
  • Amazon MSK (Kafka streaming)
  • Amazon CloudFront (content delivery network / CDN)
  • AWS Auto Scaling (EC2 Auto Scaling group capacity)
  • Amazon Bedrock (foundation-model invocations and tokens)
  • Amazon EventBridge (event rules)
  • AWS Site-to-Site VPN (VPN connections)
  • Amazon EKS (Kubernetes control plane: API server, scheduler, etcd)
  • AWS Billing (worldwide estimated month-to-date charges; opt-in profiles)

This collector queries runtime metrics from Amazon CloudWatch. It complements the AWS EC2 Compute instances integration, which exposes EC2 inventory and capacity information, and the AWS Quota integration, which exposes AWS Service Quotas. These integrations use different AWS data sources and do not replace one another.

CloudWatch coverage is defined by one or more profiles – YAML files declaring a CloudWatch namespace, an exact resource-dimension grain, supported regions, metrics, statistics, and chart template. A service can use multiple profiles when AWS publishes distinct grains, and coverage can be extended without collector code changes. See the AWS CloudWatch profile format for the complete schema and authoring rules.

Tip - Need a service that isn’t listed?

Request a profile – it’s just a YAML file, no code change. Open a feature request and attach the service’s CloudWatch metric schema, captured with this read-only command. It prints only metric and dimension names (no resource IDs, ARNs, or metric values), so the output is safe to share:

aws cloudwatch list-metrics --namespace "AWS/<Service>" --region <your-region> --output json \
  | jq -c '[.Metrics[] | {metric: .MetricName, dimensions: ([.Dimensions[].Name] | sort)}] | unique'

Replace AWS/<Service> with the service namespace (for example AWS/AmazonMQ) and <your-region> with a Region where the service runs. The exact metrics and dimensions in the output are what we need to author a correct profile quickly.

The collector compiles named credential sources, monitored targets, and ordered collection rules into an immutable runtime plan. Profiles with identifying dimensions discover available metrics with one CloudWatch ListMetrics scan per target, region, and namespace, using the least restrictive recent-activity policy required by the participating selected series, then apply each profile’s exact dimension matcher. A profile whose dimensions are all constants compiles a known static instance and skips ListMetrics. Optional resource-tag predicates are resolved with the Resource Groups Tagging API before query expansion. Each selected series has a resolved query policy: aggregation period, rolling lookback, and publication delay. GetMetricData searches the aligned rolling window for the newest complete finite datapoint, while Netdata receives the retained numeric value on every collection cycle. Each target resolves its AWS account id through sts:GetCallerIdentity; target names remain distinct execution identities even when they resolve to the same account. Rule order, then target order within each rule, resolves overlap: the first matching rule/target owns each overlapping exported metric/statistic series.

This collector is supported on all platforms.

This collector supports collecting metrics from multiple instances of this integration, including remote instances.

Every target requires cloudwatch:GetMetricData. It also requires cloudwatch:ListMetrics when any selected profile has an identifying dimension and therefore needs discovery; an all-constant profile such as billing_total is queried directly. The collector calls sts:GetCallerIdentity for account attribution, but AWS does not require an explicit permission grant for that operation. A credential source used by a target with assume_role additionally requires sts:AssumeRole for that role ARN. Resource tag filtering (rule_defaults.filters.resource_tags or rules[].filters.resource_tags) and resource tag labels (labels.resource_tags) additionally require tag:GetResources.

Default Behavior

Auto-Detection

A rule that omits profiles selects all default-enabled profiles for its targets and regions. A rule that omits metrics collects every metric exported by those profiles. When present, metrics groups exact AWS MetricNames by profile; group statistics are inherited by included metrics unless a metric supplies a replacement list. The collector emits charts only for profiles with live metrics. Discovery and the query blueprint are cached; discovery refreshes every discovery.refresh_every seconds (default 300).

Limits

  • Minimum collection interval is 60 seconds (CloudWatch’s minimum metric period).
  • Query timing resolves field-by-field from rules[].query, rule_defaults.query, nested metric query defaults, nested profile query defaults, and the built-in 10-minute publication delay. The stock daily S3 storage profile uses a conservative one-day collector policy. AWS documents that S3 storage metrics are reported once per day, but does not guarantee publication within one day.
  • A successful sparse query can replay its newest eligible CloudWatch value for up to lookback. During transient AWS failures, the retained value can be replayed longer, until a successful query replaces or expires it.
  • Query-plan preflight rejects more than 20,000 selected series, 600,000 all-due datapoints, or 40 packed GetMetricData requests before allocating AWS query structures. Billing units of up to five statistics for one structural metric are kept in one request.
  • limits.max_instances defaults to 1000 distinct final static or discovered instances after tag filtering and overlap resolution. Exceeding it rejects the refreshed query plan; instances are never silently truncated.
  • limits.max_discovery_groups defaults to 64 unique (target, region, namespace) groups per job. Compatible rules and profiles share a group. The safeguard catches accidental expansion and can be raised to the hard maximum of 100; larger collection must be split across jobs because one refresh can admit at most 100 groups that reach ListMetrics.
  • Each discovery group is independently bounded to 100 ListMetrics pages, 50,000 scanned metrics, 1,000,000 residual same-shape profile matches, and 20,000 candidate instances. Overflow fails the group without replacing its previous snapshot.
  • One discovery refresh is additionally bounded to 100 admitted ListMetrics SDK operations, 50,000 scanned metrics, 1,000,000 residual profile matches, 20,000 retained candidates, 64 MiB of conservatively weighted candidate storage, and one shared timeout. Every non-skipped group that resolves a client runs its first admitted operation before continuations share the remaining budget; skipped groups and client-resolution failures consume no operation budget. Successful replacements and failed-group carry-forward are rechecked as one bounded effective snapshot before installation. The AWS SDK can retry each admitted operation up to five wire attempts.
  • Aggregate discovery-limit, merged effective-snapshot limit, or timeout exhaustion discards the attempted refresh atomically. An existing snapshot remains active and discovery retries after discovery.refresh_every. On the first pass, any executable all-constant static profiles continue while dynamic discovery waits for its retry; without static work, total failure makes the collection attempt fail. Parent cancellation always aborts without advancing state.
  • Resources are labeled by their identifying CloudWatch dimensions (for example EC2 instance_id). Selected resource tags can additionally be attached as non-identity labels via labels.resource_tags; changing those tags updates labels without changing chart identity. (A dimension that is constant across resources, such as CloudFront’s Region=Global, is used to match and query metrics but is not turned into a label.)

Performance Impact

AWS bills CloudWatch API usage. GetMetricData is the cost driver; ListMetrics discovery falls under the free tier and then costs a fraction as much. As a rough anchor, GetMetricData is billed at roughly $0.01 per 1,000 metrics requested – confirm current CloudWatch pricing for your region. The collector sends one query per selected metric/statistic series, but for AWS billing up to five statistics requested for the same metric count as one metric request. Point-aware batching keeps each such five-statistic billing unit in one request. In normal operation each series is queried once per newly eligible effective-period window, not once per Netdata collection cycle. A transient request or result failure retries that series after one update_every; subsequent delays double within the same eligible window and are capped at its effective period. Those retries are also billable, so update_every affects failure-time cost even though it does not set normal query cadence. Cost otherwise scales with selected targets, instances, metrics, statistics beyond AWS’s grouping, periods, and lookback response work. Longer lookbacks increase requested datapoints and can disable CloudWatch’s three-hour recently-active discovery filter. The collector minimizes work with curated profiles, exact metric and resource-tag selection, exact dimension matching, single-statistic defaults, completed-sibling isolation, bounded retry backoff, billing-group-preserving batches, shared compatible discovery scans, cached discovery/query plans, and recently_active_only. To reduce cost further, narrow rules[].targets, rules[].profiles, rules[].metrics, rules[].regions, or configure resource tag filters.

The opt-in Billing profiles use a 10-minute period, so each selected single-statistic Billing series normally produces 144 billable metric requests per day before retries. For example, 200 selected Billing series produce 28,800 metric requests per day. Their 24-hour lookback reserves 144 datapoint slots per query (28,800 slots when all 200 are due), but AWS charges GetMetricData by metrics requested, not by those reserved datapoint slots. The total profile is static and performs no ListMetrics; the three dynamic Billing grains share one namespace discovery stream per target and refresh interval before pagination and SDK retries. Billing cardinality grows with services, linked accounts, and observed account/service pairs, so select only the grains you need.

Setup

You can configure the cloudwatch collector in two ways:

MethodBest forHow to
UIFast setup without editing filesGo to Nodes → Configure this node → Collectors → Jobs, search for cloudwatch, then click + to add a job.
FileIf you prefer configuring via file, or need to automate deployments (e.g., with Ansible)Edit go.d/cloudwatch.conf and add a job.

Important

UI configuration requires paid Netdata Cloud plan.

Prerequisites

Create an AWS IAM identity with CloudWatch read access

The collector needs an IAM identity (user or role) allowed to read CloudWatch metrics. It resolves the AWS account identity with GetCallerIdentity, which does not require an explicit permission grant.

Attach a policy such as:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "cloudwatch:ListMetrics",
        "cloudwatch:GetMetricData"
      ],
      "Resource": "*"
    }
  ]
}

cloudwatch:ListMetrics and cloudwatch:GetMetricData do not support resource-level permissions, so "Resource": "*" is required – this is already least-privilege for these read actions. A job that selects only profiles whose dimensions are all constants, such as billing_total, can omit cloudwatch:ListMetrics; every dynamic profile needs it. GetCallerIdentity needs no explicit grant. Scope sts:AssumeRole to the specific role ARN(s) rather than *. To enable resource tag filtering or labels, also grant tag:GetResources (it likewise requires "Resource": "*").

Define one or more named credential sources:

  • default – the AWS SDK default credential chain (environment variables, shared config/credentials files, EC2 instance profile, or EKS IRSA). Recommended when Netdata runs inside AWS.
  • static – an explicit access key ID and secret access key, plus an optional session token. Use go.d secret references rather than plaintext values.

A target can use either source directly or use it to assume one IAM role. If the role trust policy requires an external ID, the role owner supplies that value; it is not an AWS password or access key. See AWS guidance for third-party access.

Enable CloudWatch Billing metrics before selecting Billing profiles

The Billing profiles are opt-in. Before collecting them, enable Receive CloudWatch Billing Alerts in AWS Billing Preferences as described in AWS’s Billing alarm documentation.

This setup action is separate from the collector’s runtime IAM policy:

  • Enabling the preference requires the account root user or an IAM principal allowed to view Billing information. The collector identity still needs only the CloudWatch read permissions described above.
  • Once enabled, AWS says Billing metric data collection cannot be disabled. Deleting CloudWatch alarms does not disable the metric feed.
  • The first data normally appears about 15 minutes after first enablement. AWS then calculates and publishes estimates several times daily, so a working job can legitimately have no chart or can hold the last published value between updates.
  • Billing data is published only in us-east-1, represents worldwide charges, and is reported only in USD. The charts show the latest published estimate for the current month, not a forecast.
  • For consolidated billing, enable the preference in the management/payer account. That account can expose the consolidated total plus linked-account views; a standalone/member view can contain fewer grains. If the management/payer account changes, enable the preference again in the new account.
  • AWS does not publish these Billing metrics for Amazon Partner Network (APN) accounts.

The collector’s Billing profiles use a 24-hour retrieval window and sample-and-hold the newest successful value. Around the UTC month boundary, a prior-month observation can remain visible until a successful query replaces or expires it; a transient AWS failure can retain it longer under the collector’s general replay policy. Treat the value as latest-published month-to-date data, not an invoice or real-time ledger.

Configuration

Options

The following options can be defined globally or per job.

Profile file locations:

TypePath
Stock profiles/usr/lib/netdata/conf.d/go.d/cloudwatch.profiles/default/
User overrides/etc/netdata/go.d/cloudwatch.profiles/

A user profile file with the same basename as a stock profile overrides it.

GroupOptionDescriptionDefaultRequired
Collectionupdate_everyData collection interval (seconds). Must be at least 60 (CloudWatch’s minimum period).60no
autodetection_retryRecheck interval (seconds) when the job fails to start. Default 0 means no retry; set a positive value to keep retrying.0no
timeoutAWS operation timeout (seconds). Identity, resource-tag, and query operations use it for their operation scope; discovery shares one timeout across its whole refresh stage.30no
AuthenticationcredentialsList of named credential sources. Every source has a type of default (AWS SDK default chain) or static (explicit access/session credentials in type_static). Credential sources are reusable by targets and every defined source must be used.yes
credentials[].nameCredential source name referenced by targets.yes
credentials[].typeCredential source type: default uses the AWS SDK chain; static requires type_static.yes
credentials[].type_staticConfiguration used only when the credential source type is static.no
credentials[].type_static.access_key_idAWS access key ID. Required in type_static. Use a go.d secret reference such as ${env:AWS_ACCESS_KEY_ID}.no
credentials[].type_static.secret_access_keyAWS secret access key. Required in type_static. Use a go.d secret reference; do not store plaintext credentials in the file.no
credentials[].type_static.session_tokenOptional AWS session token in type_static for temporary credentials. Use a go.d secret reference.no
TargetstargetsUp to 64 named monitored AWS identities. A target uses one credential source directly or uses that source to assume one role. Targets remain distinct even when they resolve to the same AWS account.yes
targets[].nameTarget name referenced by collection rules.yes
targets[].credentialsName of the credential source used by this target.yes
targets[].assume_role.role_arnOptional IAM role ARN to assume using the target’s credential source.no
targets[].assume_role.external_idOptional value supplied by the role owner when the role trust policy requires an external ID. It is not an AWS password or access key.no
RulesrulesOrdered collection rules. Each rule selects targets, profiles, optional exact metrics, regions, and optional resource-tag filters. The earliest matching rule and target own each overlapping exported metric/statistic series.yes
rules[].nameUnique rule name used in diagnostics.yes
rules[].targetsOrdered names of monitored targets selected by this rule. Order breaks overlap ties within the rule.yes
rules[].profiles.defaultsInclude all default-enabled profiles. Defaults to true when profiles or defaults is omitted.yesno
rules[].profiles.includeProfile basenames to add explicitly. Set defaults to false to collect only this list, including profiles disabled by default. The higher-cardinality PrivateLink subnet view is privatelink_endpoint_subnet. Billing choices are billing_total, billing_service, billing_linked_account, and billing_linked_account_service.no
rules[].profiles.excludeProfile basenames to remove from the selected set. A profile cannot be both included and excluded.no
rules[].metricsOptional per-profile exact metric/statistic allowlists that narrow the profiles selected by this rule. Omit metrics to collect every metric exported by those profiles. The expanded rule supports at most 256 metric/statistic selections.no
rules[].metrics[].profileProfile basename. It must already be selected by rules[].profiles and may appear in only one metrics group per rule.yes
rules[].metrics[].statisticsOptional non-empty default AWS statistics inherited by included metrics that omit their own list. Named statistics are case-insensitive.no
rules[].metrics[].includeNon-empty list of exact, case-sensitive AWS CloudWatch MetricNames exported by this profile. Duplicate names are rejected.yes
rules[].metrics[].include[].nameExact, case-sensitive AWS CloudWatch MetricName exported by the profile.yes
rules[].metrics[].include[].statisticsOptional non-empty replacement for the group statistics. Use Average, Minimum, Maximum, Sum, SampleCount, or p<N>; named statistics are case-insensitive. A metric with no replacement must inherit a group default.no
rules[].regionsCanonical lowercase AWS region codes selected by this rule. The compiler intersects them with intrinsic profile restrictions; CloudFront and the Billing profiles support only us-east-1.yes
Query Policyrule_defaults.queryShared query timing inherited field-by-field by collection rules. Omitted fields continue to profile or built-in fallbacks.no
rule_defaults.query.periodDefault CloudWatch aggregation period from 1m through 1d, as an exact multiple of 1m. An omitted rule period inherits this value before nested metric and profile query defaults.no
rule_defaults.query.lookbackDefault rolling window searched for the newest eligible datapoint. It must be at least the effective period, an exact period multiple, and no more than 1,440 buckets.no
rule_defaults.query.publication_delayDefault collector wait after a bucket closes before it becomes eligible. This is a scheduling policy, not an AWS publication guarantee. Omission falls through to the profile value and then the built-in 10m fallback. Setting this option overrides profile-specific delays for every inheriting rule, including the stock S3 storage profile’s conservative 1d; AWS documents only that S3 storage metrics are reported once per day, so use a shorter default only after verifying each workload’s publication timing.no
rules[].queryOptional query timing overrides for this rule. Each omitted field independently inherits rule_defaults.query, then the relevant profile or built-in fallback.no
rules[].query.periodCloudWatch aggregation period for every series selected by this rule. Rate metrics are normalized using this effective period.no
rules[].query.lookbackRolling window searched for the newest complete finite datapoint. Successful queries may present the retained datapoint as current for up to this duration; longer lookbacks increase response work.no
rules[].query.publication_delayCollector wait after a bucket closes before querying it. This is a scheduling policy, not an AWS publication guarantee. Explicit 0s is allowed for metrics known to publish immediately.no
Resource Filtersrule_defaults.filters.resource_tagsJob-wide list of exact, case-sensitive AWS resource tag predicates inherited by rules that omit rules[].filters.resource_tags. All keys must match; any listed value for one key may match. The Resource Groups Tagging API performs the focused lookup and requires tag:GetResources.no
rule_defaults.filters.resource_tags[].keyExact AWS resource tag key. A filter list supports at most 50 distinct keys.yes
rule_defaults.filters.resource_tags[].valuesOne to 20 exact, case-sensitive accepted values for this key. Values for one key are ORed.yes
rules[].filters.resource_tagsPer-rule replacement for rule_defaults.filters.resource_tags. Omit it to inherit the default, provide a non-empty list to replace the default, or set [] to disable tag filtering for this rule.no
rules[].filters.resource_tags[].keyExact AWS resource tag key. A filter list supports at most 50 distinct keys.yes
rules[].filters.resource_tags[].valuesOne to 20 exact, case-sensitive accepted values for this key. Values for one key are ORed.yes
Resource Labelslabels.resource_tagsOptional AWS resource tags copied to charts as non-identity labels. This is presentation only and does not select resources. Tag values may contain personal data, so expose only keys intended for Netdata. Requires tag:GetResources.no
labels.resource_tags[].keyExact, case-sensitive AWS resource tag key.yes
labels.resource_tags[].labelOptional Netdata label key. When omitted, the AWS key is normalized (Name becomes name). Use an explicit label to avoid invalid names or collisions with identity labels such as region.no
Limitslimits.max_instancesMaximum distinct final static or discovered CloudWatch instances that emit at least one selected series after filtering and exported-series overlap resolution. Metric/statistic fan-out is not counted. Overflow rejects the refreshed plan; collection never truncates to the first N instances.1000no
limits.max_discovery_groupsMaximum unique (target, region, namespace) discovery groups compiled for the job. Compatible rules and profiles share groups. The default is an accidental-expansion safeguard; raise it only for intentional work. Valid range 1–100. Split larger collection across jobs because one refresh can admit at most 100 groups that reach ListMetrics.64no
Discoverydiscovery.refresh_everyHow often (seconds) to re-discover metrics. Minimum 60.300no
discovery.recently_active_onlyUse CloudWatch’s three-hour activity filter only when every selected series sharing a target, region, and namespace scan has publication_delay + lookback + period of 3 hours or less. Any longer-horizon participant keeps the shared scan unfiltered.yesno
Virtual NodevnodeAssociates this data collection job with a Virtual Node.no

via UI

Configure the cloudwatch collector from the Netdata web interface:

  1. Go to Nodes.
  2. Select the node where you want the cloudwatch data-collection job to run and click the :gear: (Configure this node). That node will run the data collection.
  3. The Collectors → Jobs view opens by default.
  4. In the Search box, type cloudwatch (or scroll the list) to locate the cloudwatch collector.
  5. Click the + next to the cloudwatch collector to add a new job.
  6. Fill in the job fields, then click Test to verify the configuration and Submit to save.
    • Test runs the job with the provided settings and shows whether data can be collected.
    • If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.

via File

The configuration file name for this integration is go.d/cloudwatch.conf.

The file format is YAML. Generally, the structure is:

update_every: 1
autodetection_retry: 0
jobs:
  - name: some_name1
  - name: some_name2

You can edit the configuration file using the edit-config script from the Netdata config directory.

cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
sudo ./edit-config go.d/cloudwatch.conf
Examples
Default credentials, single region

Monitor the base AWS identity in us-east-1 using the SDK default credential chain and all default-enabled profiles.

jobs:
  - name: default_credentials
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: base
        credentials: sdk_default
    rules:
      - name: base-defaults
        targets: [base]
        regions: [us-east-1]
Static credentials assume multiple roles

Use one static/session credential source to assume roles for multiple monitored targets. Store credentials in supported secret providers, not in plaintext.

jobs:
  - name: cross_account
    credentials:
      - name: bootstrap
        type: static
        type_static:
          access_key_id: ${env:AWS_ACCESS_KEY_ID}
          secret_access_key: ${env:AWS_SECRET_ACCESS_KEY}
          session_token: ${env:AWS_SESSION_TOKEN}
    targets:
      - name: production
        credentials: bootstrap
        assume_role:
          role_arn: "arn:aws:iam::[ACCOUNT]:role/[ROLE]"
          external_id: ${env:AWS_EXTERNAL_ID}
      - name: staging
        credentials: bootstrap
        assume_role:
          role_arn: "arn:aws:iam::[ACCOUNT]:role/[ROLE]"
    rules:
      - name: both-defaults
        targets: [production, staging]
        regions: [us-east-1, eu-west-1]
Different timing for metrics in one profile

Use disjoint exact metric selections when one service needs different query timing. Earlier rules own only the metric/statistic series they select, so these two Lambda rules do not shadow each other.

jobs:
  - name: lambda_split_policy
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: base
        credentials: sdk_default
    rules:
      - name: lambda-activity
        targets: [base]
        profiles:
          defaults: false
          include: [lambda]
        metrics:
          - profile: lambda
            statistics: [Sum]
            include:
              - name: Invocations
        regions: [us-east-1]
        query:
          period: 5m
          lookback: 30m
          publication_delay: 10m
      - name: lambda-latency
        targets: [base]
        profiles:
          defaults: false
          include: [lambda]
        metrics:
          - profile: lambda
            statistics: [Average, p90]
            include:
              - name: Duration
        regions: [us-east-1]
        query:
          period: 1m
          lookback: 5m
          publication_delay: 5m
AWS Billing estimated charges

Collect each available exact Billing grain independently. Billing metrics must be enabled first, are published only in us-east-1, and do not support resource-tag filters or resource-tag-derived labels. The stock profiles use a 10-minute period and 24-hour retrieval window; a rules[].query block would override only the fields it sets.

jobs:
  - name: billing_estimated_charges
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: billing
        credentials: sdk_default
    rules:
      - name: billing-grains
        targets: [billing]
        profiles:
          defaults: false
          include:
            - billing_total
            - billing_service
            - billing_linked_account
            - billing_linked_account_service
        regions: [us-east-1]
        filters:
          # Billing is not an RGTA resource. This explicitly
          # disables any inherited resource-tag filter.
          resource_tags: []

Collect endpoint-level Average statistics every minute, collect six-hour processed-byte totals independently, and opt into the higher-cardinality endpoint-by-subnet view. Both endpoint grains support the same VPC endpoint resource-tag filters and labels; subnet charts inherit their parent endpoint’s tags.

jobs:
  - name: privatelink_endpoints
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: base
        credentials: sdk_default
    rule_defaults:
      filters:
        resource_tags:
          - key: environment
            values: [production]
    rules:
      - name: endpoint-averages
        targets: [base]
        profiles:
          defaults: false
          include: [privatelink_endpoint]
        metrics:
          - profile: privatelink_endpoint
            statistics: [Average]
            include:
              - name: ActiveConnections
              - name: BytesProcessed
              - name: NewConnections
        regions: [us-east-1]
        query:
          period: 1m
          lookback: 5m
          publication_delay: 5m
      - name: endpoint-six-hour-bytes
        targets: [base]
        profiles:
          defaults: false
          include: [privatelink_endpoint]
        metrics:
          - profile: privatelink_endpoint
            include:
              - name: BytesProcessed
                statistics: [Sum]
        regions: [us-east-1]
        query:
          period: 6h
          lookback: 6h
          publication_delay: 5m
      - name: endpoint-subnets
        targets: [base]
        profiles:
          defaults: false
          include: [privatelink_endpoint_subnet]
        regions: [us-east-1]
    labels:
      resource_tags:
        - key: Name
All services including opt-in profiles

Select defaults and explicitly add every disabled opt-in profile, including the PrivateLink endpoint-by-subnet view and four Billing grains. Their cardinality, prerequisites, and cost guidance still apply.

jobs:
  - name: defaults_and_opt_in
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: base
        credentials: sdk_default
    rules:
      - name: expanded-services
        targets: [base]
        profiles:
          defaults: true
          include:
            - alb_target
            - dynamodb_operation
            - s3_requests
            - ebs_stalled_io
            - privatelink_endpoint_subnet
            - billing_total
            - billing_service
            - billing_linked_account
            - billing_linked_account_service
        regions: [us-east-1]
Filter resources by tag and add tag labels

Apply one job-wide exact tag filter, disable it for an unsupported profile, and expose selected AWS tags as mutable non-identity chart labels.

jobs:
  - name: tagged_resources
    credentials:
      - name: sdk_default
        type: default
    targets:
      - name: base
        credentials: sdk_default
    rule_defaults:
      filters:
        resource_tags:
          - key: managed-by
            values: [platform]
    rules:
      - name: filtered-defaults
        targets: [base]
        regions: [us-east-1]
      - name: unfiltered-cloudfront
        targets: [base]
        profiles:
          defaults: false
          include: [cloudfront]
        regions: [us-east-1]
        filters:
          resource_tags: []
    labels:
      resource_tags:
        - key: Name
        - key: owner
          label: resource_owner
    limits:
      max_instances: 1000
      max_discovery_groups: 64

Metrics

Charts are generated at runtime from the active service profiles. Each static or discovered AWS instance becomes a chart instance identified by its account_id, region, and the profile’s identifying dimensions (for example instance_id for EC2, or bucket_name and storage_type for S3); its contexts live under the cloudwatch. namespace. All CloudWatch metrics use the job’s configured vnode when present, otherwise the node running the collector. Individual AWS resources are distinguished by labels, not generated as separate Netdata nodes. Because CloudWatch publishes with a delay, allow a few minutes for the first data points.

Key terms:

  • Namespace – AWS’s grouping for a service’s metrics (e.g. AWS/EC2).
  • Dimension – a name/value pair that identifies a resource within a namespace (e.g. InstanceId).
  • Statistic – the CloudWatch aggregation applied per period (e.g. average, sum, maximum).
  • Profile – the Netdata YAML file that maps a namespace’s metrics to charts.
  • Partition – an isolated AWS region group (standard aws, GovCloud aws-us-gov, or China aws-cn); all regions selected for one target must share a partition, and an assumed-role ARN must match it.

The built-in profiles ship the following charts by default. Each service links to its profile – the authoritative definition of its exact metrics, statistics, dimensions, and charts:

ProfileMetric prefixDescription
Amazon EC2cloudwatch.ec2.*CPU utilization, network traffic, disk operations, status-check failures, attached-EBS status-check failures
Amazon RDScloudwatch.rds.*CPU utilization, database connections, freeable memory, free storage space, disk throughput, IOPS, latency, replica lag, PostgreSQL transaction ID usage, EBS credit balance
Classic Load Balancer (ELB)cloudwatch.elb.*request count, backend and load-balancer response codes, backend connection errors, latency, host count, spillover count
Application Load Balancer (ALB)cloudwatch.alb.*request count, target and load-balancer response codes, connection rate, active connections, processed traffic, target response time, consumed LCUs
ALB Target Healthcloudwatch.alb_target_health.*per-target-group unhealthy host count
Network Load Balancer (NLB)cloudwatch.nlb.*active and new flow counts, processed bytes and packets, consumed LCUs, TCP resets
NLB Target Healthcloudwatch.nlb_target_health.*per-target-group unhealthy host count
Amazon S3cloudwatch.s3.*bucket size, number of objects (daily storage metrics)
AWS Lambdacloudwatch.lambda.*invocations, errors and throttles, duration
Amazon SQScloudwatch.sqs.*message throughput, empty receives, queue depth, age of oldest message, sent message size
Amazon DynamoDBcloudwatch.dynamodb.*consumed and provisioned capacity, throttle events
Amazon API Gatewaycloudwatch.api_gateway.*requests, errors, latency
AWS Step Functionscloudwatch.step_functions.*executions, throttled events, execution time
NAT Gatewaycloudwatch.nat_gateway.*traffic, active connections, connection rate, errors, idle timeouts
AWS PrivateLink endpointscloudwatch.privatelink_endpoint.*endpoint-level active and new connections, processed bytes, dropped packets, and received reset packets
Amazon Kinesis Data Streamscloudwatch.kinesis.*data throughput, records, GetRecords iterator age, operation latency, throughput exceeded, PutRecords rejected
Amazon Data Firehosecloudwatch.firehose.*records, throughput, put requests, throttled records, S3 delivery freshness and success
Amazon SNScloudwatch.sns.*messages published, notifications, invalid notification filters, DLQ redrive, published message size
Amazon EBScloudwatch.ebs.*volume throughput, IOPS, queue length, idle time, burst balance
Amazon EFScloudwatch.efs.*I/O throughput, metered vs permitted throughput, percent I/O limit, burst credit balance, client connections
Amazon ECScloudwatch.ecs.*service utilization, EBS filesystem utilization, live task count
Amazon ElastiCachecloudwatch.elasticache.*CPU utilization, memory, database memory usage, current and new connections, cache hits and misses, evictions, network traffic
Amazon OpenSearch Servicecloudwatch.opensearch.*cluster status, index writes blocked, automated snapshot failures, nodes, CPU utilization, JVM memory pressure, old-gen JVM memory pressure, free storage space, search and indexing rate, search and indexing latency
Amazon DocumentDBcloudwatch.docdb.*CPU utilization, freeable memory, connections, buffer cache hit ratio, disk IOPS, latency, throughput, replica lag, cursors timed out
Amazon Redshiftcloudwatch.redshift.*health, CPU utilization, disk space used, database connections, disk IOPS, throughput, network throughput
Amazon MSKcloudwatch.msk.*broker throughput, messages in, CPU, disk used, memory, heap memory after GC, partitions, under-min-ISR partitions, connections
Amazon MSK Clustercloudwatch.msk_cluster.*active controllers and offline partitions
Amazon CloudFrontcloudwatch.cloudfront.*requests, downloaded and uploaded traffic, total/4xx/5xx error rates
AWS Auto Scalingcloudwatch.auto_scaling.*group sizing (min/max/desired/total) and instances by state (in-service, pending, standby, terminating)
Amazon Bedrockcloudwatch.bedrock.*invocations, invocation errors, token throughput, invocation and time-to-first-token latency
Amazon EventBridgecloudwatch.eventbridge.*target invocations, rule activity (matched events, triggered rules), ingestion-to-invocation latency
AWS Site-to-Site VPNcloudwatch.vpn.*tunnel traffic (in/out) and tunnel state (fraction of tunnels up)
Amazon EKScloudwatch.eks.*control-plane health: API server request rate, errors, p99 latency, and in-flight requests; etcd database size; scheduler pending pods and scheduling attempts

Each profile also carries optional metrics that are commented out to keep cost and cardinality low; uncomment a metric and its matching chart, then restart the Netdata Agent (profiles are loaded once per go.d process and cached). Stock profiles are shipped at /usr/lib/netdata/conf.d/go.d/cloudwatch.profiles/default/. To customize a service, copy its profile into /etc/netdata/go.d/cloudwatch.profiles/ (keep the same filename) and edit it – a user profile fully replaces the stock one of the same name – then restart the Agent.

These disabled opt-in profiles are collected when a rule names them in profiles.include:

ProfileMetric prefixDescription
ALB Target Groupscloudwatch.alb_target.*per-target-group host count, requests per target, response time, response codes, connection errors
DynamoDB Operationscloudwatch.dynamodb_operation.*per-operation successful request latency, system errors, throttled requests, returned items
EBS Stalled I/Ocloudwatch.ebs_stalled_io.*per-volume stalled I/O health check
S3 Request Metricscloudwatch.s3_requests.*requests, request errors, request latency, request data transfer
AWS PrivateLink endpoints by subnetcloudwatch.privatelink_endpoint_subnet.*the endpoint metrics split by subnet_id; one endpoint can produce several chart instances
AWS Billing totalcloudwatch.billing_total.*latest worldwide estimated month-to-date charge; identity labels: account_id, region
AWS Billing by servicecloudwatch.billing_service.*estimated charges by service_name
AWS Billing by linked accountcloudwatch.billing_linked_account.*estimated charges by linked_account_id when the payer/management account publishes this grain
AWS Billing by linked account and servicecloudwatch.billing_linked_account_service.*estimated charges by linked_account_id and service_name when available

The Billing profiles are exact grains rather than one wildcard: select only the views you need. All use EstimatedCharges, Maximum, USD, a 10-minute period, and a 24-hour retrieval window. AWS publishes the underlying estimate several times daily, so the collector normally re-emits the newest successful value between publications. Around the UTC month boundary, a prior-month value can remain visible until a successful query replaces or expires it, and transient AWS failures can retain it longer. region=us-east-1 identifies the CloudWatch publication/query location; the charge itself is worldwide. Billing dimensions are not AWS Resource Groups Tagging API resources, so resource-tag filters and resource-tag-derived labels do not apply. A valid job can produce no Billing chart when AWS has not published the selected grain.

PrivateLink endpoint metrics have two exact grains. The default privatelink_endpoint profile identifies one endpoint with endpoint_type, service_name, vpc_endpoint_id, and vpc_id. The opt-in privatelink_endpoint_subnet profile adds subnet_id; enable it deliberately because every endpoint can produce several subnet chart instances and seven additional metric/statistic queries per instance. Both profiles use a five-minute period, lookback, and publication delay by default. They share one AWS/PrivateLinkEndpoints discovery scan and one VPC endpoint Resource Groups Tagging API association, so endpoint resource-tag filters and labels also apply to every subnet child. The stock surface exports Average and per-second Sum views where both are useful; exact metric/statistic rules can assign different timing without shadowing siblings. These profiles do not query the separate AWS/PrivateLinkServices namespace.

At stock timing, each endpoint or opted-in subnet instance is queried once per five-minute window. Its seven selected series form five structural CloudWatch metrics for AWS billing grouping, or 1,440 metric requests per day before retries. A one-minute override runs its selected metrics five times as often; narrow profiles, metrics, regions, and resource tags when that freshness is not required.

These opt-in profiles include potentially high-cardinality data. S3 Request Metrics additionally require per-bucket request-metrics configuration in AWS and are billed at CloudWatch custom-metric rates; they collect nothing until enabled on the bucket. PrivateLink subnet cardinality grows with an endpoint’s attached subnets. The Billing service/account grains grow with the payer’s services and linked accounts; their cost guidance is described above.

Alerts

The following alerts are available:

Alert nameOn metricDescription
aws_cloudwatch_ec2_status_check_failedcloudwatch.ec2.status_check_failedEC2 status check failed on ${label:instance_id}
aws_cloudwatch_ec2_attached_ebs_status_check_failedcloudwatch.ec2.status_check_failedEC2 attached EBS status check failed on ${label:instance_id}
aws_cloudwatch_alb_target_group_unhealthy_hostscloudwatch.alb_target_health.unhealthy_hostsALB target group has unhealthy targets on ${label:load_balancer}/${label:target_group}
aws_cloudwatch_nlb_target_group_unhealthy_hostscloudwatch.nlb_target_health.unhealthy_hostsNLB target group has unhealthy targets on ${label:load_balancer}/${label:target_group}
aws_cloudwatch_ebs_stalled_io_check_failedcloudwatch.ebs_stalled_io.stalled_io_checkEBS volume stalled I/O check failed on ${label:volume_id}; requires the opt-in ebs_stalled_io profile
aws_cloudwatch_nat_gateway_port_allocation_errorscloudwatch.nat_gateway.errorsNAT Gateway port allocation errors on ${label:nat_gateway_id}
aws_cloudwatch_efs_io_limit_reachedcloudwatch.efs.io_limitEFS I/O limit reached on ${label:file_system_id}
aws_cloudwatch_efs_burst_credits_exhaustedcloudwatch.efs.burst_creditEFS burst credits exhausted on ${label:file_system_id}
aws_cloudwatch_ecs_cpu_utilizationcloudwatch.ecs.utilizationECS service CPU utilization high on ${label:cluster_name}/${label:service_name}
aws_cloudwatch_ecs_memory_utilizationcloudwatch.ecs.utilizationECS service memory utilization high on ${label:cluster_name}/${label:service_name}
aws_cloudwatch_ecs_ebs_filesystem_utilizationcloudwatch.ecs.ebs_filesystem_utilizationECS EBS filesystem utilization high on ${label:cluster_name}/${label:service_name}
aws_cloudwatch_opensearch_cluster_status_redcloudwatch.opensearch.cluster_statusOpenSearch cluster red on ${label:domain_name}
aws_cloudwatch_opensearch_cluster_status_yellowcloudwatch.opensearch.cluster_statusOpenSearch cluster yellow on ${label:domain_name}
aws_cloudwatch_opensearch_index_writes_blockedcloudwatch.opensearch.index_writes_blockedOpenSearch index writes blocked on ${label:domain_name}
aws_cloudwatch_opensearch_jvm_memory_pressurecloudwatch.opensearch.jvm_memory_pressureOpenSearch JVM memory pressure high on ${label:domain_name}
aws_cloudwatch_opensearch_cpu_utilizationcloudwatch.opensearch.cpuOpenSearch CPU utilization high on ${label:domain_name}
aws_cloudwatch_opensearch_automated_snapshot_failurecloudwatch.opensearch.automated_snapshot_failureOpenSearch automated snapshot failed on ${label:domain_name}
aws_cloudwatch_opensearch_old_gen_jvm_memory_pressurecloudwatch.opensearch.old_gen_jvm_memory_pressureOpenSearch old-gen JVM memory pressure high on ${label:domain_name}
aws_cloudwatch_elasticache_engine_cpu_utilizationcloudwatch.elasticache.cpuElastiCache engine CPU utilization high on ${label:cache_cluster_id}/${label:cache_node_id}
aws_cloudwatch_msk_active_controller_missingcloudwatch.msk_cluster.active_controllersMSK cluster has no active controller on ${label:cluster_name}
aws_cloudwatch_msk_multiple_active_controllerscloudwatch.msk_cluster.active_controllersMSK cluster has multiple active controllers on ${label:cluster_name}
aws_cloudwatch_msk_offline_partitionscloudwatch.msk_cluster.offline_partitionsMSK cluster has offline partitions on ${label:cluster_name}
aws_cloudwatch_msk_cpu_utilizationcloudwatch.msk.cpuMSK broker CPU utilization high on ${label:cluster_name}/${label:broker_id}
aws_cloudwatch_msk_data_logs_disk_usedcloudwatch.msk.disk_usedMSK broker data-log disk utilization high on ${label:cluster_name}/${label:broker_id}
aws_cloudwatch_msk_heap_memory_after_gccloudwatch.msk.heap_memory_after_gcMSK broker heap memory after GC high on ${label:cluster_name}/${label:broker_id}
aws_cloudwatch_msk_under_replicated_partitionscloudwatch.msk.partitionsMSK broker has sustained under-replicated partitions on ${label:cluster_name}/${label:broker_id}
aws_cloudwatch_msk_under_min_isr_partitionscloudwatch.msk.under_min_isrMSK broker has partitions below minimum ISR on ${label:cluster_name}/${label:broker_id}
aws_cloudwatch_rds_replica_lagcloudwatch.rds.replica_lagRDS replica lag high on ${label:db_instance_identifier}
aws_cloudwatch_rds_maximum_used_transaction_idscloudwatch.rds.maximum_used_transaction_idsRDS transaction ID usage high on ${label:db_instance_identifier}
aws_cloudwatch_rds_ebs_byte_balancecloudwatch.rds.ebs_balanceRDS EBS byte balance low on ${label:db_instance_identifier}
aws_cloudwatch_rds_ebs_io_balancecloudwatch.rds.ebs_balanceRDS EBS I/O balance low on ${label:db_instance_identifier}
aws_cloudwatch_vpn_tunnel_downcloudwatch.vpn.tunnel_stateVPN tunnel down on ${label:vpn_id}
aws_cloudwatch_sns_invalid_notification_attributescloudwatch.sns.invalid_notificationsSNS invalid notification attributes on ${label:topic_name}
aws_cloudwatch_sns_invalid_notification_bodycloudwatch.sns.invalid_notificationsSNS invalid notification message body on ${label:topic_name}
aws_cloudwatch_sns_notifications_redriven_to_dlqcloudwatch.sns.dlq_redriveSNS notifications redriven to DLQ on ${label:topic_name}
aws_cloudwatch_sns_notifications_failed_to_redrive_to_dlqcloudwatch.sns.dlq_redriveSNS notifications failed to redrive to DLQ on ${label:topic_name}

Troubleshooting

Debug Mode

Important: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.

To troubleshoot issues with the cloudwatch collector, run the go.d.plugin with the debug option enabled. The output should give you clues as to why the collector isn’t working.

  • Navigate to the plugins.d directory, usually at /usr/libexec/netdata/plugins.d/. If that’s not the case on your system, open netdata.conf and look for the plugins setting under [directories].

    cd /usr/libexec/netdata/plugins.d/
    
  • Switch to the netdata user.

    sudo -u netdata -s
    
  • Run the go.d.plugin to debug the collector:

    ./go.d.plugin -d -m cloudwatch
    

    To debug a specific job:

    ./go.d.plugin -d -m cloudwatch -j jobName
    

Getting Logs

If you’re encountering problems with the cloudwatch collector, follow these steps to retrieve logs and identify potential issues:

  • Run the command specific to your system (systemd, non-systemd, or Docker container).
  • Examine the output for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.

System with systemd

Use the following command to view logs generated since the last Netdata service restart:

journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep cloudwatch

System without systemd

Locate the collector log file, typically at /var/log/netdata/collector.log, and use grep to filter for collector’s name:

grep cloudwatch /var/log/netdata/collector.log

Note: This method shows logs from all restarts. Focus on the latest entries for troubleshooting current issues.

Docker Container

If your Netdata runs in a Docker container named “netdata” (replace if different), use this command:

docker logs netdata 2>&1 | grep cloudwatch

No metrics are collected

Check the following:

  • Permissions – every target allows cloudwatch:GetMetricData; targets selecting any profile with dynamic dimensions also require cloudwatch:ListMetrics. All-constant profiles such as billing_total skip discovery and do not need ListMetrics. GetCallerIdentity needs no explicit grant. Targets with assume_role also require sts:AssumeRole on the source identity. Resource tag filtering or labels require tag:GetResources.
  • Rulesrules[].targets, rules[].profiles, optional rules[].metrics, and rules[].regions select the expected target, service, exact metric/statistic, and region. CloudFront publishes metrics only in us-east-1; its profile enforces this automatically.
  • Resource filters – a rule that omits filters.resource_tags inherits rule_defaults.filters.resource_tags. An explicitly included profile without a safe Resource Groups Tagging API association is rejected; use filters.resource_tags: [] for a deliberate unfiltered rule.
  • Resources are active – confirm in the AWS CloudWatch console that the resources are publishing metrics.
  • Collector logs – check for authentication or API errors:
    # systemd
    journalctl -u netdata --namespace=netdata --grep cloudwatch --since "5 minutes ago"
    # non-systemd
    grep cloudwatch /var/log/netdata/collector.log
    

Missing metrics for some services

  • Profile selection – omit rules[].profiles to select defaults, or ensure the service basename appears under rules[].profiles.include and is not excluded.
  • Metric selection – omit rules[].metrics to collect every metric exported by selected profiles. When configured, verify each profile group names a selected profile, every include[].name is an exact AWS MetricName, and every metric has effective statistics from either its replacement list or the group default.
  • Daily metrics – AWS documents that S3 storage metrics are reported once per day, but does not state a publication-within-one-day guarantee. The stock profile therefore uses a conservative 1d collector delay policy, and recently_active_only is automatically disabled for its long query horizon.
  • Resource activity – some metrics only appear when the resource is actively processing data (for example, EventBridge and Bedrock publish a metric only when its value is non-zero).
  • Auto Scaling group metrics – Auto Scaling group metrics (cloudwatch.auto_scaling.*) are not published until group-metrics collection is enabled on the group (aws autoscaling enable-metrics-collection --granularity 1Minute). Amazon EKS managed node groups have it enabled by default.
  • EKS control-plane metrics – EKS control-plane metrics (cloudwatch.eks.*) are published to the AWS/EKS namespace automatically, at no additional EKS charge, only for clusters running Kubernetes 1.28 or later; older clusters do not report them. These are distinct from Container Insights / the CloudWatch Observability add-on (agent-based, billed separately).

Charts have gaps or incomplete data

CloudWatch publishes metrics with a delay.

  • Keep rules[].query.period at or above the metric’s real publication cadence. A shorter override increases billed query frequency but cannot make AWS publish more often, so it can create empty windows.
  • Set rules[].query.publication_delay when a workload publishes completed buckets later than its profile or the built-in 10m fallback.
  • Check rule_defaults.query.publication_delay before overriding an individual rule. A job-wide value replaces profile-specific delays for inheriting rules, including the stock S3 storage profile’s conservative 1d, and a shorter value can query daily data before it is published.
  • Set rules[].query.lookback to search a wider rolling window for sparse datapoints. It must be an exact multiple of the effective period.
  • A successful query presents its newest eligible datapoint on every Netdata collection cycle, so an old CloudWatch value can appear current for up to lookback. During transient AWS failures, replay can continue longer until a successful query replaces or expires it.
  • Longer lookbacks increase response work and may disable recently_active_only for the shared discovery scan.
  • Transient query failures preserve the retained value and retry after one update_every; later delays double within the same eligible window up to the effective period. A newly eligible window resets the backoff.

Discovery group limit exceeded

A discovery group is one unique (target, region, namespace) combination. Compatible rules and profiles share the same group. limits.max_discovery_groups defaults to 64 to bound accidental ListMetrics expansion.

This can result from accidental target/region/profile expansion or from a legitimate large installation. Verify the derived scope first. For intentional scale, raise the safeguard up to 100. Beyond 100, split the configuration across multiple jobs: one bounded refresh admits at most 100 ListMetrics SDK operations, and every non-skipped group that resolves a client reaches its first admitted operation before continuations. Skipped groups and client-resolution failures consume no operation budget. Splitting preserves metric coverage while keeping each job’s discovery cost, memory, and completion time bounded.

Access denied or authentication errors

  • Verify the credential source referenced by the failing target is valid and not expired.
  • For a target with assume_role, confirm its source identity is allowed to assume the role and that the role trust policy permits it. If the trust policy requires an external ID, use the value supplied by the role owner.
  • For AWS GovCloud or China partitions, ensure each target’s selected rule regions match its role ARN partition.

The observability platform companies need to succeed

Sign up for free

Want a personalised demo of Netdata for your use case?

Contact Sales