The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / oracle-database-checkpoint-not-complete ▌

Operations Guides

Oracle 'Checkpoint not complete': redo log sizing, DBWn, and log-switch stalls

The alert log says Checkpoint not complete followed by Current log# N seq# N mem# N: <path>. Foreground sessions stall on log file switch (checkpoint incomplete) at every log switch. Commits hesitate for tens of milliseconds to seconds, and the stall repeats each time LGWR wraps to the next redo log group. The wrong fix (enlarging redo logs when the real bottleneck is DBWn I/O) only buys minutes.

This guide assumes you understand the Oracle write path from how Oracle Database works in production.

What this means

LGWR cycles through online redo log groups. Before it can reuse a log group, two conditions must hold for that group: every redo byte in it must have been flushed by LGWR (trivially true once the log fills), and the dirty buffers protected by that redo must have been written to datafiles by DBWn. The second condition is the one that fails. The thread checkpoint SCN cannot advance past group N until DBWn has flushed everything covered by group N’s redo.

When DBWn has not caught up, CKPT cannot mark the checkpoint complete, and LGWR cannot wrap. Every session that issues COMMIT or any DML waits on log file switch (checkpoint incomplete). The wait is system-wide, not per-SQL. TPS drops uniformly because every write transaction is blocked at the same point.

Alert-log diagnostics vary by release, so check the release-specific strings for your patch level. The traditional Checkpoint not complete text is the canonical signal to grep.

This is not the same failure as the archiver hang. log file switch (archiving needed) and Thread N cannot allocate new log, sequence N in the alert log mean ARCn cannot copy redo to the archive destination. That requires freeing archive space, fixing the FRA, or repairing ARCn. This guide covers the DBWn path. The two have similar symptoms but different fixes.

flowchart TD
    A[LGWR fills current redo log] --> B{Checkpoint
advanced past it?} B -->|Yes, wrap| A B -->|No| C[All committing sessions wait
log file switch checkpoint incomplete] C --> D[LGWR sleeps and polls controlfile
control file sequential read] D --> E[CKPT updates datafile headers
holds controlfile enqueue] E --> F[DBWn flushes dirty buffers
to datafiles] F -->|Slow I/O or too few writers| G[Stall persists
commits hesitate] F -->|Catches up| B

LGWR does not synchronously block on DBWn. It goes into an idle sleep and polls the controlfile (producing control file sequential read waits) while waiting for checkpoint progress. CKPT, holding the controlfile enqueue while updating datafile headers during the switch, is what stalls LGWR. The wait chain to look for, if you have Diagnostics Pack, is foreground to LGWR to CKPT (controlfile enqueue) to DBWn (slow I/O).

Common causes

CauseWhat it looks likeFirst thing to check
Redo logs too smallSwitches more than 1 per minute; V$LOG_HISTORY shows high switch rateOPTIMAL_LOGFILE_SIZE in V$INSTANCE_RECOVERY
Too few redo log groupsTwo or three groups, switching every few minutes, recurring stallsCount groups in V$LOG
DBWn I/O bottleneckdb file parallel write and free buffer waits elevated; redo logs already largeV$FILESTAT.AVGIOTIM on datafile LUNs
FAST_START_MTTR_TARGET too aggressiveAggressive checkpoints, high write I/O even with large logsV$INSTANCE_RECOVERY.ESTIMATED_MTTR vs TARGET_MTTR
archive_lag_target forcing switchesSwitches at a fixed interval regardless of redo rate; common on SESHOW PARAMETER archive_lag_target
SE2 with no MTTR knobSetting FAST_START_MTTR_TARGET raises ORA-00439SELECT BANNER FROM V$VERSION

Quick checks

-- Confirm the wait event and recent stall time
SELECT EVENT, TOTAL_WAITS, TIME_WAITED_MICRO,
       ROUND(TIME_WAITED_MICRO / NULLIF(TOTAL_WAITS,0) / 1000, 2) AS avg_ms
FROM V$SYSTEM_EVENT
WHERE EVENT LIKE 'log file switch%';
-- Log switch frequency by hour
SELECT TO_CHAR(FIRST_TIME, 'YYYY-MM-DD HH24') AS hr, COUNT(*) AS switches
FROM V$LOG_HISTORY
WHERE FIRST_TIME > SYSDATE - 1
GROUP BY TO_CHAR(FIRST_TIME, 'YYYY-MM-DD HH24')
ORDER BY 1;
-- Redo log group count and size
SELECT GROUP#, BYTES/1048576 AS mb, MEMBERS, STATUS, ARCHIVED
FROM V$LOG ORDER BY GROUP#;
-- Oracle's recommended redo size; OPTIMAL_LOGFILE_SIZE populates only when FAST_START_MTTR_TARGET is non-zero
SELECT ESTIMATED_MTTR, TARGET_MTTR, OPTIMAL_LOGFILE_SIZE
FROM V$INSTANCE_RECOVERY;
-- DBWn write latency (the DBWn process I/O wait, not the session view)
SELECT EVENT, TOTAL_WAITS, TIME_WAITED_MICRO,
       ROUND(TIME_WAITED_MICRO / NULLIF(TOTAL_WAITS,0) / 1000, 2) AS avg_ms
FROM V$SYSTEM_EVENT
WHERE EVENT IN ('db file parallel write', 'free buffer waits');
-- Per-datafile average I/O time in ms (expect <10ms SSD, <20ms SAN)
SELECT f.FILE#, d.NAME, f.PHYRDS, f.PHYWRTS, f.AVGIOTIM
FROM V$FILESTAT f JOIN V$DATAFILE d ON f.FILE# = d.FILE#
ORDER BY f.AVGIOTIM DESC;
-- Edition and MTTR state (FAST_START_MTTR_TARGET is EE-only)
SELECT BANNER FROM V$VERSION WHERE BANNER LIKE 'Oracle%';
SHOW PARAMETER fast_start_mttr_target
SHOW PARAMETER archive_lag_target
# Grep the alert log for the canonical strings; add release-specific strings as needed
adrci exec="show alert -tail 1000" | grep -iE "checkpoint not complete|cannot allocate new log"

How to diagnose it

  1. Confirm the symptom class. If the alert log only shows Checkpoint not complete (not cannot allocate new log), you are on the DBWn/checkpoint path. If both appear, fix archive space first. The archive hang is the more dangerous failure: existing sessions freeze silently and new non-SYSDBA connections get ORA-00257.

  2. Compute the log switch rate from V$LOG_HISTORY. Compare to current redo log size. A widely used operational target is to switch no more often than every 15 to 30 minutes at peak; Oracle requires each online redo log file to be at least 4 MB, but the right floor is the measured redo-generation rate, not a fixed 1 GB. Anything switching more than once per minute at peak is undersized for the current write rate.

  3. Separate redo-rate problems from DBWn-throughput problems. Sample redo size from V$SYSSTAT over 60 seconds to get MB/sec. Then check db file parallel write average latency. If DBWn write latency is high (>10ms on SSD, >20ms on SAN), the bottleneck is datafile I/O, not redo log size. Adding redo log groups buys time but does not fix the underlying write-path issue.

  4. Check OPTIMAL_LOGFILE_SIZE in V$INSTANCE_RECOVERY. It is only populated when FAST_START_MTTR_TARGET is non-zero. If it is more than 2x your current redo log file size, redo log sizing is your primary lever.

  5. If you are on Standard Edition, you cannot set FAST_START_MTTR_TARGET. Attempting it raises ORA-00439: feature not enabled: Fast-Start Fault Recovery. Check archive_lag_target instead. When archive_lag_target forces a switch at a fixed interval shorter than the natural fill rate, checkpointing becomes passive and the checkpoint position does not advance fast enough. Adding more redo groups does not help in this scenario; raise or remove archive_lag_target.

  6. Walk the wait chain if you have a Diagnostics Pack license. Look at V$ACTIVE_SESSION_HISTORY or an ASH wait-chain report. The typical chain during a stall is foreground to LGWR to CKPT (controlfile enqueue) to DBWn (slow I/O). The presence of control file sequential read for LGWR is the polling signature.

  7. Confirm the fix path before changing anything. If db file parallel write is the bottleneck, the highest-leverage fix is storage and DBWn configuration, not redo size. If redo logs are clearly undersized and DBWn I/O is healthy, resize the redo logs. Often both are needed: bigger logs as immediate relief, plus a DBWn I/O investigation for the root cause.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
log file switch (checkpoint incomplete) in V$SYSTEM_EVENTDirect wait-event confirmation of the stallAny non-trivial TOTAL_WAITS in a healthy window
Log switches per hour from V$LOG_HISTORYBest early-warning before stalls happen>6 per hour sustained, >1 per minute critical
OPTIMAL_LOGFILE_SIZE in V$INSTANCE_RECOVERYOracle’s own sizing recommendation>2x current redo log file size
db file parallel write average latencyDBWn’s actual I/O time on datafiles>10ms SSD, >20ms SAN
free buffer waitsDBWn cannot keep the cache cleanAny sustained non-zero count
ESTIMATED_MTTR vs TARGET_MTTRWhether checkpoint can hit its MTTR goalESTIMATED consistently above TARGET
Alert log string Checkpoint not completeThe symptom itselfMore than a few per day on a busy system

Fixes

Add or enlarge redo log groups

The most common immediate relief. Add groups first (less disruptive than resizing existing members), then add larger members and drop the old ones once the new ones have cycled through. Oracle recommends at least three groups per thread with multiplexed members. Target a switch every 15 to 30 minutes at peak redo rate; Oracle requires online redo logs to be at least 4 MB, while the production size must follow measured redo generation.

-- Add a new group with multiplexed members
ALTER DATABASE ADD LOGFILE GROUP 4 ('/redo/redo04a.log', '/redo/redo04b.log') SIZE 1G;

-- WARNING: Only drop groups whose STATUS is INACTIVE and ARCHIVED is YES.
-- Dropping the CURRENT or ACTIVE group stalls the database.
-- Verify first: SELECT GROUP#, STATUS, ARCHIVED FROM V$LOG;
-- Oracle does not delete physical files at the OS level after the drop.
ALTER DATABASE DROP LOGFILE GROUP 1;

Tradeoff: bigger redo logs lengthen crash recovery because more redo must be applied on instance startup. This is exactly the tradeoff FAST_START_MTTR_TARGET exists to manage. Increasing redo log size without bound to mask a DBWn problem will eventually push recovery time past your RTO.

Tune FAST_START_MTTR_TARGET (Enterprise Edition only)

Set a non-zero target so Oracle self-tunes checkpoint aggressiveness. After setting it, V$INSTANCE_RECOVERY.OPTIMAL_LOGFILE_SIZE populates and you can size redo logs to match. Setting too aggressive a target increases write I/O during normal operation. Setting it too lenient lengthens recovery time.

-- Takes effect immediately; checkpoint behavior changes on the next log switch
ALTER SYSTEM SET fast_start_mttr_target = 300 SCOPE=BOTH;
-- Then re-check V$INSTANCE_RECOVERY for the optimal logfile size hint

On Standard Edition, attempting this raises ORA-00439. There is no MTTR knob on SE. Use redo log sizing and archive_lag_target discipline instead.

In 19c, LOG_CHECKPOINT_INTERVAL and LOG_CHECKPOINT_TIMEOUT are not deprecated, but Oracle recommends disabling/removing them when FAST_START_MTTR_TARGET is set. A nonzero LOG_CHECKPOINT_INTERVAL overrides FAST_START_MTTR_TARGET. FAST_START_IO_TARGET is obsolete and should not be used.

Fix DBWn write throughput

If db file parallel write is the actual bottleneck, redo log sizing is a bandage. The root cause is datafile write I/O. Checks:

  • Confirm asynchronous I/O is enabled: SHOW PARAMETER disk_asynch_io (should be TRUE) and SHOW PARAMETER filesystemio_options (should be SETALL or ASYNCH on filesystem-backed datafiles).
  • Review DB_WRITER_PROCESSES. One DBWn may not be enough on busy systems with many datafiles spread across LUNs.
  • Look at OS-level latency on the datafile LUN with iostat -x 1 5. High await on the datafile device while redo is on a separate, healthy device confirms the diagnosis.
  • Consider separating redo logs from datafiles if they currently share a LUN. Redo and datafile I/O contention is a classic misconfiguration that produces exactly this symptom.

When async I/O is misconfigured, enabling it alone (disk_asynch_io=true, filesystemio_options=setall) can eliminate the stalls without redo log changes. Check these settings first on any system where DBWn latency is unexpectedly high.

Review archive_lag_target (especially on Standard Edition)

archive_lag_target forces a log switch at a fixed interval. When it forces switches more often than redo naturally fills the logs, Oracle’s checkpointing becomes passive. The checkpoint position does not advance fast enough, and LGWR stalls when it cycles back. This is a documented failure mode on SE where MTTR tuning is unavailable. Adding redo groups does not help; raise or remove archive_lag_target and re-baseline.

Prevention

  • Track log switches per hour as a baseline metric. Set a planning alert at >6 per hour sustained and a ticket at >1 per minute. Sustained high switch frequency is the leading indicator for this failure.
  • On EE, set FAST_START_MTTR_TARGET to a non-zero value. Monitor ESTIMATED_MTTR against TARGET_MTTR. A rising gap is an early warning that DBWn cannot keep up with the current checkpoint goal.
  • Re-evaluate redo log sizing whenever redo generation rate grows. A 50% increase in redo size per second without a redo log resize will roughly halve the time to switch.
  • Monitor db file parallel write average latency as a standing DBWn health signal, not just during incidents. It should be steady week-over-week on the same storage.
  • After any storage change (LUN migration, ASM rebalance, SAN firmware update), recheck V$FILESTAT.AVGIOTIM on datafiles. Slow datafile I/O after a storage change is a common root cause for stalls that appear weeks later.
  • On SE, audit archive_lag_target against the natural redo fill rate. They need to be in the same order of magnitude. A 2-minute archive_lag_target with 30-minute natural fill is a stall waiting to happen.

How Netdata helps

Per-second collection matters here because checkpoint stalls are bursty. A 20-second stall recurring every 5 minutes can fall between 30-minute AWR snapshots.

  • Redo log switch rate and log file switch (checkpoint incomplete) wait counts are collected per second, surfacing stalls within seconds rather than at the next snapshot.
  • db file parallel write latency correlated with checkpoint-incomplete events separates a redo-sizing problem from a DBWn I/O problem without stitching together separate tools.
  • Alert log strings (Checkpoint not complete, Checkpoint lag detected, cannot allocate new log) produce distinct alerts, so the DBWn path and archiver path page different responders with different runbooks.
  • Anomaly detection on log switch frequency catches gradual drift toward undersized redo logs before the first stall.
  • ESTIMATED_MTTR vs TARGET_MTTR tracked over time shows whether DBWn can hit its checkpoint goal before the alert log complains.

For the full setup, see Oracle Database monitoring with Netdata.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.