The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / oracle-database-ora-00600-internal-error ▌

Operations Guides

ORA-00600: internal error code, arguments - triage and what to capture

ORA-00600 is Oracle’s catch-all internal error code. A server process hit an unexpected condition inside the kernel and failed an internal assertion. The database may stay up, the failing call rolls back, and the system may look normal for minutes or hours, but the engine has reported a condition you cannot fix from SQL.

The message format is ORA-00600: internal error code, arguments: [kdsgrp1], [], [], [], [], [], [], []. The first bracketed token, here [kdsgrp1], is the internal message number and the single most important input for Oracle Support. Multiple distinct bugs can assert in the same internal function, so the argument narrows the search but does not uniquely identify the bug. The full argument list plus the exact database version, down to patch level, is what Support needs to find the matching known bug.

The alert log records every ORA-00600 with a reference to a trace file in the Automatic Diagnostic Repository (ADR). That trace contains the call stack, the SQL text when applicable, and a block dump for corruption-related asserts. It is the only durable evidence that the error occurred. If the instance is restarted before the trace is captured, the on-disk trace remains, but session-level context (open cursors, bind values, PGA state) is gone.

What this means

Oracle reserves the ORA-006xx range for internal errors that should never occur in normal operation. Three codes appear together in the alert log and are routinely confused:

CodeMeaningOriginSeverity
ORA-00600Internal error code, argumentsOracle kernel assertion failed; usually a bug, occasionally corruption or hardwarePAGE
ORA-07445Exception encountered: core dumpOperating system signal (SIGSEGV, SIGBUS, and similar) caught inside Oracle codePAGE
ORA-00700Soft internal error, argumentsInternal “should not happen” condition handled gracefullyVariable

ORA-00600 and ORA-07445 are treated the same operationally: both are critical, both generate a trace, both require Support engagement. The distinction matters for triage because ORA-07445 points more often at OS-level causes (bad memory, OS bug, storage fault returning bad data) while ORA-00600 points more often at a kernel logic bug or a corrupted block being processed. Capture the same artifacts either way.

Do not bucket ORA-00600 / ORA-07445 with ORA-01555 (a tuning signal) or ORA-04031 (urgent but different from corruption). Any ORA-00600 in production is a PAGE, even a single occurrence against an otherwise healthy instance.

Common causes

CauseWhat it looks likeFirst thing to check
Oracle kernel bugRecurring ORA-00600 with same first argument on a specific SQL pattern or featureAlert log and trace; match first argument against My Oracle Support with exact version
Block corruptionORA-00600 with arguments referencing block reads, accompanied by ORA-01578V$DATABASE_BLOCK_CORRUPTION, RMAN VALIDATE, storage SMART data
Hardware fault (RAM, CPU, storage)Random ORA-00600 or ORA-07445 across unrelated SQL; intermittentdmesg for ECC/EDAC errors, storage controller logs
Bug triggered by upgradeFirst ORA-00600 shortly after upgrade or patch applyRecent change history, optimizer feature changes
Resource starvation edge caseORA-00600 under extreme load (memory pressure, very high cursor count)PGA, shared pool, open cursor counts around the time of error
Recovery issue at startupORA-00600 on instance open after a crash, for example [kcratr_nab_less_than_odr]Alert log recovery section, redo log validity

The first argument is the rough category. The remaining arguments and the trace file are the actual evidence.

Quick checks

Read-only. Run these as soon as ORA-00600 is suspected. None modify database state.

# Tail the alert log for the most recent ORA-00600 entries
adrci exec="show alert -tail 200" | grep -A 20 "ORA-00600"
# Alert log location (generic form):
#   $ORACLE_BASE/diag/rdbms/<db_unique_name>/<instance_name>/trace/alert_<instance>.log
ls -lh $ORACLE_BASE/diag/rdbms/*/*/trace/*.trc | tail -20
-- Confirm the instance is still OPEN after the error
SELECT INSTANCE_NAME, STATUS, DATABASE_STATUS, ACTIVE_STATE, SHUTDOWN_PENDING
FROM V$INSTANCE;
-- List recent critical incidents from ADR (12c+)
SELECT INCIDENT_ID, PROBLEM_KEY, ERROR_FACILITY, ERROR_NUMBER,
       ORIGINATING_TIMESTAMP, STATUS
FROM V$DIAG_INCIDENT
WHERE PROBLEM_KEY LIKE 'ORA 600%'
   OR PROBLEM_KEY LIKE 'ORA 7445%'
ORDER BY ORIGINATING_TIMESTAMP DESC
FETCH FIRST 20 ROWS ONLY;
-- Check for concurrent corruption signals
SELECT FILE#, BLOCK#, BLOCKS, CORRUPTION_TYPE
FROM V$DATABASE_BLOCK_CORRUPTION;
-- Capture exact database version, down to patch level
SELECT BANNER_FULL, CON_ID FROM V$VERSION;
# OS-level evidence of hardware fault (run as the oracle OS user, outside SQL)
# -T requires util-linux dmesg; on other systems drop -T for raw timestamps
dmesg -T | grep -iE "ecc|edac|hardware error|i/o error" | tail -30

If V$DIAG_INCIDENT is unavailable, the alert log plus adrci are sufficient. The alert log is the universal correlation anchor for Oracle incidents.

How to diagnose it

  1. Confirm the error and capture the arguments verbatim. Open the alert log, find the ORA-00600 line, and record the entire argument list, not just the first token. The entry typically reads:

    ORA-00600: internal error code, arguments: [kdsgrp1], [], [], [], [], [], [], []
    Errors in file .../trace/<instance>_ora_<pid>.trc
    

    Copy both lines. The trace path is what you package next.

  2. Verify the instance is still usable. A single ORA-00600 usually rolls back the failing statement; the session survives and the instance stays OPEN. If a background process (DBWn, LGWR, PMON) asserts, the instance will abort and restart. Check V$INSTANCE and the alert log for instance recovery messages.

  3. Pull the trace file before anything else changes. ADR preserves trace files across restarts, but if the trace filesystem fills or retention purges old files, you lose the evidence. Copy the trace off the host to a safe location immediately. Do not rely on the file still being there tomorrow.

  4. Distinguish ORA-00600 from ORA-07445 and ORA-00700. If the alert log shows ORA-07445 instead, the cause is an OS-level signal inside Oracle code (often memory or storage). If ORA-00700, it is a soft internal condition. The triage artifacts are the same; the SR routing differs.

  5. Correlate with concurrent signals. Cross-check the ORA-00600 timestamp against:

    • V$DATABASE_BLOCK_CORRUPTION for recently detected corruption
    • dmesg for ECC, EDAC, or I/O errors in a window around the timestamp
    • Recent DDL, statistics gathering, or upgrade activity in the alert log
    • Redo log switches near the time, since some ORA-00600 variants are recovery-related
  6. Package the incident for Oracle Support. ADR creates an incident automatically. Use ADRCI’s Incident Packaging Service (IPS) to bundle the trace, alert log snippet, and related files into a zip Support can ingest. On systems with Autonomous Health Framework (AHF) or Trace File Analyzer (TFA) installed, the collection path is a single command:

    # SRDC collection for ORA-00600 via TFA / AHF
    tfactl diagcollect -srdc ORA-00600
    

    Without AHF, the manual ADRCI IPS path produces the equivalent bundle:

    adrci
    IPS CREATE PACKAGE INCIDENT <incident_id>
    IPS GENERATE PACKAGE <package_id> IN /tmp/ora00600_pkg
    
  7. Open the SR with the right opening data. Include:

    • The exact ORA-00600 line including all arguments
    • Database version (full BANNER_FULL from V$VERSION)
    • Platform and OS version
    • The trace file or the package zip
    • Whether this is the first occurrence or recurring
    • Any recent changes (upgrade, patch, parameter change, statistics gathering)
  8. Do not “fix” by restarting. Restarting does not resolve a kernel bug. It loses session context, clears ASH samples, and may mask a recurrence if the trigger is workload-dependent. If the instance has crashed on its own, that is different: let SMON complete recovery and capture the post-recovery alert log.

flowchart td
  A["Alert log: ORA-00600 line"] --> B["Capture arguments + version verbatim"]
  B --> C["Locate trace file in ADR"]
  C --> D["Copy trace off host"]
  D --> E{"Instance still OPEN?"}
  E -- Yes --> F["Correlate with corruption, dmesg, recent changes"]
  E -- No --> G["Wait for SMON recovery, capture post-recovery log"]
  F --> H["Distinguish 600 vs 7445 vs 700"]
  G --> H
  H --> I["Package via ADRCI IPS or tfactl diagcollect"]
  I --> J["Open SR with arguments, version, trace"]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Alert log ORA-00600 / ORA-07445 countTrend reveals whether a bug is recurring or hardware is degradingAny non-zero value in production; rising rate is escalation
V$DIAG_INCIDENT with PROBLEM_KEY ORA 600%Programmatic access to incident history with timestampsCluster of incidents with same first argument in a short window
V$DATABASE_BLOCK_CORRUPTION rowsCorruption is a common ORA-00600 trigger; the kernel asserts when reading bad blocksAny row
dmesg ECC / EDAC / I/O errorsHardware faults often surface as ORA-00600 / ORA-07445 before they crash the hostAny new hardware error near an ORA-00600 timestamp
Instance status (V$INSTANCE)A background process asserting can abort the instanceSTATUS cycling STARTED -> MOUNTED -> OPEN unexpectedly
V$VERSION BANNER_FULLExact version is mandatory for matching known bugsVersion change should reset the ORA-00600 baseline
Redo log switch frequencySome ORA-00600 variants are recovery-relatedSwitch spikes near ORA-00600 timestamps
Recent DDL / statistics / parameter changesOptimizer and feature changes are common triggers after upgradesChange window correlation

Fixes, or what you actually do

There is no operator-side fix for an ORA-00600. The action set is narrow.

Preserve evidence and engage Support

The first action. Without the trace and arguments, Support cannot identify the bug. Not optional.

Apply the patch Support identifies

Once Support matches the arguments and version to a known bug, they will point you at a patch. Apply it in the next maintenance window. Track the bug ID against the SR.

Short-term workarounds, Support-directed only

Support may suggest a hidden parameter (_some_feature_enabled = false) or a SQL-level workaround (hint, SQL patch, disabling a feature). Treat these as risk-bearing: an underscore parameter that suppresses the assert may silently disable the feature that triggered it. Keep the SR open until a real patch is applied. Do not promote underscore-parameter workarounds into permanent configuration.

Treat the rare non-bug cases

  • Block corruption as root cause: restore and recover the affected blocks via RMAN Block Media Recovery, then investigate storage health. The ORA-00600 is the symptom; corruption is the cause.
  • Hardware fault as root cause: engage the hardware vendor. Replace DIMMs, HBAs, or disks before chasing a software bug that does not exist.
  • Recovery-related ORA-00600 at startup: not kernel bugs in the usual sense. Follow the MOS note for the specific argument. The fix is usually a recovery procedure, not a patch.

What not to do

  • Do not flush the shared pool. It will not fix a kernel bug and may cause a hard parse storm.
  • Do not restart in the hope the error goes away. It masks recurrence.
  • Do not apply underscore parameters found on a forum without a Support-confirmed bug ID.
  • Do not ignore a single occurrence. Any ORA-00600 in production is a PAGE.

Prevention

  • Patch current. Most ORA-00600 bugs are fixed in later release updates. Running an old, unpatched release is the most common reason for “mysterious” recurring internal errors.
  • Enable block checking. DB_BLOCK_CHECKING = FULL or MEDIUM adds CPU overhead but catches corruption before it triggers ORA-00600 downstream.
  • Enable block checksums. DB_BLOCK_CHECKSUM = FULL detects block damage on read.
  • Validate proactively. Schedule RMAN BACKUP VALIDATE CHECK LOGICAL DATABASE periodically to surface corruption before an ORA-00600 does.
  • Monitor hardware. ECC errors, EDAC reports, and storage SMART data are leading indicators. A failing DIMM produces ORA-00600 / ORA-07445 before it produces a clean crash.
  • Track the trend. A rising rate of ORA-00600 / ORA-07445 over weeks suggests either a regression (recent patch introduced a new bug) or hardware degradation. Use V$DIAG_INCIDENT counts as a trend signal, not just a per-event alert.
  • Stage upgrades. Most post-upgrade ORA-00600 storms come from optimizer or feature changes. Test the upgrade against a production-shaped workload and keep a rollback path.

How Netdata helps

Netdata’s value for ORA-00600 is correlation, not detection of the error itself. The error lives in the alert log; monitoring assembles the surrounding context that explains why it happened.

  • Per-second metric context around the timestamp. CPU saturation, PGA spikes, redo latency, and I/O errors in the minutes before an ORA-00600 narrow the cause from “kernel bug” to “kernel bug triggered under condition X”.
  • OS-level signals alongside database signals. dmesg ECC errors, disk latency outliers, and OOM killer activity correlate with ORA-07445 / ORA-00600 in ways a database-only view cannot.
  • Anomaly detection on incident counts. A baseline of zero ORA-00600 makes any occurrence anomalous. A baseline of occasional makes a cluster a spike.
  • Alert log scraping. Picking up ORA-00600 / ORA-07445 / ORA-01578 / “cannot allocate new log” lines from the alert log is the fastest path from “user complained” to “we already know.”
  • Correlation across instances. In RAC, an ORA-00600 on one node paired with interconnect errors on another tells a different story than an isolated assert.

Netdata’s Oracle Database monitoring with Netdata brings these signals together with per-second metrics and anomaly detection.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.