The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / how-oracle-database-works-in-production ▌

Operations Guides

How Oracle Database actually works in production: a mental model for operators

Oracle is dense, but four abstractions make most production incidents legible: the shared memory region (SGA), per-process private memory (PGA), the background processes that move data to disk, and the wait-event model that tells you where time is going. When you see log file sync climbing, or free buffer waits appearing, or the database hanging while basic health checks still pass, you should know immediately which subsystem is involved and what it competes for.

What it is and why it matters

Oracle is a multi-process architecture on Unix/Linux (multi-threaded on Windows) built around a single shared memory region called the System Global Area (SGA) and per-session private memory called the Program Global Area (PGA). In dedicated server mode, every connected session gets its own server process and PGA. Background processes run alongside, moving data between the SGA and disk.

Every Oracle incident maps back to one of these layers. A “slow database” complaint is almost always one of three root causes: LGWR cannot flush redo fast enough, DBWn cannot write dirty buffers fast enough, or a session is holding a lock others need. Knowing which layer owns the symptom is the difference between a five-minute fix and a two-hour war room.

How it works

The SGA: shared memory

The SGA is the shared workspace all sessions see. It is the largest memory consumer and the source of most tuning complexity.

ComponentPurposeWhat breaks if starved
Buffer CacheCaches data blocks from datafiles. Every read and write checks here first.Physical I/O floods storage; db file sequential read / db file scattered read waits spike
Shared PoolLibrary cache (parsed SQL plans), data dictionary cache, result cache. Hard parses require exclusive latches.library cache: mutex X waits, ORA-04031 shared pool OOM, CPU spike from parsing
Redo Log BufferCircular buffer for change vectors before LGWR flushes to online redo logs. Small (typically 16-256MB), but on the critical commit path.Redo log buffer space waits (rare; LGWR usually keeps up)
Large PoolRMAN I/O buffers, shared server session memory, parallel query message buffers.RMAN backup failures, parallel query degradation
Java Pool / Streams PoolJava stored procedures, GoldenGate integrated capture.Only relevant if these features are active

The buffer cache is the single most important cache. The shared pool is the single most fragile, because hard parsing requires exclusive latches and fragmentation leads to ORA-04031.

PGA: per-process memory

Each dedicated server process gets its own PGA for sorting, hashing, bitmap operations, and session state. PGA is managed via PGA_AGGREGATE_TARGET (a soft target) and PGA_AGGREGATE_LIMIT (a hard limit, 12c+). When PGA is insufficient for a sort or hash join, the operation spills to the temp tablespace, visible as direct path read temp and direct path write temp waits. On analytics workloads, PGA aggregate can rival the SGA in size.

Critical background processes

ProcessWhat it doesWhat breaks if it stalls
DBWnWrites dirty buffers from buffer cache to datafilesBuffer cache fills, free buffer waits appear, all DML stalls
LGWRWrites redo log buffer to online redo logsEvery COMMIT hangs (log file sync), entire application stalls
CKPTSignals DBWn to write, updates datafile headersRecovery time grows; checkpoint not complete warnings
SMONInstance recovery, coalesces free extents, cleans tempDead transactions not rolled back; temp not reclaimed
PMONCleans up failed user processes, releases their locksOrphaned locks persist indefinitely
ARCnCopies filled online redo logs to archive destinationRedo logs cannot be reused, database hangs
MMON/MMNLAWR snapshots, ASH samplingDiagnostics blind; not operationally critical
RECOResolves in-doubt distributed transactionsDistributed transaction locks not released
CJQ0/JnnnJob schedulerScheduled jobs stop executing
LCKn/LMSn/LMDRAC only: Global Cache/Enqueue ServiceRAC cluster communication fails

The connection lifecycle

  1. Client connects to the Listener (typically port 1521).
  2. The Listener spawns a dedicated server process (or routes to a shared server dispatcher).
  3. Each dedicated server gets its own PGA and one OS process.
  4. PROCESSES and SESSIONS parameters hard-limit concurrency. Hitting them causes ORA-00020.

Each dedicated server process consumes one OS process and one PGA allocation. OS limits (ulimit, pid_max) can be hit before Oracle’s own limits. The listener is a single point of failure for new connections: if it is down, existing sessions are unaffected but no new connections can be established. Background processes consume a variable number of process slots — the count depends on enabled features such as RAC or Data Guard, and is visible in V$BGPROCESS — reducing what is available for users.

The redo and undo write path

This is the mechanism that connects most background processes into a single chain. Understanding it explains why one slow link freezes everything.

Every data change generates two things:

  • A redo vector (for recovery) written to the redo log buffer in the SGA.
  • An undo record (for rollback and read consistency) written to undo segment blocks in the buffer cache, persisted to the undo tablespace by DBWn.

COMMIT triggers LGWR to flush the redo log buffer to the online redo logs. LGWR must finish writing before the commit returns to the client. This means write throughput is gated by LGWR performance, and redo log I/O is the most latency-sensitive path in the system.

flowchart LR
  S[Session DML] -->|dirty block| BC[Buffer Cache]
  S -->|redo vector| RB[Redo Log Buffer]
  S -->|undo record| BC
  RB -->|LGWR on commit| OL[Online Redo Logs]
  OL -->|ARCn after switch| AD[Archive Dest]
  BC -->|DBWn checkpoint| DF[Datafiles]
  CKPT[CKPT] -.->|signals| DBWn[DBWn]

Two consequences follow. First, if LGWR is slow, every committing session waits uniformly. Second, if ARCn cannot archive filled online redo logs, LGWR cannot switch to a new log group, and every session needing redo space freezes.

Undo has its own failure modes. Long-running queries need undo blocks to remain available for read consistency. If undo is overwritten before the query finishes, you get ORA-01555 (snapshot too old). If the undo tablespace fills with active extents, you get ORA-30036 and writes fail.

The wait event model

Oracle’s primary diagnostic framework is wait events. Every session is either ON CPU or waiting for something. The views V$SESSION_EVENT (per-session accumulated), V$SYSTEM_EVENT (system-wide since startup), and V$ACTIVE_SESSION_HISTORY (sampled, requires Diagnostics Pack) expose this.

Wait times tell you where time is being spent. Hit ratios do not. Oracle’s performance methodology since 10g explicitly de-emphasizes cache hit ratios in favor of wait event analysis. When you triage, the fastest first step is to look at the dominant wait class among active sessions:

Wait classWhat it means
CommitLGWR bottleneck. Look at log file sync and log file parallel write.
User I/OStorage. Look at db file sequential read and db file scattered read.
ConcurrencyLocks or latches. Look at enq: TX and library cache: mutex X.
ApplicationEnqueue waits from application logic. Look for uncommitted transactions.
IdleNot a performance problem. SQL*Net message from client means the database is waiting for the client.

Always filter idle waits (WAIT_CLASS != 'Idle') in performance analysis. SQL*Net message from client dominates V$SYSTEM_EVENT by count on most systems and is just the database waiting for the client to send the next request.

Where it shows up in production

ArchetypeWhat it looks likeWhich layer
LGWR cannot keep upEvery committing session waits on log file sync. The most common Oracle performance emergency.Redo path
Archive destination fullARCn cannot write, online redo logs fill, database hangs silently. Existing sessions freeze, new non-SYSDBA connections get ORA-00257.Archive path
Space exhaustionTablespace full (ORA-01653/01654), temp full (ORA-01652), undo full (ORA-30036). Cliff-edge, no graceful degradation.Storage
Lock contention cascadeOne uncommitted transaction holds row locks, sessions queue, process/session limits approached.Concurrency
Parse stormLiteral SQL causes hard parses, library cache mutex contention, shared pool fragmentation, eventually ORA-04031.Shared pool
Plan regressionOptimizer picks a bad plan after stats collection. A query goes from 10ms to 10 minutes.Optimizer
Connection exhaustionPROCESSES/SESSIONS limit hit. ORA-00020 refuses new connections.Process slots
Memory pressure / OOMSGA + PGA exceed physical RAM. Linux OOM killer targets Oracle processes. Random session deaths.Memory
RAC inter-node thrashingHot blocks bounced between instances. gc buffer busy waits dominate.RAC interconnect

The most dangerous pattern is the archive destination full scenario. It masquerades as “database up.” The instance is OPEN, the listener responds, TCP checks pass. But every session needing redo is frozen on log file switch (archiving needed), and applications with existing connections get no error. They just hang.

Common misuses of the mental model

Monitoring hit ratios instead of wait events. Buffer cache hit ratio is the most over-relied-upon Oracle metric. A 99% ratio means nothing if log file sync is 50ms. A system doing nothing but SELECT * FROM dual in a loop has a 99.99% hit ratio. A system doing massive parallel analytics might have 60% and be performing perfectly. Focus on where time is spent.

Treating all ORA- errors equally. ORA-00600 (internal error) and ORA-07445 (segfault in Oracle code) are always critical and require Oracle Support engagement. ORA-01555 (snapshot too old) is an undo tuning signal. ORA-04031 (shared pool OOM) is urgent but fundamentally different from corruption. Triage by error code.

Using instance status as the only availability signal. An instance that is OPEN but hung passes basic health checks. A SELECT 1 FROM DUAL from an existing connection pool may succeed even during a total LGWR stall because it generates no redo. A meaningful health check must execute a DML statement with a COMMIT to exercise the full write path.

Shared pool flushes as a fix. ALTER SYSTEM FLUSH SHARED_POOL causes a hard parse storm and is almost never the right action. It creates more problems than it solves. The exception is a one-time intervention for documented shared pool corruption, not a recurring maintenance task.

Static thresholds on baseline-dependent metrics. TPS, logical reads, and redo generation rate are all workload-dependent. Static thresholds generate noise. Use baseline deviation, ideally a 7-day rolling baseline by hour of day.

Signals to watch in production

SignalWhy it mattersWarning sign
log file sync average waitCommit latency. The single most important Oracle performance signal.>5ms on SSD, >20ms on SAN sustained
log file parallel write average waitLGWR’s actual I/O time. Compare with log file sync to isolate storage from scheduling.Elevated alongside log file sync means storage problem
Active sessions vs CPU coresPrimary load measure. Equivalent to load average on Linux.Active sessions sustained at more than 2x CPU cores
Redo log switch frequencyDirect indicator of redo throughput pressure.More than 6 switches per hour with current redo log sizing
Archive destination status and spaceSilent-hang risk when it fills.STATUS = ERROR or destination more than 95% full
Undo tablespace ACTIVE extentsWrite-path failure risk.ACTIVE undo more than 85% of tablespace
Session and process utilization vs limitsCliff-edge at ORA-00020.Current utilization above 85% of limit
Hard parse rateShared pool health.Hard parses above 100/sec sustained or above 5% of total parses
enq: TX waits with INACTIVE blockerLock cascade from uncommitted transactions.Any blocker with more than 10 waiters

V$ACTIVE_SESSION_HISTORY, AWR, and ADDM require the Diagnostics Pack license (Enterprise Edition only). Querying these without a license is a compliance violation that Oracle audits catch.

How Netdata helps

  • Correlating log file sync with log file parallel write across the same time window isolates redo storage latency from LGWR scheduling overhead.
  • Trending redo log switch frequency and archive destination free space together surfaces the approach to an archive hang before the database freezes.
  • Active session count tracked per second, broken down by dominant wait class, gives the fastest triage signal: Commit means LGWR, User I/O means storage, Concurrency means locks.
  • Tablespace and undo utilization trends with configurable thresholds catch cliff-edge failures (ORA-01653, ORA-30036, ORA-01555) before they hit.
  • OS-level signals (CPU run queue, I/O latency per device, memory pressure, OOM killer activity) alongside database wait events make it obvious whether a bottleneck is inside Oracle or underneath it.
  • Session and process utilization against PROCESSES limits, tracked over time, prevents the ORA-00020 connection exhaustion cliff.

See Oracle Database monitoring with Netdata for per-second metrics collection and alerting.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.