The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / oracle-database-pga-memory-pressure ▌

Operations Guides

Oracle PGA memory pressure: over-allocation, temp spills, and PGA_AGGREGATE_TARGET

Program Global Area (PGA) is the private memory each dedicated server process uses for sorting, hashing, bitmap operations, and session state. It is allocated outside the SGA, one chunk per process, and managed as an aggregate pool with a soft target (PGA_AGGREGATE_TARGET) and, from Oracle 12c onward, a hard ceiling (PGA_AGGREGATE_LIMIT). PGA pressure has two opposite failure signatures, and the operator’s job is to tell them apart quickly.

The first signature is “too little PGA”: work areas cannot fit, the optimizer still picks hash joins and sorts, and Oracle silently spills those work areas to temp tablespace. The wait events direct path read temp and direct path write temp show up, temp fills faster, and query latency rises without a clean error. The second signature is “too much PGA”: total allocation grows past the soft target toward the hard limit, and either sessions hit ORA-04036 or, if PGA plus SGA exceeds the box, the Linux OOM killer targets Oracle server processes.

Both directions can coexist on a mixed-workload instance. Raising the target fixes spills but raises OOM risk; lowering it fixes OOM risk but increases spills.

What this means

PGA_AGGREGATE_TARGET is a soft target. Oracle tries to keep the sum of all PGA below it, but individual work areas can exceed the per-session optimal when the workload demands it. When that happens, the over allocation count statistic in V$PGASTAT increments. A non-zero and growing over allocation count means the target is too small for the actual workload. It is not, by itself, an incident. Some over-allocation is expected on analytics. It becomes a problem only when it correlates with declining cache hit percentage, rising direct path read temp waits, or ORA-04036.

PGA_AGGREGATE_LIMIT (12c+) is a hard ceiling. When the instance hits it, Oracle interrupts calls or sessions with the most untunable PGA memory to get back under the limit; the affected call can fail with ORA-04036. On small VMs the default limit computation can land close to physical RAM, so ORA-04036 sometimes appears shortly after upgrade without any workload change. Check dmesg and the alert log together: OOM killer entries mean the OS killed Oracle; ORA-04036 means Oracle killed its own session before the OS got involved.

The cache hit percentage in V$PGASTAT is computed as (bytes processed * 100) / (bytes processed + extra bytes read/written for spills). Sustained below 80% means significant work is being done in temp that could fit in PGA. The target is a workload-dependent tuning knob, not an absolute. OLTP instances may run at 100% with a small target; mixed or analytical instances need a much larger target to keep the same hit percentage.

flowchart TD
    A[Sort or hash join
needs work area] --> B{Fits in PGA?} B -- Yes --> C[Optimal execution
in memory] B -- No --> D[Spill to temp
direct path read/write temp] D --> E[Cache hit pct drops
over-allocation grows] E --> F{Total PGA near
PGA_AGGREGATE_LIMIT?} F -- Yes --> G[ORA-04036
call aborted or session killed] F -- No, but SGA+PGA
exceeds physical RAM --> H[Linux OOM killer
targets Oracle PID] H --> I[Random session death
ORA-27300 / ORA-03113]

Common causes

CauseWhat it looks likeFirst thing to check
PGA_AGGREGATE_TARGET undersized for workloadover allocation count non-zero and growing, cache hit percentage below 80%, direct path read temp waits risingV$PGASTAT deltas on over allocation count and cache hit percentage
Single large operation dominating PGAOne or two sessions in V$PROCESS with PGA_ALLOC_MEM an order of magnitude above the rest; ORA-04036 on specific sessionsV$PROCESS ordered by PGA_ALLOC_MEM
PGA_AGGREGATE_LIMIT too close to physical RAMORA-04036 with no runaway query, alert log entries, dmesg OOM tracesdmesg | grep -i oom, V$PGASTAT total PGA allocated vs limit
AMM (MEMORY_TARGET) on LinuxSGA in /dev/shm, no hugepages, page table overhead, /dev/shm fillingshow parameter MEMORY_TARGET, /proc/meminfo HugePages_Total
Mixed OLTP and analytics on same instanceOLTP fine during the day, spills and ORA-04036 during batch or report windowsV$PGASTAT and temp usage over time, batch schedule
PGA leak in PL/SQLPGA_MAX_MEM per process grows monotonically and never releases after statement completionV$PROCESS.PGA_MAX_MEM over session lifetime

Quick checks

Run these read-only. None change instance state.

# Confirm whether the OOM killer has been touching Oracle processes
dmesg -T | grep -iE "out of memory|oom|killed process" | tail -30
-- Core PGA health from V$PGASTAT
SELECT NAME, VALUE
FROM V$PGASTAT
WHERE NAME IN (
  'aggregate PGA target parameter',
  'aggregate PGA auto target',
  'total PGA inuse',
  'total PGA allocated',
  'maximum PGA allocated',
  'total freeable PGA memory',
  'over allocation count',
  'cache hit percentage'
);
-- Top PGA consumers right now (12c+ FETCH FIRST; use ROWNUM <= 20 on 11g)
SELECT p.SPID, p.PGA_USED_MEM/1048576  AS pga_used_mb,
       p.PGA_ALLOC_MEM/1048576 AS pga_alloc_mb,
       p.PGA_MAX_MEM/1048576   AS pga_max_mb,
       s.SID, s.USERNAME, s.SQL_ID
FROM V$PROCESS p
LEFT JOIN V$SESSION s ON p.ADDR = s.PADDR
ORDER BY p.PGA_ALLOC_MEM DESC
FETCH FIRST 20 ROWS ONLY;
-- Temp spill waits (PGA undersize indicator)
SELECT EVENT, TOTAL_WAITS, TIME_WAITED_MICRO,
       ROUND(TIME_WAITED_MICRO/NULLIF(TOTAL_WAITS,0)/1000, 2) AS avg_ms
FROM V$SYSTEM_EVENT
WHERE EVENT IN ('direct path read temp', 'direct path write temp');
-- Confirm memory management mode
SELECT NAME, VALUE, ISDEFAULT
FROM V$PARAMETER
WHERE NAME IN ('pga_aggregate_target',
               'pga_aggregate_limit',
               'memory_target',
               'memory_max_target',
               'sga_target',
               'workarea_size_policy');

How to diagnose it

  1. Confirm direction. Read V$PGASTAT. If over allocation count is growing and cache hit percentage is below 80%, the problem is undersized PGA. If total PGA allocated is hovering near PGA_AGGREGATE_LIMIT and ORA-04036 appears in the alert log, the problem is oversized sessions relative to the limit. If dmesg shows OOM kills, the problem is total Oracle memory versus physical RAM, not PGA tuning alone.
  2. Find the heavy consumers. Sort V$PROCESS by PGA_ALLOC_MEM. A normal OLTP session is in the tens of MB. A hash join against a large build table or a large sort can be GB-scale. If one or two SPIDs dominate, the fix is at the SQL or workload level, not the target.
  3. Correlate with workload. Match the pressure window to the batch schedule, the stats gathering job, or the reporting window. PGA pressure that only appears during DBMS_STATS or during a known ETL is a workload-isolation problem first and a tuning problem second.
  4. Check the memory management mode. If MEMORY_TARGET is set, AMM is in use. On Linux that means the SGA lives in /dev/shm instead of hugepages, page table overhead grows with process count, and /dev/shm can fill independently of PGA. ASMM (SGA_TARGET plus PGA_AGGREGATE_TARGET) is the recommended production configuration on Linux.
  5. Check the advice view. V$PGA_TARGET_ADVICE predicts cache hit percentage and over allocation count at a range of target sizes. Oracle’s tuning guide enables generation by setting PGA_AGGREGATE_TARGET with automatic PGA memory management and STATISTICS_LEVEL to TYPICAL (the default) or ALL; BASIC disables it. Confirm any Oracle management-pack entitlement separately before relying on advisory data.
  6. Rule out undo cross-talk. Undo pressure and PGA pressure both surface during long-running reports, but the fixes are different. If the failing queries are reads that error with ORA-01555, the root cause is undo, not PGA. See the ORA-01555 snapshot too old guide for that path.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
V$PGASTAT over allocation countCumulative counter; growing rate means the soft target is too small for the workloadSlope increasing day over day at the same workload
V$PGASTAT cache hit percentageHow much work stayed in PGA vs spilled to tempSustained below 80%, or sudden drop correlated with a workload change
V$PGASTAT total PGA allocatedCurrent aggregate PGA, compare against target and limitSustained above PGA_AGGREGATE_TARGET or approaching PGA_AGGREGATE_LIMIT
V$PGASTAT maximum PGA allocatedHigh-water mark since startupApproaching PGA_AGGREGATE_LIMIT flags ORA-04036 risk on the next peak
V$PROCESS.PGA_ALLOC_MEM per SPIDIdentifies runaway sessions and PGA leaksSingle process holding GB-scale PGA while peers hold MB-scale
direct path read temp / direct path write temp waitsVisible cost of undersized PGAGrowing share of total DB time
Temp tablespace utilizationSpills consume temp; ORA-01652 is the cliffSee the ORA-01652 unable to extend temp guide
OS memory: SGA + PGA vs physical RAMOOM killer territoryLinux MemAvailable declining, /proc/meminfo HugePages_Free at zero
dmesg OOM entriesConfirms the OS killed Oracle, not Oracle killing itselfAny new “Out of memory: Kill process” line referencing an Oracle PID

Fixes

PGA_AGGREGATE_TARGET is undersized

Use V$PGA_TARGET_ADVICE to pick a new target that lifts cache hit percentage above 90% without forcing total PGA allocated against the hard limit. Apply the change at runtime:

ALTER SYSTEM SET PGA_AGGREGATE_TARGET = <value>M SCOPE=BOTH;

This is dynamic and reversible. Watch over allocation count (cumulative; track the rate, not the absolute) and cache hit percentage for the next maintenance window. If spill waits drop and temp utilization stabilizes, the change worked. If the rate of over-allocation does not change, the workload has a hard lower bound on PGA that target tuning alone cannot fix.

A single session or SQL is consuming too much PGA

Tuning the target will not help here. Identify the SQL_ID from V$PROCESS joined to V$SESSION, then look at the execution plan. Common patterns: a hash join with a much larger build input than the optimizer estimated, a sort on an unindexed column set, or a PL/SQL collection growing without bound. Fix the plan (statistics, SQL profile, SQL plan baseline) or partition the workload. If the consumer is a runaway batch job, killing the session is a valid short-term intervention; it will not change recurrence.

-- WARNING: this terminates the session immediately.
-- Roll back uncommitted work and may cascade to dependent sessions.
ALTER SYSTEM KILL SESSION 'sid,serial#' IMMEDIATE;

PGA_AGGREGATE_LIMIT is too close to physical RAM

Lower the limit if the OS OOM killer is firing, or raise it if ORA-04036 is firing without OS pressure. Both are dynamic:

ALTER SYSTEM SET PGA_AGGREGATE_LIMIT = <value>M SCOPE=BOTH;

When MEMORY_TARGET is not set, the documented default is normally 200% of PGA_AGGREGATE_TARGET; if that target is explicitly 0, it is 90% of physical memory minus SGA. In all cases the default is at least 2 GB and at least 3 MB × PROCESSES (5 MB × PROCESSES for Oracle RAC). Do not set PGA_AGGREGATE_LIMIT below its computed default: Oracle documents that instance startup will fail. A value of 0 disables the limit. For a PDB, the limit must be at least twice that PDB’s PGA_AGGREGATE_TARGET; violating the range can return ORA-00093, whose meaning is “parameter outside valid range.” On small VMs, the 2 GB or PROCESSES-derived floor can land close to physical RAM and produce ORA-04036 shortly after upgrade. The alert log warning “pga_aggregate_limit value is too high for the amount of physical memory” does not block startup but signals the mismatch.

AMM is in use on Linux

Move to ASMM. This is not a runtime-only change: removing MEMORY_TARGET and adopting SGA_TARGET plus PGA_AGGREGATE_TARGET requires a planned restart, and you must provision hugepages for the SGA first. Verify with /proc/meminfo that HugePages_Total covers the SGA and HugePages_Free is non-zero. Without hugepages, each Oracle server process carries page table overhead that does not appear in V$PGASTAT or V$SGASTAT but does count against physical RAM. The OOM killer sees that overhead; Oracle does not. AMM also couples SGA and PGA in ways that defeat hugepages even when they are configured, which is the core reason it is not recommended for production on Linux.

Mixed OLTP and analytics

If OLTP is fine during the day and pressure appears only during batch or reporting windows, isolate the workloads. Options include an Active Data Guard standby for reads, a separate reporting instance, Resource Manager directives, or rescheduling the batch window. Tuning PGA_AGGREGATE_TARGET for the peak analytical workload over-provisions PGA for the OLTP trough and may bring ORA-04036 or OOM risk back during the wrong window.

Prevention

  • Track over allocation count rate, not the absolute. It is cumulative since startup. The diagnostic signal is the slope, sampled at consistent workload windows.
  • Track cache hit percentage against a workload-specific baseline. A mixed workload that lives at 85% may be normal; the same instance dropping from 95% to 75% after a stats gather is not.
  • Watch V$PROCESS.PGA_MAX_MEM for monotonic growth. A session whose max PGA keeps climbing after the statement that needed it has finished is a leak, not a tuning problem.
  • Size for physical RAM, not for PGA_AGGREGATE_TARGET in isolation. The real budget is physical_RAM - SGA - OS_overhead - other_processes. PGA lives inside that envelope.
  • Use ASMM plus hugepages on Linux. AMM is workable on small systems but is not the production recommendation on Linux.
  • Validate PGA_AGGREGATE_LIMIT after every upgrade. The limit was introduced in 12c and the default computation has changed across releases. What was fine on 12.1 may behave differently on 19c.

How Netdata helps

Netdata’s per-second collection turns PGA pressure from a guessed-at tuning problem into a correlation problem. The signals worth wiring together:

  • V$PGASTAT rows (over allocation count, cache hit percentage, total PGA allocated, maximum PGA allocated) collected at per-second granularity, so a burst of over-allocation during a batch window shows up against the OLTP baseline.
  • V$PROCESS.PGA_ALLOC_MEM per SPID, so the single runaway session appears as a spike alongside the aggregate trend rather than being hidden inside it.
  • direct path read temp and direct path write temp wait time, so spills correlate with cache hit percentage drops and temp tablespace growth.
  • OS memory, MemAvailable, hugepages utilization, and dmesg OOM entries on the same timeline, so ORA-04036 in the alert log lines up with OOM killer activity on the host.
  • ML anomaly detection on total PGA allocated and cache hit percentage, which catches slow drift toward the limit before ORA-04036 fires.

Netdata’s Oracle Database monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.