The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / oracle-database-ora-04036-pga-limit-exceeded ▌

Operations Guides

ORA-04036: PGA memory used by the instance exceeds PGA_AGGREGATE_LIMIT

ORA-04036 fires when instance-wide PGA consumption crosses PGA_AGGREGATE_LIMIT, the hard cap introduced in 12c. This is not a tuning warning. The database is defending itself by killing or interrupting work.

PGA_AGGREGATE_TARGET is a soft target Oracle tries to honor. PGA_AGGREGATE_LIMIT is an enforced ceiling. Sessions can exceed the target legitimately. Crossing the limit triggers Oracle to abort the call of, and then terminate, the sessions holding the most untunable PGA. SYS and most background processes are not eligible for termination, which produces a distinct and more dangerous failure mode described below.

The Linux OOM killer may also be involved if SGA (hugepages) + PGA + OS needs exceed physical RAM, but ORA-04036 itself is Oracle’s internal mechanism. Both can fire in the same incident.

What this means

When aggregate PGA crosses PGA_AGGREGATE_LIMIT:

  • The CKPT background process checks PGA usage every three seconds.
  • Oracle first aborts the call of the session holding the most untunable PGA. Untunable memory is memory Oracle cannot release on demand: sort areas in use, PL/SQL collections, certain call memory.
  • If the instance is still over the limit after the abort, Oracle terminates the session.
  • SYS sessions and most background processes are normally not the practical enforcement target; trace/alert output identifies any top consumers that are not eligible for ORA-4036 interrupts. Treat the exact exempt set as release-dependent and read the incident dump rather than assuming a fixed process list.

The last point is the dangerous variant. When a background process is the top PGA consumer, Oracle cannot enforce the limit against it. The alert log will say something like:

PGA_AGGREGATE_LIMIT has been exceeded but some processes using the most PGA memory are not eligible to receive ORA-4036 interrupts.

In that case ORA-04036 still fires for user sessions, but the root cause keeps growing. Killing user sessions will not self-correct the instance.

In 19c and later, the incident dump produced when ORA-04036 fires includes a “Top 10 processes” section showing private memory usage per Oracle process. Read that dump before guessing.

flowchart TD
    A[Aggregate PGA rises] --> B{Crosses PGA_AGGREGATE_LIMIT?}
    B -->|No| Z[Normal operation]
    B -->|Yes| C[CKPT detects every 3s]
    C --> D[Abort call of highest-PGA session]
    D --> E{Still over limit?}
    E -->|No| Z
    E -->|Yes| F[Terminate the session]
    F --> G{Top consumer user or background?}
    G -->|User| Z
    G -->|Background| H[Not eligible for ORA-4036
PGA written to trace] H --> I[Persistent overage
user sessions keep dying]

Common causes

CauseWhat it looks likeFirst thing to check
Unbounded PL/SQL collection or LOBOne session’s PGA_MAX_MEM grows monotonically over its lifetimeV$PROCESS.PGA_MAX_MEM per session
Massive serial sort or hash joinSingle SQL doing direct path read temp / direct path write temp, one big consumerV$SQL_WORKAREA_ACTIVE, top PGA consumers
Background PGA leak (MMON, space slaves)“Not eligible to receive ORA-4036 interrupts” in alert logIncident trace “Top 10 processes” section
Too many sessions with moderate PGAtotal PGA allocated tracks session count linearlyV$PGASTAT vs V$SESSION count
PGA_AGGREGATE_LIMIT set too low for workloadORA-04036 fires at start of batch or right after restartSHOW PARAMETER pga_aggregate_limit
OS memory pressure as a confounderdmesg OOM kills alongside ORA-04036; ORA-27300, ORA-27301dmesg, /proc/meminfo

Quick checks

All read-only. Run as DBA.

-- Show PGA configuration
SHOW PARAMETER pga_aggregate_target;
SHOW PARAMETER pga_aggregate_limit;
SHOW PARAMETER memory_target;       -- AMM check
SHOW PARAMETER sga_target;
-- Aggregate PGA state from V$PGASTAT
SELECT NAME, VALUE
FROM V$PGASTAT
WHERE NAME IN (
  'aggregate PGA target parameter',
  'total PGA inuse',
  'total PGA allocated',
  'maximum PGA allocated',
  'over allocation count',
  'total freeable PGA memory',
  'cache hit percentage'
);
-- Top PGA consumers, with OS PID and SQL_ID
SELECT p.SPID,
       p.PGA_USED_MEM/1048576  AS pga_used_mb,
       p.PGA_ALLOC_MEM/1048576 AS pga_alloc_mb,
       p.PGA_MAX_MEM/1048576   AS pga_max_mb,
       s.SID, s.USERNAME, s.SQL_ID, s.PROGRAM
FROM V$PROCESS p
LEFT JOIN V$SESSION s ON p.ADDR = s.PADDR
ORDER BY p.PGA_ALLOC_MEM DESC
FETCH FIRST 20 ROWS ONLY;
-- Per-category breakdown for top offenders (12c+)
SELECT PID, SERIAL#, NAME, CATEGORY,
       ALLOCATED/1048576 AS alloc_mb,
       USED/1048576      AS used_mb,
       MAX_ALLOCATED/1048576 AS max_mb
FROM V$PROCESS_MEMORY
ORDER BY ALLOCATED DESC
FETCH FIRST 20 ROWS ONLY;
-- Sessions currently spilling to temp (PGA undersized indicator)
SELECT EVENT, TOTAL_WAITS, TIME_WAITED_MICRO
FROM V$SYSTEM_EVENT
WHERE EVENT IN ('direct path read temp', 'direct path write temp');
# Confirm ORA-04036 in the alert log and look for ineligible-process text
adrci exec="show alert -tail 500" | grep -E "ORA-04036|ORA-4036|not eligible"

# Rule out OS OOM killer as a confounder
dmesg -T | grep -iE "out of memory|oom|killed process"

How to diagnose it

  1. Confirm ORA-04036 in the alert log. Note whether the “not eligible to receive ORA-4036 interrupts” message appears alongside it. That changes the playbook.
  2. Snapshot V$PGASTAT. Compare total PGA allocated against PGA_AGGREGATE_LIMIT and PGA_AGGREGATE_TARGET. The ratio tells you whether you are nudging the limit or blowing through it.
  3. Identify the top PGA consumer in V$PROCESS and break it down with V$PROCESS_MEMORY. The CATEGORY column separates SQL, PL/SQL, OLAP, Java, and Other. A session whose PL/SQL category is huge is leaking collections. A session whose SQL category is huge is doing a large workarea operation.
  4. If the top consumer is a background process (MMON, a space background slave, DBWn), open the incident dump in the ADR. In 19c+ the dump contains a “Top 10 processes” section. Confirm which background process is the offender.
  5. Rule out OS OOM killer. If dmesg shows Oracle PIDs being killed, the real ceiling is physical RAM, not PGA_AGGREGATE_LIMIT. The fix is to free RAM (hugepages for SGA, reduce PGA footprint, move other workloads off the host), not to raise the limit.
  6. Correlate timing. Did ORA-04036 start after a stats gather, a batch window, a datapatch run, an upgrade, or a connection pool resize? PGA leaks often surface after a code change that introduced an unbounded loop.
  7. Check direct path read temp / direct path write temp trends. If they rise in step with PGA, the workload is genuinely outgrowing PGA. If PGA rises while temp spill stays flat, suspect a leak rather than legitimate demand.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
total PGA allocated / PGA_AGGREGATE_LIMITHeadroom before the hard capSustained above 80%
over allocation count in V$PGASTATPersistent over-target pressureNon-zero and growing
cache hit percentage in V$PGASTATRatio of work done in PGA vs spilled to tempBelow 80% sustained
direct path read temp / direct path write tempPGA is undersized for the workRising trend
Per-process PGA_MAX_MEMDetects slow leaks in long-lived sessionsSingle PID growing monotonically
dmesg OOM killsOS-level memory ceilingAny Oracle-related entries
Alert log: “not eligible” textBackground process is the top consumerAny occurrence

Fixes

Immediate: stop the bleeding

If a single user session is the top consumer and you can identify the offending SQL_ID or program, kill the session. Rolling back an uncommitted transaction in that session releases its PGA.

-- Disruptive: kills the target session immediately
ALTER SYSTEM KILL SESSION 'sid,serial#' IMMEDIATE;

Use this when the cause is a runaway user session, not a background process.

Short-term: raise the limit

PGA_AGGREGATE_LIMIT is dynamic and can be raised without a restart:

-- Dynamically raise the limit, both spfile and memory
ALTER SYSTEM SET pga_aggregate_limit = <new_value> SCOPE=BOTH;

Do this only if physical RAM has headroom. Raising the limit above what the host can actually hold moves the failure from ORA-04036 to the OS OOM killer, which is worse because the OOM killer does not distinguish background from foreground processes.

To disable the limit entirely as an emergency workaround, for example during a datapatch run known to spike PGA:

-- Emergency only. Removes Oracle's self-protection.
ALTER SYSTEM SET pga_aggregate_limit = 0 SCOPE=BOTH;

Setting the limit to 0 disables it. Use it only for the duration of the operation, then restore a sane value.

Medium-term: reduce demand

  • Fix the PL/SQL leak. Look for collections populated in a loop without a DELETE, session-level LOBs held across calls, or associative arrays used as caches that never age out.
  • Reduce PGA_AGGREGATE_TARGET if it has drifted higher than the workload needs. Oracle sizes per-session workareas off the target; an oversized target allows individual sessions to grab more memory before Oracle starts spilling to temp.
  • Move analytical workloads to an Active Data Guard standby (licensed) or a separate instance so a single hash join cannot starve OLTP.

Confounder: free OS memory first

If the issue is total RAM, not Oracle’s cap:

  • Verify hugepages are configured for the SGA. Without hugepages, page table overhead can consume hundreds of MB per Oracle process on large SGAs. Check /proc/meminfo for HugePages_Total, HugePages_Free, and HugePages_Rsvd.
  • Avoid MEMORY_TARGET (AMM) on Linux. It uses /dev/shm instead of hugepages and makes memory leak diagnosis harder.
  • Move monitoring agents, backup scripts, or other tenants off the Oracle host.

Background-process leak: targeted workaround

If a background process is the top consumer, raising the limit only delays the next incident. Look for known defects and underscore workarounds relevant to your version. For space-management slave pressure, the incident dump identifies the allocation category and Wnnn/SMCO consumers. If Oracle Support confirms a release defect, reducing _max_spacebg_slaves can limit the blast radius; treat that parameter as a Support-directed workaround, not a routine tuning knob.

Prevention

  • Monitor V$PGASTAT continuously. Track total PGA allocated, over allocation count, and cache hit percentage. The leading indicator is over allocation count rising before ORA-04036 ever fires.
  • Track per-process PGA_MAX_MEM. A session whose high-water mark keeps climbing across snapshots is leaking. Catch it before it becomes the top consumer.
  • Size PGA_AGGREGATE_LIMIT with explicit headroom. Peak total PGA allocated should stay below 80% of the limit. From 18c onward, MGA (Managed Global Area) is included in the PGA aggregate limit, so size the limit from measured MGA plus PGA rather than a fixed per-process multiplier. Use V$PGASTAT, incident dumps, and Oracle Support guidance for the release-specific sizing.
  • Confirm hugepages for SGA. This frees the memory Oracle would otherwise burn on page tables.
  • Avoid MEMORY_TARGET on Linux. It conflicts with hugepages and obscures leak diagnosis.
  • Validate the 19c+ “Top 10 processes” dump works in your environment. When ORA-04036 fires, you want that dump to be there.
  • Watch for the “not eligible” alert log text. It means the fix is not “kill user sessions”. It means a background process has a leak and you need Oracle Support.
  • For datapatch and other known PGA-spiking operations, temporarily raise or disable PGA_AGGREGATE_LIMIT for the duration, then restore it.

How Netdata helps

  • Per-second PGA utilization from V$PGASTAT (total PGA inuse, total PGA allocated, over allocation count, cache hit percentage) lets you see the climb toward the limit before ORA-04036 fires, not just after.
  • Correlate PGA with TPS, active sessions, and wait events. If direct path read temp waits climb alongside total PGA allocated, the workload is outgrowing PGA. If PGA climbs while TPS is flat, suspect a leak.
  • Per-process PGA tracking surfaces a single PID whose PGA_MAX_MEM grows monotonically, the signature of a slow PL/SQL or LOB leak.
  • Linux memory pressure signals on the same timeline (OOM kills from dmesg, hugepage utilization, slab, free memory) let you distinguish Oracle’s internal cap from the host’s physical RAM ceiling.
  • Anomaly detection on over allocation count and cache hit percentage flags the leading indicators of PGA pressure even when absolute values look fine.
  • Alert log parsing for ORA-04036 and the “not eligible” text routes you straight to the background-process variant when it occurs.

For the full integration, see Oracle Database monitoring with Netdata.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.