The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / oracle-database / oracle-database-ora-00020-maximum-processes-exceeded ▌

Operations Guides

ORA-00020: maximum number of processes exceeded

ORA-00020 is a connection-admission cliff-edge. At 99% of the configured PROCESSES limit, everything works. At 100%, new connections are refused and the error cascades into every monitoring tool, DBA session, and application pool. Existing sessions keep running, which is why this is often discovered late.

The symptom: applications report connection failures, the listener answers TCP but the database rejects the actual connect with ORA-00020, and your usual SYSDBA session may also be refused. The alert log fills with the error. If your connection pool retries in a tight loop, the situation worsens before it improves, because each retry is another failed admission attempt.

The fix is rarely urgent-only. Raising PROCESSES is a static parameter change that requires an instance restart. The harder work is identifying why process count climbed: a connection leak, an oversized pool, a retry storm, dead-connection accumulation, or an OS-level limit that bit before Oracle’s own.

What this means

ORA-00020 means the Oracle instance has exhausted its allocation of process state objects. The PROCESSES initialization parameter hard-limits the number of OS processes the instance can manage: dedicated server processes, shared servers, dispatchers, and background processes all draw from this pool.

  • PROCESSES is static. Changing it requires an instance restart. There is no online resize.
  • SESSIONS is derived from PROCESSES. From 12c onward the formula is approximately 1.5 * PROCESSES + 22. On 11g and earlier it is approximately 1.1 * PROCESSES + 5. Oracle silently bumps SESSIONS up if you set it below the derived value.
  • Background processes consume slots too. Expect roughly 40 to 70 process slots reserved for background processes (DBWn, LGWR, PMON, ARCn, MMON, and the rest). The user-available count is less than the raw limit.
  • Each dedicated server is one OS process plus one PGA allocation. Oracle’s limit and the OS limit both apply, and either can fail first.
  • In RAC, each instance has its own PROCESSES limit. One node can hit ORA-00020 while others sit idle if load balancing is misconfigured.
  • OS limits can bite first. The oracle user’s ulimit -u (nproc) and the kernel’s pid_max can be exhausted before Oracle’s own limit. When that happens you see fork failures and ORA-27xxx errors rather than ORA-00020.

This is a cliff-edge failure with no graceful degradation. There is no slowdown phase. Connection admission works, then it does not.

Common causes

CauseWhat it looks likeFirst thing to check
Connection pool misconfigurationProcess count climbs to the pool max and stays there; MAX_UTILIZATION plateaus at the limitApplication pool config: max connections, number of pools
Connection leakProcess count climbs monotonically and never returns to baseline after load dropsV$SESSION for long-lived INACTIVE sessions from the same MACHINE/PROGRAM
Retry stormRapid climb, application logs show connection retries, listener handler count spikesApplication retry logic and backoff policy
Dead connections (no DCD)INACTIVE sessions accumulate from crashed or firewalled clientsSQLNET.EXPIRE_TIME in sqlnet.ora
OS limit hit firstFork failures in alert log, ORA-27xxx errors, no row in V$RESOURCE_LIMIT at limitulimit -u for the oracle user; /proc/sys/kernel/pid_max
Lock contention cascadeSessions queue behind an uncommitted transaction, pool spins up new connections that also blockBLOCKING_SESSION in V$SESSION, enq: TX waits
Undersized PROCESSESSteady-state usage always near the limit, no single runaway causeV$RESOURCE_LIMIT trend, MAX_UTILIZATION history

Quick checks

Run these read-only checks to confirm where the wall is. All are safe to run on a production instance.

-- Check process and session utilization against limits
SELECT RESOURCE_NAME, CURRENT_UTILIZATION, MAX_UTILIZATION, LIMIT_VALUE
FROM V$RESOURCE_LIMIT
WHERE RESOURCE_NAME IN ('processes', 'sessions');

MAX_UTILIZATION is the high-water mark since instance startup. If it equals LIMIT_VALUE, the limit has already been hit at least once, even if current usage is lower now.

# Run as the oracle user, not as root, to see the real nproc in effect
ulimit -u

# Check the kernel pid_max
cat /proc/sys/kernel/pid_max

# Count oracle processes at the OS level (adjust the username if yours differs)
ps -u oracle -o pid= | wc -l
# Grep the alert log for ORA-00020 and OS-level fork failures
adrci exec="show alert -tail 200" | grep -E "ORA-00020|ORA-27"
-- Identify which machines and programs are holding the most sessions
SELECT MACHINE, PROGRAM, COUNT(*) AS sessions,
       SUM(CASE WHEN STATUS = 'INACTIVE' THEN 1 ELSE 0 END) AS inactive
FROM V$SESSION
WHERE TYPE = 'USER'
GROUP BY MACHINE, PROGRAM
ORDER BY sessions DESC
FETCH FIRST 20 ROWS ONLY;
# Confirm the listener is up and check handler availability
lsnrctl status LISTENER

A listener that is up but reports services as BLOCKED, or returns TNS-12516 (no available handler) or TNS-12519 (too many connections), is the same exhaustion seen from the client side.

How to diagnose it

The first decision is whether Oracle’s limit or the OS limit is the binding constraint. The second is whether the saturation is steady-state, a leak, or a spike.

flowchart TD
  A[ORA-00020 reported] --> B{Can SYSDBA connect?}
  B -- No --> C[Use sqlplus -prelim as sysdba]
  B -- Yes --> D[Query V$RESOURCE_LIMIT]
  C --> D
  D --> E{CURRENT = LIMIT?}
  E -- No --> F[OS limit hit first: check ulimit -u and pid_max]
  E -- Yes --> G[Oracle PROCESSES exhausted]
  G --> H{MAX_UTILIZATION at LIMIT for long?}
  H -- Steady at limit --> I[Undersized PROCESSES: resize]
  H -- Recent spike --> J{INACTIVE sessions accumulating?}
  J -- Yes --> K[Connection leak or no DCD]
  J -- No, mostly ACTIVE --> L[Lock cascade or retry storm]
  1. Confirm the binding constraint. Query V$RESOURCE_LIMIT. If CURRENT_UTILIZATION for processes equals LIMIT_VALUE, Oracle’s limit is the wall. If current is below the limit but you are still seeing failures, the OS limit (ulimit -u or pid_max) is the culprit. Check the alert log for ORA-27xxx errors or fork failures to confirm.

  2. Identify the slot consumers. Group V$SESSION by MACHINE and PROGRAM. One application host or one program holding a disproportionate share points the finger. Cross-reference with the application team’s pool configuration.

  3. Distinguish leak from steady-state. If MAX_UTILIZATION has been at the limit for the whole uptime and current usage is always near the limit, the parameter is undersized for the workload. If current usage climbed recently and never returned to baseline, suspect a leak. If it spiked sharply, suspect a retry storm or a lock cascade driving pool growth.

  4. Check for dead connections. Count INACTIVE sessions grouped by MACHINE. INACTIVE sessions from hosts that should have disconnected are candidates for dead-connection cleanup. Confirm whether SQLNET.EXPIRE_TIME is set in sqlnet.ora. Without it, clients that crash or get firewalled leave server processes behind as INACTIVE sessions consuming process slots indefinitely.

  5. Check for a lock contention cascade. A single idle session holding an uncommitted transaction can queue dozens of waiters. Application connection pools, seeing latency, spin up new connections that also block, exhausting PROCESSES. Query V$SESSION for EVENT LIKE 'enq: TX%' and check BLOCKING_SESSION. See the blocking sessions guide for the full chain-walking procedure.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
V$RESOURCE_LIMIT.CURRENT_UTILIZATION for processesDirect measure of slot usage against the hard limitPeak above 75% of limit
V$RESOURCE_LIMIT.MAX_UTILIZATIONHigh-water mark since instance startupEquals LIMIT_VALUE: already hit at least once
OS process count for the oracle userOS-level resource that can fail before Oracle’s limitApproaching ulimit -u
SQLNET.EXPIRE_TIME configurationEnables Dead Connection Detection to reclaim zombie slotsNot set, or set too high
Listener handler availabilityAdmission path beyond the TCP probeTNS-12516, TNS-12519, services BLOCKED
INACTIVE session countDead and leaked connection candidatesGrowing without churn
Application pool saturation errorsDemand-side pressure driving the exhaustionPool-exhausted errors in app logs

For historical trending, DBA_HIST_RESOURCE_LIMIT holds snapshots of the same utilization data over time, useful for capacity planning when AWR is licensed and the retention covers the period you care about.

Fixes

Immediate: regain access

When ORA-00020 prevents even a SYSDBA connection, use the preliminary connection mode to bypass the process limit:

# Emergency connection that does not consume a process slot
sqlplus -prelim / as sysdba

From there, SHUTDOWN ABORT is the fastest path to recovery. The consequences matter: in-flight transactions are not rolled back at shutdown; SMON performs crash recovery on the next STARTUP, which can take significant time on a busy system. This is a last-resort tool, not a routine fix.

If you can connect normally, the less disruptive path is to free slots by killing sessions. ALTER SYSTEM KILL SESSION is safe in the sense that it targets a specific session, but it is still disruptive to that session’s work. Identify targets first:

-- Identify the sessions to kill first
SELECT SID, SERIAL#, USERNAME, MACHINE, PROGRAM, STATUS, SQL_ID
FROM V$SESSION
WHERE TYPE = 'USER' AND STATUS = 'INACTIVE'
ORDER BY LOGON_TIME;

-- Kill a specific session. Disruptive to that session.
ALTER SYSTEM KILL SESSION 'sid,serial#' IMMEDIATE;

ALTER SYSTEM KILL SESSION does not always free the process slot immediately. Killed sessions can linger in a KILLED state until PMON cleans them up. In some cases an OS process persists without a corresponding V$SESSION entry, still consuming a slot. If PMON is not keeping up, the only reliable cleanup is an instance restart.

Short-term: stop the bleeding

  • Throttle retry logic. If the application retries failed connections in a tight loop, each retry is another failed admission attempt. Add backoff. This alone can break the cascade.
  • Kill leaky or idle sessions. Long-lived INACTIVE sessions from a misbehaving application host are the usual suspects.
  • Enable Dead Connection Detection. Set SQLNET.EXPIRE_TIME in sqlnet.ora to a nonzero value (minutes, per the Net Services reference). The server then probes idle connections and reclaims slots from dead clients. This prevents zombie accumulation.

Long-term: resize and re-architect

  • Raise PROCESSES. This is static. Plan a maintenance window. SESSIONS will derive upward automatically; TRANSACTIONS will too if it is unset. If you set TRANSACTIONS explicitly, raise it to match.
  • Fix the OS ulimit. Oracle on Linux recommends an nproc of 65536 or unlimited (both hard and soft) in /etc/security/limits.conf. The default soft nproc on RHEL/OEL is often 1024, which is far too low for a production database.
  • Right-size the connection pool. The sum of all application pool maximums across all hosts must fit comfortably under PROCESSES minus background overhead minus DBA headroom. Multiple application instances each opening a large pool is a common cause.
  • Consider shared server. For workloads with many idle sessions (OLTP with think time), shared server (formerly MTS) allows more sessions per process. It has tradeoffs: certain features are unavailable or restricted, and official guidance cautions against it for batch and long-running work. Evaluate before adopting.

Prevention

  • Leave 25% headroom. Peak usage should not exceed 75% of PROCESSES. The remainder covers background processes, DBA sessions during incidents, and unexpected spikes.
  • Set SQLNET.EXPIRE_TIME. Dead Connection Detection is the single most effective preventative for zombie-driven ORA-00020.
  • Tune ulimit proactively. Do not wait for fork failures. Set nproc to 65536 or unlimited for the oracle user before production load.
  • Monitor MAX_UTILIZATION as a capacity signal. If it is climbing toward the limit over weeks, resize before the cliff.
  • Walk blocking chains proactively. Lock contention cascades exhaust processes by driving pool growth. Catch the blocker early.
  • Match pool size to limit. Document the math: sum of pool maximums plus background overhead plus headroom must be under PROCESSES.

How Netdata helps

  • Per-second process and session utilization from V$RESOURCE_LIMIT catches the climb toward the limit before the cliff, not after.
  • MAX_UTILIZATION tracking surfaces the high-water mark so you know whether the limit has ever been hit, even if current usage looks fine.
  • Correlation with OS process count and ulimit distinguishes an Oracle-limit failure from an OS-limit failure without switching tools.
  • Listener handler availability signals (TNS-12516, TNS-12519, BLOCKED services) show the admission path failing before clients report it.
  • Anomaly detection on connection churn flags retry storms and leaks as they form, rather than after ORA-00020 fills the alert log.

Netdata’s Oracle Database monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Oracle Database monitoring with Netdata

Netdata monitors Oracle Database with per-second metrics and automatic dashboards. Watch wait events, redo and archive-log activity, tablespace and undo space, and session and lock activity so the failure modes in these runbooks surface before the instance hangs.