The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Observability

What Is Distributed Tracing How It Works And Use Cases

A complete guide to understanding request flows in modern distributed systems
by Netdata Team · June 10, 2025

Your team has deployed a new feature, but users are reporting that a specific action in your application is painfully slow. You look at the metrics for your frontend service, and they seem fine. You check the logs for the user authentication service, and there are no errors. The database CPU usage is normal. So where is the bottleneck? In a modern distributed architecture built with microservices, a single user request can trigger a complex chain reaction across dozens of independent services. Finding the root cause of a problem can feel like searching for a needle in a global haystack.

This is where distributed tracing becomes an indispensable tool. It moves beyond isolated logs and metrics to give you a complete, end-to-end view of a request’s journey through your entire system. By understanding what distributed tracing is and how it works, you can transform your troubleshooting process from a slow, frustrating guessing game into a fast, data-driven investigation.

Why Traditional Monitoring Fails in Modern Architectures

In the age of monolithic applications, monitoring was relatively straightforward. All the code ran within a single process on a single server. To debug an issue, you could attach a debugger, analyze a stack trace, or read through a single log file to understand the sequence of events.

The shift to distributed tracing in microservices architectures has changed everything. Applications are now composed of dozens or even hundreds of small, independently deployable services that communicate over the network. While this approach offers incredible scalability and agility, it introduces significant observability challenges:

  • Lack of Visibility: A single user request might travel through an API gateway, an authentication service, a product catalog service, a payment processor, and a notification service. No single team has a complete view of this entire flow.
  • Cascading Failures: A small issue in a downstream service (like a slow database query) can cause a cascade of timeouts and errors in all the services that depend on it.
  • Pinpointing Latency: When a request is slow, it’s difficult to determine which of the many service-to-service calls is the culprit. Is it network latency, a slow business logic operation, or a delayed external API call?

Traditional tools fall short because they look at each service in isolation. A log file from one service only tells you its part of the story. A metric like CPU usage doesn’t explain the context of the work being done. You need a way to connect the dots across service boundaries.

What is Distributed Tracing? A Deeper Look

Distributed tracing is a method used to profile and monitor applications, especially those built using a microservices architecture. It provides a holistic view of a request as it travels through all the different services and components of a system.

Think of it like tracking a package delivery. When you order something online, you get a unique tracking number. You can use this number to see every step of the package’s journey—from the warehouse, to the shipping hub, to the delivery truck, and finally to your doorstep. A distributed trace works in exactly the same way for a software request.

To make this happen, distributed tracing relies on three core concepts:

Traces, Spans, and Context Propagation

  1. Trace: A trace represents the entire end-to-end journey of a single request. It is a collection of all the operations and steps that occurred to fulfill that request, from the initial user click in the browser to the final database write. Every trace is identified by a unique Trace ID.

  2. Span: A span represents a single, named, and timed operation within a trace. Think of it as a single “stop” on the package’s journey. A span could be an HTTP call to another microservice, a database query, or a specific business logic function. Each span has its own unique Span ID, a start time, a duration, and other relevant metadata (like HTTP status codes or error messages). Spans are organized in a hierarchy, with an initial “parent span” and subsequent “child spans.”

  3. Context Propagation: This is the magic that ties everything together. When a service makes a call to another service, it injects the trace context (including the Trace ID and the parent Span ID) into the request headers. The receiving service extracts this context and uses it to create a new child span linked to the original trace. This propagation ensures that all operations related to the initial request are correctly correlated, even across process and network boundaries.

How Does Distributed Tracing Work in Practice?

Let’s walk through a simplified example of booking a movie ticket to see how distributed tracing works:

  1. Request Initiated: A user clicks the “Confirm Booking” button on your website. This action sends a request to your API Gateway.

  2. The First Span: The API Gateway is instrumented for tracing. It receives the request and, seeing no existing trace context, generates a new unique Trace ID. It creates the first or “parent” span, which we’ll call POST /bookings.

  3. First Hop and Context Propagation: The API Gateway needs to validate the user’s session. It makes an API call to the Auth Service. Before sending the request, it injects the Trace ID and the ID of its own span into the HTTP headers.

  4. The Child Span: The Auth Service receives the call. It extracts the trace context from the headers and understands it’s part of an existing trace. It creates a new “child” span, perhaps named validate-session. This span is nested under the API Gateway’s span.

  5. Continuing the Journey: After validating the session, the API Gateway calls the Booking Service (propagating the context again). The Booking Service creates its own span and then calls the Payment Service and Database Service, each of which creates its own child spans.

  6. Trace Assembly: As each of these operations completes, every service sends its span data to a central backend collector. This backend system gathers all the spans that share the same Trace ID.

  7. Visualization: The tracing tool assembles these spans into a complete, ordered trace. It’s often visualized as a timeline or flame graph. You can now see the entire request flow in one view: the sequence of calls, the duration of each operation, and the dependencies between services. If the Database Service took 2 seconds to respond, it would be immediately visible as a long bar on the graph, clearly identifying it as the source of latency.

The Key Benefits of Implementing Distributed Tracing

Adopting a distributed tracing system brings profound benefits to development and operations teams.

Drastically Reduce MTTD and MTTR

With a trace, you can instantly see the exact path of a failed or slow request. There’s no more need to manually sift through logs from ten different services. This dramatically reduces the Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR) issues, getting your services back to a healthy state faster.

Understand Service Dependencies

Distributed traces provide a real-world map of how your services actually interact. You can discover hidden dependencies, identify critical paths, and understand the performance impact that one service has on others. This is invaluable for architecture planning and optimization.

Improve Developer Collaboration

When an error occurs, the trace pinpoints exactly which service—and therefore which team—is responsible. This eliminates finger-pointing and “war room” scenarios. Teams can collaborate effectively because they are all looking at the same objective data.

Enhance the End-User Experience

By proactively identifying and resolving performance bottlenecks and errors, you directly improve the user experience. Distributed tracing helps you meet your Service Level Agreements (SLAs) and keep your users happy by ensuring your application is fast and reliable.

Distributed Tracing vs. Logging: What’s the Difference?

A common point of confusion is how distributed tracing relates to logging. They are not mutually exclusive; they are complementary tools that solve different problems.

  • Distributed Logging or centralized logging is the practice of collecting timestamped event records from individual components. A log entry tells you what happened at a specific point in time within a single service (e.g., “User login failed: invalid password”). It provides deep, granular detail about an isolated event.

  • Distributed Tracing connects events across multiple services for a single request. A trace tells you why an overall process failed by showing the full story and causal relationships. (e.g., “The checkout process failed because the Payment Service timed out while waiting for a response from a slow, third-party Fraud-Detection API”).

The best observability platforms allow you to seamlessly pivot between them. You use a trace to identify the slow or failing service, then click to view the detailed logs for that specific span to get the rich, contextual information needed to debug the problem.

Getting Started: Standards and Tools

The distributed tracing ecosystem has matured significantly, largely thanks to open standards.

The Rise of OpenTelemetry

To avoid being locked into a single vendor’s proprietary solution, the community developed standards. The two early projects, OpenTracing and OpenCensus, merged to create OpenTelemetry (OTel). Now a CNCF project, OpenTelemetry is the industry standard for generating and collecting telemetry data (traces, metrics, and logs). It provides a single set of APIs, SDKs, and tools to instrument your applications, regardless of which backend you use for analysis.

Instrumentation: Manual vs. Automatic

To generate traces, your application code needs to be “instrumented.”

  • Manual Instrumentation: Developers use OpenTelemetry SDKs in their code to explicitly start and end spans, add attributes, and record events. This offers maximum control and customization but requires more development effort.
  • Automatic Instrumentation: This is the easiest way to get started. OTel provides libraries for popular languages and frameworks that can automatically create spans for common operations like incoming HTTP requests, outgoing client calls, and database queries—often with no code changes required.

Embracing Full-System Observability

In our increasingly complex, distributed world, you can no longer afford to fly blind. Distributed application tracing is not just a tool for debugging; it’s a fundamental requirement for understanding system behavior, ensuring reliability, and delivering a high-quality user experience. It provides the narrative context that isolated metrics and logs lack, turning a chaotic sea of data into a coherent story.

By embracing standards like OpenTelemetry and integrating tracing into your observability stack, you empower your teams to build, ship, and run resilient software with confidence. To truly master your complex systems, you need a solution that brings together metrics, logs, and traces in one place.

Get started with Netdata for free today and take the first step towards achieving true end-to-end observability for your entire stack.