The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Buyer’s Guide - August 2026

Best Fluentd monitoring tools for buffer, retry, and queue health

Fluentd monitoring means watching the collector daemon itself, not the logs it ships. The signals that matter are retry_count, buffer_queue_length, and buffer_total_queued_size, because rising values mean the destination is falling behind and log loss is getting closer. This ranking grades tools on Fluentd health coverage, collection resolution, alerting, deployment fit, cost shape, and correlation with the rest of the stack.

Best Fluentd monitoring tools for buffer, retry, and queue health product interface

Why this list exists

Most pages about Fluentd monitoring quietly switch topics. They rank log management platforms that receive logs from Fluentd, such as Splunk, New Relic, Dynatrace, and Sematext, when the buyer actually asked a narrower question: how do I know the Fluentd daemon is healthy before chunks start dropping. Those are different purchases. A destination can be ingesting perfectly while one Fluentd output plugin is stuck in retry backoff and its buffer queue is climbing.

The three numbers that predict trouble are retry_count, buffer_queue_length, and buffer_total_queued_size. Retry count shows flush failures. Queue length and queued size show records that have not reached the destination. When all three rise together, Fluentd cannot flush fast enough and log loss becomes possible. Everything else is context that helps explain why.

Three dimensions decide the outcome:

  1. Resolution: per-second collection catches short buffer spikes and retry bursts that a 15-second Prometheus default scrape can miss.
  2. Time to detection: prebuilt Fluentd alerts and anomaly detection beat hand-written thresholds when the pipeline breaks at 03:00.
  3. Blast radius: the tool should run where Fluentd runs, including Kubernetes DaemonSets, VMs, bare metal, and edge, without forcing metrics to leave the network if that is a requirement.

We do not quote competitor list prices in this guide. List numbers are a poor predictor of the actual bill in this category because pricing grows with different levers: hosts, metric series, log volume, query count, or cloud events. We describe the pricing shape and link each vendor’s pricing page — and our own — so the model can be checked against your fleet. For operator background, keep our operator runbooks for Fluentd open while shortlisting.

Methodology

How we evaluated Fluentd monitoring

The shortlist was limited to tools that actually read Fluentd’s own health endpoints: the built-in monitor_agent input plugin on port 24220 serving /api/plugins.json, or the fluent-plugin-prometheus metrics endpoint on port 24231. Tools that only receive logs from Fluentd were excluded, because receiving logs does not tell you whether the collector is about to lose them.

The heaviest weights went to daemon health coverage and collection resolution. Coverage decides whether you can see flush failures and queue growth per plugin. Resolution decides whether you see the spike while it is still actionable. Alerting quality, deployment flexibility, cost predictability, ecosystem correlation, and setup effort fill out the score in that order.

Tester credit

Compiled by the Netdata team - Updated August 12, 2026

Scoring criteria

  • Fluentd daemon health coverage 25%
    Retry count, buffer queue length, queued size, and deeper flush or emit counters.
  • Collection resolution 20%
    Per-second collection beats a 15-second scrape default for short spikes.
  • Alerting and anomaly detection 15%
    Prebuilt Fluentd alerts and ML cut time-to-detection.
  • Deployment flexibility 15%
    Kubernetes DaemonSets, VMs, bare metal, edge; self-hosted versus SaaS-only.
  • Cost predictability 10%
    Flat per-node versus per-GB, per-series, or query-based growth.
  • Ecosystem and correlation 10%
    Correlate daemon health with shipped logs and infrastructure.
  • Ease of setup 5%
    Dedicated integration versus assembling plugins, scrapers, and rules.

Vendor 01 / 06 · #netdata

01

Netdata

Real-time infrastructure monitoring with a dedicated Fluentd collector, per-second collection, and built-in ML anomaly detection.

Netdata metrics dashboard showing real-time per-second charts, the general view used to monitor Fluentd daemon metrics such as buffer queue length and retry count

Best for

  • Teams that want per-second visibility into Fluentd buffer and retry behavior without building a Prometheus stack
  • SREs who want ML-assisted anomaly detection on log pipeline health
  • Fleets of any size that want predictable per-node pricing with no per-GB or per-series bills

Pricing

  • Per-node pricing: Netdata Cloud Business starts at $4.50/node/month on annual plans, and the per-node price decreases as node count grows
  • Free Community tier for small fleets up to 5 nodes
  • Paid plans include unlimited metrics, logs, users, and retention
  • Open source agent under AGPL, self-hosted or Cloud; the bill grows with the number of Netdata Agents, not data volume

Pros

  • Dedicated Fluentd collector in go.d.plugin with a default 1-second collection interval
  • Reads the monitor_agent API and tracks retry_count, buffer_queue_length, and buffer_total_queued_size per plugin
  • Built-in ML-based anomaly detection and AI-assisted root cause analysis on metric behavior
  • Supports local and remote Fluentd instances in one configuration, with HTTP auth and TLS options
  • Open source agent runs beside Fluentd on the same host, with a free Cloud tier for small fleets

Where teams pair it

  • The collector tracks 3 Fluentd metric groups; Datadog and Sysdig expose 12 to 14 metrics including emit_count, flush_time_count, slow_flush_count, rollback_count, and retry_wait for deep flush-performance analysis
  • No default alert rules ship with the Fluentd collector; every alert must be created manually, whereas some commercial platforms ship prebuilt Fluentd alerts
  • The monitor_agent endpoint must be enabled and configured manually; there is no auto-detection

Verdict

Netdata leads because it combines the fastest default resolution in this list with the three metrics that directly predict log loss, without making you assemble Prometheus, Grafana, and Alertmanager first. The ML anomaly detection is useful for buffer and retry behavior that does not respect neat thresholds. The honest caveat is depth: if you need emit_count, flush_time_count, slow_flush_count, rollback_count, and retry_wait today, Datadog and Sysdig expose more Fluentd counters. Setup also assumes monitor_agent is enabled deliberately, which is a small one-time task rather than a discovery story.

Vendor 02 / 06 · #prometheus-grafana

02

Prometheus + Grafana

The CNCF monitoring stack: Prometheus scrapes Fluentd’s Prometheus endpoint and Grafana visualizes and alerts on it.

Best for

  • Teams already standardized on the CNCF and Prometheus ecosystem
  • Kubernetes shops running Fluentd as a DaemonSet with the official Helm chart, which enables the Prometheus plugins by default
  • Buyers who want full control of their monitoring stack and do not want per-node SaaS fees

Pricing

  • Prometheus is open source and self-hosted; you run and operate it
  • Grafana OSS is open source and self-hosted; Grafana Cloud is usage-based across active metric series, data points per minute, per-GB log ingestion, and per-active-user fees
  • On the Cloud tier the bill grows with series count, scrape frequency, and log volume

Pros

  • The Fluentd project’s officially recommended monitoring approach; both Fluentd and Prometheus are CNCF projects
  • fluent-plugin-prometheus exposes the standard metric set including retry_count, retry_wait, emit_count, and buffer_total_bytes
  • PromQL supports rate(), max_over_time(), and custom alerting on buffer growth and retry behavior
  • Large ecosystem across Grafana dashboards, Alertmanager, Thanos, and VictoriaMetrics compatibility
  • Open source and self-hosted with no per-metric license cost, though operations are yours

Cons

  • Default 15-second scrape interval is far coarser than per-second collection, so short buffer spikes can be missed
  • Requires assembling and operating Prometheus, Grafana, and Alertmanager yourself
  • fluent-plugin-prometheus must be installed and configured inside Fluentd before anything is collected
  • No built-in ML or anomaly detection; alerting is threshold-based PromQL you write yourself

Verdict

This is the default path for a reason: Fluentd documents it, the metric coverage is broad, and PromQL can express rate(fluentd_output_status_retry_count[1m]) and max_over_time(fluentd_output_status_retry_wait[1m]) directly. It beats Netdata on raw metric breadth when fluent-plugin-prometheus is configured well. It loses on time to value and resolution: the 15-second default scrape, the in-Fluentd plugin requirement, and DIY alerting mean more moving parts before the first useful page. Choose it when Prometheus is already the house standard.

Vendor 03 / 06 · #datadog

03

Datadog

SaaS observability with a built-in Fluentd check that reads monitor_agent and correlates daemon health with shipped logs.

Best for

  • Organizations already using Datadog for infrastructure and APM
  • Teams that want one SaaS pane for Fluentd daemon metrics plus the logs Fluentd ships
  • Enterprises needing commercial support and a broad integration catalog

Pricing

  • Per-host pricing for infrastructure monitoring
  • Per-GB pricing for log ingestion and indexing, and per-metric pricing for custom metrics
  • Infrastructure, APM, and logs are billed as separate modules, so the bill grows with host count, log volume, and custom metric cardinality

Pros

  • Built-in Fluentd check in the Datadog Agent, with nothing extra to install on Fluentd servers
  • Collects 14 Fluentd metrics including emit_count, flush_time_count, slow_flush_count, rollback_count, and tracked_file_count
  • fluentd.is_ok service check reports whether Fluentd and its monitor_agent are running
  • Correlates Fluentd daemon health with the logs Fluentd ships and the rest of the stack
  • Official Fluentd output plugin for shipping logs to Datadog

Cons

  • SaaS-only; Fluentd metrics leave your network
  • Per-host plus per-GB log pricing makes the bill grow with both fleet size and log volume
  • Collection cadence follows the Datadog Agent check interval rather than per-second streaming

Verdict

Datadog has the deepest commercial Fluentd coverage here: 14 metrics, a service check, and native correlation with the log stream. That is genuinely broader than Netdata’s three Fluentd metric groups. The trade-offs are structural. It is SaaS-only, the cadence is not per-second, and the pricing model couples hosts, log volume, and custom metrics in a way that can surprise teams as Fluentd throughput grows. It is the easiest enterprise answer and one of the easiest bills to misjudge.

Vendor 04 / 06 · #sysdig

04

Sysdig

Kubernetes-first monitoring with an out-of-the-box Fluentd integration, prebuilt dashboard, and nine Fluentd alert rules.

Best for

  • Kubernetes-centric teams monitoring Fluentd as part of the logging DaemonSet
  • Organizations that want prebuilt Fluentd alerts without writing PromQL
  • Security-conscious shops already using Sysdig Secure

Pricing

  • Per-host or per-node licensing, quote-based, with no published list prices
  • The bill grows with the number of monitored hosts and cloud log events processed

Pros

  • Out-of-the-box Fluentd integration that scrapes fluent-plugin-prometheus endpoints with no exporter to install
  • Collects 12 Fluentd metrics including retry_wait, num_errors, slow_flush_count, and buffer_available_space_ratio
  • Ships 9 prebuilt Fluentd alerts covering high retry ratio, high retry wait, buffer queue length increasing, low buffer available space, and more
  • Prebuilt Fluentd dashboard
  • Works automatically with the official Fluentd Helm chart and the OpenShift Logging Operator

Cons

  • Kubernetes-centric; host-based Fluentd deployments need manual Prometheus scrape configuration
  • Quote-based per-host pricing requires a sales conversation
  • SaaS platform; metrics leave your network

Verdict

Sysdig is the strongest packaged Kubernetes experience in this list. Twelve metrics, a ready dashboard, and nine Fluentd alerts are exactly what teams mean when they say they do not want to hand-roll monitoring. It clearly beats Netdata on prebuilt Fluentd alerting. The limits are just as clear: it is SaaS, quote-based, and much less automatic outside Kubernetes, where host-based Fluentd needs manual scrape setup. For OpenShift or DaemonSet-heavy estates, it belongs on the shortlist.

Vendor 05 / 06 · #telegraf-influxdb

05

Telegraf + InfluxDB

Open source metric pipeline: Telegraf’s Fluentd input reads monitor_agent into InfluxDB for storage, querying, and alerting.

Best for

  • InfluxDB users who want Fluentd metrics in the same time series database as the rest of their stack
  • Teams comfortable assembling open source pipelines with Grafana for dashboards
  • Edge and IoT deployments where Telegraf is already the collection agent

Pricing

  • Telegraf is open source and self-hosted; you run and operate it
  • InfluxDB OSS is open source and self-hosted; InfluxDB Cloud is usage-based across data written, query count, storage, and data out
  • On the Cloud tier the bill grows with data volume, query frequency, and retention

Pros

  • Dedicated Fluentd input plugin reading /api/plugins.json, available since Telegraf v1.4.0
  • Collects 14 buffer and output fields including emit_count, flush_time_count, slow_flush_count, rollback_count, and buffer_available_buffer_space_ratios
  • Open source, self-hosted, and works with any InfluxDB or compatible output
  • Telegraf’s broader plugin ecosystem enables correlated monitoring of the hosts Fluentd runs on

Cons

  • No built-in dashboards or alerting; you build them in Grafana, Kapacitor, or another front end
  • The plugin_id tag is random after each Fluentd restart, creating high-cardinality series unless you add @id to every Fluentd plugin
  • Collection interval follows Telegraf’s global interval setting, not per-second streaming

Verdict

Telegraf plus InfluxDB is a serious option when InfluxDB is already the metrics home. The Fluentd input is broad, with 14 buffer and output fields that rival Datadog’s coverage. It beats Netdata on field depth and stays open source and self-hosted. The catch is assembly: dashboards, alerts, and cardinality hygiene are operator work, and the random plugin_id tag after restarts can punish anyone who skips @id discipline. Good pipeline, thin product.

Vendor 06 / 06 · #victoriametrics

06

VictoriaMetrics

Prometheus-compatible time series database that scrapes Fluentd’s Prometheus endpoint at scale with lower storage overhead.

Best for

  • High-scale Prometheus workloads where storage cost is a concern
  • Teams replacing Prometheus with a more storage-efficient compatible backend
  • Managed-cloud buyers who want Prometheus compatibility without operating Prometheus

Pricing

  • Open source and self-hosted; you run and operate it
  • Enterprise subscription adds support and advanced features
  • VictoriaMetrics Cloud is a managed SaaS with capacity-tier pricing

Pros

  • Prometheus-compatible scraping of fluent-plugin-prometheus endpoints with no configuration changes
  • More efficient storage than Prometheus for high-cardinality Fluentd metrics
  • Open source self-hosted option plus a managed Cloud option
  • vmagent handles service discovery and scraping at scale

Cons

  • No Fluentd-specific integration, dashboard, or alerts; you build everything from raw Prometheus metrics
  • Requires fluent-plugin-prometheus to be installed and configured in Fluentd first
  • No built-in alerting UI; pairs with Grafana and Alertmanager or vmalert

Verdict

VictoriaMetrics is not a Fluentd monitor so much as a better backend for Fluentd metrics at scale. If the problem is Prometheus storage blowup from high-cardinality Fluentd series, it earns its place. If the problem is knowing which Fluentd output is backing up right now, it ships nothing Fluentd-specific to help. It ranks last here because setup, dashboards, and alerting are all DIY, even though the storage engine underneath is strong.

Frequently Asked Questions