Langfuse vs. Datadog Agent Observability
This guide outlines the key differences between Langfuse and Datadog Agent Observability. All Datadog facts were checked against public Datadog sources in September 2026.
TL;DR
- Open source, yours to run. Langfuse is one MIT-licensed product, identical on Cloud, self-hosted, and air-gapped. Datadog Agent Observability is SaaS only: no self-hosting, not even on Datadog's own government sites.
- An open data plane. Every trace lands in ClickHouse, open source and in production for a decade, the same engine on Cloud and self-hosted; self-hosters query it in SQL, and Cloud exports to your own bucket. Datadog spans live in a closed, Datadog-operated store you reach through its UI and a cursor-paged export API.
- 3 years of history instead of 15 days, for less. Datadog keeps traces and spans for 15 days on every plan; 30 to 90 days is a paid per-span add-on arranged through your account team, and nothing beyond 90 days is published. At 500k traces/month, Langfuse Pro is $621 with 3 years of history where Datadog Pro is $875 with the 90-day add-on.
- The whole improvement loop in the UI, for Python and JS/TS. Code evaluators run inside Langfuse on live observations, experiments start from the UI, and CI gates block regressions; prompt management serves both SDKs. Datadog runs experiments and prompt retrieval from the Python SDK only, and online code evaluations are computed in your code and submitted through its API.
- Many teams run both. Datadog for APM and infrastructure, Langfuse for AI engineering, from the same OpenTelemetry instrumentation.
Choose Langfuse if you want a platform built for AI engineering rather than an add-on to an APM suite: one MIT-licensed product on Cloud, self-hosted, and air-gapped, proven at 21 of the Fortune 50, with code evaluators, experiments, and prompt versions in one UI, and 3 years of raw traces in ClickHouse you can query, export, and turn into datasets. On the worked example that costs less than 90 days of history on Datadog. You accept that correlation with APM, infrastructure, logs, and RUM stays in your ops tool.
Choose Datadog Agent Observability if you already run Datadog and AI diagnosis must correlate with APM, infrastructure, logs, and RUM in one account, or procurement wants one vendor. You accept SaaS-only hosting, 15-day trace retention with paid add-ons to 90 days, and experiments and prompt retrieval that are Python-only.
Open source and distribution
Langfuse is open source (MIT). Self-hosting is a first-class deployment mode and uses the same codebase as Langfuse Cloud. Datadog Agent Observability is proprietary SaaS. Traces are stored on Datadog-hosted sites. The product is unavailable on Datadog's government sites (US1-FED, US2-FED).
| Langfuse | Datadog Agent Observability | |
|---|---|---|
| License | MIT | Proprietary SaaS |
| GitHub stars | N/A (platform not on GitHub) | |
| Self-hosting | Every tier, including free | Not offered. Traces are stored on Datadog-hosted sites |
| Air-gapped operation | Supported | Not offered |
| Data residency | Cloud EU, US, Japan; HIPAA region; any region via self-hosting | Nine hosted sites (US, EU, Japan, Australia, UK, US-gov); Agent Observability is unavailable on the two gov sites |
| Published Cloud Enterprise | $2,499/month | Custom |
History and data plane
ClickHouse has been open source since June 15, 2016 (Apache 2.0) and is today the most popular open-source analytical database, with 2,000+ contributors. Langfuse Cloud and self-hosted run the same ClickHouse architecture, so you can move between them without locking into a closed store.
Langfuse works directly with ClickHouse's core team as part of ClickHouse. Every LLM call, tool execution, and agent step is written once to a wide, immutable observations table (v4 GA) built for high-throughput ingestion and fast analytical reads, which is how Pro keeps 3 years of raw traces queryable. High-volume reads use the Observations and Metrics APIs, scheduled exports land Parquet in your own bucket, and self-hosters query ClickHouse in SQL.
Datadog Agent Observability stores spans in Datadog's closed store. You reach them through the Trace Explorer, dashboards on derived metrics, and an Export API that pages up to 5,000 spans per request; there is no direct database access and no self-hosted deployment of the store.
Datadog Agent Observability retains traces and spans for 15 days on every plan, including evaluation scores attached to them. A paid add-on extends this to 30, 60, or 90 days and experiments to 6, 9, or 12 months; add-ons are not available on the free tier and are arranged through your account team rather than enabled in the UI. Nothing beyond 90 days is published. Aggregated ml_obs.* metrics are kept for 15 months, so long-term dashboards survive after the underlying spans expire, but the raw traces you would replay or turn into a dataset do not.
| Langfuse | Datadog Agent Observability | |
|---|---|---|
| Data store | ClickHouse, launched 2016 (OSS), Apache 2.0 | Datadog-operated proprietary store, SaaS only |
| Maturity | ~10 years open source; 2,000+ contributors; SQL hiring pool | Closed; internals not published |
| Cloud / self-host | Same ClickHouse engine on Cloud and self-hosted | No self-hosted option |
| Getting data out | SQL on ClickHouse when self-hosting; Observations and Metrics APIs; scheduled export to your bucket | Export API, cursor-paged, up to 5,000 spans per request |
| Default trace retention | 30 days Hobby, 90 days Core, 3 years Pro and Enterprise | 15 days on Free, Pro, M2M, and annual |
| Extended trace retention | Included in the plan window; project retention policies on Pro+ | Paid add-on to 30, 60, or 90 days via your account team; not on Free. Rates in Pricing |
| Experiment retention | Same window as traces | 15 days on Free and on-demand; 90 days on committed plans; add-ons to 6 / 9 / 12 months |
| Dataset retention | Same project data plane | Current version 3 years; previous versions 90 days, reset when used |
| Aggregated metrics | Computed from the retained traces, 3 years on Pro | ml_obs.* metrics kept 15 months after the spans expire |
Pricing
Run your own numbers. The Langfuse vs Datadog pricing model is an editable Google Sheet with every list price in this section. Set traces, observations per trace, LLM calls per trace, scores, and retention to see where the costs cross over.
Langfuse bills units: one unit is a trace, an observation, or a score. Datadog Agent Observability bills LLM inference spans (one call to an LLM provider). Tool, workflow, agent, embedding, and retrieval spans are free on Datadog's published pricing, and retention beyond 15 days is a second meter charged on every LLM span.
| Langfuse | Datadog Agent Observability | |
|---|---|---|
| Free tier | 50,000 units/month, 2 users, 30 days | 40,000 LLM spans/month, 15-day retention, full feature access |
| Paid entry | Core $29/month (100,000 units, unlimited users) | Pro $160/month billed annually ($200 month-to-month) for the first 100,000 LLM spans |
| Usage rate | $8/100k units after the included 100k, graduating to $6/100k at 50M+ | $3.50 per 10k additional LLM spans (annual); $4.20 M2M; $5.00 on-demand |
| Retention | 90 days Core, 3 years Pro and Enterprise | 15 days included; 30 / 60 / 90 days at $1.50 / $3.00 / $4.00 per 10k LLM spans |
| Seats | None on Core and above | None published for Agent Observability |
| Evaluations | Scores bill as units | Eval LLM calls bill as LLM spans; no separate eval product fee |
| Enterprise | $2,499/month | Custom |
Worked example
500,000 traces/month · 10 observations per trace · 2 LLM calls per trace · 1 score per trace · 5 users.
| Base | Usage | Retention | Total/month | |
|---|---|---|---|---|
| Langfuse Pro (3-year retention) | $199.00 | $422.00 | $0.00 | $621.00 |
| Datadog Pro (90-day retention) | $160.00 | $315.00 | $400.00 | $875.00 |
Public annual list prices:
- Langfuse Pro: 6,000,000 units (traces + observations + scores), 100k included → $199 base + $72 (100k–1M at $8/100k) + $350 (1M–6M at $7/100k) = $621, 3-year history included.
- Datadog Pro: 1,000,000 LLM spans → $160 base + $315 (900,000 spans beyond the included 100k, at $3.50/10k) + $400 (90-day retention add-on at $4.00 per 10k spans, applied to all 1,000,000) = $875. With the default 15-day retention and no add-on, Datadog is $475. Month-to-month Pro is $200 + $4.20/10k.
The crossover depends on your span mix and how much history you need: Datadog bills only LLM-provider calls, Langfuse bills every observation. Different volumes or fewer LLM calls per trace? The editable sheet reproduces this math from your inputs.
See Langfuse pricing for current rates, start free, or talk to us about Enterprise volumes.
Tracing and stack context
Both platforms capture hierarchical traces of LLM applications and agents, including token usage, latency, errors, and cost.
Datadog Agent Observability correlates an LLM span with APM, infrastructure, logs, and RUM in the same Datadog account. Those signals come from the rest of Datadog.
Langfuse records a trace per request, with typed observations for LLM calls, tools, and retrieval. Multi-turn chats group into sessions. Agent runs render as agent graphs. Cost is attributed on each generation.
| Feature | Langfuse | Datadog Agent Observability |
|---|---|---|
| SDKs | Python and JS/TS on OpenTelemetry; other languages via the OTLP endpoint | Python, Node.js, and Java; HTTP API for other languages |
| Auto-instrumentation | 100+ integrations, including OpenAI, LangChain, LlamaIndex, Vercel AI SDK, LiteLLM, CrewAI | Supported LLM providers and agent frameworks |
| OpenTelemetry | OTel-native SDKs + OTLP endpoint | OTLP intake for GenAI semantic conventions 1.37+ and OpenInference |
| Platform context | Sessions, users, environments, releases, agent graphs | Correlation with Datadog APM, infrastructure, logs, and RUM |
| Production insights | Dashboards, Pulse outliers, alerts, Assistant on all Cloud plans | Patterns topic clustering (sampled to 10k spans or 5k traces per run), Insights anomalies, Bits Evals in preview |
| Sensitive data | SDK and OTel masking before data leaves your process | SDK span processors before send; Sensitive Data Scanner after ingest |
Evaluation and experiments
Langfuse evaluations include LLM-as-a-judge evaluators, code evaluators on live observations and experiments, custom scores via SDK/API, annotation queues, and dataset experiments on Cloud and self-hosted. CI/CD gates fail a pipeline on experiment results.
Datadog provides LLM-as-a-judge evaluators from a template or your own prompt at span, trace, and session scope, annotation queues, end-user feedback, and external evaluations submitted through the Python or Node.js SDK or the API. Experiments are started from the Python SDK with function- or class-based evaluators; datasets and run comparison live in the UI. Both platforms run judges on your own connected model accounts. Eval LLM calls bill as LLM spans on Datadog and scores bill as units on Langfuse.
| Feature | Langfuse | Datadog Agent Observability |
|---|---|---|
| LLM-as-a-judge | Yes on observations and experiments | Templates or your own prompt; span, trace, or session scope |
| Session-level evaluation | Score the session via SDK or annotation; run a judge on the observation that carries the full conversation | Managed judge runs once, 30 minutes after the last span; later spans are excluded |
| Deterministic online code evals | Yes, Python or TypeScript run inside Langfuse | Run in your code, submit via SDK or API; function evaluators in experiments |
| Experiments | UI and SDK in Python and JS/TS, baseline comparison | Python SDK; datasets and run comparison in the UI |
| CI/CD gates | GitHub Action | Via SDK / custom |
| Annotation queues | UI + API | Yes |
| Where evals run | Cloud or self-hosted, including air-gapped | Datadog SaaS |
Prompt management
Edit the prompt in the UI. The running app fetches it. In Langfuse, pass the fetched prompt into the generation to link it, so each version gets cost, latency, and scores. Datadog's Prompt Management registry serves prompts to Python applications only; Prompt Tracking records which prompt a span used without serving it.
| Langfuse | Datadog Agent Observability | |
|---|---|---|
| Who fetches | Python and JS/TS, then cached | Python (ddtrace>=4.13), LLMObs.get_prompt() |
| How you aim a version | Labels (production, staging, custom) | DD_ENV, numeric version, Feature Flags targeting |
| One prompt includes another | Prompt references | Not supported; concatenate in application code |
| Where it runs | Cloud or self-host | Datadog SaaS |
Using Langfuse and Datadog together
Keep Datadog for APM and infrastructure and add Langfuse for AI engineering.
The Python SDK sets up OpenTelemetry when you initialize the client. The default filter exports Langfuse SDK spans, gen_ai.* attributes, and known LLM instrumentors.
from langfuse import Langfuse
langfuse = Langfuse()To keep Langfuse spans out of Datadog, pass an isolated TracerProvider. JS/TS still registers a LangfuseSpanProcessor on your NodeSDK. A collector can also fan out one export to both backends.
Which should you choose
Langfuse when you want a platform built for AI engineering rather than an add-on to an APM suite: MIT self-hosting on an open ClickHouse store, 3 years of raw history, and the AI engineering loop in the UI for Python and JS/TS. Datadog Agent Observability when you already run Datadog and AI diagnosis must correlate with APM, infrastructure, logs, and RUM in one account. Both when SRE keeps Datadog and the AI team uses Langfuse. See Using Langfuse and Datadog together.
Why teams move away from Datadog Agent Observability
- SaaS only. Datadog Agent Observability cannot be self-hosted and is not offered on Datadog's government sites. Langfuse runs the same application self-hosted, including air-gapped, and moves to Cloud later without a migration.
- 15 days of raw traces. Datadog keeps traces and spans for 15 days on every plan. Longer retention is a paid add-on to 90 days, arranged through your account team, with nothing published beyond that. Failures older than the window cannot be replayed or turned into dataset items. Langfuse Pro keeps 3 years on ClickHouse, and in the worked example still costs less than Datadog Pro with the 90-day add-on.
- The improve loop is Python-first. Datadog documents experiment runs and prompt retrieval for the Python SDK; online code evaluations are computed in your code and submitted through the SDK or API. Langfuse runs code evaluators inside the platform, experiments from the UI, and CI gates for Python and JS/TS, on Cloud and self-hosted.
Why teams choose Langfuse
- Canva built a multi-agent support system for a product with 250 million monthly active users. Help Assistant is Java via OpenTelemetry; Omni Agent is Python via the Langfuse SDK. LLM-as-a-judge evaluators score both systems across 15–20 metrics, and domain experts build evaluators without engineering help. They self-hosted first, then moved to Cloud.
- SumUp rolled AI support to 35+ markets and deflected almost 50% of conversations, cutting external BPO cost by 30%. They started self-hosted for the PoC, then moved to Langfuse Cloud.
- Merck runs about 80 of 300+ GenAI use cases on Langfuse with 200+ people building on it. They self-host for data sovereignty and provision via API.
Switching from Datadog Agent Observability
Move live instrumentation and durable assets. Trace history usually stays in Datadog. Most teams see first traces in Langfuse the same day.
- Instrumentation: initialize the Langfuse SDK or send OpenTelemetry to the OTLP endpoint.
- Datasets, prompts, and evaluators: import datasets via CSV, SDK, or API. Recreate prompts and judges in prompt management and evaluators.
- Trace history: stays in Datadog for its retention window.
Talk to us if you want help planning a migration.
Start free: Cloud or self-host
Langfuse Cloud Hobby includes 50k units/month, no credit card. Or self-host the MIT stack. Explore the example project first. Already on Datadog Agent Observability? See Switching from Datadog Agent Observability.
FAQ
Is Langfuse an alternative to Datadog Agent Observability?
Yes. Langfuse is an open-source alternative to Datadog's AI observability SKU. See Which should you choose.
Can I use Langfuse and Datadog at the same time?
Yes. Keep Datadog APM and add Langfuse for LLM traces and the AI quality loop. See Using Langfuse and Datadog together.
How does Datadog Agent Observability pricing compare to Langfuse?
Datadog Agent Observability bills LLM inference spans (Pro $160/month billed annually for 100k spans, September 2026) plus a per-span retention add-on beyond 15 days. Langfuse bills units from $29/month, retention included. On the rates in our editable public model, 500k traces / 10 observations / 2 LLM calls / 1 score per trace is $621 Langfuse Pro with 3-year history vs $875 Datadog Pro with the 90-day retention add-on, or $475 at Datadog's default 15 days. Details in the worked example.
Can I self-host Datadog Agent Observability?
No. Agent Observability is Datadog-hosted SaaS. Langfuse self-hosts on every tier.
Did joining ClickHouse change the Langfuse product?
No. Langfuse already ran on ClickHouse. The announcement states the MIT license, self-hosting, Cloud endpoints, and roadmap stay the same, with more capacity to ship. Cloud Core and Pro remain self-serve; OSS self-host has no sales motion. Details in History and data plane and clarifications.
How do I migrate from Datadog Agent Observability to Langfuse?
Point live traffic at Langfuse and recreate datasets, prompts, and evaluators. History usually stays in Datadog. See Switching from Datadog Agent Observability.
This comparison is out of date? Please raise a pull request with up-to-date information.