Langfuse v4: up to 165× faster · Read more

Langfuse vs. Datadog Agent Observability

This guide outlines the key differences between Langfuse and Datadog Agent Observability. All Datadog facts were checked against public Datadog sources in September 2026.

TL;DR

Choose Langfuse if you want a platform built for AI engineering rather than an add-on to an APM suite: one MIT-licensed product on Cloud, self-hosted, and air-gapped, proven at 21 of the Fortune 50, with code evaluators, experiments, and prompt versions in one UI, and 3 years of raw traces in ClickHouse you can query, export, and turn into datasets. On the worked example that costs less than 90 days of history on Datadog. You accept that correlation with APM, infrastructure, logs, and RUM stays in your ops tool.

Choose Datadog Agent Observability if you already run Datadog and AI diagnosis must correlate with APM, infrastructure, logs, and RUM in one account, or procurement wants one vendor. You accept SaaS-only hosting, 15-day trace retention with paid add-ons to 90 days, and experiments and prompt retrieval that are Python-only.


Open source and distribution

Langfuse is open source (MIT). Self-hosting is a first-class deployment mode and uses the same codebase as Langfuse Cloud. Datadog Agent Observability is proprietary SaaS. Traces are stored on Datadog-hosted sites. The product is unavailable on Datadog's government sites (US1-FED, US2-FED).

LangfuseDatadog Agent Observability
LicenseMITProprietary SaaS
GitHub starsLangfuse GitHub starsN/A (platform not on GitHub)
Self-hostingEvery tier, including freeNot offered. Traces are stored on Datadog-hosted sites
Air-gapped operationSupportedNot offered
Data residencyCloud EU, US, Japan; HIPAA region; any region via self-hostingNine hosted sites (US, EU, Japan, Australia, UK, US-gov); Agent Observability is unavailable on the two gov sites
Published Cloud Enterprise$2,499/monthCustom

History and data plane

ClickHouse has been open source since June 15, 2016 (Apache 2.0) and is today the most popular open-source analytical database, with 2,000+ contributors. Langfuse Cloud and self-hosted run the same ClickHouse architecture, so you can move between them without locking into a closed store.

Langfuse works directly with ClickHouse's core team as part of ClickHouse. Every LLM call, tool execution, and agent step is written once to a wide, immutable observations table (v4 GA) built for high-throughput ingestion and fast analytical reads, which is how Pro keeps 3 years of raw traces queryable. High-volume reads use the Observations and Metrics APIs, scheduled exports land Parquet in your own bucket, and self-hosters query ClickHouse in SQL.

Datadog Agent Observability stores spans in Datadog's closed store. You reach them through the Trace Explorer, dashboards on derived metrics, and an Export API that pages up to 5,000 spans per request; there is no direct database access and no self-hosted deployment of the store.

Datadog Agent Observability retains traces and spans for 15 days on every plan, including evaluation scores attached to them. A paid add-on extends this to 30, 60, or 90 days and experiments to 6, 9, or 12 months; add-ons are not available on the free tier and are arranged through your account team rather than enabled in the UI. Nothing beyond 90 days is published. Aggregated ml_obs.* metrics are kept for 15 months, so long-term dashboards survive after the underlying spans expire, but the raw traces you would replay or turn into a dataset do not.

LangfuseDatadog Agent Observability
Data storeClickHouse, launched 2016 (OSS), Apache 2.0Datadog-operated proprietary store, SaaS only
Maturity~10 years open source; 2,000+ contributors; SQL hiring poolClosed; internals not published
Cloud / self-hostSame ClickHouse engine on Cloud and self-hostedNo self-hosted option
Getting data outSQL on ClickHouse when self-hosting; Observations and Metrics APIs; scheduled export to your bucketExport API, cursor-paged, up to 5,000 spans per request
Default trace retention30 days Hobby, 90 days Core, 3 years Pro and Enterprise15 days on Free, Pro, M2M, and annual
Extended trace retentionIncluded in the plan window; project retention policies on Pro+Paid add-on to 30, 60, or 90 days via your account team; not on Free. Rates in Pricing
Experiment retentionSame window as traces15 days on Free and on-demand; 90 days on committed plans; add-ons to 6 / 9 / 12 months
Dataset retentionSame project data planeCurrent version 3 years; previous versions 90 days, reset when used
Aggregated metricsComputed from the retained traces, 3 years on Proml_obs.* metrics kept 15 months after the spans expire

Pricing

Run your own numbers. The Langfuse vs Datadog pricing model is an editable Google Sheet with every list price in this section. Set traces, observations per trace, LLM calls per trace, scores, and retention to see where the costs cross over.

Langfuse bills units: one unit is a trace, an observation, or a score. Datadog Agent Observability bills LLM inference spans (one call to an LLM provider). Tool, workflow, agent, embedding, and retrieval spans are free on Datadog's published pricing, and retention beyond 15 days is a second meter charged on every LLM span.

LangfuseDatadog Agent Observability
Free tier50,000 units/month, 2 users, 30 days40,000 LLM spans/month, 15-day retention, full feature access
Paid entryCore $29/month (100,000 units, unlimited users)Pro $160/month billed annually ($200 month-to-month) for the first 100,000 LLM spans
Usage rate$8/100k units after the included 100k, graduating to $6/100k at 50M+$3.50 per 10k additional LLM spans (annual); $4.20 M2M; $5.00 on-demand
Retention90 days Core, 3 years Pro and Enterprise15 days included; 30 / 60 / 90 days at $1.50 / $3.00 / $4.00 per 10k LLM spans
SeatsNone on Core and aboveNone published for Agent Observability
EvaluationsScores bill as unitsEval LLM calls bill as LLM spans; no separate eval product fee
Enterprise$2,499/monthCustom

Worked example

500,000 traces/month · 10 observations per trace · 2 LLM calls per trace · 1 score per trace · 5 users.

BaseUsageRetentionTotal/month
Langfuse Pro (3-year retention)$199.00$422.00$0.00$621.00
Datadog Pro (90-day retention)$160.00$315.00$400.00$875.00

Public annual list prices:

  • Langfuse Pro: 6,000,000 units (traces + observations + scores), 100k included → $199 base + $72 (100k–1M at $8/100k) + $350 (1M–6M at $7/100k) = $621, 3-year history included.
  • Datadog Pro: 1,000,000 LLM spans → $160 base + $315 (900,000 spans beyond the included 100k, at $3.50/10k) + $400 (90-day retention add-on at $4.00 per 10k spans, applied to all 1,000,000) = $875. With the default 15-day retention and no add-on, Datadog is $475. Month-to-month Pro is $200 + $4.20/10k.

The crossover depends on your span mix and how much history you need: Datadog bills only LLM-provider calls, Langfuse bills every observation. Different volumes or fewer LLM calls per trace? The editable sheet reproduces this math from your inputs.

See Langfuse pricing for current rates, start free, or talk to us about Enterprise volumes.


Tracing and stack context

Both platforms capture hierarchical traces of LLM applications and agents, including token usage, latency, errors, and cost.

Datadog Agent Observability correlates an LLM span with APM, infrastructure, logs, and RUM in the same Datadog account. Those signals come from the rest of Datadog.

Langfuse records a trace per request, with typed observations for LLM calls, tools, and retrieval. Multi-turn chats group into sessions. Agent runs render as agent graphs. Cost is attributed on each generation.

FeatureLangfuseDatadog Agent Observability
SDKsPython and JS/TS on OpenTelemetry; other languages via the OTLP endpointPython, Node.js, and Java; HTTP API for other languages
Auto-instrumentation100+ integrations, including OpenAI, LangChain, LlamaIndex, Vercel AI SDK, LiteLLM, CrewAISupported LLM providers and agent frameworks
OpenTelemetryOTel-native SDKs + OTLP endpointOTLP intake for GenAI semantic conventions 1.37+ and OpenInference
Platform contextSessions, users, environments, releases, agent graphsCorrelation with Datadog APM, infrastructure, logs, and RUM
Production insightsDashboards, Pulse outliers, alerts, Assistant on all Cloud plansPatterns topic clustering (sampled to 10k spans or 5k traces per run), Insights anomalies, Bits Evals in preview
Sensitive dataSDK and OTel masking before data leaves your processSDK span processors before send; Sensitive Data Scanner after ingest

Evaluation and experiments

Langfuse evaluations include LLM-as-a-judge evaluators, code evaluators on live observations and experiments, custom scores via SDK/API, annotation queues, and dataset experiments on Cloud and self-hosted. CI/CD gates fail a pipeline on experiment results.

Datadog provides LLM-as-a-judge evaluators from a template or your own prompt at span, trace, and session scope, annotation queues, end-user feedback, and external evaluations submitted through the Python or Node.js SDK or the API. Experiments are started from the Python SDK with function- or class-based evaluators; datasets and run comparison live in the UI. Both platforms run judges on your own connected model accounts. Eval LLM calls bill as LLM spans on Datadog and scores bill as units on Langfuse.

FeatureLangfuseDatadog Agent Observability
LLM-as-a-judgeYes on observations and experimentsTemplates or your own prompt; span, trace, or session scope
Session-level evaluationScore the session via SDK or annotation; run a judge on the observation that carries the full conversationManaged judge runs once, 30 minutes after the last span; later spans are excluded
Deterministic online code evalsYes, Python or TypeScript run inside LangfuseRun in your code, submit via SDK or API; function evaluators in experiments
ExperimentsUI and SDK in Python and JS/TS, baseline comparisonPython SDK; datasets and run comparison in the UI
CI/CD gatesGitHub ActionVia SDK / custom
Annotation queuesUI + APIYes
Where evals runCloud or self-hosted, including air-gappedDatadog SaaS

Prompt management

Edit the prompt in the UI. The running app fetches it. In Langfuse, pass the fetched prompt into the generation to link it, so each version gets cost, latency, and scores. Datadog's Prompt Management registry serves prompts to Python applications only; Prompt Tracking records which prompt a span used without serving it.

LangfuseDatadog Agent Observability
Who fetchesPython and JS/TS, then cachedPython (ddtrace>=4.13), LLMObs.get_prompt()
How you aim a versionLabels (production, staging, custom)DD_ENV, numeric version, Feature Flags targeting
One prompt includes anotherPrompt referencesNot supported; concatenate in application code
Where it runsCloud or self-hostDatadog SaaS

Using Langfuse and Datadog together

Keep Datadog for APM and infrastructure and add Langfuse for AI engineering.

The Python SDK sets up OpenTelemetry when you initialize the client. The default filter exports Langfuse SDK spans, gen_ai.* attributes, and known LLM instrumentors.

from langfuse import Langfuse

langfuse = Langfuse()

To keep Langfuse spans out of Datadog, pass an isolated TracerProvider. JS/TS still registers a LangfuseSpanProcessor on your NodeSDK. A collector can also fan out one export to both backends.


Which should you choose

Langfuse when you want a platform built for AI engineering rather than an add-on to an APM suite: MIT self-hosting on an open ClickHouse store, 3 years of raw history, and the AI engineering loop in the UI for Python and JS/TS. Datadog Agent Observability when you already run Datadog and AI diagnosis must correlate with APM, infrastructure, logs, and RUM in one account. Both when SRE keeps Datadog and the AI team uses Langfuse. See Using Langfuse and Datadog together.


Why teams move away from Datadog Agent Observability

  • SaaS only. Datadog Agent Observability cannot be self-hosted and is not offered on Datadog's government sites. Langfuse runs the same application self-hosted, including air-gapped, and moves to Cloud later without a migration.
  • 15 days of raw traces. Datadog keeps traces and spans for 15 days on every plan. Longer retention is a paid add-on to 90 days, arranged through your account team, with nothing published beyond that. Failures older than the window cannot be replayed or turned into dataset items. Langfuse Pro keeps 3 years on ClickHouse, and in the worked example still costs less than Datadog Pro with the 90-day add-on.
  • The improve loop is Python-first. Datadog documents experiment runs and prompt retrieval for the Python SDK; online code evaluations are computed in your code and submitted through the SDK or API. Langfuse runs code evaluators inside the platform, experiments from the UI, and CI gates for Python and JS/TS, on Cloud and self-hosted.

Why teams choose Langfuse

  • Canva built a multi-agent support system for a product with 250 million monthly active users. Help Assistant is Java via OpenTelemetry; Omni Agent is Python via the Langfuse SDK. LLM-as-a-judge evaluators score both systems across 15–20 metrics, and domain experts build evaluators without engineering help. They self-hosted first, then moved to Cloud.
  • SumUp rolled AI support to 35+ markets and deflected almost 50% of conversations, cutting external BPO cost by 30%. They started self-hosted for the PoC, then moved to Langfuse Cloud.
  • Merck runs about 80 of 300+ GenAI use cases on Langfuse with 200+ people building on it. They self-host for data sovereignty and provision via API.

Switching from Datadog Agent Observability

Move live instrumentation and durable assets. Trace history usually stays in Datadog. Most teams see first traces in Langfuse the same day.

Talk to us if you want help planning a migration.


Start free: Cloud or self-host

Langfuse Cloud Hobby includes 50k units/month, no credit card. Or self-host the MIT stack. Explore the example project first. Already on Datadog Agent Observability? See Switching from Datadog Agent Observability.


FAQ

Is Langfuse an alternative to Datadog Agent Observability?

Yes. Langfuse is an open-source alternative to Datadog's AI observability SKU. See Which should you choose.

Can I use Langfuse and Datadog at the same time?

Yes. Keep Datadog APM and add Langfuse for LLM traces and the AI quality loop. See Using Langfuse and Datadog together.

How does Datadog Agent Observability pricing compare to Langfuse?

Datadog Agent Observability bills LLM inference spans (Pro $160/month billed annually for 100k spans, September 2026) plus a per-span retention add-on beyond 15 days. Langfuse bills units from $29/month, retention included. On the rates in our editable public model, 500k traces / 10 observations / 2 LLM calls / 1 score per trace is $621 Langfuse Pro with 3-year history vs $875 Datadog Pro with the 90-day retention add-on, or $475 at Datadog's default 15 days. Details in the worked example.

Can I self-host Datadog Agent Observability?

No. Agent Observability is Datadog-hosted SaaS. Langfuse self-hosts on every tier.

Did joining ClickHouse change the Langfuse product?

No. Langfuse already ran on ClickHouse. The announcement states the MIT license, self-hosting, Cloud endpoints, and roadmap stay the same, with more capacity to ship. Cloud Core and Pro remain self-serve; OSS self-host has no sales motion. Details in History and data plane and clarifications.

How do I migrate from Datadog Agent Observability to Langfuse?

Point live traffic at Langfuse and recreate datasets, prompts, and evaluators. History usually stays in Datadog. See Switching from Datadog Agent Observability.

This comparison is out of date? Please raise a pull request with up-to-date information.


Was this page helpful?