01 · Compliance
Compliance monitoring
Run evals that check AI outputs against loaded regulations and policies. Catch non-compliant responses across back-office and customer-facing flows before they ship.
Industries · Financial services
Observe and evaluate AI agents across your institution. Give engineering, platform, and risk teams visibility into production behavior, with deployment in Langfuse Cloud or your own infrastructure.
Standardize observability and evals across teams: Bring agents across frameworks, models, and gateways into a shared view of traces and evaluations.
Ship reliable agents with clear deployment quality gates: evaluate model and agent changes against datasets before release.
Retain execution history: keep traces of AI executions as evidence for supervisory and regulatory reviews.
Flag potential policy violations: with evals that score outputs against your policies and regulations.
Self-host or air-gap: deploy with no internet access, lock to internal users via VPN to cater for data sensitivity needs.


“Generative AI will only earn enterprise trust when we can see what's happening under the hood. Langfuse enables us to track every prompt, response, cost, and latency in real time, turning black-box models into auditable, optimizable assets.”
Walid Mehanna, Chief Data & AI Officer at Merck
200+
users, engineers, PMs and domain experts
80+
use cases built on Langfuse
1
central platform team
Customer story · Self-hosted


How a European neobroker runs self-hosted Langfuse in production.
Read story →Customer story · Spend management
How Ramp auto-improves agents on Langfuse.
Read story →Customer story · Banking
How DKB automates 20,000 customer conversations daily.
Coming soonNever in the inference path
Langfuse observes your model and tool calls; it does not proxy them. Docs →
Deployment where your data must stay
Langfuse Cloud in the EU, US or Japan, self-hosted in your VPC, or fully air-gapped. Docs →
Multiple layers of data redaction
Client-side and server-side PII protections before anything is stored. Docs →
Enterprise access controls
SSO and role-based access control scoped to organizations and projects. Docs →
Audit logs
Record both LLM activity and developer actions for internal and supervisory review. Docs →
Managed cloud in the EU, US or Japan.
Your app→OTel SDK→Langfuse Cloud · EU
Platform and risk teams in financial services are looking for a complete observability and eval suite. Langfuse covers the full surface, and we are open about where the work lives in your pipeline rather than in our UI.
Test before release. Monitor production. Review with your experts.
Golden datasets built from production traces, versioned experiments with baseline comparison, and a release gate that fails the pull request on regression.
Docs →Deterministic sampling of live traffic, LLM-as-judge and code evaluators, and threshold alerts to Slack, webhooks, or GitHub Actions when quality drifts.
Docs →Annotation queues for subject-matter experts, corrected outputs, and one-click promotion of failures into a permanent regression set.
Docs →Score Analytics measures agreement between human labels and model judges (Cohen's Kappa, F1, Pearson, Spearman) so you can defend the judge to model risk.
Docs →Control changes, protect sensitive data, and manage access.
Immutable versions, staging and production labels, protected labels for separation of duties, and full audit history.
Docs →Masking in the SDK before data leaves your application, with trace structure preserved for debugging.
Docs →Organizations and projects as the data boundary, project-level RBAC, OIDC SSO with domain enforcement, SCIM provisioning, audit logs, and a metrics API for chargebacks.
Docs →Connect your models and manage deployment in your infrastructure.
OTLP ingest from your AI gateway, judge models pinned to your own Bedrock, Azure OpenAI, Vertex, or OpenAI connection.
Docs →Dashboards and widgets managed through the API and CLI, versioned in Git, deployed identically to dev, staging, and prod.
Docs →Documented self-hosting on AWS with Terraform and Helm, a published release cadence, and autoscaling guidance.
Docs →Use cases
01 · Compliance
Run evals that check AI outputs against loaded regulations and policies. Catch non-compliant responses across back-office and customer-facing flows before they ship.
02 · Compliance
Support anomaly detection and investigation workflows — reduce cost of service and risk with better observability of agent/tool behavior.
03 · Investing
Observe and improve multi-step advisory flows, score outcomes, and keep an audit trail suitable for model risk review. Evaluate your system on representative cases before it runs in production.
04 · Risk
Trace agents that combine identity, fraud and credit-bureau checks into a risk summary. Score decisions, flag drift, and keep the auditable trail that credit and model risk teams require.
05 · SupportDKB · 20,000 conversations / day
Raise autonomy rate of AI support agents: the share of cases resolved without human intervention. Trace edge cases, score good/bad runs, and iterate so agents cover more of the long tail safely.
06 · Engineering
Trace Claude Code, Codex, Cursor, OpenCode, and GitHub Copilot — no proxy required. See cost per developer and model, replay failed sessions, and search across the org.
Start with deployment gates and execution history, then see how teams like Trade Republic run Langfuse in production.
Evaluate model and prompt changes against datasets and policy evals before release.
Trace, score and iterate on a support agent to raise its autonomy rate safely.
How a European neobroker runs self-hosted Langfuse in production. Read or watch on YouTube.
Yes. You can deploy Langfuse in your own cloud, VPC, or on-premises infrastructure. Langfuse supports deployments without public internet access (air-gapped); features such as LLM-as-a-judge evaluations need a model endpoint reachable within your environment. Our team can help you assess the setup for your deployment and security requirements. Self-hosting, networking documentation.
Talk through deployment options, compliance needs, and how teams like Trade Republic use Langfuse in production.
No credit card required · Free tier available · Self-hosting option