Langfuse Blog
The latest updates from Langfuse. See Changelog for more product updates.
Highlights
engineering
Yes, you can copy our eval setup
Evals need to be custom to your app. At least some of them. A good chunk you can, and should be copying. Here is our setup for our docs chatbot.
Jul 22, 2026
·Annabell
engineering
Designing the runtime for Langfuse code evaluators
Code evaluators let you score traces with your own Python or TypeScript code. A look at the execution model behind them: the requirements, the options we rejected, and the security stance we adopted.
Jun 22, 2026
·Tobias
engineering
AI is eating the AI engineering loop
The full AI engineering loop can technically be automated now. But that doesn't mean it should. Here is what we think you should hand to agents, and what you should keep doing yourself.
Jun 9, 2026
·Lotte
All Posts
update
Langfuse August Update
Langfuse v4, new evaluator setup flow, multi-modal evaluators, CLI 1.0, responsive timeline and more
Aug 31, 2026
·Marc
update
Langfuse July Update
Chart any table, dashboards via API, CLI, and MCP, new graph view modes, media previews and more
Jul 31, 2026
·Marc
guide
Building Deployment Gates for LLMs and AI Agents in Financial Services
How we used Langfuse datasets, experiments, and the public API to build automated pass/fail deployment gates for LLMs and AI agents in financial services.
Jul 15, 2026
·Doneyli
engineering
How we use agents to review production infrastructure
How repo-owned agent workflows help us review incidents, infra cost, security findings, and bugs in production.
Jun 5, 2026
·Max
update
Langfuse May Update
Code Evaluators, full-text search, Langfuse MCP, Experiments in CI/CD and more
May 31, 2026
·Marc
update
Langfuse April Update
Japan Cloud Region, new Experiments, Langfuse Academy, LLM-as-a-Judge API and more
Apr 30, 2026
·Marc
announcement
Langfuse Cloud 日本リージョンを開始しました
Langfuse Cloud 日本リージョンを公開しました。LLM のトレースや評価データを日本国内に保管したいチーム向けの専用クラウドリージョンです。
Apr 27, 2026
·Marc, Clemens, Max
announcement
Langfuse Cloud Japan Region
Langfuse Cloud Japan is live. A dedicated cloud region hosted in Japan for teams that need their LLM observability data to stay in Japan.
Apr 27, 2026
·Marc, Clemens, Max
engineering
Classifying User Intent with Categorical LLM-as-a-Judge
A guide on how to set up a categorical LLM-as-a-judge evaluator to classify user intent. Follow along with how we applied this to our demo application.
Apr 14, 2026
·Lotte
engineering
The Rage Clicks of LLM apps: High-Signal Production Monitoring for AI Customer Support Agents
How to use LLM-as-a-judge to detect when your users say "f**k".
Apr 1, 2026
·Annabell
update
Langfuse March Update
Agent Skill, Langfuse CLI, boolean and categorical LLM-as-a-Judge scores, Kiro integration, and more
Mar 31, 2026
·Marc
engineering
We Used Autoresearch on Our AI Skill, It Taught Us to Write Better Tests
We applied Karpathy's autoresearch to optimize our Langfuse prompt migration skill — and got a lesson in why the target function matters more than the optimizer.
Mar 24, 2026
·Lotte
engineering
How We Built an Agent Skill to Synthesize what Langfuse Users want
We built an AI agent skill that synthesizes GitHub issues, support tickets, and meeting notes into a weekly digest, and used Langfuse to monitor and improve it.
Mar 13, 2026
·Lotte
engineering, architecture
Simplifying Langfuse for Scale
A deep dive into how we moved Langfuse to an observations-first data model.
Mar 10, 2026
·Steffen, Valeriy, Max
update
Langfuse February Update
Observation-centric data model, faster UI, Observations API v2 and Metrics API v2 out of beta, faster evaluation workflows
Feb 28, 2026
·Marc