Langfuse v4: up to 165ร— faster ยท Read more
September 30, 2026

Langfuse September Update

Jev-as-a-judge, Langfuse Assistant sandbox, rebuilt experiment results, evaluator backfills and more

Picture Marc KlingenMarc Klingen

September was about evaluations at scale: Jev-as-a-judge to score every observation for a fraction of a cent, backfills for data you already have, and an Assistant that runs code over thousands of observations.

Jev-as-a-judge

Jev-as-a-judge

You can now use TypeSafe's Jev as a judge in Langfuse evaluators. Jev does not generate text. You give it a state and typed questions, and it returns typed answers with probabilities. Each question becomes a score on the observation.

Jev is priced by input tokens only, at $0.042 per million, and answers come back in well under a second. That means you can score every observation instead of a sample, and alerts fire fast enough to matter. Jev does not replace LLM judges: when you need written reasoning, keep using LLM-as-a-Judge, and run both if it helps. New to decision models? Annabell and Hassieb explain how Jev works in Using TypeSafe's Jev for evals.

To get started quickly, use the ready-made User Conversation Signal template. In one pass, it scores a chat for rephrases, corrections, human hand-off requests, repeats, error quotes, frustration, and success confirmation. The walkthrough shows how to set it up.

โ†’ Read more

Langfuse Assistant: analyze thousands of observations

Langfuse Assistant

The Langfuse Assistant now works through thousands of observations at once. It fetches them and runs code in a sandbox instead of loading every payload into the model context. Ask it to cluster the main error types of the last seven days, count how often each tool was called, or find what the slowest 100 traces have in common. It also builds datasets and dashboards on request.

Every run happens in the background. Keep exploring Langfuse or start a second conversation while the first one is still working. Anything that changes data or configuration still waits for your approval.

โ†’ Read more

Rebuilt experiment results

Rebuilt experiment results

We rebuilt the experiment results view around scores. The results grid leads with score columns, the chart grid is now one compact metric strip, and compared runs are color-coded.

A score-by-run matrix and a worse-than-baseline filter get you from an aggregate to the exact items that regressed and their traces. The comparison run is pre-picked and grouped, and side-by-side comparison of two runs is streamlined.

โ†’ See how it works

Evaluator backfills

Evaluator backfills

When you attach an evaluator to a rule, you can also run it on past observations to score observations that arrived before the rule existed. Pick a window from the last 24 hours up to six months, set a maximum item count, and sampling rate. Langfuse shows the matching count and the estimated LLM cost before the run starts. For a one-off run over hand-picked traces, batch evaluation from the observations table is still the right tool.

โ†’ Read more

Langfuse walkthrough video

Langfuse walkthrough video

New Introduction to Langfuse: I demo observability, prompt management, and evaluation end-to-end in one video.

โ†’ Watch demo

Reminder: Langfuse v4 cutover

Langfuse Cloud becomes v4-only on November 16, 2026. Check the Migration Assistant in the sidebar. If it lists required actions for your project, finish them before the cutover. Self-hosted deployments set their own timeline.

โ†’ Upgrade guide

Fixes & improvements

  • Feat: Redesigned session timeline with observation actions and public links
  • Feat: Filter and search bar on the Scores and Users tables (v4.35.0)
  • Filters: has: and -has: operators to filter on metadata key presence (v4.38.0)
  • UI: Persistent trace header with tags; graph is now a view on the tree switch (v4.38.0)
  • UI: Cached input tokens and cost as traces table columns, plus cost source and waterfall in the breakdown tooltip (v4.36.0)
  • UI: 100-row page size on tracing tables (v4.46.0)
  • Evals: Create scores from the annotation sidebar (v4.43.0)
  • Evals: Override the variable mapping in batch evaluations (v4.28.1)
  • Prompts: Automations can filter prompt events by label and by tags (docs)
  • API: providedModelName is now model in observation APIs; prompt name prefix filtering; optional startTime on GET /observations/{id} (v4.42.0)
  • Security: Revoke your own active sessions; filter project members by role (v4.43.0)
  • MCP: Evaluator testing tool and batchUpsertDatasetItems (v4.37.0)
  • Self-hosting: Vertex AI for the Assistant and other instance AI features; custom ClickHouse cluster migrations (docs)
  • SDK: Python SDK captures reasoning_effort and verbosity for OpenAI calls and masks pydantic secret values (v4.15.6); default exporter batches are capped at 64 MiB for more reliable exports (v4.16.0)
  • and many more!

Customer story: Ramp

Customer story: Ramp

Ramp's engineering team traces its coding agents with self-hosted Langfuse. Their Inspect agent accounts for about 70% of merged PRs at Ramp, across roughly 1,500 PRs a day. Reflect, their agent that monitors other agents, scores those traces, groups similar ones, and proposes fixes for a human to accept or reject: 10 to 20% fewer tokens, 15% fewer tool calls, and 30% faster sessions.

"We wanted something built for agents as users first. And that means API first. An agent should never get stuck waiting for a human because of a deficiency in the API. This is what Langfuse is." โ€“ David Traina, Data Platform at Ramp

โ†’ Watch the story


Was this page helpful?