Catching conversation signals in Langfuse
Score every chat agent conversation for frustration, corrections, follow-ups, and resolution — in one Jev pass, at a fraction of LLM-as-a-judge cost.
Chat agents can handle millions of user interactions every day. A volume, impossible for a human to review manually. Jev-like models now open an opportunity to review production traces for signals of failure at scale. Something that can get costly with LLMs and is hard to solve with deterministic code-based evals.
About user signals
There are different types of signals in chat applications that can be used to monitor and improve AI applications.
- You can ask for explicit user feedback, like thumbs up or down
- You can track behavior, like dismissing a chat
- You can track factual outcomes of a chat, like a completed booking
- You can track conversation signals, like a user being upset or asking follow ups.
Especially conversation signals help you steer your attention to the relevant traces from production as soon as manual review is not feasible anymore. They tell you something is going wrong. See the table below for a typical list of conversation signals.
![]()
Established approaches to catch these signals were especially centered around LLM-as-a-judge. With scale, the caused costs were primarily contained through sampling. Jev-like models now allow a full pass on all production traces.
For more on the signal taxonomy (explicit, behavioral, conversation, and outcome), see Capturing signals in the Langfuse Academy. To set up Jev evaluators in Langfuse, see Jev as a judge.
Catching conversation signals with Jev
Jev-like models offer the advantage that the same state or context can be reused to answer multiple questions in parallel. This means the same conversation can concurrently be evaluated for follow-up questions, user frustration, a user disagreeing with the agent, and if the user seems to be happy with the outcome.
Let us look at an example conversation in a support resolution. The user wants to rebook a flight. Over the course of the interaction she has to first correct the AI, then ask a follow up question, but then finally gets the request resolved. Multiple things are true here. The conversation outcome was successful. Nevertheless, the AI made a mistake first, and was apparently not effective enough in their answers to not cause a follow-up by a user.
![]()
With Jev, these signals can be caught concurrently in one pass, at the fraction of the cost of using frontier LLM models. All dimensions can be modelled as individual noul questions, to get separate verdicts on the different conversation signals.
Using Jev in Langfuse
While evals should be custom to your application, catching conversation signals is universal. The Langfuse template gallery now offers a full Jev based batch evaluator template to catch production signals.
Go to the Evaluator tab
Go to the Evaluator tab and click on New evaluator.
Choose the User Conversation Signal template
Choose the ‘User Conversation Signal’ template.
![]()
Add your TypeSafe connection
Add your Jev connection if you have not already.
Choose the observations to run on and test
Choose the observations to run on and test.
![]()
Save and set rule for evaluators to run on
Save and set rule for evaluators to run on.
The scores will then appear on incoming traces. You can also opt to do a pass on historic traces. To visualize the scores you can create dashboards, and create saved views in your tracing table.
Ready to get started with Langfuse?
Join thousands of teams building better LLM applications with Langfuse's open-source observability platform.
No credit card required · Free tier available · Self-hosting option