Langfuse v4: up to 165ร— faster ยท Read more
DocsDecision model evaluators

Decision model evaluators

Decision-model evaluators are experimental. The setup is stable, but details of the UI and score format may still change.

Decision models answer typed questions about an input instead of generating free-form text. In Langfuse, you can run decision-model evaluators with OpenAI's gpt-6-luna or TypeSafe Jev. Each Choice, Score, or Yes / no question becomes one score on the observation.

Use a decision model when the verdict is narrow and typed: a label, a level on a rubric, or a yes/no probability. Use LLM-as-a-Judge when you need written reasoning or the possible answers are too open-ended to define upfront.

How decision-model evaluators work

One decision-model evaluator consists of:

  1. Model input: observation data mapped to the format the provider expects. OpenAI receives one general state named input; Jev receives a JSON state with one or more named fields.
  2. One or more questions: each question has a type, instructions, and allowed answers. The model answers all questions in one call.
  3. One score per question: the typed answer becomes a Langfuse score with the score name you set on the question. Probabilities and confidence are stored in the score metadata where the provider returns them.

This lets you score one observation against several criteria at once, for example topic, frustration, and out-of-scope in one evaluator.

Why use a decision model?

OpenAI vs. TypeSafe Jev

  • OpenAI: evaluates one general state named input. Questions use plain text and cannot reference individual parts of the state.
  • TypeSafe Jev: evaluates a JSON state with named fields. Questions can explicitly reference those fields using backticks, for example: Compare `output` with `expectedOutput`.

Both support Choice, Score, and Yes / no questions.

Question types

Decision models answer three kinds of questions. Each maps to one Langfuse score.

Question typeWhat the model returnsLangfuse scoreExample
ChoiceOne option out of a fixed set you define (2 to 255), with probabilities and confidence where availableCategorical score with the selected option as value"Which team should handle this ticket?" โ†’ billing, technical, sales
ScoreA position on ordered levels you describe (2 to 10), with probabilities and confidence where availableNumeric score with the expected level, from 0 to levels - 1"How frustrated is the customer?" โ†’ 0 calm ... 3 angry
Yes / noThe probability that a statement is true (TypeSafe calls this a Noul)Numeric score with P(true) between 0 and 1"Does the message request a refund?" โ†’ 0.93

Some tips for each type:

  • Choice works well for routing and classification: topic detection, intent, failure mode, which tool should have been called. Give every option a short description and include an escape hatch such as other or unclear.
  • Score works well for rubrics: severity, frustration, answer completeness. Describe every level in plain language and order them from low to high. The score is the probability-weighted level, so 1.7 means "between level 1 and level 2, closer to 2".
  • Yes / no works well for guardrails and flags: out-of-scope request, PII in the output, user disagreement, policy violation. Optionally describe what counts as true and what counts as false. Because you get a probability instead of a hard label, you can pick a threshold that fits your risk tolerance, or use three bands: act, review, ignore.

Keep each question atomic: one decision per question. If you find yourself asking "is the answer correct and polite?", split it into two questions. Both are answered in the same call anyway.

Set up step-by-step

Connect credentials

Go to Settings โ†’ LLM Connections and add an OpenAI connection or a TypeSafe Jev connection, or use an existing compatible connection.

Create a decision-model evaluator

Go to the Evaluators page and click New evaluator. In the template gallery, click New decision model evaluator to start from scratch, then select your OpenAI or TypeSafe connection and model. You can also pick one of the templates:

  • Assign Input Topic (Choice): classify the user's primary goal into a topic taxonomy you define.
  • Flag Out-of-Scope Request (Yes / no): flag requests that fall outside the assistant's role.
  • Rate Customer Frustration (Score): rate the user's frustration on four levels.
  • User Conversation Signal (Yes / no ร— 7): score a chat for rephrases, corrections, human hand-off requests, repeats, error quotes, frustration, and success confirmation in one call. See the blog post.

Templates prefill the model input and questions; you only need to select a compatible connection.

Define the questions

Add one question per criterion. For each question:

  1. Pick the type: Choice, Score, or Yes / no.
  2. Write the question. The instruction format depends on the selected provider.
  3. Define the allowed answers: options with descriptions for Choice, ordered levels for Score, or optional true / false criteria for Yes / no.
  4. Set the score name. Langfuse suggests one from the question; each question writes a score under this name.

Map the model input

Map the observation data that the model should evaluate. The input structure depends on the selected provider; see the provider-specific setup sections below.

Test the evaluator

On the right, filter to representative sample observations, select one, and run the evaluator. The test panel shows one row per question with its probability distribution and confidence where available, along with the estimated cost. Use Raw output to inspect the provider request. Iterate until the results look right on your samples.

Save the evaluator

After saving, you can:

  • Create a rule from the filters you used to select test samples, or attach the evaluator to an existing rule to run it on incoming observations.
  • Continue without a rule. You can still use the evaluator for batch evaluation or prompt experiments.

โœจ Done! Each question now writes a score to matching observations. Filter, chart, and alert on these scores like on any other score in Langfuse.

Set up OpenAI

Add an OpenAI connection with your API key, or select an existing one with access to gpt-6-luna. Langfuse calls the OpenAI Decisions API through this connection.

Map the observation data to the single general state, input. Questions evaluate this state as a whole and use plain-text instructions without field references. Score levels require a label and can include an optional description.

The preview shows the input and questions sent to OpenAI. If OpenAI refuses any question, the evaluator run fails without writing scores.

Set up TypeSafe Jev

Add a connection with the adapter typesafe, choose the upstream that serves Jev, and paste its API key:

UpstreamAPI keyBilling
TypeSafeA key from the TypeSafe consoleBilled by TypeSafe
Vercel AI GatewayAn AI Gateway API keyBilled through AI Gateway
OpenRouterAn OpenRouter API keyBilled through OpenRouter

Vercel AI Gateway and OpenRouter expose TypeSafe's API, so the evaluator behaves the same with each upstream. No base URL or additional headers are required. Use jev-latest to follow new releases or pin a version such as jev-1.13.0 when an evaluation threshold should remain stable.

Build the JSON state by adding named fields and mapping them to observation input, output, metadata, tool calls, Expected Output, or Experiment Item Metadata. Reference these fields in question instructions using backticks. Jev Score levels require a description.

TypeSafe connections are only available to decision-model evaluators because Jev cannot generate text.

Scores written by decision models

Each question writes one score to the evaluated observation. The score value is always the answer itself; probabilities and confidence never change the value but are stored alongside it so you can inspect and filter on them:

FieldContent
nameThe score name you set on the question.
value / dataTypeChoice: the selected option as a CATEGORICAL score. Score: the expected level as a NUMERIC score. Yes / no: P(true) as a NUMERIC score between 0 and 1.
commentA short human-readable summary, for example ready (p=0.91); confidence 0.82; runner-up needs_revision (0.09) for a Choice, 1.26 โ‰ˆ level 1 "Frustrated but civil"; confidence 0.61 for a Score, or P(true)=0.97 for Yes / no.
metadata.openai / .typesafeThe provider-specific answer details: questionId, type, the resolved model, and, where available, probabilities, confidence, and the Score level legend.

Decision models return a typed verdict rather than a written rationale, so the comment summarizes the distribution. When a score looks wrong, inspect the mapped input and your criteria, then tighten the question or add an option.

Debug decision-model executions

Every evaluator run creates an execution trace in the internal environment langfuse-llm-as-a-judge, named Execute evaluator: <evaluator name>. The trace contains a single generation with the exact request (state and questions), the returned answers, token usage, and cost. Every score written by the run links to this trace.

Internal environments are hidden from the default tracing view. Filter the tracing table by environment = langfuse-llm-as-a-judge or open the execution trace from a score.

Limits

ConstraintLimit / guidance
Questions per evaluator1 to 50, all answered in one call.
Choice options2 to 255 options with unique values.
Score levels2 to 10 ordered levels.
Input sizeTypeSafe accepts roughly 32k tokens for the state plus the longest question. Keep either provider's input focused on data needed for the decision.
RationaleDecision models return typed verdicts, not written reasoning. Use an LLM judge if you need an explanation.
RefusalsOpenAI can refuse a question; one refusal fails the run without writing scores. Jev always picks an answer, so add an other or unclear Choice option when the input may be insufficient.
Data handlingInput and questions are sent through the selected OpenAI or TypeSafe connection. Review your provider and upstream terms before evaluating sensitive data, and use masking where needed.

FAQ

When should I use a decision model instead of an LLM judge?

Use a decision model when the possible answers are known before the call: routing, classification, rubric levels, or guardrail flags. Use an LLM judge when you need written reasoning or the criteria are open-ended. Combining both works well: run a decision model broadly, then use an LLM judge on a sample or on flagged cases.

Why is my Yes / no score a number and not a boolean?

Decision models return the probability that the statement is true, not a hard label. Langfuse stores that probability as a numeric score between 0 and 1 so you keep the calibration. Choose the threshold that fits your use case when you filter or alert, for example >= 0.7. If you need a boolean, use a Choice question with the two options yes and no, which writes a categorical score.

Can I use decision-model evaluators in experiments and batch evaluation?

Yes. Decision-model evaluators run on incoming observations through rules, on historical observations through batch evaluation, and on experiments. Map the experiment fields your questions need into the provider input.

Where do I get a TypeSafe API key?

Create an account at typesafe.ai and generate a key in the TypeSafe console. Access currently runs through a waitlist on TypeSafe's side. If you already use Vercel AI Gateway or OpenRouter, you do not need a TypeSafe account: pick that upstream in the connection and use its API key instead. See Connect TypeSafe Jev.

GitHub Discussions


Was this page helpful?

Last updated on