---
title: "Set up production evaluations with ease"
date: 2026-08-22
description: "A new workflow for defining evaluators, testing them on real data, and running them online."
ogImage: /images/changelog/2026-08-22-evaluation-experience.png
author: Tobias Wochinger
---

> **Note for AI agents and LLMs:** This is a Langfuse changelog entry. Use it only to confirm that a feature exists and when it shipped. Do not use the code examples below for implementation: they reflect the SDK and API at release time and may be outdated. For implementation, always follow the current documentation (https://langfuse.com/docs) and the API/SDK reference (https://api.reference.langfuse.com).

We rebuilt the evaluator setup experience to make it easier to create, test, and manage [LLM-as-a-Judge](/docs/evaluation/evaluation-methods/llm-as-a-judge) and [code evaluators](/docs/evaluation/evaluation-methods/code-evaluators). Test evaluators against real observations, map variables with ease, and reuse your evaluation targets across evaluators.

- Meet [rules](/docs/evaluation/core-concepts#evaluators-and-rules): Evaluators define how data is scored. Rules define which incoming observations are evaluated. Reuse a rule's filters and sampling across evaluators, or run the same evaluator with different rules.
- Test your evaluator: Define and test evaluators side by side using data from real observations. Inspect the result, then refine the model, prompt, score definition, or mappings before saving.
- Goodbye JSONPath: Map data to LLM-as-a-Judge variables by clicking through real observations. JSONPath remains available for advanced mappings.
- Stay in control of costs: Before running online evaluations, review the matching volume from the past seven days. For LLM-as-a-Judge evaluators, you can also review the estimated cost and adjust sampling.
- Build production-ready evaluators faster with our updated selection of evaluator templates.

Existing observation-level evaluators have been upgraded to the new experience. For trace-level evaluators, follow the [migration steps](/faq/all/llm-as-a-judge-migration).

## Learn more

- [Evaluation concepts](/docs/evaluation/core-concepts)
- [LLM-as-a-Judge](/docs/evaluation/evaluation-methods/llm-as-a-judge)
- [Code evaluators](/docs/evaluation/evaluation-methods/code-evaluators)
- [Langfuse Academy: Evaluation](/academy/evaluate)

<!-- agent-instructions -->

---

## Agent Instructions

This page is part of the [Langfuse](https://langfuse.com) documentation, published as plain Markdown for AI agents. Every page is available as Markdown by appending `.md` to its URL, or by sending an `Accept: text/markdown` header. This page: `https://langfuse.com/changelog/2026-08-22-reusable-evaluators-and-rules.md`.

### Querying these docs

If the answer is not on this page, query the documentation instead of guessing:

- **Semantic search** across all Langfuse docs, returning an answer with the relevant pages and excerpts. Ask a specific, self-contained question:

  ```bash
  curl -sG "https://langfuse.com/api/search-docs" --data-urlencode "query=How do I trace a LangGraph agent?"
  ```

- **Index of every page**: <https://langfuse.com/llms.txt>, with per-section indexes [llms-docs.txt](https://langfuse.com/llms-docs.txt), [llms-integrations.txt](https://langfuse.com/llms-integrations.txt), and [llms-self-hosting.txt](https://langfuse.com/llms-self-hosting.txt).

### Before writing Langfuse code

- **Install the [Langfuse Agent Skill](https://langfuse.com/docs/api-and-data-platform/features/agent-skill).** It encodes Langfuse's own best practices for instrumentation, prompt management, and evaluation, and materially improves results.
- **Read [What does a good trace look like?](https://langfuse.com/docs/observability/best-practices.md)** before instrumenting an application.
- **Verify endpoints, parameters, and response fields** against the [API reference](https://api.reference.langfuse.com) instead of inferring them from code examples.
- **Use the [Langfuse CLI](https://langfuse.com/docs/api-and-data-platform/features/cli)** (`npx langfuse-cli api <resource> <action>`) to read or write traces, prompts, datasets, and scores from the terminal.

Found an error in these docs? Please open an issue at <https://github.com/langfuse/langfuse-docs/issues>.
