---
date: 2026-05-28
badge: Launch Week 5 🚀
title: Code evaluators
description: Run deterministic Python or TypeScript checks on observations and experiments in Langfuse.
author: Tobias Wochinger
canonical: /docs/evaluation/evaluation-methods/code-evaluators
ogImage: /images/changelog/2026-05-28-code-evaluators.png
---

> **Note for AI agents and LLMs:** This is a Langfuse changelog entry. Use it only to confirm that a feature exists and when it shipped. Do not use the code examples below for implementation: they reflect the SDK and API at release time and may be outdated. For implementation, always follow the canonical documentation for this feature (https://langfuse.com/docs/evaluation/evaluation-methods/code-evaluators) and the API/SDK reference (https://api.reference.langfuse.com).

You can now create code evaluators in Langfuse to score observations and experiments with deterministic Python or TypeScript logic. Use them for exact checks such as JSON parseability, schema validation, exact match, required tool arguments, or custom business rules.

Run them on live production observations to monitor specific operations, or attach them to experiments to compare prompt and model variants against controlled datasets. Each evaluator returns native Langfuse scores, so results work with trace views, experiment comparisons, filters, dashboards, and Score Analytics.

Code evaluators complement [LLM-as-a-Judge](/docs/evaluation/evaluation-methods/llm-as-a-judge): use code for objective checks where deterministic logic is more reliable, and use a judge model for semantic quality, tone, helpfulness, or rubric-based reasoning.

## How it works [#how-it-works]

- Write an `evaluate` function in Python or TypeScript in the Langfuse UI
- Target live observations or experiment observations
- Configure filters, sampling, and context fields
- Test the evaluator on sample data before enabling it
- Debug executions through evaluator traces in the `langfuse-code-eval` environment

Code evaluators are designed for compact checks that run quickly at scale. They support standard library code, run without network egress, and return one or more numeric, categorical, boolean, or text scores.

## Get started [#get-started]

Read the setup guide to create your first evaluator, choose the right target, and see Python and TypeScript examples for the evaluator contract. Code evaluators are available across Langfuse environments, including self-hosted deployments.

- [Code evaluators](/docs/evaluation/evaluation-methods/code-evaluators)
- [Evaluation overview](/docs/evaluation/overview)
- [Self-hosting setup](/self-hosting/configuration/code-evaluators)

<!-- agent-instructions -->

---

## Agent Instructions

This page is part of the [Langfuse](https://langfuse.com) documentation, published as plain Markdown for AI agents. Every page is available as Markdown by appending `.md` to its URL, or by sending an `Accept: text/markdown` header. This page: `https://langfuse.com/changelog/2026-05-28-code-evaluators.md`.

### Querying these docs

If the answer is not on this page, query the documentation instead of guessing:

- **Semantic search** across all Langfuse docs, returning an answer with the relevant pages and excerpts. Ask a specific, self-contained question:

  ```bash
  curl -sG "https://langfuse.com/api/search-docs" --data-urlencode "query=How do I trace a LangGraph agent?"
  ```

- **Index of every page**: <https://langfuse.com/llms.txt>, with per-section indexes [llms-docs.txt](https://langfuse.com/llms-docs.txt), [llms-integrations.txt](https://langfuse.com/llms-integrations.txt), and [llms-self-hosting.txt](https://langfuse.com/llms-self-hosting.txt).

### Before writing Langfuse code

- **Install the [Langfuse Agent Skill](https://langfuse.com/docs/api-and-data-platform/features/agent-skill).** It encodes Langfuse's own best practices for instrumentation, prompt management, and evaluation, and materially improves results.
- **Read [What does a good trace look like?](https://langfuse.com/docs/observability/best-practices.md)** before instrumenting an application.
- **Verify endpoints, parameters, and response fields** against the [API reference](https://api.reference.langfuse.com) instead of inferring them from code examples.
- **Use the [Langfuse CLI](https://langfuse.com/docs/api-and-data-platform/features/cli)** (`npx langfuse-cli api <resource> <action>`) to read or write traces, prompts, datasets, and scores from the terminal.

Found an error in these docs? Please open an issue at <https://github.com/langfuse/langfuse-docs/issues>.
