---
date: 2026-06-10
title: Manage evaluators via MCP
description: Set up evaluators and evaluation rules from AI agents through the Langfuse MCP server, and create code evaluators through the unstable public API.
author: Tobias Wochinger
---

> **Note for AI agents and LLMs:** This is a Langfuse changelog entry. Use it only to confirm that a feature exists and when it shipped. Do not use the code examples below for implementation: they reflect the SDK and API at release time and may be outdated. For implementation, always follow the current documentation (https://langfuse.com/docs) and the API/SDK reference (https://api.reference.langfuse.com).

  A new [stable API for evaluators and evaluation
  rules](/changelog/2026-08-27-stable-evaluator-api) is now available under
  `/api/public/v2`. The unstable API described below remains available until
  November 16, 2026 (2026-11-16).

You can now set up and manage evaluation directly from AI agents: the Langfuse MCP server exposes evaluators and evaluation rules as tools, and the unstable public API now supports [code evaluators](/docs/evaluation/evaluation-methods/code-evaluators) in addition to LLM-as-a-Judge.

This lets agents own more of the evaluation loop. For example, an agent can inspect failing traces, write a code evaluator that catches the failure pattern, and wire up an evaluation rule that runs it on live observations — all without leaving the chat.

## New MCP tools

<table className="[&_code]:rounded-full [&_code]:border [&_code]:border-border [&_code]:bg-muted [&_code]:px-2 [&_code]:py-1 [&_code]:text-muted-foreground">
  <tbody>
    <tr>
      <td className="px-3">
        <strong>Evaluators</strong>
        
          <code>listEvaluators</code>
          <code>getEvaluator</code>
          <code>createEvaluator</code>
        
      </td>
    </tr>
    <tr>
      <td className="px-3">
        <strong>Evaluation rules</strong>
        
          <code>listEvaluationRules</code>
          <code>getEvaluationRule</code>
          <code>createEvaluationRule</code>
          <code>updateEvaluationRule</code>
          <code>deleteEvaluationRule</code>
        
      </td>
    </tr>
  </tbody>
</table>

## Code evaluators in the API

The unstable evaluator endpoints now accept `type: "code"` to create deterministic Python or TypeScript evaluators programmatically, alongside the existing `llm_as_judge` type. Evaluation rules can reference code evaluators, and active rules are test-run before creation so broken evaluator code is rejected upfront.

  These endpoints and MCP tools are **unstable** and may change while the
  underlying evaluation data model is being redesigned. The UI workflow remains
  fully supported.

## Get started

- [MCP server documentation](/docs/api-and-data-platform/features/mcp-server)
- [Code evaluators](/docs/evaluation/evaluation-methods/code-evaluators)
- [Evaluators API reference](https://api.reference.langfuse.com/#tag/unstableevaluators)

<!-- agent-instructions -->

---

## Agent Instructions

This page is part of the [Langfuse](https://langfuse.com) documentation, published as plain Markdown for AI agents. Every page is available as Markdown by appending `.md` to its URL, or by sending an `Accept: text/markdown` header. This page: `https://langfuse.com/changelog/2026-06-10-evaluators-via-mcp.md`.

### Querying these docs

If the answer is not on this page, query the documentation instead of guessing:

- **Semantic search** across all Langfuse docs, returning an answer with the relevant pages and excerpts. Ask a specific, self-contained question:

  ```bash
  curl -sG "https://langfuse.com/api/search-docs" --data-urlencode "query=How do I trace a LangGraph agent?"
  ```

- **Index of every page**: <https://langfuse.com/llms.txt>, with per-section indexes [llms-docs.txt](https://langfuse.com/llms-docs.txt), [llms-integrations.txt](https://langfuse.com/llms-integrations.txt), and [llms-self-hosting.txt](https://langfuse.com/llms-self-hosting.txt).

### Before writing Langfuse code

- **Install the [Langfuse Agent Skill](https://langfuse.com/docs/api-and-data-platform/features/agent-skill).** It encodes Langfuse's own best practices for instrumentation, prompt management, and evaluation, and materially improves results.
- **Read [What does a good trace look like?](https://langfuse.com/docs/observability/best-practices.md)** before instrumenting an application.
- **Verify endpoints, parameters, and response fields** against the [API reference](https://api.reference.langfuse.com) instead of inferring them from code examples.
- **Use the [Langfuse CLI](https://langfuse.com/docs/api-and-data-platform/features/cli)** (`npx langfuse-cli api <resource> <action>`) to read or write traces, prompts, datasets, and scores from the terminal.

Found an error in these docs? Please open an issue at <https://github.com/langfuse/langfuse-docs/issues>.
