---
title: Agentic access
sidebarTitle: Agent Access
description: Let AI agents work with scores, datasets, experiments, evaluators, and annotation queues through the Agent Skill, CLI, or MCP server.
---

# Agentic access to evaluation

AI agents can help investigate quality issues and operate evaluation workflows in Langfuse. There are different ways for agents to access your data:

## Choose an access method

| Agent capabilities                         | Recommended access                                                                                                                                                                                                                    |
| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Can install tools and run shell commands   | Install the [Langfuse Agent Skill](/docs/api-and-data-platform/features/agent-skill). It teaches the agent Langfuse workflows and uses the [Langfuse CLI](/docs/api-and-data-platform/features/cli) to query and update project data. |
| Cannot install tools or run shell commands | Connect the [Langfuse MCP server](/docs/api-and-data-platform/features/mcp-server) to expose Langfuse operations as tools.                                                                                                            |
| Runs as part of a script or CI/CD pipeline | Use the [Langfuse CLI](/docs/api-and-data-platform/features/cli) directly or call the [Public API](/docs/api-and-data-platform/features/public-api).                                                                                  |

## Example workflows

Ask your agent to:

- Find low-scoring observations and add representative examples to a dataset
- Create or update score configurations and record scores
- Review experiment results and identify regressions
- Set up evaluators and evaluation rules
- Create and manage annotation queues for human review

## Work across Langfuse

Agents can also [investigate production behavior](/docs/observability/features/agentic-access) and [manage prompts](/docs/prompt-management/features/agentic-access) in Langfuse.

## Related guides and blog posts

- [Calibrate LLM-as-a-judge with the Langfuse skill](/guides/llm-as-a-judge-calibration-skill) — Use an agent-guided workflow to compare an evaluator with human labels and improve its prompt.
- [Evaluating AI Agent Skills](/blog/2026-02-26-evaluate-ai-agent-skills) — See how Langfuse datasets, tracing, and an agent SDK can be used to iteratively evaluate and improve a skill.
- [Headless Langfuse from your coding agent](/guides/videos/headless-langfuse) — Analyze production traces, build a dataset, and configure evaluations without leaving your coding agent.

<!-- agent-instructions -->

---

## Agent Instructions

This page is part of the [Langfuse](https://langfuse.com) documentation, published as plain Markdown for AI agents. Every page is available as Markdown by appending `.md` to its URL, or by sending an `Accept: text/markdown` header. This page: `https://langfuse.com/docs/evaluation/agentic-access.md`.

### Querying these docs

If the answer is not on this page, query the documentation instead of guessing:

- **Semantic search** across all Langfuse docs, returning an answer with the relevant pages and excerpts. Ask a specific, self-contained question:

  ```bash
  curl -sG "https://langfuse.com/api/search-docs" --data-urlencode "query=How do I trace a LangGraph agent?"
  ```

- **Index of every page**: <https://langfuse.com/llms.txt>, with per-section indexes [llms-docs.txt](https://langfuse.com/llms-docs.txt), [llms-integrations.txt](https://langfuse.com/llms-integrations.txt), and [llms-self-hosting.txt](https://langfuse.com/llms-self-hosting.txt).

### Before writing Langfuse code

- **Install the [Langfuse Agent Skill](https://langfuse.com/docs/api-and-data-platform/features/agent-skill).** It encodes Langfuse's own best practices for instrumentation, prompt management, and evaluation, and materially improves results.
- **Read [What does a good trace look like?](https://langfuse.com/docs/observability/best-practices.md)** before instrumenting an application.
- **Verify endpoints, parameters, and response fields** against the [API reference](https://api.reference.langfuse.com) instead of inferring them from code examples.
- **Use the [Langfuse CLI](https://langfuse.com/docs/api-and-data-platform/features/cli)** (`npx langfuse-cli api <resource> <action>`) to read or write traces, prompts, datasets, and scores from the terminal.

Found an error in these docs? Please open an issue at <https://github.com/langfuse/langfuse-docs/issues>.
