---
title: "07 Evaluate a Change"
description: "Learner guide: 07 Evaluate a Change"
---
# 07 Evaluate a Change

Workshop material is maintained in the public [`langfuse/langfuse-workshop`](https://github.com/langfuse/langfuse-workshop) repository. Use the repository for the runnable app, checkpoint branches, and local setup.

[View this Markdown file](https://github.com/langfuse/langfuse-workshop/blob/main/docs/instructor/07-evaluation.md)

Learner guide: [07 Evaluate a Change](/workshop/learner/07-evaluation)

## Instructor notes

- Make learners inspect run 1 before changing anything. The change should respond to evidence, not vibes.
- Keep the iteration deliberately small: one prompt rule, one rerun, one comparison.
- Emphasize regressions. The most useful comparison is often the item that got worse.
- Remind learners that both experiment scores are platform-side now, so a missing score usually means "still pending" or "evaluator target mismatch," not a bug in `scripts/run-dataset.ts`.

## Demo rhythm

1. Read low-scoring items from the first run.
2. Add or promote a new prompt version.
3. Run `npm run dataset:run` again.
4. Compare both runs side by side and decide whether the change is worth shipping.

## Watch for

- Learners changing both model and prompt at the same time, making the comparison hard to interpret.
- Learners editing the prompt in the UI but forgetting they need a new saved version promoted to the label the app fetches.

<!-- agent-instructions -->

---

## Agent Instructions

This page is part of the [Langfuse](https://langfuse.com) documentation, published as plain Markdown for AI agents. Every page is available as Markdown by appending `.md` to its URL, or by sending an `Accept: text/markdown` header. This page: `https://langfuse.com/workshop/instructor/07-evaluation.md`.

### Querying these docs

If the answer is not on this page, query the documentation instead of guessing:

- **Semantic search** across all Langfuse docs, returning an answer with the relevant pages and excerpts. Ask a specific, self-contained question:

  ```bash
  curl -sG "https://langfuse.com/api/search-docs" --data-urlencode "query=How do I trace a LangGraph agent?"
  ```

- **Index of every page**: <https://langfuse.com/llms.txt>, with per-section indexes [llms-docs.txt](https://langfuse.com/llms-docs.txt), [llms-integrations.txt](https://langfuse.com/llms-integrations.txt), and [llms-self-hosting.txt](https://langfuse.com/llms-self-hosting.txt).

### Before writing Langfuse code

- **Install the [Langfuse Agent Skill](https://langfuse.com/docs/api-and-data-platform/features/agent-skill).** It encodes Langfuse's own best practices for instrumentation, prompt management, and evaluation, and materially improves results.
- **Read [What does a good trace look like?](https://langfuse.com/docs/observability/best-practices.md)** before instrumenting an application.
- **Verify endpoints, parameters, and response fields** against the [API reference](https://api.reference.langfuse.com) instead of inferring them from code examples.
- **Use the [Langfuse CLI](https://langfuse.com/docs/api-and-data-platform/features/cli)** (`npx langfuse-cli api <resource> <action>`) to read or write traces, prompts, datasets, and scores from the terminal.

Found an error in these docs? Please open an issue at <https://github.com/langfuse/langfuse-docs/issues>.
