OpenAI Codex tracing with Langfuse
What is Codex? Codex is OpenAI's agentic coding tool. It runs in your terminal (and IDE), understands your codebase, edits code, runs commands, and can spawn subagents to work on tasks in parallel.
What is Langfuse? Langfuse is an open-source AI engineering platform. It helps teams trace agentic applications, debug issues, evaluate quality, and monitor costs in production. Use Langfuse Cloud (hosted by Langfuse, free tier, no infrastructure to run) or self-host it.
Install via the Codex plugin marketplace
The easiest way to set this up is the Langfuse Codex Plugin. Add the marketplace and enable the plugin:
codex plugin marketplace add langfuse/codex-observability-pluginThen install the tracing plugin, enable its hook, and set your Langfuse keys (full steps below). Requires Node.js 22+ and Codex 0.143+.
What can this integration trace?
Using Codex's plugin hooks, this integration reads the transcript Codex writes for each session and sends it to Langfuse. You can monitor:
- User inputs: every prompt you send to Codex, including any images you attach. Images are stored as Langfuse media objects and rendered in the trace.
- Model responses: assistant messages and reasoning summaries for each model call in a turn
- System prompts: Codex's base instructions, its
developer-role messages, and the injected<environment_context>, sent with every model call and measured on the turn so you can see when the system prompt changed - Tool calls: shell commands (
exec_command,shell), file edits (apply_patch), MCP tools, web search, and subagent spawns (spawn_agent), with their inputs, outputs, and error status - Token usage: input, output, cached, and reasoning tokens per model call, so you can monitor cost
- Model metadata: model, provider, reasoning effort, and the Codex CLI version behind each turn
- Subagents: subagent threads resolved from their own transcripts and nested under the spawning turn
- Skills: a
skill:<name>tag on every turn that invoked a skill, whether you invoked it explicitly or the agent picked it up itself - Sessions: all turns from one Codex session grouped together for replay
- Timing: accurate, backdated start and end times for every step
Interrupted turns, where you cancel Codex mid-response, are uploaded as well and flagged as interrupted. Prompts, tool inputs, and tool outputs are captured in full and are never truncated.
How it works
Codex provides a plugin system with hooks that run custom commands at lifecycle points. This integration uses the Stop hook, which runs after each Codex turn.
- The plugin registers a
Stophook that runs each time Codex finishes a turn. - The hook reads Codex's session transcript (the rollout file).
- Turns are reconstructed and converted into Langfuse traces using the Langfuse TypeScript SDK.
- All turns from the same session are grouped using a shared
session_id. - A small sidecar file records which turns were already uploaded, so resuming a session never creates duplicates.
Tracing is opt-in via the TRACE_TO_LANGFUSE environment variable. The hook fails open, so if anything goes wrong it logs and exits without blocking your Codex session.
Quick start
Set up Langfuse
- Sign up for Langfuse Cloud or self-host Langfuse.
- Create a new project and copy your API keys from the project settings.
Add the plugin marketplace
Add the Langfuse marketplace via the Codex CLI:
codex plugin marketplace add langfuse/codex-observability-pluginInstall and enable the plugin
Install the tracing plugin from the marketplace:
codex plugin add tracing@codex-observability-pluginEnable hooks and the tracing plugin globally in ~/.codex/config.toml, or only for a specific project in <project>/.codex/config.toml:
[features]
hooks = true
[plugins."tracing@codex-observability-plugin"]
enabled = trueVerify the feature with codex features list, which shows hooks and its effective state. Older Codex releases used a plugin_hooks key, which has since been removed and no longer has any effect.
When Codex first runs the plugin hook, approve the Langfuse Stop hook if Codex asks for permission. Codex stores hook trust separately from plugin installation, and it records trust against the current hook hash, so a plugin update can ask you to review the hook again. If you previously trusted the hook but it remains inactive, make sure this generated hook-state entry is enabled:
[hooks.state."tracing@codex-observability-plugin:hooks/hooks.json:stop:0:0"]
enabled = trueSet your Langfuse credentials
Tracing stays off until TRACE_TO_LANGFUSE is "true", so you opt in explicitly. Add your credentials to your shell profile (~/.zshrc, ~/.bashrc, or ~/.bash_profile):
export TRACE_TO_LANGFUSE="true"
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com" # 🇪🇺 EU regionAlternatively, create a JSON config file at ~/.codex/langfuse.json (global) or <project>/.codex/langfuse.json (per-project):
{
"enabled": true,
"public_key": "pk-lf-...",
"secret_key": "sk-lf-...",
"base_url": "https://cloud.langfuse.com"
}Configuration is resolved as defaults, then the global config file, then the project config file, then environment variables, with environment variables taking precedence. LANGFUSE_CODEX_* variables override the matching standard LANGFUSE_* variables, so you can scope credentials to Codex.
Restart Codex and verify tracing
Fully restart Codex after changing its global configuration, then start a new Codex session. Global configuration applies to new sessions in every project; existing sessions do not load the hook retroactively.
For a reliable test, send two short messages. The Stop hook uploads each completed turn, while the latest turn is finalized on the next hook invocation.
Run Codex as usual:
cd your-project
codexView traces in Langfuse
Open your Langfuse project to see the captured traces. Search for Codex Turn and widen the time range if needed; Langfuse timestamps may be displayed in UTC. The structure mirrors how Codex actually works:
- Turn trace (
Codex Turn): one trace per turn, from your prompt to the final answer, captured as an agent observation. - Generations (
LLM): one per model response in the turn. Each shows the input it received, including the system prompt, the model's reasoning and text, the tool calls it requested, and token usage. - Tool spans (
exec_command,apply_patch,spawn_agent, …): nested under the generation that triggered them, with input, output, and error status. Failed commands are flagged as errors. MCP tools appear asserver.tool, and a command that loads a skill appears asskill:<name>. - Subagents (
Codex Subagent Turn, withLLM Subagentgenerations): subagent threads are nested under the spawning turn so you can follow parallel work in one place. - Sessions: all turns from the same Codex session are grouped via
session_id. Open the Sessions tab to replay the full run.
Environment variables
| Variable | Description | Required |
|---|---|---|
TRACE_TO_LANGFUSE | Set to "true" to enable tracing | Yes |
LANGFUSE_PUBLIC_KEY | Your Langfuse public key (pk-lf-...) | Yes |
LANGFUSE_SECRET_KEY | Your Langfuse secret key (sk-lf-...) | Yes |
LANGFUSE_BASE_URL | Langfuse host. EU: https://cloud.langfuse.com, US: https://us.cloud.langfuse.com, Japan: https://jp.cloud.langfuse.com, HIPAA: https://hipaa.cloud.langfuse.com | No (defaults to EU) |
LANGFUSE_TRACING_ENVIRONMENT | Environment label for the traces (e.g. production) | No |
LANGFUSE_CODEX_USER_ID | User attached to every trace, shown as the user in Langfuse. Defaults to the Codex auth email, if one is found. | No |
LANGFUSE_CODEX_TAGS | Tags for all traces (JSON array or comma-separated) | No |
LANGFUSE_CODEX_METADATA | JSON object of metadata to attach to all traces | No |
LANGFUSE_CODEX_SKILL_TAGS | Tag traces with skill:<name> for every skill invoked in the turn (default true) | No |
LANGFUSE_CODEX_TRACE_SEED | Seed that makes trace IDs predictable, so a caller can derive a run's trace ID before the trace exists. Use a unique seed per session, otherwise sessions collide. | No |
LANGFUSE_CODEX_DEBUG | Set to "true" for verbose logging to stderr | No |
LANGFUSE_CODEX_FAIL_ON_ERROR | Set to "true" to make upload errors fail the hook instead of failing open. Useful together with debug logging while testing. | No |
The credential variables also accept a LANGFUSE_CODEX_ prefix (for example LANGFUSE_CODEX_PUBLIC_KEY), which takes precedence over the standard variable.
Every variable has an equivalent key in langfuse.json, using the lower-case name without the prefix. For example LANGFUSE_CODEX_SKILL_TAGS becomes "skill_tags", and TRACE_TO_LANGFUSE becomes "enabled".
Troubleshooting
No traces appearing in Langfuse
- The plugin isn't installed or enabled. Run
codex plugin add tracing@codex-observability-plugin, then confirmhooks = trueunder[features]and that thetracing@codex-observability-pluginplugin is enabled in~/.codex/config.toml. Runcodex plugin listandcodex features listto check both. - The Stop hook is disabled. Approve the Langfuse Stop hook when Codex prompts you. If the hook has already been trusted, verify that its generated entry under
[hooks.state]hasenabled = true. A plugin update changes the hook hash and can require a fresh review. - Tracing isn't turned on.
TRACE_TO_LANGFUSEmust be the exact string"true"and visible to the Codex process, unless you enabled tracing in~/.codex/langfuse.json. Also verify the public key starts withpk-lf-. - Restart and test again. Fully restart Codex, start a new session, and send two short messages before checking Langfuse.
- Enable debug logging. Set
LANGFUSE_CODEX_DEBUG=trueto log to stderr and surface the actual cause. AddLANGFUSE_CODEX_FAIL_ON_ERROR=trueto turn a silent upload failure into a visible hook error.
Authentication errors
Verify your API keys are correct and that LANGFUSE_BASE_URL matches the region your keys belong to:
- EU region:
https://cloud.langfuse.com - US region:
https://us.cloud.langfuse.com - Japan region:
https://jp.cloud.langfuse.com - HIPAA region:
https://hipaa.cloud.langfuse.com
Data privacy
When enabled, the plugin uploads completed Codex transcript data to Langfuse, including prompts, attached images, assistant messages, reasoning summaries, system prompts, tool inputs and outputs, model metadata, and token usage. This content is captured in full and is not truncated, so do not enable tracing for sessions containing data you do not want stored in Langfuse.
Resources
- Langfuse Codex Plugin (GitHub)
- OpenAI Codex documentation
- Langfuse Sessions
- Langfuse TypeScript SDK
- Tracing coding agents with Langfuse
Last updated on