Langfuse v4: up to 165× faster · Read more
IntegrationsCodex

OpenAI Codex tracing with Langfuse

What is Codex? Codex is OpenAI's agentic coding tool. It runs in your terminal (and IDE), understands your codebase, edits code, runs commands, and can spawn subagents to work on tasks in parallel.

What is Langfuse? Langfuse is an open-source AI engineering platform. It helps teams trace agentic applications, debug issues, evaluate quality, and monitor costs in production. Use Langfuse Cloud (hosted by Langfuse, free tier, no infrastructure to run) or self-host it.

Install via the Codex plugin marketplace

The easiest way to set this up is the Langfuse Codex Plugin. Add the marketplace and enable the plugin:

codex plugin marketplace add langfuse/codex-observability-plugin

Then install the tracing plugin, enable its hook, and set your Langfuse keys (full steps below). Requires Node.js 22+ and Codex 0.143+.

What can this integration trace?

Using Codex's plugin hooks, this integration reads the transcript Codex writes for each session and sends it to Langfuse. You can monitor:

  • User inputs: every prompt you send to Codex, including any images you attach. Images are stored as Langfuse media objects and rendered in the trace.
  • Model responses: assistant messages and reasoning summaries for each model call in a turn
  • System prompts: Codex's base instructions, its developer-role messages, and the injected <environment_context>, sent with every model call and measured on the turn so you can see when the system prompt changed
  • Tool calls: shell commands (exec_command, shell), file edits (apply_patch), MCP tools, web search, and subagent spawns (spawn_agent), with their inputs, outputs, and error status
  • Token usage: input, output, cached, and reasoning tokens per model call, so you can monitor cost
  • Model metadata: model, provider, reasoning effort, and the Codex CLI version behind each turn
  • Subagents: subagent threads resolved from their own transcripts and nested under the spawning turn
  • Skills: a skill:<name> tag on every turn that invoked a skill, whether you invoked it explicitly or the agent picked it up itself
  • Sessions: all turns from one Codex session grouped together for replay
  • Timing: accurate, backdated start and end times for every step

Interrupted turns, where you cancel Codex mid-response, are uploaded as well and flagged as interrupted. Prompts, tool inputs, and tool outputs are captured in full and are never truncated.

How it works

Codex provides a plugin system with hooks that run custom commands at lifecycle points. This integration uses the Stop hook, which runs after each Codex turn.

  1. The plugin registers a Stop hook that runs each time Codex finishes a turn.
  2. The hook reads Codex's session transcript (the rollout file).
  3. Turns are reconstructed and converted into Langfuse traces using the Langfuse TypeScript SDK.
  4. All turns from the same session are grouped using a shared session_id.
  5. A small sidecar file records which turns were already uploaded, so resuming a session never creates duplicates.

Tracing is opt-in via the TRACE_TO_LANGFUSE environment variable. The hook fails open, so if anything goes wrong it logs and exits without blocking your Codex session.

Quick start

Set up Langfuse

  1. Sign up for Langfuse Cloud or self-host Langfuse.
  2. Create a new project and copy your API keys from the project settings.

Add the plugin marketplace

Add the Langfuse marketplace via the Codex CLI:

codex plugin marketplace add langfuse/codex-observability-plugin

Install and enable the plugin

Install the tracing plugin from the marketplace:

codex plugin add tracing@codex-observability-plugin

Enable hooks and the tracing plugin globally in ~/.codex/config.toml, or only for a specific project in <project>/.codex/config.toml:

[features]
hooks = true

[plugins."tracing@codex-observability-plugin"]
enabled = true

Verify the feature with codex features list, which shows hooks and its effective state. Older Codex releases used a plugin_hooks key, which has since been removed and no longer has any effect.

When Codex first runs the plugin hook, approve the Langfuse Stop hook if Codex asks for permission. Codex stores hook trust separately from plugin installation, and it records trust against the current hook hash, so a plugin update can ask you to review the hook again. If you previously trusted the hook but it remains inactive, make sure this generated hook-state entry is enabled:

[hooks.state."tracing@codex-observability-plugin:hooks/hooks.json:stop:0:0"]
enabled = true

Set your Langfuse credentials

Tracing stays off until TRACE_TO_LANGFUSE is "true", so you opt in explicitly. Add your credentials to your shell profile (~/.zshrc, ~/.bashrc, or ~/.bash_profile):

export TRACE_TO_LANGFUSE="true"
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com" # 🇪🇺 EU region

Alternatively, create a JSON config file at ~/.codex/langfuse.json (global) or <project>/.codex/langfuse.json (per-project):

{
  "enabled": true,
  "public_key": "pk-lf-...",
  "secret_key": "sk-lf-...",
  "base_url": "https://cloud.langfuse.com"
}

Configuration is resolved as defaults, then the global config file, then the project config file, then environment variables, with environment variables taking precedence. LANGFUSE_CODEX_* variables override the matching standard LANGFUSE_* variables, so you can scope credentials to Codex.

Restart Codex and verify tracing

Fully restart Codex after changing its global configuration, then start a new Codex session. Global configuration applies to new sessions in every project; existing sessions do not load the hook retroactively.

For a reliable test, send two short messages. The Stop hook uploads each completed turn, while the latest turn is finalized on the next hook invocation.

Run Codex as usual:

cd your-project
codex

View traces in Langfuse

Open your Langfuse project to see the captured traces. Search for Codex Turn and widen the time range if needed; Langfuse timestamps may be displayed in UTC. The structure mirrors how Codex actually works:

  • Turn trace (Codex Turn): one trace per turn, from your prompt to the final answer, captured as an agent observation.
  • Generations (LLM): one per model response in the turn. Each shows the input it received, including the system prompt, the model's reasoning and text, the tool calls it requested, and token usage.
  • Tool spans (exec_command, apply_patch, spawn_agent, …): nested under the generation that triggered them, with input, output, and error status. Failed commands are flagged as errors. MCP tools appear as server.tool, and a command that loads a skill appears as skill:<name>.
  • Subagents (Codex Subagent Turn, with LLM Subagent generations): subagent threads are nested under the spawning turn so you can follow parallel work in one place.
  • Sessions: all turns from the same Codex session are grouped via session_id. Open the Sessions tab to replay the full run.

Environment variables

VariableDescriptionRequired
TRACE_TO_LANGFUSESet to "true" to enable tracingYes
LANGFUSE_PUBLIC_KEYYour Langfuse public key (pk-lf-...)Yes
LANGFUSE_SECRET_KEYYour Langfuse secret key (sk-lf-...)Yes
LANGFUSE_BASE_URLLangfuse host. EU: https://cloud.langfuse.com, US: https://us.cloud.langfuse.com, Japan: https://jp.cloud.langfuse.com, HIPAA: https://hipaa.cloud.langfuse.comNo (defaults to EU)
LANGFUSE_TRACING_ENVIRONMENTEnvironment label for the traces (e.g. production)No
LANGFUSE_CODEX_USER_IDUser attached to every trace, shown as the user in Langfuse. Defaults to the Codex auth email, if one is found.No
LANGFUSE_CODEX_TAGSTags for all traces (JSON array or comma-separated)No
LANGFUSE_CODEX_METADATAJSON object of metadata to attach to all tracesNo
LANGFUSE_CODEX_SKILL_TAGSTag traces with skill:<name> for every skill invoked in the turn (default true)No
LANGFUSE_CODEX_TRACE_SEEDSeed that makes trace IDs predictable, so a caller can derive a run's trace ID before the trace exists. Use a unique seed per session, otherwise sessions collide.No
LANGFUSE_CODEX_DEBUGSet to "true" for verbose logging to stderrNo
LANGFUSE_CODEX_FAIL_ON_ERRORSet to "true" to make upload errors fail the hook instead of failing open. Useful together with debug logging while testing.No

The credential variables also accept a LANGFUSE_CODEX_ prefix (for example LANGFUSE_CODEX_PUBLIC_KEY), which takes precedence over the standard variable.

Every variable has an equivalent key in langfuse.json, using the lower-case name without the prefix. For example LANGFUSE_CODEX_SKILL_TAGS becomes "skill_tags", and TRACE_TO_LANGFUSE becomes "enabled".

Troubleshooting

No traces appearing in Langfuse

  1. The plugin isn't installed or enabled. Run codex plugin add tracing@codex-observability-plugin, then confirm hooks = true under [features] and that the tracing@codex-observability-plugin plugin is enabled in ~/.codex/config.toml. Run codex plugin list and codex features list to check both.
  2. The Stop hook is disabled. Approve the Langfuse Stop hook when Codex prompts you. If the hook has already been trusted, verify that its generated entry under [hooks.state] has enabled = true. A plugin update changes the hook hash and can require a fresh review.
  3. Tracing isn't turned on. TRACE_TO_LANGFUSE must be the exact string "true" and visible to the Codex process, unless you enabled tracing in ~/.codex/langfuse.json. Also verify the public key starts with pk-lf-.
  4. Restart and test again. Fully restart Codex, start a new session, and send two short messages before checking Langfuse.
  5. Enable debug logging. Set LANGFUSE_CODEX_DEBUG=true to log to stderr and surface the actual cause. Add LANGFUSE_CODEX_FAIL_ON_ERROR=true to turn a silent upload failure into a visible hook error.

Authentication errors

Verify your API keys are correct and that LANGFUSE_BASE_URL matches the region your keys belong to:

  • EU region: https://cloud.langfuse.com
  • US region: https://us.cloud.langfuse.com
  • Japan region: https://jp.cloud.langfuse.com
  • HIPAA region: https://hipaa.cloud.langfuse.com

Data privacy

When enabled, the plugin uploads completed Codex transcript data to Langfuse, including prompts, attached images, assistant messages, reasoning summaries, system prompts, tool inputs and outputs, model metadata, and token usage. This content is captured in full and is not truncated, so do not enable tracing for sessions containing data you do not want stored in Langfuse.

Resources


Was this page helpful?

Last updated on