e665300d6b
Salvaged from PR #83437 by @erosika, with adopted fixes from @bgodlin (#81054), @aldoeliacim (#82332), @nftpoetrist (#42326), @rodboev (#39653), @FnExpress (#64292, supersedes #32175 by @db-aeon), @Per0-1 (#61166), @NaMinhyeok (#64797), and @liuhao1024 (#43130). Widens the bundled Langfuse plugin from 6 to 11 hooks and fixes two attribution bugs. Also adopts shutdown/atexit lifecycle fixes and composes 8 prior community PRs with interaction-fix follow-ups. Model attribution: on_pre_llm_request and on_post_llm_call now prefer the wire value (request body model, response model) over the agent attribute, which goes stale after /model switch or provider fallback. Cost total: both cost paths now send a summed total alongside the per-type breakdown, since Langfuse does not derive calculatedTotalCost from cost_details keys. Subscription-included routes send no cost keys at all. New coverage: api_request_error closes failed generations with ERROR level; on_session_finalize/on_session_end close dangling traces for tool-only and interrupted turns; subagent_start/subagent_stop trace delegated children as spans; MoA advisor fan-out emits one generation per advisor priced at the advisor's own model. Capture modes: HERMES_LANGFUSE_CAPTURE=metadata|sanitized|full (default sanitized). Sanitized mode redacts secret patterns before truncation. Adopted lifecycle fixes: shutdown client at session finalize when reason=shutdown (not on session rotation); atexit finalizer ends open root spans for short-lived processes; root context manager exited to prevent interpreter-teardown TypeError; TOCTOU on _get_langfuse() fixed with lock; reasoning_content surfaced in traces; system prompt included in generation input for Anthropic/Codex/Bedrock; SDK v3 update_trace replaces set_trace_io. Closes #29482, #43129, #72661. Supersedes #81054, #82332, #42326, #39653, #64292, #32175, #61166, #64797, #43130. Partially addresses #67544 (capture modes + secret redaction; user_id remains open).
84 lines
2.9 KiB
Markdown
84 lines
2.9 KiB
Markdown
# Langfuse Observability Plugin
|
|
|
|
This plugin ships bundled with Hermes but is **opt-in** — it only loads when
|
|
you explicitly enable it.
|
|
|
|
## Enable
|
|
|
|
Pick one:
|
|
|
|
```bash
|
|
# Interactive: walks you through credentials + SDK install + enable
|
|
hermes tools # → Langfuse Observability
|
|
|
|
# Manual
|
|
pip install langfuse
|
|
hermes plugins enable observability/langfuse
|
|
```
|
|
|
|
## Required credentials
|
|
|
|
Set these in `~/.hermes/.env` (or via `hermes tools`):
|
|
|
|
```bash
|
|
HERMES_LANGFUSE_PUBLIC_KEY=pk-lf-...
|
|
HERMES_LANGFUSE_SECRET_KEY=sk-lf-...
|
|
HERMES_LANGFUSE_BASE_URL=https://cloud.langfuse.com # or your self-hosted URL
|
|
```
|
|
|
|
Without the SDK or credentials the hooks no-op silently — the plugin fails
|
|
open.
|
|
|
|
## Verify
|
|
|
|
```bash
|
|
hermes plugins list # observability/langfuse should show "enabled"
|
|
hermes chat -q "hello" # then check Langfuse for a "Hermes turn" trace
|
|
```
|
|
|
|
Generation observations include the Hermes system prompt when the provider
|
|
uses a separate `system` param (Anthropic Messages API). Open an **LLM call**
|
|
child span to inspect `role: system` (truncated via `HERMES_LANGFUSE_MAX_CHARS`).
|
|
|
|
## Optional tuning
|
|
|
|
```bash
|
|
HERMES_LANGFUSE_ENV=production # environment tag
|
|
HERMES_LANGFUSE_RELEASE=v1.0.0 # release tag
|
|
HERMES_LANGFUSE_SAMPLE_RATE=0.5 # sample 50% of traces
|
|
HERMES_LANGFUSE_MAX_CHARS=12000 # max chars per field (default: 12000)
|
|
HERMES_LANGFUSE_CAPTURE=sanitized # content capture mode (see below)
|
|
HERMES_LANGFUSE_DEBUG=true # verbose plugin logging
|
|
```
|
|
|
|
## Capture modes
|
|
|
|
`HERMES_LANGFUSE_CAPTURE` controls how much *content* (prompts, responses,
|
|
tool arguments/results) is exported. Structural metadata — IDs, roles, tool
|
|
names, token usage, cost, timing — is always captured in every mode.
|
|
|
|
| mode | behavior |
|
|
|------|----------|
|
|
| `metadata` | No content. Each content field is replaced by a shape/size stub (`{"omitted": true, "type": "text", "chars": N}`). |
|
|
| `sanitized` | **(default)** Content is exported after secret-pattern redaction (API keys, tokens, JWTs, private keys, `password=`-style assignments) and truncation. Redaction runs *before* truncation. |
|
|
| `full` | Raw content, truncated only. Explicit opt-in — traces will contain whatever passed through the conversation, including injected memory and file contents. |
|
|
|
|
The active mode is recorded on every trace as `metadata.capture_mode`.
|
|
|
|
Note: `sanitized` is pattern-based defense in depth, not a DLP guarantee.
|
|
For personal sessions or shared Langfuse projects, prefer `metadata`.
|
|
|
|
## Error + shutdown coverage
|
|
|
|
- Failed model requests (`api_request_error` hook) close their generation
|
|
with `level=ERROR`, status code, retry counters, and a capture-mode-scrubbed
|
|
error message. Non-retryable failures also finish the turn trace.
|
|
- Session end/finalize closes any still-open traces for that session and
|
|
flushes queued events, so interrupted or tool-only turns don't dangle.
|
|
|
|
## Disable
|
|
|
|
```bash
|
|
hermes plugins disable observability/langfuse
|
|
```
|