Commit Graph

556 Commits

Author SHA1 Message Date
xxxigm a86569bd11 test(teams-pipeline): cover quoted-user Graph @odata.id paths
Lock in users('{id}')/onlineMeetings('{id}') parsing, job creation from that
notification shape, and replay that re-reads a stored transcript id.
2026-08-21 12:15:36 +05:30
Teknium 4ea69d9d2c feat: keyless web tier becomes a 5-vendor round-robin ring (adds Tavily, Firecrawl, Keenable)
Fresh installs with zero web credentials now rotate web_search/
web_extract across FIVE vendors' public free tiers — Exa, Parallel,
Tavily, Firecrawl, Keenable — instead of a 2-vendor 50/50 split, with
next-in-line ring failover on rate limits (multi-hop until a vendor
serves or the ring is exhausted; served_by marks the actual vendor).

- plugins/web/keenable/: new bundled provider (search via /v1/search,
  fetch via /v1/fetch; keyed Bearer or keyless with the mandatory
  X-Keenable-Title app header). Credit: integration proposed by
  Ilya Gusev (Keenable) in #49758; Free/Paid picker rows included.
- keyless_mcp: tavily/firecrawl/keenable keyless search+extract
  wrappers, _KEYLESS_RING + per-process round-robin cursor (seeded by
  the random session id, advances per unpinned request), pinned-vendor
  entry (pin = start there; rotation off), paid-pinned vendors excluded
  from the ring entirely.
- Tavily/Firecrawl providers route keyless traffic through the ring;
  both are now default-on ring members (no longer selection-gated).
- web_tools/registry: keenable in backend sets, auto-detect, availability
  probes; _keyless_preference() delegates to the ring cursor.
- KEENABLE_API_KEY in OPTIONAL_ENV_VARS; docs updated (ring semantics).

Live E2E: all 10 vendorXcapability paths (5 search + 5 extract) served
real results keyless; rotation cycled all five vendors over 5 dispatch
calls; double-throttle failover walked exa->parallel->tavily.
2026-08-20 00:17:25 -07:00
LeonSGP43 f51e61136a fix(web): explicit Firecrawl selection works keyless against the public cloud API
Salvaged from #50659 by @LeonSGP43 onto current main (the client
resolver was rewritten for strict-selection semantics since the PR;
reapplied the keyless mode as a third client_mode inside the new
resolver). An explicit firecrawl selection with no FIRECRAWL_API_KEY /
FIRECRAWL_API_URL now routes through a minimal REST client (v2 search +
scrape, no Authorization header) instead of erroring. Unconfigured
installs never route here — the keyless path requires the explicit
selection. Fixes #49912.
2026-08-19 22:54:28 -07:00
Gille cefeed4ca8 fix(a2a): expose schemas through tool describe 2026-08-20 10:41:12 +05:30
Brooklyn Nicholson b4f978d983 fix(nous): treat "takes no reasoning parameter" as a definitive no
Both the wire path and the picker only consulted the catalog's
`mandatory` flag, so a route the Portal lists as accepting no reasoning
parameter at all still got sent a disable, and still offered a Thinking
toggle in the model picker.

For a route it serves, the aggregator's own catalog outranks the
models.dev inference: `supports_reasoning: false` now suppresses the
disable on the wire and drops reasoning controls from the picker
entirely, so there is no disable left to describe.
2026-08-19 23:28:14 -05:00
Brooklyn Nicholson d39a031329 fix(nous): stop dropping "thinking off" on Portal models that can honor it
reasoning: {enabled: false} is the only shape the Portal honors, and the
profile refused to send it for every model. Sending nothing means the
upstream default instead, which on a thinking-first route like
deepseek/deepseek-v4-pro (catalog: default_effort high) is thinking ON — so
turning thinking off kept billing reasoning tokens on every turn.

The blanket omission was over-broad. The Portal only rejects a disable on
reasoning-mandatory routes ("Reasoning is mandatory for this model"), which
its catalog flags per model, so that flag now gates the omission. Models the
catalog can't speak to keep the old behavior rather than risk the 400.

extra_body.thinking, DeepSeek's own disable shape, is not forwarded upstream
by the Portal and is not an option here.
2026-08-19 22:14:56 -05:00
Teknium aebab05f9e Merge pull request #90313 from NousResearch/feat/keyless-web-search-fallback
feat: web search works keyless on fresh installs (Parallel + Exa free tiers)
2026-08-19 19:47:12 -07:00
Teknium 1fa66f2577 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	website/docs/user-guide/configuration.md
2026-08-19 19:36:12 -07:00
Axmr1 4511ba49dd fix(image_gen/openai-codex): do not save progressive partial frames as finals
Codex Responses streams can emit partial_image_b64 previews without a final
image_generation_call.result. The provider treated any b64 as success and could
let a partial overwrite a coexisting final in the same payload, delivering
smeared intermediates as finished GPT Image 2 outputs.

Request partial_images=0, prefer final over partial in extraction, fail closed
(with one content-agnostic retry) unless source=final, and surface image_source
plus pixel_size for QA.
2026-08-19 19:35:07 -07:00
Teknium f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Jeffrey Quesnelle 612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
Teknium 08b7fad3a5 test: registry zero-credential resolution may return a keyless-capable provider
test_no_config_no_credentials_returns_none pinned 'resolved provider
must be is_available()' — stale now that the keyless tier resolves
Parallel/Exa with is_available()=False + is_keyless_available()=True.
Accept keyed OR keyless-capable results (env-leak detection intact).
2026-08-19 15:38:08 -07:00
Teknium aba96d5251 feat(image-gen): route live-catalog models to the Image API; merge picker catalogs; docs
Follow-ups on top of the salvaged #82631 surface:

- _select_surface: an unknown model id found in the live /images/models
  catalog now ROUTES to the dedicated Image API instead of only logging a
  hint — without this, a model picked from the live picker that postdates
  the curated snapshot would fall onto chat-completions and fail. Curated
  defaults stay pinned to chat (no behaviour change for existing setups);
  offline probes still fall back to chat. _HINTED_MODELS removed.
- list_models (OpenRouter): union of the live GET /images/models catalog
  (43 models today) and the chat-completions image models, deduped,
  defaults first; curated metadata wins for known ids, API names for the
  rest. Nous Portal (no /images route) keeps its chat-only catalog.
  Offline fallback: static chain + curated Image API snapshot.
- Tests updated/added: unknown-id routing (flipped from the hint-only
  pinning test), non-catalog id stays on chat, merged-picker union/dedupe/
  order, Nous exclusion.
- Docs: image-generation.md gains the OpenRouter Image API section and an
  editing-support row.

Live-verified: picker lists 43 models; generation succeeded through the
dedicated API on google/gemini-3.1-flash-lite-image and on the previously
unreachable black-forest-labs/flux.2-klein-4b (config-selected, no kwarg).
2026-08-19 14:44:36 -07:00
AI Staff d6e6e8b602 feat(plugins): add OpenRouter Image API surface to openrouter image_gen backend 2026-08-19 14:44:36 -07:00
Alex Fournier 31402f630b fix(relay): complete native plugin cutover
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski 88300217c2 feat(relay): activate configured dynamic plugins
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Teknium 9c6ecf2ca7 feat(image-gen): OpenRouter image picker lists every live image-output model; xAI edits honor dispatched model
- plugins/image_gen/openrouter: list_models() now queries the endpoint's
  /models catalog filtered to output_modalities containing "image"
  (per-backend 5-min cache, 10s timeout, static 2-model chain as offline
  fallback; openrouter/auto* router pseudo-models excluded). Every image
  model OpenRouter serves — including future releases — is selectable in
  `hermes tools` with no code change. Applies to Nous Portal too via the
  shared provider class.
- plugins/image_gen/xai: forward the dispatched model kwarg into
  _resolve_edit_model() so an explicitly selected edit-capable model is
  honored on /images/edits (extends the salvaged #55893 fix to the edit
  path; text-only models still fall back to quality).
- Tests: OpenRouter live-catalog filtering/exclusions/order, offline
  fallback, cache single-fetch; xAI edit-kwarg forwarding incl. the
  text-only-hijack negative case.

Live-verified against openrouter.ai: 9 image-output models returned and
rendered, matching the public models?output_modalities=image listing.
2026-08-19 02:04:29 -07:00
srojk34 008d469991 fix(xai): forward image_gen.model kwarg to _resolve_model in generate() 2026-08-19 02:04:29 -07:00
Teknium ac0a8cd281 feat(image-gen): xAI Grok image catalog goes live-driven; grok-imagine-image-2.0 selectable
- plugins/image_gen/xai: merge the live /v1/image-generation-models catalog
  (5-min cache, 10s timeout, static-table fallback when offline/unauth)
  into the picker so new xAI Imagine models appear automatically the day
  they launch, with generic metadata until curated text is added.
- Add grok-imagine-image-2.0 to the curated static table (typography/
  layout-aware model, API-available since Aug 8 2026).
- Edits honor an explicitly selected image-input-capable model
  (e.g. grok-imagine-image-2.0) instead of always forcing
  grok-imagine-image-quality; quality remains the default edit baseline.
- Tests: hermetic autouse fixture keeps unit runs offline; new coverage
  for live-merge, unknown-future-model selection, offline fallback, and
  edit-model resolution. Docs model table updated (en + zh-Hans).

Live-verified: /image-generation-models returns grok-imagine-image,
grok-imagine-image-2.0, grok-imagine-image-quality; real generation with
2.0 succeeded end to end.
2026-08-19 01:19:37 -07:00
zhuermu c6b680c445 fix(image-gen): handle stale OpenRouter model defaults 2026-08-19 01:19:37 -07:00
Teknium c7d0f6c35f fix(providers): honor a custom base_url over models_url in fetch_models
Follow-up to the salvaged CommandCode signature fix: accepting base_url
but ignoring it left custom endpoints (user-configured model.base_url /
COMMANDCODE_BASE_URL proxies) fetching the public catalog instead of the
configured one. Reviewer dansigma flagged this on PR #88851.

Class-wide fix, not a CommandCode patch:

- providers/base.py: a caller base_url that DIFFERS from the profile's
  default now wins over models_url. Equality with the default means "not
  customised" (callers pass base_url unconditionally, defaulting to the
  profile's own URL) and keeps models_url as the endpoint, preserving the
  OpenRouter-style split-catalog behavior.
- commandcode: _fetch_commandcode_models() takes the endpoint override;
  both profile overrides forward base_url.
- Tests: base-class precedence (custom beats models_url, default does
  not), CommandCode redirect via live local HTTP server incl. claude-*
  filter, and default-echo hitting the canonical endpoint. All verified
  to fail against the pre-fix implementation (sabotage run).
2026-08-18 14:27:36 -07:00
greyvito 7072fc4f87 fix(providers): resolve plugin-registered provider profiles in get_provider
Plugin-only providers (commandcode, tencent-tokenhub, ...) are absent from
models.dev and HERMES_OVERLAYS, so resolve_provider_full returned None and
/model switches failed with "Unknown provider ..." even though the picker
lists them (CANONICAL_PROVIDERS auto-extends from the same registry).

Fall back to providers.get_provider_profile() before giving up, mapping the
profile api_mode to the ProviderDef transport.
2026-08-18 14:27:36 -07:00
greyvito f5ea3fa9cb fix(commandcode): accept base_url kwarg in fetch_models overrides
The model picker's generic live-fetch path (hermes_cli/models.py
provider_model_ids) calls profile.fetch_models(api_key=..., base_url=...).
Both CommandCode overrides only accepted api_key/timeout, so every picker
open raised TypeError, which was silently swallowed, leaving the provider
with zero models.

Match the base ProviderProfile.fetch_models signature (base_url kwarg) and
add a regression test asserting both profiles accept it.
2026-08-18 14:27:36 -07:00
Teknium d03fe2adf4 fix: detect base64-encoded transcript markers in Graph ids
The getAllTranscripts resourceData.id from the field report is a base64url
blob whose DECODED payload ends in "-TranscriptV2" while the encoded form
contains no readable marker, so the substring heuristic in
looks_like_transcript_id missed it. Add a best-effort base64 decode hint so
degraded notifications (no @odata.id) are still refused with the clear
guidance error instead of a cryptic Graph 400.
2026-08-18 12:37:40 -07:00
kyssta-exe db004d1801 fix(teams-pipeline): support organizer-scoped meeting lookup (#83422) 2026-08-18 12:37:40 -07:00
xxxigm 79c5380f41 test(teams-pipeline): cover getAllTranscripts meeting-id parsing
Pin that transcript notifications keep the onlineMeeting id from @odata.id, and that meeting GET uses the organizer-scoped Graph path.
2026-08-18 12:37:40 -07:00
Teknium aa8ceed4b6 fix(memory): keep a stale holder's late close() from evicting a fresh registry entry
Follow-up to the #88347 salvage: after release_all_under() force-closes a
profile's shared connection, a store re-created on the same path registers
a fresh entry under the same key. A stale holder that later calls close()
would pop that fresh entry (its refs were transferred nowhere), letting a
third store open a second connection to the same database — exactly the
multi-writer contention the shared registry exists to prevent. close()
now evicts the registry entry only when it is still its own.
2026-08-17 16:33:52 -07:00
liuhao1024 4f354c27b7 fix(profiles): release memory-store handles before rmtree on profile delete
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.

MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.

Fixes #88347
2026-08-17 16:33:52 -07:00
Johann 26d8bf567c feat: add CommandCode provider plugin
Add first-class CommandCode provider with dual API mode support:

profile commandcode (chat_completions):
  20+ models via OpenAI-compatible endpoint
  DeepSeek, Qwen, Kimi, GLM, MiniMax, StepFun, Mimo, Gemini, GPT
  Default: deepseek/deepseek-v4-pro (1M context)

profile commandcode-anthropic (anthropic_messages):
  Claude models via Anthropic Messages-compatible endpoint
  Default: claude-sonnet-4-6 (1M context)

Changes:
- plugins/model-providers/commandcode/ — provider plugin
  - __init__.py: dual ProviderProfile classes with fetch_models
  - plugin.yaml: manifest
- agent/anthropic_adapter.py: recognize api.commandcode.ai as Bearer auth
- tests/plugins/model_providers/test_commandcode_profile.py: 28 tests
- tests/providers/test_plugin_discovery.py: bump profile count 34→36

171 provider tests pass (28 new, 0 regressions)
2026-08-17 02:56:17 -07:00
kshitij b202293125 Merge pull request #87604 from ehz0ah/fix/openviking-env-reliability 2026-08-17 12:19:22 +05:30
Teknium f06c41522e fix(image_gen): disable default-on upscaling everywhere — opt-in only
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.

Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.

- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
  now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
2026-08-16 11:12:05 -07:00
ehz0ah 2854fab46c docs(openviking): correct environment handling explanations
Clarify that the Desktop backend can add Hermes venv packages to PYTHONPATH and that current .env loaders use the last duplicate value.
2026-08-17 01:44:09 +08:00
ehz0ah 38175b8c22 fix(openviking): preserve non-UTF-8 env bytes on update 2026-08-17 01:41:49 +08:00
Drexuxux dde7075d6c fix(openviking): read .env BOM-tolerantly when rewriting credentials
f1ea4a56c ("cover the remaining setup-time .env reads with utf-8-sig",
following 75afc47ba for mem0/hindsight) swept this class; openviking's
_write_env_vars was missed and still reads with strict utf-8.

It copies every existing line through on each update, so the read decides
whether a credential update lands:

  BOM'd .env  -> the first key never matches, so the old line survives and
                 the new value is appended as a duplicate. .env loaders keep
                 the first occurrence, so the update silently does nothing.
  cp1252 .env -> UnicodeDecodeError aborts setup outright.

Read exactly like the canonical hermes_cli/config.py save_env_value
(utf-8-sig + errors="replace"). A plain UTF-8 file rewrites byte-identically.

Scope: hermes_cli/memory_setup.py has the same read but is already the
subject of #30281 / #60587, so it is left alone here.

(cherry picked from commit 175c6852c2c255b3219575b5de0b1b70f1f0efcb)
2026-08-17 01:41:49 +08:00
kyssta-exe 4fdaadd907 fix(openviking): strip PYTHONPATH from autostarted server child env (#78153)
(cherry picked from commit 7afd99155667cde480c0ab4ee31e242dab849d40)
2026-08-17 01:41:49 +08:00
Teknium fe0a56ed16 fix(nemo_relay): bound plugin Relay marks so a wedged native pipeline cannot stall the agent
The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".

The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.

Changes:
- plugins/observability/nemo_relay: the wrapper always passes
  timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
  costs one telemetry span, never the agent; it also sets scope_errored
  (so close_session skips the ATIF export for the wedged session) and
  warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
  reaches the host (fails on the pre-fix code — sabotage-verified),
  a TimeoutError flags the session and disables its export, and the
  generic error path keeps its scope_errored contract.
2026-08-15 14:31:32 -07:00
David Metcalfe 08f32a6335 fix(kanban): replace native browser dialogs with in-app ConfirmDialog
Migrates 8 of 12 native dialog call sites in the kanban dashboard plugin
to the SDK's ConfirmDialog primitive (added in PR #50550):
  - moveTask, moveSelected, applyBulk, deleteTask, deleteSelected,
    archiveBoard, removeAttachment, doPatch

The 4 remaining carve-outs (window.prompt for completion summary,
window.alert for missing summary, cli_hint clipboard fallback) are
documented inline — the host's ConfirmDialog hardcodes onClick → unmount,
preventing the keep-open-across-validation behavior the completion-summary
form needs. Followup: upstream a `disabled` prop to ConfirmDialog and
rebuild the completion body using host Dialog components.

New architecture:
  - useKanbanDialogs(t) — Promise-based dialog state machine at
    KanbanPage scope. request({kind, ...}) returns {confirmed, summary?}.
  - KanbanDialog component — renders ConfirmDialog from SDK for kind=confirm.
  - performMoveTask(taskId, newStatus, count, summary) — extracted shared
    dispatch path for single + bulk moves (optimistic UI + PATCH/POST +
    error recovery).
  - requestDialog prop threading — KanbanPage → BoardSwitcher,
    TaskDrawer → TaskDetail → doPatch/AttachmentsSection. Every call
    site has a defensive fallback to window.confirm if the prop is
    missing (verified by test_dashboard_done_actions_prompt_for_completion_summary
    counting the cancel guards + destructive:true markers in the bundle).

New host i18n keys (web/src/i18n/en.ts + types.ts):
  - kanban.confirmDoneMany / confirmArchiveMany / confirmBlockedMany
  - kanban.trash.confirmTitle / confirmManyTitle

Tests:
  - Replaced bundle-string-only completion-summary test with behavioral
    coverage: bundle cancel-guard count + destructive marker count, plus
    backend tests that confirm cancel preserves old status and confirm
    dispatches the expected PATCH/DELETE body.
  - Removed the SDK_CONTRACT_VERSION snapshot test from
    web/src/plugins/registry.test.ts (forbidden by AGENTS.md
    "Don't write change-detector tests"; the two remaining tests in that
    file already cover the new SDK surface behaviorally).

Closes #50547 (consumers of #50550).

Cross-vendor re-review: Gemini 3.5 Flash + GPT-OSS 120B (both SHOULD-FIX,
no remaining BLOCKERs after these fixes).
2026-08-15 00:33:32 -07:00
Teknium edd73daaf4 test(kanban): regression for idle-board WS disconnect detection (#77833) 2026-08-15 00:32:53 -07:00
Tachi d1df111ccd fix(update): restore Hermes Tools dependencies 2026-08-14 22:33:44 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Greg Gibeau c600fd46bd fix(memory): complete discovery and registration parity for out-of-tree providers
Builds on the three salvaged commits: adds the sources and integration points
they leave out, so a pip-installed memory provider is not a second-class
citizen next to a directory install.

Discovery
- Project-local providers (./.hermes/plugins/<name>/), gated on
  HERMES_ENABLE_PROJECT_PLUGINS exactly as PluginManager gates its own project
  scan. Completes the four sources CONTRIBUTING.md and AGENTS.md already
  promised; memory was the only discovery system missing two of them.
- find_provider_dir() now resolves a package entry point to its directory.
  This is load-bearing: config_schema.py (the dashboard panel) and cli.py (the
  `hermes <provider>` subcommands) are read from disk rather than imported, so
  without a directory a pip-installed provider silently lost both.
- list_memory_provider_names() includes entry-point providers, so they appear
  in the dashboard's memory.provider dropdown.

Resolution stays import-free. hermes_cli.plugins.resolve_module_origin() is
extracted from _resolve_module_source() (added by the salvaged #76567) and
shared, so discovery walks a module's file layout instead of importing it.
find_provider_dir() is called from the dashboard and from argparse setup, long
before the operator has chosen a provider — importing every installed candidate
would execute third-party code on the strength of a package being present.
A test asserts the resolution leaves no side effects and no sys.modules entry.

Registration
- PluginContext gains register_memory_provider(). Memory was the only provider
  category without one; context engine, image gen, video gen, web search,
  browser, TTS, transcription, secret source, dashboard auth and platform all
  have one.
- _ProviderCollector delegates unknown register_* calls to a real
  PluginContext instead of carrying three hand-written no-ops. It silently
  dropped register_tool/register_hook, and had no register_auxiliary_task at
  all — despite PluginContext.register_auxiliary_task documenting a memory
  provider (hindsight's pre-retain dedup) as its worked example. It can no
  longer drift behind PluginContext.
- A raise after register_memory_provider() no longer costs the provider. The
  loader caught it into a debug log, discarded the registered instance, and
  fell through to "instantiate any MemoryProvider subclass" — returning a
  different, unconfigured provider. A silent downgrade that looked like
  success, and the exact outcome of calling register_auxiliary_task.

Activation is unchanged: still gated on memory.provider naming the plugin, and
covered by a test so the real PluginContext cannot start requiring
plugins.enabled — that would break every existing user-installed provider.

Verified end to end against a real third-party provider (kainappsinc/elephant)
installed by pip alone, with no directory copy: it appears in the dropdown,
resolves its directory, loads with its tools, and renders its dashboard panel.

Closes #40101.
2026-08-13 11:49:14 -07:00
spfcraze 32238f9942 fix(honcho): resolve peers host keys via profile_host_key (underscore form) (#76414)
_all_profile_host_configs() built per-profile host keys inline as
f"{HOST}.{profile}" ("hermes.work") while profile_host_key() — used by
honcho status/enable/sync and the runtime memory plugin — produces the
underscore form ("hermes_work"). The lookup always missed, so
'hermes honcho peers' showed "(not set)" / leaked the raw malformed key
into the AI-peer column for every non-default profile. Profile names
needing sanitization (dots/spaces) were doubly broken.

Verified live: with hosts["hermes_work"] populated, cmd_peers showed
'work ... hermes.work' before the fix and 'work ... hermes' after.

Tests: host keys match the writer form, sanitized profile names resolve,
peers output shows populated identities with no key leak, and clean
fallback for profiles without a block.
2026-08-13 23:43:15 +05:30
kshitij 8b243dff62 fix: security + efficiency review fixes for salvaged PR #74379
1. Use open_credentialed_url() instead of bare urlopen() in
   templates.py apply_template() and probe_existing_customization().
   Both send Authorization: Bearer headers; bare urlopen forwards
   credentials on cross-origin redirects. The codebase has
   open_credentialed_url() in hermes_cli/urllib_security.py that
   strips credentials on cross-origin redirects — used by 4 other
   modules.

2. Guard unavailable_reason() with the dedup set check before
   calling it. The gateway builds a fresh AIAgent per message, so
   without this guard unavailable_reason() (which calls _load_config()
   → stat + file read + JSON parse, and _check_local_runtime() →
   importlib probes) runs on every gateway turn for an unavailable
   provider, even though the warning is deduped after the first.

3. Move INDICATOR_GLYPH from Hindsight's eye emoji to a generic
   brain (🧠) in core (agent/memory_provider.py). Hindsight overrides
   with its own _HINDSIGHT_GLYPH (👁️) in recall_status() and
   _emit_saving_indicator(). Other memory providers no longer inherit
   Hindsight's brand mark as the default glyph.
2026-08-13 23:15:25 +05:30
Ben 34c727c5c2 feat(hindsight): memory provider improvements — recall_sync, retain_source, setup templates, memory indicators, error hints
Bundles previously-separate Hindsight/memory PRs into a single review surface:
- opt-in synchronous recall (recall_sync) — recall the injected memory in-turn instead of next-turn prefetch (#5820)
- actionable error when local_embedded runtime is missing — tells the user which package to install (#7718)
- default retain_source to 'hermes' so every stored memory self-identifies its provenance
- offer a starter memory template during hermes memory setup, plus warn before overwriting an already-configured bank
- warn when a configured memory provider reports unavailable (#2765)
- deterministic 'recalled N memories' recall indicator — Hermes itself emits a status line when auto-recall injects memory
- 'saving to memory' retain indicator — emitted the moment a turn is dispatched to the writer

Authored by @benfrank241 (ben.bartholomew@vectorize.io).
Salvaged from PR #74379.
2026-08-13 23:15:25 +05:30
kshitij ace830134e fix: reuse redact_sensitive_text, fix leaky abstraction, fix test data
Follow-up fixes from /hermes-pr-review + /simplify-code on PR #83437:

1. Replace _redact_secrets with agent.redact.redact_sensitive_text(force=True)
   — the plugin's 11-pattern list was a strict subset of the 50+ patterns in
   agent/redact.py. Secrets like Stripe keys, Google API keys, GitLab tokens,
   HuggingFace tokens, DB connection strings, and Telegram bot tokens would
   all leak through the plugin's list but are caught by the existing redactor.
   Added pk-lf- (Langfuse public key) to _PREFIX_PATTERNS in agent/redact.py.

2. Remove dead 'not isinstance(client, object)' check in on_session_finalize —
   always False for any Python value.

3. Fix MoAClient.last_reference_metrics() to call the public
   self.chat.completions.last_reference_metrics() instead of reaching into
   the private _last_reference_metrics attribute via getattr.

4. Deduplicate _coerce_request_messages call in on_pre_llm_request — pass
   pre_coerced=input_messages to _messages_for_langfuse_input to avoid
   double-coercion + double _capture_content serialization per API request.

5. Add HERMES_LANGFUSE_CAPTURE to OPTIONAL_ENV_VARS in hermes_cli/config.py
   for consistency with the other HERMES_LANGFUSE_* env vars.

6. Fix test_sanitized_mode_redacts_secrets test data — the old samples
   ('sk-abc...1234', 'sk-ant...1234', 'Authorization: Bearer ***') were too
   short to match the regex thresholds and never actually tested redaction.
   Updated to realistic-length secrets and changed assertions to check that
   the output differs from input (redact_sensitive_text masks rather than
   inserting the literal string 'REDACTED').
2026-08-13 23:10:16 +05:30
kshitij e665300d6b feat(langfuse): widen tracing to errors, sessions, subagents, and MoA fan-out
Salvaged from PR #83437 by @erosika, with adopted fixes from @bgodlin (#81054),
@aldoeliacim (#82332), @nftpoetrist (#42326), @rodboev (#39653), @FnExpress
(#64292, supersedes #32175 by @db-aeon), @Per0-1 (#61166), @NaMinhyeok (#64797),
and @liuhao1024 (#43130).

Widens the bundled Langfuse plugin from 6 to 11 hooks and fixes two
attribution bugs. Also adopts shutdown/atexit lifecycle fixes and composes
8 prior community PRs with interaction-fix follow-ups.

Model attribution: on_pre_llm_request and on_post_llm_call now prefer the
wire value (request body model, response model) over the agent attribute,
which goes stale after /model switch or provider fallback.

Cost total: both cost paths now send a summed total alongside the per-type
breakdown, since Langfuse does not derive calculatedTotalCost from
cost_details keys. Subscription-included routes send no cost keys at all.

New coverage: api_request_error closes failed generations with ERROR level;
on_session_finalize/on_session_end close dangling traces for tool-only and
interrupted turns; subagent_start/subagent_stop trace delegated children as
spans; MoA advisor fan-out emits one generation per advisor priced at the
advisor's own model.

Capture modes: HERMES_LANGFUSE_CAPTURE=metadata|sanitized|full (default
sanitized). Sanitized mode redacts secret patterns before truncation.

Adopted lifecycle fixes: shutdown client at session finalize when
reason=shutdown (not on session rotation); atexit finalizer ends open root
spans for short-lived processes; root context manager exited to prevent
interpreter-teardown TypeError; TOCTOU on _get_langfuse() fixed with lock;
reasoning_content surfaced in traces; system prompt included in generation
input for Anthropic/Codex/Bedrock; SDK v3 update_trace replaces set_trace_io.

Closes #29482, #43129, #72661.
Supersedes #81054, #82332, #42326, #39653, #64292, #32175, #61166, #64797, #43130.
Partially addresses #67544 (capture modes + secret redaction; user_id remains open).
2026-08-13 23:10:16 +05:30