- plugins/image_gen/openrouter: list_models() now queries the endpoint's
/models catalog filtered to output_modalities containing "image"
(per-backend 5-min cache, 10s timeout, static 2-model chain as offline
fallback; openrouter/auto* router pseudo-models excluded). Every image
model OpenRouter serves — including future releases — is selectable in
`hermes tools` with no code change. Applies to Nous Portal too via the
shared provider class.
- plugins/image_gen/xai: forward the dispatched model kwarg into
_resolve_edit_model() so an explicitly selected edit-capable model is
honored on /images/edits (extends the salvaged #55893 fix to the edit
path; text-only models still fall back to quality).
- Tests: OpenRouter live-catalog filtering/exclusions/order, offline
fallback, cache single-fetch; xAI edit-kwarg forwarding incl. the
text-only-hijack negative case.
Live-verified against openrouter.ai: 9 image-output models returned and
rendered, matching the public models?output_modalities=image listing.
Adapter ingress derives a session key BEFORE the runner stamps
source.profile in _make_profile_message_handler, so the namespace fell
back to the active profile and every bot in a multiplexed gateway
produced agent:main:<platform>:<chat>. A Telegram private chat reports
the user's own id as chat.id, identical for every bot, so two profiles
sharing one human collapsed onto a single lane: _pending_text_batches,
_active_sessions, the busy-session guard and _post_delivery_callbacks are
all keyed on that string. A day of production logs across two bots shows
60 flushes, none carrying the secondary profile's namespace.
set_owner_profile records credential ownership on the adapter and
_session_key_profile resolves the namespace as source.profile ->
_owner_profile -> the session store's resolver, so a secondary adapter
keys into its own namespace even before the source is stamped. Stamped
sources keep priority, so relay/connector ingress, which routes per event
rather than per credential, is unchanged. _configure_profile_adapter
installs the owner alongside the other handlers, covering startup and
reconnect.
Every candidate is type-checked as a non-blank str, and every attribute
read goes through getattr: adapters are routinely built without
BasePlatformAdapter.__init__, and a duck-typed session store returns a
truthy non-string that would otherwise be interpolated into the key as
agent:<MagicMock ...>:.
Also routes the four call sites that passed no profile at all (feishu
media batches, raft, slack _session_key_for_source, telegram photo
batches) through the same resolver.
test_multiplex_busy_input_mode's secondary-adapter busy case seeded
_active_sessions with the unstamped agent:main: key, asserting the
pre-fix collapse. It now seeds the lane the profile-owned adapter
actually derives.
A primary adapter has no owner and an unstamped source, so it resolves
exactly as before; with multiplex_profiles off the resolver returns None
and every key is byte-identical to today's.
- plugins/image_gen/xai: merge the live /v1/image-generation-models catalog
(5-min cache, 10s timeout, static-table fallback when offline/unauth)
into the picker so new xAI Imagine models appear automatically the day
they launch, with generic metadata until curated text is added.
- Add grok-imagine-image-2.0 to the curated static table (typography/
layout-aware model, API-available since Aug 8 2026).
- Edits honor an explicitly selected image-input-capable model
(e.g. grok-imagine-image-2.0) instead of always forcing
grok-imagine-image-quality; quality remains the default edit baseline.
- Tests: hermetic autouse fixture keeps unit runs offline; new coverage
for live-merge, unknown-future-model selection, offline fallback, and
edit-model resolution. Docs model table updated (en + zh-Hans).
Live-verified: /image-generation-models returns grok-imagine-image,
grok-imagine-image-2.0, grok-imagine-image-quality; real generation with
2.0 succeeded end to end.
Three local-environment leaks made tests red locally while green on CI:
- tests/conftest.py: blank HERMES_REAL_HOME and TERMINAL_HOME_MODE per
test. The terminal tool injects both into subprocess envs, so any
pytest run launched from a Hermes session inherits them and the
hermes_constants home-resolution helpers prefer HERMES_REAL_HOME over
the monkeypatched HOME (4 failures in
test_subprocess_home_isolation.py).
- test_modal_sandbox_fixes.py: reset the import-time _YOLO_MODE_FROZEN
flag and pin approval mode to manual in _isolate_approval_state().
HERMES_YOLO_MODE=1 in the launching shell froze True at collection
time and every guard auto-approved (2 failures).
- test_noninteractive_git.py: strip GIT_ASKPASS/VS Code askpass vars in
the fail-fast clone E2E. noninteractive_git_env() intentionally keeps
a working askpass helper, but this test asserts the no-helper path;
under VS Code the helper blocks on the editor until the 30s timeout
(1 failure).
Verified: all 49 tests in the three files pass both in a plain dev
shell (with HERMES_YOLO_MODE=1, HERMES_REAL_HOME, and VS Code askpass
set) and inside an unshare -rn network namespace.
hermes update treated a failed desktop pack as non-fatal and still printed
✓ Update complete!, so Windows users kept running an old Hermes.exe after a
"successful" update. Withhold the success banner, surface the stale app in
the summary, and write .update_exit_code=1 for gateway watchers.
Supersedes #88359, #87984.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Prove apply_layout and read_window_below stay direct when MCP tools
activate the bridge, and that a HUD turn still gets its surface note.
Co-authored-by: fangliquan <fangliquan@qq.com>
_is_other_profile only allowed empty/current, so a ?profile=default
save skipped the session.info broadcast on the process whose config
it just wrote. Compare the resolved target to the process HERMES_HOME.
Ten behavior tests for target discovery, stable-selector ordering, step
paging, recovery hints, and the self-containment contract the preview
injection depends on. Docs cover data-tour markup and the curated-tour
entry point alongside the tool itself.
The subprocess-based exclusivity test invoked hermes_cli.main in the CI
environment where startup exits 1 before the manual guard runs. Enforce
the conflict in argparse itself (mutually exclusive group, exit 2 at parse
time) and test the parser directly; the manual guard stays for programmatic
namespace fills.
The Bot Mode teammate-DM protocol told agents to inline the message into a
double-quoted shell argument: quotes truncated the body and $(...)/backticks
executed on the sender's machine. The protocol now writes the message to a
temp file and delivers it via a new 'hermes chat --query-file' flag (or '-'
for stdin); 'hermes peer dm' already accepted stdin and the peer recipe now
uses it. No shell pass touches the body at any point.
Supersedes the tool-based approach in #89077 — same bug, fixed with a CLI
flag + protocol rewrite instead of a new model tool.
Co-authored-by: mehmetkr-31 <mehmetkr-31@users.noreply.github.com>
Some authorization servers and WAFs reject httpx's default User-Agent on
the OAuth token endpoint. mcp_servers.<name>.oauth.user_agent now stamps a
custom User-Agent onto the two token-endpoint requests (authorization-code
exchange and refresh) on both provider construction paths. Opt-in,
per-server, token requests only — never MCP traffic or discovery, and no
other headers are configurable. Empty/null/non-string values are ignored.
Completes the second half of #75576 (the CIMD half landed via #89566).
The _MAX_RESERVED_SOCKETS cap applied to pinned CIMD sockets too, so under
heavy concurrency an ephemeral-reservation churn could close a parked pinned
socket before _wait_for_callback adopted it, silently reopening the
port-stealing window the pin exists to prevent (#22161). Eviction now skips
the pinned range; it is already bounded by _CIMD_PORTS.
Follow-up to the #84050 salvage.
After a standing goal auto-paused on turn-budget exhaustion, every
surface's /goal resume handler only flipped the persisted state back to
active (and reset turns_used) and rendered an acknowledgement — nothing
re-entered the conversation loop, so the goal sat idle until the user
sent another ordinary message.
Fix the whole class by scheduling the canonical
GoalManager.next_continuation_prompt() through each surface's existing
input path after a successful resume:
- Desktop/TUI (tui_gateway/methods_tools.py command.dispatch): return a
sendable {type: "send"} dispatch with the continuation as the message,
a "Continuing now" notice, and display "/goal resume" so the
transcript shows the concise invocation instead of the model-facing
scaffolding. No-goal keeps the exec response.
- Classic CLI (hermes_cli/cli_commands_mixin.py): put the continuation
on _pending_input, same as the /goal <text> kickoff.
- Messaging gateway (gateway/slash_commands.py): enqueue a continuation
MessageEvent through the adapter FIFO — the same path the post-turn
judge uses — so queued real user messages preempt naturally and the
pause/clear stale-continuation cleanup recognizes it.
Also correct the now-misleading gateway.goal.resumed copy ("Send any
message to continue…") across all 17 locale files.
Regression tests cover exact budget exhaustion → resume on the real CLI
handler, the real gateway handler (including the
_is_goal_continuation_event guard contract), and the TUI
command.dispatch boundary; verified each fails on the pre-fix code.
Fixes#75362
The agent could reveal single panes (focus_pane) but had no way to arrange
the workspace as one act. apply_layout closes that gap: a desktop_ui tool
that emits layout.apply over the existing bridge, resolved in the renderer
against the layouts contribution registry — the same list the layout picker
reads — so core presets (default/focus/terminal-deck/quad), plugin presets,
and user-saved presets are all addressable by id. Active session only, same
as pane.reveal: a background turn never rearranges the user's desktop.
The `questions` parameter had a full description, but the top-level
tool description still described three single-question modes and never
mentioned batching. The model decides how to call a tool from that
description, so it kept asking one question per call.
The description now states that 2-5 independent questions can go in
one call and that one batched call is preferred over a chain of
single-question calls. The parameter description also tells the model
to put a short batch title in the still-required top-level question.
Two schema tests pin the contract: the description names the batch
capability, and the questions parameter stays optional with the
MAX_QUESTIONS cap.
Shift-Tab walks backwards through the questions, with wrap, the same
way Tab walks forward. A locked answer now renders on its own indented
line in a distinct color under its question, instead of an arrow
suffix on the status line, so the current answers stay readable while
the cursor moves.
A re-visited question restores its earlier state: a choice answer puts
the cursor back on that choice, a typed answer highlights the Other
row and shows the typed text next to it. Enter on an answered Other
switches to freetext with the composer prefilled with the earlier
text, so the user edits instead of retyping. Answer metadata records
how each answer was produced to drive the restore.
The clarify callback accepts a questions list and renders a batch
panel. The batch panel shows all questions as a status list with one
expanded active question. Enter locks the active answer and moves to
the next unanswered question. Tab cycles questions for any-order
answering. Locked answers stay editable until the batch completes. A
timeout returns the locked partial answers with a timed_out flag. The
single-question panel is unchanged.
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).
Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.
Both agent-side clarify dispatch sites forward the questions arg.
The clarify tool gets an optional questions parameter (2-5 independent
questions, issue #18450). Batch-capable platform callbacks receive the
normalized list in one call and reply with per-question answers. Legacy
callbacks are looped one question at a time. The loop stops on timeout
so the user is not asked the remaining questions after they walk away.
Locked answers survive a timeout: the result carries them plus a
timed_out flag, and unanswered entries have an empty user_response.
The single-question path is byte-identical to the previous behavior.
detect_install_method reads the stamp against an allowlist. The
allowlist held "nixos" but not "home-manager", and a stamp that names
home-manager gave "unknown". The managed path (step 3) returned the
correct name, so the gap was invisible: it appeared only for an install
that carries a stamp.
An install with the value "unknown" gets "hermes update" as its update
guidance. That command is the one command a managed install refuses, so
the user gets a dead end.
The test for this was also environment-dependent. It called the real
get_project_root(), and it passed here only because this worktree
carries no stamp. A checkout from the curl installer carries a "git"
stamp, and the assertion then failed for the contributor and not for
us. The test now detects against a temporary install tree.
The new test stamps each managed system and asserts the value that
comes back. With the allowlist reverted, the home-manager case fails
with "assert 'unknown' == 'home-manager'". The nixos case passes,
because that name was already in the allowlist.
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.
The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.
This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:
- the options
- the renderers for config.yaml, .env and the documents
- the activation body
- the command lines of the processes
nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.
`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.
The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.
The change also makes four corrections that apply to both modules:
- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
had only the gateway. But Hermes Desktop and the web dashboard connect
to a different process, so six of the configurations in public repos
add a second unit by hand. serve and dashboard are one entry point with
one flag of difference, and you can run only one of them. Thus the
option is an enum. The NixOS module asserts against container mode with
a backend, and does not make a unit that cannot start.
- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
installs into the working directory, which is correct for AGENTS.md but
wrong for SOUL.md and memories/. Hermes reads those files from
HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
made a workspace file that Hermes never loaded as the identity. The
documentation said this in prose, but two directory diagrams showed the
opposite. This change corrects both. A key in either option can now
contain subdirectories.
- `documents` needs an explicit `workingDirectory`. The default of that
option is bad on both modules. It is the home directory of the user on
Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
workspace files without a directory therefore gets a place that the
user did not select. The place is also different on each module. The
modules now refuse that combination.
The test is on the priority of the option and not on its value. An
option that nothing sets keeps the priority of its own default, and
each definition from a user is stronger. Thus a directory with the same
text as the default still counts as a selection, and so does a
mkDefault. A comparison of values detects neither case.
- Each activation writes .env again from a base in the Nix store, and
does not add to the file that exists. Thus a second activation cannot
put the same secret in the file two times, and a removed
environmentFile goes away. environmentFiles keeps the type `listOf
str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
the Nix store, which all users can read.
- HERMES_MANAGED and the .managed marker now hold the name of the system
that manages the install. Thus a refusal says "managed by home-manager"
and not "managed by NixOS", and `hermes update` gives the Nix guidance
for both shapes. The CLI does not print a rebuild command for each
system. It names the owner, and the user knows their own tool. A bare
`true` and an empty marker still mean NixOS, so this does not change an
existing install.
Verification. Six new checks, all built:
nixos-module evaluates the module with evalModules and the
NixOS module list. It asserts both units, one
HERMES_HOME, and that the module refuses
container mode with a backend.
home-manager-module evaluates the module with the
homeManagerConfiguration function of
home-manager. The process assertions run against
systemd units on Linux and launchd agents on
Darwin.
module-option-parity asserts that each shared option is on both
modules, and that the two exclusion lists name
only options that exist.
env-file-assembly runs the real .env script and checks the
contents, the mode, that a second run gives the
same bytes, and that a removed file goes away.
workspace-files-need-a-directory
checks that the module refuses `documents`
without a directory, and accepts a directory
that has the same text as the default.
service-argv runs each command line that the modules build
through the real parser of the CLI, with one
sentinel flag added, and requires that argparse
refuses only the sentinel.
`nix flake check` passes, with 21 checks in total.
The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.
Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:
- a lost --no-open
- a backend that runs the gateway
- an overwritten config.yaml
- documents in the wrong directory
- a different HERMES_HOME on the two processes
- a lost HERMES_HOME export
- a missing backend unit
- a removed assertion
- an .env file that grows at each activation
- an install that reports NixOS
- an empty .managed marker
- an option on the NixOS module only
- a stale entry in an exclusion list
- a renamed subcommand
- an unknown flag
- the workspace-files assertion always passes
- the assertion compares values instead of priorities
- an off-by-one that lets an untouched default through
- the assertion also fires for hermesHomeFiles
- a mkDefault no longer counts as a selection
- the Home Manager module stops wiring the assertion
- the NixOS module stops wiring the assertion
The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.
These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:
- test_git_probe_tree_kill.py (2 tests)
- test_update_import_guard.py (1 test)
- test_telegram_media_read_timeout.py (2 tests)
- test_teams.py (a collection error)
Closes#9056
# Conflicts:
# hermes_cli/main.py
# hermes_cli/update_cmd.py
# hermes_cli/web_server.py
Some providers reject the structured-output request field with a hard
400. The error classifier marks a 400 as non-retryable, so one rejected
field failed the whole auxiliary call. Session titles stayed derived
forever (#82816), and no fallback fired.
Three rejection shapes are covered, from live reports:
- vLLM gateways translate response_format into guided_grammar and fail
when the grammar backend is absent (compile_grammar_error: No module
named 'xgrammar').
- Some OpenAI-compatible endpoints answer "This response_format type
is unavailable now".
- Anthropic-compatible gateways that predate structured outputs reject
the translated field: "output_config: Extra inputs are not
permitted". The documented case is the bedrock-mantle Messages
endpoint.
The fix is reactive, the same pattern as the temperature and
max_tokens rungs: when the provider rejects the field, retry once
without it. Callers tolerate an unconstrained reply — the title prompt
demands bare JSON and _extract_title_text has a loose-JSON fallback —
so the call succeeds with prompt compliance instead of failing. The
retry only fires when the request carried the field, and both the sync
and async paths get the same rung.
Closes#82816
The adapter builds the Messages body from a fixed allow-list of kwargs.
A caller that passes response_format as a top-level kwarg (the OpenAI
SDK call shape) got it dropped on the floor. The request succeeded, but
the schema contract silently became prompt compliance. No in-tree
caller uses this shape today. The pin-test makes sure that a future
refactor cannot open this leak again.
The top-level kwarg gets the same output_config.format translation as
the extra_body shape. When a caller sends both shapes, the extra_body
value wins because every in-tree caller uses that shape.
Pin-test pattern from PR #85626 review follow-up.
Co-authored-by: Matt McClean <mmcclean@amazon.com>
Plugin structured completions (plugin_llm.complete_structured) build an
OpenAI Chat Completions response_format payload in extra_body. The
anthropic_messages transport forwarded it verbatim, and strict
Anthropic-compatible gateways reject it with HTTP 400:
response_format: OpenAI Chat Completions structured-output shape is
not supported. Use output_config.format = {"type": "json_schema", ...}
Observed live: every discord-thread-autotitle structured call failed
for 2+ days (1,600+ logged errors) once the main provider became an
anthropic_messages gateway.
Fix: _translate_anthropic_response_format converts
- json_schema -> output_config.format = {type: json_schema, schema: S}
- json_object -> permissive object schema (SDK 0.87.0 has no
schema-less JSON mode)
merging into any existing output_config (adaptive-thinking effort
coexists) and excluding response_format from the raw extra_body
passthrough alongside the existing reasoning exclusion. The async
adapter delegates to the sync adapter via asyncio.to_thread and is
covered by a test. Non-Anthropic transports are unchanged.
The salvaged fix persisted the recovery note unconditionally — correct for
the synthesized empty auto-resume turn, but a user who typed real text while
resume was pending would get the [System note: ...] scaffold persisted as
their own words, leaking scaffolding into the durable transcript (the same
class as #81841 on the assistant side).
_prepare_resume_pending_message now persists the note only when the original
message is blank; real text persists verbatim while the model still receives
the wrapped note. Whitespace-only counts as blank. Tests cover all three
shapes.
The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.
Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.
Projection-side approach credit: @JoaoMarcos44 (PR #88996).
Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (#88955).
Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused #81841.
Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.
Refs #88955
Review follow-up on the salvaged #88965 work:
- test_goal_command_slow_db_init_still_persists: drop the 4s slow-init
loop-gap harness (wall-clock gap assertions on shared CI runners are
their own flake class; loop-freeze bounds are already covered in
test_goals_db_bootstrap_off_loop.py). The persistence contract keeps
its discriminating power by shrinking the monkeypatched init window
(0.2s) under a 0.8s slow init — the window-only path still fails it.
- test_slow_construction_does_not_block_the_loop: monkeypatch both
bootstrap windows down (0.3s/0.05s) and shrink the blocking init from
6s to 1.5s; the two-window contract is what's under test, not the
production constants. Adds an elapsed ordering assertion so the
kick-vs-in-flight window distinction stays pinned.
Combined wall time for the pair: ~10.5s -> ~1.4s.
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).
Two changes, one per caller shape:
- Async callers (_get_goal_manager_for_event,
_get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
the heartbeat poller) warm the SessionDB cache off-loop through the
context-preserving executor before constructing the manager (shared
_warm_goals_session_db helper). The loop never blocks and the first
write lands at any init duration. A bare to_thread would lose the
per-turn profile home override under multiplex; the executor hop
keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
cannot await, so the bootstrap windows stay: the call that starts the
bootstrap waits a one-time init window (1.5s) instead of the short
per-call window (0.25s), giving healthy cold inits room to land while
a contended migration still degrades to None with only a bounded
one-time stall. The bootstrap thread binds the caller's home as a
contextvar override so a multiplexed worker cannot cache the default
profile's DB under another profile's key.
save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).
Independent diagnosis + measurement by jackulau (#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
Third CI hit today for tests/gateway/test_goal_verdict_send.py (twice on
salvage PRs, once on main's own push run), always the same shape:
adapter.sends == [] after the full drain.
Mechanism (reproduced, not log-read): the tests call GoalManager.set() on
the event-loop thread. _get_session_db() refuses to construct SessionDB on
a loop thread (loop-liveness guard from the 2026-08-14 crash-loop fix) and
waits only _DB_BOOTSTRAP_LOOP_WAIT_S=0.25s for the background bootstrap.
On a loaded CI runner the init overruns that window, set() degrades to a
silent no-op by design, the goal never exists, and the continuation path
correctly does nothing — so no amount of drain-waiting helps (the #88975
de-flake addressed a different, downstream race).
Fix: the hermes_home fixture pre-warms the SessionDB cache from its sync
context (direct construction path), so the loop-thread set() always finds
a cached DB. Production behavior untouched.
Proof: injecting a 0.4s-slow SessionDB.__init__ reproduces the exact CI
failure on the old fixture and passes 2/2 with the pre-warm.
Adds a reviewable in-app install path for Hermes plugins:
hermes://plugin/install?repo=owner/repo (and Settings -> Plugins ->
Install from Git) opens a confirmation modal showing the repo identity
and source links, shallow-clones to probe for agent and/or desktop
plugin artifacts, lets the user pick components, then installs — agent
side through the gateway's new plugins.manage `install` action (wrapping
the existing dashboard_install_plugin), desktop side through a new
Electron git-install module with subdir-escape guards, a 60s clone
timeout, non-interactive git env, and insecure-scheme warnings. Never
auto-installs; hybrid repos get one dialog. Legacy plugin-agent /
plugin-desktop deeplinks route into the same modal.
Salvaged from PR #82735 by @serefyarar (net diff applied onto current
main as a single authored commit; the branch carried merge commits).
The preview screenshot PNG from the original branch was intentionally
not carried over — images live in PR bodies, not the repo.
Follow-up to the salvaged CommandCode signature fix: accepting base_url
but ignoring it left custom endpoints (user-configured model.base_url /
COMMANDCODE_BASE_URL proxies) fetching the public catalog instead of the
configured one. Reviewer dansigma flagged this on PR #88851.
Class-wide fix, not a CommandCode patch:
- providers/base.py: a caller base_url that DIFFERS from the profile's
default now wins over models_url. Equality with the default means "not
customised" (callers pass base_url unconditionally, defaulting to the
profile's own URL) and keeps models_url as the endpoint, preserving the
OpenRouter-style split-catalog behavior.
- commandcode: _fetch_commandcode_models() takes the endpoint override;
both profile overrides forward base_url.
- Tests: base-class precedence (custom beats models_url, default does
not), CommandCode redirect via live local HTTP server incl. claude-*
filter, and default-echo hitting the canonical endpoint. All verified
to fail against the pre-fix implementation (sabotage run).
Plugin-only providers (commandcode, tencent-tokenhub, ...) are absent from
models.dev and HERMES_OVERLAYS, so resolve_provider_full returned None and
/model switches failed with "Unknown provider ..." even though the picker
lists them (CANONICAL_PROVIDERS auto-extends from the same registry).
Fall back to providers.get_provider_profile() before giving up, mapping the
profile api_mode to the ProviderDef transport.
The model picker's generic live-fetch path (hermes_cli/models.py
provider_model_ids) calls profile.fetch_models(api_key=..., base_url=...).
Both CommandCode overrides only accepted api_key/timeout, so every picker
open raised TypeError, which was silently swallowed, leaving the provider
with zero models.
Match the base ProviderProfile.fetch_models signature (base_url kwarg) and
add a regression test asserting both profiles accept it.
PR #67934 marked auto-discovered catalogs by writing two sentinel keys
INSIDE the user-facing ``models`` mapping of custom provider entries:
``__discovered_model_catalog__`` (written by
_save_discovered_models_to_config) and ``__explicit_model_allowlist__``
(injected by _normalize_custom_provider_entry). Every consumer of that
mapping — pickers, selectors, gateway/agent readers, and the user's own
config.yaml — had to know to filter those keys, and any site that
didn't listed them as phantom model IDs (``__discovered_model_catalog__``
showing up as a selectable "model"). The v11→v12 config migration and
the ACP session-state test caught exactly that leak on main.
Replace the in-mapping sentinels with a single entry-level flag:
- ``models_discovered: true`` now sits next to ``models``/``base_url``
on the provider entry; the models mapping stays a clean
``{model_id: metadata}`` dict with no reserved keys.
- _save_discovered_models_to_config writes the new shape and refreshes
catalogs it previously discovered (entry-level flag or legacy
sentinel) instead of treating them as user-curated metadata.
- _normalize_custom_provider_entry no longer injects
``__explicit_model_allowlist__``; a dict-shaped models mapping counts
as an explicit allowlist exactly when the entry is NOT marked
models_discovered.
- _models_config_is_allowlist takes the discovered flag as a parameter
(new helper _entry_models_discovered resolves it, including the
legacy in-mapping sentinel); all call sites updated
(model_switch.py, model_setup_flows.py, acp_adapter/server.py).
- Backward compat, no config version bump: configs written by a
pre-fix Hermes (sentinels inside models) still read correctly —
``__discovered_model_catalog__: true`` is treated as
models_discovered, both sentinel keys are stripped from model
listings, and the next discovery save migrates the entry to the
clean shape. Covered by a new regression test.
Also restore ``except Exception:`` on the pre-existing guards this PR
had narrowed to specific exception tuples (the resolve_runtime_provider
fallback in switch_model, the picker discovery/cache guards in
list_authenticated_providers, _get_model_config_dict, and
_credential_fingerprint). Those guards were intentionally broad on
main — a failed resolution or probe must degrade to the fallback path,
never crash the model switch. Guards the PR introduced for its own new
probe code keep their authored tuples.
The ACP new_session payload also goes back to
probe_current_custom_provider=False, matching the contract main's
test_new_session_returns_authenticated_cross_provider_model_state pins
(session opens must not block on live-probing the current custom
endpoint).
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.
Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.
Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.
Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root
Fixes#65828