The model picker's generic live-fetch path (hermes_cli/models.py
provider_model_ids) calls profile.fetch_models(api_key=..., base_url=...).
Both CommandCode overrides only accepted api_key/timeout, so every picker
open raised TypeError, which was silently swallowed, leaving the provider
with zero models.
Match the base ProviderProfile.fetch_models signature (base_url kwarg) and
add a regression test asserting both profiles accept it.
PR #67934 marked auto-discovered catalogs by writing two sentinel keys
INSIDE the user-facing ``models`` mapping of custom provider entries:
``__discovered_model_catalog__`` (written by
_save_discovered_models_to_config) and ``__explicit_model_allowlist__``
(injected by _normalize_custom_provider_entry). Every consumer of that
mapping — pickers, selectors, gateway/agent readers, and the user's own
config.yaml — had to know to filter those keys, and any site that
didn't listed them as phantom model IDs (``__discovered_model_catalog__``
showing up as a selectable "model"). The v11→v12 config migration and
the ACP session-state test caught exactly that leak on main.
Replace the in-mapping sentinels with a single entry-level flag:
- ``models_discovered: true`` now sits next to ``models``/``base_url``
on the provider entry; the models mapping stays a clean
``{model_id: metadata}`` dict with no reserved keys.
- _save_discovered_models_to_config writes the new shape and refreshes
catalogs it previously discovered (entry-level flag or legacy
sentinel) instead of treating them as user-curated metadata.
- _normalize_custom_provider_entry no longer injects
``__explicit_model_allowlist__``; a dict-shaped models mapping counts
as an explicit allowlist exactly when the entry is NOT marked
models_discovered.
- _models_config_is_allowlist takes the discovered flag as a parameter
(new helper _entry_models_discovered resolves it, including the
legacy in-mapping sentinel); all call sites updated
(model_switch.py, model_setup_flows.py, acp_adapter/server.py).
- Backward compat, no config version bump: configs written by a
pre-fix Hermes (sentinels inside models) still read correctly —
``__discovered_model_catalog__: true`` is treated as
models_discovered, both sentinel keys are stripped from model
listings, and the next discovery save migrates the entry to the
clean shape. Covered by a new regression test.
Also restore ``except Exception:`` on the pre-existing guards this PR
had narrowed to specific exception tuples (the resolve_runtime_provider
fallback in switch_model, the picker discovery/cache guards in
list_authenticated_providers, _get_model_config_dict, and
_credential_fingerprint). Those guards were intentionally broad on
main — a failed resolution or probe must degrade to the fallback path,
never crash the model switch. Guards the PR introduced for its own new
probe code keep their authored tuples.
The ACP new_session payload also goes back to
probe_current_custom_provider=False, matching the contract main's
test_new_session_returns_authenticated_cross_provider_model_state pins
(session opens must not block on live-probing the current custom
endpoint).
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.
Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.
Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.
Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root
Fixes#65828
_write_manifest still used a hand-rolled mkstemp + atomic_replace,
so every sync reset .bundled_manifest to mkstemp's 0600, dropping a
group-readable or shared mode the operator had set. Replace the block
with utils.atomic_write_text(preserve_mode=True) — the same shared
writer and mode-preservation contract PR #86255 applied to the skill
manager's document writes.
This is the remaining half of PR #14410 by @sgaofen, who reported the
manifest mode reset first; the skill-manager half of that PR was
superseded by the atomic_write_text refactor and #86255.
Port the stale PR #61665 behavior to current main. The original two-file contribution is by yungchentang; this candidate preserves its scoped Codex OAuth fallback intent.
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
The getAllTranscripts resourceData.id from the field report is a base64url
blob whose DECODED payload ends in "-TranscriptV2" while the encoded form
contains no readable marker, so the substring heuristic in
looks_like_transcript_id missed it. Add a best-effort base64 decode hint so
degraded notifications (no @odata.id) are still refused with the clear
guidance error instead of a cryptic Graph 400.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
test_real_aiagent_builds_section_once_and_keeps_it_out_of_static_prefix
builds the system prompt twice and asserts byte equality, but the prompt
embeds build_coding_workspace_block() — live git status/log output. A git
call failing between the two builds (xdist contention in CI) makes the
Branch/Recent-commits lines differ and fails the test on unrelated PRs
(first seen on #89027's run: diff showed only '- Branch: (detached HEAD)'
and recent-commit lines).
Mechanism reproduced locally: with coding posture on and cwd inside the
checkout, failing git calls on the second build only => first != rebuilt;
with the snapshot pinned, identical git failure => byte-equal. The real
block's byte-stability is coding_context's own contract; this test is
about plugin sections.
Per review: even with an active sandbox env, spilled tool results
belong in $HERMES_HOME/cache/spillover with the other Hermes-owned
caches — not the sandbox temp dir as primary storage.
- Host-side write happens first on every backend; local/no-env
sessions reference the host path directly (unchanged).
- cache/spillover joins the auto-mount/sync cache-dir list
(credential_files._CACHE_DIRS), so docker bind-mounts it and
modal/ssh/daytona file-sync it. Remote references use the
translated in-sandbox path after a readability probe.
- Probe failure (persistent containers created before spillover
joined the mount list, translation failures) falls back to the
previous in-sandbox temp-dir copy, so nothing regresses.
The steer-survives-budget tests pinned the inline 'Truncated:' fallback
shape, which only occurred because env=None persistence was broken.
Now that host-side spillover succeeds, budget enforcement produces a
<persisted-output> block instead. Assert the actual contract — the
oversized payload was replaced (persisted OR truncated) — via a shared
helper, not which replacement shape was used.
Sessions that never ran a terminal command (MCP-only, cron, gateway)
have no active sandbox environment, so maybe_persist_tool_result()
got env=None and fell through to the inline-truncate fallback --
a 467K MCP result was cut to a ~1.3K preview with no file written
('Full output could not be saved to sandbox').
Now the host-side cases (env=None or the local backend) write the
spill file directly to $HERMES_HOME/cache/spillover/<id>.txt,
alongside the other Hermes-owned caches instead of littering /tmp.
Remote backends (docker/ssh/modal/daytona) keep the in-sandbox
env.execute() write since read_file resolves in-sandbox there.
Cleanup: the gateway housekeeping loop prunes spillover hourly with
the other media caches, and a once-per-process best-effort prune on
first spill covers CLI-only installs.
test_session_chat_stream_treats_pre_existing_poisoned_row_as_no_model
asserted mock_run.call_args right after the 200 status, but the stream
handler runs _run_agent inside asyncio.create_task(_run_and_signal())
and response.prepare() returns the 200 before that task necessarily
starts — on loaded CI runners call_args was still None (TypeError:
cannot unpack non-iterable NoneType). Draining the SSE body (resp.text())
joins the stream end, which guarantees the runner task completed.
test_goal_verdict_send used fixed asyncio.sleep(0.05) waits before
asserting on sends/enqueues produced by spawned tasks; replaced with a
bounded _drain_until() poll (5s cap, returns as soon as the condition
holds) so the asserts stay exact without the fixed-delay race.
These three tests red-flagged unrelated main pushes and PR runs on
Aug 18 (runs 32099139396, 32101135100, 32106224479, 32101070064).
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.
Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.
Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
test_update_hangup_protection pinned the exact stdout of
_print_update_completion; the new branch+HEAD suffix (parked-branch guard)
broke that pin. The two receipt tests assert the action-identity contract,
not the branch display, so they now neutralize _branch_head_suffix — the
suffix behavior itself is covered by test_update_parked_branch_guard.py.
Live incident 2026-08-17: the source checkout was parked on a stale feature
branch (claude-code-inspired/local-terminal-memory-limit, days behind main),
left there by earlier tooling. 'hermes update' autostashed, refreshed lazy
backends, synced skills, and printed '✓ Code updated!' / '✓ Update complete!'
while the checkout stayed on the stale branch with none of main's new code.
Two sessions burned time on 'the fix is missing' confusion.
- Parked-branch guard: auto-switch back to the update target ONLY when the
parked branch is clean and fully merged (git cherry origin/<target> shows
nothing unmerged); the checkout then STAYS on the target instead of being
re-parked. Otherwise: loud CODE UPDATE SKIPPED block naming the branch,
behind-count, and resolution commands; exit 1; branch untouched.
- The up-to-date (commit_count == 0) path no longer switches back to a
fully-merged parked branch either.
- Post-pull gate additionally refuses to print '✓ Code updated!' when HEAD
ends up attached to a non-target branch.
- Summary lines now carry the actual branch + HEAD short-sha:
'✓ Update complete! [main @ 30fcf9580]' — drift visible at a glance.
- New config toggle updates.auto_switch_parked_branch (default true).
- Real-git-fixture regression tests (init/clone/branch, no subprocess
mocks): clean+merged auto-switch, dirty skip, unmerged skip, cherry-picked
equivalence, config opt-out, unverifiable ref, on-main fast path,
up-to-date no-repark, summary branch/sha assertions.
Backend: /api/status now carries a stable random install_id persisted once
under the root HERMES_HOME, shared by every profile of the install.
Desktop: roster enumeration captures it per connection (TTL-cached probe),
buildAgentRoster collapses same-install rows with a deterministic canonical
pick (active > local > ssh > remote > cloud > earliest), the @name-device
handle rule runs after the collapse, and Settings → Gateways shows a
display-only 'Same backend as' hint. Backends without install_id bypass the
collapse (fully backward compatible).
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
Review finding on the off-loop bootstrap: returning None on every cold-
cache loop-thread call silently dropped the first goal/heartbeat
persistence op even when the DB was perfectly healthy. The loop-thread
path now waits up to 250ms on the bootstrap event - a healthy init
(tens of ms) completes inside the window and the caller gets the real
DB; a contended init (the crash-loop scenario) exceeds it and degrades
to None with a bounded, watchdog-safe stall.
Only the initial SELECT of _dedupe_legacy_system_prompts was guarded;
a 'database is locked' on any per-row write propagated out, aborted
schema init, left the schema version below 25, and made every later
SessionDB.__init__ re-enter the same migration against the same
contended DB - the second half of the enterprise crash-loop report.
The per-row loop now catches OperationalError, logs once, and returns.
Partial migration is safe by design: the legacy system_prompt column
is the documented read fallback for unmigrated rows, and the next
schema init resumes where the contention stopped. Tests prove rows
migrated before the failure stay migrated, the remainder stays
readable, and a later run completes it.
SessionDB.__init__ runs schema init, and a migration against a contended
state.db blocks for seconds. The goal/heartbeat path reached it
synchronously on the gateway's event-loop thread (GoalManager() ->
load_goal -> _get_session_db -> SessionDB()), so a contended DB starved
the loop-liveness watchdog, which hard-exited with code 75 and the
supervisor restarted straight back into the same state - an unbounded
crash loop reported from an enterprise fleet.
_get_session_db now detects a running loop on the calling thread: on a
cache miss it kicks a one-shot background bootstrap thread and returns
None immediately (every caller already degrades gracefully on None);
the cached instance serves all later calls. Worker threads construct
inline as before, with a lock-guarded cache so a bootstrap race keeps
one instance and closes the loser. The heartbeat module shares this
boundary via the same _get_session_db.
Field report (enterprise, v0.20.0): cron jobs delivering text + PDF/image
attachments to Slack DMs deliver both on scheduled ticks but text-only on
manual `hermes cron run <job-id>`. Same box, same token, same scopes —
the divergence is process context and error visibility, not credentials.
Three defects, one bug class (attachment failures invisible + policy
divergence between the gateway process and standalone processes):
1. Standalone lane swallowed warnings: platform standalone senders
(Slack files_upload_v2, Discord, ...) report per-file upload failures
in result['warnings'] while returning success=True for the delivered
text leg. _deliver_result only read result['error'], so the run was
marked ok and the attachment vanished without a trace. Warnings now
surface into delivery_errors (and the job's last_error).
2. Live-adapter lane swallowed media failures: _send_media_via_adapter
logged failures at WARNING and returned None. It now returns per-file
error strings and _deliver_result records them — text-delivered-but-
attachment-failed is a visible partial failure on BOTH lanes.
3. Media-policy env bridge was gateway-only: gateway.strict /
media_delivery_allow_dirs / trust_recent_files were translated from
config.yaml to the env vars validate_media_delivery_path reads ONLY in
gateway startup. A CLI-process manual run filtered attachment paths
under a different policy — in strict/allowlisted deployments the exact
reported symptom (scheduled delivers, manual drops, silently). The
translation now lives in gateway/media_policy.apply_media_policy_env
(idempotent, env-wins, never raises); gateway startup delegates to it
and _deliver_result applies it before filtering. Attachments dropped
by the policy filter are also reported in the run status instead of
only a stderr WARNING.
On v0.20.0 specifically the failure was double-blind: the pre-9cf2cbd382
isinstance(resp, dict) gates meant upload failures were undetectable in
the sender AND unsurfaced by the scheduler. 9cf2cbd382 (in 2026.8.13)
fixed detection; this fixes visibility and policy parity.
8 new tests (tests/cron/test_media_delivery_parity.py): warnings→errors,
clean-delivery control, media-reaches-sender control, live-adapter
failure/dropped-path reporting, bridge helper semantics, strict+allowlist
end-to-end in a non-gateway process, and the .env-strict/config-allowlist
split that reproduces the field symptom. Mutation check: disabling the
warnings loop and the bridge fails exactly the 2 guarding tests.
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):
- Incomplete-validator findings are now PRESERVED as partial evidence;
only the validator's pass/fail verdict is excluded from the advisory
verdict. A report with findings from an incomplete check no longer
reads as clean.
- Clean-report wording is now "no findings from completed checks"
whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
discard-pinning test, plus the completed-checks wording case.