Clarify prompts share the same emitted-while-detached failure class the
pending-approval replay fixed: `clarify.request` rides `_block()`'s pending
registry, so a client whose transport was down when the event fired never
sees the question and the agent thread stays parked until timeout.
Widen the resume snapshot the same way:
- tui_gateway/server.py: `_live_session_payload` now carries
`pending_clarify` — a read-only snapshot of the clarify prompt still
blocking the session, scoped to the owning runtime sid. The registry stays
authoritative; the embedded request_id resolves via clarify.respond.
- Desktop resume paths (`use-session-actions`) restore the parked clarify
into the clarify store (multi_select preserved) and flag needsInput,
mirroring restorePendingApproval on both the activate and resume paths.
- pending_approval replay now also forwards the queue-injected request_id so
the restored prompt responds with exact-request correlation.
- Tests: server-side replay + scoping test; harmonized the #82087 replay
test with the request_id `_ApprovalEntry` now injects.
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
Desktop/TUI busy-input `steer` mode escalated any fall-through message
(steer() rejected, raised, or a non-steerable multimodal payload) into a
hard interrupt of the live turn. AIAgent.interrupt() also clears the
pending steer buffer, so a burst of user messages sent while the agent
was busy could be silently destroyed: earlier successfully-steered
messages were dropped from the buffer and the live turn was killed.
Steer-mode fall-throughs now keep pure queue semantics: preserved FIFO
in queued_prompt/queued_prompts and drained on turn end, per the
existing steer contract. Only explicit `interrupt` mode still fires
_interrupt_busy_session. No synthetic user messages are injected
mid-loop; accepted steers continue through the sanctioned OOB
steer-marker path.
Regression tests cover: rejected steer queues without interrupting,
steer exception falls back to queue, multimodal payload queues, a mixed
burst preserves accepted steers plus the queued fall-through, and a
fall-through burst drains all texts FIFO after turn end.
Fixes#86134
A mid-turn correction (Desktop session.redirect / busy-input interrupt redirect)
must not leave a server-queue self-copy of the live inflight user prompt.
Otherwise post-turn _drain_queued_prompt restarts that original text as a
fresh agent turn after Q completes (#84417).
Scrub text-only self-duplicates of inflight_turn.user on successful
redirect/steer, refuse admitting them in _enqueue_prompt, rewrite merged
"{P}\n\n{Q}" slots to Q-only, bump _queued_prompt_generation on compression
session rotation, and restore the claimed queue envelope when generation
cancels mid-drain. Stabilize profile-scoped agent-build unit tests under CI load.
Fixes#84417
Address review on #83878:
- Permanent fatal fences all hold producers and discards pending maps on
teardown instead of re-populating a queue that can never drain.
- Any hold created while connected schedules a tracked redispatch (cancel-
after-pop no longer orphans until a future reconnect).
- Redispatch failures re-hold current + remainder without tight-looping.
Regression coverage for the three residual paths, plus the interaction
with OOF-156's connect-failure classification: the retryable network
path (telegram_connect_error) must NOT clear the hold queue — reconnect
is precisely what drains it; only non-retryable fatals discard.
The disconnect drop-guard (#55971) correctly prevents dispatch into a
torn-down session. Destroying the event was wrong: by enqueue/flush time
python-telegram-bot has already acked the update and advanced the polling
offset, so Telegram never redelivers. Result: silent permanent loss, no
log, no error.
Hold inbound events (text/photo/media-group) when the drop-guard fires,
salvage pending batch maps on teardown, cancel+await the redispatch task
in the delivery cancel map (lifecycle-tracked), and redispatch from
_mark_connected after reconnect. Cap the hold queue (default 64), dedupe
by object identity, discard on non-retryable fatal. Cancel-after-pop in
flush paths also holds.
Distinct from #72037 (cancel-after-pop during follow-up supersession) and
#81528 (boundary discard). Tests use delay=0 and entered/release Events —
no wall-clock races; includes production terminal-step coverage.
Restores pass-through behavior for cron delivery and react/unreact that
was lost when d409f6748 routed them through resolve_send_target. Stored
cron job targets the channel directory doesn't recognize (e.g.
telegram:ops-room on a fresh install, photon group GUIDs) used to go to
the adapter verbatim; after d409f6748 they were silently dropped. Same
for react on platform-native ids.
Adds an opt-in pass_unresolved_references flag to resolve_send_target,
passed only by cron and react. Model-facing send tool stays strict.
Plugin platforms with a parser stay strict for all callers. The optional
validator still has the final say over passed-through ids.
Follow-up fixes on salvage:
- Update test_cron_relay_delivery_guards.py mock lambdas to accept **kw
(file added to main after PR branch point; lambdas didn't accept the
new keyword argument)
- Consolidate duplicated pass-through blocks into _pass_through_unresolved
local helper
Fixes#85128
Co-authored-by: Adolanium <Adolanium@users.noreply.github.com>
ActualProfile.fetch_models() overrides ProviderProfile's default
implementation with its own Actual-specific base_url resolution
(ACTUAL_BASE_URL env var, hosted-vs-local normalization), but called raw
urllib.request.urlopen(req, timeout=timeout) directly instead of the base
class's open_credentialed_url(). Every other provider either uses the
base class default or forwards to it via super() and gets
SafeCredentialRedirectHandler for free — Actual is the only provider that
attaches a Bearer token to its own Request object and opens it with the
stdlib's default redirect handling, which forwards every header,
including Authorization, across a cross-origin redirect.
Actual's own feature surface makes the trigger realistic: ACTUAL_BASE_URL
is a first-class, documented way to point this provider at a self-hosted
or local-offline endpoint (see the local-loopback no-auth path already
handled elsewhere in this provider), so a misconfigured or compromised
endpoint 302-ing to another host leaks ACTUAL_API_KEY to it.
Fix: import and call the same open_credentialed_url() the base class
uses, keeping Actual's own URL-resolution logic unchanged.
Adds an end-to-end regression test using two real local HTTP servers (no
mocking of the security module itself) — one redirects, the other
records the Authorization header it receives — mirroring
test_urllib_security.py's own redirect tests. Also repoints the existing
fetch_models test's mock from urllib.request.urlopen to
hermes_cli.urllib_security.open_credentialed_url, since fetch_models no
longer calls the former. Mutation-verified: the new redirect test fails
on pre-fix code with the Authorization header observed at the redirect
target.
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:
- the requests-based metadata/pricing probe
(agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
(hermes_cli/models.py::probe_api_models)
A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.
This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:
- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
ssl_ca_cert before falling back to the env vars. Callers with no base_url
keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
passes it through open_credentialed_url, which gains an ssl_context seam on
the cloned secure opener. Unmatched or public endpoints pass None and keep
urllib's default policy.
Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.
For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.
Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.
Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
Follow-up on the #83854 salvage: prepend $HERMES_HOME/bin ahead of the
venv and user-local bin dirs, matching the managed-first Browser Use
CLI resolution policy — the worker resolves the same canonical binary
the agent process does.
The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).
Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.
Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
Adds the full MCP setup surface as profile-scoped gateway RPCs so a
desktop client (Bot Mode's bot editor, the core Capabilities tab) can
add/configure/test/authenticate/remove MCP servers for ANY profile, not
just the launch profile:
- mcp.servers.list (profile) -> configured servers (transport, auth,
oauth_tokens_present, enabled, tool names; no secret values)
- mcp.servers.add (profile, name, config|preset, bearer_token?) -> reuses
mcp_config._apply_mcp_preset / _save_mcp_server / _save_bearer_auth_token
- mcp.servers.set_api_key (profile, name, value, env_var?) -> http auth
header template or stdio env ref, via save_env_value
- mcp.servers.test (profile, name) -> _probe_single_server + oauth state
- mcp.servers.remove (profile, name)
- mcp.servers.oauth.start/poll (profile, name[, session_id]) -> mirrors the
PROVIDER oauth session/poll model (not the FastAPI dashboard flow): a
background worker drives the same interactive machinery 'hermes mcp login'
uses, capturing the browser redirect on a local loopback listener. Client
opens auth_url via openExternal and polls until status=='approved'.
All handlers are profile-scoped via set_hermes_home_override in try/finally
(mirrors skills.manage). Shared helpers live in tui_gateway/mcp_rpc_helpers.py
and are aliased onto server.py's namespace so the rebound handler bodies
(HandlerRegistry.install) can resolve them — a plain def in methods_tools is
unreachable post-rebind. Reuses hermes_cli/mcp_config.py throughout; no config
logic duplicated; no raw yaml near config.yaml (config-read-guard safe).
Tests: tests/tui_gateway/test_mcp_profile_rpcs.py, 8 E2E against real temp
HERMES_HOME profiles asserting add/list/set_api_key/remove land in the RIGHT
profile's config.yaml and not the launch profile's. 8/8. Registration +
live mcp.servers.list verified in an imported gateway.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
test_run_prompt_submit_requeues_all_unstarted_notifications_with_real_threading
failed twice in one hour on CI slices for two UNRELATED PRs (86371,
86374) with `assert set() == {proc_batch_2, proc_batch_3}`. Root cause:
session.init/create tests earlier in the file start real per-session
notification poller daemon threads and never stop them. Those pollers
outlive their test and keep polling the PROCESS-GLOBAL
process_registry.completion_queue, stealing-and-requeuing the target
test's events mid-assertion so its bounded drain loop can starve.
Reproduced: with 30 leaked foreign-session pollers injected via a
sabotage conftest, the target test fails standalone ~1 in 3 runs with
the exact CI assertion; with the reap fixture active it passed 8/8
under the same sabotage.
Fix:
- tui_gateway/server.py: _start_notification_poller registers
(stop_event, thread) in module-level _notification_pollers (pruned of
dead threads on each spawn; threads get a stable
tui-notif-poller-<sid> name for debugging).
- tests/test_tui_gateway_server.py: autouse fixture sets every
registered live poller's stop event after each test and joins them
under ONE shared 3s budget (the poller loop wakes at least every
0.5s), so no poller survives into the next test. No per-thread
timeout, no session-dict mutation — a first draft that mutated
session state and joined per-thread hung the file; full-file runtime
with this version is 15.2s vs 13.2s baseline.
Per hermes-sweeper review suggestion on #50242: the existing tests only
covered the unset case (override is None after the call). Add a test
where a caller already holds an override — _persist_live_session_system_prompt
must build the prompt under the session's profile and then restore the
caller's override via the reset token, not clear it to None.
Fixes#50233
_persist_live_session_system_prompt rebuilds the system prompt after a
live model switch (/model), but _start_agent_build's finally block has
already reset set_hermes_home_override by then. load_soul_md() and
build_skills_system_prompt() call get_hermes_home() which falls back to
the root ~/.hermes, loading the wrong SOUL.md identity and skills for
the session's profile.
Fix: set_hermes_home_override(session["profile_home"]) before calling
agent._build_system_prompt() in _persist_live_session_system_prompt,
and reset it in a finally block. Also upgrade the failure log from
DEBUG to WARNING so silent profile-wrong-prompt issues are visible.
The first-prompt lazy-build path is unaffected — _run_prompt_submit
already re-sets the override before run_conversation. The /model slash
worker path does not, which is the gap this fixes.
Co-Authored-By: Claude <noreply@anthropic.com>
Normal prompt turns bind session['profile_home'] via set_hermes_home_override
before run_conversation, but the two ephemeral RPC paths (prompt.background,
preview.restart) spawn a fresh AIAgent on a new thread where the HERMES_HOME
ContextVar does not propagate — so a background/preview turn under a
non-default profile ran against the wrong home. Re-bind for the duration of
the ephemeral turn and restore in finally, mirroring the normal prompt turn.
Surgically reapplied from PR #50777 (handlers moved to methods_prompt.py
since the PR was authored; handler bodies rebind onto server.py globals, so
the original pattern transplants verbatim). Includes the contributor's
regression tests unchanged.
Two tests asserted resolve+reload events but acquired the lease
immediately (no wait). After gating the reload behind _lease_waited,
these tests must simulate a contended wait via on_wait(0.0) to
exercise the resolve+reload path.
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
Presence-only _delegate_from/_branched_from checks stopped the lease walk on
continuations that copied a delegate's model_config, so the first refresh
after rotation missed the parent-key lease and hard-interrupted. A failed
get_session probe also skipped acquire entirely. Walk the lineage inside
the write transaction and treat a probe error as contended, not a fresh
session.
Honor interrupts while waiting for admission, stop the turn when refresh
loses the lease, poll once per second under contention, and test dead-PID
reclaim.
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers
The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.
The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.
Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:
* /auth/native/authorize now accepts a supports_password provider:
register the pending broker authorization as usual, then 302 the
system browser to the interactive /login form with the opaque
broker_state in the gateway's PKCE cookie (the same server-controlled
channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
broker handle, a successful credential check completes the pending
authorization exactly like the /auth/callback native branch — mint
the one-time loopback code, return the loopback redirect (validated
loopback-only at authorize time) as `next`, clear the PKCE cookie,
and set NO session cookies. A lapsed broker is a clean 400 telling
the user to restart sign-in; a failed credential attempt leaves the
pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
session provider is registered (previously only for non-password
providers), so the desktop selects the system-browser strategy for
password-only gateways.
Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.
Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(dashboard-auth): bind native password completion to the authorize-time provider
Review follow-ups for #75808:
* /auth/password-login now enforces that body.provider matches the
provider recorded in the server-set PKCE cookie by
/auth/native/authorize before completing a pending native
authorization. /login renders a form for every session provider, so
without this a native flow started for provider A could be completed
with provider B's credentials, binding B's session into A's pending
entry. The mismatch is rejected BEFORE credential verification (no
session minted, no oracle) and preserves both the pending entry and
the cookie, so the user can still submit the correct provider's form.
Covered by a two-password-provider E2E regression test.
* Update the two docs spots that still said password-only providers do
not advertise native_pkce (website desktop-native-signin guide and the
auth_flows type comment in web/src/lib/api.ts).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: map contributor email for #75808 (buffpesos)
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
The messaging gateway multiplexes profiles over ONE shared launch-home
state.db, binding the profile per turn via the HERMES_HOME ContextVar
(copy_context into the worker thread). _agent_home derived the home from
db_path unconditionally, so on that lane the launch home stomped the
correctly-bound profile — deterministically inverting the leak #86313
fixed (found by @kshitijk4poor's post-merge probe; @helix4u flagged the
plugin-metadata half).
- _agent_home: bound override wins; session_db home is the unbound-thread
fallback
- _plugin_session_info: profile_name derives from _agent_home too
- full-prompt wiring regression (SOUL + skills + profile line on a bare
thread with the bot's DB) — reverting any call-site wire fails it;
multiplex, bare-thread, and plugin-metadata cases each pinned;
sabotage-verified both new tests fail against the merged behavior
- skills LRU cap 8 -> 32 (key is now per-profile x platform)
`hermes gateway setup` writes `gateway.platforms` as a LIST of
enabled platform names (e.g. `- telegram`), not a dict. Treat any
non-dict shape as "no per-platform overrides" instead of crashing
on `.get()` for every incoming turn (#83185).
Co-authored-by: SeashoreShi <seashore.shi@gmail.com>
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.
Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.
- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
rejected only when overwriting an existing regular file (new-PDF
creation stays allowed, matching the upstream split guard). Wired into
write_file_tool and patch_tool (replace + V4A Update/Add headers;
Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
end-to-end write_file/patch coverage incl. bytes-untouched assertions.
load_soul_md resolved the home ambiently, so a build thread that lost the
HERMES_HOME ContextVar read the launch profile's SOUL.md into another
profile's prompt — same class as the skills-index leak fixed in #86313.
load_soul_md and build_context_files_prompt now accept home_override, and
build_system_prompt_parts passes the agent's own home (from session_db)
at both SOUL call sites. Ambient behavior unchanged when no override.
Review feedback: the previous tests mocked get_hermes_home,
get_default_hermes_root and _resolve_active_profile_name, so they
checked template rendering but never exercised the relationship that
causes the defect — _resolve_active_profile_name returns a named profile
only when the active home is already <root>/profiles/<name>, which is
precisely why appending that suffix doubled it.
Replace them with tests that set a real HERMES_HOME under a tmp root and
mock no resolver. They assert the chain first
(_resolve_active_profile_name == "coder", get_hermes_home == the profile
dir, get_default_hermes_root == the root), then the rendered prompt.
Adds the default-profile branch the same way.
This also removes the order-dependent failures noted in the PR
description. Module-attribute monkeypatching stops taking effect once
other files in tests/agent/ have run, which is what made the mocked
tests fail in a full-directory run; the un-mocked tests are immune.
tests/agent/ now reports 87 failures — identical to the baseline on
unmodified main, with none in this file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The named-profile branch of the "Active Hermes profile" hint built its
paths by appending `/profiles/{active_profile}` to `get_hermes_home()`.
But `_resolve_active_profile_name()` returns a non-default name *only*
when `get_hermes_home()` has already resolved under `<root>/profiles/`
— that is how it derives the name in the first place. Both scoping
mechanisms (a `HERMES_HOME=<root>/profiles/<name>` env var and the
multiplexer's `set_hermes_home_override` contextvar) satisfy that, so
the suffix always doubled.
The same branch used `get_hermes_home()` for the *default* profile's
data pointers, where the root was intended — placing them inside the
active profile.
On a real 4-profile install the hint rendered:
reads and writes ~/.hermes/profiles/via/profiles/via/
default profile's data lives at ~/.hermes/profiles/via/skills/
against actual paths of `~/.hermes/profiles/via/` and `~/.hermes/skills/`.
Use the session home directly as the profile home, and
`get_default_hermes_root()` for the root pointers. Every named-profile
session was shipping a prompt that named nonexistent directories and
mislabeled this profile's own skills/plugins/cron/memories as the
default profile's — the exact cross-profile confusion the hint and
`classify_cross_profile_target` exist to prevent.
The default-profile branch is unchanged.
Fixes#72894
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On a correctly bound profile session get_hermes_home() returns the profile
dir itself, so relative_to(home/'profiles') never matched and every profile
misreported as 'default' (with wrong paths in the profile hint text). Use
get_default_hermes_root() for both the name derivation and the default-data
root string. Adds regression tests for the bound-profile and root-home cases;
sabotage-verified the bound-profile test fails against the old resolution.
A bot profile's system prompt could list the DEFAULT profile's ~80
skills and print 'Active Hermes profile: default' — while the live
skills_list() correctly showed the bot's real (often empty) set. The
agent plans against that index, so a false inventory makes it claim
capabilities it doesn't have, waste context tokens, and lose trust.
Root cause (confirmed empirically): the skills-prompt builder and the
active-profile line resolve the home through get_hermes_home(), which
reads a HERMES_HOME ContextVar. ContextVars do NOT propagate into
threading.Thread, so an agent build running on a thread that didn't
bind the profile's home falls back to the launch (default) home and
builds default's index. A bare no-override thread builds default's
full 7621-char block; the same thread with the fix builds empty.
Fix: resolve the agent's OWN home from its dedicated _session_db.db_path
(ground truth, ContextVar-independent) and pass it explicitly:
- build_skills_system_prompt(skills_dir_override=...) scopes the index,
the disk snapshot, and external-dir resolution to that home
- the active-profile line derives the profile name from the same home
Both fall back to ambient resolution when no db is present, so the CLI
and default-profile paths are unchanged.
Regression tests: an empty bot profile yields an empty skills block on
a bare thread even with ambient HERMES_HOME bound to a skills-rich
default; agent-home resolution from session_db.db_path. 3/3.
Port t_3778a491's in-flight stale-claim guard, absent from origin/main.
_submit_with_guard adds a job id to _running_job_ids before the future
that owns its release exists. Anything that hangs or dies between the
add and pool.submit (EAGAIN thread exhaustion on a substrate spike, or a
wedged SessionDB.__init__ on a stale sqlite flock) leaks the claim; every
later tick short-circuits with 'already running - skipping' silently - no
execution row, no last_error, no counter - until the gateway process
restarts. This wedged 4 recurring no_agent router/watchdog jobs (verdict-
router, wake-scanner, auto-review-router, blocked-task-notifier) for ~1h47m
on 2026-08-14 (t_20e23f84), cleared only by manual force-run.
- Record claim timestamp + pending-future sentinel in the same critical
section as the add; replace sentinel with the owning future after submit.
- sweep_stale_inflight() runs every tick (even idle) and force-releases
claims older than max(2*interval, 30m floor) with no live future: WARNING
cron.inflight.forced_release, get_inflight_guard_stats() counter, JSONL
record, and mark_job_run(success=False) so the wedge surfaces as last_error.
- Wrap the pre-future init (create_execution/copy_context) so an exception
there releases the claim immediately instead of leaking it.
- Finite-repeat jobs are released without mark_job_run so a forced release
never consumes a one-shot budget.
Scheduler-internal only: no provider/model routing, no credentials, no
spend, no guardrail weakening, no cron permission widening.
Tests: tests/cron/test_inflight_stale_guard.py (18), plus regression tests
for the recurring EAGAIN re-dispatch and the create_execution/pool-submit
leak paths. Full tests/cron/: 616 passed.
Follow-up on the #84795 salvage: the sequential deadline gets its own
resolver key. Unset, it inherits the concurrent batch deadline (same
value, same HERMES_CONCURRENT_TOOL_TIMEOUT_S bridge) so the two executor
paths cannot drift by default; set, it can be tuned or disabled
independently. Documented in cli-config.yaml.example; 5 contract tests.
Deliberately NOT on run_bounded_sync: the executors extend deadlines
dynamically during human approval waits (authorization-gate excluded
seconds) — the shared primitive is fixed-deadline. Noted in the docstring.