Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Perplexity Computer's July update let its agent manage sessions
conversationally from any surface — pin, archive, rename, fork — treating
session organization as operational infrastructure rather than a GUI
nicety. Hermes already has the durable pinned flag in state.db (Desktop
sidebar writes it; auto-archive honors it), but no CLI access existed:
GUI-only management was a single point of failure and blocked scripting
(issue #52955).
- hermes sessions pin <id...> / unpin <id...>: set/clear the durable keep
flag via SessionDB.set_session_pinned (whole compression lineage,
prefix resolution, multi-id, exit 1 on any miss)
- hermes sessions pinned [--json]: list all pinned conversations via the
include_pinned back-fill (old pins can't fall off a paging window);
--json enables backup/restore scripting
- docs: user-guide/sessions.md section
- tests: 6 tests covering prefix resolution, multi-id partial failure,
pinned-only filtering, JSON shape, empty hint
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:
- tools/checkpoint_manager.py: per-project agent-write ledger
(sha256 of every landed write_file/patch), safe_restore_plan()
classifier, and restore(safe=True) that reverts only agent-authored
changes, deletes agent-created files, and preserves user hand-edits.
Empty ledger (pre-existing stores) falls back to the classic full
restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
tweaks, agent-created file removal, empty-ledger fallback.
Sibling-test blast radius: main's TestInstallResolution fake_core stubs pin
the old _install_plugin_core signature; the scanning feature adds the
scan_decision_cb kwarg, so the real cmd_install call site now passes it.
Widen the three stubs to accept it.
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.
- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
pattern engine. Exempts the documented provider-plugin patterns (own
requires_env API-key reads, HTTP calls with keys) on code files while
keeping true threat signals (foreign credential-store access, reverse
shells, destructive/persistence/obfuscation patterns, prompt injection
in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
or --force), dangerous=blocked (--force does NOT override). Re-scan on
`hermes plugins update`; a dangerous updated tree is deactivated until
the user reviews the findings. Dashboard install path returns
structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
clones.
UTF-16 text files (Windows Notepad .txt, PowerShell > redirects) were
refused as binary: the terminal env decodes stdout as UTF-8 with
errors=replace, so their content arrived mangled with U+FFFD and
tripped the binary guard.
ShellFileOperations.read_file now probes the raw bytes via the
backend's Python when the binary guard fires: a BOM or the zero-byte
parity heuristic (derived from VS Code's encoding sniffer, tolerant of
mixed Latin/CJK content) identifies UTF-16 LE/BE, and the file is
transcoded to UTF-8 with CRLF normalized and the BOM stripped. Real
binaries (zeros at both parities), binary extensions, files over
10 MiB, and legacy 8-bit encodings (GBK, Big5) still refuse — a wrong
silent guess is worse than a clear refusal. Works on every shell
backend (local/docker/ssh) since the probe runs via python3 -c.
Tests run against a real LocalEnvironment (E2E, no mocks); sabotage
run confirmed 6/9 fail without the fix.
MCP tool results carry a server _meta mapping (exposed as .meta by the
Python SDK) alongside structuredContent. Servers return namespaced
machine-readable contracts there (validated payloads, browser-handoff
URLs); Hermes previously dropped the field entirely, so that data was
invisible to the agent.
Now _meta is included in the JSON tool output, after filtering
protocol-reserved keys per the MCP spec's key-name rules: a prefix is
reserved when a modelcontextprotocol or mcp label is followed by at
least one more label (modelcontextprotocol.io/..., tools.mcp.com/...).
Vendor namespaces with a trailing reserved word (com.example.mcp/...)
and unprefixed keys pass through. Non-serializable metadata drops the
extras rather than failing the call.
Unicode TAG characters (U+E0000-U+E007F) render as nothing in terminals
and chat UIs but are fully visible to LLM tokenizers, making them an
ASCII-smuggling prompt-injection channel for untrusted MCP servers.
- tools/ansi_strip.py: new strip_unicode_tags() with fast path; unlike
goose we preserve valid emoji tag sequences (U+1F3F4 base + tag spec +
U+E007F cancel), so regional flags survive.
- tools/mcp_tool.py: applied at every MCP text ingestion point — tool
result text blocks, embedded resource text, read_resource contents,
get_prompt message content, and tool descriptions entering the schema.
- tests/tools/test_unicode_tag_strip.py: smuggled-instruction vectors,
goose's test vector, emoji-tag-sequence preservation, ZWJ untouched.
Gemini 3+ models require explicit tool call IDs on functionCall /
functionResponse parts in replayed history; without them parallel tool
calls can be rejected or mispaired. The native adapter now:
- threads the model id into request building and includes ids for
Gemini >= 3 (version-gated: 2.x rejects unexpected id fields)
- preserves provider-returned functionCall.id on both non-streaming
and streaming responses instead of always minting a random one
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
Port from anomalyco/opencode#40869: parallel tool calls hitting the same
dangerous-command gate each enqueued their own _ApprovalEntry and fired
their own notify_cb — the user got N identical prompts and had to
/approve N times while the agent sat wedged.
_await_gateway_decision now detects an already-pending identical
approval (same command text + pattern-key set) in the session queue and
waits on the leader's event via _await_coalesced_leader instead of
re-prompting. Followers adopt session/always (persistence would auto-pass
a re-check anyway) and deny/timeout (re-asking a just-declined command is
prompt spam); a single-use 'once' makes the follower issue a fresh
prompt. Pre/post approval hooks fire with coalesced=True for followers.
Port from anomalyco/opencode#40707: connection-establishment and DNS
failure messages wrapped in generic exceptions (RuntimeError from local
shims, MCP bridges, SDKs re-raising without chaining) fell through to
FailoverReason.unknown, which misses the retry loop's eager transport
fallback — the full retry budget burned against a dead endpoint before
provider fallback.
New _CONNECTION_MESSAGE_PATTERNS (connect refused, no route, network
unreachable, DNS phrasings across Python/glibc/macOS/Node, fetch failed,
Envoy upstream connect error) classify as retryable timeout via
_classify_by_message, mirroring _TIMEOUT_MESSAGE_PATTERNS. Mid-stream
disconnect strings are deliberately excluded — they keep their
_SERVER_DISCONNECT_PATTERNS routing (large-session compression).
'✓ Update complete!' now shows what the update actually delivered:
'✓ Update complete! (v0.19.4 → v0.20.0)' when the pyproject version
changed, '(v0.20.0)' when commits landed within one release, and the
plain message when the version cannot be read. Reads the on-disk
pyproject.toml (not importlib.metadata, which still describes the old
install after a pull). Applied to both the git and Windows-ZIP paths.
macOS reports editor workspaces as /var/... while sessions are stored
under /private/var/... (same for /tmp vs /private/tmp), so the lexical
normpath comparison in _normalize_cwd_for_compare treated them as
different directories and ACP history filters silently dropped a
workspace's own sessions.
Canonicalize with os.path.realpath; nonexistent paths (e.g.
WSL-translated Windows drives on a Linux host) keep the previous
lexical behavior since realpath(strict=False) is lexical for them.
The casing/hidden/literal probes already ran the widened search to produce
their counts, then threw away the paths and returned a hint-only warning.
Strong models pivot in one turn; weak models spiral — the A/B eval measured
qwen3-coder-30b going 3.3 -> 9.3 turns on err_case_search, retrying casing
variants the probe had already resolved.
All three probes (case-insensitive, hidden/gitignored, literal-vs-regex) now
include up to 5 matched paths (+N more) in the warning via a shared tally
helper. Fixes the class, not the site.
Closes#80522
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.
When delegation.provider/model is explicitly pinned, the child no longer
inherits the parent's fallback chain: a mid-run auth/429 failure on the
pin previously rerouted the quiet-mode child onto parent fallback models
with no surfaced signal. Same treatment as the existing override_provider
OpenRouter filter-clearing — explicit pins are honored or fail loudly.
Also upgrades the pinned delegation.command-missing-from-PATH case from
warning + silent transport fallback to a loud spawn refusal, both at
credential preflight and in _build_child_agent.
Fixes#80450 (tracker #79686 audit item).
The 'Configure auxiliary models' menu under 'hermes model' now includes a
Delegation entry so the delegate_task subagent model is discoverable and
configurable interactively, instead of requiring hand-edited
delegation.provider / delegation.model keys in config.yaml.
Delegation is not an auxiliary_client task — subagents are full child
agents resolved via tools/delegate_tool.py — so the picker entry writes to
the top-level delegation.* section rather than auxiliary.*. 'auto' (inherit
the parent agent) is persisted as empty strings, never the literal 'auto',
which delegate_tool would try to resolve as a provider name. 'Reset all to
auto' clears only the four delegation routing fields and preserves
non-routing settings like max_concurrent_children.
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:
agent:
empty_response_guard:
enabled: true # false = legacy fixed 3-retry behaviour
cost_threshold_usd: 0.25 # per-attempt cost that halves the budget
- hermes_cli/config_defaults.py: new documented subsection under agent
(additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
section to (enabled, threshold) with fail-open tolerance for
malformed values; guard_enabled()/_cost_threshold_usd() now read the
init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
covering malformed sections, YAML string booleans, bad thresholds,
and a DEFAULT_CONFIG sync check; new integration test proving
enabled:false restores the legacy 1+3-call behaviour.
Requested by isak-ialogics on PR #75115.
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).
Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.
New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:
- Deterministic-empty detection: two consecutive empty attempts with
usage present, output_tokens == 0 (reasoning tokens count as output),
and identical (model, provider, finish_reason) skip the remaining
retries and go straight to the fallback chain — a different model may
well answer. Missing usage, nonzero output, or any signature change
keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
included/subscription routes are untouched.
At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.
Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.
Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.
Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.
Refs NS-503.
Port from Kilo-Org/kilocode#12698: report signal-terminated commands with
a human-readable note instead of a bare numeric exit code.
Kilo's fix settles a signal-killed process as the conventional 128+signum
exit code so its bash tool stops hanging. Hermes already produces numeric
codes for signal deaths (subprocess -signum, or the shell's 128+signum),
but the model saw a bare exit_code=-9 or 137 and burned turns
mis-diagnosing (137 = OOM kill being the most common). This adapts the
idea to Hermes' existing exit-code semantics tier:
- _interpret_signal_exit(): maps negative codes (definite signal death)
and the 128+signum band (hedged with 'usually') to a note naming the
signal and its likely cause, wired into _interpret_exit_code() ahead of
the per-command semantics table.
- Curated signal table (SIGKILL/SIGSEGV/SIGTERM/SIGABRT/...) so ambiguous
application exit codes are never mislabeled; uncurated 128+N codes stay
silent, SIGINT is excluded (executor's interrupt-marker path owns
rc=130).
- Notes surface via the existing exit_code_meaning result field.
E2E verified against real SIGSEGV/SIGKILL processes.
Kanban worktree workspaces were never removed by anything: _cleanup_workspace
preserved them by design, the CLI startup pruner explicitly defers t_* trees
to 'hermes kanban gc', and gc only sweeps scratch — so every worktree task
leaked its checkout forever (measured ~130GB on one estate).
- _cleanup_workspace now dispatches worktree workspaces to a new
_cleanup_worktree_workspace, which removes the worktree and its
auto-generated wt/<task-id> branch only when the tree is clean AND every
commit is reachable from a remote-tracking ref (reusing cli.py's
_worktree_is_dirty / _worktree_has_unpushed_commits predicates). Any
doubt preserves the worktree. dir workspaces stay untouched.
- The #33774 active-children deferral now covers worktree parents, and
_try_cleanup_parent_workspaces reaps deferred worktree parents when the
last child reaches a terminal state.
- archive_task reaps workspaces too; tasks archived without completing
previously leaked forever.
- 'hermes kanban gc' gains a backstop sweep for archived worktree tasks
that predate these hooks.
Tests: tests/hermes_cli/test_kanban_worktree_teardown.py (10 cases: removal,
dirty/unpushed/custom-branch/main-checkout/non-git preservation, complete/
archive integration, deferred-parent handoff).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to #80451. /api/status cleared gateway_platforms whenever the
gateway process was down — correct for a clean stop (stale 'connected'
states are noise) but wrong for startup_failed, where the fatal entries
ARE the diagnosis: per-profile credential collisions and auth failures
(multiplex '<profile>:<platform>' keys) that the single exit_reason
string cannot express. #80451's writer-identity and freshness filters
already drop entries from other/older processes, so preserving
fatal-state entries here cannot leak another gateway's live state.
Live-validated shape: a real multiplex gateway (2 secondary profiles,
rejected tokens) persists telegram / alpha:telegram / beta:telegram
fatals in gateway_state.json; /api/status previously reported {} for
platforms while state was startup_failed.
Rework of the #88049 inline early-return per review:
- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
checks the explicit XAI_API_KEY first and falls back to OAuth. The key
is read through tools.tool_backend_helpers.resolve_provider_secret
(config -> profile secret scope -> env/.env -> credential pool) so the
preferred path enforces the same scope policy as the existing fallback
branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
root cause for /v1/tts 403s (#87045, supersedes the inline shape in
#87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
the shared resolver; added coverage for the flag's OAuth fallback,
HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
wins (metered billing implication).
When a paid XAI_API_KEY is configured alongside xAI SuperGrok OAuth,
_resolve_xai_bearer() took the OAuth path unconditionally. The OAuth
credential authorizes /v1/responses but answers in a degraded Grok
explanatory mode with no citations, while the API key returns real
posts - so every x_search query silently degraded (#88040).
Prefer the explicit API key when set (same shape as the TTS fix for
#87045 in #87081), keeping OAuth as the fallback when no API key is
configured. The shared resolver and every other xAI call site are
unchanged.
When check_gateway_lifecycle refuses a cron script that lives on a
FileProvider path, the generic error implied the job contained a dangerous
gateway lifecycle command. Surface the real reason instead — the script
lives on a cloud-synced path whose evicted placeholder could hang the
preflight scan — while staying fail-closed. Regression test asserts the
cron-script scan path blocks without opening the file and that the message
names the cloud-synced path rather than a lifecycle command.
Widen #88052 per review:
- The walk-level short-circuit only protected _contains_unsafe_gateway_action;
the sibling caller _read_script_for_scanning still opened cloud-resident cron
scripts and could hang preflight. Move the check into _read_referenced_script,
the shared choke point, so every caller fails closed without opening.
- Generalize _is_apple_file_provider_path -> _is_cloud_placeholder_path: detect
~/Library/CloudStorage (Dropbox/OneDrive/Google Drive third-party FileProvider
domains) alongside iCloud's Library/Mobile Documents.
- Regression tests: CloudStorage lexical path blocked without open; the choke
point itself refuses cloud paths with os.open forbidden.
Cmd/Ctrl+Shift+B worktree flows on a remote gateway route through the
backend's /api/git mirror (hermes_cli/web_git.py), but that mirror had
drifted behind the Electron-local git ops the same UI drives locally, so
the flows broke exactly and only on remote connections:
- Convert-a-branch: the picker offers remote-tracking refs, and the
Electron op turns "origin/feature" into a local tracking branch. The
mirror ran `git worktree add <dir> origin/feature` verbatim, which
either fails or detaches HEAD. It now resolves the ref's remote via
git (never assuming "origin"), fetches best-effort, and creates the
worktree with `--track -b <short-name>`.
- branch_list omitted remote-tracking refs entirely and never set the
`isRemote` flag the renderer's HermesGitBranch contract requires —
the convert picker on a remote gateway couldn't reach a teammate's
branch and mislabeled every row's action.
- Branching off an `origin/…` base silently wired the new branch to the
remote upstream; the mirror now passes `--no-track` like the Electron
op does.
Renderer side, replace the silent degradation with a capability gate:
when a remote backend predates the /api/git worktree routes, worktree
creation failed with an opaque "Expected JSON … got HTML" toast. The
route-missing shapes now surface a clear "update the Hermes backend"
message (isGitEndpointMissingError, mirroring the sidebar batch-endpoint
detector); real git errors still pass through untouched.
Sibling audit (documented, no code change needed): repo status / review /
file-diff / git-root / default-cwd already route through desktopGit()'s
REST bridge or /api/fs on remote; repo scan is deliberately a no-op there.
Stale comments claiming "empty/false on a remote backend" in projects.ts
and coding-status.ts updated to describe the backend-routed reality.
Fixes#81724
#88048 documented the token-writer self-pin (bound-method thread target +
strong atexit hook) as a permanent contract: "__del__ never runs for
exactly the instances that leak". #88063 then removed both pins (idle
writer retirement + weakref atexit hook), making abandoned handles
eventually collectible.
Reword the __enter__ docstring and the context-manager test module
docstring to describe the pin as historical motivation, note the #88063
behavior, and keep the guidance that owners close deterministically.
No code changes.
The freshness window (updated_at >= live process create_time - 2s) had a
P1 boundary hole: a stale failure written by the PREVIOUS process
immediately before a fast restart landed inside the slack and was
aggregated; if that platform was then removed, the new process never
replaces the entry and NAS stays degraded indefinitely.
Replace clock heuristics with persisted writer identity:
- write_runtime_status now stamps every platform entry with the writing
process's (writer_pid, writer_start_time) — the same PID-reuse
fingerprint the liveness checks use, so a recycled PID never
masquerades as the original writer.
- The aggregation ownership filter requires exact equality between an
entry's stamp and the profile's validated live gateway process
(get_runtime_status_running_pid + _get_process_start_time). No slack,
no timestamps. Legacy entries without a stamp fail closed.
- Writer stamps are process recon (same class as the auth-gated
gateway_pid) and are stripped from all /api/status projections, both
active-profile and merged cross-profile entries.
Near-boundary regression test: prior-process entry stamped 100ms before
restart is excluded; recycled-pid-different-fingerprint excluded;
legacy no-stamp excluded; current-process entry kept.
Gateway startup deliberately preserves plain platform entries in
gateway_state.json across restarts, and the active-profile endpoint
compensates by filtering against current configuration. The cross-profile
aggregation copied raw maps, so a fatal entry for a platform the operator
had since disabled/removed could keep NAS reporting the instance degraded
indefinitely.
The aggregation has no cheap per-profile config context (platform sets
depend on tokens in each profile's .env behind its secret scope), so use
freshness instead: an entry is aggregatable only when its updated_at is
at/after the live gateway process's create time (validated PID via
get_runtime_status_running_pid + psutil create_time; the record's own
start_time field is a PID-reuse fingerprint in clock ticks, not a
timestamp). Config changes require a restart to take effect, so
restart-anchored freshness is exactly the config filter's semantics.
Fail closed: unparseable timestamps or no live process exclude the entry
— a false 'degraded forever' is the worse failure mode.
- /api/status now folds LIVE independent per-profile gateways' platform
failures (gateway_mode == 'multiple', the OOF-3 deployment mode) into
gateway_platforms under the validated <profile>:<platform> grammar, so
NAS fleet health sees them without a schema change. ?profile= requests
stay unmerged (single-profile view).
- Namespaced-key validation no longer fails open: colon-containing keys
are grammar-checked even when configured-platform loading throws.
- Platform key segment now accepts hyphens, matching plugin platform IDs
(plugins/platforms/<dir> names, e.g. foo-bar).
Scoped credential locks (Telegram bot token, Discord bot token, etc.) are
machine-global, but the conflict error only reported the holder's PID:
Telegram bot token already in use (PID 559). Stop the other gateway first.
On multi-profile hosts (e.g. hosted instances running 13 profiles), a bare
PID gives the operator no way to tell WHICH profile owns the credential —
the exact failure mode observed on zerocool-9781, where the 'default'
profile was misconfigured with the same bot token as 'lead-gen-outreach'
and logged an unattributable conflict every ~5 minutes (4,602 rows).
Fix:
- acquire_scoped_lock() now stamps a 'profile' label on lock records,
inferred from the process HERMES_HOME (<root>/profiles/<name> layouts,
'default' for the root home). Omitted when not inferable.
- New scoped_lock_owner_label() resolves the owning profile from a lock
record: prefers the explicit field, falls back to inferring from the
persisted hermes_home for locks written before the field existed.
Labels are validated against the profile-id grammar before use (lock
files are plain JSON on disk and the label flows into log lines and a
suggested CLI command).
- _acquire_platform_lock() conflict message now names the owning profile
and gives the correct remedy:
Telegram bot token already in use by the 'lead-gen-outreach' profile
gateway (PID 559). Stop that gateway first
(hermes --profile lead-gen-outreach gateway stop).
Records with no attribution signal keep the original PID-only wording.
Testing:
- New TestScopedLockOwnerLabel suite covering label inference (named,
Docker, root/default, unknown layouts), grammar validation, explicit-
field preference, hermes_home fallback, and legacy/malformed records.
- acquire_scoped_lock tests for profile stamping and omission.
- Adapter-level tests for profile-attributed, legacy-home-inferred, and
PID-only conflict messages.
- 76/76 targeted gateway tests pass; broad gateway suite failures are
baseline-identical (verified via git stash comparison). Ruff clean.
Independent review of the prior commit found the cache-invalidation key
alone doesn't fix the reported #88023 dead path: slash.exec runs as a
_LONG_HANDLER on the pool with a copied context, and no binding of
_HERMES_HOME_OVERRIDE happens between the transport read and the handler
body, so get_skill_commands() there always fell back to the process-level
HERMES_HOME regardless of which profile's session issued the request.
Bind the session's own profile_home around the get_skill_commands() check,
mirroring the same bind/reset-in-finally pattern already used at every
other per-turn HERMES_HOME scoping site (e.g. server.py's prompt-turn and
system-prompt-rebuild paths). This makes the #88023 dead path actually
reachable by the fix instead of only exercising the cache primitive in
isolation.
Switching Desktop profiles mid-session changes HERMES_HOME but not the
platform scope, so get_skill_commands() kept serving the previous
profile's skill list. A skill only available under the new profile then
looked like a cache miss to callers such as slash.exec, which fall
through to the slash_worker dead path (#88023).
MEDIA_TAG_CLEANUP_RE (and MEDIA_EXTENSIONLESS_TAG_RE) only recognized
ASCII terminators after a MEDIA:<path> tag. Chinese-language agent
output naturally writes MEDIA:D:\...\zhibao.pdf(782.6 KB)or ...pdf:内容 —
the full-width punctuation failed the trailing lookahead and the
attachment was silently dropped (cron even reported 'delivered') (#88038).
Both lookaheads now accept a CJK full-width terminator set (()〈〉《》:,。;
!?、curly quotes【】) alongside the ASCII set. The #68773 adjacent-tag
splitting guard is covered by a regression test.
A SessionDB handle cannot be released by dropping the last reference.
Once its background token writer starts, the instance pins ITSELF two
ways: the writer thread's target is a bound method, and
queue_token_counts registers atexit.register(_drain_token_queue_at_exit),
which only close() unregisters. A dropped-but-pinned handle keeps its
state.db/-wal/-shm descriptors for the life of the process, and __del__
never runs for it, so the existing safety net is dead code for exactly
the instances that leak.
That is why owning call sites are expected to close explicitly, in those
words, in the ownership comments in run_agent.py and
tui_gateway/methods_session.py. This adds the ergonomic half of that
contract so an owner can scope a handle and be exception-safe by
construction:
with SessionDB(path) as db:
db.append_message(...)
Purely additive. __enter__ returns self, __exit__ closes and returns
False so a caller's exception always propagates, and close() is already
idempotent, so a scope that closes early still exits cleanly. Nothing
changes for callers that already close directly.
Four regressions cover the scope closing the handle, __enter__ returning
the instance itself, the failure path closing while still propagating,
and an early close leaving the exit clean. They assert on the
sqlite_safe_read tracking registry rather than raw descriptor counts,
matching test_session_db_read_conn_pool.py, because SQLite's unix VFS
parks a closed descriptor on a per-inode reuse list and makes raw counts
lag the real connection count.
Refs #88033
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).
Split the unreachable path by durable operator intent:
- desired_state == "stopped" (written only by the s6 lifecycle commands;
the same intent signal container-boot reconciliation trusts) -> drop
the fire with 200 + a structured log line, mirroring NAS's own
instance_stopped drop. Jobs are not lost: the Chronos provider
reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
file without desired_state) -> keep the retryable 503, now stamped
with Retry-After: 60 so a scheduler that honors it spaces retries
past the wake/restart window instead of exhausting them inside it.
The gateway's own pass-through 503s (draining) get the same hint.
The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.
gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.