Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).
The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.
Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).
agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).
toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.
providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.
agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.
model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.
Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog). A
profiled session with ~130 children was carrying ~1000 threads. All
of these timers now run on a single process-wide daemon thread.
- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
one Condition-driven daemon thread. schedule(fn, interval) -> handle;
handle.cancel(wait=) blocks for an in-flight run like the old join.
A callback returning False stops itself; a raising callback is logged
at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
_lease_refresh_interval; lease-lost / refresh-error interrupt paths
and the stop-event fencing are unchanged; the join(timeout=1.0) is now
cancel(wait=1.0) so the interrupt clear still runs after any in-flight
tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
schedule(); the poll body is _tick(), same sampling state machine.
Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132. At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:
1. bind_subagent_parent() stored the agent strongly in the
`hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
for its own turn, and every asyncio Handle/Future scheduled during
that turn (LSP reader loops, kernel pipe transports) snapshots the
Context — 56 live Contexts held 14 finished children after the bench.
The ContextVar now holds a weakref (non-weakrefable doubles fall back
to a closure); get_active_subagent_parent() dereferences it.
2. AIAgent.close() cleared _session_messages but not the
_db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
every successful DB flush) nor _streamed_assistant_text_parts, so the
agent — kept alive by (1) — retained every message dict. close() now
drops both.
The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.
Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
The gateway broadcast and turn-failure explanation printed a literal
~/.hermes/state.db; now that the line is a command the operator is meant to
run as-is, interpolate _default_db_path() so profile / HERMES_HOME installs are
pointed at the store that actually failed. Also: split a comment that a merge
fused onto the logger line in session_lost_and_found.py, and fix an inverted
test docstring.
Review blocker on e62940d: every state-db guidance site printed
hermes sessions recover --source <db>
but cmd_sessions rejects that shape with exit 2 ("--output is required
unless --inspect-only is used") before any snapshot is taken — the user
follows the instruction during a corruption incident and gets nothing.
All five state-db sites now print the established two-stage operator
contract (the same shape `sessions repair` failure output and
docs/state-db-recovery.md already use):
hermes sessions recover --source <db> --inspect-only
hermes sessions recover --source <db> --output recovered-state.db
with the stop-the-gateway precondition stated for the gateway/turn
banners, and --inspect-only leading in the hermes_state refusal strings
(inspection before writing anything).
New TestEmittedCommandsSatisfyCliContract dispatches the exact emitted
flag shapes through the real cmd_sessions and asserts they pass the
contract gate (rc != 2) on a scratch DB, plus a premise test pinning
that the v1 no-flag shape is still rejected with rc 2 — so a guidance
string can never again pass a source-substring test while the command
it prints deterministically fails.
Noted for merge order: #101423 and #101168 also touch
hermes_cli/session_recovery.py. They are complementary recovery-integrity
work, not duplicates of this guidance/gate fix; whichever lands second
should rebase and rerun the lost_and_found + session-recovery suites.
(cherry picked from commit 34dc59a284509e76a0342c36d03a2a437aa8a3b9)
Refs #100368. The forensics thread established that a sqlite3 CLI with
the WAL-reset opener bug (fixed 3.51.3+ / backports 3.50.7 / 3.44.6;
Debian/Ubuntu system shells 3.45.1/3.46.1 are in the vulnerable band)
unlinks the live -wal/-shm pair when pointed at a live state.db whose
writer's DMS lock has been cancelled, splitting the store into two
concurrent generations whose acknowledged writes vanish while both
report integrity_check ok. Hermes' own corruption banners instructed
exactly that command.
- gateway corruption broadcast, run_agent corrupt-cause explanation,
hermes_state repair-budget and forensic-backup refusals, and the
kanban manual-recovery hint now route operators to
`hermes sessions recover --source <db>` (which snapshots the damaged
bundle before any shell touches it) and warn against a raw sqlite3
shell on the live file
- find_sqlite3_cli() now refuses a WAL-reset-vulnerable shell for the
page-level salvage lane even on the snapshot, reusing the canonical
gate from hermes_cli.sqlite_runtime so the embedded runtime and the
salvage shell can never disagree
- find_sqlite3_cli_refusal() records why a shell was refused so the
lost_and_found lane can tell the operator exactly what to install
instead of a generic "not found"
- regression tests cover the version gate (vulnerable/fixed matrix, the
mirror check), every refusal reason, and each guidance site
Test plan:
- scripts/run_tests.sh tests/hermes_cli/test_sqlite3_cli_salvage_gate.py
tests/test_state_db_repair_loop_cap.py
tests/run_agent/test_corruption_recovery_guidance.py
tests/hermes_cli/test_session_recovery_lost_and_found.py
tests/hermes_cli/test_session_recovery.py tests/test_sqlite_wal_reset_gate.py
tests/hermes_cli/test_sqlite_runtime.py - 91 passed, 1 skipped locally
(cherry picked from commit e62940d1021e80e9b7d6423ced1cbdfe7dd0c37d)
The stampede fix added a `stale_access_token` hint to
resolve_nous_runtime_credentials() so a process whose bearer just 401'd
adopts a token a sibling already rotated instead of re-POSTing the shared
grant — but only the credential-pool caller passed it. The main agent's
401 path (run_agent._try_refresh_nous_client_credentials), the auxiliary
client rebuild, and the proxy adapter all called force_refresh=True with
no hint, so `_already_rotated_by_peer` could never fire: N subagents
hitting hourly expiry still issued N serialized refreshes, each one
invalidating the token a sibling had just adopted.
Live 12-process A/B against a fake Portal: 12 refresh POSTs / 9 distinct
final tokens before, 1 POST / 1 token after.
`_record_streamed_assistant_text` grew the turn's visible text with `+=`
on an attribute. Python only grows a string in place when the target is a
local variable, so this copied the whole text on every delta.
The loop runs once per streamed token, so a reply of length N costs about
N squared in copying. A 200 KB answer arriving in 4-character deltas
moves several billion characters and burns seconds of CPU in the loop the
file itself calls the hottest one in the agent.
The text is now kept as a list of pieces and joined when read. Reading
happens at turn end and on interrupt, not per delta, so the whole turn is
linear in the length of the reply. `_fire_stream_delta` used to join on
every token just to ask if the text was empty. That check now looks at
the parts list.
`_current_streamed_assistant_text` becomes a property over that list, so
the seven readers and the call sites that clear it between turns keep
working unchanged. Reading does not collapse the pieces, because a delta
landing between the join and the write back would be lost.
Measured with 8-character deltas: adding 20k deltas to an already long
text took 3.1 times as long as the first 20k before, and 0.9 times after.
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).
Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).
Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.
Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.
Fixes#100661Closes#97766 (overflow-force idea; the bundled continuation changes were not taken)
Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.
Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".
Refs #90837, #90950, #97940, #89332, #45383
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.
Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
Ollama cloud models (model name contains ':cloud') run generation on
Ollama's hosted server; the local 11434 endpoint is only a transparent
proxy that forwards finish_reason faithfully. _is_ollama_glm_backend()
was matching these models because the proxy listens on the same port
as local Ollama, causing _should_treat_stop_as_truncated() to rewrite
a correct finish_reason='stop' into 'length'.
The 4-attempt continuation loop then injects a synthetic user nudge
that reasoning-capable GLM models spend their output budget deliberating
over, producing unpunctuated tails that re-trigger _has_natural_response_ending()
rejection -- a self-reinforcing loop that always exhausts retries.
Fix: add an early return in _is_ollama_glm_backend when the model name
contains ':cloud'. Local GLM inference (no ':cloud') is still caught
by the existing port/URL/provider checks.
Fixes#98406
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:
1. _is_ollama_glm_backend() matched "ollama" in base URL, which
included Ollama Cloud (ollama.com). The hosted service correctly
reports finish_reason and is not affected by the local Ollama
stop-reason bug. Exclude "ollama.com" before the substring check.
2. _handle_session_chat_stream() hardcoded "partial": False in the
assistant.completed SSE event instead of reading result.get("partial").
The WebUI could not detect truncation and rendered partial responses
incorrectly (showing only the continuation instead of the full text).
Read the partial flag from the agent result, matching the pattern
used by other SSE paths in the same file.
Fixes#72316
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).