Commit Graph

1351 Commits

Author SHA1 Message Date
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium eb8a30cc2f simplify(compat): skills_sync/skill_manager/kanban/computer_use — drop 37 re-exports/aliases, repoint 8 callers + 9 test files 2026-09-03 13:27:01 -07:00
Teknium 8e1a5b9f46 simplify(compat): tools/wake_word — drop 5 re-exports, repoint 1 test 2026-09-03 13:24:36 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium fb14bc4e11 review-fix(whitespace): strip trailing whitespace and EOF blank lines introduced by this PR
Trailing-whitespace-only edits so 'git diff --check BASE HEAD' is clean
(16 diagnostics across 11 files). No code changes.
2026-09-03 09:31:54 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 561b053f79 perf(agents): run per-child timers on one shared scheduler thread
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog).  A
profiled session with ~130 children was carrying ~1000 threads.  All
of these timers now run on a single process-wide daemon thread.

- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
  one Condition-driven daemon thread.  schedule(fn, interval) -> handle;
  handle.cancel(wait=) blocks for an in-flight run like the old join.
  A callback returning False stops itself; a raising callback is logged
  at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
  scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
  idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
  stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
  _lease_refresh_interval; lease-lost / refresh-error interrupt paths
  and the stop-event fencing are unchanged; the join(timeout=1.0) is now
  cancel(wait=1.0) so the interrupt clear still runs after any in-flight
  tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
  schedule(); the poll body is _tick(), same sampling state machine.

Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132.  At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
2026-09-03 02:44:24 -07:00
Teknium c96568f66c perf(delegation): finished delegate children no longer pin their transcripts in the parent heap
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:

1. bind_subagent_parent() stored the agent strongly in the
   `hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
   for its own turn, and every asyncio Handle/Future scheduled during
   that turn (LSP reader loops, kernel pipe transports) snapshots the
   Context — 56 live Contexts held 14 finished children after the bench.
   The ContextVar now holds a weakref (non-weakrefable doubles fall back
   to a closure); get_active_subagent_parent() dereferences it.

2. AIAgent.close() cleared _session_messages but not the
   _db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
   every successful DB flush) nor _streamed_assistant_text_parts, so the
   agent — kept alive by (1) — retained every message dict. close() now
   drops both.

The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.

Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
2026-09-03 02:35:37 -07:00
kshitijk4poor 914d8a0bd6 fix(recovery): name the real state.db in the copy-pasteable recovery banners
The gateway broadcast and turn-failure explanation printed a literal
~/.hermes/state.db; now that the line is a command the operator is meant to
run as-is, interpolate _default_db_path() so profile / HERMES_HOME installs are
pointed at the store that actually failed. Also: split a comment that a merge
fused onto the logger line in session_lost_and_found.py, and fix an inverted
test docstring.
2026-09-03 11:28:21 +05:30
sal a15f96450b fix(recovery): make the printed salvage command satisfy the real CLI contract
Review blocker on e62940d: every state-db guidance site printed

  hermes sessions recover --source <db>

but cmd_sessions rejects that shape with exit 2 ("--output is required
unless --inspect-only is used") before any snapshot is taken — the user
follows the instruction during a corruption incident and gets nothing.

All five state-db sites now print the established two-stage operator
contract (the same shape `sessions repair` failure output and
docs/state-db-recovery.md already use):

  hermes sessions recover --source <db> --inspect-only
  hermes sessions recover --source <db> --output recovered-state.db

with the stop-the-gateway precondition stated for the gateway/turn
banners, and --inspect-only leading in the hermes_state refusal strings
(inspection before writing anything).

New TestEmittedCommandsSatisfyCliContract dispatches the exact emitted
flag shapes through the real cmd_sessions and asserts they pass the
contract gate (rc != 2) on a scratch DB, plus a premise test pinning
that the v1 no-flag shape is still rejected with rc 2 — so a guidance
string can never again pass a source-substring test while the command
it prints deterministically fails.

Noted for merge order: #101423 and #101168 also touch
hermes_cli/session_recovery.py. They are complementary recovery-integrity
work, not duplicates of this guidance/gate fix; whichever lands second
should rebase and rerun the lost_and_found + session-recovery suites.

(cherry picked from commit 34dc59a284509e76a0342c36d03a2a437aa8a3b9)
2026-09-03 11:28:21 +05:30
sal 5d9a2110ba fix(recovery): stop pointing sqlite3 .recover guidance at the live state.db
Refs #100368. The forensics thread established that a sqlite3 CLI with
the WAL-reset opener bug (fixed 3.51.3+ / backports 3.50.7 / 3.44.6;
Debian/Ubuntu system shells 3.45.1/3.46.1 are in the vulnerable band)
unlinks the live -wal/-shm pair when pointed at a live state.db whose
writer's DMS lock has been cancelled, splitting the store into two
concurrent generations whose acknowledged writes vanish while both
report integrity_check ok. Hermes' own corruption banners instructed
exactly that command.

- gateway corruption broadcast, run_agent corrupt-cause explanation,
  hermes_state repair-budget and forensic-backup refusals, and the
  kanban manual-recovery hint now route operators to
  `hermes sessions recover --source <db>` (which snapshots the damaged
  bundle before any shell touches it) and warn against a raw sqlite3
  shell on the live file
- find_sqlite3_cli() now refuses a WAL-reset-vulnerable shell for the
  page-level salvage lane even on the snapshot, reusing the canonical
  gate from hermes_cli.sqlite_runtime so the embedded runtime and the
  salvage shell can never disagree
- find_sqlite3_cli_refusal() records why a shell was refused so the
  lost_and_found lane can tell the operator exactly what to install
  instead of a generic "not found"
- regression tests cover the version gate (vulnerable/fixed matrix, the
  mirror check), every refusal reason, and each guidance site

Test plan:
- scripts/run_tests.sh tests/hermes_cli/test_sqlite3_cli_salvage_gate.py
  tests/test_state_db_repair_loop_cap.py
  tests/run_agent/test_corruption_recovery_guidance.py
  tests/hermes_cli/test_session_recovery_lost_and_found.py
  tests/hermes_cli/test_session_recovery.py tests/test_sqlite_wal_reset_gate.py
  tests/hermes_cli/test_sqlite_runtime.py - 91 passed, 1 skipped locally

(cherry picked from commit e62940d1021e80e9b7d6423ced1cbdfe7dd0c37d)
2026-09-03 11:28:21 +05:30
Teknium d68a4c01a7 refactor(run_agent): restore main() docstring verbatim (fire renders it as --help output) 2026-09-02 22:51:39 -07:00
Teknium 55385a8111 refactor(run_agent): collapse defensive layers in close/todo-hydration/listing paths; -31 LOC 2026-09-02 21:53:57 -07:00
Teknium 116ca1db1e fix: sibling Nous 401 recovery adopts a peer's refresh instead of rotating again
The stampede fix added a `stale_access_token` hint to
resolve_nous_runtime_credentials() so a process whose bearer just 401'd
adopts a token a sibling already rotated instead of re-POSTing the shared
grant — but only the credential-pool caller passed it. The main agent's
401 path (run_agent._try_refresh_nous_client_credentials), the auxiliary
client rebuild, and the proxy adapter all called force_refresh=True with
no hint, so `_already_rotated_by_peer` could never fire: N subagents
hitting hourly expiry still issued N serialized refreshes, each one
invalidating the token a sibling had just adopted.

Live 12-process A/B against a fake Portal: 12 refresh POSTs / 9 distinct
final tokens before, 1 POST / 1 token after.
2026-09-02 21:22:07 -07:00
Teknium a208c541f1 refactor(run_agent): fold codex/sanitization forwarders into staticmethod aliases, pack multi-line signatures/calls; -198 LOC 2026-09-02 21:04:11 -07:00
Teknium 7a56696c8d refactor(run_agent): keep review-defer/engine-end helpers module-level so bound-method test harnesses still work 2026-09-02 18:56:20 -07:00
Teknium b45961212c refactor(run_agent): compact docstrings/comments, collapse remaining boolean ladders and redundant locals 2026-09-02 18:35:08 -07:00
Teknium 477a9b46e3 refactor(run_agent): phase helpers for close()/main(), shared engine-hook/quiet helpers, collapsed defensive layers 2026-09-02 18:27:32 -07:00
Adolanium 7caee2898b perf(agent): stop rebuilding the streamed reply text on every delta
`_record_streamed_assistant_text` grew the turn's visible text with `+=`
on an attribute. Python only grows a string in place when the target is a
local variable, so this copied the whole text on every delta.

The loop runs once per streamed token, so a reply of length N costs about
N squared in copying. A 200 KB answer arriving in 4-character deltas
moves several billion characters and burns seconds of CPU in the loop the
file itself calls the hottest one in the agent.

The text is now kept as a list of pieces and joined when read. Reading
happens at turn end and on interrupt, not per delta, so the whole turn is
linear in the length of the reply. `_fire_stream_delta` used to join on
every token just to ask if the text was empty. That check now looks at
the parts list.

`_current_streamed_assistant_text` becomes a property over that list, so
the seven readers and the call sites that clear it between turns keep
working unchanged. Reading does not collapse the pieces, because a delta
landing between the join and the write back would be lost.

Measured with 8-character deltas: adding 20k deltas to an already long
text took 3.1 times as long as the first 20k before, and 0.9 times after.
2026-09-03 04:46:41 +05:30
Teknium 0ab2e9672c refactor(run_agent): extract VisionMessagePrepMixin and ReasoningParamsMixin 2026-09-02 13:29:40 -07:00
Teknium 9576b49c0b refactor(run_agent): extract TurnFacadeMixin (run_conversation/chat admission wrapper) 2026-09-02 13:29:40 -07:00
Teknium f0526e6d74 refactor(run_agent): extract CompressionFacadeMixin; collapse 10 more lazy forwarders 2026-09-02 13:29:40 -07:00
Teknium 45e88aff38 refactor(run_agent): re-wrap comment blocks left with dangling fragments 2026-09-02 13:29:40 -07:00
Teknium 1e778ae7c6 refactor(run_agent): re-wrap compacted docstrings, restore lost issue-ref rationale 2026-09-02 13:29:40 -07:00
Teknium e564e60a3d refactor(run_agent): extract SessionPersistenceMixin (agent/session_persistence.py); _DB_PERSISTED_MARKER sourced from context_compressor 2026-09-02 13:29:40 -07:00
Teknium 81abe4799a refactor(run_agent): extract InterruptControl/TurnExplainers/ActivityTracking/RateLimitCredits mixins 2026-09-02 13:29:40 -07:00
Teknium ce6b4d1e16 refactor(run_agent): extract ApiErrorSummaryMixin (agent/api_error_summary.py) 2026-09-02 13:29:40 -07:00
Teknium 64cc8cf0d3 refactor(run_agent): extract ApiRequestHooksMixin (agent/api_request_hooks.py) 2026-09-02 13:29:40 -07:00
Teknium c274125690 refactor(run_agent): extract StreamDeliveryMixin and StatusOutputMixin 2026-09-02 13:29:40 -07:00
Teknium fd54e7faec refactor(run_agent): extract ClientLifecycleMixin (agent/client_lifecycle.py) + agent/lazy_forward.py 2026-09-02 13:29:40 -07:00
Teknium 5423336e8c refactor(run_agent): collapse 39 lazy forwarder methods into _forward() descriptors; __init__ forwards via locals() 2026-09-02 13:29:40 -07:00
Teknium 1c21f09558 refactor(run_agent): route default-header chain -> _ROUTE_DEFAULT_HEADERS dispatch table 2026-09-02 13:29:40 -07:00
Teknium 7048763e99 refactor(run_agent): unify non-vision message preprocessors, drop dead forwarders 2026-09-02 13:29:40 -07:00
Teknium 30e060211c refactor(run_agent): compact verbose comments and docstrings in AIAgent facade
Hand-reviewed compaction only; code is AST-identical. Rationale-bearing
sentences kept (issue-ref rationale restored in a follow-up commit).
2026-09-02 13:29:40 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
kshitijk4poor 2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
salch-cred bbed304536 fix(agent): exempt :cloud GLM models from stop->length truncation rewrite
Ollama cloud models (model name contains ':cloud') run generation on
Ollama's hosted server; the local 11434 endpoint is only a transparent
proxy that forwards finish_reason faithfully.  _is_ollama_glm_backend()
was matching these models because the proxy listens on the same port
as local Ollama, causing _should_treat_stop_as_truncated() to rewrite
a correct finish_reason='stop' into 'length'.

The 4-attempt continuation loop then injects a synthetic user nudge
that reasoning-capable GLM models spend their output budget deliberating
over, producing unpunctuated tails that re-trigger _has_natural_response_ending()
rejection -- a self-reinforcing loop that always exhausts retries.

Fix: add an early return in _is_ollama_glm_backend when the model name
contains ':cloud'.  Local GLM inference (no ':cloud') is still caught
by the existing port/URL/provider checks.

Fixes #98406
2026-09-01 23:27:10 -07:00
JonthanaHanh 5360886f54 fix(gateway): exclude Ollama Cloud from GLM truncation detection; propagate partial flag (#72316)
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:

1. _is_ollama_glm_backend() matched "ollama" in base URL, which
   included Ollama Cloud (ollama.com). The hosted service correctly
   reports finish_reason and is not affected by the local Ollama
   stop-reason bug.  Exclude "ollama.com" before the substring check.

2. _handle_session_chat_stream() hardcoded "partial": False in the
   assistant.completed SSE event instead of reading result.get("partial").
   The WebUI could not detect truncation and rendered partial responses
   incorrectly (showing only the continuation instead of the full text).
   Read the partial flag from the agent result, matching the pattern
   used by other SSE paths in the same file.

Fixes #72316
2026-09-01 23:27:10 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium 9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
kshitijk4poor 45bd48ff67 fix: initialize affinity_token before the turn-lease early returns
The turn-lease timeout/interrupt paths return from inside the try block
before set_affinity_scope() runs; the finally then read an unassigned
local -> UnboundLocalError. This was the cause of the 4 red
cross-process lease tests on PR #97158's CI.

(cherry picked from commit 05a3c8a4caa59831b01d13f3950e61385aadbfa5)
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
kshitijk4poor f20bbfa40d fix: a declined liveness abort must not cancel a pending compression
Closes the #99758 review P1 (andrexibiza): with a generation claim in
play, `interrupt()` called `_admit_hard_cancel()` BEFORE the claim was
validated at the final mutation edge, and the production
`CompressionCommitFence.cancel_before_commit()` irreversibly sets
`_cancelled = True` whenever no commit has started. So a watchdog abort
that ultimately DECLINED (real progress landed in the window, claim went
stale) had already killed the recovered turn's legitimate pending
compression commit — `begin_commit()` refuses a cancelled fence forever.
Generation authority covered interrupt publication but not the
compression-fence mutation that preceded it.

Split hard-cancel admission into two halves:

- `_wait_for_compression_commit()` runs pre-claim and is NON-mutating:
  it only blocks when `commit_in_flight` is true (the started-commit
  branch of the production fence waits for `finish_commit` without
  cancelling), so the interrupt still publishes only after an in-flight
  SessionDB mutation has finished — exactly as before.
- `_cancel_pending_compression_commit()` runs AFTER
  `_consume_claim_and_publish_first_state()` survives, so the
  destructive pending-commit cancellation can never outlive a stale
  claim. If a commit crossed its boundary in between, it is no longer
  fence-cancellable and completes on its own.

Regression coverage (both use the real `CompressionCommitFence`):

- `test_declined_abort_does_not_cancel_pending_compression_commit`:
  parks the interrupt at the claim-reservation release, lands real
  progress (G+1), lets the interrupt decline, then proves
  `fence.begin_commit()` still admits. Red on the pre-fix tree
  (mutation-checked: the fence was left cancelled).
- `test_declined_abort_parks_and_leaves_fence_operational`: the
  in-flight-commit window variant — activity lands while the interrupt
  waits on a started commit; the interrupt declines and a fresh
  `begin_commit()` still admits afterwards.
- The round-6 witness (`...resumes_inside_interrupt_publication`) now
  models an in-flight commit (`commit_in_flight = True`) so its park
  point stays inside the pre-claim wait, matching the new admission
  shape.

Also updates the `interrupt()` docstring for the deferred destructive
cancellation.
2026-09-01 03:19:59 +05:30
kshitijk4poor c394b005fc fix: publish watchdog settlement only after the abort commits
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.

- Split the surface: `_surface_stall` is now observational only ("no
  progress for Ns; attempting recovery"), and the definitive
  aborted/lease-stopped settlement moves to a new
  `_surface_committed_abort` that runs only after `_commit_abort`
  succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
  turn whose aborts keep declining no longer re-logs an ERROR and
  re-warns the user every poll interval.
- Add the committed-path regression test
  (`test_watchdog_publishes_definitive_settlement_only_after_commit`)
  and extend the declined-path witness
  (`...resumes_during_warning`) to assert no committed-abort or
  definitive pre-commit claim appears when the abort is vetoed. Both
  fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
  unconditionally, no generation claim — losing the lease means the
  process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
  WHY, drop the round numbering), and drop the dead
  `cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.

On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
2026-09-01 03:19:59 +05:30