Commit Graph

14240 Commits

Author SHA1 Message Date
fangliquanflq 61645cde82 fix(agent): exempt clarify from sequential tool deadline
Clarify waits on a human for up to 3600s or unlimited. The generic sequential timeout was aborting that wait at 420s and leaving the prompt and worker active.
2026-08-15 02:21:39 +05:30
fangliquan 82a1b5a115 fix(agent): suppress late timeout observer events 2026-08-15 02:21:39 +05:30
fangliquanflq ededa8c4f1 fix(agent): bound sequential tool calls 2026-08-15 02:21:39 +05:30
algf 2912093e06 fix(dashboard): redraw TUI after PTY reattach 2026-08-14 13:51:26 -07:00
Teknium 20e5d51bea fix(browser): managed-first browser-use CLI resolution
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.

- _find_cli(): probe order flipped to managed bin -> PATH ->
  user-level tool dir (then uvx across the same order). A user's own
  uv tool install can no longer shadow the Hermes-managed copy with a
  drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
  install — only the managed copy does, so selecting any backend
  provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
  and always delegates to install_cli(), the single owner of the
  managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
  HERMES_HOME/bin/browser-use counts as installed.

Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.

Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
2026-08-14 13:51:07 -07:00
Brooklyn Nicholson f696380d07 test(desktop-update): guard the Windows hand-off python.exe contract
Source-level regression for the Windows Desktop update self-lock: assert every
Invoke-HermesStep call in scripts/desktop-update/windows.ps1 drives $pythonExe
via `python.exe -m hermes_cli.main`, never the hermes.exe shim. Driving the
update through the shim keeps hermes.exe mapped as a running image, so uv's
final `pip install -e .` shim rewrite fails with os error 32 and the update can
never complete. Runs on Linux CI (no PowerShell execution needed).

Co-authored-by: Sascha Haase <sascha.haase@textiletsg.com>
Co-authored-by: adamcap926 <adamcap926@users.noreply.github.com>
2026-08-14 15:48:31 -05:00
andyst-dev 898d786ad7 fix(update): count real behind commits in SSH fast path
Fixes #84851

The SSH fast path in _check_via_local_git compared only the exact tip SHA
of local HEAD against upstream main. When a local carried commit makes the
SHAs differ, it returned 1 ('behind') without checking ancestry, so an
ahead-of-origin checkout was misreported as '1 commit behind' — nudging the
user to run 'hermes update' and wipe their carried work. Fall back to
git rev-list --count HEAD..origin/main (mirroring the full-clone path) when
the SHAs differ.
2026-08-14 13:46:33 -07:00
AlexMnrs ca1b3f8705 fix(windows): avoid locale-sensitive update timestamps 2026-08-14 15:41:38 -05:00
teknium1 f79440e0f4 feat: /loop — recurring in-session wakeups (Claude Code parity)
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).

Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.

Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.

Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
  process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
  post-turn tick completion, and a supervised loop_wakeup_watcher that
  injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
  notification-poller wakeup driver + post-turn completion in the turn
  dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
  ticks defer until it finishes, pauses, or parks; real user input always
  wins over both

Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
2026-08-14 13:40:19 -07:00
Teknium d8d7cc068d fix(update): stop reporting bogus 'Found 9980 new commit(s)' on shallow installs
The hermes update APPLY path still ran an unconditional
rev-list --count HEAD..origin/<branch> — on a depth-1 installer checkout
that walks the truncated graph and reports the entire remote ancestry
(#53479's 'Found 9980 new commit(s)' on Windows 11). The zero/nonzero gate
stays (a 0 count is trustworthy on any graph); when the count is positive
on a shallow repo, recover the real number via the GitHub compare API
(added in PR #86257) and print count-free wording when that fails.
ahead_by==0 (local-ahead) falls through to the up-to-date path.

Completes the class fix from PR #86257 on its last remaining site.
2026-08-14 13:36:42 -07:00
kimyxx d29abb7e6b fix(browser): discover browser-use from user-level tool directories
Desktop/TUI workers can spawn with a minimal PATH that omits
~/.local/bin, the default location where uv tool install links the
browser-use binary. _find_cli() then failed to resolve an installed
CLI and Browser Use mode silently fell back to the built-in tools.

Probe the user-level tool dir (~/.local/bin on POSIX, APPDATA/uv/bin
on Windows) between PATH and the managed HERMES_HOME/bin, for both
the browser-use binary and the uvx fallback.

Salvaged from PR #83788 by @kimyxx onto current main; tests adapted
and extended with precedence and uvx coverage.
2026-08-14 13:32:03 -07:00
Brooklyn Nicholson 7de5634a2f test: attach RPCs complete while the agent is still building
Behavior contracts, not timings: each handler must return with the session's
agent_ready event still unset, the staged image must still reach the turn,
and an unknown session must still be rejected. Verified to fail against the
unfixed resolver (4 failed, 90s of real stalls) rather than only passing
against the fix.
2026-08-14 13:24:40 -07:00
Teknium 9442a718da fix(update-check): recover the real behind-count via the GitHub compare API
The honesty half (no fabricated counts) leaves shallow installs permanently
count-less. The compare API knows the full graph regardless of local clone
depth: GET /repos/<o>/<r>/compare/<current>...<target> returns ahead_by —
exactly the behind count the shallow boundary lost.

- hermes_cli/banner.py: _github_compare_behind() (bounded, unauthenticated,
  best-effort); wired into _check_via_rev and the shallow branch of
  _check_via_local_git. ahead_by==0 with differing tips = local-ahead => 0.
- hermes_cli/update_cmd.py: hermes update --check shallow path prints the
  exact count when recoverable, presence-only wording otherwise.
- apps/desktop/electron/update-count.ts: compareApiUrl() +
  parseCompareBehindCount() pure helpers; main.ts fetches the count when
  resolveBehindCount() returns null, and the SSH-official passive path stops
  fabricating behind:1 (uses compare API + updateAvailable flag).
- apps/desktop/src/lib/version-status.ts: updateAvailable now applies to the
  client target too, so a shallow desktop install shows '(update)' instead of
  nothing (or the old frozen '(+1)').

Fixes #84591; CLI siblings of #78253 / #53479 behavior.

E2E: live compare API returned 61/62 for real 61/62-commit gaps and 0 for the
reversed (local-ahead) pair; real shallow-clone fixture (depth-1 clone +
depth-1 fetch, merge-base broken) recovers the exact count with the API and
falls back to the honest sentinel offline.
2026-08-14 13:09:44 -07:00
Teknium 3bd98ec4bf fix(cli): redraw on terminal focus regain + docs for scrollback rebuild config
Completes the duplicated-chrome class fix:
- Focus-in (CSI I) now routes through the same rate-limited full-redraw
  recovery as Ctrl+L//redraw, clearing ghost prompt/composer copies after
  Alt+Tab / tab switches (focus-regain variant reported on #60920, #25337)
- Document display.cli_rebuild_scrollback_on_redraw in configuration.md
- Register the new default in hermes_cli/config_defaults.py (moved from
  the pre-refactor config.py location the salvaged commit targeted)
2026-08-14 13:09:37 -07:00
HunterSThompson c3d7f6eefd fix(cli): recover prompt_toolkit paint after tmux attach
Same-width SIGWINCH (typical tmux attach) skipped screen clear and
left previous_screen inconsistent, crashing redraw with
'cell' object has no attribute 'char'. Always clear on resize
recovery and retry _output_screen_diff with previous_screen=None
on AttributeError/TypeError.
2026-08-14 13:09:37 -07:00
angeon 88a1a9fd95 fix(cli): let redraw recovery rebuild scrollback 2026-08-14 13:09:37 -07:00
angeon 0e1cba326b fix(cli): honor persisted status bar visibility 2026-08-14 13:09:37 -07:00
halaprix a22fbba340 fix(cli): don't replay transcript on the session's first benign SIGWINCH
The resize recovery treated the first SIGWINCH of a session as a width
change (no prior width to compare against), running the Ctrl+L-style
viewport clear + _OUTPUT_HISTORY replay. The 2J clear preserves
scrollback, so everything in the deque printed a second copy below the
still-visible original. After --continue/--resume the deque holds the
whole "Previous Conversation" recap plus the first live exchange, so a
benign resize signal (GNOME Terminal tab bar appearing, monitor-scale
change, focus events) duplicated the entire conversation.

Seed the width baseline when the resize hook is installed, and replay
only on an observed width change. The baseline is read from app.output
— get_app() at install time is still the DummyApplication whose
DummyOutput reports a fake 80 columns, which would turn the first real
signal back into a phantom width change. A real initial maximize or
restore still differs from the seeded width and is still recovered
(#49120 behavior preserved; verified in a pty harness both ways).

Fixes #65293
2026-08-14 13:09:37 -07:00
Alli 6625c72c92 fix(cli): use _suspend_output_history for interrupt marker instead of clearing _OUTPUT_HISTORY
The original fix for #60920 cleared _OUTPUT_HISTORY in _recover_terminal_after_interrupt
to prevent the interrupt marker from being replayed on redraw. This approach:
- Discarded legitimate scrollback content unnecessarily
- Broke /redraw and Ctrl+L replay for any content after an interrupt

Instead, the interrupt marker is now printed via _cprint inside a
_suspend_output_history() context so it never enters _OUTPUT_HISTORY.
_recover_terminal_after_interrupt no longer needs to clear history — the
marker was never recorded, so _replay_output_history cannot duplicate it.

Also adds:
- _show_interrupt_marker flag to cleanly separate marker rendering from
  response construction
- Focused regression tests covering the marker recording suppression,
  history preservation after recovery, replay cleanliness, and flag logic

Fixes: #60941
2026-08-14 13:09:37 -07:00
Teknium 3885c1096a fix(gateway): stop internal bookkeeping writes from advancing the session activity clock
set_session_metadata() and advance_compression_session()'s repoint both
stamped entry.updated_at = now. updated_at is the user-activity clock that
drives idle/daily reset policy and the restart-resume freshness gate
(suspend_recently_active, #85709), so a background metadata write (e.g.
Slack thread watermark) or a background compression repoint on a long-idle
session could make it look freshly active and get it falsely resume_pending
after a gateway restart.

These are the last internal stamp sites after 784f733cf (recover) and
5462f689b (touch_activity gating): drop the stamps, keep the durable save.

Follow-up to #85895 (closed) — credit @GodsBoy for the report-side push and
@chelsealong for the analysis on #85709.
2026-08-14 13:08:39 -07:00
Teknium 1169fb50a4 fix: install Browser Use CLI for every browser backend except Camofox
The Browser Use CLI 3.0 is the primary driver engine for all browser
backends except Camofox, but only the explicit 'Browser Use' picker row
ran the install hook. Local Browser, Browserbase, Firecrawl, and the
Nous-managed cloud rows left the CLI uninstalled, so those selections
depended on the uvx zero-install fallback (first-use PyPI download
inside the tool-call timeout) or silently downgraded to the built-in
browser tools where uvx was unavailable.

- Extract the install logic into _ensure_browser_use_cli() and run it
  from the agent_browser/browserbase post_setup branch too (Firecrawl
  and the Nous cloud row both declare post_setup: browserbase).
- Camofox is untouched: Firefox-based, no CDP surface, cannot be driven
  by the CDP-only browser-use harness.
- Failure stays non-fatal: uvx fallback, then built-in tools.

Tests pin the contract: every browser post_setup key except camofox
attempts the install; camofox never does; install failure never raises.
2026-08-14 13:05:14 -07:00
kshitij d6a5cb9725 Merge pull request #85147 from kshitijk4poor/feat/unified-deadline-layer
feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
2026-08-15 01:10:40 +05:30
Teknium a90d5369f7 feat(gateway): dump wedged worker stacks when the turn reaper fires
When the inactivity reaper interrupts a timed-out turn, the interrupt
frees the blocked frame — destroying the only evidence of where the
turn was wedged. The Aug 2026 zombie-turn incident (WhatsApp session,
Relay-corrupted scope stack) wedged every turn for exactly the 1800s
timeout somewhere between 'Turn ended' and run_sync returning, and the
wedge point was unprovable post-mortem.

The reaper now logs the stack of every thread with turn-machinery
frames BEFORE interrupting, so the next occurrence names the exact
blocked line. Best-effort, bounded (8 threads, 25 frames), pure
in-process, never raises into the reaper.
2026-08-14 11:22:14 -07:00
Teknium afe09c7942 fix(gateway): exclude wedged turns from the restart after-turn wait
A turn idle past agent.gateway_timeout (the same threshold the turn
reaper uses) no longer defers an in-band restart. Restart is usually
the remedy for a wedged turn; waiting restart_after_turn_timeout on
one inverts the graceful path's purpose — a wedged WhatsApp turn
pinned 'hermes update' in draining until SIGTERM was sent manually
(Aug 2026). stop()'s bounded drain interrupts wedged turns instead.

gateway_timeout=0 (unbounded turns) disables wedge detection; cron
and API-server work has no per-turn activity clock and is never
counted as wedged; unreadable activity summaries fail open.
2026-08-14 11:22:05 -07:00
Soheil Fakour 0c6761c511 fix(gateway): restart_after_turn_timeout default 6h -> 30min (#79133)
The 21600s (6h) default shipped in #77184 makes an interactive
'hermes gateway restart' block for up to six hours when a turn wedges
(hung tool call, wedged event loop, stuck provider stream) — the exact
scenario the cap exists for. The intent (don't force-kill an agent
mid-turn) is sound, but the default must be a safety valve for hung
agents, not a target latency.

Lower to 1800s (30 min): still protects the overwhelming majority of
long autonomous turns (tool calls have their own timeouts well below
that), keeps worst-case interactive restart latency human-tolerable, and
users running very long unattended turns can raise it in config.yaml.

RED: new contract test fails on old 21600 default. GREEN: 4/4.
2026-08-14 11:22:05 -07:00
joaomarcos 11c5aae104 fix(compaction): gate checkpoint replay/prune on current request eligibility
A captured native-compaction checkpoint lives in the persisted
codex_reasoning_items sidecar, but the wire restructure that follows it
(prune_pre_checkpoint_items) ran unconditionally: the native gate only
decided whether context_management went into the request, and no signal
from it ever reached _chat_messages_to_responses_input.

So a single checkpoint kept deleting every pre-checkpoint item from all
later requests — after a mid-session swap out of the gpt-5.6 family,
after compression.enabled: false, after the rejection kill switch, and
after a session resume that reloads the sidecar from state.db. The model
receiving the opaque blob was no longer the one able to decode it, and
nothing was logged.

Thread a single native_compaction_eligible boolean, derived from the same
value that gates the context_management field, into the converter. When
ineligible: do not replay type: "compaction" items and do not prune. Safe
because native compaction never truncates Hermes' local history, so the
fallback still carries the full conversation.

All Responses call sites are covered: build_kwargs and convert_messages
derive the flag via _native_compaction_active, the auxiliary/compression
client is explicitly ineligible, and the converter defaults to False
(pre-feature wire) so future call sites are safe by construction.

Fixes #85914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 10:39:25 -07:00
Austin Pickett 29d0cc2602 fix(dashboard): treat Ctrl+C serve shutdown as a clean exit (supersedes #52970) (#85711)
* fix(dashboard): suppress Ctrl+C shutdown traceback

* fix(dashboard): extend clean Ctrl+C exit to the Windows serve branch

The Windows loop-factory branch (and its pre-0.36 asyncio.run fallback)
runs under the same uvicorn capture_signals() re-raise as the POSIX
path, so console Ctrl+C leaked the identical KeyboardInterrupt
traceback there. Guard both serve calls with the same clean-exit
contract, keeping the import-resolution try/except comment accurate
(genuine serve-time errors still propagate).

Also ports the reworded POSIX-test docstring (the serve path is no
longer 'byte-for-byte unchanged'), wraps the POSIX KI test in
pytest.fail so a regression reports red instead of aborting the pytest
session, and adds the windows_only sibling test.

Extends #52970 to the whole bug class.

* chore: map contributor email for @wangs1203

* test(dashboard): actually exercise the pre-0.36 Windows fallback KI contract

Copilot review caught that patching uvicorn._compat.asyncio_run with
raising=False makes the import succeed, so _runner is non-None and the
extra asyncio.run patch never covered the fallback. Split it out: a
dedicated windows_only test halts the _compat import (None in
sys.modules) so the fallback branch is genuinely selected, then asserts
the same clean-KI contract on bare asyncio.run.

---------

Co-authored-by: Emiya·Leon <wangs.coder@gmail.com>
2026-08-14 12:02:05 -04:00
kshitij 1b1975781f test: use tmp_path fixture for import test HERMES_HOME
The subprocess import test was creating .tmp-hermes-exec-ask-import/
in the repo root without cleanup. Switch to pytest's tmp_path fixture
so the temp directory is auto-cleaned and never appears as untracked.
2026-08-14 17:48:28 +05:30
xxxigm d6b4083f41 test(approval): cover CLI EXEC_ASK leak and fix slash-worker Path mock
Regression tests for silent pending_approval when ask-mode leaks into
interactive CLI, plus a Path-typed hermes_constants mock so the
slash-worker profile_home test survives per-file isolation.
2026-08-14 17:48:28 +05:30
Laura f84f1dd132 fix(tests): return Path from get_hermes_home mock in slash_worker profile test
test_slash_worker_accepts_profile_home mocks hermes_constants with
get_hermes_home=MagicMock(return_value="/tmp/hermes_test"), a str. In
production get_hermes_home() returns a Path, and hermes_state.py's
module-level DEFAULT_DB_PATH = get_hermes_home() / "state.db" does path
division. Under the str mock that becomes str / str, so importing
tui_gateway.server inside the patch raises TypeError and the test fails on
every main run (slice 4). Wrap the mock return in Path(...) so it matches the
real return type. Test-only; no production code change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-14 16:53:40 +05:30
webtecnica 56a41715dc fix: persist provider on model switch and add billing_provider fallback
Salvage of #79604 (webtecnica) + #85721 (pierrenode), combined and
rebased onto current main with simplify-code findings folded in.

#79604: update_session_model() wrote the model name to sessions.model
but never persisted the provider into model_config. On resume, the
runtime recombined the persisted model with the config.yaml primary
provider (which may not serve that model), producing auth errors.
Fix: add optional provider parameter to update_session_model, merged
into model_config via the shared _merge_model_config_json helper (not
hand-rolled SQL). Wire both gateway /model call sites to pass
result.target_provider.

#85721: session_gateway_runtime() had no billing_provider fallback.
A CLI session that never ran /model has no gateway_runtime or
top-level provider in model_config — billing_provider (written on
every session's first accounted API call) is the only durable record.
Fix: add billing_provider as the last-resort fallback in
session_gateway_runtime(), filtering bare billing buckets (auto/custom)
that are not routable identities.

Simplify-code findings addressed:
- Use _merge_model_config_json instead of 40 lines of branched SQL
- Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated
  as a set in tui_gateway/server.py)
- Merge None-filtering from #85920 with the billing_provider fallback
  into one coherent return path

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-08-14 14:39:08 +05:30
Teknium 16b54e2a0f fix(tests): resolve cost-guard fixture collision between pricing distrust and gpt-5.5-pro confusion nudge (#85970)
54cc39aa15 (distrust foreign pricing for custom providers) tested with
openai/gpt-5.5-pro fixtures; 83d373aae6 (salvaged #70324) made that exact
id warn unconditionally as a known-confusion model. Each was green alone;
together the distrust tests fail on every main run (slice 6).

Use a neutral fixture id for the distrust tests and add a regression test
pinning the composed behavior: the id-keyed nudge survives custom-provider
pricing distrust.
2026-08-14 02:03:38 -07:00
lkz-de 83d373aae6 fix(cli): guard expensive startup model overrides
Run the expensive-model warning for explicit startup `-m` / `--provider`
overrides before the chat loop starts, and fail closed for non-interactive
invocations that select an expensive or known-confusing model.

Also classify Nous paid-model 404s that say credits are required as billing
exhaustion so they fail fast with billing guidance.

Tested:
- scripts/run_tests.sh tests/hermes_cli/test_cli_startup_model_cost_guard.py tests/hermes_cli/test_model_cost_guard.py tests/agent/test_error_classifier.py -- --tb=short -q
2026-08-14 01:31:31 -07:00
Teknium f2678b8706 test: registry cost-guard passthrough uses a models.dev-trusted provider
The custom-provider pricing-trust fix makes provider="test" (not a
models.dev provider) correctly silent — use anthropic in the fixture.
2026-08-14 01:31:26 -07:00
Dustin Persek 54cc39aa15 fix(models): don't trust foreign catalog pricing for custom/unknown providers
Custom providers (custom:xxx) serve their own pricing; models.dev stores
OpenRouter prices for the same model ids. The cost guard fired on that
foreign pricing and blocked composer/CLI model switches on custom
providers with a wildly wrong warning (#54348).

expensive_model_warning now only trusts model_info/models.dev pricing
when the provider maps to a models.dev provider and the info's
provider_id matches, and only consults the pricing-entry lookup when
the billing route is known. Salvaged from #54422; the PR's desktop-hook
half predates the use-model-controls rewrite and is superseded by the
hook's existing rollback handling.
2026-08-14 01:31:26 -07:00
Turgut Kural 9a8af40192 fix(cli): run /model expensive-confirm off the main thread (#79401)
The command path (/model <name> --provider <p>) called
_confirm_expensive_model_switch() inline. That modal blocks its calling
thread on a response queue (see _prompt_text_input_modal); on the
prompt_toolkit main thread the TUI event loop freezes, the modal never
renders, and the switch silently cancels after the 120s timeout — the
user sees a frozen terminal and 'Model switch cancelled.' without ever
seeing the warning. The picker path already dispatched confirm+apply on
a worker thread; the command path now mirrors that contract.

Extract the inline confirm+apply block into
_confirm_and_apply_cli_model_switch() (preserving --once restore and
persist semantics) and dispatch it on a daemon thread when a TUI app is
present, keeping the synchronous path for non-interactive/test use.

Tests: new test_model_switch_confirm_thread.py pins (a) confirm runs off
the main thread when _app is present, and (b) the no-app path stays
synchronous. Existing _StubCLI helpers forward to the extracted method.
2026-08-14 01:31:21 -07:00
dsad 858a6008be fix(gateway): keep process notification routing off-loop 2026-08-14 01:21:52 -07:00
Teknium 7619564fbd fix(gateway): gate background-process completions on spawning-session boundary
Plain type=completion events built in _run_process_watcher carried only
session_key (chat/thread routing) with no spawning-session stamp, so after
/new (or a session switch) a completion notification from the OLD session
was injected into the chat's NEW session. Main already solved this exact
class for async delegations via the _classify_completion_target pre-flight
(_USER_BOUNDARY_END_REASONS drop on user-closed sessions, deliver on
idle-ends, follow the compression-tip chain), but the gate only ran for
type=async_delegation events.

Kernel salvage of #16455:

- Stamp the spawning conversation's session-db id (HERMES_SESSION_ID via
  session-scoped env) on the ProcessSession and the pending_watchers entry
  at spawn time in tools/terminal_tool.py; persist it through the process
  registry checkpoint/restore so recovered watchers keep the stamp.
- Thread the stamp into the completion_evt built by _run_process_watcher
  (watcher entry first, ProcessSession fallback for recovered watchers).
- In _deliver_completion_notification, run the SAME pre-flight classifier
  for stamped type=completion events: terminal -> drop with a log (output
  stays available via process(action='log')), retry -> False so the
  watcher re-polls, deliver -> proceed. The policy has exactly one owner
  (_classify_completion_target); nothing is forked. Unstamped legacy
  events keep today's deliver-always behavior, and the async-delegation
  path is untouched.

Based on the session-boundary approach from #16455 by @Tosko4 (original PR
was over-scoped across adapters/slash-commands/cron; this lands the kernel
only).

Tests: completion from a /new-closed session is dropped; completion after
an idle-end still delivers; unstamped legacy event delivers; retry verdict
returns retryable False without adapter injection; async_delegation gate
unchanged; stamp survives checkpoint recovery.
2026-08-14 01:09:42 -07:00
Teknium b9e7bead13 fix(gateway): coalesce same-tick async-delegation completions into one turn (#70300)
The async-delegation watcher drained the completion queue as a batch but
then delivered each event as its own synthetic turn, flooding the session
when a fan-out of background subagents finished together. Builds on the
per-process completion batching salvaged from PR #71898 (thanks
@yuzilongleif-collab) which coalesces concurrent _run_process_watcher
completions behind a short per-route fan-in window.

This commit adds the async-delegation half: group the drained batch by
full routing key (session_key + parent_session_id + platform/chat/thread/
user) and inject ONE consolidated turn per group. Durable-ack handling
stays honest: sibling rows are claimed up front via claim_event_delivery;
rows another consumer owns are excluded from the consolidated text (no
double-delivery); sibling claims are acknowledged only after adapter
acceptance and released (still pending) on failure. Events for different
sessions never coalesce, and a single-event group rides the existing
per-event path unchanged (latency and text identical).

Tests: 3 same-tick events -> exactly one adapter.handle_message carrying
all 3 results with all 3 durable rows delivered; 2 sessions -> 2 turns;
single-event path unchanged; failed batch releases claims and retries;
foreign-claimed sibling excluded and left pending.
2026-08-14 01:08:25 -07:00
yuzilongleif-collab a96cd10349 fix(gateway): force-redact coalesced completion output 2026-08-14 01:08:25 -07:00
yuzilongleif-collab 7536655b8f fix(gateway): own completion batch task lifecycle 2026-08-14 01:08:25 -07:00
yuzilongleif-collab cf09a30a9a fix(gateway): coalesce concurrent process completions 2026-08-14 01:08:25 -07:00
Teknium 8dc9401d7e fix(gateway): persist internal synthetic turns typed as internal_notification (#82888)
Async-delegation batch completions and background watch notifications
re-enter the gateway as synthetic MessageEvent(internal=True) turns via
_inject_watch_notification, but were persisted as bare role='user' rows —
indistinguishable from real user input in transcripts and the desktop UI.

Thread the event's internal flag through to persistence: when
event.internal is set, the turn's persisted user row is stamped
display_kind='internal_notification' (the existing DB-only presentation
sidecar used by auto_continue / model_switch rows). Wired through
_run_agent → _run_agent_inner → TurnContext → run_conversation's
persist_user_display_kind, and onto the three gateway-side fallback user
rows (transient failure, no-new-messages, pre-run crash), whose
append_to_transcript writer now forwards display_kind/display_metadata
to SessionDB.append_message.

Invariants preserved: role stays 'user' (alternation untouched), no new
injections, no past-context mutation, and display_kind is already popped
from every provider-bound copy in conversation_loop, so replayed sessions
never leak the marker to the API.

Regression tests: internal turn marked, real user turn unmarked, fallback
rows marked/unmarked per event, and a DB round-trip proving replay keeps
role/content intact while the provider copy drops the marker.
2026-08-14 01:08:18 -07:00
dsad 668396c2e1 test(security): add tests for notification-path redaction
Add TestNotificationRedaction class with two tests:

1. test_completion_notification_redacts_secret — verifies _move_to_finished
   redacts API keys in completion notifications before enqueueing
2. test_watch_match_notification_redacts_secret — verifies _check_watch_patterns
   redacts secrets in watch_match notifications before enqueueing

These tests cover the gap identified in #43025 where the explicit process
tool path (poll/log/wait) was redacted but the automatic notification
delivery path was not.
2026-08-14 01:08:13 -07:00
spfcraze d2d7766750 fix(tools,gateway): format watch_overflow events instead of dropping them
format_process_notification had no case for watch_overflow_tripped /
watch_overflow_released, so a watch-pattern notification flood surfaced
as '[IMPORTANT: Background process  exited (exit code ?)]' — a phantom
exit notification for a process that never existed — while the actual
'watch flood, N notifications suppressed' summary in the event's
message field was silently dropped. The gateway delivery path was
worse: _drain_gateway_watch_events retained only watch_match and
watch_disabled, discarding overflow events entirely before formatting.
Route both event types through the message field in the shared
formatter and the gateway formatter, and retain them in the gateway
drain.
2026-08-14 01:08:13 -07:00
Lidang-Jiang 9b554758bd fix(gateway): honor notification-off watch reinjection
Signed-off-by: Lidang-Jiang <lidangjiang@gmail.com>
2026-08-14 01:08:13 -07:00
fangliquanflq 8cf9e8a61b fix(compression): ignore background process notifications 2026-08-14 01:08:13 -07:00
Teknium 9166530942 feat(models): unify selection-time guards into one registry across all surfaces
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.

Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
2026-08-14 01:06:13 -07:00
Beto de Paola a06f1d7617 feat(models): warn on data-training tiers at model selection
muse-spark-1.2-contributor is heavily discounted BECAUSE Meta uses your
prompts and completions to train future models. Selecting it for the price
without realising the data trade-off is a footgun.

Add hermes_cli/model_data_policy_guard.py (mirrors model_cost_guard):
data_training_warning(model_id, provider, base_url) -> DataTrainingWarning|None,
driven by a vendor-agnostic rule table. The status is not machine-readable on
/v1/models or models.dev, so the v1 rule keys on the documented '-contributor'
model id (fires regardless of provider, so it also covers custom/gateway
routes). Message mirrors Meta's pricing-doc language and figures
(https://dev.meta.ai/docs/pricing-rate-limits/).

Wire it into the CLI model picker's confirm flow (auth.py) as a [y/N]
disclosure, chained after the expensive-model cost guard. Fires only on the
contributor tier; silent on muse-spark-1.1/1.2 and all other models.
2026-08-14 01:06:13 -07:00
kshitij 9cb456a9b9 fix: unify route dict or-None discipline in /model persist
The route dict in _persist_model_switch_to_session used  filtering
(omits falsy values) while the top-level keys used  (writes
explicit None to trigger deletion in _merge_model_config_json). This
asymmetry meant stale keys from a previous /model switch survived in
the nested gateway_runtime dict even after the fix in #85261 that
properly deleted them from the top-level keys.

Fix: build the route dict with  and derive the top-level keys
from **route so both shapes always use identical deletion semantics.
Also filter None values in session_gateway_runtime's reader since
gateway_runtime is replaced as a whole dict (not deep-merged), so None
values written by the persist path survive in the nested dict.

Found by /simplify-code 3-reviewer review on #85261 (all 3 reviewers
converged on the route dict asymmetry as the verdict-relevant finding).
2026-08-14 13:18:53 +05:30