Commit Graph

14240 Commits

Author SHA1 Message Date
pierrenode a1ddb54840 fix(gateway): tag the loop-liveness and heartbeat-poll tasks as permanent supervised watchers (#84558)
#84327 excluded _spawn_supervised's permanent watchers (session-expiry,
kanban, reconnect, the scale-to-zero watcher itself, ...) from
_scale_to_zero_has_live_background_work() via a _hermes_supervised_watcher
tag, because counting them made an armed gateway consider itself busy
forever and never go dormant.

Two more permanent, infinite-loop tasks are added to _background_tasks
OUTSIDE _spawn_supervised and were untagged:

- _loop_heartbeat_task (loop_heartbeat_forever, #66892): a `while True`
  loop started unconditionally in start() on every gateway boot. Extracted
  the inline spawn block into _start_loop_heartbeat_task() so it's
  independently testable, matching the existing _start_heartbeat_poller()
  pattern.
- _heartbeat_poll_task (_poll_loop in _start_heartbeat_poller): also a
  `while True` loop, started the first time a session registers a
  heartbeat watch, and then permanent for the rest of the process.

Because _loop_heartbeat_task starts on every boot, it alone made
_scale_to_zero_has_live_background_work() return True forever on every
armed instance, regardless of the #84327 fix -- confirmed empirically
against the real method with the exact untagged-task shape this task has.

Two new regression tests spawn each task through its real production
entry point and assert the busy check returns False; both fail against
the unfixed code (missing method / real assertion failure).

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-08-20 20:24:40 +10:00
Ben Barclay 6cd1ed2e78 test(relay): rename misnamed precedence test; document the flat-key fallback nuance
test_flat_key_wins_over_subblock asserted the OPPOSITE of its name (the
sub-block wins, matching _relay_slack_extra). Rename to what it proves.
Also note in _resolve_cron_surface_mode why its fallback differs from
_relay_slack_extra's all-or-nothing sub-dict: the flat key is the legacy
staging shape, and a flat knob applies to every fronted platform, gated
only by the per-platform D6 capability check.
2026-08-20 20:13:29 +10:00
Ben Barclay 162b23c3e2 fix(relay): D6 in_channel capability gate resolves the destination platform's descriptor
RelayAdapter.supports_inchannel_continuable is a scalar adopted from the
PRIMARY identity's handshake descriptor, but one RelayAdapter fronts N
platforms and the connector advertises the bit per platform. Reading the
scalar for every logical platform both leaked a Slack-primary True onto
other fronted platforms (activating the flat surface their descriptor
never advertised) and suppressed a non-primary platform's advertised
True (forcing thread mode on capable Slack behind a Discord primary).

Add supports_inchannel_continuable_for_platform(platform): resolves the
platform's own negotiated descriptor via descriptor_for_platform (the
same Phase 1.5 seam max_message_length uses), scalar fallback only when
the per-platform descriptor is unavailable. The scheduler's D6 gate
prefers the query when the adapter provides it; native adapters keep
the class-attribute path byte-identically.

Tests: two-platform descriptor matrix (primary-True no-leak,
non-primary-True honored, unknown-platform scalar fallback).
2026-08-20 20:12:52 +10:00
Ben Barclay 79c39025c0 fix(relay): format hints resolve the DESTINATION platform, and stamp on send_for_platform
Two gaps in the block-formatting hint stamping:

1. Wrong descriptor: _format_hints gated on self.descriptor — the PRIMARY
   identity's scalar — while one RelayAdapter fronts N platforms. A
   Slack-primary adapter stamped Slack hints onto known Discord chats; a
   Discord-primary adapter suppressed hints for Slack chats whose own
   negotiated descriptor advertised the bit. Resolve per destination:
   send/edit use _descriptor_for_chat (the same seam max_message_length
   already uses) plus the chat's logical platform for the config
   sub-block; the knob lookup is now per-logical-platform
   (platforms.relay.extra.<platform>.*) instead of hardwired to slack.

2. Missing lane: send_for_platform — the scheduled/persisted-home lane
   (gateway/delivery.py), i.e. the CRON delivery path, the flagship
   consumer of the in_channel brief — never stamped hints at all. Stamp
   there too, resolving descriptor_for_platform(logical) off the
   transport; the scalar descriptor is used only when it belongs to that
   exact platform (fail closed).

Tests: Slack-primary/Discord-chat no-leak, Discord-primary/Slack-chat
still-stamps, send_for_platform stamps for capable platform and stays
clean for incapable — all against a two-platform negotiated-descriptor
transport. Existing single-platform suite unchanged and green.
2026-08-20 20:10:20 +10:00
Ben Barclay 20c56f82b8 fix(cron): in_channel thread-flatten uses the seed's gate (origin_target)
The seed was decoupled from the mirror opt-in (in_channel is the
continuation surface regardless of attach_to_session), but the
thread-id-clearing gate above it still read mirror_this_target. With the
advertised default config (attach_to_session=false, cron.mirror_delivery
unset) and an origin carrying a real thread_id, the brief kept delivering
INTO the origin thread while the flat (thread_id=None) session got
seeded — brief and continuation surface in different places, so a plain
reply never saw it.

Flatten on the same gate as the seed: origin_target (with the existing
live_adapter_ready guard). Fan-out/broadcast targets are unaffected.

Test drives _deliver_result with a thread-carrying origin and default
knobs, asserting on the routed DeliveryTarget.thread_id — RED on the old
gate, GREEN now.
2026-08-20 20:07:44 +10:00
Teknium e69b8e561d feat: consolidate 'hermes version' into 'hermes --version', remove the subcommand
'hermes --version' (and -V) now prints the full version report — banner
version line with upstream SHA, install directory, authoritative install
method, Python and OpenAI SDK versions, and update status — making the
separate 'hermes version' subcommand redundant. The subcommand is removed.

- _startup_fast.print_fast_version_info() is now THE canonical version
  printer: static lines print instantly from stdlib probes, then the
  banner label, install-method resolver, and update check lazy-import
  after the first line is on screen (each degrades gracefully).
- main.py _print_version_info() delegates to it (used by /version in the
  CLI chat surface and the --version flag path); the old duplicate
  implementation is deleted.
- hermes_cli/subcommands/version.py removed; parser wiring, subcommand
  sets, console-engine extraction entry, and tests updated. Hermes
  Console keeps a 'version' command wired to the shared printer.
- Termux fast paths now include update status too (previously
  check_updates=False).
- Docs/i18n, CONTRIBUTING, SECURITY, and nix checks updated to
  'hermes --version'.
2026-08-20 03:05:46 -07:00
Teknium d425658d27 Merge pull request #90688 from NousResearch/feat/keyed-failure-one-shot-rescue
feat: failing keyed web backends rescue onto the keyless ring for one call, never sticky
2026-08-20 03:05:40 -07:00
Ben Barclay 8b6cf434cb fix(cron): carry Slack workspace scope_id into continuable seed keys
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.

Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.

Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
2026-08-20 20:04:49 +10:00
Teknium 1647d030cf fix(update): call out GitHub rate limiting/outage on fetch 429 instead of a generic failure
A GitHub-side HTTP 429 during 'hermes update' printed only
'Failed to fetch updates from origin.' — and the curl
'unable to access ... returned error: 429' shape even matched the
network-error branch, blaming the user's connection for a GitHub
outage.

- new _classify_fetch_failure(): 429/rate-limit -> 'GitHub is rate
  limiting requests or having an outage — try again in 5 minutes';
  5xx -> outage message with githubstatus.com; ordered BEFORE the
  generic 'unable to access' network check
- both fetch-failure sites (update apply + --check) now share the
  classifier via _print_fetch_failure(), and both always print the
  first raw stderr line so the wire error stays diagnosable
- tests: classifier matrix + E2E against a live local HTTP server
  returning 429 through real git

Fixes #89287
2026-08-20 02:45:57 -07:00
Teknium aca40d1d63 fix(sessions): error paths return non-zero exit codes (delete/rename/prune/import) 2026-08-20 02:06:05 -07:00
Teknium d1eefe6acc feat: keyed web backends get a one-shot keyless rescue on failure — never sticky
When the chosen/keyed backend fails a web_search or web_extract call
(bad key, upstream outage, 5xx, raised exception), that single call
retries on the keyless free-tier ring instead of erroring. The next
call attempts the chosen backend again — no sticky failover, no state.
Resolves the keyed half of #78984/#32159 (keyless half landed in the
ring PR).

- tools/web_tools.py: _rescue_eligible (keyed ring vendors + non-ring
  backends eligible; keyless-mode calls excluded — they already walked
  the ring), _rescue_search/_rescue_extract (search annotates
  rescued_from + backend_error naming the original failure and the
  retry-next-call semantics; extract rescues only whole-batch failures,
  partial failures pass through untouched; rescue failure preserves the
  ORIGINAL backend error with the rescue note appended)
- both dispatchers wrap the provider call: failure-results AND raised
  exceptions rescue; ineligible paths re-raise unchanged
- web.keyless_rescue config key (default true; implicitly off when
  keyless_fallback is off); docs updated

Live E2E: keyed Tavily with an invalid key 401'd and the call was
served by the real ring with the rescue annotation; a second call
re-attempted Tavily first (statelessness proven); whole-batch extract
rescue returned real page content. 13 new tests; 67 green across the
keyless suites.
2026-08-20 02:04:32 -07:00
Teknium 0a8a4cdb3d fix(backup): friendly error on unwritable output path instead of raw traceback 2026-08-20 01:47:06 -07:00
Teknium 138e482ed0 fix(config): set coerces negatives/whitespace/null and rejects malformed keys 2026-08-20 01:46:54 -07:00
Teknium 76653a8eba fix(sessions): prune/archive spare pinned sessions by default (data loss) 2026-08-20 01:46:20 -07:00
Teknium f796239c6b Merge pull request #90572 from NousResearch/feat/keyless-tavily-firecrawl-failover
feat: keyless web tier is now a 5-vendor free rotation (Exa/Parallel/Tavily/Firecrawl/Keenable) with ring failover + honest doctor readiness
2026-08-20 01:46:16 -07:00
Teknium 3f67921fca fix(cli): /undo typo no longer quits the CLI; +3 slash-command papercuts 2026-08-20 01:46:08 -07:00
Teknium 33e64fd83c test: capture worktree messages through the _cprint route 2026-08-20 01:45:56 -07:00
Teknium e7e3885113 test: pin ring entry vendor in provider-routing tests (ring rotation made direct-callable mocks stale) 2026-08-20 00:20:17 -07:00
Teknium 90e477d3ed Merge remote-tracking branch 'origin/main' into feat/keyless-tavily-firecrawl-failover 2026-08-20 00:18:29 -07:00
Teknium 4ea69d9d2c feat: keyless web tier becomes a 5-vendor round-robin ring (adds Tavily, Firecrawl, Keenable)
Fresh installs with zero web credentials now rotate web_search/
web_extract across FIVE vendors' public free tiers — Exa, Parallel,
Tavily, Firecrawl, Keenable — instead of a 2-vendor 50/50 split, with
next-in-line ring failover on rate limits (multi-hop until a vendor
serves or the ring is exhausted; served_by marks the actual vendor).

- plugins/web/keenable/: new bundled provider (search via /v1/search,
  fetch via /v1/fetch; keyed Bearer or keyless with the mandatory
  X-Keenable-Title app header). Credit: integration proposed by
  Ilya Gusev (Keenable) in #49758; Free/Paid picker rows included.
- keyless_mcp: tavily/firecrawl/keenable keyless search+extract
  wrappers, _KEYLESS_RING + per-process round-robin cursor (seeded by
  the random session id, advances per unpinned request), pinned-vendor
  entry (pin = start there; rotation off), paid-pinned vendors excluded
  from the ring entirely.
- Tavily/Firecrawl providers route keyless traffic through the ring;
  both are now default-on ring members (no longer selection-gated).
- web_tools/registry: keenable in backend sets, auto-detect, availability
  probes; _keyless_preference() delegates to the ring cursor.
- KEENABLE_API_KEY in OPTIONAL_ENV_VARS; docs updated (ring semantics).

Live E2E: all 10 vendorXcapability paths (5 search + 5 extract) served
real results keyless; rotation cycled all five vendors over 5 dispatch
calls; double-throttle failover walked exa->parallel->tavily.
2026-08-20 00:17:25 -07:00
Teknium c2f5d2da21 test: vary marathon-turn fixture args — identical calls now legitimately dedupe to stubs 2026-08-20 00:16:22 -07:00
Teknium 761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
kshitijk4poor cef999c5cf refactor(agent): consolidate uncompressed-overflow guard to one warn site + re-arm
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
  unconditionally computed request_pressure_tokens (zero marginal cost,
  covers turn-start AND mid-turn growth) — drops the duplicate every-turn
  estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
  under the window, so warn -> /compress -> regrow warns again (the dedup
  was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
  len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
  ~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
  sites (context_compressor always exists; its context_length property
  hard-floors positive; the fallback would have been a synchronous
  network probe mid-turn that also bypassed config overrides) and the
  undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
  fake reimplementation left them uncovered) and add dedup, re-arm,
  no-rearm-while-over, and multimodal-gate coverage.
2026-08-20 12:35:20 +05:30
joaomarcos db5d5dffea fix(agent): guard against uncompressed session overflow when compression is disabled (#89297)
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.

Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
2026-08-20 12:35:20 +05:30
Teknium fc9b4186a0 fix(worktree): pruner reaps rebase-merged trees via PR state; worktree add survives disk contention
Two gaps behind the recurring 'hermes -w timed out after 30 seconds':

1. Rebase-merge leak: git cherry only catches patch-identical commits.
   Salvage flows routinely change the diff (conflict resolution, follow-up
   commits), so 12 of 22 'unpushed' trees on the incident box had MERGED
   PRs and were preserved forever. The pruner now falls back to
   'gh pr list --head <branch> --state merged' — authoritative, memoized
   on (branch, head_sha) with True-only caching, fail-safe to preserve.

2. Creation timeout 30s -> 120s: the ~10k-file checkout measured 113s at
   near-zero CPU under multi-agent disk contention vs 1.2s idle. 30s
   killed legitimate creates and threw away completed work.
2026-08-20 00:00:28 -07:00
Teknium f309f92d30 feat: hermes worktree list/prune — attended reclaim for accumulated worktrees and merged branches
The startup pruner is deliberately conservative (unattended, pre-banner),
so real installs accumulate what it can never touch: trees preserved for
untracked-only scratch, and orphaned local branches beyond the two
auto-generated prefixes it deletes. A measured multi-agent box: 35 trees /
15GB / 244 local branches, 120 of them fully merged.

New attended surface (hermes_cli/worktree_gc.py + worktree_cmd.py):
- hermes worktree list — audit every tree: age, size, verdict, reason,
  plus deletable-branch count
- hermes worktree prune [--dry-run|--trees-only|--branches-only]
- /worktree prune [--dry-run] — same engine in-session; never touches the
  session's own active tree
- startup escalation: one WARNING when .worktrees/ exceeds 10 trees or
  5GB, naming the reclaim commands (silence is how boxes hit 15GB)

Safety invariants (shared with the startup pruner via cli.py primitives):
tracked modifications and unique unpushed commits never deleted at any
age; live-locked trees untouched; branch deletion gated on worktree
removal success; untracked-only scratch ARCHIVED to
~/.hermes/archive/worktree-prune/ before its tree is reaped.

Branch GC is content-gated, not name-gated: any local branch fully merged
or git-cherry patch-equivalent upstream is safe to delete (rebase merges
rewrite SHAs, so --merged alone misses the dominant leak); unique-commit,
checked-out, protected, and stale-base (>50 ahead) branches are kept.
Classification is parallel (8 workers) — 244 branches audit in ~64s live.

git timeouts degrade to keep (returncode 124) instead of crashing the
audit — live-verified failure on a 746MB .git repo.

16 behavior-contract tests against real git fixtures; live dry-run on the
production repo: 12 trees reclaimable, 120 branches deletable, 0 false
positives among kept trees.
2026-08-19 23:55:20 -07:00
kshitijk4poor 446da3ef56 fix(cli): widen lock-bit coverage to alias installers, legacy nav keys, and PUA functional keys
Follow-up to the salvaged #89676 + #90291 lock-bit fixes: extract a shared _lock_variants() helper and cover the sites both PRs missed - install_shift_enter_alias / install_ctrl_enter_alias / install_cmd_backspace_alias CSI-u spellings, legacy CSI-letter and CSI-tilde navigation twins derived from the existing table for ALL modifiers 1-16 (not just plain/shift), plain F1-F4 SS3 fallback, unmodified CSI-u keys (Tab/Enter/Space/Backspace), and kitty PUA functional keys (keypad, F13-F24, Ignore range) under lock bits. 8 new tests.
2026-08-20 12:16:20 +05:30
liuhao1024 118dbe871f fix(cli): map kitty CSI-u lock-bit variants so key combos survive NumLock
kitty and ghostty OR the CapsLock (64) / NumLock (128) state into the
CSI-u modifier parameter. With NumLock on, Ctrl+C arrives as
ESC[99;133u (5 + 128) instead of ESC[99;5u; the alias table had no
entry for it, so every key combo leaked as literal text like
[127;133u (#89651). Install every CSI-u alias with the lock-bit
variants (+64/+128/+192); the xterm modifyOtherKeys encoding never
carries lock bits, so the ESC[27;N;CP~ form is left untouched. The
Esc-key registration now covers modifier 1 as well (1+128=129 is a
lone Esc with NumLock on).
2026-08-20 12:16:20 +05:30
Teknium 01c3bd4c81 test: explicit keyless Firecrawl selection asserts the keyless cloud route
The salvaged #50659 behavior makes 'firecrawl selected, no creds' a
WORKING keyless state, so the old expectation (hard error naming
FIRECRAWL_API_KEY) is stale. The test now mocks httpx and asserts the
request routes to api.firecrawl.dev with results returned — still
proving keyless Tavily can't silently take over, which was the test's
point. Also stops the test making a real network call in CI.
2026-08-19 23:20:52 -07:00
Teknium a92412ede1 test: mock load_config_readonly in memory status gate tests
check_memory_requirements() reads the readonly config loader; the salvaged
tests only patched load_config, so the gate saw the real config.
2026-08-19 23:15:01 -07:00
Cassie Gray b38c40319d fix(cli): align hermes memory status and docs with memory tool gate 2026-08-19 23:15:01 -07:00
kshitijk4poor 45f11263bd fix(tui): skip the kitty protocol push for Ghostty in the Ink TUI too
Widen the cli.py Ghostty exception to the sibling sites the review found: the Ink TUI pushes CSI >1u at raw-mode entry (App.tsx), on alt-screen exit, and on the extended-keys re-assert path (ink.tsx) for every EXTENDED_KEYS_TERMINALS entry including ghostty - same Alt-stripping bug. New skipKittyKeyboardProtocol() helper in terminal.ts gates the ENABLE push at all 3 sites; the DISABLE (pop) stays unconditional since popping an empty stack is a spec no-op. Also fix the cli.py comment citing the modifyOtherKeys encoding where the kitty CSI-u form (ESC[127;3u) is what the broken path expected, dedupe the quadruplicated Ghostty comment, and update the stale 'mirroring the Ink TUI' docstring. 7 new vitest cases.
2026-08-20 11:39:03 +05:30
kshitij 1a8fea3ce2 fix(cli): skip Kitty keyboard protocol push for Ghostty, use modifyOtherKeys only
Ghostty's Kitty disambiguate-mode implementation strips the Alt modifier
from the Backspace key — Option+Backspace arrives as bare \x7f instead of
the expected \x1b[27;3;127~, breaking backward-kill-word.  This was a
regression introduced when PR #87630 re-added the CSI >1u Kitty protocol
push for all allowlisted terminals including Ghostty.

Under modifyOtherKeys mode (CSI >4;2m), Ghostty correctly sends
\x1b[27;3;127~ for Option+Backspace, which the alias table in
pt_input_extras already maps to (Escape, ControlH) = backward-kill-word.

Fix: for Ghostty only, push just modifyOtherKeys and skip the Kitty
protocol push.  All other terminals (iTerm2, WezTerm, kitty, tmux, VS Code)
still get the full dual-protocol push.

Ghostty upstream tracking: discussion #9560, issue #9895 (cmd+backspace
variant of the same root cause).
2026-08-20 11:39:03 +05:30
kshitijk4poor b7e12decc6 fix(agent): route relay-wrapped output-cap 429s into the output-cap handler
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
  provider fallback (a deterministic request-shape failure that failover
  cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
  clamp+compress recovery as plain output-cap 400s.

Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
2026-08-20 11:37:01 +05:30
ekinnee 99c980f466 fix(model-metadata): parse 'exceeds model maximum output tokens' cap errors
Recognizes the DeepSeek/OpenAI-compatible relay wording
  max_tokens (98304) exceeds model's maximum output tokens (65536)
in both parse_available_output_tokens_from_error (returns the cap) and
is_output_cap_error (keeps the 400 out of the compression death-loop).

Salvaged from PR #72283; the conversation_loop early-clamp block was
dropped in favor of routing through the existing output-cap handler
(follow-up commit).
2026-08-20 11:37:01 +05:30
Dhruv Modi 7f2733b71c fix(model-metadata): converge output-cap retry on vLLM
Fixes the retry loop that spins forever when a vLLM server rejects a
request for having a max_tokens too big for what is left of the context
window.

The catch is that vLLM does not tell you how big your prompt actually is
in that situation. It works the number backwards from the constraint it
just failed, so you get:

    "requested 65536 output tokens and your prompt contains at least
     36865 input tokens, for a total of at least 102401 tokens"

That 36865 is just window + 1 - requested, and the total is always
exactly window + 1. Subtracting it from the window hands back
requested - 1 every single time, whatever the real prompt size is.

parse_available_output_tokens_from_error believed it and returned
requested - 1. conversation_loop then takes off its 64 token safety
margin and retries, which walks the cap down 65 tokens at a time while
the reported input walks up by the same 65:

    65536 -> 65471 -> 65406 -> 65341

Three attempts is the default budget, so the session gives up with
"Context length exceeded" having closed 195 tokens of a roughly 28000
token gap. Compression cannot save it either, because the input was
never the problem, which is why the compressor keeps refusing with
"summary would have GROWN".

This is also what is behind the unexplained "input-token drift" in
issue #61761. The input is not drifting. It is a derived number, and it
moves because we moved max_tokens.

So when that shape shows up (the "at least" wording, plus a budget that
works out to exactly requested - 1), halve the requested cap instead. It
is still guaranteed to sit under whatever was just rejected, and it
converges on the first retry: 65536 -> 32768, which next to a real 36865
token prompt comes to 69633 against a 102400 window.

Nothing else moves. A measured input is still trusted, and a genuine
input overflow still returns None so the caller falls through to
compression the way it always did.

The existing test asserted the bogus 65535, so it is updated. Added
tests for the measured input path, and for the retry actually
converging.
2026-08-20 11:37:01 +05:30
kshitijk4poor 0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs 5c03fbedc6 fix(config): recognize memory nudge interval 2026-08-20 11:36:54 +05:30
MindDragonLabs fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30
liuhao1024 62016a1b0a fix(compression): count a would-grow refusal as an ineffective strike
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).

Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.

Fixes #88568
2026-08-20 11:36:54 +05:30
Ben Barclay 5a17b1f41d Merge pull request #89584 from victor-kyriazakos/relay-ws-hardening
fix(relay): rc.4 relay transport + inbound fixes — dedupe replays, fail pending on drop, fail fast mid-redial, WAN keepalive
2026-08-20 16:06:23 +10:00
Teknium 797bc4bf9b Merge remote-tracking branch 'origin/main' into feat/keyless-tavily-firecrawl-failover 2026-08-19 23:04:36 -07:00
Slobaka 6ff341c4d6 fix(doctor): web readiness reflects the selected provider's real state (#78412)
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).

Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
  mode is a working state, not a misconfiguration (zero-config installs
  and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
  explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
  doctor processes saw an empty registry and warned on everything)

E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
2026-08-19 23:03:58 -07:00
Brooklyn Nicholson 4dcefed089 fix(tui_gateway): scan remote git roots and scope projects.* to the focused profile
A remote desktop cannot crawl the host disk, and projects.* always read the
launch profile's stores, so switching profiles left the wrong tree on screen.
Bind the requested profile's HERMES_HOME and session db for the whole family,
and add scan:true so repos with no Hermes sessions still appear.

Co-authored-by: Chen Jin <Enough1122@users.noreply.github.com>
Co-authored-by: Ryan Weddle <weddle@gmail.com>
Co-authored-by: izumi0uu <izumi0uu@gmail.com>
Co-authored-by: webtoolbox <1911826+webtoolbox@users.noreply.github.com>
2026-08-20 01:02:17 -05:00
Teknium 481bc9391e fix(memory): profile-only config gets narrow USER_PROFILE_GUIDANCE instead of the full memory block
With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
2026-08-19 22:59:17 -07:00
HexLab98 a969c5a93d test(memory): cover the disabled built-in memory surface
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).

Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.

The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
2026-08-19 22:59:17 -07:00
kshitij 37fa4a7c63 chore: remove unused import pytest from test file
Follow-up cleanup from simplify-code review on PR #90521 salvage.
2026-08-20 11:28:26 +05:30
liuhao1024 fbca706789 fix(telegram): log the first confirmed getUpdates progress per generation
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).

Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.

Fixes #90504
2026-08-20 11:28:26 +05:30
Teknium 2eb5217af9 feat: keyless free-tier failover + Tavily/Firecrawl salvage integration
- Cross-vendor failover: when Exa's or Parallel's keyless free tier
  returns a rate-limit-shaped error, the request retries once on the
  other vendor's free endpoint (search + whole-batch extract). Result
  notes served_by; a peer pinned to its paid tier is never used;
  non-throttle errors never fail over.
- Docs: failover note + Tavily/Firecrawl keyless-when-selected rows.
- Firecrawl keyless test expectations aligned with the keyless tier.
2026-08-19 22:58:05 -07:00
Gille 188d47919a feat(computer-use): expose screenshots for chat delivery 2026-08-19 22:58:02 -07:00