Commit Graph

2878 Commits

Author SHA1 Message Date
Teknium 27f32bd50b test: exercise output-cap removal across native and child surfaces 2026-09-07 06:15:43 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 65f033a1a2 fix(execute-code): teach the working helper import contract
Slim adaptation of #83772 to the current schema and failure-hint table.
Generated helpers are module exports on every execution path, not globals.
Correct schema, recovery hints and CLI tip rather than injecting names or
changing the execution boundary. Two registry-driven invariants reproduce
both misleading instructions on main and execute the corrected guidance.

Additional tool fix discovered during campaign #104904.
Original diagnosis and correction: @yuzilongleif-collab (#83772).

Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-09-07 06:12:24 -07:00
Teknium 57c60f2e0c fix: explain the launchctl registration restriction without inventing KeepAlive 2026-09-07 06:05:51 -07:00
Teknium 746b14b900 test: use explicit UTF-8 in file sync fixtures 2026-09-07 06:02:41 -07:00
liuzikaii b4e0f4a7bb fix(file-sync): hash the uploaded snapshot instead of mutable host files 2026-09-07 06:02:41 -07:00
Teknium f8c9e93dad fix: round-trip checkpoint path bytes without text translation 2026-09-07 06:00:46 -07:00
liuzikaii d77df6674a fix(checkpoints): preserve literal paths in Git filename output 2026-09-07 06:00:46 -07:00
Teknium a3ad585fd9 fix: limit skill update change to unusable local installs 2026-09-07 05:59:43 -07:00
Teknium 6798a9b8a4 fix: bound skill update wait budget and lingering fetch workers 2026-09-07 05:59:43 -07:00
Teknium 2079e4f08d fix: skip lock entries replaced by non-directory files 2026-09-07 05:59:43 -07:00
Teknium 36b0b6c9f2 fix: enforce complete fetch deadlines and inherit request context 2026-09-07 05:59:43 -07:00
liuhao1024 47887693c6 fix(skills): skip orphaned hub entries and bound per-fetch time in update checks
check_for_skill_updates() fetched every lock-file entry remotely, even
when the entry's install directory no longer existed, and each fetch had
no wall-clock bound — a few dead sources turned a routine
`hermes skills update` into a multi-minute stall (#104291).

- Entries whose recorded install_path resolves but does not exist are
  reported as "orphaned" and skipped without a remote fetch;
  unresolvable paths keep the previous fetch behavior.
- Each fetch now runs under a daemon helper thread with a hard timeout
  (default 30 s) and degrades to "unavailable" when abandoned.
- `hermes skills check` prints a removal hint for orphaned entries.

Fixes #104291
2026-09-07 05:59:43 -07:00
Teknium 76af5ebf09 test: release notification fixtures without writable stdin 2026-09-07 05:57:26 -07:00
Teknium 231828cdac test: synchronize child exit with notification admission 2026-09-07 05:57:26 -07:00
Teknium cfe07df09f test(agent): assert owner-scoped teardown instead of bulk cleanup 2026-09-07 04:38:59 -07:00
Teknium 0d8a1575c5 refactor(mcp): keep the passive status RPC, drop the SDK contract and reason codes
mcp.servers.status now rides the shared _mcp_rpc decorator (profile scope, 4064,
5024 with the real message) instead of a hand-rolled try/finally with a blanket
except. Drop the _MCPConnectErrorText str subclass and reason taxonomy: the
existing status/error fields already carry the state, and a whitelist on the RPC
keeps error text out of the wire. The Desktop connections.health contribution
contract is held back until its consumer plugin is public. Tests trimmed to the
scope invariants (per-profile runtime visibility, scoped shutdown clears only its
own status, launch runtime never leaks into another profile).
2026-09-06 13:18:19 -07:00
Joey a6699d60f4 feat(mcp): expose profile-scoped cached connection health 2026-09-06 13:18:19 -07:00
Teknium b167e81750 fix: legacy Bot Mode section in SOUL.md no longer taxes every session or shadows the live roster
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.

- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
  the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
  prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
2026-09-06 13:08:31 -07:00
Teknium d0090a147c Merge remote-tracking branch 'origin/main' into fix/async-batch-task-failure-notice
# Conflicts:
#	tools/delegate_tool_dispatch.py
2026-09-06 12:07:50 -07:00
Teknium 6d9c166455 Merge pull request #103529 from NousResearch/fix/terminal-auto-background
fix(terminal): an over-cap foreground timeout runs as a tracked background process instead of being refused (454 refusals in one run)
2026-09-06 12:06:04 -07:00
Teknium fdd140557b Merge pull request #103551 from NousResearch/fix/tool-output-caps
feat(file_tools): write_file flags a whole-file rewrite that mostly re-sends what is on disk (661 such writes, ≈$155, in one run)
2026-09-06 12:05:53 -07:00
Teknium cd658b6a46 Merge pull request #103492 from NousResearch/fix/approval-scanner-quoted-subshell
fix(approval): a grep inside "$(...)" no longer trips the hardline malformed block, and a quoted substitution body keeps its command boundaries (546 false blocks; review-found bypass closed)
2026-09-06 12:03:29 -07:00
Teknium e5f8420be8 Merge pull request #103513 from NousResearch/fix/subagent-context-cap
fix(delegation): compression_threshold_tokens is opt-in (default off, children keep the 500K ratio trigger); validate the value
2026-09-06 12:02:49 -07:00
Teknium 3d5831fa59 Merge pull request #103486 from NousResearch/fix/nested-delegate-deadline-and-summary-budget
fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum
2026-09-06 12:02:43 -07:00
kshitijk4poor 9158bd8e0d refactor(file-ops): one _run_rg_bounded owns the native-vs-shell transport choice
Three call sites each repeated `if native: _run_rg_native(...) else: _exec(... | head -n N)`.
The choice now lives in _run_rg_bounded; callers pass the words, the bound, and the
one thing the native lane cannot express (a cd prefix → native_ok=False). The grep/find
pipeline keeps its explicit shell form because of the column cap.

Test file: one module-scoped LocalEnvironment instead of thirteen (~0.8 s each).
2026-09-07 00:28:46 +05:30
kshitijk4poor ef5a534c9a fix(file-ops): native rg runner honours deadline and /stop while rg is silent
The first version checked the deadline only after a line arrived, so an rg
that produced nothing for 60 s (huge tree, no hits yet) pinned the caller
past the timeout and ignored the interrupt flag that the shell path honours
via _wait_for_process. Drain on a daemon thread; the waiter owns deadline
(124) and interrupt (130) and kills the process group, so no rg or child
survives the return. Probe: silent 10 s process, timeout=2 → 2.0 s / 124;
interrupt at 0.5 s → 0.5 s / 130; zero stray processes afterwards.
2026-09-07 00:28:46 +05:30
kshitijk4poor 83063aeb87 perf(file-ops): run rg natively for search_files on local POSIX hosts
read_file already bypasses the backend shell on a local POSIX environment
(_read_file_native); search_files still paid two bash spawns per call — the
`test -e` existence probe and `set -o pipefail; rg ... | head -n N` — plus one
per zero-match probe. Measured on macOS against the repo's tools/ tree:
content search 85 ms → 15 ms, no-match search (three probes) 179 ms → 44 ms,
file-name search 73 ms → 12 ms; raw `rg` argv is ~15 ms, so the remainder was
transport.

Same gate and kill switch as reads (_native_read_enabled: LocalEnvironment,
not win32, HERMES_NATIVE_FILE_READ=0 disables). The argv builders and
_parse_search_output are unchanged and shared: _run_rg_native shlex-splits
the already-quoted words, streams stdout and stops after fetch_limit lines
like `head` would, and reports exit 0/1/2 (124 with partial output on
timeout) so the parser sees the shell contract. grep/find fallbacks, remote
backends, Windows and the multi-root cd form keep the shell path.

Two shell-observer tests in test_search_zero_match_and_multipath pin the
shell lane explicitly; they assert on command text, not behaviour.
2026-09-07 00:28:46 +05:30
kshitijk4poor 83467c28f7 test: drop the duplicate #90322 file — TestSearchHints already json.loads the truncated output 2026-09-07 00:20:11 +05:30
kshitijk4poor 8ed990e70f test: trim #90322 regression to one invariant, drop the hint-splitting workaround
The credentials test no longer has to strip trailing text before json.loads;
the new file keeps the single behaviour contract (truncated output round-trips
through json.loads and carries the next offset) — the existing TestSearchHints
cases already cover offset arithmetic.
2026-09-07 00:20:11 +05:30
liuhao1024 478e66475b fix(tools): keep truncated search_files output pure JSON
When results were truncated, search_files appended the pagination hint
as plain text after the serialized payload ("{...}\n\n[Hint: ...]"),
so the tool result was no longer parseable JSON — downstream tool-message
handling on providers strict about tool-content formatting could reject
or mishandle it, contributing to 400 upstream errors in sessions with
truncated search output (#90322).

Move the hint into the payload as a structured _hint field, matching the
existing _omitted/_warning side-channel convention in the same function.
The model-facing guidance (explicit next offset) is unchanged.

Fixes #90322
2026-09-07 00:20:11 +05:30
jango ed406f929d fix(browser): make the real-profile attach daemon reapable (#100855, salvage #103284)
The `hermes-real-profile` agent-browser daemon (the attach lane for consented
real-profile browsing) ran with plain `_build_browser_env()`: no
`AGENT_BROWSER_SOCKET_DIR`, so it lived in agent-browser's default dir and no
reap path could see it. A wedged daemon + headless Chrome survived 47h across
two gateway restarts (#100855), and on macOS the genuine Chrome binary it held
made "Chrome won't open" for the user.

Give the attach lane the same contract every other lane already has:
`_prepare_session_socket_dir()` (per-session socket dir + `owner_pid` claim)
and `_agent_browser_command_env()`, and add the named dir to the orphan
reaper's scan. The existing `_reap_socket_dir` then applies its owner-liveness
and start-time-fingerprint rules unchanged; the daemon is listed as tracked so
the untracked-idle escape hatch never fires under a live user (per-task `rp_*`
sessions drive it over `--cdp`, so its own dir shows no activity). The daemon-side idle
timeout is NOT inherited: Chrome is launched by Hermes, not the daemon, so a
self-exiting daemon would leave Chrome holding the copy dir while the next
attach re-runs the snapshot overlay over it.

When a reaped daemon's Chrome (Hermes-launched, own process group) still
holds the copy dir, `_real_profile_cdp` re-attaches to it instead of running
the snapshot overlay over a live profile. DevToolsActivePort outlives a
crashed Chrome and its port can be recycled, so the file's browser id must
match `/json/version` before it is trusted; an attach failure on a live
Chrome fails closed rather than overlaying.

Tests: attach/get/close commands carry the reaper-visible socket dir and
owner_pid and no idle timeout; a dead-owner real-profile daemon is reaped by
`_reap_orphaned_browser_sessions`; a surviving Chrome is re-attached, never
overlaid (all red on main).
2026-09-06 23:31:05 +05:30
jango b2b026dc22 fix(bot-mode): retain message_agent across tool rebuilds (#102864) 2026-09-06 23:11:12 +05:30
Teknium dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
Teknium c5594ec4b3 fix(delegation): a crash mid-unit keeps the children that already finished
Since #104299 a background delegate_task call is split into completion units
(one per `group`, one per ungrouped task). A multi-child unit still joined on
all its children before anything was written durably, so an owner crash
between the first and last child lost the finished work and replayed the whole
unit as "outcome unknown" — the restart-granularity gap that #104233 (Xipong's
#76228/#76229 direction) solved with a second row per child.

Each finished child of a detached unit is now recorded on the unit's OWN row
(`record_unit_child` → result_json {results, partial}) as its future lands;
the real result overwrites it at finalize. `recover_abandoned_delegations`
replays recorded children with their real summaries and marks only the
unfinished ones unknown, naming the count. No new rows, no new consumer shape.
`task_indexes` is persisted so recovery knows a split unit's members.
2026-09-06 09:21:58 -07:00
Teknium 03e9f4caee Merge pull request #104284 from NousResearch/fix/nous-anthropic-chat-wire
fix(nous): route anthropic/* over chat/completions by default; the native Portal route re-writes the cache on 14-20% of calls (nous.anthropic_wire)
2026-09-06 07:17:48 -07:00
Teknium b1403d521c Merge pull request #104168 from NousResearch/fix/subagent-cache-ttl-5m
fix(delegate): subagents never inherit the 1h prompt-cache tier (2x write price for retention they never use)
2026-09-06 07:17:43 -07:00
Teknium 028fe2c4c8 feat(delegation): per-task completion groups — ungrouped subagents return as they finish
A background delegate_task call used to be ONE async unit: the runner joined on
every child and a single consolidated message re-entered the conversation when
the SLOWEST finished. Fifteen independent PR reviews therefore waited on the
fifteenth before the parent could act on the first.

Each task now carries an optional `group`. `_units_of` partitions the call's
children into units — one per distinct group, one per ungrouped task — and each
unit is dispatched to the async registry on its own, so its results re-enter the
conversation as soon as THAT unit is done. Tasks that must be compared or merged
share a group and still return together.

Capacity is unchanged: every unit of one call joins the first unit's pool slot
(`slot_key` in `async_delegation._dispatch`), so splitting never consumes more of
`delegation.max_concurrent_children` than the call did. Unit ids suffix the
call's id (`deleg_xxxx-1`, `-2`, …) so live transcripts stay under one dir; the
completion block names the group and notes that sibling units report separately;
`active_task_count` counts a unit's own tasks.
2026-09-06 07:17:30 -07:00
Xipong 25a80b3bb6 fix(skills): exclude generated runtime caches from bundled ownership hashes 2026-09-06 07:17:03 -07:00
Teknium 335ecf9f4a test: the two remaining native-wire contracts select nous.anthropic_wire=native explicitly
test_nous_anthropic_fallback_uses_the_messages_wire and
test_nous_child_rederives_api_mode_from_model describe the native wire, which
is now opt-in; select it in the test the same way the wire-contract suite
does, so both keep guarding the flip-back.
2026-09-06 05:55:18 -07:00
Teknium 3a15c39e8e fix(banner): deferred update notice keeps prompt_toolkit routing; drop dead console arg; PTY harness
Follow-up on the cherry-picked fix: _defer_update_notice no longer takes the Rich console it
can never safely use (the notice always lands after patch_stdout owns stdout), the unit test
asserts the contract at prompt_toolkit's boundary (an ANSI fragment reaches
print_formatted_text, with no raw ESC/markup in the visible text) instead of patching our own
cprint, and evals/cli_deferred_notice.py drives the real CLI under a Linux PTY with a FIFO-gated
update cache to A/B the late-notice cases (garbled on base, clean after).
2026-09-06 05:36:26 -07:00
turingcat a366a0bb17 fix: route deferred update notice through cprint to avoid ANSI garbling
The deferred update notice ("N commits behind") was rendered via
console.print which writes ANSI escapes directly to stdout. Under
patch_stdout (active in the interactive CLI), the StdoutProxy strips
ESC bytes, leaving visible [1;33m...[0m artifacts.

Fix: render Rich markup to an ANSI string via a captured Console, then
route through cprint (which uses prompt_toolkit ANSI parser) to
bypass the StdoutProxy mangling.

This is the same class of bug as #2262 (fixed by #2448 for agent._print_fn
and display.py), but the _defer_update_notice callsite in banner.py was
not covered by that fix.

Fixes #83969
2026-09-06 05:36:26 -07:00
Teknium 3009efe75e fix(tools): read_file dispatch honors the advertised 2000-line default when the model omits limit
#76996 raised DEFAULT_READ_LIMIT, the read_file_tool signature and the schema
default from 500 to 2000, but the registry dispatch handler `_handle_read_file`
kept its own literal `args.get("limit", 500)`. Since models omit `limit` on most
reads, every dispatched read still stopped at 500 lines with `truncated: true`,
so the default flip never reached production traffic.

All three sites now read the single DEFAULT_READ_LIMIT constant, so the schema,
the Python default and the dispatch fallback cannot drift again.

Live probe (registry.dispatch on a 1500-line file, no limit): before 501 lines /
truncated=true; after 1501 lines / truncated=false.
2026-09-06 05:32:49 -07:00
Teknium 0cb996d977 fix(delegation): number delegation batches per conversation, not per process
`format_batch_tag` handed out `set N` ordinals from one process-wide table,
so every conversation on a shared backend and every child's nested fan-out
advanced the same counter. A user's second wave of 15 lanes rendered as
`[set 20 · 13/15]`, which reads like 20 batches were spawned.

Scope the ordinal table by the parent conversation (`parent_agent.session_id`)
and thread the parent through the three render sites (batch header,
completion lines, child tree-line prefix via the shared session_ref). The
first fan-out in a conversation is `set 1`, the next `set 2`; sibling
conversations and nested child batches no longer inflate it.

Invariant test proven red on origin/main, green with the fix.
2026-09-06 05:03:20 -07:00
kshitijk4poor 3513a3b922 test(process_registry): pin that the availability probe argv omits OOMPolicy too
Requested in the review of #102357: the spawn builder was pinned, the probe
was not, and the probe is where the rejection was cached as "unavailable".
2026-09-06 14:57:33 +05:30
kshitijk4poor 04885c25cf test(process_registry): pin that scope argv never emits OOMPolicy
The spawn_local systemd test asserted the property was present; it now
asserts the invariant the fix establishes (no OOMPolicy= on a transient
scope), so a reintroduction fails here instead of on older systemd hosts.
2026-09-06 14:57:33 +05:30
kshitijk4poor 10e8756cf4 test(memory): drop assertions implied by the disk-unchanged check 2026-09-06 14:53:13 +05:30
Halldrix 7602b33d79 fix(memory): point empty-batch refusal at single remove() calls (#103419)
Review feedback: the sanctioned deliberate wipe is repeated single-op
remove() calls, not a manual file edit — direct the model there.
2026-09-06 14:53:13 +05:30
Halldrix faa46ed2c7 test(memory): pin batch refusal to empty a non-empty store (#103419) 2026-09-06 14:53:13 +05:30
Teknium 0edb6b928a fix(delegate): subagents never inherit the 1h prompt-cache tier
A delegated child copies the parent's prompt_caching.cache_ttl. The 1h tier
is priced for a person who steps away between turns (2x write vs 1.25x for
5m); a subagent calls every few seconds for minutes and is gone, so it paid
2x on every tool result and never collected the retention. Live measurement
(Sep 5, 40 concurrent Fable 5.1 children via OpenRouter, 1h markers on the
wire): cache writes billed at $20/M against $12.51/M for the same run at
5m, ~60% of a write-dominated bill.

_apply_child_cache_ttl runs right after child construction: 1h -> 5m,
5m stays, disabled stays disabled, parent untouched. Tests: unit (markers
on the wire drop ttl; parent's 1h layout still differs) and the real spawn
path with cache_ttl: 1h in a temp HERMES_HOME.
2026-09-06 02:20:49 -07:00