Commit Graph

366 Commits

Author SHA1 Message Date
teknium1 08a2e7dbcc refactor(cli): vim mode is a config key only — drop the /vim slash command
Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
2026-09-12 22:00:02 -07:00
Teknium 7b037f0efa feat(cli): opt-in git_branch status-bar field (⎇ current branch)
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.

Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.

Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
2026-09-12 21:47:53 -07:00
teknium1 de2d6a1b93 fix(config): a fresh process recovers the last good config.yaml instead of running on defaults
CI / Detect affected areas (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / Detect affected areas (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
Nix flake check / Detect affected areas (push) Has been cancelled
Build Skills Index / build-index (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / OS-specific tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Installer tests (push) Has been cancelled
CI / Rust tests (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Docs Site (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / Check uv.lock (push) Has been cancelled
CI / Check no committed infographics (push) Has been cancelled
CI / Profile artifact check (push) Has been cancelled
CI / Check no case-colliding filenames (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest-32-core) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-latest-32-arm-core) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest-32-core) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-latest-32-arm-core) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
Nix flake check / nix flake check (push) Has been cancelled
Build Skills Index / trigger-deploy (push) Has been cancelled
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).

Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.

Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
2026-09-12 16:17:04 -07:00
Teknium bf51fee548 docs(multiplex): make the multiplexed-gateway page match what the code does
The "one gateway for all profiles" section had drifted from the runtime. Each
claim was re-verified at its defining symbol on current main and rewritten to
the behaviour users will actually see:

- named-profile guard: only `gateway run` refuses (exit 78 / EX_CONFIG,
  systemd RestartPreventExitStatus; launchd KeepAlive still retries);
  `start`/`install` do not refuse themselves and the message lands in the
  service log; `--force` is a `run`-only flag
  (hermes_cli/gateway.py::_guard_named_profile_under_multiplexer, _cmd_start)
- same (platform, token) in two profiles: the duplicate adapter is parked as
  fatal/duplicate_credential and the gateway keeps running — it was described
  as a fail-fast startup error (gateway/run_adapters.py::_refuse_duplicate_claim)
- status surfaces: one gateway_state.json under the default home with
  `<profile>:<platform>` entries + served_profiles; nothing is written under a
  secondary home (the page claimed a per-profile runtime_status.json), and
  `hermes status` does not list served profiles — `gateway list`,
  `-p X gateway status` and /api/status do (gateway/status.py::write_runtime_status,
  hermes_cli/status.py::_render_gateway, hermes_cli/gateway.py::_cmd_status)
- API_SERVER_KEY in a secondary .env auto-enables api_server and trips the
  port-binding skip; document the `enabled: false` pin
  (gateway/config_env.py::_api_server / _enable_from_env)
- allowlist: it is a start-time snapshot, and the Desktop backend's cron
  ticker enumerates every local profile regardless of it
  (hermes_cli/web_server.py::_start_desktop_cron_ticker)
- routed-profile cron via the shared bot: only when the profile has no live
  adapter of its own, and a route carrying guild_id never matches a cron
  target because delivery matches on chat_id/thread_id only
  (cron/scheduler_provider.py::tick_adapters_for,
  cron/scheduler_preflight.py::SharedRouteAdapters.get)
- add a "What is isolated per profile" table (credentials, authorization,
  endpoints, media denylist, MCP child env, outbound egress, session
  namespace, logs, terminal) describing behaviour, not PR numbers
- configuration.md: the ${VAR} scoping paragraph now says where it applies
  and links to the table
2026-09-11 15:51:11 -07:00
Teknium 338bf9ea9a fix(cli): banner/TUI/dashboard update checks go through the GitHub API, cached 24h
The Python passive check (`banner.check_for_updates`, used by the CLI
banner, `hermes --tui`, every `tui_gateway` spawn and the dashboard's
/api/hermes/update/check) also ran `git fetch origin main` on every cache
miss, and never cached an inconclusive result so a flaky line retried on
every start. Same GitHub complaint, same fix:

- remote tip via GET /repos/{slug}/commits/main (vnd.github.sha), local tip
  via rev-parse, exact count + changelog via the compare API when they
  differ. HTTPS `ls-remote` remains only as the fallback when the API is
  unreachable or the origin isn't on GitHub.
- cache TTL 6h -> 24h, failures cached 1h; the cache is keyed on HEAD so
  `hermes update` invalidates it immediately.
- the dashboard's "what's changed" list comes from the memoized compare
  payload (`upstream_commits_behind`) instead of `git log HEAD..origin/main`,
  which was stale without a fetch.

Tests rewritten to the new contract: passive checks must not run
`git fetch`/`ls-remote` for a GitHub origin; the daily cache invalidates
when HEAD moves and re-asks after the failure window.
2026-09-10 18:15:54 -07:00
Teknium bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium 06dc51d62d fix: verification evidence ledger is inert while verify_on_stop is off
The ledger in verification_evidence.db exists only to feed the verify-on-stop
guard, but the recorder kept running on every foreground terminal command and
every file edit after #53552 turned the guard off by default. Users who never
opted in still accumulated a multi-MB database (7 MB / 4.6k rows on one install).

Every ledger entry point (record_terminal_result, record_verify_run,
mark_workspace_edited, verification_status) now checks verify_on_stop_enabled()
first and returns without opening or creating the database when the guard is
off. verification_status reports {"status": "disabled"} in that case; no client
consumes the verification.status RPC yet, so nothing downstream changes.

Existing ledger tests pin HERMES_VERIFY_ON_STOP=1 since they exercise the ledger
itself; the new test proves the off path never creates the file (red on base).
2026-09-09 02:35:41 -07:00
Ayush Nangia 3204bfa5e2 docs+defaults: declare delegation.fallback_providers in config surfaces
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.

Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
2026-09-08 02:26:05 +05:30
Teknium 03f3b09222 fix(tui-gateway): subagent lifecycle survives display.tool_progress=off
`_on_tool_progress` bailed on the tool-progress gate before dispatching
`subagent.*`, so a Desktop/TUI user who hid tool-call chrome also lost the
subagent rows in the status stack and spawn tree. Subagent lifecycle is
application state (like `todo.updated`, clarify and MCP consent cards,
which already bypass the gate); the gate now applies only to the optional
progress chrome (reasoning previews, MoA rows, tool.generating).
2026-09-07 11:25:20 -07:00
Teknium 93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium 3a42722c84 test: verify SSH update checks with real PTY authentication controls 2026-09-07 08:21:24 -07:00
Teknium cba612e128 docs: explain same-model background review reasoning inheritance 2026-09-07 05:59:24 -07:00
Teknium 8d24bc24e1 fix: wait sixty seconds before provider silence notices 2026-09-07 05:24:15 -07:00
Teknium 18955c35b5 fix(sessions): repair closed lanes through canonical disconnect cleanup
The real closed-WebSocket resident variant still wedges after the missing-timer repair. Re-enter existing transport cleanup before rearming orphan timers, preserving viewer transfer and the reconnect/delegation fence rather than deleting registry rows from a stale snapshot. Extend the same invariant to both dead-transport shapes. Live class investigation informed by #104710; no direct-vouch reclaim machinery imported.
2026-09-07 04:55:32 -07:00
fangliquanflq 919584e971 fix(sessions): rearm lost detached chat teardown timers
Salvage #104704 (e9423d2d0bbe3e795c5eaccb86a913f1d95444ba, 5fa2b98fc833fc4e1ee7f1aaa7eb45cdd3cd7cde). Reuse guarded orphan teardown instead of deleting ownership fences. Real two-backend WebSocket probe reproduces the missing-timer wedge on base and proves reconnect/delegation protection and recovery. Add reusable probe and user documentation.
2026-09-07 04:55:32 -07:00
Teknium 22c5684b98 fix(agent): clear resumed stream wait status without synthetic reasoning 2026-09-07 02:37:13 -07:00
Teknium 77f79ae831 fix(agent): keep active Codex reasoning out of wait warnings
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.

Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.

Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
2026-09-07 02:37:13 -07:00
Teknium dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
kshitijk4poor 9c4c548cd5 docs(terminal): document docker_snap_compat 2026-09-05 21:00:19 +05:30
kshitijk4poor f77e2a57ed docs(dashboard): document ssh_isolated_idle_grace_s 2026-09-05 20:44:51 +05:30
kshitijk4poor 33d6dab04c docs(ssh): document the AcceptEnv prerequisite for env passthrough over SSH 2026-09-05 15:54:39 +05:30
Teknium 40da0fd52f feat(delegation): subagents compress at an absolute context cap (delegation.compression_threshold_tokens, default 200K)
A delegate_task child inherits the compression threshold as a RATIO of the
model window. On a 1M-window model at the run's configured 0.85 that is an
850K-token trigger: in the 1,393-agent refactor run 1,373 of 1,375 children
never compressed once, 62% of all API calls carried >150K of context, and the
calls above 200K carried ~$10.9k of the $19.3k bill (58% of it cache WRITES,
i.e. re-sending a 300-800K prefix on every call). A sawtooth replay of the
logged calls with a 200K cap / 65K floor cuts context spend by ~49% (~$7k).

Children are brief-driven and disposable; they re-read their brief and the
files they touch, so a large window buys them little. New
delegation.compression_threshold_tokens (default 200000) is applied to the
child's ContextCompressor right after construction as the lower of it and
any global compression.threshold_tokens; the parent's own trigger is
untouched. 0 disables the subagent-specific cap. The compressor applies
threshold_tokens_cap on first window resolution, so this is byte-equivalent
to the user having set compression.threshold_tokens for the child.

Live through the real spawn path (_build_child_agent, real imports, temp
HERMES_HOME, 1M-window model): main child trigger 500,000 / branch 200,000;
parent 500,000 on both.

Tests (3): default caps a 1M child at 200K; the cap is the lower of the
delegation and global values and never raises a small-window child's
trigger; 0 disables and an already-resolved trigger is re-clamped.
Docs: delegation.md, configuration.md.
2026-09-05 01:07:09 -07:00
Teknium f1ccf436a2 feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:

- search: POST https://api.perplexity.ai/search (documented Search API),
  search_context_size=low so `snippet` stays description-sized;
  results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
  route behind `pplx content snippets` (the CLI's `content fetch` is
  deprecated upstream). web_extract has no query, so the URLs' path words
  serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
  backend set + credential ladder + availability probe (web_tools),
  registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
  dump key lists, nous_subscription direct-credential detection, setup
  summary, test conftests, docs.

Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
2026-09-04 07:17:00 -07:00
Teknium d0b7cec0b8 fix(prompt): Muse Spark gets tool-use enforcement + execution guidance on defaults (#96550)
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.

Co-authored-by: Edder Talmor <talmoredder@gmail.com>
2026-09-03 00:58:32 -07:00
kshitijk4poor 52e070f4cb docs: document context_file_read_timeout in configuration.md 2026-09-03 03:15:31 +05:30
Teknium 0fd9218e5a docs: ${VAR} config refs resolve per-profile under a multiplexed gateway 2026-09-02 06:19:03 -07:00
Teknium 70dc1606c6 test(cli): pin OSC 9 / Warp OSC 777 bell emitters; docs + contributor mappings
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
  only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
  of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
2026-09-02 06:17:10 -07:00
Teknium 552159d222 feat(cli,tui): collapse bell_on_clarify/approval into display.bell_on_prompt
One key covers every blocking prompt modal: clarify (single + batch),
dangerous-command approval (incl. computer_use), sudo password, and
secret capture. CLI gets a _ring_bell() helper shared with
bell_on_complete; TUI rings on clarify/approval/sudo/secret .request
events (isTTY-gated). 'hermes config' Bell summary shows both flags.
2026-09-02 05:34:35 -07:00
Turgut Kural 3082a34669 feat(cli,tui): add display.bell_on_approval + fix eslint error
- display.bell_on_approval (default false): same BEL mechanism as
  bell_on_complete, rings when a dangerous-command approval prompt
  opens (_approval_callback / approval.request event). Complements
  bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
  (if without braces) that failed the CI JS & TS checks job.
2026-09-02 05:34:35 -07:00
Turgut Kural ef6d3367a6 feat(cli,tui): add display.bell_on_clarify — terminal bell on clarify prompts
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
2026-09-02 05:34:35 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 25d954c2cf fix(guardrails): hard stops catch replays, never legitimate iteration
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:

- Edit -> re-run is progress. A successful mutating call (write_file/patch,
  a green terminal/execute_code, browser actions, job/message/cron/memory/
  skill mutations) marks progress for every failing signature still being
  counted this turn; the next identical retry restarts its streak instead
  of accumulating toward exact_failure_block_after. A pure replay never
  mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
  (terminal, execute_code, process pollers, browser_navigate, web_extract)
  same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
  task loops with a live parent/client and do real edit -> re-run work.

Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
  unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
  this commit:        COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
2026-09-02 00:26:57 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
João Vitor Cunha 384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
Teknium cb446a5bed fix(config): canonicalize platforms.<name>.<display_setting> at one chokepoint; mirror on get/unset (#71047 Problem A)
Follow-up to @zoser69's #78111 cherry-pick:
- lift the redirect into _redirect_platform_display_key() and apply it
  BEFORE _validate_config_key / type coercion, so the unknown-key hint and
  the string-vs-bool coercion both see the canonical path
- widen to the sibling surfaces: config get resolves the canonical key
  (previously echoed the dead top-level value — the misleading half of the
  report) and config unset removes the canonical leaf
- regression tests: get mirrors gateway resolve_display_setting, unset
  removes the redirected leaf, note printed, helper touches ONLY
  OVERRIDEABLE_KEYS (connection keys / 4-segment / already-canonical
  paths untouched)
- docs: configuration.md per-platform section names the canonical CLI
  path and the accepted shorthand
2026-09-02 00:07:15 -07:00
Teknium 9de9d7613c fix(compression): keep hygiene turn-hold worker's commit admission so thinking-model summaries are adopted, not burned
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.

Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):

- CompressionCommitFence gains mark_commit_watermark_fenced() /
  commit_watermark_fenced; compress_context marks the fence right after
  capturing get_active_message_watermark() under the durable compression
  lock (#75316/#87484) — the property that makes a LATE commit safe: rows
  appended after compression start survive both commit paths verbatim as
  cloned concurrent tail (archive_and_compact watermark= and
  publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
  the detached worker (already kept alive via
  _defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
  user's turn proceeds on the uncompressed transcript at the same 10s
  budget, and the summary is adopted at the worker's own watermark-fenced
  commit boundary. Unfenced workers are cancelled exactly as before —
  never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
  block preflight adoption via the same-session cooldown); re-attempt
  spacing is covered by the durable compression lock
  (_session_has_compression_in_flight). If the worker ends WITHOUT
  committing, a done-callback restores the flat non-escalating 60s
  retry-after; a successful adoption resets the hygiene failure streak.
  The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
  to describe deferred adoption and the thinking-model case;
  config_defaults.py comment updated. Knob stays config.yaml-only.

Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
  test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
  UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
  path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
  start watermark; the fence still gates admission and unfenced/late
  results are discarded.

New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
  turn still released at the budget, no cooldown while running,
  streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
  turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.

Fixes #97963
2026-09-01 23:56:23 -07:00
Teknium 92fa0845ee docs(compression): note the summary stream now closes at context_total_ceiling_seconds on every aux wire 2026-09-01 23:56:06 -07:00
Teknium e71352aef3 docs: document model.streaming escape hatch (#80789 salvage follow-up) 2026-09-01 22:14:06 -07:00
Teknium c5b99a3ee5 fix(agent): delegated children and cron turns stream again — inline, no worker
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).

Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.

should_use_direct_api_call() itself is untouched; only what it routes to.

Live A/B (real SSE server, real AIAgent.run_conversation):
  before: subagent/cron wire stream=None, request on conversation thread
  after:  subagent/cron wire stream=True, request on conversation thread
          cli unchanged (stream=True, spawned worker)
  inline stale detector kills a one-chunk-then-silence stream at budget;
  AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.

Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
2026-09-01 21:42:19 -07:00
Lakshya Agarwal 89ca5e614b fix(tavily): update Tavily provider documentation 2026-09-01 10:56:49 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium 67de93862c fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).

The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).

Fixes #98028
Fixes #100325
2026-09-01 10:52:23 -07:00
kshitijk4poor c394b005fc fix: publish watchdog settlement only after the abort commits
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.

- Split the surface: `_surface_stall` is now observational only ("no
  progress for Ns; attempting recovery"), and the definitive
  aborted/lease-stopped settlement moves to a new
  `_surface_committed_abort` that runs only after `_commit_abort`
  succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
  turn whose aborts keep declining no longer re-logs an ERROR and
  re-warns the user every poll interval.
- Add the committed-path regression test
  (`test_watchdog_publishes_definitive_settlement_only_after_commit`)
  and extend the declined-path witness
  (`...resumes_during_warning`) to assert no committed-abort or
  definitive pre-commit claim appears when the abort is vetoed. Both
  fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
  unconditionally, no generation claim — losing the lease means the
  process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
  WHY, drop the round numbering), and drop the dead
  `cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.

On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
2026-09-01 03:19:59 +05:30
LucidPaths 0943702c55 fix(gateway): keep long turns controllable without blocking Telegram 2026-08-31 12:21:01 -07:00
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Teknium 4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
Teknium bacb90fe20 feat(delegation): honor delegation.request_overrides on all three resolution branches with explicit-over-runtime merge precedence
Completes the #90953 salvage on post-#98237 main:

- New _merge_request_overrides helper defines the precedence contract:
  explicit delegation.request_overrides merges OVER runtime/parent-derived
  overrides — explicit top-level keys win; extra_body is deep-merged one
  level so runtime extra_body keys survive unless redefined. Inputs are
  copy.deepcopy'd so transport-side mutation can't leak into config or the
  provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
  provider-alongside-base_url runtime overrides instead of being a separate
  return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
  delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
  (previously only when override_provider was set), enabling the inherit
  branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
  proofs, explicit-over-runtime precedence on the provider-alongside-base_url
  path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
  the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
2026-08-29 19:13:23 -07:00
Teknium 86a2fdc634 feat(tui): status rule shows cache-hit %, latency, t/s and honors display.status_bar.fields
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
  avg_latency_s, avg_tps (reads the same per-call deque history from
  agent/conversation_loop.py; keys omitted when no data — Codex
  app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
  segments (breakpoints 96/104/110 cols, lowest priority — they shed
  first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
  filters TUI segments too: cache_hit, latency, tps, duration,
  compressions, bg_tasks, bg_subagents, voice, battery, title,
  context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
  events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
2026-08-29 19:12:24 -07:00
Teknium 9e017428ba fix: unify status-bar field keys, docs, and tests for salvaged cluster
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
  second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
  title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
  and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
  baseline resets, title badge gating
2026-08-29 18:34:51 -07:00
liuhao1024 fb786d2f5b feat(cli): add display.status_bar.fields config for customizing status bar
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.

Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.

When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.

total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.

Closes #41909
2026-08-29 18:34:51 -07:00