Commit Graph

219 Commits

Author SHA1 Message Date
Ayush Nangia 3204bfa5e2 docs+defaults: declare delegation.fallback_providers in config surfaces
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.

Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
2026-09-08 02:26:05 +05:30
Teknium 93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium 1491342661 style: remove blank line left by snapshot setting removal 2026-09-07 08:08:41 -07:00
Teknium 7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium c89f3b8800 fix(delegation): one completion per call by default; queued units no longer stalled; tell the model results land between turns
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:

1. delegation.independent_completions (new, default false). #104299 made every
   ungrouped task its own completion message, so a 15-task call woke the
   orchestrator up to 15 times; one chain received 132 notices and answered
   130 of them with "already incorporated". A multi-task call now returns as
   ONE consolidated message unless the flag is on; `group` is inert until then.

2. Queued units were killed before they started. Units of one call share a
   pool slot but the executor was still sized by slots, so with 15 units live
   a new unit queued behind a full pool; the stale monitor's clock ran from
   dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
   when its thread finally came up (13 such lanes in one session). The
   executor now grows to the number of live units and the stall clock arms
   when the runner actually starts.

3. The tool text said "do not wait or poll — just continue" without saying
   that completions are delivered only BETWEEN turns. A model that never ends
   its turn (one 203-minute turn, 717 API calls) never received 40 finished
   results. Tool description, dispatch note and completion header now say to
   finish independent work, give a one-line status, and end the turn.
2026-09-07 06:46:54 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 4dcdb5e896 fix(cli): allow pinned installs to disable passive update checks
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.

Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
2026-09-07 06:13:16 -07:00
HoneyTyagii 381d6064d7 fix(desktop): preserve custom Linux launcher entries when opted out
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
2026-09-07 04:55:17 -07:00
Teknium 08b140d14e fix: gpt-6 Astra on Codex OAuth gets the 85% compaction autoraise
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
2026-09-06 23:08:45 -07:00
Teknium b167e81750 fix: legacy Bot Mode section in SOUL.md no longer taxes every session or shadows the live roster
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.

- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
  the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
  prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
2026-09-06 13:08:31 -07:00
Teknium dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
Teknium 45646f4d09 feat(nous): anthropic_wire=auto decides a session's wire from its first response
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.

`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.

GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.

Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.

Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
2026-09-06 09:11:04 -07:00
Teknium 54e24ae1fa fix(nous): route anthropic/* over chat/completions by default (nous.anthropic_wire)
Portal serves anthropic/* on two routes. The native /v1/messages wire, which
Hermes has used since 02d5e23085, re-writes the previous turn's prompt cache
on 14-20% of consecutive calls in concurrent tool loops; the chat/completions
route does not. Measured 2026-09-06, 20 concurrent sessions x 6 tool calls on
Fable 5.1, same account, same hour, same first-party pin:

  Nous /v1/messages          180 pairs, 25 stuck (13.9%); 3 earlier 40-session
                             runs 15.0 / 17.3 / 20.2%; unchanged by the pin
  Nous /v1/chat/completions  320 pairs (2 runs), 0 stuck
  OpenRouter direct, pinned  161 pairs, 0 stuck

"stuck" = cache_read on call k+1 equals cache_read on call k instead of
call k's prompt size: the last write was not visible, the turn (~28K) was
written twice. Every stuck pair had byte-identical system, tools and
per-message shas, so it is not client-side mutation. At $10-20/M for
cache writes that is 15-20% of a fan-out's write bill.

The cause is inside the portal's native route (every element of its
outgoing request reproduces clean from outside; NousResearch/api#227
carries the diagnostics). Until it is fixed, anthropic/* rides
chat/completions. nous.anthropic_wire: native opts back in. Cost of
chat: prior-turn thinking travels as OpenAI-style reasoning fields, and
cache_control scopes are translated by the portal's adapter.

Tests: new test_nous_anthropic_wire_default.py reads the knob through the
real config loader in a temp HERMES_HOME and through
resolve_runtime_provider (both settings); the existing native-wire
contract suite selects native via an autouse fixture so the wire keeps
working for the flip-back. Live: a real AIAgent on this branch,
provider=nous + anthropic/claude-fable-5.1, dispatches to
chat/completions with the session_id intact.
2026-09-06 05:45:44 -07:00
kshitijk4poor 256bd1adc9 feat(docker): terminal.docker_snap_compat opt-out for snap-packaged Docker under AppArmor (#9730)
On hosts where Docker ships as a snap (Ubuntu cloud images / Azure VMs), the
snap's AppArmor confinement turns two sandbox hardening flags into a dead
container at start: `--init` fails with "exec /sbin/docker-init: operation not
permitted" and `--security-opt no-new-privileges` then fails every exec the
same way ("exec /usr/bin/sleep: operation not permitted"). This is snapd
LP#1908448 — not probeable from the client, and docker_extra_args cannot remove
flags we add.

`terminal.docker_snap_compat: true` drops exactly those two flags; cap-drop ALL,
the tmpfs hardening, PID limits and the privdrop caps are unchanged, and a
warning is logged at container start. Bridged everywhere the other docker_*
keys are (CLI env map, gateway env map, `hermes config set` sync, terminal_tool
env read, the shared container_config shaper, DEFAULT_CONFIG).
2026-09-05 21:00:19 +05:30
Teknium ec4c1e0c98 fix(delegation): validate compression_threshold_tokens; state that it caps the trigger, not the payload
Independent review: a YAML `true` coerced to int 1 and gave every child a
one-token compression trigger; "200k" silently disabled the default cap.
Values are now validated: an int >= 16000 is used, 0/false/null disable on
purpose, anything else warns and falls back to the 200K default so a typo
never costs money. Config comment and docs say it caps the compaction
trigger, not a hard request-size limit.

Test: true / "200k" / 5 -> default; 0 / false / None -> disabled; valid ints
pass.
2026-09-05 06:29:28 -07:00
Teknium 40da0fd52f feat(delegation): subagents compress at an absolute context cap (delegation.compression_threshold_tokens, default 200K)
A delegate_task child inherits the compression threshold as a RATIO of the
model window. On a 1M-window model at the run's configured 0.85 that is an
850K-token trigger: in the 1,393-agent refactor run 1,373 of 1,375 children
never compressed once, 62% of all API calls carried >150K of context, and the
calls above 200K carried ~$10.9k of the $19.3k bill (58% of it cache WRITES,
i.e. re-sending a 300-800K prefix on every call). A sawtooth replay of the
logged calls with a 200K cap / 65K floor cuts context spend by ~49% (~$7k).

Children are brief-driven and disposable; they re-read their brief and the
files they touch, so a large window buys them little. New
delegation.compression_threshold_tokens (default 200000) is applied to the
child's ContextCompressor right after construction as the lower of it and
any global compression.threshold_tokens; the parent's own trigger is
untouched. 0 disables the subagent-specific cap. The compressor applies
threshold_tokens_cap on first window resolution, so this is byte-equivalent
to the user having set compression.threshold_tokens for the child.

Live through the real spawn path (_build_child_agent, real imports, temp
HERMES_HOME, 1M-window model): main child trigger 500,000 / branch 200,000;
parent 500,000 on both.

Tests (3): default caps a 1M child at 200K; the cap is the lower of the
delegation and global values and never raises a small-window child's
trigger; 0 disables and an already-resolved trigger is re-clamped.
Docs: delegation.md, configuration.md.
2026-09-05 01:07:09 -07:00
Teknium f1ccf436a2 feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:

- search: POST https://api.perplexity.ai/search (documented Search API),
  search_context_size=low so `snippet` stays description-sized;
  results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
  route behind `pplx content snippets` (the CLI's `content fetch` is
  deprecated upstream). web_extract has no query, so the URLs' path words
  serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
  backend set + credential ladder + availability probe (web_tools),
  registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
  dump key lists, nous_subscription direct-credential detection, setup
  summary, test conftests, docs.

Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
2026-09-04 07:17:00 -07:00
Teknium 0a5164cebe compat(plugins): tell users which installed plugins break on 2026-09-14, and stop loading them after
hermes_cli/plugin_compat.py is now the single source of truth for the compat window:
  COMPAT_REMOVAL_DATE = 2026-09-14; scan_plugin() statically finds `from F import n`, `import F` + `F.n`,
  alias forms and string targets against compat_manifest.json; compat_report() aggregates over the user's
  ENABLED external (non-bundled) plugins; disable_reason() decides the loader's skip.

Surfaces (all read from that one report):
  * CLI: yellow block under the banner naming plugins + date + `hermes plugins compat` (red + DISABLED after)
  * `hermes plugins compat [--json] [path]`: file:line, old -> new per hit; exit 1 while anything remains;
    `path` lets a plugin author scan their own checkout
  * `hermes doctor`: "Plugin import paths (removed Sep 14, 2026)" section next to the xAI retirement check
  * `hermes update`: post-update notice alongside the FTS/curator notices
  * Desktop: compat_report() writes HERMES_HOME/.plugin-compat-report.json (deleted when clean); Electron
    shows ONE warning dialog per distinct report after the backend is up and persists the dismissal in
    userData/plugin-compat-dismissed.json. A new affected plugin, or the date passing, is a new report.

From the date, PluginManager skips a hitting external plugin before importing it, with the reason in
LoadedPlugin.error ("uses N import path(s) removed on 2026-09-14; run `hermes plugins compat` ...") — the
same path a plugin with a broken register() takes, so nothing else is affected. Escape hatch:
plugins.allow_deprecated_imports: true (config_defaults), which only helps until the compat commit is
actually reverted.

Docs: COMPAT_MANIFEST.md (removal date, what-happens table, author instructions), plugin dev guide section.
Tests: tests/test_plugin_compat_notice.py (scanner forms, report scope, date gate + escape hatch, summary
text, report file lifecycle, loader skip via a real PluginManager), electron/plugin-compat-notice.test.ts
(show once, re-show on a different set or on the date passing, malformed file ignored).

Live A/B on this box with a demo plugin on old paths: before the date it loads and the banner/doctor/report
name it; with today=2026-09-14 it is skipped with the reason and the banner turns red; with the escape
hatch it loads again.
2026-09-04 01:28:31 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium c0373cb95f refactor(hermes_cli): config_defaults — comment reflow to 100 cols (every word preserved, checker-proven), OPTIONAL_ENV_VARS via per-category _prov/_tool/_msg/_skill/_setting/_base_url factories (dicts byte-identical incl. key order), drop blank lines before comment blocks 2026-09-03 00:04:05 -07:00
Zane Chee b1bf099f14 fix(computer-use): distinguish doctor and gateway environments 2026-09-02 21:49:23 -07:00
Teknium e3d78021f3 refactor(hermes_cli): codex_runtime_switch synonym table + _migration_lines; context_switch_guard/credential_lifecycle defensive collapse + doc compaction; config_defaults drop key-restating section comments (values byte-identical) 2026-09-02 20:31:42 -07:00
Teknium 61c83bafef refactor(config_defaults): compact comments by hand (keep rules/invariants, drop history) 2026-09-02 15:56:54 -07:00
Teknium 2d0d797ceb refactor(config_defaults): build auxiliary task blocks via _aux() 2026-09-02 15:23:53 -07:00
Teknium 86100e5321 refactor(config_defaults): build OPTIONAL_ENV_VARS via _env() and collapse scalar sub-dicts 2026-09-02 15:21:52 -07:00
kshitijk4poor cfe88a1f7d feat(config): context_file_read_timeout key + narrow reader catch
Expose the read deadline as a top-level config.yaml key beside
context_file_max_chars (same load_config_readonly resolution shape), default
5s, documented in context-files.md. Narrow the reader thread's catch from
BaseException to Exception: control-flow exceptions can't originate inside
read_text on a worker thread, and re-raising one would bypass the sites'
except Exception / except (OSError, UnicodeDecodeError) handlers.
2026-09-03 03:15:31 +05:30
Teknium 5d4b97939e refactor(hclib): config — config/config_migrations/tools_config/toolset_* dispatch tables and dedupe 2026-09-02 14:42:18 -07:00
olopez25 c6c8c74c70 Move the keepalive interval to config.yaml and tighten the schedule guard
Addresses review feedback on #84928.

The tick interval was exposed as HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS.
AGENTS.md reserves .env for credentials and puts behavioural thresholds in
config.yaml, so the knob moves to `nous.keepalive_interval_seconds`,
following the existing `vertex:` section's precedent for non-secret
provider settings. The env var is dropped rather than bridged: it was never
released, so nothing depends on it. Adding a key to a new section is handled
by the deep-merge, so no _config_version bump is required.

test_keepalive_interval_fits_inside_the_token_lifetime asserted
`900 < 899 * 4 - 120`, which is true for any realistic interval and could
never fail. It also tested the wrong value: the configured constant is only
a ceiling, while the schedule that ships is the derived tick. Replaced with
an assertion over the derived tick for each observed lifetime, which does
fail if the derivation constants regress -- verified against both
TICKS_PER_LIFETIME=1 and MIN_INTERVAL_SECONDS=5000.

Also adds coverage for an unreadable config.yaml, which must fall back to
the module default rather than take the keepalive thread down.

pytest tests/hermes_cli/test_nous_auth_keepalive.py -> 9 passed

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 10:22:30 -07:00
Teknium 73f68362b3 fix(sessions): auto-prune state.db by default (90d) and gate VACUUM on freelist ratio (#54189)
Flip the state.db retention defaults per Teknium's decision on #54189:

- sessions.auto_prune: false -> true. A stock install now prunes ENDED
  sessions inactive for retention_days at CLI/gateway/cron startup
  (at most once per min_interval_hours). Open, pinned and mid-turn
  sessions are never deleted; the only open rows touched are stale
  automation sessions (#100903 sweep), which are closed, not deleted,
  and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
  the file: PRAGMA freelist_count / page_count must exceed 25%
  (AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
  min_vacuum_interval_days throttle. Pruning a few small sessions on a
  dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
  Unknown ratio (pragma read failure) falls back to the time throttle.

Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.

Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
2026-09-02 07:26:52 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
jinglun010 7aff724e56 feat(desktop): add display.resume_last_session config toggle (#60812)
Config default (true) plus the Appearance-settings strings for a
"Reopen Last Chat on Launch" switch. Salvaged from PR #60816 onto
current main (defaults moved to config_defaults.py since the PR).
2026-09-02 05:56:54 -07:00
Teknium 552159d222 feat(cli,tui): collapse bell_on_clarify/approval into display.bell_on_prompt
One key covers every blocking prompt modal: clarify (single + batch),
dangerous-command approval (incl. computer_use), sudo password, and
secret capture. CLI gets a _ring_bell() helper shared with
bell_on_complete; TUI rings on clarify/approval/sudo/secret .request
events (isTTY-gated). 'hermes config' Bell summary shows both flags.
2026-09-02 05:34:35 -07:00
Turgut Kural 3082a34669 feat(cli,tui): add display.bell_on_approval + fix eslint error
- display.bell_on_approval (default false): same BEL mechanism as
  bell_on_complete, rings when a dangerous-command approval prompt
  opens (_approval_callback / approval.request event). Complements
  bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
  (if without braces) that failed the CI JS & TS checks job.
2026-09-02 05:34:35 -07:00
Turgut Kural ef6d3367a6 feat(cli,tui): add display.bell_on_clarify — terminal bell on clarify prompts
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
2026-09-02 05:34:35 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 45b0d8cab5 feat(gateway): one gateway.trust_env key controls aiohttp proxy-env honoring at every adapter site (#48820 bug 3)
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.

- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
  resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
  auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
  teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
  bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.

Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
2026-09-02 04:13:02 -07:00
Teknium 00a7115a02 fix(cron): make cron push-notify configurable (cron.delivery.notify) and surface UNVERIFIED live deliveries in cron list/doctor
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.

An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.

Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
2026-09-02 00:56:52 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
Teknium 9de9d7613c fix(compression): keep hygiene turn-hold worker's commit admission so thinking-model summaries are adopted, not burned
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.

Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):

- CompressionCommitFence gains mark_commit_watermark_fenced() /
  commit_watermark_fenced; compress_context marks the fence right after
  capturing get_active_message_watermark() under the durable compression
  lock (#75316/#87484) — the property that makes a LATE commit safe: rows
  appended after compression start survive both commit paths verbatim as
  cloned concurrent tail (archive_and_compact watermark= and
  publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
  the detached worker (already kept alive via
  _defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
  user's turn proceeds on the uncompressed transcript at the same 10s
  budget, and the summary is adopted at the worker's own watermark-fenced
  commit boundary. Unfenced workers are cancelled exactly as before —
  never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
  block preflight adoption via the same-session cooldown); re-attempt
  spacing is covered by the durable compression lock
  (_session_has_compression_in_flight). If the worker ends WITHOUT
  committing, a done-callback restores the flat non-escalating 60s
  retry-after; a successful adoption resets the hygiene failure streak.
  The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
  to describe deferred adoption and the thinking-model case;
  config_defaults.py comment updated. Knob stays config.yaml-only.

Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
  test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
  UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
  path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
  start watermark; the fence still gates admission and unfenced/late
  results are discarded.

New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
  turn still released at the budget, no cooldown while running,
  streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
  turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.

Fixes #97963
2026-09-01 23:56:23 -07:00
Ben Barclay 180291162f feat(telemetry): opt-in shared-metrics exporter (#95278)
feat(telemetry): opt-in shared-metrics exporter
2026-09-02 08:35:36 +10:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium 67de93862c fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).

The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).

Fixes #98028
Fixes #100325
2026-09-01 10:52:23 -07:00
Teknium f709bd88b6 feat(skills): render the configured create dir in every instruction that names the path
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
2026-09-01 07:32:45 -07:00
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30
Kshitij Kapoor 44eeef9a83 fix(gateway): apply startup-watchdog config to the already-armed handle
/simplify-code quality+efficiency reviewers (converged, verified): on
the standard 'hermes gateway run' path the argv fast-path arms BEFORE
run_gateway's config bridge executes, and arm_startup_watchdog() is
idempotent — so gateway.startup_watchdog: false and
startup_watchdog_timeout_seconds were dead knobs (env bridged, live
handle untouched). run_gateway now applies the config to the live
handle: disarm on disable; disarm+re-arm on a bridged config timeout so
the fresh handle covers the remaining pre-loop startup with the
configured deadline. config_defaults comment updated to match reality.

E2E (real module, fast-path armed first): disable path disarms the live
handle; timeout path re-arms a fresh handle at 123s.
2026-08-31 14:01:39 -07:00
Kshitij Kapoor d2c3c38e98 fix(gateway): config.yaml surface for the startup watchdog + precise argv arming
Review follow-ups on the salvaged #89750:

- gateway.startup_watchdog / gateway.startup_watchdog_timeout_seconds in
  config_defaults, bridged to the internal HERMES_STARTUP_WATCHDOG env
  vars in run_gateway() (the argv fast-path arms before config can load,
  so env remains the mechanism; config.yaml is the user-facing surface
  per policy — explicit env values still win as operator override).
- hermes_cli/main.py argv sniff now requires the ADJACENT token pair
  'gateway run' instead of independent membership, so unrelated commands
  mentioning both words can't arm a 300s hard-exit timer; profile-flagged
  invocations (-p work gateway run) still arm.
2026-08-31 14:01:39 -07:00
Teknium fb9b2c893f feat(agent): escalate repeated transcript-sanitiser heals with a one-time user notice (#96870)
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):

- Per-session heal counters (heal events + messages healed) tracked by the
  repair path in agent_runtime_helpers.py, session totals preserved across
  10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
  configurable via agent.sanitizer_heal_escalation_threshold in
  config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
  (events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
  the conversation loop through _emit_warning (status callback -> gateway
  status message / CLI print). Never injected into conversation context or
  the wire copy: prompt caching, role alternation, and durable history are
  untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
  the /debug share // hermes debug report, and the config key surfaced in
  hermes dump overrides. errors.log carries the ERROR line for `hermes logs
  errors`.
2026-08-31 13:11:41 -07:00
LucidPaths 0943702c55 fix(gateway): keep long turns controllable without blocking Telegram 2026-08-31 12:21:01 -07:00