Commit Graph

27169 Commits

Author SHA1 Message Date
Teknium aaa34b0e08 fix(desktop): model picker no longer hardcodes --global; one persist policy for every surface (#90235)
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).

Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).

Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
  session-only rule: when neither model.default nor model.provider is
  configured yet, persist. This preserves #86414's first-pick motivation
  for CLI, gateway and Desktop alike. With a default configured, a plain
  pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
  Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
  lists --global.
- Docs: desktop.md picker note + slash-commands /model row.

Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
2026-09-02 05:33:33 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 2e25d179cb chore(contributors): map ifastcc email (PR #34308 co-author) 2026-09-02 05:33:13 -07:00
Aldo 6ff9426d22 fix(anthropic): track the current fast-mode model matrix (Opus 4.8 / Opus 5)
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):

- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
  only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
  requests silently run at standard speed and bill standard rates
  (usage.speed: 'standard'). Today's allowlist therefore shows 4.6
  users a fast toggle that does nothing, while denying it to the two
  models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
  select fast inference via the model field and are explicitly
  excluded from the param gate.

Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.

## How to test

scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q

113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
2026-09-02 05:33:13 -07:00
Teknium 4d02c78102 refactor(email): single _normalize_security helper, loopback-scoped verify warning, tests + docs
Follow-up to the #99641 salvage:
- One module-level _normalize_security() (ssl/tls/implicit -> tls, starttls,
  plain/none -> plain; unknown -> WARNING + secure default) replaces the three
  copies of the alias set; _connect_imap/_connect_smtp/_standalone_send all
  compare against the canonical value. Unknown modes no longer raise.
- _tls_context(verify, host) is module-level and shared by all sites; when
  verification is disabled for a non-loopback host it logs a WARNING.
- _esecret_bool: an unset/empty env var now yields the caller's default
  (previously is_truthy_value('') returned False, silently disabling TLS
  verification whenever EMAIL_*_TLS_VERIFY was unset).
- Documented surface is platforms.email.extra.{imap,smtp}_security and
  {imap,smtp}_tls_verify in config.yaml; env vars remain an internal bridge
  and are NOT added to plugin.yaml (optional_env feeds hermes setup prompts).
- Docs: Proton Mail Bridge / local relays recipe in user-guide/messaging/email.md.
- Tests: starttls builds IMAP4 then .starttls(); unknown mode falls back to
  tls/starttls with verification still on.
2026-09-02 05:32:31 -07:00
Alessandro Lamberti 92a9864517 feat(email): configurable IMAP/SMTP transport security (tls/starttls/plain) and TLS verify toggle
Adds EMAIL_IMAP_SECURITY / EMAIL_SMTP_SECURITY and EMAIL_IMAP_TLS_VERIFY /
EMAIL_SMTP_TLS_VERIFY (env or platforms.email.extra.*) so the adapter can
talk to local relays such as Proton Mail Bridge (IMAP 1143 / SMTP 1025 with
STARTTLS and a self-signed certificate) instead of hardcoding IMAP4_SSL and
SMTP+STARTTLS with a verified default context.

Salvaged from #99641 (adapter.py only).
2026-09-02 05:32:31 -07:00
Teknium ab1f81ce04 chore(contributors): map umit.ediz@hotmail.com -> Edizzier 2026-09-02 05:32:20 -07:00
Teknium 44a57921c5 fix(providers): hide phantom -cn picker rows lit only by shared intl keys; give alibaba-token-plan-cn its own key var
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
  as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
  ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
  keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
  ones it shares with its non-CN sibling, unless that CN provider is the
  configured model.provider. With only the shared key: one row, not two;
  DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
2026-09-02 05:32:20 -07:00
Edizzier 805498e6df fix(providers): give alibaba-coding-plan-cn its own API key env var
ALIBABA_CODING_PLAN_CN_API_KEY is checked first for the China Coding Plan
endpoint (mirroring kimi-coding-cn), so the intl and CN rows no longer
light off the same key. Fixes #101122.
2026-09-02 05:32:20 -07:00
Teknium a81e503866 fix(kanban): reap notify subscriptions for stale blocked tasks too
purge_stale_done_notify_subs only matched status='done', so a task the
circuit breaker parked in 'blocked' kept its notify-sub rows forever on
boards that never archive. Widen the predicate to done OR blocked while
keeping the existing age clause; backlog/ready cards are idle, not
abandoned, and stay exempt (test_gc_spares_reopened_task_even_when_old).
Watcher comment/log and docs updated to say done/blocked.

Closes #100955

Co-authored-by: itsflownium <itsflownium@users.noreply.github.com>
2026-09-02 05:32:10 -07:00
itsflownium 97baa38885 test(kanban): pin that stale blocked-task notify subs are purged
Adapted from #101103: a task parked in blocked past the retention window
must have its notify subscriptions reaped like a stale done task.
2026-09-02 05:32:10 -07:00
Teknium 5c6cbbc1be fix(loops): pause /loop --until on a blocked verdict; trim redundant gate condition and duplicate test
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
2026-09-02 05:32:01 -07:00
itsflownium ae63db2304 chore: map contributor email 2026-09-02 05:32:01 -07:00
itsflownium 1bd9fce6cb fix(kanban): judge unachievable goals as blocked, never done 2026-09-02 05:32:01 -07:00
Teknium d3e2ace1dd fix(profiles): profile delete refuses to kill another profile's gateway (#89315)
`hermes profile delete` read the target profile's gateway.pid raw and
SIGTERMed it. When that pid file was poisoned by a sibling profile's gateway
(the #89315 shape), deleting profile A killed profile B's running gateway.

- gateway/status.py: `_pid_record_belongs_to_profile()` helper — a pid
  record whose recorded home differs from the expected profile home is not
  ours; legacy records without a home prove nothing and are left alone.
- hermes_cli/profiles.py: `_stop_gateway_process` refuses (and says so)
  when the record belongs to another profile; still stops its own gateway.

The stop/restart paths in hermes_cli/gateway.py did not need a guard:
`get_running_pid()` already filters cross-profile records and unlinks the
poisoned pid file before any kill can happen — verified live; the test for
that path now pins the real contract (returns False, other process alive,
poisoned pid file gone).

Live repro (unpatched main): `_stop_gateway_process(tim_home)` -> "Gateway
stopped (PID ...)" and the OTHER profile's process exits -15. After: "Refusing
to stop PID ..." and the process stays alive. 8 tests; sabotage (guard
removed) fails 1.
2026-09-02 04:53:59 -07:00
Ayush Nangia 3a980a431b fix(desktop): keep drafts editable while connecting 2026-09-02 04:36:51 -07:00
kshitijk4poor e9160625dc fix(state): also disable SQLite's internal close-time checkpoint on quarantine (py3.12+)
Skipping the explicit PRAGMA wal_checkpoint(PASSIVE) in close() left
sqlite3.Connection.close() running SQLite's own last-connection PASSIVE
checkpoint, which still checkpoints the WAL and unlinks -wal/-shm on a
structurally corrupt file (E2E: the -wal vanished on close despite the
quarantine). Python 3.12+ exposes SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE via
Connection.setconfig(); arm it in _halt_db_corrupt so the WAL image
survives close() for forensics/recovery. On 3.11 the switch does not
exist; the docstring and docs now say so instead of claiming sqlite3
cannot reach it at all.

Follow-up to #101095; flagged by JoaoMarcos44 on #101093.
2026-09-02 16:57:21 +05:30
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
leomcamilo d8616f1c88 chore: map contributor email for leocamilo@me.com 2026-09-02 16:57:21 +05:30
Teknium dbb6acd333 test(state): non-contention errno table, repair-lock sibling, in-process deferred-FTS retry via housekeeping tick
Regression coverage for the #100130 salvage, all against real SessionDB
files and a real child process holding the flock:

* errno table for `is_advisory_lock_contention` (EAGAIN/EWOULDBLOCK/EACCES
  contend; ESTALE/ENOTSUP/ENOLCK/EIO fail fast); no misleading "held by
  another process" line on the fast-fail path; `_cross_process_repair_lock`
  shares the filter (sibling site).
* `retry_deferred_fts_recovery`: open under a live holder -> stale; retry
  returns in <2s with a 30s admission budget (timeout=0); rate limit +
  60s->120s backoff engaged; holder dies -> same instance recovers, triggers
  restored, breadcrumb cleared; no-op when not stale / read-only.
* `_start_gateway_housekeeping` tick (real loop, 50ms interval) recovers a
  stale shared-registry SessionDB with no direct call and no extra thread.

Backoff floor: a monkeypatched 0s base interval must not zero the doubled
interval (min 1s), so the cap math is testable.

Sabotage run (source at origin/main, these tests): 16 failed / 35 passed,
including 30s timeouts on the fast-fail tests.
2026-09-02 04:15:02 -07:00
HexLab98 c5138618f7 test(state): cover deferred FTS retry, leftover lock files, and WAL interpreter identity 2026-09-02 04:15:02 -07:00
teknium1 fd05029430 fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:

* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
  EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
  ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
  environment failures that polling cannot fix — `_acquire_db_flock` and
  both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
  now defer immediately with the real errno instead of burning the full
  120s / holder timeout and then logging a fake "held by another process".

* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
  open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
  lock) stayed `_fts_stale` — LIKE-only search — until the process
  reopened state.db. Short-lived CLIs reopen every run; the gateway opens
  once and stays up for days, so the deferral was effectively permanent
  (#100108). The retry runs from the EXISTING gateway housekeeping tick
  (`_start_gateway_housekeeping`, 60s) against the shared SessionDB
  instances via `hermes_state_registry.live_shared_session_dbs()`:
  non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
  bounded backoff 60s -> 1h, no new thread, still fails closed on live
  holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.

* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
  line can be matched to the interpreter that actually linked it
  (#100108 point 3).

Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 04:15:02 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
Teknium 45b0d8cab5 feat(gateway): one gateway.trust_env key controls aiohttp proxy-env honoring at every adapter site (#48820 bug 3)
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.

- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
  resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
  auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
  teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
  bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.

Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
2026-09-02 04:13:02 -07:00
Teknium bc71b8bc95 fix(compression): anchor on the LAST intent row — newer user turn outranks older steer (#100053 follow-up)
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.

Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
2026-09-02 04:12:12 -07:00
finn763 f40333a80e fix(agent): preserve busy steer during compression and avoid replaying historical user request
Compression with display.busy_input_mode: steer embeds the follow-up
as an out-of-band marker inside the latest role=tool result. The
post-compression user-turn preservation path only classified
non-scaffolding role=user rows as real intent, so a compressed
transcript that contained no role=user row would discard the steer
and clone an older historical role=user message as the new active
turn, re-activating a previously consumed request.

Fix _ensure_compressed_has_user_turn to (1) treat a compressed
transcript that already carries a steer marker as having user intent,
and (2) prioritize the latest steer payload from the original
transcript over historical user cloning, inserting it as a proper
role=user turn via _insert_real_user_anchor. This preserves the
actual current intent exactly once and never turns history into new
input.

Closes #100053
2026-09-02 04:12:12 -07:00
kshitijk4poor 6f1733ca22 test(state): reset the single-flight _opening map in the registry fixture
The _clean_registry fixture clears _generations and _retired between
tests; the new _opening map needs the same reset so a test that aborts
mid-construction cannot leave a stale opening event that stalls the
next test's cold acquire.
2026-09-02 16:37:21 +05:30
fangliquanflq 61635e1b53 fix(state): single-flight shared database opens 2026-09-02 16:37:21 +05:30
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium 6e7c7c7da9 fix(desktop): a bot row click always lands on the Bot Chat the row previews
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.

- roster-actions: the open-tab shortcut may front only the canonical chat
  (registry id or lineage tip, via a new onlyStoredIds allowlist on
  focusWorkspaceOwnerSessionTile); anything else resolves the registry and
  opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
  menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
  the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
  none reach this window's stream), re-open it in place so the transcript
  refreshes instead of waiting for an app restart (#99393 class).

Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.
2026-09-02 03:41:44 -07:00
hermes-seaeye[bot] 254158f453 fmt(js): npm run fix on merge (#101150)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 10:12:08 +00:00
Teknium 3e63367175 feat(desktop): status bar can show live cache-hit rate and tokens/sec (off by default)
Two new right-click-toggleable status bar items, mirroring the CLI/TUI
Pantheon status bar upgrades: prompt-cache hit rate ("87%") and rolling
output throughput ("42 t/s"). Both are hidden by default and enabled from
the bar's existing 'Show in status bar' context menu, like the context meter.

Renderer-only: the tui_gateway already emits cache_hit_pct and avg_tps in
every session.usage tick and message.complete payload, so the items ride
the same UsageStats the context meter reads — no new RPC, no polling.
Labels show a placeholder until the backend has data, never self-hide.
2026-09-02 03:06:20 -07:00
Teknium 5d4aa4fcb2 fix(desktop): group chat rooms are serial again; keep only the push-woken turn poll
#101112 made round members take their turns concurrently. That changed what
a group chat IS: later speakers in a round no longer saw earlier speakers'
replies, so bots answered the user independently instead of building on
each other. Group rooms are serial round-robin by design — this restores the
pre-#101112 round engine (group-rounds.ts, group-chat.ts, group-chat-view.tsx,
their tests, and the docs) byte-for-byte.

What stays from #101112: the per-turn poll wakes on the member session's
terminal frame (message.complete / error via host.onEvent) instead of
sleeping a fixed 2s between session.resume reads; 5s timer kept as backstop.
That is a pure latency fix with no change to room semantics.

Live A/B (real tui_gateway over WS, 4 members, one serial round):
2s poll 32.5s -> push-woken 22.5s. The remaining time is model latency.

Refs #92760
2026-09-02 02:54:46 -07:00
kshitijk4poor bcb412cd9d refactor(cron): share the fire-path telemetry recorder
_record_timezone_migration_catchup was a line-for-line clone of
_record_persisted_error_recovery (counter bump, bounded recent list,
best-effort jsonl append). Extract _append_telemetry_record and route
both through it; one shared history cap replaces the two per-counter
constants. Also correct the "distinct from catch_up_occurrences" comment:
a migrated row that is also past its grace window increments both.

No behavior change; both recorders write the same entries to the same
files.
2026-09-02 15:13:46 +05:30
deinte 187251800b fix(cron): don't silently skip a due run after a timezone-offset migration
Upgrading from a UTC-scheduling build to one that honours the profile
timezone (Europe/Brussels) left daily cron jobs sitting in jobs.json with
pre-migration instants — e.g. next_run_at "2026-09-02T04:00:00+00:00" for
expr "0 4 * * *". _ensure_aware normalizes that to 06:00+02, which the
expression excludes, so the stale-expression guard (#93049) read it as a
direct jobs.json edit, logged exactly that, and re-anchored to tomorrow
without firing. The due occurrence disappeared with no error anywhere.

The guard only asked "is the stored instant an occurrence of the current
expr?", never "why not?" — and the two possible answers demand opposite
actions. Add _classify_stale_cron_next_run, which distinguishes them by
whether normalization itself moved the wall clock:

  * expr_edit           — wall clock unchanged (or the stored wall clock is
                          not an occurrence either): the instant is genuinely
                          excluded by the current expression. Re-anchor
                          without firing, exactly as before.
  * timezone_migration  — the stored value's own wall clock IS a legal
                          occurrence and it only left the lattice because
                          _ensure_aware converted it to a different offset.
                          Fall through and fire the overdue run once.

Because every value written by this build carries the configured offset, a
real expr edit leaves the wall clock untouched and can never be reclassified
as a migration, so the #93049 protection is intact. At-most-once is
unchanged: the fire flows through the normal due path and the usual
advance_next_run / mark_job_run re-anchor rewrites next_run_at in the
current offset, so the legacy instant is never read again. Future local
wall-clock occurrences are untouched — not-yet-due rows never reach the
guard, and the #28934 offset-repair branch still runs first for a
still-future stored wall clock.

The migration case is classified explicitly rather than retried broadly: it
logs cron.timezone_migration.catch_up with the stored and normalized
instants plus both offsets, and increments a probe-visible counter
(get_timezone_migration_catchup_stats, timezone_migration_catchups.jsonl)
kept separate from catch_up_occurrences so an operator can tell "the upgrade
backlog is draining" from "runs are missing their grace window".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 15:13:46 +05:30
kshitijk4poor d96b180eda chore: map contact@danteschrauwen.be -> deinte (PR #101090 salvage) 2026-09-02 15:13:46 +05:30
Teknium fb5023950e perf(desktop): group chat rooms answer in the time of one bot, not the sum of all
Bot Mode group rooms were slow by construction: the round engine ran every
member's turn one after another, and each turn found out its bot had finished
by re-reading session.resume on a fixed 2s timer. A 4-bot room paid
4 x (model latency + up to 2s) per round, serially.

- group-rounds: members of a round now take their turns concurrently
  (Promise.all). Rounds stay serial so bots still build on each other's
  replies. Each member's delta is computed at its own turn start and its
  watermark advances only to the pre-turn log length, so sibling replies
  that land while it thinks are delivered next round exactly once; a
  member's own replies are excluded from its delta by author (they are
  already in its session). Message cap enforced per round; stop path
  interrupts every member mid-turn (room.turn -> room.turns map).
- group-turns: the poll wakes on the member session's terminal frame
  (message.complete / error via host.onEvent), then re-checks at 250ms
  until session.running clears. The timer poll stays as a 5s backstop for
  hosts without the event tap. Feature-detected; node test harness unaffected.
- group-chat-view: "X is thinking..." lists every member mid-turn.
- docs: bot-mode.md describes concurrent rounds + push-woken replies.

Live A/B (real tui_gateway over WS, 4 members, one round, same model):
serial+2s poll 35.0s -> concurrent+push 8.6s; every turn woke on the event.

Refs #92760
2026-09-02 02:36:26 -07:00
hermes-seaeye[bot] fbd40c907d fmt(js): npm run fix on merge (#101107)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 08:51:02 +00:00
Teknium 7ff8aae8fb gateway: warn once when an explicit platforms.<x>.enabled: false overrides env credentials
De-risking for the #48820 behaviour change: before this branch, credentials
in the environment force-enabled twelve platforms regardless of an explicit
enabled: false in config.yaml. Now that the explicit disable wins, users who
relied on the old override would see the platform go dark with no trace.

_enable_from_env (and Slack's inline copy) now emit ONE WARNING per platform
per process when the platform is explicitly disabled AND its env credentials
are present, naming the platform, the winning key
(platforms.<x>.enabled: false), the env var(s) being ignored, and the remedy.
A plain disable with no credentials, an enabled platform, and the env-only
(no YAML opinion) path stay silent; repeated config reloads do not repeat it.
_ENV_ENABLE_CREDENTIALS maps every _enable_from_env platform to its
triggering env var(s); a test pins that the map covers every routed branch.

Docs: messaging/index.md gains a 'Disabling a platform whose credentials are
still in .env' section with the exact warning text.

Live repro (real load_gateway_config on a temp HERMES_HOME with
platforms.weixin/telegram.enabled: false + WEIXIN_TOKEN/TELEGRAM_BOT_TOKEN in
env): before — both stayed disabled with zero log output; after — one
WARNING each ('Platform 'weixin' is explicitly disabled by
platforms.weixin.enabled: false ... (WEIXIN_TOKEN, WEIXIN_ACCOUNT_ID) will
NOT start its adapter ...'), none for the enabled homeassistant, none on the
second load.
2026-09-02 01:50:45 -07:00
Teknium 867e4158f0 fix(gateway): honor explicit platforms.<x>.enabled: false over env credentials (#48820)
Twelve credential-presence branches in _apply_env_overrides (weixin,
whatsapp_cloud, homeassistant, email, sms, dingtalk, feishu, wecom,
wecom_callback, bluebubbles, qqbot, yuanbao) force-set enabled = True
unconditionally, so a user's explicit `platforms.<x>.enabled: false` in
config.yaml was silently overridden whenever the platform's token/secret
lived in .env. Telegram/Discord/Slack/Signal/Matrix already routed through
_enable_from_env, which honors the `_enabled_explicit` marker written by
load_gateway_config.

Route all twelve sites through the same helper. Credentials are still wired
into the (disabled) PlatformConfig so send-only tooling keeps working —
the same contract Slack and api_server already follow.

Live repro (real load_gateway_config against a temp HERMES_HOME, yaml
`enabled: false` + creds in env): 12/13 platforms flipped to enabled=True
on main; 0/13 after the fix (telegram control unchanged).

Bug 2 of #48820. Fix direction from @JoaoMarcos44 in #48852 (surgically
reapplied on current main — the June branch no longer applies).

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-02 01:50:45 -07:00
hermes-seaeye[bot] 6840bb02e8 fmt(js): npm run fix on merge (#101102)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 08:42:10 +00:00
Teknium aa80626764 fix(tui-gateway): adopt late compute-host compress acks instead of a false 120s timeout (#97948)
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.

- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
  registered when the waiter times out; control.ack/control.error/error
  frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
  crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
  `compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
  cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
  applies the metadata mirror and emits the same `session.info` a normal
  compress does plus the existing `status.update kind=compacted` edge; a
  late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
  on waiter timeout answer `status: pending` (not 5019) and register the
  late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
  `status: 'pending'` renders as an info notice, not `error:`; the
  `compacted` status edge rehydrates an idle active session's transcript
  (mid-turn compaction still defers to the turn settle path).

Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.

Refs #97948

Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
2026-09-02 01:35:59 -07:00
Teknium 9bc249c7e5 fix(state): report closed stale-open count from auto-maintenance, document the sweep (#54189)
Follow-up on top of the salvaged #94095 commit:
- maybe_auto_prune_and_vacuum() now returns 'closed' (stale open state-owned
  sessions marked ended) alongside 'pruned', so entrypoints can report the
  reconciliation without parsing logs.
- Docstring explains the two-window lifecycle (close now, delete after a
  further retention window).
- Regression test: cron/kanban/subagent rows with ended_at NULL are closed on
  pass 1 and deleted on pass 2; a telegram row is never touched.
- website/docs sessions.md documents the automatic stale-open sweep.
2026-09-02 01:20:47 -07:00
Alex 0aa84bb3fd fix(state): reap stale state-owned sessions safely 2026-09-02 01:20:47 -07:00
Teknium 2f44998353 fix(dashboard-auth): a non-JWT bearer is "not my token", not "provider unreachable" (#94558)
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.

Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.

Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).

This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.

Refs #94558
2026-09-02 01:15:58 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
kshitijk4poor 2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
kshitijk4poor 26f0de23cf fix(agent): clone the /refine snapshot too, not just the automatic review
Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).

Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
2026-09-02 13:28:41 +05:30
notkisk 619ca3011e fix(agent): isolate background review snapshots 2026-09-02 13:28:41 +05:30