Commit Graph

33064 Commits

Author SHA1 Message Date
Ahmett101 dc8044c086 fix(state): quarantine a state.db whose page 0 is not SQLite, not just a zeroed one
Startup only quarantined a 0-byte / all-NUL state.db. A file whose first page
was clobbered with record bytes (#102198) went straight to sqlite3.connect,
which raised "file is not a database" and deleted the -wal sidecar — the one
piece of evidence that could have been recovered.

`has_invalid_sqlite_header_preopen` generalises the zeroed probe (zeroed is a
subset of "no SQLite header"; same live-connection contract, never raises).
`quarantine_invalid_state_db` moves the file AND its -wal/-shm aside as
`state.db.<zeroed|notadb>-<ts>-<pid>.bak` before anything opens it; a fresh
DB is created as before.

Re-authored on current main (the quarantine helpers moved to
hermes_state_dbfile.py in d15c61b5dc); one regression test proves the
notadb case is quarantined with its sidecars and the new DB passes
integrity_check.

Refs #102198 (the write-after-SIGTERM that clobbers page 0 is not addressed
here; this preserves the evidence instead of destroying it).
2026-09-09 18:25:11 +05:30
kshitijk4poor 9e6c4100cb fix(agent): close the interrupted tool tail on the overflow terminal
The overflow-terminal path ends the turn without reaching finalize_turn, so
a transcript that overflowed right after a tool batch ended on a raw tool
result; strict providers reject the next user turn (tool -> user). Close it
with the same final text, mirroring the truncated-tool-call terminal above.

Also: classify once before either log so an overflow no longer emits a
"so the loop can continue" WARNING followed by the contradicting "NOT
seeding" one; reset the stale-streak breaker once for both branches; drop
the "compression could not recover it" wording (this path never reached
compression); trim the test file to the three tests that bind behaviour
(stream -> terminal stub; 413 stays non-terminal; the terminal ends the
turn, closes the tool tail, carries compression_exhausted).
2026-09-09 17:45:12 +05:30
ca-shrimp 3b0e81459b fix(agent): carry compression_exhausted bit; scope overflow terminal to context_overflow
Address review P1s (andrexibiza) on #106266:

1. The overflow-terminal exit in recover_from_truncation now forwards the
   #98722 typed compression_exhausted bit (partial_result/end_turn gained the
   flag) so the gateway resets/moves future input to a clean session instead
   of leaving the bloated durable session authoritative for the next turn.

2. _overflow_terminal is scoped to FailoverReason.context_overflow ONLY.
   payload_too_large (413) has its own byte-scored recovery owner
   (turn_overflow._recover_payload_too_large, #88960/#47339) that must not be
   bypassed; a post-delta 413 keeps its normal continuation stub. Regression
   covers both lanes.

Tests: unit asserts result compression_exhausted=True on the marker; a
real streamed partial hitting a 413 payload-too-large error keeps content and
is not terminal. 50 streaming/continuation/gateway regressions pass.
2026-09-09 17:45:12 +05:30
ca-shrimp ece584b8f5 fix(agent): don't seed continuation stub after a context-overflow stream death
When a stream delivered text and then died on a context-overflow /
payload-too-large error, the partial content (often tens of KB) was seeded as
a length-continuation stub, growing the transcript monotonically. In a
session whose transcript cannot be compressed back under budget
(protect_last_n covers everything -> no_progress, or the summary would
itself be larger -> would_grow), every later request is larger than the one
that just failed — an unrecoverable loop where the user sees a 30+ minute
fake hang and the only remedy is killing the session (#106260).

classify_api_error already labels these errors context_overflow /
payload_too_large (should_compress=True). _partial_stream_stub now returns
an EMPTY stub marked _overflow_terminal for that class instead of seeding
the recovered text, and recover_from_truncation treats the marker as
terminal: the turn ends via the recovery contract with a clear message
(start /new) and the transcript is not polluted with the partial.

Normal partials (network stall, output-cap truncation, tool-call drops) are
unchanged — only the overflow error class changes behavior.

Tests: stub marker + empty content; a real streamed partial hitting a
'maximum context length' error returns the terminal stub; recover_from_
truncation ends the turn (no fragment/nudge appended) on the marker while a
normal stub still runs the continuation path. 67 streaming/continuation
regressions pass.
2026-09-09 17:45:12 +05:30
hermes-seaeye[bot] ed2d821021 fmt(js): npm run fix on merge (#106554)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-09 12:04:16 +00:00
hermes-seaeye[bot] 9e7b4eca47 fmt(js): npm run fix on merge (#106539)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-09 11:58:08 +00:00
kshitijk4poor 809654d488 refactor(desktop): keep main's boot classification and share the ws-URL guard
Follow-up to the renderer-dial retry:

- Drop the `!connectionDescriptorResolved` gate so `bootFailureIsRetryable()`
  still decides every failure main can see (#82679 contract preserved for
  post-descriptor mint failures); the renderer-owned dial is OR'd in as the
  one failure main cannot classify. The budget check now short-circuits
  before the IPC, as on main.
- Four booleans → one `stage` variable; the `isGatewayReauthRequired` term
  is dropped (reauth is raised before the dial stage, so it was unreachable).
- `isGatewayWebSocketUrl` moves to `apps/shared/src/json-rpc-gateway.ts`
  and `JsonRpcGatewayClient.connect()` uses it, so the hook's "valid dial"
  predicate cannot drift from what `connect()` accepts.
- Tests: keep the three that bind behaviour (first dial retries under a
  stale ready snapshot; invalid URL terminal; post-connect failure
  terminal), drop the four that pinned the removed gate or duplicated the
  existing #82679 bound test. 46/46 green; reverting the hook to main makes
  the retry test fail.
2026-09-09 17:19:36 +05:30
cervantesh 60e90ed6e7 fix(desktop): retry transient remote gateway boot dials 2026-09-09 17:19:36 +05:30
nftpoetrist 511633be90 fix(agent): stop the run-budget wrap-up notice from mutating a persisted tool row
_maybe_inject_run_budget_wrapup() appends its wrap-up notice to the newest
role:"tool" message in place, with no _DB_PERSISTED_MARKER check. Its sibling,
_maybe_inject_iteration_budget_warning(), got exactly this guard added in the
same recent saga (turn_iteration_prep.py), with the comment "an older turn may
already be cached."

The reachability is structural, not an edge case: _maybe_inject_run_budget_wrapup
is only ever called from prepare_iteration(), at the START of the next iteration
-- strictly after tool_executor.py's _flush_session_db_after_tool_progress has
already flushed and marked the previous iteration's tool row persisted. So every
successful injection was mutating an already-persisted row: the wire request for
that turn carried the notice, but the durable transcript never did, diverging
replay from the live bytes and invalidating the provider's prompt-cache prefix
from that row onward.

Fix:
- Add the same _DB_PERSISTED_MARKER guard to _maybe_inject_run_budget_wrapup,
  scoped to the specific tool row the reversed scan lands on (not just
  messages[-1], since this function -- unlike its sibling -- scans backward for
  the newest tool row rather than only checking the tail).
- Wire _maybe_inject_run_budget_wrapup into _flush_session_db_after_tool_progress
  (pre-flush), mirroring exactly how _maybe_inject_iteration_budget_warning is
  wired in both places. Without this, the guard alone would make the notice stop
  firing in the common case, since prepare_iteration's call site almost always
  hits an already-persisted row -- the pre-flush call site is what actually lets
  it land in durable bytes.

Verified empirically: read the real call graph (tool_executor.py's three
_flush_session_db_after_tool_progress call sites cover every tool-completion
path) to confirm the guard's premise, then added an end-to-end test using a real
AIAgent + SessionDB that flushes and checks the persisted row for the notice
text. Mutation-verified: reverting the two production files drops exactly the 2
new/updated assertions (28 pass, 2 fail); reapplying restores green (30 passed).
Also ran the sibling iteration-budget-warning and /steer suites (71 passed) to
check for interaction regressions -- none.
2026-09-09 17:04:30 +05:30
Teknium bbf91b7650 docs(cron): describe snapshot-as-pin instead of the fail-closed drift guard
The user guide said an unpinned job "fails closed" on a global model change
and documented cron.model_drift_guard; both are gone. It now explains that the
job keeps running on its creation snapshot and how to move it (per-job pin or
cron.model).
2026-09-09 04:32:13 -07:00
Teknium bcff0a920e test(cron): replace fail-closed drift tests with snapshot-as-pin invariants
The old tests encoded the skip (agent never constructed, [drift_skip] text,
alert-once bit). New invariants, proven red on origin/main: an unpinned job
with model_snapshot=A / provider_snapshot=P runs with AIAgent(model=A) and
resolve_runtime_provider(requested=P) after the global default moved to B/Q;
an explicit job pin still beats the snapshot; cron.model / cron.model_provider
still beat the snapshot; a legacy job without a snapshot follows the global
default. test_cron_drift_alert_once.py is deleted with the bit it tested; the
config-notice and impact-summary tests drop guard_enabled / model_drift_guard;
the desktop toast test asserts the informational copy.
2026-09-09 04:32:13 -07:00
Teknium 7e4d02fef5 fix(cron): unpinned jobs run on their creation-snapshot model instead of failing closed
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).

The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.

Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".

Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).

The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
2026-09-09 04:32:13 -07:00
Teknium 48465c3933 fix(profiles): --clone-all no longer copies cron jobs into the new profile
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.

Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.

--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
2026-09-09 04:27:24 -07:00
Teknium fd6434b3b3 chore: map contributor email for #47026 salvage 2026-09-09 03:52:47 -07:00
Teknium e1838c5b5a fix: trim opencode-go 422 salvage to the invariant set
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
2026-09-09 03:52:47 -07:00
ericmaddox bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Yuan Chenglu (袁成路) 6a0f519d1e fix(opencode-go): set supports_vision_tool_messages=False for Xiaomi MiMo backend
## Problem

When using the opencode-go provider with Xiaomi MiMo models (e.g.
mimo-v2.5, mimo-v2.5-pro), the Hermes agent intermittently fails with:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

This occurs specifically when tool results contain multipart content
with image_url parts (e.g. browser screenshots). The opencode-go relay
forwards these as-is to the Xiaomi MiMo backend, which rejects list-type
tool message content while still accepting multimodal user messages.

## Root Cause

The OpenCodeGoProfile inherits supports_vision_tool_messages=True from
ProviderProfile (the default). When this flag is True, the agent sends
tool results with image parts directly to the model. However, Xiaomi
MiMo's API rejects this format:

> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."
>   — providers/base.py, line 73

The direct 'xiaomi' provider profile already correctly sets this to
False (plugins/model-providers/xiaomi/__init__.py, line 13), but the
opencode-go relay profile was missing this safeguard.

The relevant code path is in run_agent.py:_tool_result_content_for_active_model()
(line 4543), which checks _provider_supports_vision_tool_messages() when
deciding whether to embed images in tool-result messages.

## Fix

Add supports_vision_tool_messages=False to the OpenCodeGoProfile
instantiation in plugins/model-providers/opencode-zen/__init__.py.

This single-line change prevents tool-result images from being sent
as multipart content to the MiMo backend, while preserving the model's
image recognition capability through user messages and vision tool
invocations (both of which use different code paths unaffected by this
flag).

## Testing

Verified with the mimo-v2.5 model via opencode-go provider:

1. Browser tool + screenshot recognition
   → Navigated to https://www.baidu.com, took screenshot, identified
     top 3 trending topics from the image
   → Result: PASSED, recognized all topics correctly

2. Direct image as user message
   → Sent a screenshot PNG directly via --image flag, asked model to
     describe the content
   → Result: PASSED, model correctly read text from the image

3. Provider profile verification
   → Confirmed get_provider_profile('opencode-go').supports_vision_tool_messages
     returns False at runtime
   → Result: PASSED

4. No regression on non-MiMo models
   → opencode-zen provider retains supports_vision_tool_messages=True
     (unaffected)

---

fix(opencode-go): 为 Xiaomi MiMo 后端设置 supports_vision_tool_messages=False

## 问题描述

使用 opencode-go provider 搭配 Xiaomi MiMo 模型(如 mimo-v2.5、
mimo-v2.5-pro)时,Hermes agent 间歇性地抛出以下错误:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

该错误发生在工具返回结果包含 image_url 类型的 multipart 内容的场景下
(如浏览器截图)。opencode-go 中继层将这些内容原样转发给 Xiaomi MiMo
后端,而 MiMo 接受多模态用户消息,但拒绝 list-type tool message 内容。

## 根因分析

OpenCodeGoProfile 继承了 ProviderProfile 的默认值
supports_vision_tool_messages=True。当此标志为 True 时,agent 会将含
图片的工具结果直接发送给模型。但 Xiaomi MiMo API 拒绝此格式:

> providers/base.py 第 73 行注释明确指出:
> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."

直接的 'xiaomi' provider profile 已正确设置了该值为 False
(plugins/model-providers/xiaomi/__init__.py 第 13 行),但
opencode-go 中继 profile 遗漏了这一安全设置。

相关代码路径:run_agent.py 的 _tool_result_content_for_active_model()
方法(第 4543 行),该方法通过检查
_provider_supports_vision_tool_messages() 来决定是否在 tool-result
消息中嵌入图片。

## 修复方案

在 plugins/model-providers/opencode-zen/__init__.py 的
OpenCodeGoProfile 实例化中添加 supports_vision_tool_messages=False。

这一行改动阻止了 tool-result 图片以 multipart 格式发送给 MiMo 后端,
同时通过用户消息和 vision tool 调用的路径(使用不同代码路径,不受
此标志影响)保留了模型的图像识别能力。

## 测试验证

使用 mimo-v2.5 模型通过 opencode-go provider 验证:

1. 浏览器截图 + 图像识别
   → 导航至 https://www.baidu.com,截取首页截图,从图片中识别出
     热搜榜前三条
   → 结果:通过,正确识别所有热搜话题

2. 用户消息直接传图
   → 通过 --image 参数直接发送截图 PNG,要求模型描述图片内容
   → 结果:通过,模型正确读取图片中的文字

3. Provider profile 运行时验证
   → 确认 get_provider_profile('opencode-go')
     .supports_vision_tool_messages 在运行时返回 False
   → 结果:通过

4. 非 MiMo 模型无回归
   → opencode-zen provider 保持 supports_vision_tool_messages=True
     不受影响

## 修改文件

  plugins/model-providers/opencode-zen/__init__.py (+5 lines)

Signed-off-by: Yuan Chenglu (袁成路) <ycl_pj@163.com>
2026-09-09 03:52:47 -07:00
Teknium 416b2bd281 chore: map contributor emails for salvaged #76158 (mengyuyuan) and #105470 (webtecnica) 2026-09-09 03:52:03 -07:00
Teknium c2524736d1 fix(desktop): a late startup locale read never overrides the user's in-session pick
With the startup config read now retried (previous commits, #76158), a retry
can resolve after the user has already chosen a language in Settings and
would repaint the stale on-disk value over it. Mark the explicit pick in a
ref and let every retry outcome (success or exhausted fallback) defer to it.

Idea and guard from #105470 (webtecnica), folded onto the #76158 base.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 03:52:03 -07:00
mengyuyuan 5353fc9b32 fix(desktop): bound i18n locale retries and keep English fallback
Address review feedback on the startup-race fix:

- Restore the permanent-failure contract: a rejected config load settles
  on English (DEFAULT_LOCALE) so the UI stays usable, as enforced by
  context.test.tsx.
- Bound the retry loop to MAX_LOCALE_RETRIES (10) at 3s intervals instead
  of retrying forever — permanent auth/backend failures reach a terminal
  English state instead of scheduling requests indefinitely.

Tests: add coverage for transient-failure recovery (first attempt rejects,
retry applies persisted display.language) and for bounded-budget exhaustion
(no retry fires after the budget is spent).
2026-09-09 03:52:03 -07:00
mengyuyuan 140eaf9e03 fix(desktop): retry i18n locale load after backend startup race
The renderer mounts before the desktop's own backend is ready, so the
first GET /api/config times out after 60s. The catch handler fell back
to DEFAULT_LOCALE (en) permanently, leaving the UI stuck in English
until a manual language switch, even when config.yaml persisted
display.language: zh.

Retry with a 3s delay on failure instead of locking the default locale.
The backend comes up, the next attempt succeeds, and the persisted
language takes effect. Cleanup clears the pending retry timer on unmount.
2026-09-09 03:52:03 -07:00
Teknium 4b6c5bee4c docs(backup): document hermes backup --keep 2026-09-09 03:33:14 -07:00
buihongduc132 c1ff9390f6 fix(backup): prune old hermes-backup-*.zip after each run, keep last 3
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
2026-09-09 03:33:14 -07:00
Teknium bf28c2fe1f fix(desktop): local endpoint probes ignore HTTP(S)_PROXY; non-2xx names the status (#63472)
httpx honours the env/system proxy (on Windows, the registry ProxyServer
even with no *_PROXY vars) but never the bypass list, so a system proxy
(Clash, corporate) answered 127.0.0.1 probes from both Desktop validators
with its own error page. That parsed as models=[] and the GUI said
"advertised no models at /v1/models" for a llama.cpp server the CLI
(urllib, honours <local>) saw fine.

Local endpoints (loopback, LAN, Tailscale via is_local_endpoint) now
probe with trust_env=False; public endpoints keep honouring env proxies.
A reachable endpoint answering non-2xx with no model list reports
"<url> answered HTTP <status>." instead of an empty catalog, so the
onboarding card stops telling the user to start a model.

Reimplemented on the decomposed router (the original patched
web_server.py before the split). Diagnosis and fix direction by
Solitud1nem in #63656; Windows registry-proxy confirmation by
Ulysses-Gaia on #63472.

Live repro (real loopback server, HTTP_PROXY=http://127.0.0.1:9):
  before  ok=False reachable=False 'Could not reach .../v1/models'
  after   ok=True  models=['Qwen3.6-35B-A3B-Q5_K_M.gguf']

Co-authored-by: Solitud1nem <76743883+Solitud1nem@users.noreply.github.com>
2026-09-09 03:33:06 -07:00
Teknium 734461d213 fix(models): same-URL custom endpoints stop evicting each other's cached catalog; no-probe picker opens revalidate
Two picker-freshness defects in cached_fetch_api_models():

1. The disk cache row was keyed on base_url only, with the credential
   fingerprint stored inside the row. N custom_providers entries sharing
   one proxy URL with different keys (#106184) took turns overwriting the
   single slot; every sibling then failed the fingerprint check, got an
   empty catalog, and disappeared from the Desktop pickers (which hide
   zero-model rows). Key on url#fingerprint so each credential owns a row.

2. cache_only opens (Desktop model.options without refresh) served a
   past-TTL row for up to 7 days without ever revalidating, so a model
   loaded on a non-current local endpoint stayed invisible until the user
   found "Refresh Models". Serve the stale row AND spawn the same
   off-thread SWR refresh the blocking path uses; the caller still never
   waits on the network.

Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
  GUI no-probe open  before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
                     after  {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
2026-09-09 03:33:06 -07:00
briandevans f4c55323fa fix(cli): back up config.yaml before --reset overwrites it
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.

The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.

Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.

Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.

Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).

Refs #3522
2026-09-09 03:32:58 -07:00
Teknium bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium 06dc51d62d fix: verification evidence ledger is inert while verify_on_stop is off
The ledger in verification_evidence.db exists only to feed the verify-on-stop
guard, but the recorder kept running on every foreground terminal command and
every file edit after #53552 turned the guard off by default. Users who never
opted in still accumulated a multi-MB database (7 MB / 4.6k rows on one install).

Every ledger entry point (record_terminal_result, record_verify_run,
mark_workspace_edited, verification_status) now checks verify_on_stop_enabled()
first and returns without opening or creating the database when the guard is
off. verification_status reports {"status": "disabled"} in that case; no client
consumes the verification.status RPC yet, so nothing downstream changes.

Existing ledger tests pin HERMES_VERIFY_ON_STOP=1 since they exercise the ledger
itself; the new test proves the off path never creates the file (red on base).
2026-09-09 02:35:41 -07:00
Teknium 677e8ed8a4 fix(desktop): SSH remote backend stops following the host's sticky active_profile
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".

Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.

Live repro (before/after, temp HERMES_HOME with active_profile=foo):
  serve --isolated --ssh-session-token-file ...   hermes_home=<root>/profiles/foo -> <root>
  same + --profile foo                             hermes_home=<root>/profiles/foo (unchanged)
  serve (no token file, user CLI)                  hermes_home=<root>/profiles/foo (unchanged)
2026-09-09 02:34:59 -07:00
Teknium d0df324862 fix(tui-gateway): /review shows its reviewer in the Desktop subagent stack
`slash.exec` runs on the RPC pool, outside any turn, so `/review` dispatched the
reviewer with no HERMES_UI_SESSION_ID and no steer authority bound. delegate_task
registered the child with `owner_session_id=None`, `subagent.list` (owner-scoped)
returned nothing for the parent session, and the Desktop status stack's 5s
snapshot poll reconciled the live `subagent.start` row away — the user saw only
"Review started. Results will return here." with no subagent card.

Bind the same session identity a turn binds (`_set_session_context(...,
ui_session_id=sid)` + `_current_runtime_session_record`) around `start_review`,
and clear it after. The reviewer now registers under the parent sid with the
request transport as authority, so `subagent.list`, steer, stop and the Desktop
roster all see it.

Live repro (tui_gateway stdio, real OpenRouter reviewer):
  before: registry owner_session_id=None, owner_transport=NoneType;
          subagent.list -> {"subagents": []}
  after:  owner_session_id=<parent sid>, owner_transport=StdioTransport;
          subagent.list -> [{"goal": "Review recent work", "status": "running", ...}]
2026-09-09 02:00:21 -07:00
Teknium 13c580422c fix(auth): carry pool-row lineage into the provider-block heal
With account-identity matching gone, the providers.<id> block consolidation
only fired on shared token material. A historical fork (same copied pool-row
id, profile rotated, both pairs diverged) then healed the pool row into root
but left root's providers.openai-codex block on the spent pair; root's next
load_pool() re-seeds its device_code row FROM that block and undid the heal.

_HealPass now records that a profile pool row matched root by copied id or
shared tokens and passes that verdict to _heal_forked_provider_block, which
accepts it as lineage proof. No account-identity guessing is restored; an
independent same-account grant (no id/token match) is still left alone.

Follow-up to simpolism's #106177.
2026-09-09 01:46:03 -07:00
simpolism 73f9de0c3a fix(auth): preserve independent same-account OAuth grants 2026-09-09 01:46:03 -07:00
nftpoetrist 9e0dc4319a fix(state): guard vacuum() and optimize_fts() against quarantined SessionDB handles
A quarantined/replaced/split-generation handle must never run a full-file rewrite or an FTS5
'optimize': both read damaged or foreign pages and commit the result back, turning contained,
diagnosable corruption into an amplified one. Same guard _execute_write applies to every write.

Salvaged from #102092 onto current main: the _try_wal_checkpoint half landed via #106315's
_quarantine_reason(), so only the two rewrite sites remain.
2026-09-09 13:17:19 +05:30
kshitijk4poor 26f4a674e0 fix(agent): a /steer row is human input for every user-turn predicate
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.

Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
2026-09-09 13:08:25 +05:30
kshitijk4poor bb2c961e9a fix(state): VACUUM is gated by the same quarantine rule as the checkpoints
Follow-up to #106315. vacuum() ran PRAGMA wal_checkpoint + VACUUM + wal_checkpoint(TRUNCATE) on
self._conn with no quarantine check; the only guard it inherited (optimize_fts raising
DeletedWalGenerationError) was swallowed by its own try/except and the rewrite proceeded on the
split-brain handle. Mutation on main: vacuum() returned 2 and rewrote pages after the write stop.
2026-09-09 13:06:34 +05:30
kshitijk4poor b530a482d8 fix(gateway): the ephemeral delete goes to the adapter that sent the final
Follow-up to #106316. send_final_ledgered resolved the live adapter internally and
_send_final_text resolved it a second time for _schedule_ephemeral_delete; a reconnect between
the two sent the delete to a transport that never owned result.message_id (the ownership rule
_final_delivery_adapter documents). The bracket now returns (result, adapter).

The queued lane carried the ledger identity through MessageEvent.message_id while the PR added
ledger_message_id for exactly that; it now uses the typed field, and the ledger read is
getattr-tolerant of duck-typed events (a missing attribute was swallowed as "ledger skipped").
2026-09-09 13:06:23 +05:30
kshitijk4poor 91433c8466 fix(loop): the turn-boundary export skips preflight-timeout envelopes and stops re-anchoring the persist index
Follow-up to #106312. _preflight_timeout_result carries the prior history without this turn's
user row (#7100); with a repeated prompt ("continue") the verbatim scan resolved to the
historical copy and exported it as this turn's proven boundary — the exact relabeling the export
exists to prevent. Nothing is exported for that envelope now.

The trailing `agent._persist_user_message_idx = idx` ran after finalize_turn had already flushed
the transcript, so it never influenced a persist and the next turn reset it: dead state, removed.
2026-09-09 12:55:43 +05:30
kshitijk4poor 7f3e0bb59a test(agent): the worker-start-failure test intercepts the callback worker again
`kwargs.get("target") is sched._run_callback` is always False (a bound method is a fresh object
per access), so the fake never returned Boom and the _dispatch failure branch went untested;
the test passed on the normal worker. Compare with == and assert the interception happened
(mutation: retiring the handle on start failure now fails the test).

The thread-count assertions sampled while per-fire workers were still live; quiesce every
handle with cancel(wait=) before sampling so the count is deterministic (AGENTS.md: timing tests
must not assume a quiet runner).

Follow-up to #106308.
2026-09-09 12:42:02 +05:30
kshitijk4poor 8077206073 fix(cli): keep the no-stored-model early return ahead of the route read
CI: tests/cli/test_cli_resume_command.py builds bare HermesCLI objects without .model; the
refactor read self.model before the stored-model check the contributor's code made first.
2026-09-09 12:41:07 +05:30
kshitijk4poor 32273b8118 refactor(cli): one stored_session_route for interactive and one-shot resume
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.

main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
2026-09-09 12:41:07 +05:30
liuhao1024 8aa773af89 fix(cli): restore stored session runtime and reopen ended rows on oneshot resume
Review fixes (#105957):

- A resumed one-shot ignored the session's stored model/provider runtime:
  _resolve_model_and_provider()/resolve_runtime_provider() ran before
  _load_resume_target(), which only loaded the session id + transcript, so an
  ambient config (e.g. openrouter/ambient-model) served the resumed transcript
  instead of the stored route (custom:stored/stored-model). The stored runtime
  is now applied before runtime resolution, with the same contract as the
  interactive _restore_session_model(): stored model/provider/base_url/api_mode
  replace the ambient choice unless --model was passed explicitly, and a
  changed provider drops the ambient api_key so resolution re-fetches
  credentials for the restored endpoint.

- Passing the resumed id to AIAgent did not reopen the already-ended session
  row: end_session() only writes rows whose ended_at is null and the
  existing-row upsert never clears the end fields, so the resumed turn was
  recorded under a session that stayed closed and its new lifecycle boundary
  was lost. _load_resume_target() now reopens the row (best effort), same as
  the interactive resume does before continuing.
2026-09-09 12:41:07 +05:30
liuhao1024 5ff6cb0edb fix(cli): keep the resolved session id when a resumed oneshot session is empty
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
2026-09-09 12:41:07 +05:30
liuhao1024 86d606ca34 fix(cli): honor --resume in one-shot mode (#105892)
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).

Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.
2026-09-09 12:41:07 +05:30
kshitijk4poor c402bcf1d8 fix(profiles): a stored <root>/profiles/<name> home names its profile even when the root carries no markers
CI: tests/test_tui_gateway_server.py::test_ensure_session_db_row_stamps_profile_name used a bare tmp
root; profile_name_for_home fell through to None and the row was stamped default. The stored home
is authoritative (its owner resolved it), so the profiles/<name> shape is sufficient.
2026-09-09 12:39:39 +05:30
kshitijk4poor a39fedff34 fix(tui): a real named profile "hermes" is not swallowed by the legacy-basename alias
"hermes" matches the profile-id regex, so canonicalising it unconditionally at the RPC
boundary misrouted a genuine <root>/profiles/hermes to the default profile. Alias only
when no such named profile exists; ".hermes" can never be a real id and stays aliased.

Also: profile_name_for_home collapses its duplicated pre/post-resolve block into one
loop over (path, resolved path) and drops the bare "parent named profiles" fallback
that bypassed named_profile_home's root check; _profile_home goes back to main's
single resolve() comparison; the symlink-loop assertion in the target-unavailable
test is no longer wrapped in a try/except that could silently skip it.
2026-09-09 12:39:39 +05:30
joaomarcos 81ff1a4022 test(tui): add coverage for custom default roots, real session db stamping, and sibling isolation 2026-09-09 12:39:39 +05:30
joaomarcos e1b0ee6f03 fix(tui): fail closed on unavailable profile targets matching custom root basenames 2026-09-09 12:39:39 +05:30
joaomarcos 96128b0673 fix(tui): preserve names for custom profile homes 2026-09-09 12:39:39 +05:30
joaomarcos 298893df66 fix(tui): resolve default profile session names 2026-09-09 12:39:39 +05:30
kshitijk4poor 7dc796463d fix(agent): a persisted /steer row survives the next prompt's alternation repair; typed for history
Both steer sites now build the row through one helper, prompt_builder.steer_user_row:
a role:user row with display_kind="steer" and no leading blank lines. The alternation
repair (_merge_consecutive_users) skips a steer-typed prev row, so a run that ended
right after a steered batch (Ctrl-C, interrupt) does not get the next real prompt
merged INTO the already-persisted steer row — which would have rewritten it in place
and re-broken live≠replay parity, the exact class this PR fixes.

TUI/desktop history projects the steer row as the user's own words instead of the
model-facing marker wrapper; 'steer' joins the display_kind union. The compression
anchor scan keeps its tool-row branch for transcripts persisted before this change and
its docstring says so.
2026-09-09 12:21:28 +05:30