Commit Graph

24482 Commits

Author SHA1 Message Date
Dhanesh Purohit ca28a69ada fix(state): apply macOS write barriers on every state.db repair connection
state.db corrupted twice in two days with the torn-b-tree signature —
repeated "2nd reference to page", "Rowid out of order", and long runs of
"never used" pages in messages (rootpage 5) and idx_messages_session.

macOS fsync() guarantees neither data-on-platter nor write ordering, which
_enforce_macos_synchronous_full already documents: a rewrite interrupted by
process or OS termination leaves half-written b-tree pages. The mitigation
is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was
applied only through apply_wal_with_fallback(). The repair path opened
state.db with a bare sqlite3.connect() six times and then ran REINDEX,
VACUUM and writable_schema surgery through it — the operations that rewrite
nearly every page of the file — with no barrier at all.

- _connect_repair_durable() routes every repair/probe connection through the
  barriers. Applying them is best-effort by necessity: SQLite loads the
  schema before any statement, so on a malformed schema even
  PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is
  precisely this helper's input. _reapply_durability_barriers() retakes them
  before REINDEX and VACUUM, once the schema parses and they can stick.
- verify_state_db_integrity() adds the proactive check that was missing.
  Repair only ever ran reactively, after a caller already hit a malformed
  error, so a database torn in pages no query happened to touch stayed live
  and kept accepting writes. On 2026-08-19 that gap was 11 hours across two
  restarts that both reported a clean start. Size-aware: degrades to an O(1)
  probe above 2 GiB rather than pegging a CPU at startup.

Also restores two fixes lost when `hermes update` reset the tree to
origin/main before they were committed:

- _db_fingerprint keys the repair ledger on dev+inode+size instead of
  size+mtime_ns. The old form was justified as "stable for a file nothing
  can successfully write to"; that premise is false, because on FTS
  corruption this module deliberately keeps canonical writes enabled with
  FTS detached. mtime churned on every write, so each pass re-keyed the
  ledger and reset the counter to 1 — the cap could never be reached and the
  damaging surgery could retry forever.
- _live_writer_holds_db() refuses surgery while another connection holds the
  database. The cross-process lock only serialises repairers against each
  other; it says nothing about the gateway, Desktop or a CLI. Rewriting
  b-tree pages under a concurrent writer is what spread the 2026-08-18/19
  damage out of the FTS shadow tables and into the canonical ones. Fails
  open, so it cannot strand the self-heal path it protects.

The guard's own tests built a two-table toy schema, so every repair aborted
on "no such table: sessions" before reaching the guards under test — the
assertions were passing over a code path that never ran. They now build
through a real SessionDB.

Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure.
Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>
2026-08-22 04:02:29 +05:30
AlexFucuson9 f5adeed3dc fix(cli): guard empty message text in _display_resumed_history
text.splitlines() returns [] for empty strings. Accessing msg_lines[0]
then raises IndexError, making session resume crash when the session
contains a message with empty or whitespace-only text (e.g. reasoning-only
turns, tool-only assistant messages).

Guard with `or [""]` in all three branches (user, assistant_last,
regular assistant) so an empty message renders as a blank line.

Fixes #59265

Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
2026-08-22 03:59:28 +05:30
kshitijk4poor 2ea5287da0 fix(state): only flag uninspectable Hermes processes as holders
The cmdline fallback was matching every system daemon with an
unreadable fd table (init, systemd-journald, dockerd, etc.), causing
FTS rebuilds to be skipped on every Linux system. Add _looks_like_hermes
filter so only processes whose cmdline contains Hermes markers are
flagged — matching @jackulau's suggestion of 'uninspectable AND
identifiable as another Hermes process.'
2026-08-22 03:56:13 +05:30
kshitijk4poor 1c59daaace fix(state): use /proc readlinks + cmdline fallback for holder detection
Address review feedback from @jackulau on PR #90871:

1. psutil.open_files() silently drops '(deleted)' WAL sidecar entries
   on Linux because isfile_strict() stats the literal path including
   the suffix and fails. Switch to direct /proc/<pid>/fd readlinks
   which preserve the '(deleted)' suffix so _canonical can match.

2. psutil.process_iter() converts AccessDenied to None, which
   or-() skips silently — the fail-closed branch never runs. For the
   root-gateway vs user-desktop topology in the issue, the fd table is
   unreadable but /proc/<pid>/cmdline is world-readable. Add a cmdline
   fallback that flags uninspectable processes.

Also keep the psutil path for macOS/BSD (no '(deleted)' convention).
2026-08-22 03:56:13 +05:30
kshitijk4poor c45e2b19c3 fix(state): guard gateway FTS rebuild + comment early flag-set
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.
2026-08-22 03:56:13 +05:30
fangliquanflq fc72d6c716 fix(state): defer FTS rebuild under foreign WAL holders 2026-08-22 03:56:13 +05:30
Teknium 334bcbac93 fix(desktop): error card honors the classifier's retry verdict + failing-session identity (review feedback)
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
  verdict) next to failure_reason; error_surface prefers it and only falls
  back to the reason set for older results. Fallback set corrected to match
  classify_api_error (auth, format_error, billing_unverified now
  non-retryable).
- The descriptor carries the failing session's provider/model captured at
  classification time; Copy error details prefers them over the foreground
  composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
  the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
  requests/aiohttp so other adapter SDKs don't misclassify as gateway.
2026-08-21 15:24:03 -07:00
Teknium 50f1e414bc polish(desktop): rename error-card action to 'Copy error details'
'Copy diagnostics' was dev-speak; match the familiar OS-error phrasing.
All five locales + docs updated.
2026-08-21 15:24:03 -07:00
Teknium 3903428a72 Revert "feat(desktop): error card offers Nous support link on Portal-auth sessions"
This reverts commit 31872bfcf555cedb2501122a75e29328c0e90e80.
2026-08-21 15:24:03 -07:00
Teknium e3d46bb5fb feat(desktop): error card offers Nous support link on Portal-auth sessions
Sessions running on provider 'nous' get a 'Nous support' action on the
failed-turn card, opening the portal help hub
(https://portal.nousresearch.com/help — docs, Discord, GitHub) in the
external browser. All five locales + docs updated.
2026-08-21 15:24:03 -07:00
Teknium 892790f980 fix(desktop): error card renders router-free threads without crashing
useNavigate() throws outside a <Router>; streaming.test.tsx renders the
thread bare. Move the Settings deep-link into a SwitchProviderAction child
gated on useInRouterContext(), which is safe in any tree.
2026-08-21 15:24:03 -07:00
Teknium 98f6fc549a feat(desktop): failed turns name the failing layer with recovery actions
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.

Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
2026-08-21 15:24:03 -07:00
Teknium e26d91dc11 feat(bot-mode): message_agent tool — structured, Bot-Chat-only agent-to-agent DMs
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.

Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).

The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
2026-08-21 15:23:51 -07:00
Teknium 04acfb9673 fix: remove function-level 'import time as _time' that shadowed the module import
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
2026-08-21 15:23:49 -07:00
Teknium 70151dd549 feat(update): updates.parked_branch_strategy gates the in-place merge; switch stays the default
Adapts the in-place branch update from PR #89507 (@willfrombr) onto the
switch-by-default behavior: the deterministic switch path remains the
default so non-interactive updates (desktop, gateway, cron) never dead-end
on a merge conflict, and deliberate custom-branch users opt in with
updates.parked_branch_strategy: update_in_place. --switch-branch overrides
the in-place strategy for one run (deep feature branches that must not
accumulate update merge commits). Docs + config comments + tests cover
all three routes.

Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
2026-08-21 15:23:49 -07:00
Willian Santos 4fad27a101 feat(update): --switch-branch opts an unmerged branch out of the in-place merge
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.

--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.

Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.

Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Willian Fernandes Santos 91096bb2f0 feat(update): update branches carrying unmerged commits in place instead of skipping
The parked-branch guard (8ce8ffd429) distinguishes checkouts by what the
branch carries, then treats both non-clean cases the same: a stale
fully-merged leftover is switched back to the target (correct), but a
branch with unmerged commits — a branch someone is actually working on —
gets CODE UPDATE SKIPPED and exit 1. For anyone running a maintained
custom branch on top of main, every update now refuses, and the guidance
('checkout main') abandons their branch.

The guard's own reason codes already separate the cases, so use them:

- fully merged      -> switch back to the target (unchanged)
- unmerged:N        -> update the branch IN PLACE: fetch, then bring
                       origin/<target> into the checkout. Fast-forward
                       when possible; on divergence, a true merge behind
                       a pre-update safety tag, stopping cleanly on
                       conflict. The checkout never moves; local commits
                       survive; the running code advances.
- dirty/unverifiable/opted out -> skip loudly (unchanged)

The post-pull success gate learns that an in-place update legitimately
ends on a non-target branch: origin/<target> was merged INTO the checkout,
so refusing to claim success there would fail every update that did
exactly the right thing.

Guard tests updated: the unmerged case now asserts the in-place outcome —
target code arrives (b.txt from c3), the branch's own commit survives, and
HEAD never moves. 18/18 guard tests, 20/20 with the diverged-update suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Teknium bbbc50acc2 fix: hermes update no longer strands non-interactive updates on a parked branch with unmerged commits
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.

Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
2026-08-21 15:23:49 -07:00
Teknium 2dcb623956 chore: map salvage contributor emails 2026-08-21 15:02:29 -07:00
Teknium 42dd219f46 test: trim salvage of #65076 to a lean regression set
Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
2026-08-21 15:02:29 -07:00
Vinay Shah 5cb7b521cf refactor(bedrock): make resolve_bedrock_runtime_region the single region chokepoint
Follow-up structural pass on the review fix:

- Runtime provider, auxiliary resolution, model validation
  (hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
  and the Mantle URL/SigV4 fallbacks all resolve their region through
  resolve_bedrock_runtime_region() — one canonical implementation of the
  config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
  initializing client_kwargs unconditionally at the top of the else
  branch; the Mantle kwargs hook is a documented no-op for non-Mantle
  base URLs.
2026-08-21 15:02:29 -07:00
Vinay Shah 41ca67c5b1 fix(bedrock): align auxiliary region resolution with runtime + document Mantle route
Address review feedback on #65076:

- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
  config-first region resolution (bedrock.region in config.yaml, then
  AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
  runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
  branch) to the new helper. Previously it derived its region with bare
  resolve_bedrock_region() (env-first), so when config.yaml pinned
  bedrock.region to a different region than the ambient AWS env, auxiliary
  calls (compression, memory, vision) left the primary runtime's region.
  Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
  Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
  for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
  uses the OpenAI-compatible endpoint, which the Mantle route made stale.
  Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
  Converse), the Mantle auth model (bearer token or SigV4), and add the
  GPT-5.5/5.6 model IDs to the models table.
2026-08-21 15:02:29 -07:00
Vinay Shah 16476fad10 feat(bedrock): add OpenAI GPT-5.6 family (Sol/Terra/Luna) to Mantle Responses routing
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:

- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
  so runtime resolution, auxiliary calls, and MoA slots all take the
  SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
  Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
  BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
  additions do not require test surgery; add routing, picker, and
  context-length coverage for the 5.6 family.

Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
2026-08-21 15:02:29 -07:00
Nathaniel Branscum e57d55fc7a fix(moa): keep Bedrock slots on provider runtime
Preserve the Bedrock provider identity for MoA reference and aggregator slots so Bedrock OpenAI Responses models use the aws_sdk/SigV4 runtime instead of being downgraded to a generic custom endpoint. Add regression coverage for Bedrock GPT-5.5 MoA slots.
2026-08-21 15:02:29 -07:00
Nathaniel Branscum e5b96fcb10 feat(bedrock): support OpenAI Responses models
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
2026-08-21 15:02:29 -07:00
kshitijk4poor bad2ed866c fix(telegram): honor the direct-messages-topic alias in the fresh-final gate
prefers_fresh_final_streaming read only the raw direct_messages_topic_id
key; the adapter's canonical accessor _metadata_direct_messages_topic_id
also accepts the documented telegram_direct_messages_topic_id alias
(treated as equivalent in gateway/delivery.py), so an alias-only lane
would still flatten tables. Route the gate through the accessor and pin
the alias with a regression (mutation-checked: raw-key gate fails it).
Also reshape the happy-path endpoint assertion into the actual invariant
(sendRichMessage present, no rich draft frames) instead of a frozen call
list. Surfaced during review of PR #91436.
2026-08-22 03:31:54 +05:30
HexLab98 a5ca9c06e6 test(telegram): cover DM-topic table streaming after draft degradation
Pins integer topic routing on send_draft, a successful topic stream
that finalizes through sendRichMessage, and the reporter path where
sendMessageDraft and in-place rich edits both fail — the persistent
payload must still be the raw pipe table, not convert_table_to_bullets.
2026-08-22 03:31:54 +05:30
HexLab98 194729c95f fix(telegram): keep DM-topic tables on sendRichMessage when drafts degrade
#91241 stopped root-DM tables collapsing to bullets by keeping native
draft transport when rich_drafts is off. Private Telegram topics still
reject sendMessageDraft (string thread ids, forum-style thread fields),
so the stream consumer falls back to edit-in-place. Telegram then
rejects a rich edit of that plain MarkdownV2 preview and format_message
permanently rewrites pipe tables into bullet lists — the remaining
report after that merge.

Route drafts through the same integer topic kwargs as send(), and on
that degraded topic path prefer a fresh sendRichMessage (then delete
the preview) instead of the table-to-bullets formatter.
2026-08-22 03:31:54 +05:30
Teknium 30f9955a44 fix(zai): GLM-5.3 low/medium reasoning effort reaches the wire instead of clamping to high
GLM-5.3 accepts a graded low/medium/high/max reasoning_effort scale
(verified live in #91789: monotonic reasoning-token scaling, no 400s),
but the effort mapper reused GLM-5.2's two-level vocabulary, silently
rewriting low/medium to high. Adds GLM53_EFFORTS/GLM53_OVERRIDES and a
per-model vocabulary pick in the zai plugin; 5.2 keeps its high/max
clamp. Closes #91789. Also covers the gap noted when closing #86947
(credit @santhanakrishnan-d and @terje1965 for the graded-scale finding).
2026-08-21 14:58:37 -07:00
kshitijk4poor 3841910cee fix(telegram): widen cancellation-shielded stop to sibling paths
The network-error reconnect path (PR #91524) was the only site converted
from asyncio.wait_for to _await_with_thread_deadline.  The same
cancellation-shielding vulnerability exists at two more updater.stop()
sites:

- Conflict-retry path: asyncio.wait_for could hang forever if PTB/AnyIO
  cleanup swallowed CancelledError, stalling the conflict-retry ladder.
  Now uses _await_with_thread_deadline and escalates to fatal on timeout
  (same reasoning: cannot safely reuse an Updater whose lifecycle lock
  may still be held).

- Conflict-exhausted fatal path: asyncio.wait_for could hang before the
  fatal notification fired.  Now uses _await_with_thread_deadline; the
  timeout handler already proceeds to fatal notify, so no behavior change
  beyond the deadline mechanism.

All three asyncio.wait_for(updater.stop()) sites now use the
thread-deadline helper consistently.
2026-08-22 03:20:48 +05:30
Good Chang 9e36774d77 fix(telegram): rebuild after cancellation-shielded stop
Use the existing wall-clock deadline helper for updater.stop() during network recovery. If PTB cleanup remains cancellation-shielded past the deadline, escalate to retryable fatal recovery so the runner builds a fresh adapter instead of calling start_polling() while the old Updater may still hold its lifecycle lock.

Add regression coverage with stop() swallowing cancellation while holding the same lock start_polling() needs, and verify the old Updater is never reused.
2026-08-22 03:20:48 +05:30
Jack Lau dd03471858 fix(cron): nudge review of escaped-run failures too
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.

The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.

Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.

Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.

Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.

Fixes #88655
2026-08-22 03:19:08 +05:30
Sam Campbell 2fb1e62b2e fix(gateway): multiplex refusal must exit EX_CONFIG (78), not 1
`_guard_named_profile_under_multiplexer` correctly refuses a named-profile
gateway while the default gateway is multiplexing — starting a second one would
double-bind that profile's platforms. The refusal is right; its exit code was
not.

The refusal is decided entirely by configuration (`multiplex_profiles` plus the
allowlist), so it is permanent: no number of retries can change the answer.
Exiting 1 made it look transient to a service manager.

That matters because this module generates the systemd unit, and the template
pairs `Restart=always` / `RestartSec=5` with `StartLimitIntervalSec=0` — it
deliberately trades systemd's generic start-rate limiter for the specific
`RestartPreventExitStatus=GATEWAY_FATAL_CONFIG_EXIT_CODE` backstop declared
three lines below it. Returning 1 left that backstop unarmed with the limiter
already disabled, so a correct, permanent refusal became an unbounded restart
loop. Observed on a host running `multiplex_profiles: true` with a leftover
per-profile unit: 136 refusals in ~13 minutes, stopped only by hand.

`GATEWAY_FATAL_CONFIG_EXIT_CODE` (78, EX_CONFIG) is this codebase's existing
answer for exactly this case — `gateway/restart.py` documents it as the fatal
configuration error that the s6 finish script translates into 125 "permanent
failure" (#51228). This adopts that contract rather than inventing one, so the
fix also works on s6 hosts, not just systemd.

After: one refusal, `status=78/CONFIG`, `NRestarts=0`, unit settles in `failed`.

Also strengthens the two guard tests. They asserted
`pytest.raises(SystemExit, match="1")`, but `match=` is a regex search over
`str(exc)`, so it passed for 1, 21, 100 and 111 alike — it read like an exit-code
assertion while pinning nothing. They now assert
`excinfo.value.code == GATEWAY_FATAL_CONFIG_EXIT_CODE`. The exit code is the
contract here: it is the only thing that tells a supervisor the failure is
permanent.
2026-08-22 03:10:32 +05:30
kshitijk4poor ac64f8a7e7 chore: AUTHOR_MAP — add samtcam@gmail.com → samclams
For PR #91806 salvage (multiplex refusal exit code fix).
2026-08-22 03:10:32 +05:30
Teknium bc8f49618c chore: map contributor email for S-Claw 2026-08-21 14:39:02 -07:00
xthezealot d422f7103e fix(backup): don't hang forever on locked SQLite sources
hermes backup freezes mid-archive when a .db file under HERMES_HOME is
locked by another process — e.g. a live Chromium profile database held
with an exclusive lock by a running browser. sqlite3.Connection.backup()
retries SQLITE_BUSY indefinitely and never honors the connection's busy
timeout, while a plain statement on the same source fails cleanly after
~5s with "database is locked".

Fixes:
- Probe the source with a cheap read before snapshotting, so a locked
  database fails fast instead of hanging the whole backup.
- Add a watchdog that interrupts the source connection after 15 minutes
  as a last resort for pathological cases.
- Exclude browser-profiles/ from full backups: the CDP browser profile is
  live, regenerable (cache + re-login), and unsafe to snapshot while
  running. On a real install this cut the backup from 28,396 files /
  1.1 GB to ~4,000 files / 548 MB, completing in ~33s instead of hanging.

The pre-update automatic backup shares this code path and was equally
at risk.

Adds a regression test that holds an EXCLUSIVE transaction in a separate
process and asserts _safe_copy_db returns False in bounded time.
2026-08-21 14:39:02 -07:00
Simon f9849c43a2 fix(backup): don't nest state-snapshots/ into full backups
`hermes backup` already skips `backups/` so a full zip never re-ships
earlier pre-update zips. `state-snapshots/` (written by `hermes backup
--quick`, `/snapshot create`, and the pre-update safety net) has the same
shape — every retained snapshot holds its own copy of state.db — but was
not in `_EXCLUDED_DIRS`, so a full backup shipped the DB once per
retained snapshot on top of the live one.

Two places hit this in practice:

- `hermes update` in `full` mode takes the quick snapshot *before* the
  full zip, so the pre-update zip always nests the snapshot it just made
  (state.db twice in every pre-update-*.zip).
- Any recurring `hermes backup --quick` (default keep=20) makes a daily
  `hermes backup` grow by roughly one compressed state.db per retained
  snapshot; a 750 MB state.db with two snapshots on disk pushed a daily
  zip from 1.8 GB to 2.3 GB.

Add `_QUICK_SNAPSHOTS_DIR` to `_EXCLUDED_DIRS` (moving the constant up
next to the exclusion rules so there is one source of truth). Both walk
sites and `_should_exclude` share the set, so `hermes backup`, the
pre-update zip and the auto-backup path all pick it up. Restoring
snapshots after a machine move was never the point of the full backup —
`profiles.py` already excludes `state-snapshots/` from `--clone-all` for
the same reason.

Tests: unit case next to the `backups/` one, plus two end-to-end cases
that use the real `create_quick_snapshot` producer and assert the zip
carries exactly one state.db (full backup and pre-update-order).
2026-08-21 14:39:02 -07:00
Teknium bd93a5f316 feat(models): free models show star + -100% in the model picker discount column
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
2026-08-21 14:38:41 -07:00
Teknium 098a7acd4a add openclaww@gmail.com to contributors 2026-08-21 14:38:32 -07:00
Teknium 907da145b6 feat(models): glm-5.3 replaces glm-5.1 in the OpenRouter and Nous Portal catalogs
Follow-up to the salvaged GLM-5.3 support commit: drop z-ai/glm-5.1 from
both curated lists per Teknium's direction (glm-5.2 keeps the 'default'
tag), and regenerate the docs manifest. glm-5.1 remains available via
live discovery and on out-of-scope surfaces (zai plugin, setup defaults,
opencode-go) — named leftovers, not silently swept.
2026-08-21 14:38:32 -07:00
openclaw 01d8562fce fix(zai): add GLM-5.3 support — 1M context window, model lists, reasoning_effort
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.

- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
  context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
  2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
  by the endpoint, HTTP 200)
2026-08-21 14:38:32 -07:00
Teknium 1bf8bd2c7d feat(models): 'ox alpha' now finds x-preview-f-free in every model picker
The OpenCode Zen wire slug for the Ox Alpha stealth model is opaque
(x-preview-f-free); users searching the picker for 'ox' or 'ox-alpha'
found nothing. Adds the search alias across all four synced alias
tables (CLI, desktop, web, TUI) plus tests. Wire id is unchanged and
still what renders and gets sent to the provider, matching the k3 →
kimi-k3 precedent. No canonical-dedup collision with opencode-go's
keyed ox-alpha-free slug.
2026-08-21 14:38:19 -07:00
ethernet 76f6ba3706 feat(nix): give Home Manager a programs module and the desktop app
Home Manager separates an installation from a daemon. This module put
both under `services.hermes-agent`, and `installPackage` added a program
to the PATH from a service module.

`programs.hermes-agent` now installs the command line application and
the desktop application. `services.hermes-agent` keeps the state, the
configuration and the daemons, and stays the authority: the new module
reads `hermesHome` and the backend address from it. A person can enable
one without the other, which is a machine with an application and no
gateway, or a headless gateway with no display.

The desktop application needs this split to work correctly. A launcher
that starts from the desktop menu reads no shell profile, thus the
HERMES_HOME that `home.sessionVariables` exports reaches an interactive
shell only. Home Manager writes `systemd.user.sessionVariables` to
environment.d, and this module puts no HERMES_HOME there, because that
file applies to each user unit. The application then opens ~/.hermes
while the services use `hermesHome`, and the person sees no sessions and
no keys. Thus the launcher carries the value itself, through a new
`extraEnv` argument on the desktop package.

The application also gets the Nix agent package, with
HERMES_DESKTOP_HERMES. The usual distribution of the Electron
application carries its own Hermes runtime and downloads more at the
first start. `hermesDesktop` is a passthru of the agent and pins
`finalAttrs.finalPackage`, so an override of `extraPythonPackages` or
`extraDependencyGroups` reaches both. One machine thus has one runtime.

`backend.sessionTokenFile` connects the application to the backend of
the service. Without it the module runs `hermes serve` and the
application starts a backend of its own, which gives two backends on one
HERMES_HOME. The backend reads the file into
HERMES_DASHBOARD_SESSION_TOKEN. The launcher reads the same file into
HERMES_DESKTOP_REMOTE_TOKEN, beside a HERMES_DESKTOP_REMOTE_URL that
names the address of the service.

Measurements against a live `hermes serve` on loopback show why that
shape is the correct one:

- `_resolve_session_token()` reads HERMES_DASHBOARD_SESSION_TOKEN, and
  `_has_valid_session_token` accepts that value as a Bearer credential.
  A request without it gets 401, and a request with the wrong value
  gets 401.
- The /api/ws socket accepts a query parameter only. A header gets 403,
  and `?token=` connects. Hermes Desktop builds exactly that URL, in
  `apps/desktop/electron/connection-config.ts`. Thus a test of the HTTP
  leg alone is a false positive.
- `resolveDesktopRemoteRoute` throws when the URL is set and the token
  is not. Thus the two variables travel together or not at all.

The token enters no Nix store path. `makeWrapper --set` and a systemd
`Environment=` value both write a literal into the store, which all
users can read. Thus each side reads the file at start time. The
launcher does it through a new `extraRun` argument on the desktop
package, and the backend through the launcher script that
`backend.waitFor` already uses. launchd has no EnvironmentFile, so a
script is the one shape that works on Linux and on Darwin.
`backendArgv` gives the plain argv only when nothing must run before
the backend.

`services.hermes-agent.installPackage` is removed. It defaulted to true,
so a person who never named it still got the command line. A silent
removal thus gives them a machine with no `hermes` and no message. The
module refuses a configuration that sets it, and the text names the
exact replacement for the value they gave.

Checks:

- the launcher carries HERMES_HOME
- the launcher reports HERMES_MANAGED only when the services own the
  configuration, because no activation writes a marker without them
- the launcher pins the agent package that `programs.enable` installs
- the launcher names the backend of the service, and gives a token
  beside the URL
- the backend reads the session token
- each side reads the file at start time, and the token is no `--set`
  value
- `programs.enable` alone starts no service
- `installPackage` is refused, with a message that names the
  replacement, and its absence evaluates

Each check reads the wrapper of the real package, and not an option
value. Each one was tested with a mutation that breaks the behavior it
asserts.
2026-08-21 17:00:30 -04:00
Minsang Lee 0287dfb0c2 fix(bot-mode): a bot row opens the conversation you were last having
Clicking a bot in the roster always reopened its pinned canonical Bot Chat.
Start a new conversation with bot A, click bot B, click back to A — the new
conversation was gone, replaced by the pinned transcript. A bot row is a
workspace entry point, so it has to land on the live conversation.

Two independent causes, both fixed here:

1. The pin overrode newer work.
   `openBotCanonicalChat` opened the pin unconditionally. It now prefers the
   bot's freshest VISIBLE session — but only AFTER `profiles.list` has
   verified through `preferred_session` that the pin is alive and is a real
   canonical Bot Chat. That ordering matters: with a dead or unverified pin,
   adopting the profile's latest row would claim an unrelated user
   conversation as the bot's chat, and the hide sweep would then hide it.
   The existing "no pin" / "dead pin" safety tests cover exactly that and
   still pass. The pin keeps owning plumbing (creation, hide sweep, DM
   delivery); it just stops shadowing newer conversations.

   Guards on the candidate (`newerVisibleBotChat`): the canonical chat can
   never shadow itself, an empty draft never displaces a real conversation,
   and a gateway that omits `message_count` is treated as real history
   rather than discarded.

2. The workspace did not follow the bot.
   The three `host.openSession` calls on the bot path relied on the SDK
   default `keepAllProfilesScope: true`, so `$activeGatewayProfile` stayed on
   whatever profile was active before the click. Sessions created afterwards
   were then filed under the previous bot's profile — measured: four new
   chats started from three different bots all persisted into one profile's
   state.db. Clicking a bot IS a profile switch, so these pass `false`.

Note on the call shape: `previewSession` is `bot.preferred_session || last`,
so on a pinned bot it resolves to the PIN (preview identity must match click
identity). Feeding that as the "newer" candidate makes the whole preference
dead code — it always sees the pin and short-circuits on "same id". The
freshest visible session therefore arrives as its own argument. The first
attempt at this fix had that bug and passed its tests, which is why
`bot-row-opens-latest.test.mjs` mirrors the production call site argument for
argument rather than constructing a convenient one.

Tests: 362 pass (was 348). Each new guard was verified by sabotage — reverting
any one of the three behaviours above makes the suite fail (1, 3, and 1 tests
respectively), so none of them is a test that passes either way.
2026-08-21 13:42:54 -07:00
Teknium d9d967e07a fix(cli): Linux hermes.desktop entry launches instead of silently dying on system python (#90292)
resolve_exec_command wrote the repo hermes script (env-python shebang)
straight into Exec=; spawned by the DE that shebang escapes the venv and
dies on the first import, invisibly (Terminal=false, entry rewritten
every launch). A python-script launcher whose shebang points outside the
running interpreter's env now gets Exec={sys.executable} {script} desktop;
native binaries, bash wrappers, and venv-shebang scripts are untouched.
2026-08-21 13:41:39 -07:00
unsupportedpastels ac8dff4fbc fix(compression): auto-raise Daybreak Codex threshold 2026-08-21 13:19:07 -07:00
unsupportedpastels f8e5949f61 fix(model_metadata): add Daybreak Codex 900K context 2026-08-21 13:19:07 -07:00
Teknium 67af79d7e1 feat(models): stealth/ox-alpha free model in the Nous Portal catalog
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.

Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
2026-08-21 13:06:42 -07:00
ethernet 9815319d5f refactor(desktop): derive the tab hover close button from the close verb
PaneTab gated its hover close button on two independent inputs: the
onClose verb, and a showCloseButton prop that TreeGroup fed from a
showCloseButton flag on the pane contribution. The middle-click and
Meta-click gestures read only onClose. A tab could therefore close on a
pointer gesture and advertise no control for it.

The flag had no user that hideOnly did not already cover. Both setters
also set hideOnly: true, which removes every close gesture:

- the sessions pane (app/contrib/controller.tsx),
- the Bots pane (plugins/hermes-bots/plugin.js).

The flag was an opt-out marker with no reachable effect, so this change
deletes it instead of teaching it to track the gestures. onClose alone
now decides both shapes. A tab that closes shows the button. A tab
without the verb shows nothing. To make a tab uncloseable, give it no
close verb.

hideOnly and uncloseable keep their meaning. They gate the verb, and
both shapes follow the verb together.

The DialogContent and SheetContent prop of the same name is a different
prop and stays. It has no close verb to derive from, and one caller
changes it while the dialog is open.

Tests: the new tab-close-affordance test renders the real TreeGroup and
asserts that button presence equals middle-click closure. It covers
hideOnly chrome, a plain side pane, the uncloseable workspace, and a
session tile. It reads closure from the layout tree, not from a spy, so
a wired-up mock cannot pass it. A regression that hides the button on a
closeable tab fails two of the four cases. The compiler rejects the
deleted prop, so the test carries no fixture for it. The pane-tab unit
test moves off the deleted prop.

Verified with the full apps/desktop vitest suite, npm run typecheck, and
npm run lint. Two electron process-spawn tests fail on this machine.
They also fail on a clean tree, and they do not touch the pane shell.
2026-08-21 16:04:25 -04:00
Teknium 1575116629 fix(update): pre-update snapshots now cover every profile, not just the invoking one (#66140)
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.

- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
  snapshot set, per-file 1GiB cap, and keep policy as the invoking
  profile (no partial tier, no new restore-coherence class), each into
  the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
  runs the #34600 cron-loss safety net per profile against its OWN
  snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
  profile's (best-effort, receipt-recorded); post-update cron restore
  extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
  behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
  jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
2026-08-21 13:01:35 -07:00