Commit Graph

22714 Commits

Author SHA1 Message Date
briandevans f7de2ca416 fix(desktop): order project-tree lanes by recency in the overview, not alphabetically
_build_repos emptied each lane's sessions array for the overview (hydrate=False)
payload BEFORE _sort_lanes ran. _lane_sort_key derives a lane's activity from
max(_session_time(s) for s in group["sessions"]), so with the rows already gone
every non-trunk lane scored activity=0.0 and the sort key
(is_trunk, is_kanban, -activity, label) collapsed to alphabetical-by-label. The
documented intent — branches and linked worktrees sort by most-recent activity,
then label — was silently defeated on the projects.tree RPC that feeds the
desktop sidebar overview, while the drill-in path (hydrate=True) kept the rows
and sorted correctly. Any repo with two or more non-trunk lanes showed a
different order in the overview than when opened.

Move the session-clearing to after _sort_lanes/_disambiguate_labels so the sort
reads real recency. Lane counts are still captured before clearing, so
sessionCount and the slim overview payload are unchanged — only the order is
fixed, and the overview now matches the drill-in.
2026-08-15 00:33:11 -07:00
Tranquil-Flow 0d07fe63f9 fix(projects): use _branch_lane_id for non-git folders to prevent duplicate lanes (#53329)
_place_by_heuristic used the raw path as the lane key for non-git
project folders, while the desktop overlay independently computed
::branch::main for the same session (since git_branch was null).
The ID mismatch caused duplicate lanes — one from the backend with
the folder name, one from the overlay labeled 'main'.

Use _branch_lane_id(path, DEFAULT_BRANCH_LABEL) so the backend's
lane key matches the overlay's expected ::branch::main scheme,
eliminating the duplicate lane.
2026-08-15 00:33:11 -07:00
Teknium 2fecf392a2 chore: map contributor email for Hangzian 2026-08-15 00:33:01 -07:00
Teknium 94ce8396e8 fix(sessions): release active-session leases against their acquisition registry
A gateway active-session lease is acquired against the root HERMES_HOME,
but release_active_session()/transfer_active_session() re-resolved the
registry path from the *current* HERMES_HOME. Under native multiplex a
routed turn runs agent cleanup inside _profile_runtime_scope, so the
release looked under the named profile while the root entry stayed
alive — after max_concurrent_sessions routed turns every new session was
rejected with 'Hermes is at the active session limit' (#85431).

Pin state/lock paths on the lease at acquisition time and prefer them on
release and transfer. Fixes #85431.
2026-08-15 00:33:01 -07:00
rainbowgits aba4934274 fix(agent): omit unsupported metadata on Relay scope.pop
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
2026-08-15 00:33:01 -07:00
Adolanium b9672ea24e fix(pets): remove the non-PNG base draft after hardening
generate_base_drafts hardens every base draft to a transparent PNG with
_harden_transparency. When the provider returned a non-PNG file (webp,
jpg, or gif), the hardened PNG is saved under a new path and the original
draft is left in cache/images. Nothing prunes that directory outside the
gateway housekeeping loop, so a CLI or desktop draft round leaks one
original per non-PNG draft.

Remove the original after a successful hardening when the output path
differs from the input. PNG inputs (including mixed-case suffixes like
.PNG) are hardened in place so a case-insensitive filesystem cannot
treat with_suffix(".png") as a different file and unlink the output.
2026-08-15 00:33:01 -07:00
Adolanium 7de5a65906 fix(pets): delete row strips after extracting their frames
Each hatch generates one row strip per state into cache/images and
extracts the animation frames from it, but never removes the strip. The
only cleanup for that directory runs in the gateway housekeeping loop,
which a CLI, desktop, or cron hatch never starts, so the strips
accumulate for good.

Drop the strip after every attempt once its frames are decoded into
memory, including failed or retried attempts, so a hatch no longer grows
the image cache without bound.
2026-08-15 00:33:01 -07:00
andy c99a45b28e fix(browser): reap leaked agent-browser daemons whose owner is still alive
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:

1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
   loop started, so a leak appearing after boot could never be recovered.

2. `owner_alive is True` skipped unconditionally. In-memory session tracking
   is lost on any exception path between spawn and registration, but the
   owner PID stays up — so such a daemon was skipped forever.

The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.

Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.

Changes:

- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
  the cleanup loop. Cycle 0 preserves the existing startup reap.

- When the owner is alive but the session is untracked, fall back to idle
  age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
  `max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.

- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
  socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
  there, making it a last-activity marker that survives hermes restarts and
  does not depend on in-memory bookkeeping surviving an exception path. It
  scans directory entries rather than reading the directory mtime alone:
  command names repeat, and rewriting an existing `_stdout_click` updates
  that file's mtime but not the directory's, so a dir-mtime-only check would
  report a busy session as idle and reap it.

Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.

Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:33:01 -07:00
Reksely 3d73821e9d fix(agent): stop thread output descriptor leaks 2026-08-15 00:33:01 -07:00
Teknium aa8e92f516 test: pin document.hasFocus in status timer test (useViewedInterval gating) 2026-08-15 00:32:53 -07:00
Teknium edd73daaf4 test(kanban): regression for idle-board WS disconnect detection (#77833) 2026-08-15 00:32:53 -07:00
Teknium f535a4f284 test(desktop): pin document.hasFocus in background-sync backstop tests 2026-08-15 00:32:53 -07:00
Teknium 106207c7fb chore: map contributor emails for attribution audit 2026-08-15 00:32:53 -07:00
Teknium 67a1c1ed1a fix(dashboard): use a fixed sidebar cache TTL (no HERMES_* env var for non-secret config) 2026-08-15 00:32:53 -07:00
AlexDev_ bee0d45ddf perf(desktop): pause background UI work while unfocused 2026-08-15 00:32:53 -07:00
Christopher 5bceb3e84b fix(dashboard): add idle back-off to PTY pump loop (#42627) 2026-08-15 00:32:53 -07:00
Tugrul Guner 6d0d748aa0 chore: add contributor email mapping 2026-08-15 00:32:53 -07:00
Tugrul Guner 44665783a9 fix(kanban): detect WS client disconnect on idle board, prevent zombie poll tasks
Closes #77833

stream_events() only detected client disconnect via send_json()
raising WebSocketDisconnect. When no events were pending (idle
board), send_json was never called, so the poll loop ran forever
even after the client disconnected — leaking one poll task per
disconnect.

Fix: race ws.receive() with a timeout matching the poll interval.
On timeout, poll the DB as before. On disconnect, exit cleanly.
2026-08-15 00:32:53 -07:00
Lucas Oliveira f0cfe5a56f perf(dashboard): bound multi-profile sidebar polling 2026-08-15 00:32:53 -07:00
Gabriel Atkinson 868c400e54 fix(tui): stop disabled pet cell polling 2026-08-15 00:32:53 -07:00
Teknium 1c23c2a50b chore: map contributor email for attribution audit 2026-08-15 00:32:44 -07:00
Teknium 4f29374662 fix(gateway): send-once spritesheet semantics for pet.info (#54730)
pet.info accepts knownRevision; when it matches the active sheet's
revision the multi-MB spritesheetBase64 is elided and
spritesheetUnchanged=true is returned. The desktop floating pet passes
the revision it already holds and keeps its cached bytes, so backstop
refreshes no longer resend ~3.2MB frames over the WS (write-loop stalls,
disconnect storms). Legacy callers omitting knownRevision get the full
payload unchanged.
2026-08-15 00:32:44 -07:00
thatssoheil 8052d5dd24 test(pets): lock the quoted-false behavior on the CLI surfaces
Review follow-up (final round PASS with a repeated suggestion): the pets
CLI regression test only exercised real bools, leaving the exact bug this
fix shipped untested. Add a quoted-'false' test driving _has_active_pet
(now False) and toggle_pet_display (now takes the ENABLE branch,
distinguished by the 'no pets installed' error instead of err=None from
the old wrong-way disable). Verified RED on the pre-fix pets.py and GREEN
on the fix.
2026-08-15 00:32:44 -07:00
thatssoheil ce02f0ab8a fix(pets): cover the pets CLI (doctor, has-active, /pet toggle) too
Second review follow-up: hermes_cli/pets.py still read display.pet.enabled
with bare bool()/truthiness in _cmd_doctor (misreported quoted 'false' as
enabled in 'hermes pets doctor'), _has_active_pet (quoted 'false' treated
as active, so /pet install skipped the selection prompt), and
toggle_pet_display (/pet toggle flipped the WRONG way). All three now go
through is_truthy_value(default=False); the module imports the shared
helper at the top.
2026-08-15 00:32:44 -07:00
thatssoheil 343068eb10 fix(petdex): cover the CLI pet pane and status-line signature too
Review follow-up: _pet_resolve_config (cli.py, the CLI pet pane gate) and
_pet_sig (tui_gateway/server.py, the status-line 'off' signature) read
display.pet.enabled with bare bool/truthiness, so a quoted 'false' still
enabled the mascot on the CLI surface. Route both through is_truthy_value
(default False) like the other three sites.
2026-08-15 00:32:44 -07:00
thatssoheil 9e8828999d fix(petdex): quoted 'false' now disables display.pet.enabled everywhere
Three bare bool() reads of display.pet.enabled (deep-merged config, so a
hand-edited quoted YAML value lands as the string 'false'): the pet.cells
gate, the pet.gallery enabled echo, and the shared pet-state helper.
bool('false') is True, so a quoted value kept the mascot enabled against
the operator's explicit intent.

All three now go through utils.is_truthy_value (default False, matching
DEFAULT_CONFIG). Regression test drives pet.gallery with a quoted 'false'
config and asserts enabled=False; verified RED on the old code.
2026-08-15 00:32:44 -07:00
Lavínia Beghini 28d2de1845 fix(desktop): surface isolated tool failures in pet state 2026-08-15 00:32:44 -07:00
Adolanium 5eb1d2b0aa fix(desktop): restore pet.info backstop after live-sync regression
Event-capable backends no longer polled pet.info, so a cold-start
fail-open enabled:false left the mascot hidden until Settings re-seeded
the store. Keep a slow backstop and short startup retries.
2026-08-15 00:32:44 -07:00
fanfan343 d474ba5615 test(desktop): cover bicubic smoothing in pet sprite 2026-08-15 00:32:44 -07:00
fanfan343 fd0872214b fix(desktop): smooth-scale pet sprite frames
Petdex spritesheets are 192x208px illustration frames, not pixel art.
The desktop canvas drew them with imageSmoothingEnabled=false
(nearest-neighbour), so zoomed pets looked blocky. Enable bicubic
smoothing so scaled frames stay clean at any zoom.

High-DPI backing-store sizing is addressed separately in #83276.
2026-08-15 00:32:44 -07:00
NorethSea 34e05c32ec fix(desktop): render pet sprite at devicePixelRatio for sharp HiDPI output
DPR-sized canvas backing store separated from CSS footprint, tracking
zoom/display changes live. Rendering-fix subset of PR #75307; the
overlay-placement feature portion is out of scope here.

Covers the devicePixelRatio half of #83216.
2026-08-15 00:32:44 -07:00
Teknium 18bac64044 chore: map contributor email for icemeng 2026-08-15 00:32:35 -07:00
thatssoheil bceac696ae fix(desktop): artifacts page timestamps render 1970 and local images fail
All three artifact timestamp sources (message.timestamp,
session.last_active, session.started_at) are epoch SECONDS — the
transcript reader and session-date-groups both multiply by 1000 — but
the collector passed them straight to new Date() (ms), so every
artifact rendered as 1970-01-21. Normalize seconds to ms once at
collection; the Date.now() fallback stays ms.

Local file artifacts (e.g. D:\ComfyUI\output\*.png) fell through to
mediaExternalUrl() which yields a file:// URL the renderer cannot
load. Route through the desktop fs bridge whenever it exists —
readDesktopFileDataUrl already dispatches remote REST vs local
Electron internally (#83380).
2026-08-15 00:32:35 -07:00
Johnny Silverhand 021950ac81 fix(desktop): stop artifact over-indexing
Require explicit provenance for tool-result artifacts while preserving assistant links, MEDIA deliveries, generated outputs, file mutations, and browser screenshots. Normalize persisted Unix-second timestamps at collection time and retain millisecond fallbacks.

Consolidates current-main-compatible work from #41156 and #48577.

Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>

Co-authored-by: tt-a1i <53142663+tt-a1i@users.noreply.github.com>
2026-08-15 00:32:35 -07:00
icemeng 96d6db1993 fix(artifacts): convert DB timestamps from seconds to ms for Date()
The Artifacts view reads message.timestamp, session.last_active, and
session.started_at from the SQLite database, which stores all timestamps
as Unix epoch seconds (REAL). These values were passed directly to
JavaScript's Date() constructor, which expects milliseconds — causing
every artifact timestamp to display as January 1970 dates.

Fix by multiplying the database value by 1000 to convert seconds to
milliseconds at the storage point, using nullish coalescing (??) instead
of logical OR (||) so that valid zero timestamps are not skipped.
2026-08-15 00:32:35 -07:00
Teknium 647949332b docs(desktop): document Settings → Connections (multi-connection registry)
Covers the named-source registry from #86679: forced unique device names,
@profile-device disambiguation, add/edit/remove/test, automatic v1 import,
cloud-via-discovery, encrypted token storage, and the staged rollout note.
2026-08-15 00:32:22 -07:00
Teknium fbaea9bddc feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable) (#86797)
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)

Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.

- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
  existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
  archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
  incl. the compression-lineage recursive CTE); list_sessions_rich gains
  include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
  from every listing path (and the REST sidebar endpoints inherit it with no
  change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
  hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
  when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
  set_session_hidden; _session_response exposes it.

Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).

* fix: teach lost-and-found recovery about the 55-column sessions layout

Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.

---------

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:31:37 -07:00
Teknium d2672a349b feat(gateway): optional profile param on cron.manage RPC (#86796)
cron.manage resolved its jobs store from the process HERMES_HOME, so a profile
whose cron lives in ~/.hermes/profiles/<name>/cron/ was invisible to the default
gateway (and any bot/plugin querying per-profile routines saw 'no cron jobs').

Add an optional 'profile' param that scopes the whole action via
set_hermes_home_override, exactly mirroring the adjacent skills.manage handler:
resolve get_profile_dir(profile), 404 (err 4064) if missing, override in a
try/finally that always reset_hermes_home_override. Omitted/None keeps the
launch-profile behavior, so existing callers are unaffected. cronjob() itself is
unchanged (it already keys off HERMES_HOME).

Enables the Hermes-Bot-Mode plugin to show a bot's real routines
(NousResearch/Hermes-Bot-Mode#37). Needs a SERVE-backend gateway restart to take
effect live. 2/2 in the new focused test.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:25:00 -07:00
adikpb 688abc585f test(vision): assert no max_tokens cap in browser and video aux kwargs
Sweeper follow-up: the browser-screenshot and video kwargs captures now
also assert max_tokens is absent, protecting the central auxiliary
no-cap policy against refactors that would restore the hardcoded caps.
2026-08-15 12:53:37 +05:30
adikpb ec470d9db2 test(vision): assert vision aux calls carry no max_tokens cap
Covers the max-tokens-knob contract: vision call_kwargs omit max_tokens
entirely (configured values, defaults, and even an explicit
auxiliary.vision.max_tokens config entry must never be forwarded), so
providers use their full output budget.
2026-08-15 12:53:37 +05:30
adikpb dcc2f3de1d fix(vision): stop capping aux vision output with hardcoded max_tokens
The vision tools' call_kwargs hardcode max_tokens caps (2000 for
vision_analyze/browser_vision, 4000 for video analysis), truncating
descriptions of complex images at the cap. The centralized aux client
already omits max_tokens by default (#34845) so providers use their
model max output; these three call sites were the leftovers that
bypassed that policy.

Remove the hardcoded caps entirely — the aux client handles the
mandatory-max_tokens Anthropic wire via _resolve_anthropic_messages_max_tokens
(model output ceiling) and Gemini native omits maxOutputTokens (65K ceiling),
so no wire needs an explicit cap.
2026-08-15 12:53:37 +05:30
Teknium ce996d4057 feat(delegation): raise max_concurrent_children default 3 -> 10 (+migration) (#86745)
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/​investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.

Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.

- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
  default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.

Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-15 00:20:32 -07:00
kshitij 8b58f9f68f test(bedrock): pin stream-path cap omission; document truthiness edge
Self-review follow-up: cover call_converse_stream's max_tokens=None path
(same builder, previously unpinned) and document why the shim reads the
caller cap with truthiness rather than 'is None' (parity with the
Anthropic shim's reading).
2026-08-15 12:47:27 +05:30
kshitij 5ef52273cd fix(bedrock): let aux calls omit the Converse maxTokens cap
The Bedrock Converse shim hardcoded 'else 4096' when the caller passed no
max_tokens, so auxiliary vision descriptions stayed capped at 4096 tokens
on the Bedrock wire even after #75253 removed the vision call sites' own
caps (#10809 was only partially fixed there).

Converse's inferenceConfig.maxTokens is optional; when omitted, Bedrock
defaults to the model's maximum allowed output. Thread an explicit
max_tokens=None through build_converse_kwargs/call_converse to omit the
field, and drop an all-empty inferenceConfig from the wire request
entirely. The 4096 default is unchanged for every existing caller (main
transport passes params.get('max_tokens', 4096) explicitly), so only
no-cap aux calls opt in.

Surfaced during review of #75253.
2026-08-15 12:47:27 +05:30
EvanProgramming 30c469b153 fix(gateway): spare pidfile-less Scheduled-Task gateways from the orphan reaper on Windows (#83683)
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).

Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.

The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.

Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.

Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
2026-08-15 12:04:29 +05:30
hermes-seaeye[bot] 0807673e1f fmt(js): npm run fix on merge (#86751)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-15 06:22:02 +00:00
Teknium 0904f50e3e fix(desktop): address connections-registry review findings
Review fixes from #86679 comments (trevorgordon981, helix4u, kshitijk4poor):

- Edit inheritance: mergeConnectionInput preserves fields the editor does
  not carry (cloud org, ssh remoteHermesPath/remoteProfile) so a rename no
  longer wipes them. When the payload carries an ssh host string, stored
  user/port are NOT inherited — the composite host field is authoritative,
  fixing the stale user/port resurrection on edit.
- Token hygiene: tokens only persist on token-auth remotes; switching an
  entry to oauth (or cloud) clears the stale envelope.
- Plain-text opt-in: the panel now surfaces the same consent dialog as
  Settings -> Gateway on keyring-less machines (registry list exposes
  secureTokenStorage; save retries with allowPlainTextToken after consent).
- Registry test isolation: hermes:connections:test builds the probe directly
  from the registry entry instead of coercing against v1 connection.json —
  no more inheriting the v1 global token for a different host, and the local
  entry now probes the app-managed backend (never v1 remote/ssh state, so
  the test button can no longer trigger a v1 file write).
- 'local' id reserved at the validation boundary: a crafted IPC payload can
  no longer replace the local entry via upsert.
- Cloud creation hidden in the editor (a dialable cloud entry comes from the
  Cloud sign-in/discovery flow); migrated cloud entries stay editable.
- First-run migration write is guarded: a failed write keeps the migrated
  registry in memory instead of hard-failing every connections IPC call.
- uniqueLabel(): single label-dedup helper — counts up instead of "X 2 2",
  clamps 253-char migrated URL-host labels under LABEL_MAX; used by
  normalizeRegistry and both migration paths.
- UI copy: staged-rollout note replaces the "side by side" claim; test
  failure toast leads with the failure wording; dropped unused i18n keys.

Tests: +9 pure-module cases (reserved id, token-drop rules, merge
inheritance, ssh host precedence, uniqueLabel); electron+settings suites
1355 passed.
2026-08-14 23:06:23 -07:00
Teknium b54b0521dc feat(desktop): multi-connection registry — named agent sources (schema v2 + IPC + Settings UI)
First slice of multi-source agent support: the desktop can now persist ANY
number of named backends (local runtime, remote gateways, Hermes Cloud
instances, SSH hosts) side by side instead of one global connection plus
per-profile overrides.

- electron/connection-registry.ts: pure v2 registry module — required
  case-insensitively-unique labels (device names), @name-device handle rule
  for duplicate profile names across sources (agentHandle), defensive
  normalizeRegistry for corrupt files, one-time v1→v2 migration that imports
  the global block + per-profile overrides (deduped by URL/host) and leaves
  connection.json untouched for older builds.
- main.ts: connections.json storage beside connection.json (same secret
  posture: safeStorage-encrypted tokens, 0600, tighten-before-parse, mtime
  cache) + hermes:connections:* IPC (list/save/remove/set-primary/test).
  Test maps registry entries onto the existing testDesktopConnectionConfig
  probe stack — no new probe code.
- Settings → Connections: manage the registry (add/edit/remove/test/make
  primary) with forced naming; local entry is non-removable; removing the
  primary retargets to local. en + zh locales.

Storage-level only by design: routing/pool generalization to composite
(connection, profile) keys, the multi-source roster, plugin SDK surface, and
fan-out updates land as follow-up PRs.
2026-08-14 23:06:23 -07:00
kshitij e3fab0437e refactor(cache): never-raising scope resolver shared by both call sites
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.

- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
  returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
  the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
  shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
2026-08-15 11:09:56 +05:30
kshitij 96cdf19a0b refactor(cache): fold self-review findings on the rotation-scope fix
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
  _session_db re-resolves instead of staying pinned to the physical id);
  _persist_disabled agents (background-review forks that never get a DB row)
  memoize the fallback instead of re-querying the lineage per API call;
  module docstring cross-references get_conversation_root and why the two
  lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
  _prompt_cache_scope_for_agent(agent) call to a single local above the
  OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
  don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
  key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
2026-08-15 11:09:56 +05:30