A wall-clock step-back (NTP) can leave the marker's mtime in the future,
which would push the idle window out by the step size. Clamp with
min(mtime, now); _last_inbound_at has the same exposure but is at least
bounded by process uptime. Review comment on #100830.
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.
The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
Collapses the three test-refinement commits from #100614 (423f7f8703,
157db13c76, ffa72dd67f): model recovery pressure by provider-call state,
assert the rebuilt-oversized retry fails closed, keep the compacted
history compressible (user+assistant summary rows).
The `${{ false && (...) }}` if-expression on the reusable-workflow
e2e-desktop job made GitHub's workflow parser fail at startup (annotation:
"An unexpected error has occurred"), so every ci.yaml run on main and
every PR dispatched 0 jobs. Discriminator on this branch: pre-24f5a60
blob → 26 jobs; paren-relocated && form → 0; bare if: false → 25.
Lane stays disabled; comment documents the re-enable line.
session-tile-owner-route.test.ts asserted against the TEXT of
session-tile.tsx, so it passed on a broken implementation whose call site
merely looked right, and failed on this refactor, which changed nothing the
tile actually does. AGENTS.md bans the pattern outright.
Extracted the ladder as tileOwnerRoute() and replaced the three regex
assertions with six that call it: tile route wins, row falls back, hint
falls back, targetProfile carries through, a bare profile narrows away, an
untagged session stays ambient. 12s of source-matching becomes 1.1s of
behavior.
The coalescing key now carries the owner, so two backends that both expose a
session called `parent` get their own create instead of the second caller
receiving the first's child. Both creates are held open, which is the only
state the key guards — a sequential version passes even with the owner
stripped out.
Co-authored-by: Ahmett101 <ahmet.tunc@gmail.com>
TileChat re-renders per streamed token, and the owner ladder spread three
session arrays before scanning them on each one. Subscribe to the atoms it
actually reads and memoise the lookup on them, so it recomputes when the
tile store or a session list changes rather than per frame.
branchStoredSession looked its parent up in $sessions alone, and
branchCurrentSession did the same. A conversation reachable through a
profile-scoped project tree has no row there, and when it appears in both
places the flat Recents copy is the ownerless one — so the lookup returned
the row that cannot route, and the branch created its child on whichever
backend happened to be active.
cachedSessionRow spans Recents, cron, messaging and the project tree, and
prefers the self-describing candidate. One ladder, used by both branch
entry points and by resolveStoredSession.
Co-authored-by: evan-bradford <evan-bradford@users.noreply.github.com>
A scoped projects.tree / projects.project_sessions response is built from
ONE profile's state.db, so the request scope is authoritative even for
legacy rows whose persisted profile_name is NULL. Without the stamp those
rows reach the renderer ownerless, and every owner lookup off them — the
branch path included — falls back to whichever backend is active.
Co-authored-by: evan-bradford <evan-bradford@users.noreply.github.com>
Review follow-up on #93959:
1. Partial-failure window: if the row commits but the transcript copy or
title write fails, the durable-but-empty child defeated the lazy
first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
the renderer fail-latched on a transcript-less session again. The seed
block now compensates: delete just this child so the lazy path can
retry cleanly. Disk-full is exempt — deleting data on a full disk makes
things worse.
2. Silent degradation: the best-effort catch now logs at WARNING with
exc_info instead of DEBUG, so a regression in this user-facing path is
observable without enabling debug logs.
Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.
A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:
- Row created in the PARENT's profile-scoped state.db, stamped with
_branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
from pending_title — the branch lands in the parent's lineage instead
of falling back to a message-preview name.
Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.
Fixes#93959
Routing the create and stamping the optimistic row still left a
remote-owned branch child flickering into "Couldn't open this session —
Session keeps losing its backend runtime right after resuming". The RPCs
were right and every resume succeeded; the OWNING SOCKET was the
casualty. Three gaps, one cause — nothing durable named the owner:
- openSessionTile persisted ownerRoute only for workspaceMode==='bots',
so the branch tile pinned nothing in the gateway keep-set
(openTileGatewayScopes / foregroundSessionScopes). The pruner closed
the owner socket, the backend orphan-reaped the draft runtime,
session.reclaimed unbound the tile, resume re-armed and succeeded on a
fresh socket the next recompute closed again — until the resume-storm
breaker (#93892) latched the error card at TILE_RESUME_STORM_LIMIT.
Persist the route for sessions-mode tiles whose opener knows the exact
owner, and stop a route-less re-scope from clobbering it.
- forkBranch, unlike both sibling routed creates in the same file, never
called setSessionOwnerHint/holdSessionOwnerUntilForeground — so in the
gap between session.branch returning and the tile publication landing,
no keep-set rung named the owner and a prune could reap the just-minted
draft runtime before the first prompt. Add both, mirroring the
siblings; the hold retires once the tile's own route covers the scope.
- resetTileRuntimeBindings preserved cross-connection runtimes only for
bot tabs, so a flapping sibling connection (an SSH source re-dialing)
dropped the branch tile's healthy binding on every reconnect, re-arming
resume each time — the same storm by another path. Preserve any
owner-routed tile; the owner's own reconnect still rebinds.
Also keep open tiles' rows in sessionsToKeep: a branch child is a draft
the aggregator cannot return until its first turn persists it, so the
next background refresh silently dropped the optimistic "draft: branch
#N" row and the sidebar showed no trace of the branch until first send.
Each fix verified RED by reverting its line. Live acceptance against a
remote-owned parent branched from another connection: draft row visible
immediately and stable across refreshes, tile carries the owner route,
first turn accepted and completed by the owning backend, and no
storm/error card through the full 120s storm window — before the fixes
the card latched at ~20s.
Routing the branch create to the parent's owning connection was only half the
job. The child then landed in the sidebar as a row that lied about who owned
it, so the chat pane spun forever on "draft: branch #1" and never hydrated —
the create was right, the row was wrong.
upsertOptimisticSession stamps the row's profile from $activeGatewayProfile and
omits connection_id entirely when no owner is passed (utils.ts:1318-1342), and
it also skips setSessionOwnerHint. The branch call site passed no owner, so the
child got NEITHER a row tag NOR a hint. resumeSession's owner ladder starts at
`capturedOwner || getSessionOwnerHint(storedSessionId)` and forkBranch calls it
without a capturedOwner, so the missing hint alone was enough to send the
resume to whichever backend happened to be active. Pass the parent's route as
the owner argument, restoring both mechanisms. The two sibling routed creates
in this file already did exactly this.
The tile path had the same defect one rung further out. A branch of a session
that is not the open chat opens a tile instead of resuming, and
SessionTileChrome resolved its owner from the tile route alone. openSessionTile
is called for a branch child with no workspaceScope, and session-states.ts only
persists a tile ownerRoute in bots mode, so that tile had no owner at all and
its model + composer RPCs fell back to the ambient socket. Use the same
tile-route-then-row ladder its sibling in session-tile-actions.ts already uses,
resolved per render so it cannot go stale against the tile store, the
recents/cron/messaging rows, or the hint map, with only the resulting identity
memoised on primitives.
An untagged parent row still reproduces the previous ambient behaviour exactly,
so single-connection users are unaffected.
Verified end to end against two real gateways: a session owned by a remote
connection, branched through the actual sidebar context menu in a running dev
app. The remote gateway served the create (ws closed ... messages=11
detached_sessions=1) and the resulting row polled stable at connection_id =
the remote for the full 8s window. Before the fix the same gesture produced a
row with no connection_id.
Branching a session owned by one connection while another is active created
the child on the wrong backend — or nowhere — while the sidebar still painted
an optimistic row. That row pointed at an id no backend owned, so hydration
retried, exhausted, and armed the stranded-session overlay: "Couldn't load
this session. The connection to this session failed and automatic retries
gave up." Retry re-ran the same mis-route, so it never recovered.
branchStoredSession and branchCurrentSession resolved only the parent's
PROFILE and called ensureGatewayProfile(profile), then dispatched through the
ambient requestGateway. A profile name does not identify a backend once
several connections expose the same name, so both the parent transcript read
and the session.create/session.branch RPC landed on whichever socket happened
to be active. removeSession, twelve lines away, already routed by
(connectionId, profile) via SessionOwnerScope — branch simply never got the
same treatment.
Reuse that existing contract: derive the exact owner from the parent row with
sessionOwnerRouteFromRow, activate it with ensureGatewayAgent, and dispatch
via requestGatewayForAgent. getAllSessionMessages takes the same owner scope
so the transcript read cannot silently come back empty and abort the branch
as "nothing to branch" before any create is attempted. Both arms of forkBranch
(session.branch for the open chat, session.create for a sidebar right-click)
are covered.
An untagged parent row — the single-backend case — keeps the previous
profile-only path exactly, so behaviour is unchanged for users with one
connection.
Tests: three call-site regressions asserting the create rides the owning
(connection, profile) socket, that the transcript read carries the same owner
scope, and that an untagged parent still uses the ambient socket. Plus an
integration test that mocks nothing inside the router — the real
requestGatewayForAgent runs against a fake Electron bridge and transport, so
a regression that re-collapses a registry route onto the ambient socket fails
even if the call-site assertions still pass.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
The everything-flow's client leg read `$updateStatus.get() ?? await
checkUpdates()`, so a cached row always won. That row can be up to a
poll interval (30 minutes) old and is captured before the backend leg
runs, so a cached "already current" skipped the client apply entirely —
the stale-GUI gap the flow exists to close.
Re-check first and fall back to the pre-flow snapshot when the live
check can't answer. `checkUpdates()` resolves with an error status
rather than rejecting and overwrites the atom with it, so the snapshot
is taken before any leg runs.
Co-authored-by: Dhana <227747512+whoisdhana@users.noreply.github.com>
The update entry points chose their target from the connection mode, so
every surface in remote mode acted on the backend — including the ones
showing the client's own status.
The macOS "Check for Updates…" app-menu item is the clearest case: it
sits next to "About Hermes" and is the OS-standard way to update THIS
app, but on a Mac connected to a remote Linux backend it checked the
Linux box. The backend was already current, so the action reported
nothing and did nothing, and the desktop app drifted months behind with
no error and no updater log to explain it. The update-available toast
had the same split: a client check raised it, clicking it opened the
backend's overlay, which has no target switcher and no way back.
Surfaces bound to one target now name it; only genuinely generic
commands still take the connection-mode default, so the command
palette's remote-mode backend target and the everything-flow are
unchanged.
Co-authored-by: BerneYue <14088768+yuexiongHNU@users.noreply.github.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
Co-authored-by: clayduncan <234173110+clayduncan@users.noreply.github.com>
Co-authored-by: Dhana <227747512+whoisdhana@users.noreply.github.com>
When the bundle was swapped under a running process, the About banner sent the
user to the installer — a download and a reinstall for a state that a plain
restart repairs, and the reason reinstalling never helped these reports.
Report bundleSwapPending on hermes:version and give that case its own copy and
a "Restart Hermes" button. It gets its own headline too: reusing "App build out
of date" over a body that says the app is already installed repeats the
contradiction with the Updates card that the banner is supposed to resolve. The
installer link stays for the genuinely-stale-bundle case.
Packaged builds only — a dev `--build-only` rewrites the stamp under a running
`npm start`, and that is a rebuild the developer asked for, not a torn install.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
A user who reopens Hermes while an update is running lands on the boot gate,
which is what it is for. But the updater swaps the packaged bundle on disk
after `hermes update` exits, and its `open` leg only focuses this already-
running process, so nothing ever loads the new build. The parked instance then
passes the gate and boots the new runtime under the old renderer — the "App
build out of date" banner immediately after a fully successful update, over an
Updates card that says "You're on the latest version" and so offers no remedy.
Compare the install stamp this process loaded at boot with the one on disk when
the gate clears. On positive proof of a swap — different commit, or a different
builtAt at the same commit — relaunch instead of starting a backend. Detection
fails quiet like bundle-skew, so a swap that never happened (the Windows
locked-binary case) is unchanged. A one-shot argv flag makes a relaunch loop
impossible and a 15s failsafe falls back to the old behavior.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
Co-authored-by: aeonsong <aeonsong@users.noreply.github.com>
detectBundleSkew() trusted `git rev-list --count <stamp>..HEAD -- apps/desktop`
outright, which claims skew in two states where the install is not torn.
Ancestry: `A..HEAD` only measures how far HEAD is ahead of A when A is an
ancestor of HEAD. A ZIP-fallback update rewrites the tree onto a synthetic
root, so the stamp still resolves but is unreachable; the range degenerates to
HEAD's own history and reports a permanent >= 1 while apps/desktop is
byte-identical. Ask `merge-base --is-ancestor` first and go quiet unless it
answers yes.
Scope: the pathspec counted every file under apps/desktop/, so a docs- or
e2e-only commit produced a banner promising missing UI features that do not
exist. Count only the paths that reach the shipped app.
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
Co-authored-by: kokhlo <kokhlo@users.noreply.github.com>
Review finding: dashboard_client_last_seen() discarded the marker once it
was >= 45s old, so after the client disconnected the gateway fell back to
its own (much older) _last_inbound_at and suspended ~45-75s later, not
idle_timeout later as the PR claimed. The staging release leg had in fact
shown 46s.
The marker mtime is a timestamp of real inbound; is_idle already judges
recency. Drop the staleness cutoff entirely and always fold the raw mtime
into the inbound clock (max with _last_inbound_at). An old marker is
harmless: it is outside idle_timeout just like an old _last_inbound_at.
Also:
- touch the marker immediately after ws.accept(), before the ready/skin
setup, so a client waiting on a slow ready frame is still visible
- make the newer-message-wins test discriminating (timeout 10s, chat 5s,
marker 40s: choosing the marker would read idle)
- add < idle_timeout / >= idle_timeout / ancient-marker boundary tests
Re-validated live on hermes-agent-stg-test-6698: last client frame
03:18:11Z -> going dormant 03:20:17Z (126s = 120s timeout + watcher tick)
while the gateway's own inbound clock was 368s stale. Mutation check vs
origin/main files: run.py hunk reverted -> 4 fail, ws.py reverted -> 3
fail, "prefer marker over max" -> 1 fail.
* fix(linux): install Lanczos-resized panel icons, not a PNG in scalable
Cinnamon's panel is ~24px. v2026.8.31 dropped the 1024px asset into
hicolor/scalable (SVG-only), so the Mint panel nearest-neighbor scaled
it into a mangled blob. Decode the PNG and write 24/32/48/256 rasters;
undecodable bytes still copy into one indexed dir. Drop the leftover
scalable file.
* test(linux): cover resized hicolor panel icons and stale scalable cleanup
Pin that a decodeable PNG lands as 24×24/256×256 rasters (not scalable),
a leftover scalable copy from v2026.8.31 is deleted, and truncated
PNGs still fall back to an indexed copy.
`/pr-triage [[ … [412 lines] … ]]` dispatched the LABEL: the paste
expansion was computed but only consumed by /queue, so every other
command received "[412 lines]" as its argument and the agent
faithfully reported the paste as truncated.
prepareSlashSubmission names the split the branch actually needs — the
transcript keeps the collapsed label, the dispatch carries the full
text. Image tokens stay as labels, since the gateway already holds
those files in attached_images.
parseSlashCommand split the whole line on `\s+` and rejoined with a
single space, so `/pr-triage <pasted diff>` reached the skill as one
run-on line. Only the separator between the command name and its
argument belongs to the parser; everything after it is the user text
and now survives verbatim.
Every consumer already re-splits the arg it receives, so subcommand
parsing (`/cron add`, `/model x --global`) is unchanged.
The gateway's idle predicate only stamped _last_inbound_at for messaging
inbound, so a scale-to-zero instance suspended under an open desktop app /
web dashboard / TUI. The client's reconnect loop then re-poked the
Fly-proxied hostname, autostart resumed the box, and the instance flapped
suspend -> proxy-wake every ~60s (13 of 72 active opted-in prod instances
on 2026-09-02, with [PC05] connect timeouts visible to the user).
The dashboard runs in a separate process on hosted instances, so the
signal crosses over as a marker file under HERMES_HOME/state:
- tui_gateway/ws.py touches state/dashboard_clients.heartbeat on every
/api/ws connect and inbound frame (clients gateway.ping every 15s),
throttled to one write per 5s per process.
- gateway/scale_to_zero.py: dashboard_client_last_seen() reads the mtime;
a marker older than 45s (the clients' heartbeat deadline) is "client
gone", a missing marker is "no client" (not fail-awake, or nothing
would ever sleep), an unreadable marker fails awake.
- gateway/run.py folds that into seconds_since_last_inbound, so an
attached client gets exactly the same idle_timeout grace after it
disconnects as a chat message does. No new conjunct, _last_inbound_at
itself is not mutated.
Validated as a hot patch on hermes-agent-stg-test-6698 with a real
/api/ws client pinging every 15s: machine held awake for 6.5 min of zero
proxied traffic (was 2 min), then suspended within ~50s of the client
exiting. Tests exercise the real seams (temp HERMES_HOME, the runner's
_scale_to_zero_is_idle composition, tui_gateway.ws.handle_ws) and were
mutation-checked: reverting either half or dropping the staleness check
fails 2-3 of them.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
PR re-review caught that cli-config.yaml.example and the setup wizard
still described the OLD day-stamp gate ("period starts on or after the
day you opted in"). The actual gate (CONSENT_GATE_SQL) requires the
entire package period to be contained in one recorded consent window -
stricter, and privacy-significant at opt-in/revocation boundaries: a
package straddling a revocation starts after opt-in yet is correctly
held back. Both surfaces now state the containment rule literally;
docs A.1 already did.
prompt.submit now retries a completed failed build (installing a fresh
unset agent_ready before rebuilding), so a no-op _start_agent_build stub
leaves the patient wait blocking forever — the turn thread outlived the
test and the whole file hit the per-file timeout on CI. Stub a faithful
failing build instead: set agent_error, fire the session's current
ready event. The pinned contract (visible failure, no silent drop) is
unchanged.
The two #57911 regression tests hand-encoded the remote workspaceCwdKey()
localStorage literal; seed it through the exported setCurrentCwd() like the
sibling isolation test so a key-shape change can't silently strand them.
(/simplify-code reuse finding)
The remote remembered cwd is no longer a bare-new-session default after
the #57911 fix; the comment listed it as priority 2. Follow-up from
salvage self-review of PR #57926.
A bare-new-session (Cmd+N without a project scope) in remote mode used to
inherit the remembered cwd under workspaceCwdKey()'s remote variant — i.e.
the last project the user attached to on that backend. Symptom: Cmd+N landed
on the wrong project's workspace (e.g. a configuration discussion thread
opened in the tradingview directory).
The remote branch in workspaceCwdForNewSession is the entire bug class:
- It overrides the configured-default-pre-attaches rule, so users who
configured an explicit default didn't get it under remote mode either.
- The bare-new-session path has no notion of "remember where I was" —
that's startSessionInWorktree's job. The remembered cwd is only meant
to assist resume/restore (ensureDefaultWorkspaceCwd, which keeps its
remote-keyed sticky seed and is unaffected by this change).
Drop the mode === 'remote' early return and route both modes through
getConfiguredDefaultProjectDir(), matching the proposed fix in the issue.
Test: the existing 'keeps remote workspace memory separate from local and
other remotes' case set the local key then switched connection to remote,
so under the buggy code it returned '' regardless — it never gated the
regression. Add a remote-key repro and a symmetric explicit-default test
so the fix is observable from the suite.
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).
Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.
Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes#97994.
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.
Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.