Commit Graph

33615 Commits

Author SHA1 Message Date
teknium1 9f9ec647f5 fix(gateway): re-check the live store before broadcasting the state.db warning
_send_session_db_warning_notifications() broadcast the error recorded at startup
without asking whether it was still true. A startup `database is locked` routinely
clears while the adapters are still connecting (another profile's open, a `hermes
sessions` one-shot, a slow SMB lock release), so the home channels were told the
store was unavailable when it had already healed — and a warning that is wrong once
is ignored the next time it is right.

The RecoverableHandleCache opener already clears `_session_db_init_error` on
recovery; the broadcast now drives one open attempt through it first and only warns
when the error is still standing. A store that is genuinely still down warns exactly
as before.

Fixes #108031.
2026-09-11 06:23:23 -07:00
teknium1 045704ce00 test(state): port the #107411 dentry tests to the shared WAL-generation harness
main replaced the file-local _make_db/_require_wal/_unlink_sidecars helpers with
tests/hermes_state/_wal_generation_harness (make_db/require_wal/lose_sidecars) after
PR #107411 branched; the cherry-picked tests referenced the old names and failed at
collection with NameError. Fixture-only change; the assertions are unchanged.
2026-09-11 06:23:23 -07:00
nikkoxgonzales 4d9ec9be50 fix(db): survive DELETE fallback failure in WAL setup 2026-09-11 06:23:23 -07:00
gaoanze 031d3760d0 fix(state): distinguish retired WAL recovery guidance
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.

Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
ca-shrimp 3a3ec6eeaf fix(state): serialize replaced/generation probe with close() in _execute_write
A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (#105567).

Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.

Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.

Fixes #105567
2026-09-11 06:23:23 -07:00
chelsealong d6b036b726 fix(state): compare fd identity against the watched path, not st_nlink
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
2026-09-11 06:23:23 -07:00
chelsealong 84a3c4de74 fix(state): require nlink==0 before treating a /proc fd as an unlinked WAL sidecar
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.

Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
2026-09-11 06:23:23 -07:00
Teknium 7dec81568e test(desktop): trim salvaged pool tests to invariants
Drop two PrimaryProfilePin cases that only restate the constructor
defaults and blank-string normalisation, and the wiring-routing test that
froze POOL_LIMITS_SETTINGS_ROUTE to a literal string — a snapshot of the
constant, not a behaviour contract. The two kept pin tests cover the bug
(a live primary keeps answering for its booted profile after the stored
preference moves; teardown releases the pin), and the notifications tests
cover the toast action end-to-end.
2026-09-11 06:23:18 -07:00
Mabolla 19cff347ad fix(desktop): wait for an evicted backend to release its slot
Await LRU teardown in each pooled backend creation path so replacement wakes do not race an exiting child for the hard spawn slot.
2026-09-11 06:23:18 -07:00
Mabolla 919dea9020 fix(desktop): await LRU teardown before queued wake
Wait for selected stale backend processes to exit before a replacement profile wake enters the bounded spawn queue.
2026-09-11 06:23:18 -07:00
joaomarcos a863bbcc28 fix(desktop): make pool slot timeouts actionable 2026-09-11 06:23:18 -07:00
Alexandru Ionescu d27180ba7f fix(desktop): pin primary backend routing to its booted profile
`primaryProfileKey()` re-read active-profile.json on every call. The rail's
live workspace switch rewrites that file via `hermes:profile:remember`
WITHOUT re-homing the primary, so after a switch the routing table disagreed
with the running process: a request for the profile the primary actually
booted as (e.g. "default") no longer matched `primaryProfile` in
`resolveProfileBackendRoute`, fell through to the pool, and spawned a second
backend for the same HERMES_HOME.

The duplicate was keepalive-fresh so LRU eviction spared it, it burned a pool
slot, and with the default cap of 3 every further profile queued and failed
with `Local backend start for "<profile>" timed out while waiting for a free
slot` (repro in desktop.log: "default" spawned as a pool backend while the
primary "default" was still running; coder/qwen then timed out for 20+ min).

Snapshot the launch profile in `startHermes()` (PrimaryProfilePin.pin) and
release it in `resetHermesConnection()` so the next start follows the stored
preference again. The pin is a tiny pure module with tests; main.ts only owns
the file read and the two call sites.
2026-09-11 06:23:18 -07:00
Teknium d1bd7b00df fix(desktop): dispatch probe falls back to /api/status on pre-/api/health remotes
Switching the pooled dispatch probe to /api/health (salvaged from #97914)
would 404 on every dispatch against a remote older than 0.19, retire the
tunnel and reconnect forever - the same storm #107997 describes, moved to
old backends. Fall back to /api/status on an explicit 404 exactly the way
the boot readiness probe already does (backend-health.ts). The legacy
fallback idea and its test are taken from #101976 (@edosulai); the rest of
that PR (timeout-tolerance streak, ServerAlive SSH options) is not adopted.

Co-authored-by: Edo Sulaiman <edosulai@icloud.com>
2026-09-11 06:22:35 -07:00
bennybuoy ee8257d601 fix(desktop): probe /api/health on pooled SSH dispatch, not /api/status
Cold /api/status through a Windows no-mux SSH forward routinely exceeds
the 2.5s dispatch budget, so Desktop retires a live tunnel and respawns.
Use the cheap /api/health route (5s, same as DEFAULT_HEALTH_PROBE_TIMEOUT_MS).
Background liveness still probes /api/status at 10s.
2026-09-11 06:22:35 -07:00
KoNit-K 8ab08968f4 fix(desktop): give pooled remote dispatch probes the cold-start liveness budget
A 2500ms dispatch probe is shorter than quiet-box hermes serve cold-start (~6-8s), so a just-woken pooled backend always fails and reconnects. Reuse REMOTE_LIVENESS_TIMEOUT_MS (10s) for that probe.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 06:22:35 -07:00
Teknium 902eccae3f test(state): trim holder-scope suite to the two behaviour contracts
Keep the two invariant tests from #105428: a genuine second instance
whose argv is scoped to another HERMES_HOME is not a holder of our
state.db; argv that names our state.db stays flagged. The helper-shape
tests were change-detectors on `_argv_scoped_to_other_home` internals.

The salvaged commit is re-authored to TaoMasterCoder's GitHub noreply
identity: the original `devops@77hub.com` is a shared org address
(misconfigured local git, not malice).

Refs #107440
2026-09-11 06:21:48 -07:00
TaoMasterCoder 085623a11a fix(state): scope uninspectable-holder fallback to the instance's own HERMES_HOME
Field evidence (2026-09-07, production host): the host runs two independent
Hermes instances - a main gateway (user ubuntu, HERMES_HOME=/home/ubuntu/.hermes)
and a demo gateway (user demo, HERMES_HOME=/home/demo/.hermes). Because the
demo process is owned by another user, its /proc/<pid>/fd table is unreadable
and foreign_state_db_holders() falls back to cmdline + _looks_like_hermes().
The demo argv matches Hermes patterns exactly, so it was flagged as an
uninspectable holder of the MAIN instance's state.db even though lsof proved
0 open handles on it. Result: _recover_stale_fts was deferred 42 times across
6 gateway restarts, the fts_stale breadcrumb never cleared, and FTS
self-repair stayed permanently disabled.

#92419 removed substring false positives (journalctl/grep mentioning hermes);
a genuine second instance with a DIFFERENT HERMES_HOME was still misjudged.
Add _argv_scoped_to_other_home(): when the argv of an uninspectable Hermes
process proves it lives under a different /.hermes home (or a state.db
sidecar under a different parent) AND no token references our state.db,
sidecars, or home, do not count it as our holder. Applied to all three
uninspectable branches (holder + two descriptor paths). Ambiguous argv
without absolute-path tokens remains fail-closed, preserving the
conservative intent.

References #92401
2026-09-11 06:21:48 -07:00
Teknium d41e906598 test(hosted-rooms): pin the coordination store off the master state.db; label sidecars
Invariant test for the #103489 salvage: a profile gateway and the root
gateway resolve the same hosted-room coordination file at the install
root, and that file is never the root SessionDB (`state.db`). Every
profile gateway starts the hosted-room worker unconditionally, so the
old mapping made each profile process a long-lived writer on the master
session store — the simultaneous-restart corruption in #102120 and the
main offender in the #103339 / #103490 fleet reports.

Also: WAL-fallback labels say `shared-state.db`, and test fakes take the
store path from `default_db_path()` instead of hardcoding `state.db`.

Reported-by: #102120 #103339 (@RChina) #103490
Refs #102120 #103339 #103490
2026-09-11 06:21:48 -07:00
Teknium cfeccbdf34 fix(hosted-rooms): local_profiles skips the profile-delete tombstone dir
`hermes profile delete` leaves `profiles/.deleted/<name>` behind.
`HostedRoomService.local_profiles()` fed every subdirectory name to the
roster, so `.deleted` failed `validate_roster`'s identifier check and
`plan_next_task` raised on every cycle for every room until the
directory was removed by hand (#106847, bug 2). Skip dot-dirs and
tombstoned profiles, using the same `named_profile_is_deleted` predicate
`hermes_cli.profiles` uses to list live profiles.

Refs #106847
2026-09-11 06:21:48 -07:00
Teknium c9b71ebaf7 fix(dashboard): startup schema reconcile opens state.db read-only first
The dashboard's `_eager_reconcile_own_session_db` did an unconditional
writable `acquire()` at every startup. When the gateway shares that
state.db the dashboard became a second long-lived writable SessionDB
owner: a close-time WAL checkpoint plus a possible FTS rebuild in
`_init_fts`, the two-writer vector behind the corruption reports in
#107688 and #100896 ("5 live SessionDB handles" precursor, gateway +
dashboard both holding the WAL).

Route the startup reconcile through `_open_session_db_at_path(...,
read_only=True)`, which already bootstraps a missing store and heals a
stale/malformed schema through exactly ONE writable open before
reopening read-only. A healthy store now gets zero writable opens from
the dashboard while the #79531/#80037 "bring schema current before the
first poll" contract is kept (existing heal test unchanged).

Live repro (healthy store, count writable SessionDB.__init__ calls from
the startup worker): before=1 after=0.

Reported-by: #107688, #100896 (@kokhlo diagnosis)
Refs #107688 #100896
2026-09-11 06:21:48 -07:00
Halldrix b02f3aa00b fix(doctor): refuse the WAL checkpoint while a live writer holds state.db (#103339)
doctor --fix ran a raw sqlite3.connect + wal_checkpoint(PASSIVE) with no
holder guard. Under a running gateway that second connection joins the
live WAL and its close-time handling is the second-writer corruption
class #103339 tracks. Route the checkpoint through live_writer_holds_db
and skip with an actionable finding when the database is held.

[Salvage note: the original PR also flipped the repair probe's
DatabaseError branch to fail closed; that hunk is dropped here because
it refuses every header-destroyed DB (the exact class repair exists
for) — the probe's own connect raises before the EXCLUSIVE/BEGIN
statements ever run.]
2026-09-11 06:21:48 -07:00
joaomarcos d294a655f3 fix(cron): serialize lifecycle script reads with sqlite locks 2026-09-11 06:21:48 -07:00
Teknium 4e743b06a2 chore: contributor mapping for RikETS 2026-09-11 06:21:48 -07:00
Rik 0e422e0ece fix(gateway): isolate hosted-room state in shared-state.db — profile gateways must not write master state.db
gateway.hosted_rooms.default_db_path() resolved profile gateways
(~/.hermes/profiles/<name>/) to the ROOT state.db, making every profile
gateway process a long-lived writer on the master session store. Across a
6-gateway fleet this was the recurring multi-writer corruption vector
(WAL/FTS collisions during bot-gateway restart storms, observed 2026-09-03
with four state.db corruptions in five weeks fleet-wide; upstream refs
#100896, #100313, #102827).

Move hosted_room* tables to a dedicated root shared-state.db so profile
gateways never open state.db writable. Single-gateway installs are
unaffected (same file semantics, different filename); multi-profile fleets
stop sharing the master store.

Verified live in production: deployed 2026-09-04, zero I/O-error /
reopen-race events across 9h+ of active multi-gateway load (previously
~1 failure per 2-3 min).

Co-authored-by: Rik <rik.sedd@gmail.com>
2026-09-11 06:21:48 -07:00
Teknium 52fb4e0877 fix(desktop): read Bot Mode avatars over the active socket, not one dial per bot
useRoster repaints every 5s and hands pullServerAvatars the active-source
rows. Since the multi-source merge (ed20a6f01a) every such row is
sourceScoped, so the avatar sync branch chose requestForBot and dialed each
bot's OWN backend to read a profile-directory PNG: a fresh WebSocket with
one JSON-RPC message, torn down at refcount 0, per bot per tick (#99336's
"ws accepted / ws closed messages=1" every ~5.2s on background profiles),
and for every bot with no running backend a pool spawn that waits out
POOL_SLOT_WAIT_MS (30s) and is re-queued by the next paint, forever, once
the pool is full (#102913's per-bot "waiting for a free local slot ...
timed out" cadence). The loop was self-sustaining because the plugin's own
160px face raster is deliberately not parked in $botMeta, so the empty
image slot re-fetched it on every tick.

Assets are files under the profile directory; the gateway that just
answered profiles.list reads them for any of its profiles. Route the three
avatar RPCs through host.request like the roster query itself, and remember
face-only answers so a row is fetched once, not once per tick.

Not changed: relay.ts (its loops dedupe to one route per registered
connection and return early below two connections, so a single-connection
desktop never issues a relay RPC), and useRoster's own profiles.list, which
already rides the active socket via requestForBot({name}).
2026-09-11 06:21:24 -07:00
Teknium e4080b381b fix(desktop): roster avatar sync no longer dials a backend per registered profile at launch
The first Bot Mode roster paint after launch ran pullServerAvatars over every
row, and for a source-scoped row (every row on a local-primary desktop once
host.agents annotates the roster) each profiles.get_asset / set_asset went
through requestForBot -> host.requestProfile -> requestGatewayForAgent, i.e. a
(connectionId, profile) secondary that spawns that profile's pooled backend.
With ~60 registered profiles and 3 warm slots this queued 56 background spawns
at boot; each queued dial then rode reconnectSecondary's backoff until the
stall budget parked it (#107969), so the pool never drained and desktop.log
filled with "waiting for a free local slot" (#102978).

The active gateway's own profiles.list already produced these rows by reading
every local profile directory; get_asset/set_asset are the same directory
reads addressed by name. Route both through host.request on the active socket
with the row's backend profile name (route.targetProfile, so managed aliases
still resolve). No secondary socket, no pool slot, no spawn.

Cross-connection (remoteSource) rows never reached this path: pullServerAvatars
is fed activeSourceRoster, which filters them out.
2026-09-11 06:21:24 -07:00
kshitijk4poor ff59fcd710 fix(state): skip the display-trigger drop+recreate when the triggers already match
DEFERRED_INDEX_SQL unconditionally DROPs and re-CREATEs the four display
triggers on every open so a changed trigger body rolls out; DROP TRIGGER IF
EXISTS on an existing trigger takes the write lock, so a settled database
still blocked behind a sibling's transaction after the three statement
gates. Compare each trigger's stored sql against the desired CREATE text and
run the pair only when they differ. CREATE ... IF NOT EXISTS forms are
lock-free on an existing object and run as written. The no-writes test now
also traces DROP TRIGGER / ALTER.
2026-09-11 06:21:00 -07:00
kshitijk4poor c4c7e102aa refactor(state): tighten generation-stamp comment to the why 2026-09-11 06:21:00 -07:00
kshitijk4poor b8d7977f0d test(state): trim #101881 tests to two invariants (zero writes on settled open; open not blocked by sibling write lock) 2026-09-11 06:21:00 -07:00
John Paul Soliva ef8d682200 perf(state): stop taking the state.db write lock to open a database that needs no writes
Every read-write SessionDB open issued three writes that usually change
nothing, and a write statement takes the database write lock even when it
matches no rows:

- _ensure_db_file_generation's INSERT OR IGNORE into state_meta. The stamp
  is minted once per FILE, so every open after the first inserted nothing.
- the NULL-`active` heal, UPDATE messages SET active = 1 WHERE active IS
  NULL, which matches nothing on a healthy database.
- the fts_storage_version stamp, which re-wrote the same value on every
  open of an already-optimized database.

The connection is opened with timeout=1.0, so each blocked write costs a
full busy timeout while a sibling process holds the write lock, and the
open path's patience loop can ultimately give up and raise.

Gate all three on a read. Measured on a real 99 MB state.db (239
sessions, 6470 messages) with a sibling holding the write lock: 2117-2136
ms -> 3.6-7.1 ms. On an already-optimized database the unpatched open does
not merely stall, it raises `database is locked`; patched it completes in
3.7-6.8 ms. With a sibling running 200 ms write transactions in a loop
(n=20 opens): p50 894.6 -> 9.6 ms, p90 1094.7 -> 13.9 ms. A settled
database now issues zero main-database write statements to open.

The reads cost nothing measurable: the state_meta probe is a primary-key
seek (2.0 us), the messages probe is 1.7 us on the modern NOT NULL column
(unsatisfiable constraint, short-circuited) and 0.4-0.6 us at 300k rows on
a legacy default-less column via the partial index that already exists for
exactly this predicate. An uncontended open is unchanged.

Semantics are preserved. INSERT OR IGNORE still resolves the first-opener
race inside SQLite and racers still converge on the winner's token via the
re-read; the application_id gate and the PASSIVE-only checkpoint are
untouched; the heal is still considered on every startup, as #60108
deliberately made it, with only the write now conditional on a read
proving there is something to repair.

Read-first also fixes a correctness bug. Under contention the generation
block was abandoned by its `except sqlite3.Error` handler, so a process
ended up with no generation token at all even though the value was already
on disk and a plain read would have returned it -- and that token feeds
the deleted-WAL and replaced-file guards added by #101221. The heal's
`except OperationalError: pass` likewise skipped the repair silently, so
the unconditional form did not even deliver the unconditional repair it
advertised whenever it mattered most.

The probe deliberately does not use INDEXED BY: that hint raises
OperationalError("no query solution") against the modern NOT NULL column,
and the existing handler would swallow it, disabling the repair forever.
2026-09-11 06:21:00 -07:00
Teknium dbc5d7c60b fix(desktop): warmAgent goes through the guarded prewarm resolver too
host.warmAgent — the (connection, profile) sibling of warmProfile that
bot-row.tsx fires on pointerEnter for multi-source roster rows — still
dialed openGatewayForAgent directly, so a pointer sweep across a mixed
roster kept spawning at pointer speed past maxBackends on the registry
path even after warmProfile was guarded. Same bug class as #103631,
different door.

prewarmProfileBackend now takes an optional connectionId: the
active-profile no-op, the 60s throttle (keyed by the pool scope key) and
the pool-saturation skip apply unchanged, and the dial picks
openGatewayForAgent for a scoped source. One resolver owns every
speculative warm in the app; the real click still spawns on demand.
2026-09-11 06:20:30 -07:00
Abdulrahman Jahfali b23559877a fix(desktop): route host.warmProfile through the guarded prewarm resolver
Plugin rosters warm profile backends on pointerEnter with no dwell of
their own. warmProfile dialed openGatewayForProfile directly, bypassing
the pool-saturation guard, hover dwell, and per-profile throttle that
prewarmProfileBackend enforces for the built-in rail — so a pointer
sweep across a roster could spawn past maxBackends and leave the next
profile's real spawn queued until the 30s slot timeout, surfacing as a
profile surface that hangs forever while every other profile renders.

Delegate to prewarmProfileBackend so every speculative warm shares one
resolver and one policy, as the design guide requires. The real click
still spawns on demand; only the speculative head start is gated.
2026-09-11 06:20:30 -07:00
brooklyn! 8706517544 style(desktop): lint and format the salvaged titlebar files 2026-09-11 08:00:29 -05:00
abundantbeing 837e4b0942 feat(desktop): let Appearance choose left or right for titlebar app actions
Settings, Layout, and HUD default to the right so tabs keep the left
titlebar. Appearance has a Left/Right control for people who want the
previous left cluster.

(cherry picked from commit 7fe3175e475eb0ea81198bb69a0250662f940b17)
2026-09-11 08:00:29 -05:00
abundantbeing a09368fcd6 fix(desktop): pin settings layout and HUD back to the right titlebar
Keep panel tabs in the titlebar moved those app actions next to the
sidebar toggle, which ate the tab strip. Put them back on the right
edge. Sidebar toggle stays left. Fixes #107351.

(cherry picked from commit 7f4460a7f6028cf384506733a5bfa52273792d64)
2026-09-11 08:00:29 -05:00
Teknium 75a6fac052 fix(desktop): focus_pane un-minimizes the tree zone for every revealer
`revealDesktopPane` drove files/review/sessions/terminal through their
store setters only. Those are same-value no-ops when the pane's `$open`
already reads true while the user minimized its zone from the header
chevron, so the `focus_pane` tool reported success over an invisible
pane (#106009; class noted by @worryfreeaa). Route every tree-backed
revealer through `revealTreePane` after its own setter, which clears
`minimized` and fronts the pane.

(cherry picked from commit 690a1a75108ec63ce5cb1a638543969b0877dbaa)
2026-09-11 08:00:29 -05:00
hermes-seaeye[bot] aabad7b042 fmt(js): npm run fix on merge (#108217)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-11 12:59:42 +00:00
hermes-seaeye[bot] efca39279a fmt(js): npm run fix on merge (#108214)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-11 12:53:17 +00:00
Siddharth Balyan 0591da2ba6 Guided first launch: review fixes from #107985 and the free-tier chip badge (NS-848, NS-855) (#108211)
* fix(desktop): centralize guide handoff receipt reads

Resolve the guide receipt key and value together in setup-profile. Use the helper at all four read sites so connection scoping follows one implementation.

* fix(desktop): recover from unreadable handoff receipts

Memoize receipt reads and show Retry only for the error phase. Quarantine corrupt data before retrying, and resolve the guide identity when the failed request did not retain it so a fresh build can start.

Cover preservation of corrupt data and removal from the active receipt key with an invariant test.

* fix(desktop): validate persisted onboarding phases from one list

Derive OnboardingPhase and persisted-value validation from the same phase list so future phases survive relaunch. Verify every persisted phase reloads and an unknown value falls back to idle.

* fix(desktop): share window centering arithmetic

Extract centeredBounds and use it for onboarding boot and window growth. Keep the existing work-area clamps and coordinate rounding unchanged.

* fix(desktop): compute progress steps inline

Remove the ineffective ProgressCard memo because streaming flushes replace the messages array. Keep the same transcript scan and rendered steps.

* fix(desktop): center the free-tier status chip detail

Wrap the model label and sign-in badge in an inline flex span with a shared gap. This centers the badge beside the model text without changing other status-bar details.

* fix(desktop): derive the guide receipt key in one place

The Retry path spelled the key derivation out again because the read helper throws on a corrupt receipt before it can return the key. A separate guideHandoffReceiptKey serves both the reader and the quarantine, so the derivation has one home again.

* fix(desktop): keep the free-tier badge at its intended leading

Badge declares leading-none, but the class merger drops it behind the size variant's font-size class, so the badge inherits a 1.5 leading and renders 16px tall next to an 11px label. That height, not the inline alignment, is what read as a detached badge. Restating leading-none on the chip's badge brings it to 11.6px, inside the label's cap height. The Badge component itself is left alone; every other badge in the app has the same dropped leading and that is a separate decision.
2026-09-11 12:46:48 +00:00
brooklyn! dee30d123d fix(desktop): scope approval hints and omit zero message counts 2026-09-11 07:38:41 -05:00
brooklyn! b08a26791f fix(desktop): focus opened sessions and restore closed tab positions 2026-09-11 07:38:41 -05:00
brooklyn! 22b4b49aa7 fix(desktop): preserve composer selection across model picking 2026-09-11 07:38:41 -05:00
brooklyn! e6aa2e9fe4 fix(desktop): satisfy Electron permission type check 2026-09-11 07:37:59 -05:00
YuhGuan 244a42f636 fix(desktop): media-range review follow-ups — ignore multi-range as a whole, fstat the handle being streamed, 404 for missing files
- Multi-range requests (any comma) now fall back to a full 200 instead of silently serving only
  the first part (RFC 7233 permits ignoring Range); documented + pinned by a test.
- Open the file first and fstat that handle, then stream from the same FileHandle, so
  Content-Length and the bytes delivered come from one open file (no stat/stream race).
  Handing the FileHandle (not the raw fd) to createReadStream avoids a double close.
- ENOENT/ENOTDIR return a 404 Response instead of rejecting; covered by a test.
2026-09-11 07:37:59 -05:00
Youssef 360c1ee836 fix(desktop): allow HTML5 video fullscreen through permission handlers
The custom setPermissionCheckHandler only allowed media/audioCapture/
videoCapture, which made Electron deny the 'automatic-fullscreen'
permission consulted during HTML5 video requestFullscreen(). The
request handler's isMediaCapturePermission() also returned false for
'fullscreen'. Result: the native fullscreen button on <video controls>
in chat silently did nothing.

Allow 'fullscreen' + 'automatic-fullscreen' in both handlers.

Verified with a minimal Electron repro using Hermes' exact handlers:
requestFullscreen() failed with 'TypeError: Permissions check failed'
before; works after. User-verified in the packaged desktop app.
2026-09-11 07:37:59 -05:00
brooklyn! f36b6b0ce6 fix(desktop): cancel bootstrap manifest work during quit 2026-09-11 07:33:17 -05:00
brooklyn! 7b9808e43b fix(desktop): complete quit through one bounded teardown barrier
Let settled remote sessions continue their first quit. Fence late local starts and join existing local and SSH drains without cancelling managed update recovery.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>

Co-authored-by: ChanPark03 <parkchan0302@gmail.com>
2026-09-11 07:33:17 -05:00
brooklyn! 8607d6a52a fix(desktop): retain and cancel owned backend lifecycle work
Track pending starts, pre-claim children, and teardown removed from routing. Bound cleanup and cancel setup/update waits while preserving settled remote descriptors.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>

Co-authored-by: ChanPark03 <parkchan0302@gmail.com>
2026-09-11 07:33:17 -05:00
brooklyn! 556d19c2d5 fix(desktop): preserve refresh errors in media and verify auth recovery over HTTP
Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>

Co-authored-by: Zeus-Deus <github.commits@widow.cc>
2026-09-11 07:32:34 -05:00
brooklyn! 22751c8fd9 fix(desktop): preserve remote auth through refresh failures and login races
Salvage native refresh coordination and cookie fallback without losing forced bearer rotation or replaying REST mutations.

Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>

Co-authored-by: Zeus-Deus <github.commits@widow.cc>
2026-09-11 07:32:34 -05:00