`run_tests.sh` defaults to twice the core count, and the value this branch
started with came from a rule of thumb of 1.5x cores plus a measurement on a
16-core machine. A sweep on the real runner disagrees with both.
Run 32549672063 on the 96-core runner (EPYC 7763, 377GB) timed the whole suite
at six worker counts, two repetitions for each. A warmup run came first, and
retries were off:
workers x cores rep 1 rep 2 mean
48 0.5x 138s 139s 138s
96 1.0x 127s 126s 126s <- fastest
144 1.5x 130s 134s 132s
192 2.0x 132s 133s 132s
240 2.5x 140s 139s 140s
288 3.0x 143s 142s 142s
One worker for each core wins. Both repetitions agree on the order.
The shape is the more useful result. The range is 126s to 142s across a 6x
range of worker counts. The suite has sufficient concurrency at this machine
size, so nothing above the core count buys anything. The remaining time
belongs to the slowest individual files and to the setup. A future gain must
come from those, and not from this number.
The sweep ran from a temporary workflow that this branch does not keep.
The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.
1. Every pytest subprocess shared one temp root.
pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.
scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.
Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.
2. The config read guard walked directories that other tests were writing.
tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.
The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.
The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.
3. A PTY test waited for a file to exist, and not for its content.
tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.
The test now waits for the expected bytes, with a bounded deadline.
Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.
4. A dialog close timer outlived the test that started it.
ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:
ReferenceError: window is not defined
at resolveUpdatePriority (react-dom-client.development.js:1308)
at dispatchSetState
at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574
The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.
Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.
Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.
Verification:
- The affected Python files and the tests of the runner itself pass under
scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
"window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
before this change. The two cleanup effects carry an eslint-disable line for
the ref-mirror rule. They write a timer handle, and not a mirror of a
reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
no python3 outside the nix store, and the test uses the literal `python3`.
The child exits 127 there. The fix rests on the 25-run measurement above and
on CI.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
Addresses @helix4u's review on #92020:
- Consent notice now matches the real --nous contract: full logs up to
512KB each, likely conversation content/tool outputs/file paths, viewable
by Nous staff AND allowlisted Discord moderators (all 5 locales).
- Client-supplied text (error_context + extra_files) rides _redact_log_text
— the same upload-safe redactor as backend logs (secrets + email masking),
not the weaker bare secret pass; regression test covers both.
- ok:true without view_url or id becomes a structured failure; a returned
id without a link renders an upload-ID fallback the user can quote.
- Generation guard in the store: dismissal is immediate in every phase
(incl. mid-upload); a stale completion can no longer resurrect or
overwrite the dialog. Cancel button never disabled.
The old contract WAS the bug (#70337): exe deleted by the swap, then
rebuilt from scratch. With the release-dir graft the exe survives the
swap; the test now asserts survival + original bytes.
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.
- ZIP fallback now keys on git ACTUALLY having failed
(_should_zip_fallback_on_update_error): a dependency-install failure
after a successful pull can't be fixed by re-downloading source and
would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
(-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
Chromium's native selection copy serializes the selection as text/html
with every element's computed color inlined. Copied from a dark theme,
body text lands on the clipboard as near-white (the app ink computes to
color(srgb 0.902 0.929 0.953 / 0.94)); pasted into a light-background
target such as an email, it is invisible.
The renderer never writes rich text itself, so this payload can only
come from Chromium's serializer — which runs after copy handlers decline,
meaning clipboardData reads back empty inside the event. The new guard
therefore decides from the live DOM: it scores the computed ink of the
selected text against the rendered theme mode, and only when they are
opposite schemes does it own the payload, writing text/plain plus a
tag-structured text/html with no paint declarations.
Structure (headings, lists, tables, links, bold/italic, code layout)
survives; colors come from the paste target's defaults. A generic
font-family anchor (sans-serif, monospace inside code) keeps receivers
that convert HTML to rich text on their own compose font instead of the
Times browser default. Same-scheme copies and selections starting inside
editable fields pass through untouched.
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.
De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
Diagnostic run showed the venv shim makes every spawn a launcher/worker
chain: the child's direct parent is its own launcher (python.exe
child_scan.py), and the gateway-argv process is the grandparent. The
probe now finds the gateway ancestor by argv — the same way the pause
machinery would — and asserts THAT pid is visible to the scan.
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.
#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.
15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
On-demand workflow (fires only on wine2e/** pushes, never on PRs/main)
that runs a live venv-holder E2E on windows-latest: real spawned
processes with Hermes argv shapes, real detection/classification/
message code against the live process table. Tests pin CORRECT behavior
for the cluster issues (#90778 mislabeling, #78089 long-path exemption,
#87594 ancestor-exclusion, #81774 serve premise), so unfixed bugs fail
on the runner — empirical premise-check before the consolidation fix.
The composer middleware is now identification-only: it resolves the
user's @tags against the live roster and annotates the draft with who
they refer to (profile, friendly title, device for cross-connection
rows). The agent decides whether to contact them and does it through
its message_agent tool — one send path, composed messages only.
Deleted the renderer's entire parallel delivery transport:
deliverRemoteRosterMentions / pollRemoteDmReply /
ensureRemoteCanonicalChat and the injected shellout instructions
('[@mention handoff — run hermes -p …]' and 'Desktop is delivering …
over Connections'). This retires the whole invocation bug class at the
source instead of sanitizing it: no verbatim user text is ever
forwarded by the renderer (#91397), and no shell command is ever
composed from prompt text (#91304, #91339 shape).
Tests: mention-identification.test.mjs replaces the two delivery-era
files — identification note shape, no-shellout/no-delivery containment
(sabotage-verified: re-adding a renderer delivery call fails 2 tests),
poisoned-title inertness, pass-through for unknown @s, and a source
contract pinning the deleted machinery. hide-bots + roster-cache-key
harnesses re-pinned to the new contract. 390/390 green.
The two waitForHermesReady cloud-503 tests froze now() at 0, so the
readiness loop never crossed its deadline — the vitest electron project
hung for the full 20-minute CI budget. Advance the clock per poll like
the sibling readiness tests do.
Follow-up on the #85373 salvage: the portal and Discord URLs move out of
the localized hint prose into dedicated action buttons (URLs live in code,
translations can't drift them), matching the layered error card's
action-row idiom from #91493. Overlay test updated to the button contract;
all five locales updated.
eslint --fix output: blank lines before statements and the import-order
spacing in connection-config.test.ts that the check:lint gate rejects.
Formatting only — no logic change.
The electron boot path now classifies a Nous Cloud 502/503/504 at both the
OAuth ticket-mint and readiness boundaries and carries isCloudBackendDown /
statusCode through DesktopBootProgress, but the renderer never consumed the
structured signal — a cloud-backend failure fell into the generic remote-
failure recovery copy.
Make BootFailureOverlay branch on isCloudBackendDown: lead with the
cloud-specific title/description, drop the local-only Repair action, and
surface the actionable portal / Local-mode / Discord guidance (the electron
factory's full message is still shown in the error box).
Adds the cloudDown i18n keys (en + ar/ja/zh/zh-hant) and a regression test
asserting the cloud-down recovery renders and Repair is dropped.
The original implementation classified 502/503/504 only inside the readiness
loop, but for OAuth-backed Cloud connections the WebSocket-ticket mint runs
before waitForHermesReady. A server fault there was wrapped by
gatewayTicketFailure into a generic message and the Cloud-down classifier was
never reached. This closes that boundary and fixes a latent regex defect.
- isServerSideHttpError: structured-first (err.statusCode for 502/503/504),
legacy 'NNN:' prefix as fallback, non-Error inputs rejected. Also fixes the
committed '\d' (double-escaped, matched a literal backslash) that made the
function never detect a status prefix.
- makeNousCloudBackendDownError: single factory for the actionable Cloud-down
error (isCloudBackendDown/statusCode/detail/cause), shared by both the
ticket-mint boundary and readiness exhaustion.
- main.ts: run the Cloud classifier at mintGatewayWsTicket before the
gatewayTicketFailure wrap; 401/403 still route to reauth.
- connection-config.ts: gatewayTicketFailure preserves an integer statusCode
from the source error; auth semantics unchanged.
- boot-progress/IPC: carry isCloudBackendDown and statusCode through
DesktopBootProgress so the renderer overlay (a PR-body promise) can key on
the structured result rather than re-classifying the message string.
Tests: backend-health (structured detection, non-Error rejection, factory
shape/cause/guards, legacy fallback), connection-config (statusCode preserve,
401/403 reauth, integer-only copy), and an OAuth ticket-mint integration
regression (Cloud 503 -> actionable Cloud-down; 401 -> reauth). Connection-
config suite 80/80 green; backend-health sync tests green; the async readiness
loop tests cannot run on this host (pre-existing local-run limitation) and are
the CI gate. PR #85373 (#85335).
When a Hermes Desktop connects to a Nous-managed cloud agent
(*.agents.nousresearch.com) and that backend returns HTTP 502/503/504,
the previous error message was the opaque generic 'Hermes backend did
not become ready: 503: ...' with no guidance that the cloud server
itself is down.
Add isServerSideHttpError and isNousCloudAgentUrl helpers and use them
in waitForHermesReady to detect this exact scenario. When triggered,
throw an error with the hostname, status code, and recovery paths:
check the Nous Portal, switch to Local mode, or reach out on Discord.
Also adds a isCloudBackendDown flag and statusCode property on the
thrown error so the renderer overlay can render specialized UI if desired.
Mirror the credits_tracker caveat in _is_free_model (a paid stealth/
model would bypass the free_only gate and paid-lane warning) and fix
the stale _warn_paid_lane_once docstring.
- credits_tracker: trim inline comment block (duplicated docstring) and
correct its safety claim - a paid model under stealth/ would fail
closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
recognize stealth/ prefix (same bug class as #91843: free_only=true
wrongly skipped the OpenRouter fallback and the paid-lane warning
fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
my-stealth/model not).
Stealth-preview SKUs (e.g. stealth/ox-alpha) are free-tier but carry no
:free suffix, so is_free_tier_model() returned False for them. On gateway
sessions (which never run the model picker's pricing fetch), the free-model
suppression of the credits.depleted banner never engaged, and any response
carrying paid_access:false triggered a false "Credit access paused" notice.
Add stealth/ prefix detection to is_free_tier_model() as a zero-network
signal, same design as the existing :free suffix check. Fail-open to
False (banner still shows) if the prefix changes — recoverable noise,
never a masked depletion on a paid model.
Closes#91843
Two leftovers from the #91852 descope (integrity-check tests removed but their
scaffolding stayed):
- Drop the now-unused `import pytest` (orphaned when the verify_state_db_integrity
tests that used it were removed; no markers/raises/fixtures remain in this file).
- Rename the `# Defect 1:` section label to just `# Repair-path write durability`
— the sibling "Defect 2" section was descoped out, leaving the numbering dangling.
Test-only, no behavior change. tests/test_state_db_write_durability.py: 4 passed.
Follow-up to the salvaged repair-durability commit. Scope corrections so this
PR ships only the reachable, non-competing, WAL-mode-correct half:
- Drop verify_state_db_integrity() + its 4 tests. Zero production callers here
(dead code); the caller lives in the follow-up that wires it into
SessionStore._open_session_db_for_active_scope() (PR #91754). The function
moves with its wiring.
- Drop the _db_fingerprint change (size:mtime_ns -> dev:ino:size) + its 3
ledger tests. This is competing work: PR #88425 (salvage of @jirathip-k's
#88224) already fixes the same size:mtime_ns budget-reset bug with a
content-sample + volatile-header-mask that also handles the DELETE-mode
commit-counter case, and carries @jirathip-k's diagnosis/credit. Landing a
second, divergent fingerprint contract would stomp that lineage. Fingerprint
stays with #88425; this PR reverts _db_fingerprint to main's form.
- Mark test_repair_refuses_while_another_connection_holds_the_db requires_wal.
_live_writer_holds_db detects an out-of-process holder via the WAL-index
exclusive lock, absent in journal_mode=DELETE (used on WAL-reset-vulnerable
SQLite <3.51.3 incl. CI's 3.50.4, and on NFS/SMB). The test failed there;
the conftest requires_wal gate auto-skips it. DELETE-mode limitation is now
documented on the guard docstring: repair is serialised only by the
cross-process repairer lock there. The reported incident was in WAL mode.
- Map dhanesh@users.noreply.github.com -> dhanesh (contributors/emails) so the
attribution CI gate passes.
Net: this PR is repair-connection durability barriers + the live-writer guard.
addresses @andrexibiza's #90747 review (dead-code verifier + fingerprint
interlock with #88425).
state.db corrupted twice in two days with the torn-b-tree signature —
repeated "2nd reference to page", "Rowid out of order", and long runs of
"never used" pages in messages (rootpage 5) and idx_messages_session.
macOS fsync() guarantees neither data-on-platter nor write ordering, which
_enforce_macos_synchronous_full already documents: a rewrite interrupted by
process or OS termination leaves half-written b-tree pages. The mitigation
is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was
applied only through apply_wal_with_fallback(). The repair path opened
state.db with a bare sqlite3.connect() six times and then ran REINDEX,
VACUUM and writable_schema surgery through it — the operations that rewrite
nearly every page of the file — with no barrier at all.
- _connect_repair_durable() routes every repair/probe connection through the
barriers. Applying them is best-effort by necessity: SQLite loads the
schema before any statement, so on a malformed schema even
PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is
precisely this helper's input. _reapply_durability_barriers() retakes them
before REINDEX and VACUUM, once the schema parses and they can stick.
- verify_state_db_integrity() adds the proactive check that was missing.
Repair only ever ran reactively, after a caller already hit a malformed
error, so a database torn in pages no query happened to touch stayed live
and kept accepting writes. On 2026-08-19 that gap was 11 hours across two
restarts that both reported a clean start. Size-aware: degrades to an O(1)
probe above 2 GiB rather than pegging a CPU at startup.
Also restores two fixes lost when `hermes update` reset the tree to
origin/main before they were committed:
- _db_fingerprint keys the repair ledger on dev+inode+size instead of
size+mtime_ns. The old form was justified as "stable for a file nothing
can successfully write to"; that premise is false, because on FTS
corruption this module deliberately keeps canonical writes enabled with
FTS detached. mtime churned on every write, so each pass re-keyed the
ledger and reset the counter to 1 — the cap could never be reached and the
damaging surgery could retry forever.
- _live_writer_holds_db() refuses surgery while another connection holds the
database. The cross-process lock only serialises repairers against each
other; it says nothing about the gateway, Desktop or a CLI. Rewriting
b-tree pages under a concurrent writer is what spread the 2026-08-18/19
damage out of the FTS shadow tables and into the canonical ones. Fails
open, so it cannot strand the self-heal path it protects.
The guard's own tests built a two-table toy schema, so every repair aborted
on "no such table: sessions" before reaching the guards under test — the
assertions were passing over a code path that never ran. They now build
through a real SessionDB.
Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure.
Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>
text.splitlines() returns [] for empty strings. Accessing msg_lines[0]
then raises IndexError, making session resume crash when the session
contains a message with empty or whitespace-only text (e.g. reasoning-only
turns, tool-only assistant messages).
Guard with `or [""]` in all three branches (user, assistant_last,
regular assistant) so an empty message renders as a blank line.
Fixes#59265
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
The cmdline fallback was matching every system daemon with an
unreadable fd table (init, systemd-journald, dockerd, etc.), causing
FTS rebuilds to be skipped on every Linux system. Add _looks_like_hermes
filter so only processes whose cmdline contains Hermes markers are
flagged — matching @jackulau's suggestion of 'uninspectable AND
identifiable as another Hermes process.'
Address review feedback from @jackulau on PR #90871:
1. psutil.open_files() silently drops '(deleted)' WAL sidecar entries
on Linux because isfile_strict() stats the literal path including
the suffix and fails. Switch to direct /proc/<pid>/fd readlinks
which preserve the '(deleted)' suffix so _canonical can match.
2. psutil.process_iter() converts AccessDenied to None, which
or-() skips silently — the fail-closed branch never runs. For the
root-gateway vs user-desktop topology in the issue, the fd table is
unreadable but /proc/<pid>/cmdline is world-readable. Add a cmdline
fallback that flags uninspectable processes.
Also keep the psutil path for macOS/BSD (no '(deleted)' convention).
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Sessions running on provider 'nous' get a 'Nous support' action on the
failed-turn card, opening the portal help hub
(https://portal.nousresearch.com/help — docs, Discord, GitHub) in the
external browser. All five locales + docs updated.
useNavigate() throws outside a <Router>; streaming.test.tsx renders the
thread bare. Move the Settings deep-link into a SwitchProviderAction child
gated on useInRouterContext(), which is safe in any tree.
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.
Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.
Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).
The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
Adapts the in-place branch update from PR #89507 (@willfrombr) onto the
switch-by-default behavior: the deterministic switch path remains the
default so non-interactive updates (desktop, gateway, cron) never dead-end
on a merge conflict, and deliberate custom-branch users opt in with
updates.parked_branch_strategy: update_in_place. --switch-branch overrides
the in-place strategy for one run (deep feature branches that must not
accumulate update merge commits). Docs + config comments + tests cover
all three routes.
Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.
--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.
Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.
Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>