The HUD has no in-app browser, so a click tried to paint a webview
into the transparent overlay (OAuth). Hand those links to the OS,
mount the context menu, and skip preview-tile docking.
Ignore-mouse cannot restore on X11, so a visible band that still has
pointer-events:none swallows clarify options and links. Held prompts
and solid-window bands now take the pointer without composer focus.
- Anthropic classifier counts signature_delta and citations_delta payloads
(content-bearing delta types the transport emits — relay_llm.py handles
both) so signed-thinking/cited-text generation keeps ticking the fence.
- Fix the stale 'per streamed event' comment above the Anthropic
on_stream_event lambda (left over from the conflict resolution).
- Rename test_completed_response_without_stream_payload_does_not_tick to
test_completed_response_ticks_only_terminal_signals — the old name
contradicted its own assertion (dispatch + shim ticks are expected).
Follow-up for the salvaged substantive-progress fix, adapted to main's
three-hook architecture:
- Keep the dispatch tick and the completed-response shim tick (both
deliberate on main; one-shot terminal events that cannot defeat an
inactivity timeout) — update the two PR assertions accordingly.
- Pin the end-to-end bug: keepalive/empty-role chunks must leave
CompressionCommitFence stale, a substantive token must refresh it.
- Pin the fast-lane telemetry contract (#96945/#96963):
time_to_first_progress_ms records on the first frame of any kind via
the split-out _notify_aux_timing_response helper.
Second follow-up for salvaged PR #94547, folding in review findings from the
duplicate-PR cluster (#47015, #55054, #58476, #72977, #73685 all fix the
same 401) and the sweeper review of #73685:
- Replace the dot-anchored suffix predicate with exact-match against the
existing _ALLOWED_TEAMS_SERVICE_HOSTS allowlist (two of the five
duplicate PRs converged on this independently). Any Azure customer can
register <name>.trafficmanager.net profiles, so suffix matching was not
safe. Also requires https on the default port — :444 on an allowlisted
host no longer receives the bearer (sweeper finding on #73685).
- Stream _fetch_attachment_bytes through _read_httpx_body_with_limit
instead of buffering response.content — the shared inbound media cap now
applies to authenticated downloads too (sweeper finding: a lying
Content-Length must not OOM the gateway).
- Serialize token refresh with a lazily-bound asyncio.Lock so concurrent
attachments share one STS POST (review finding on #94547).
- Token expiry now uses time.monotonic() (from #55054) — wall-clock jumps
can't extend a stale token.
- Tests updated: exact-allowlist predicate (lookalike/subdomain/port/scheme
negatives), streaming fake client, and a concurrent-cold-cache lock test.
Mutation-checked: suffix match, silent drop, no-lock, and unbounded
buffer each fail a test.
Follow-up for salvaged PR #94547 (Sibbern's Bot Framework attachment auth):
- The host check used bare endswith('trafficmanager.net') /
endswith('botframework.com'), which matched attacker lookalike hosts
(evil-trafficmanager.net) — sending the bot's bearer token off-platform.
Same threat model _ALLOWED_TEAMS_SERVICE_HOSTS already documents. Now a
single dot-anchored predicate, _is_botframework_attachment_host(),
used by both _fetch_attachment_bytes and the _on_message image branch
(was copy-pasted in two places).
- BF images whose bytes fail image validation were silently dropped with
no log (cache_media_bytes returns None; old path logged). Add the
missing else-warning, mirroring the document branch.
- Record cached_m.media_type instead of the raw content_type so the
MessageEvent MIME matches what was actually cached.
- Init _bf_token_cache in __init__ (was masked by getattr).
- Tests: host predicate dot-anchoring (attacker lookalikes blocked),
BF routing vs generic helper, bearer attach + attacker-host block end
to end, token acquire + cache reuse, token-failure degradation, and
the silent-drop regression guard. Mutation-checked: reverting the
dot-anchor, the else-warning, or the token cache each fails a test.
Inline/pasted images in Teams arrive with a contentUrl on
smba.trafficmanager.net (/v3/attachments/...). Unlike file uploads
(pre-authenticated SharePoint downloadUrls), these connector URLs
require the bot's own bearer token; fetching them anonymously fails
with 401 Unauthorized and the image is silently dropped.
- add _get_botframework_token(): client-credentials token for
https://api.botframework.com/.default, cached until ~5 min before
expiry
- _fetch_attachment_bytes(): attach the token when the attachment
host is *.trafficmanager.net / *.botframework.com; SharePoint and
other URLs remain auth-free as before
- route image/* attachments with Bot Framework contentUrls through
the authenticated fetch instead of cache_image_from_url (which
sends no Authorization header)
SSRF guards unchanged. Verified on a live Teams personal-scope bot:
pasted images previously logged '[teams] Failed to cache image
attachment: 401 Unauthorized' and now cache and deliver correctly.
Reviewer findings from the formal /simplify-code pass, all verified:
1. Scanner blind spots (HIGH): five more locked pure readers were
invisible to the v1 gate — get_compression_fallback_streak and
get_compression_ineffective_count hid behind `conn = self._conn`
aliasing; list_gateway_sessions, find_session_by_origin and
search_sessions hid behind SQL held in variables/f-strings (the
scanner required unknown == 0 to flag). All five converted to
_read_ctx(); the scanner now (a) tracks self._conn aliases and
(b) flags lock blocks with NO proven write instead of silently
skipping unprovable SQL. Sabotage self-check extended to pin all
three detection classes (literal, alias, variable-SQL) plus a
mixed variable-SQL writer that must NOT fire.
2. get_meta reverted to the writer lock: its inline comment (present
on main) documents a real read-your-writes dependency —
fts_rebuild_step reads rebuild progress that a pooled WAL reader
cannot see mid-transaction. The blanket conversion had overridden
a documented design decision; it is now the single justified
_ALLOWED_LOCKED_READERS entry, replacing the dead
_enter_fts_fail_open entry (whose lock block counts 3 writes and
never needed allowlisting).
Strengthened scanner on pre-conversion main: 43 violations
(39 pure-read + 4 no-proven-write). This branch: zero.
Suites: gate 2/2; tests/test_hermes_state.py + tests/state/ 331
passed (same 3 pre-existing FTS-rebuild reds as clean main); 433
passed across the 12 consumer suites of the five newly-converted
methods (compression anti-thrash, session search, status, scheduler).
Pattern-C architectural fix (write-lock contention), completing what
#90734 started: that PR fixed the four UNLOCKED readers racing the
writer connection; this one fixes the 39 LOCKED pure readers convoying
every concurrent turn's persistence behind the global writer lock, and
adds the gate that stops the class from re-entering.
The gateway shares ONE SessionDB across every agent. A read-only query
under `with self._lock:` blocks all concurrent writers for its
duration; under WAL, _read_ctx() serves the same read from a pooled
read-only connection with no lock at all (non-WAL falls back to the
locked writer byte-for-byte, so DELETE-journal installs are unchanged).
Converted (SELECT-only bodies, mechanical `with self._lock:` →
`with self._read_ctx() as conn:`): gateway routing loaders, session/
message counters, titles, compression tip/lineage/cooldown readers,
telegram topic bindings, prune candidate scans, meta readers, resume
resolution — 39 methods, verified per-method that every statement is a
SELECT and every conn use stays inside the with-block. Read-modify-
write methods and checkpoint/maintenance PRAGMAs stay on the writer
lock (their read is ordered against their own write).
Gate: tests/state/test_no_locked_readers_gate.py — an AST scanner over
SessionDB that fails CI when a pure-read method body takes the writer
lock, with a sabotage self-check proving the scanner fires. On
pre-conversion main it reports 38 violations; on this branch, zero.
Measured (2 writer threads + 3 reader threads, 3s, WAL):
before: 78.5k reads, p99 2.49ms, max 160.7ms (readers convoy)
after: 200.4k reads, p99 0.53ms, max 25.5ms (2.6x throughput,
4.7x better p99, 6x better tail; writes unchanged)
Honest caveat: on runtimes where WAL is refused (the currently-bundled
SQLite 3.46 trips the WAL-reset-vulnerability gate → journal=DELETE),
_read_ctx() falls back to the identical locked-writer path and this
change is behavior-neutral by construction; the win applies to WAL
installs (legacy WAL DBs, fixed runtimes, wal-configured operators).
Suites: tests/test_hermes_state.py 243 passed; combined state sweep
540 passed — the only 3 reds are the pre-existing
test_fts_runtime_rebuild failures, verified failing on clean
origin/main before this change.
ColorDepth is lazy so MagicMock prompt_toolkit stubs can still import
cli. The redraw/resize kitty re-queue no-ops when the pet pane was
never initialized.
Reuse kitty Unicode placeholders plus after_render write_raw so
prompt_toolkit's screen-diff can host the same crisp sprite as the TUI.
Re-queue the transmit after Ctrl+L / resize so the image comes back.
Co-authored-by: Sam Foreman <saforem2@gmail.com>
The rename and settings dialogs stay mounted while closed, so they render
with a null board on every pass. Their mutation callbacks read `board!.slug`,
and the React Compiler lifts a callback's property reads into its render-time
dependency check — so the read escaped the closure and dereferenced null
immediately on mount, taking the whole contribution down behind its error
boundary.
The non-null assertion never guarded anything; it erases at compile time.
Resolve the slug null-safely in the component body instead, which is also
the form the compiler can hoist safely.
Only the bare-lambda shape is affected: the inline `useMutation({ mutationFn })`
this replaced memoized on the whole `board` object and kept the read inside
the closure, so the regression arrived with the extraction into
`useBoardWrite`, not with the feature.
Per-item menus across the app say "Rename", "Delete", "Export",
"Archive" — the row already names what you are acting on. A handful of
places had drifted to verb+noun or Title Case, so the same action read
differently depending on where you found it.
Sessions and projects now say "Rename…" like profiles and the file tree
already did. The per-profile context menu says "Export…"; the noun stays
on the profiles-list button and the native file-dialog title, which
stand alone. Bots drop "Delete Group" and "Edit Profile" for "Delete"
and "Edit…". Title Case gives way to sentence case in the file menu,
review tree, and model menu.
Nouns are kept wherever they carry weight: dialog titles, icon-button
tooltips, "Remove worktree" (its menu also has a plain "Remove"), and
"Open Bot Chat", which names the canonical session titled exactly that.
The board switcher could create and configure boards but not move,
rename, or remove one. Rename technically existed, buried as a field
inside "Settings…", which is why it read as missing; it now has its own
entry and the settings dialog is left owning scope alone.
Delete archives rather than erases — the board's directory moves to
boards/_archived/ and the toast names the path — and never appears for
`default`, which the backend refuses to remove.
The three dialogs had grown three copies of the same shell, the same
"invalidate the list and close" mutation tail, and the same name field,
so those are shared now instead of parallel-implemented.
Plugins had no sanctioned way to ask for a file path — the OS door
carried notify, openExternal, revealPath and writeClipboard, so anything
needing a dialog had to reach around the SDK for window.hermesDesktop.
pickSavePath and pickOpenPath wrap the existing selectSavePath /
selectPaths IPC with the door's usual contract: resolve null when the
bridge is missing or the user cancels, never throw at the plugin.
POST /boards/{slug}/export and POST /boards/import, so the desktop and
dashboard can drive board transfer. Both exchange filesystem paths
rather than bytes, the same contract profile export/import uses and for
the same reason: the client runs its native save/open dialog on the
machine that hosts the backend, so a path is all either side needs, and
a board carrying a few hundred megabytes of attachments never has to
cross the renderer heap.
Also covers the pre-existing rename (PATCH) and delete endpoints, which
had no tests — including that delete archives to a restorable directory
and refuses to touch `default`.
`hermes kanban boards export|import` moves a board between machines:
tasks, comments, links, history, and attachments in one .tar.gz.
Two things make this more than a tar of the board directory. The
database is live — kanban runs in WAL mode, so a filesystem copy loses
whatever still sits in the -wal sidecar and tears if the dispatcher
commits mid-copy; export goes through SQLite's online-backup API
instead. And rows carry machine-local state: claims, worker PIDs,
absolute workspace and attachment paths, session ids, and the gateway
chat ids subscribed to task events. Shipping those verbatim is how an
imported board arrives holding a claim owned by a process on someone
else's laptop, or starts pushing task events into a stranger's Telegram
thread. Everything machine-local is stripped on export and re-stripped
on import, since an archive is untrusted input.
Imports always land as a NEW board, auto-suffixing the slug on
collision, so an import can never merge into or overwrite a board that
is already there. Tasks whose workspace was a directory or git worktree
on the source machine are parked in triage rather than left for the
dispatcher to claim and burn into the failure breaker.
Profile export/import owns the only hardened tar handling in the tree:
GNU-format writing (PAX fractional mtimes make macOS Archive Utility
throw "Error 94"), plus an extractor that rejects absolute paths, `..`
components, and non-regular members.
Kanban board transfer needs exactly that, and a second copy is how the
weaker of two extractors eventually ships. Move the four helpers to
hermes_cli/archive_safe and point profiles at them; no behavior change
beyond dropping a provably-unreachable fallback in the root-listing
helper, whose condition is a strict subset of the comprehension above it.
Live A/B eval (old flat vs operations[] on qwen3.8-27b / gpt-5.6-terra /
claude-sonnet-5) caught a real regression: on a fuzzy-match miss the flat
path returns file_preview so the model can self-correct, but the batch
wrapper rebuilt the error dict and dropped every field except error/
failed_index. Sonnet, recovering blind, probed the file by writing and
reverting placeholder patches for 8+ turns (50k tokens vs 15k on the flat
arm). Batch failures now merge through all non-error fields from the
failing op's result.
An unreadable root self-heals on a 3s timer, so the probe runs for as long as
the pane is open. Every forced reload cleared `rootError`, emptied `data` and
dropped `resolvedCwd` before reading, so each tick rendered "unreadable" →
blank → "unreadable" and flickered the header's project name with it. A local
ENOENT resolves well inside the 180ms skeleton delay, so the gap paints as a
bare blank frame rather than a loading state.
Re-reading the same root now probes underneath what is on screen; only a
different root, or the same path from a different backend, clears first.
An unnamed `session.info` was treated as describing whatever the pane had
selected. The gateway stamps `stored_session_id: session_key or ""`, so every
not-yet-persisted session emits one, and `broadcast_session_info` / the
approvals loop re-emit for every live session at once. An unscoped event
applies exactly when no session is active, so with nothing selected each of
those repointed `$currentCwd` and claimed it for the null selection — the file
tree, coding rail and statusbar painted a folder no selected conversation
owned, until the next `releaseWorkspaceCwdOwner` dropped the claim and they
un-painted it.
Require the event to be bound to the pane's own runtime before an absent id
reads as the selection. The case the allowance exists for — a lazy session that
is the pane's runtime but is not persisted yet — still adopts and owns its cwd.
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)
* ci: retrigger after runner startup_failure on rerun attempt
Composer drag added renderer CSS-pixel deltas onto a window AppKit
clamps to the current display, so the bar could not follow the cursor
onto a second monitor (and drifted on mixed-DPI Windows). Track the OS
cursor in main and lift that clamp.
Co-authored-by: Biotrioo <biotrioo@protonmail.com>
The JS alignment suite pinned 24.0.0 as a supported Node; with the
engines arm raised to ^24.11.0 (babel 8 requires >=24.11), 24.0.0 and
24.10.x are now correctly rejected and 24.11+/24.18+ accepted.
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.
- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
(install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
locked dependency's engines.node, and the installer gates must encode
the same floors as the manifest — so the next babel-style floor bump
turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:
error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.
Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate
* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing
* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal
* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback
* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization
Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.
1. The watcher polls only the ROOT store.
`_handoff_watcher` resolves `self._session_db` with no profile scope, which
always yields the root `state.db`. But `/handoff` run under
`hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
store. Nothing ever reads it, so the CLI times out after 60s while the
gateway is alive and connected. The watcher now iterates
`[(None, None), *secondary_profiles]` and polls each inside
`_profile_runtime_scope`.
2. The destination session key is built without the profile namespace.
`_process_handoff` called `build_session_key()` with no `profile=`,
producing `agent:main:...` while that profile's own adapter routes organic
inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
reads.
3. Delivery uses the PRIMARY profile's adapter and config.
`self.adapters` holds only the default profile's adapters (secondaries live
in `self._profile_adapters[name]`) and `self.config` only the default's home
channel. A secondary profile's handoff was therefore sent by the wrong bot,
to the wrong chat, while persisting the right session key and reporting
`handoff_state='completed'` — a false positive that looks correct in the
database and is wrong on the wire.
Two robustness fixes in the same path:
4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
turn plus delivery; awaiting it inline meant one slow handoff stopped the
watcher from even polling the other profiles. With the CLI's 60s deadline, a
valid handoff could time out purely because another profile's was ahead of
it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
never claimed twice, and a bounded drain on shutdown.
5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
one in-process dispatch, so a row still in that state at startup belongs to
a gateway that died mid-dispatch. It can never reach a terminal state, and
`request_handoff` refuses new requests unless the state is
NULL/completed/failed — that session could never hand off again, silently.
`reclaim_stale_running_handoffs()` now fails those rows once per store at
watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
may already have switched the session key and dispatched, so a blind retry
risks double delivery.
Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.
Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
FIRST_PAINT_BUDGET 20 + BACKFILL_STEP 60 prepended the rest of a 600-unit
page across ~10 visible commits. A 290-unit step keeps the interruptible
commits and removes the strobe.
Brand-new drafts are empty on purpose. A routed session the list already
knows has messages must not drop the loader just because a runtime id is
bound — that is the blank frame during an unproven warm hold and a cold
switch.
A compressed runtime cache is a legal tail, not display history. Publishing
it on session switch then replacing it with the persisted lineage is the
warm-path flicker. Gate that paint on persisted-display provenance and keep
the previous/empty view until REST authority lands.
Co-authored-by: xrbs00 <178640517+xrbs00@users.noreply.github.com>
The old text read as if --yes answers yes to every prompt. It accepts the config-migration and stash-restore prompts but skips the fork-upstream prompt without adding a remote (#97052 review); say so.