The board switcher could create and configure boards but not move,
rename, or remove one. Rename technically existed, buried as a field
inside "Settings…", which is why it read as missing; it now has its own
entry and the settings dialog is left owning scope alone.
Delete archives rather than erases — the board's directory moves to
boards/_archived/ and the toast names the path — and never appears for
`default`, which the backend refuses to remove.
The three dialogs had grown three copies of the same shell, the same
"invalidate the list and close" mutation tail, and the same name field,
so those are shared now instead of parallel-implemented.
Plugins had no sanctioned way to ask for a file path — the OS door
carried notify, openExternal, revealPath and writeClipboard, so anything
needing a dialog had to reach around the SDK for window.hermesDesktop.
pickSavePath and pickOpenPath wrap the existing selectSavePath /
selectPaths IPC with the door's usual contract: resolve null when the
bridge is missing or the user cancels, never throw at the plugin.
POST /boards/{slug}/export and POST /boards/import, so the desktop and
dashboard can drive board transfer. Both exchange filesystem paths
rather than bytes, the same contract profile export/import uses and for
the same reason: the client runs its native save/open dialog on the
machine that hosts the backend, so a path is all either side needs, and
a board carrying a few hundred megabytes of attachments never has to
cross the renderer heap.
Also covers the pre-existing rename (PATCH) and delete endpoints, which
had no tests — including that delete archives to a restorable directory
and refuses to touch `default`.
`hermes kanban boards export|import` moves a board between machines:
tasks, comments, links, history, and attachments in one .tar.gz.
Two things make this more than a tar of the board directory. The
database is live — kanban runs in WAL mode, so a filesystem copy loses
whatever still sits in the -wal sidecar and tears if the dispatcher
commits mid-copy; export goes through SQLite's online-backup API
instead. And rows carry machine-local state: claims, worker PIDs,
absolute workspace and attachment paths, session ids, and the gateway
chat ids subscribed to task events. Shipping those verbatim is how an
imported board arrives holding a claim owned by a process on someone
else's laptop, or starts pushing task events into a stranger's Telegram
thread. Everything machine-local is stripped on export and re-stripped
on import, since an archive is untrusted input.
Imports always land as a NEW board, auto-suffixing the slug on
collision, so an import can never merge into or overwrite a board that
is already there. Tasks whose workspace was a directory or git worktree
on the source machine are parked in triage rather than left for the
dispatcher to claim and burn into the failure breaker.
Profile export/import owns the only hardened tar handling in the tree:
GNU-format writing (PAX fractional mtimes make macOS Archive Utility
throw "Error 94"), plus an extractor that rejects absolute paths, `..`
components, and non-regular members.
Kanban board transfer needs exactly that, and a second copy is how the
weaker of two extractors eventually ships. Move the four helpers to
hermes_cli/archive_safe and point profiles at them; no behavior change
beyond dropping a provably-unreachable fallback in the root-listing
helper, whose condition is a strict subset of the comprehension above it.
Live A/B eval (old flat vs operations[] on qwen3.8-27b / gpt-5.6-terra /
claude-sonnet-5) caught a real regression: on a fuzzy-match miss the flat
path returns file_preview so the model can self-correct, but the batch
wrapper rebuilt the error dict and dropped every field except error/
failed_index. Sonnet, recovering blind, probed the file by writing and
reverting placeholder patches for 8+ turns (50k tokens vs 15k on the flat
arm). Batch failures now merge through all non-error fields from the
failing op's result.
An unreadable root self-heals on a 3s timer, so the probe runs for as long as
the pane is open. Every forced reload cleared `rootError`, emptied `data` and
dropped `resolvedCwd` before reading, so each tick rendered "unreadable" →
blank → "unreadable" and flickered the header's project name with it. A local
ENOENT resolves well inside the 180ms skeleton delay, so the gap paints as a
bare blank frame rather than a loading state.
Re-reading the same root now probes underneath what is on screen; only a
different root, or the same path from a different backend, clears first.
An unnamed `session.info` was treated as describing whatever the pane had
selected. The gateway stamps `stored_session_id: session_key or ""`, so every
not-yet-persisted session emits one, and `broadcast_session_info` / the
approvals loop re-emit for every live session at once. An unscoped event
applies exactly when no session is active, so with nothing selected each of
those repointed `$currentCwd` and claimed it for the null selection — the file
tree, coding rail and statusbar painted a folder no selected conversation
owned, until the next `releaseWorkspaceCwdOwner` dropped the claim and they
un-painted it.
Require the event to be bound to the pane's own runtime before an absent id
reads as the selection. The case the allowance exists for — a lazy session that
is the pane's runtime but is not persisted yet — still adopts and owns its cwd.
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)
* ci: retrigger after runner startup_failure on rerun attempt
Composer drag added renderer CSS-pixel deltas onto a window AppKit
clamps to the current display, so the bar could not follow the cursor
onto a second monitor (and drifted on mixed-DPI Windows). Track the OS
cursor in main and lift that clamp.
Co-authored-by: Biotrioo <biotrioo@protonmail.com>
The JS alignment suite pinned 24.0.0 as a supported Node; with the
engines arm raised to ^24.11.0 (babel 8 requires >=24.11), 24.0.0 and
24.10.x are now correctly rejected and 24.11+/24.18+ accepted.
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.
- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
(install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
locked dependency's engines.node, and the installer gates must encode
the same floors as the manifest — so the next babel-style floor bump
turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:
error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.
Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate
* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing
* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal
* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback
* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization
Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.
1. The watcher polls only the ROOT store.
`_handoff_watcher` resolves `self._session_db` with no profile scope, which
always yields the root `state.db`. But `/handoff` run under
`hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
store. Nothing ever reads it, so the CLI times out after 60s while the
gateway is alive and connected. The watcher now iterates
`[(None, None), *secondary_profiles]` and polls each inside
`_profile_runtime_scope`.
2. The destination session key is built without the profile namespace.
`_process_handoff` called `build_session_key()` with no `profile=`,
producing `agent:main:...` while that profile's own adapter routes organic
inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
reads.
3. Delivery uses the PRIMARY profile's adapter and config.
`self.adapters` holds only the default profile's adapters (secondaries live
in `self._profile_adapters[name]`) and `self.config` only the default's home
channel. A secondary profile's handoff was therefore sent by the wrong bot,
to the wrong chat, while persisting the right session key and reporting
`handoff_state='completed'` — a false positive that looks correct in the
database and is wrong on the wire.
Two robustness fixes in the same path:
4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
turn plus delivery; awaiting it inline meant one slow handoff stopped the
watcher from even polling the other profiles. With the CLI's 60s deadline, a
valid handoff could time out purely because another profile's was ahead of
it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
never claimed twice, and a bounded drain on shutdown.
5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
one in-process dispatch, so a row still in that state at startup belongs to
a gateway that died mid-dispatch. It can never reach a terminal state, and
`request_handoff` refuses new requests unless the state is
NULL/completed/failed — that session could never hand off again, silently.
`reclaim_stale_running_handoffs()` now fails those rows once per store at
watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
may already have switched the session key and dispatched, so a blind retry
risks double delivery.
Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.
Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
FIRST_PAINT_BUDGET 20 + BACKFILL_STEP 60 prepended the rest of a 600-unit
page across ~10 visible commits. A 290-unit step keeps the interruptible
commits and removes the strobe.
Brand-new drafts are empty on purpose. A routed session the list already
knows has messages must not drop the loader just because a runtime id is
bound — that is the blank frame during an unproven warm hold and a cold
switch.
A compressed runtime cache is a legal tail, not display history. Publishing
it on session switch then replacing it with the persisted lineage is the
warm-path flicker. Gate that paint on persisted-display provenance and keep
the previous/empty view until REST authority lands.
Co-authored-by: xrbs00 <178640517+xrbs00@users.noreply.github.com>
The old text read as if --yes answers yes to every prompt. It accepts the config-migration and stash-restore prompts but skips the fork-upstream prompt without adding a remote (#97052 review); say so.
Review follow-up on #97052 (helix4u): a fork with no upstream remote whose HEAD matches origin/main used to print plain "Already up to date!" under --yes even though official main was never consulted, so an unattended stale fork looked current. _sync_with_upstream_if_needed now returns whether the official upstream was actually checked, and the commit_count == 0 completion line says "Up to date with your fork (official repo not checked)." when it was not. Skip-as-decline semantics are unchanged: no prompt, no remote mutation, no decline marker. Caller-level regression test added for the fork + no-upstream + --yes + HEAD==origin/main path; helper tests now pin the return contract.
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.
Closes#60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
GROUP_CHAT_MAX_ROUNDS and its four siblings carry over at the values
plugin.js shipped, so no rebase inherits a behavior change on top of a
rewrite. Making them configurable is live contributor work — #92213 for
per-room limits, #96842 for config plus a token budget — and both want the
same single seam, so say so where the constants are instead of adding a
config hook this PR has no consumer for.
Creating a bot opened its chat with the workspace fields omitted, because
they were spread only when the caller passed a staleness probe — and the
create path is the one caller that has none. The composer reads that scope to
stand its branch rail down in a companion chat, so a just-created bot showed
the git rail until the next click reopened the same row scoped. Live-verified
on Linux against a real backend, and carried over from the old plugin.js
rather than introduced by the rewrite.
The probe answers whether to navigate. What the session IS never depended on
it: a freshly minted Bot Chat is a bot's chat no matter who asked for it. With
the gate gone all four openers in this file are the same call, so they become
one.
The intro a new bot is born with was the first line of its forever-chat and
shipped in English, so a non-English user met their bot in a foreign language
and the bot's reply followed the prompt's language. It now resolves through
the plugin bundle in all four locales. Attribution — the other half of #91827
— still needs the lazy or silent birth that issue proposes, since
prompt.submit IS the user-turn API; the intro itself stays, per AGENTS.md.
The rest is the class the review named rather than only the lines it cited:
every user-visible string the group room and the bot-scoped cron pane own now
lives in the bundle. Where core already ships the vocabulary in every locale —
Remove, weekday names, Daily/Hourly — the plugin reuses it instead of shipping
a second, worse translation. The frequency and weekday option lists stop being
module consts frozen at import, which pinned whichever locale loaded first.
Prompts addressed to a model, cron syntax, and the 'You' author marker stay
hardcoded on purpose, each for a stated reason, recorded in the bundle header.
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
* refactor(read_file): capability-gate the anydoc format list; PDF coverage teaching lives in the response-time warning (426 -> 244/291 tok/call)
* feat(read_file): bundle firecrawl-anydoc 0.2.4 in core, typed NeedsOcrError handling, config-gated hosted OCR with local-OCR-first guidance
* refine(read_file): NEEDS-OCR warning hints at checking for an OCR skill without naming one; hosted_ocr knob unadvertised (maintainer-directed)
* simplify(read_file): drop the anydoc schema gate — bundled core dep makes absence a broken install, not a variant; formats stated unconditionally (263 tok/call)
* feat(read_file): PDF wording upgrades to 'scanned or text' when a trusted hosted-OCR route exists (direct key or explicit config; nous gateway excluded until Parse proxy works)
* simplify(read_file): FIRECRAWL_API_KEY is the ONLY hosted-OCR gate — nous gateway route removed (Parse proxy broken), config true no longer unlocks; false still disables
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.
Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The plan stated the policy set should only ever subtract from a list. That is
wrong when an allowlist names a model the curated manifest lacks, which empties
the picker instead of narrowing it — the behaviour fixed in 117e7fef88.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.
When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.
Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The usage anchor (#97206) now trusts provider-reported usage. These two
tests simulated a tool-heavy near-overflow turn while the shared fixture
reported a 12-token prompt — the anchored pressure check honestly
concluded no pressure. Give the scenario 18K anchored prompt tokens so
the tests pin the same compaction decision they always did.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).
The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.
Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):
1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
stray but leaves the declaring assistant message carrying the now
unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
displaced result still exists in the transcript, so the id survives
the set-subtraction, no stub is injected, and the payload 400s.
Fix both layers so every path is order-independent:
- repair_message_sequence: new Pass 2 prunes tool_calls that have no
result in the immediately-following tool run (matching on id or
call_id, same superset rule as Pass 1). If pruning empties the turn
(no content/reasoning left), the whole message is dropped rather than
sending an empty assistant message. Codex interim turns are exempt,
as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
rolling positional walk that drops results not immediately following
their declaring assistant (including results appearing BEFORE their
call) and injects stub results for positionally-uncovered calls even
when a mispositioned result exists elsewhere.
Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.