install-button.png is now cut from run 31449192962's LOSSLESS
welcome-screen.png (blue-text extents x465-558 y449-459 plus 3px margin,
verified complete '[ INSTALL ]' with no foreign pixels) instead of the
H.264 recording whose chroma subsampling made every video-sourced crop
miss the live screen. Tolerance drops from *60 to *20 accordingly.
PixelSearch is deleted outright: color-hunting matched the blue title
(run 31446691812) and the progress view's stage text (run 31447405319)
before it ever matched the button. Click-landed detection reuses the
same ImageSearch (button visible = not clicked). The TEMP
stop-after-screenshot exit is removed - the AHK path is live again.
The ImageSearch reference must be cropped from a lossless capture of the
real screen: the ffmpeg recording is H.264/yuv420p and its chroma
subsampling shifts glyph pixels enough that a video-sourced crop never
matches live rendering. Taken before the helper starts so no tooltip or
click marker contaminates it; lands in the log-dir artifact.
Run 31447405319 is the big win and the bug in one log: uv seeding
worked, the Install click LANDED, and 6 stages ran (clone off the fake
remote via SSH rewrite, venv, all Python dependencies) - then the AHK
loop, still hunting for 'the button', PixelSearch-matched the PROGRESS
view's own blue stage text (left column, x~106-115), decided the UI
'did not advance' 10 times, threw, and the driver killed a healthy
install mid-node-deps.
Fixes, all sourced from that recording: narrow the scan band to the
center third so the left-column stage text can never match; verify a
click by the blue vanishing AT THE CLICK POINT (a 24px box) instead of
anywhere in the band; and never throw from the click loop - the
authoritative failure signal is the driver's 'bootstrap FAILED' log
abort, and the marker deadline caps a wedged UI.
Run 31447045981: the click landed (manifest received, stages ran) but
Stage-Uv failed with 'uv installed but not found at ...\bin\uv.exe'.
GitHub windows runners ship uv preinstalled WITH an astral install
receipt; astral's cargo-dist installer then updates the receipt's
location in place and ignores UV_INSTALL_DIR, so Install-Uv's managed
copy never appears. Seed HERMES_HOME\bin\uv.exe from the runner's uv
before launching the installer - Install-Uv short-circuits on an
existing managed uv, and 'user already has a managed uv' is a
legitimate install state, not a bypass.
Also abort the run the moment the tailed bootstrap log says
'bootstrap FAILED': the failure screen waits on a human Retry, and the
AHK helper would otherwise idle out its whole 25-minute marker
deadline (and its blue-text retry loop hammers the Retry button,
re-running doomed installs - observed in run 7).
Run 31446691812: every attempt logged 'blue text at 220, 330' - exactly
the 45%-height scan boundary, which lands inside the HERMES AGENT title
(title bottom ~47% of window height; button ~62%, measured from the run
2/5 recordings). The click-landed check then correctly reported no
advance, ten times. Raise the boundary to 55%, between the two.
Run 31446292343 disproved the z-order theory: the recording shows the
installer frontmost, red click-marker dots painting on it, and the button
rendered - yet ImageSearch missed on all 5 attempts. The reference crop is
the problem: it came from an H.264/yuv420p recording whose chroma
subsampling smears glyph edges. Diffing the crop against run 5's OWN
recording of the same screen gives max 8 shades/channel (matches easily),
so the crop is video-faithful but not screen-faithful, and no tolerance
fixes that reliably.
Keep ImageSearch as the first try, but fall back to PixelSearch for the
button text's saturated blue (~0x3B82F6, variation 90) in the window's
lower half - the only blue there (the title sits in the upper third).
Verify the click landed by the blue vanishing (the progress view replaces
the button); retry up to 10 times.
Run 31445907233's recording shows the runner session's maximized console
covering the installer for the whole run: WinWait matches by title
regardless of z-order, but ImageSearch reads screen pixels, so the Install
button was never visible to it. WinActivate + WinMoveTop before every
attempt; run 31443096241 already proved the same reference crop renders
match-ably when the window is frontmost.
Run 31445244722's recording shows the published Hermes-Setup.exe renders
'[ INSTALL ]' as flat blue text on off-white - nothing like the solid-blue
'Install Hermes ->' reference from the dev-build era, so ImageSearch never
matched. Replace the reference with a crop of the real button taken from
that recording (tolerance *60 to absorb H.264 drift, click retried across
animation frames), and drop launch-button.png entirely: completion now
polls the installer's own bootstrap-complete marker
(.hermes-bootstrap-complete, see paths.rs likely_bootstrap_marker), which
cannot go stale with a UI restyle.
AutoHotkey64 is a GUI-subsystem exe: spawned without -NoNewWindow it has
no console, FileAppend('*') throws '(6) The handle is invalid' on the first
Log call, and OnError's own Log rethrows inside the handler - the script
hangs with the error tooltip painted over the installer and Install is
never clicked (confirmed from the run 31443096241 screen recording; ahk.log
was never created because the stdout write preceded the file write).
Wrap the stdout append in try (the log file is the record) and spawn the
helper with -NoNewWindow so its live lines reach the job log.
windows sibling of install-e2e-run.yml. no bubblewrap on windows, so the
git proxying is git's own transport rewrite: an isolated GIT_CONFIG_GLOBAL
with multi-valued url.<file://fake.git>.insteadOf for both hardcoded repo
URLs, so the published Hermes-Setup.exe's install.ps1 clone, hermes update's
fetch, and the desktop's ls-remote all land on a local bare repo whose main
the driver controls - installer and updater run verbatim.
one run: seed fake.git from the checkout, force fake main to the newest
release tag, drive the real published bootstrap installer with AutoHotkey
(GUI, no headless mode), promote fake main to HEAD, then apply the desktop
app's builtin update route (scripts/desktop-update.ps1 -NoUi when the
installed base ships it, staged hermes-setup.exe --update otherwise) and
assert HEAD == target with a working hermes.
TODO routes: bare hermes update, and re-running the bootstrap installer
over the existing checkout.
The floating HUD inherited the default Hermes title from index.html.
Set it explicitly in main and the renderer so the OS window label
matches the mode.
* test(desktop): stress long agent sessions in the multitab perf scenario
--tools seeds every transcript with settled tool rounds and drives the live
stream as a working agent turn (tool calls opened and completed between text
chunks), and each tile reveal is timed to next paint (reveal_max_ms) so deep
transcripts report their mount cost.
* perf(desktop): hold the transcript window cut steady while streaming
A fresh weight-walk per store flush slid the cut forward one message at a
time, ~30x/s, and every slide re-indexed the whole windowed transcript —
each row rendered a different message and the runtime repository took its
O(window) rebuild path instead of the one-message update. advanceTranscriptWindow
anchors the cut to a message id and re-cuts once per ~half page of new
content instead of once per flush.
* perf(desktop): stable rows, stepped backfill, pane-shared render budget
Three thread-list fixes for long streaming sessions: memoize the visible-
groups slice and each turn row so a budget-cut advance no longer re-renders
every mounted turn per streamed token; raise the first-paint backfill in
BACKFILL_STEP slices (one bounded commit per frame) instead of a single
20-to-600 transition whose commit landed as a 780ms freeze mid-stream; and
share RENDER_BUDGET across mounted panes so a 4-way grid mounts a quarter
page per pane instead of 4x the fibers.
evaluate_after_turn() calls judge_goal() which makes a synchronous
HTTP request to the auxiliary LLM. Running it on the event-loop
thread blocks Discord heartbeats for 10-40s, causing connection
flaps and gateway instability.
Offload to the default thread-pool executor so the event loop
stays responsive during evaluation.
Adds the comment-based hotspot convention (no new primitives) across three
guidance surfaces:
- KANBAN_GUIDANCE worker lifecycle: new step 7 — when a file keeps colliding
with siblings or appears in other cards' recent comments, leave a
'hotspot: <path> — <reason>' kanban_comment and repeat it in completion
metadata so the orchestrator can decompose the file first.
- kanban.md (en + zh-Hans): 'Collision hotspots in parallel campaigns'
subsection — the convention, the orchestrator response (2+ flags on one
path => dedicated decomposition card before queuing more work touching
it), and the cross-link to merge-reconciler for conflicts that already
happened.
- merge-reconciler SKILL.md Pitfalls: repeated conflicts on the same file
across rounds are a hotspot signal — flag for decomposition rather than
serially reconciling.
Live-verified: guidance renders once via real import (6152 chars); hotspot
comment round-trips through add_comment -> list_comments -> worker context
on an isolated HERMES_KANBAN_DB; kanban tools, review-surfaces, and
merge-reconciler skill tests green (45 passed).
Design decisions belong to the orchestrator: decide naming schemes,
schemas, file formats, and API shapes before fanning out; never let two
subtree cards decide the same question; stamp every decision into each
dependent card body since workers cannot see sibling context. Mirrored
in the kanban docs (en + zh-Hans) with an exporter/importer worked
example, and bounded KANBAN_GUIDANCE size with an invariant test.
Teach the kanban reviewer to vary its inspection lens per review round
instead of repeating the same framing: round 1 reads the artifact cold
before the implementer narrative, round 2 checks out and empirically
executes the work, round 3+ audits strictly against the original
acceptance criteria and every prior request_changes item. The round is
derived from the changes_requested entries already visible in the
reviewer's worker context (live-verified against build_worker_context
across two real request_review/request_changes rounds on an isolated
board). Also adds a lens-variation note for parallel delegate_task
review fan-outs. Contract test updated with section order and lens
assertions.
Ancestor-reopen descendant invalidation previously lived only in the
dashboard plugin (_set_status_direct), so board semantics diverged by
surface and the retraction was silent: completed work snapped back to
todo and live workers were killed with no operator-visible signal.
Move it into kanban_db.invalidate_descendants_for_parent_reopen as THE
single domain implementation (recursive-CTE discovery and per-run
_retry_status_for_run handling preserved). It composes under a caller's
open transaction via write_txn(allow_nested=True) — the ancestor flip
and the descendant retractions must commit atomically — and opens its
own transaction standalone. The dashboard shim now delegates; the CLI
deliberately has no done-reopen verb (reopen-review is review-phase
only), so the DB-layer function being the single implementation is the
fix, documented in its docstring.
Non-silent: every invalidated descendant gets a descendant_invalidated
event ({ancestor, prior_status, new_status, resume_status}), the legacy
status event for existing live-feed consumers, and a task comment
naming the reopened ancestor. Running descendants keep the termination
behavior (a child building on a retracted premise is wasted spend), but
the events/comment are committed BEFORE the kill, which routes through
_terminate_reclaimed_worker — the same helper the reclaim paths use.
consecutive_failures resets to 0 on invalidated descendants: operator-
initiated invalidation is a deliberate fresh start, deliberately the
opposite of the review-loop rule (reopen_review_task preserves the
counter, #35072) so the autonomous review loop can't launder its own
failure streak.
Regression: DB-function reopen demotes done descendants with events +
comments; running descendant's audit trail is durable before its worker
dies; counter resets; dashboard and DB paths produce identical task
states, event kinds, and comment counts.
request_changes and reopen_review_task no longer reset
consecutive_failures (and last_failure_error) to 0 — review transitions
are neither success nor failure signals, so the circuit-breaker counter
is preserved (not incremented either), mirroring unblock_task (#35072).
Only complete_task's success path clears the counter.
Regression: counter=1 survives a full request_review -> request_changes
-> re-request cycle; a crash after request_changes accumulates to 2 and
trips a failure_limit=2 breaker; complete_task still resets to 0.
request_review on a running task under a live claim now requires the
caller to prove ownership (expected_run_id, the unchanged worker path)
or pass an explicit force=True override (CLI --force; dashboard human
actions pass force=True) instead of silently clearing claim_lock /
worker_pid of a live run.
Failures now carry distinct diagnostic reasons via with_reason=True
(mirroring request_changes' tuple pattern): live-claim refusal,
malformed re-review provenance, unsatisfied parents, unknown task, and
CAS miss. Tool/CLI handlers surface the specific reason instead of the
generic 'unknown id or not in running/ready'.
Regression tests: live-claim refusal + force/worker paths; malformed
provenance gets a distinct reason and explicit reviewer= recovers.
Thread lane= into check_respawn_guard. For review-lane dispatch the
active_pr and recent_success rules are skipped: a fresh PR URL comment
(and often a recent completed run) is the precondition of the canonical
review handoff, not a duplicate-work signal. Rate-limit cooldown and
the auth-blocker check still apply in every lane.
Regression: a review task with a <24h PR comment is spawned by dispatch
while a ready-lane task with the same comment stays deferred; a
rate_limited latest run still defers the review lane.
Plain write_txn raises loudly on nesting again (the historical main
invariant); composition primitives (create_task, add_comment) opt in
with allow_nested=True for savepoint semantics. create_swarm activates
the swarm root with an inline blocked->done CAS flip + synthesized run
+ event instead of nesting complete_task, so complete_task's post-commit
side effects (workspace cleanup, failure-counter clear, recompute_ready)
can no longer fire under an open outer transaction; recompute_ready now
runs after the outer commit. recompute_ready docstring corrected.
Regression: plain nesting raises; allow_nested composes and an outer
rollback discards inner work with no side effects fired.
Add a non-terminal "review" status so a worker that finished implementation
can hand off for human review without abusing kanban_block. The old
kanban_block(reason="review-required: ...") convention routed the handoff
through the unblock-loop breaker, so a normal review -> changes -> review
cycle was falsely escalated to triage.
- kanban_db: request_review (running/ready -> review, non-block, emits
review_requested), reopen_review_task (review -> ready/todo, review_reopened),
complete_task accepts review -> done, and a review_dispatch gate (default off,
shared by the dispatcher loop and the gateway health probe).
- kanban_request_review worker tool + `request-review` / `reopen-review` CLI
verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS.
- Gateway notifier wakes the origin subscriber on review_requested and
block_loop_detected; the subscription survives until done/archived, so every
review cycle re-notifies.
- Dashboard PATCH + bulk route the review transitions (request_review /
reopen_review_task) and render the review column.
- goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator.
- Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests.
needs_input / failed are unchanged: they still route through kanban_block,
still count toward block_recurrences, and still escalate to triage.
An unset browser.backend ("") now resolves to Browser Use mode whenever
the browser-use CLI is runnable (installed binary or uvx); otherwise the
built-in browser tools are kept so browsing never silently breaks.
Camofox setups always keep the built-in tools (no CDP surface), and
backend: off (including YAML 1.1 bare off -> False) forces the built-in
stack. hermes tools row highlighting follows the same effective-mode
resolution, and tests/tools/ pins CLI discovery off so host uvx installs
can't flip built-in-browser tests.
The upstream ahujasid/blender-mcp and ahujasid/ableton-mcp GitHub repos
were hijacked on 2026-08-08: the maintainer (@sidahuj) publicly reported
his account was compromised and ownership stripped, and both repos now
redirect to an attacker-controlled org (MCPBlender, created the same
day, pushing new commits since).
Although our catalog pinned blender-mcp==1.6.4 from PyPI (pre-compromise,
sha256 verified unchanged), the server is only half the bridge: the
manifest's post-install instructions and the optional skill directed
users to download addon.py — arbitrary Python executed inside Blender —
from the now-compromised GitHub repo (the raw URL currently 404s, and
the addon ships in no PyPI artifact). There is no trustworthy source
for the addon half, so the entry cannot be installed safely end-to-end.
Removing the catalog entry and skill entirely until the maintainer
confirms account recovery; re-adding is a follow-up PR once upstream
is verified clean.
- optional-mcps/blender/: removed
- optional-skills/creative/blender-mcp/: removed
- docs: catalog rows, sidebar entry, skill pages (en + zh-Hans) removed
- cross-references in unreal-mcp and kanban-video-orchestrator cleaned
'Binary file - use appropriate tools' names a recovery the model may
not have — in a file-only toolset it thrashed for 41 turns / 178 tool
calls / 1.5M tokens on a PNG-behind-.txt (readtool eval, qwen3.8-max)
hunting for tools that did not exist. Name the type instead: 25 magic
signatures (images, archives, executables, media, SQLite), ftyp check
for ISO media, size in human units. 'Binary file (PNG image data,
4.1 KB) - cannot display as text.' answers what-is-this in one read.
Both ShellFileOperations refusal sites (read_file + read_file_raw) use
the shared describe_binary_file(); the extension-based guard keeps its
extension message (an extension is a claim; only sniffed content earns
a type name).
Each test slice uploads an artifact with the same file name,
test_durations.json. The save-durations job downloaded the 12
artifacts with merge-multiple, so all extractions wrote to one
path in parallel. This caused two faults:
- A race between two extractions wrote two JSON documents into
one file. The merge step then failed with 'JSONDecodeError:
Extra data' (run 31382130252).
- On green runs, the last write erased the other 11 slices. The
merged cache held ~230 of ~2760 file durations.
Remove merge-multiple so each artifact extracts into its own
directory, and point the glob at durations/*/test_durations.json.
A local merge of the 12 real artifacts from the failed run gives
2761 durations.
The Hugging Face agent-harness registry matches standard-var values
EXACTLY against the harness id. Our registry id is 'hermes-agent'
(huggingface.js agent-harnesses.ts), so AI_AGENT=hermes was counted as
'unknown' — fixed at both entry points.
Remote terminal backends (Docker/SSH/Modal/Daytona/Singularity/Vercel)
never inherit the Hermes process env, and the cross-session leak guard
deliberately strips HERMES_SESSION_* from subprocess envs in engaged
multi-session hosts — so hf/huggingface_hub traffic from those shells was
unattributable. _wrap_command now exports AI_AGENT/HERMES_AGENT inside
every wrapped command with ${VAR:-default} semantics (outer harness is
never clobbered), and the snapshot dump excludes both names so a baked
value can never shadow a later outer harness.
E2E: verified against real huggingface_hub 1.27.0 detect_agent() with a
cached registry — 'hermes-agent' detected via AI_AGENT and via
HERMES_SESSION_ID; old 'hermes' value reproduced the 'unknown' bug.
CLI and gateway entry points now set AI_AGENT=hermes (the emerging
cross-agent standard read by e.g. huggingface_hub agent detection) and
HERMES_AGENT=true, via setdefault so an outer harness is never
clobbered.
Adds skills/autonomous-ai-agents/merge-reconciler — a bundled skill teaching
a neutral third-party agent to resolve git merge conflicts between two
agents' branches: gather both diffs + intents, classify each hunk
(disjoint-intent / same-question-different-answer / superseded), resolve
under an impartiality contract, verify, and hand back a per-hunk summary.
Procedure was live-tested end-to-end against a real conflict fixture.
Includes contract tests (tests/skills/test_merge_reconciler_skill.py) and a
kanban docs cross-reference (en + zh-Hans): assign a third neutral profile a
reconciliation card with both conflicted cards as parents.