install-button.png is now cut from run 31449192962's LOSSLESS
welcome-screen.png (blue-text extents x465-558 y449-459 plus 3px margin,
verified complete '[ INSTALL ]' with no foreign pixels) instead of the
H.264 recording whose chroma subsampling made every video-sourced crop
miss the live screen. Tolerance drops from *60 to *20 accordingly.
PixelSearch is deleted outright: color-hunting matched the blue title
(run 31446691812) and the progress view's stage text (run 31447405319)
before it ever matched the button. Click-landed detection reuses the
same ImageSearch (button visible = not clicked). The TEMP
stop-after-screenshot exit is removed - the AHK path is live again.
The ImageSearch reference must be cropped from a lossless capture of the
real screen: the ffmpeg recording is H.264/yuv420p and its chroma
subsampling shifts glyph pixels enough that a video-sourced crop never
matches live rendering. Taken before the helper starts so no tooltip or
click marker contaminates it; lands in the log-dir artifact.
Run 31447405319 is the big win and the bug in one log: uv seeding
worked, the Install click LANDED, and 6 stages ran (clone off the fake
remote via SSH rewrite, venv, all Python dependencies) - then the AHK
loop, still hunting for 'the button', PixelSearch-matched the PROGRESS
view's own blue stage text (left column, x~106-115), decided the UI
'did not advance' 10 times, threw, and the driver killed a healthy
install mid-node-deps.
Fixes, all sourced from that recording: narrow the scan band to the
center third so the left-column stage text can never match; verify a
click by the blue vanishing AT THE CLICK POINT (a 24px box) instead of
anywhere in the band; and never throw from the click loop - the
authoritative failure signal is the driver's 'bootstrap FAILED' log
abort, and the marker deadline caps a wedged UI.
Run 31447045981: the click landed (manifest received, stages ran) but
Stage-Uv failed with 'uv installed but not found at ...\bin\uv.exe'.
GitHub windows runners ship uv preinstalled WITH an astral install
receipt; astral's cargo-dist installer then updates the receipt's
location in place and ignores UV_INSTALL_DIR, so Install-Uv's managed
copy never appears. Seed HERMES_HOME\bin\uv.exe from the runner's uv
before launching the installer - Install-Uv short-circuits on an
existing managed uv, and 'user already has a managed uv' is a
legitimate install state, not a bypass.
Also abort the run the moment the tailed bootstrap log says
'bootstrap FAILED': the failure screen waits on a human Retry, and the
AHK helper would otherwise idle out its whole 25-minute marker
deadline (and its blue-text retry loop hammers the Retry button,
re-running doomed installs - observed in run 7).
Run 31446691812: every attempt logged 'blue text at 220, 330' - exactly
the 45%-height scan boundary, which lands inside the HERMES AGENT title
(title bottom ~47% of window height; button ~62%, measured from the run
2/5 recordings). The click-landed check then correctly reported no
advance, ten times. Raise the boundary to 55%, between the two.
Run 31446292343 disproved the z-order theory: the recording shows the
installer frontmost, red click-marker dots painting on it, and the button
rendered - yet ImageSearch missed on all 5 attempts. The reference crop is
the problem: it came from an H.264/yuv420p recording whose chroma
subsampling smears glyph edges. Diffing the crop against run 5's OWN
recording of the same screen gives max 8 shades/channel (matches easily),
so the crop is video-faithful but not screen-faithful, and no tolerance
fixes that reliably.
Keep ImageSearch as the first try, but fall back to PixelSearch for the
button text's saturated blue (~0x3B82F6, variation 90) in the window's
lower half - the only blue there (the title sits in the upper third).
Verify the click landed by the blue vanishing (the progress view replaces
the button); retry up to 10 times.
Run 31445907233's recording shows the runner session's maximized console
covering the installer for the whole run: WinWait matches by title
regardless of z-order, but ImageSearch reads screen pixels, so the Install
button was never visible to it. WinActivate + WinMoveTop before every
attempt; run 31443096241 already proved the same reference crop renders
match-ably when the window is frontmost.
Run 31445244722's recording shows the published Hermes-Setup.exe renders
'[ INSTALL ]' as flat blue text on off-white - nothing like the solid-blue
'Install Hermes ->' reference from the dev-build era, so ImageSearch never
matched. Replace the reference with a crop of the real button taken from
that recording (tolerance *60 to absorb H.264 drift, click retried across
animation frames), and drop launch-button.png entirely: completion now
polls the installer's own bootstrap-complete marker
(.hermes-bootstrap-complete, see paths.rs likely_bootstrap_marker), which
cannot go stale with a UI restyle.
AutoHotkey64 is a GUI-subsystem exe: spawned without -NoNewWindow it has
no console, FileAppend('*') throws '(6) The handle is invalid' on the first
Log call, and OnError's own Log rethrows inside the handler - the script
hangs with the error tooltip painted over the installer and Install is
never clicked (confirmed from the run 31443096241 screen recording; ahk.log
was never created because the stdout write preceded the file write).
Wrap the stdout append in try (the log file is the record) and spawn the
helper with -NoNewWindow so its live lines reach the job log.
windows sibling of install-e2e-run.yml. no bubblewrap on windows, so the
git proxying is git's own transport rewrite: an isolated GIT_CONFIG_GLOBAL
with multi-valued url.<file://fake.git>.insteadOf for both hardcoded repo
URLs, so the published Hermes-Setup.exe's install.ps1 clone, hermes update's
fetch, and the desktop's ls-remote all land on a local bare repo whose main
the driver controls - installer and updater run verbatim.
one run: seed fake.git from the checkout, force fake main to the newest
release tag, drive the real published bootstrap installer with AutoHotkey
(GUI, no headless mode), promote fake main to HEAD, then apply the desktop
app's builtin update route (scripts/desktop-update.ps1 -NoUi when the
installed base ships it, staged hermes-setup.exe --update otherwise) and
assert HEAD == target with a working hermes.
TODO routes: bare hermes update, and re-running the bootstrap installer
over the existing checkout.
Design decisions belong to the orchestrator: decide naming schemes,
schemas, file formats, and API shapes before fanning out; never let two
subtree cards decide the same question; stamp every decision into each
dependent card body since workers cannot see sibling context. Mirrored
in the kanban docs (en + zh-Hans) with an exporter/importer worked
example, and bounded KANBAN_GUIDANCE size with an invariant test.
Teach the kanban reviewer to vary its inspection lens per review round
instead of repeating the same framing: round 1 reads the artifact cold
before the implementer narrative, round 2 checks out and empirically
executes the work, round 3+ audits strictly against the original
acceptance criteria and every prior request_changes item. The round is
derived from the changes_requested entries already visible in the
reviewer's worker context (live-verified against build_worker_context
across two real request_review/request_changes rounds on an isolated
board). Also adds a lens-variation note for parallel delegate_task
review fan-outs. Contract test updated with section order and lens
assertions.
Ancestor-reopen descendant invalidation previously lived only in the
dashboard plugin (_set_status_direct), so board semantics diverged by
surface and the retraction was silent: completed work snapped back to
todo and live workers were killed with no operator-visible signal.
Move it into kanban_db.invalidate_descendants_for_parent_reopen as THE
single domain implementation (recursive-CTE discovery and per-run
_retry_status_for_run handling preserved). It composes under a caller's
open transaction via write_txn(allow_nested=True) — the ancestor flip
and the descendant retractions must commit atomically — and opens its
own transaction standalone. The dashboard shim now delegates; the CLI
deliberately has no done-reopen verb (reopen-review is review-phase
only), so the DB-layer function being the single implementation is the
fix, documented in its docstring.
Non-silent: every invalidated descendant gets a descendant_invalidated
event ({ancestor, prior_status, new_status, resume_status}), the legacy
status event for existing live-feed consumers, and a task comment
naming the reopened ancestor. Running descendants keep the termination
behavior (a child building on a retracted premise is wasted spend), but
the events/comment are committed BEFORE the kill, which routes through
_terminate_reclaimed_worker — the same helper the reclaim paths use.
consecutive_failures resets to 0 on invalidated descendants: operator-
initiated invalidation is a deliberate fresh start, deliberately the
opposite of the review-loop rule (reopen_review_task preserves the
counter, #35072) so the autonomous review loop can't launder its own
failure streak.
Regression: DB-function reopen demotes done descendants with events +
comments; running descendant's audit trail is durable before its worker
dies; counter resets; dashboard and DB paths produce identical task
states, event kinds, and comment counts.
request_changes and reopen_review_task no longer reset
consecutive_failures (and last_failure_error) to 0 — review transitions
are neither success nor failure signals, so the circuit-breaker counter
is preserved (not incremented either), mirroring unblock_task (#35072).
Only complete_task's success path clears the counter.
Regression: counter=1 survives a full request_review -> request_changes
-> re-request cycle; a crash after request_changes accumulates to 2 and
trips a failure_limit=2 breaker; complete_task still resets to 0.
request_review on a running task under a live claim now requires the
caller to prove ownership (expected_run_id, the unchanged worker path)
or pass an explicit force=True override (CLI --force; dashboard human
actions pass force=True) instead of silently clearing claim_lock /
worker_pid of a live run.
Failures now carry distinct diagnostic reasons via with_reason=True
(mirroring request_changes' tuple pattern): live-claim refusal,
malformed re-review provenance, unsatisfied parents, unknown task, and
CAS miss. Tool/CLI handlers surface the specific reason instead of the
generic 'unknown id or not in running/ready'.
Regression tests: live-claim refusal + force/worker paths; malformed
provenance gets a distinct reason and explicit reviewer= recovers.
Thread lane= into check_respawn_guard. For review-lane dispatch the
active_pr and recent_success rules are skipped: a fresh PR URL comment
(and often a recent completed run) is the precondition of the canonical
review handoff, not a duplicate-work signal. Rate-limit cooldown and
the auth-blocker check still apply in every lane.
Regression: a review task with a <24h PR comment is spawned by dispatch
while a ready-lane task with the same comment stays deferred; a
rate_limited latest run still defers the review lane.
Plain write_txn raises loudly on nesting again (the historical main
invariant); composition primitives (create_task, add_comment) opt in
with allow_nested=True for savepoint semantics. create_swarm activates
the swarm root with an inline blocked->done CAS flip + synthesized run
+ event instead of nesting complete_task, so complete_task's post-commit
side effects (workspace cleanup, failure-counter clear, recompute_ready)
can no longer fire under an open outer transaction; recompute_ready now
runs after the outer commit. recompute_ready docstring corrected.
Regression: plain nesting raises; allow_nested composes and an outer
rollback discards inner work with no side effects fired.
Add a non-terminal "review" status so a worker that finished implementation
can hand off for human review without abusing kanban_block. The old
kanban_block(reason="review-required: ...") convention routed the handoff
through the unblock-loop breaker, so a normal review -> changes -> review
cycle was falsely escalated to triage.
- kanban_db: request_review (running/ready -> review, non-block, emits
review_requested), reopen_review_task (review -> ready/todo, review_reopened),
complete_task accepts review -> done, and a review_dispatch gate (default off,
shared by the dispatcher loop and the gateway health probe).
- kanban_request_review worker tool + `request-review` / `reopen-review` CLI
verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS.
- Gateway notifier wakes the origin subscriber on review_requested and
block_loop_detected; the subscription survives until done/archived, so every
review cycle re-notifies.
- Dashboard PATCH + bulk route the review transitions (request_review /
reopen_review_task) and render the review column.
- goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator.
- Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests.
needs_input / failed are unchanged: they still route through kanban_block,
still count toward block_recurrences, and still escalate to triage.
An unset browser.backend ("") now resolves to Browser Use mode whenever
the browser-use CLI is runnable (installed binary or uvx); otherwise the
built-in browser tools are kept so browsing never silently breaks.
Camofox setups always keep the built-in tools (no CDP surface), and
backend: off (including YAML 1.1 bare off -> False) forces the built-in
stack. hermes tools row highlighting follows the same effective-mode
resolution, and tests/tools/ pins CLI discovery off so host uvx installs
can't flip built-in-browser tests.
'Binary file - use appropriate tools' names a recovery the model may
not have — in a file-only toolset it thrashed for 41 turns / 178 tool
calls / 1.5M tokens on a PNG-behind-.txt (readtool eval, qwen3.8-max)
hunting for tools that did not exist. Name the type instead: 25 magic
signatures (images, archives, executables, media, SQLite), ftyp check
for ISO media, size in human units. 'Binary file (PNG image data,
4.1 KB) - cannot display as text.' answers what-is-this in one read.
Both ShellFileOperations refusal sites (read_file + read_file_raw) use
the shared describe_binary_file(); the extension-based guard keeps its
extension message (an extension is a claim; only sniffed content earns
a type name).
The Hugging Face agent-harness registry matches standard-var values
EXACTLY against the harness id. Our registry id is 'hermes-agent'
(huggingface.js agent-harnesses.ts), so AI_AGENT=hermes was counted as
'unknown' — fixed at both entry points.
Remote terminal backends (Docker/SSH/Modal/Daytona/Singularity/Vercel)
never inherit the Hermes process env, and the cross-session leak guard
deliberately strips HERMES_SESSION_* from subprocess envs in engaged
multi-session hosts — so hf/huggingface_hub traffic from those shells was
unattributable. _wrap_command now exports AI_AGENT/HERMES_AGENT inside
every wrapped command with ${VAR:-default} semantics (outer harness is
never clobbered), and the snapshot dump excludes both names so a baked
value can never shadow a later outer harness.
E2E: verified against real huggingface_hub 1.27.0 detect_agent() with a
cached registry — 'hermes-agent' detected via AI_AGENT and via
HERMES_SESSION_ID; old 'hermes' value reproduced the 'unknown' bug.
CLI and gateway entry points now set AI_AGENT=hermes (the emerging
cross-agent standard read by e.g. huggingface_hub agent detection) and
HERMES_AGENT=true, via setdefault so an outer harness is never
clobbered.
Adds skills/autonomous-ai-agents/merge-reconciler — a bundled skill teaching
a neutral third-party agent to resolve git merge conflicts between two
agents' branches: gather both diffs + intents, classify each hunk
(disjoint-intent / same-question-different-answer / superseded), resolve
under an impartiality contract, verify, and hand back a per-hunk summary.
Procedure was live-tested end-to-end against a real conflict fixture.
Includes contract tests (tests/skills/test_merge_reconciler_skill.py) and a
kanban docs cross-reference (en + zh-Hans): assign a third neutral profile a
reconciliation card with both conflicted cards as parents.
cryptography 48.0.1 carries three advisories (GHSA-m2h6-j472-rp4c,
GHSA-jwv3-5hgf-82ww, CVE-2026-69247). msal and alibabacloud-tea-openapi
cap cryptography below 49, so the bump needs an override-dependencies
entry in [tool.uv] to take effect.
The cap is conservative, not a real limit: we installed tea-openapi
against cryptography 50 and its client ran with no errors.
This override only governs `uv lock` / `uv sync`. The lazy-install
path does not read [tool.uv] and can still downgrade the pin; the next
commit closes that path.
aiohttp moves to 3.14.3 in the same pass, for GHSA-9548-qrrj-x5pj.
Reframe (per review): browser.backend: browser-use is now a DRIVER over
whatever browser source is configured, not a competing backend choice.
- browser_exec resolves its CDP endpoint through the same chain the
built-in tools use: BU_* env override > BROWSER_CDP_URL/browser.cdp_url
(/browser connect) > the configured cloud provider via browser_tool's
_get_session_info() — sharing the per-task session cache, expiry
replacement, inactivity reaper, and atexit cleanup instead of
duplicating them. Live-validated against Browserbase (session created,
driven, reaped) and gateway-provisioned Browser Use cloud browsers.
- Direct-API Browser Use configs skip provider resolution (the CLI talks
to their cloud natively via BU_AUTOSPAWN); the Nous-gateway variant
resolves through the provider, so subscribers get CLI mode without a
raw BROWSER_USE_API_KEY.
- Camofox: only true fallback — Firefox-based, custom HTTP API, no CDP
surface (its own health probes fail on CDP-schema calls). Active
Camofox setups keep the built-in browser tools even with
backend: browser-use set.
- hermes tools picker: provider rows and the Browser Use row are no
longer mutually exclusive; selecting a provider keeps the driver
choice, and both rows highlight when composed.
- Docs updated for driver-over-source semantics.
Camofox is selected via CAMOFOX_URL env var, not browser.cloud_provider —
so a Camofox user with a stray BROWSER_USE_API_KEY in .env matched the
legacy-migration predicate (cloud_provider unset + key present) and got
silently flipped into CLI mode, losing browser_* / Camofox entirely
(browser_exec cannot drive Camofox: its HTTP API exposes no CDP endpoint,
and the browser-use harness is CDP-only against Chromium).
is_legacy_browser_use_cloud_config() now defers to is_camofox_mode().
Follow-ups on the salvaged Browser Use CLI integration (PR #66476):
- browser_exec runs model-written Python on the host. Strip it at
tool-definition time for sessions whose resolved toolsets exclude
'terminal' so terminal-less surfaces (locked-down messaging configs)
don't silently regain host code execution through the browser toolset.
Session-level gate in model_tools, not a check_fn (check_fn results are
TTL-cached process-wide across sessions).
- Replace the live 'browser-use skill' schema fetch with a pinned helpers
digest: no third-party version-drifting text in the prompt, byte-stable
schema across machines. A/B benchmarked (108 runs, opus-4.8 + kimi-k3,
6 multi-step web tasks x 3 arms x 3 reps): pinned digest matches the
full skill dump 36/36 vs 36/36 at ~equal tokens; both cut total task
tokens ~60% vs the legacy browser_* toolset.
- Docs note for the terminal gate; contributor mapping for salvage.
The SDK-support flag is now bound lazily (startup-latency change); a
by-value module-level import freezes the pre-bind False. Read it off the
module after _ensure_mcp_sdk() so the test observes the real support
state — same contract, lazy-aware.
Cold CLI time-to-banner was ~1.8s (hermes) / ~2.8s (hermes -w). The banner
path was paying for work the session doesn't need before first input:
- aux availability probes built REAL OpenAI/httpx clients (openai import
~0.3s + SSL context) just to answer check_fns. New aux_probe_mode()
returns a cache-excluded stub; resolution policy unchanged.
- tools/mcp_tool imported the mcp SDK (~260ms, mcp.types pydantic model
construction) at module import even with zero MCP servers configured.
SDK import is now lazy behind _ensure_mcp_sdk(); _MCP_AVAILABLE is a
find_spec probe so every existing gate/test keeps its semantics.
- banner blocked 500ms on the update-check prefetch; now waits 50ms and
defers the warning line to a daemon thread (prints above the prompt).
- banner recomputed get_tool_definitions + skills scan + git state every
launch; now snapshotted to ~/.hermes/cache/banner_snapshot.json keyed on
(config.yaml, .env, checkout rev, toolsets) and replayed on warm launches
with a background refresh. Agent tool list is still computed fresh.
- _resolve_active_context_length probed the Nous portal /models (~200ms
network) per launch; the tool-search gate now prefers the on-disk
context cache when present.
- schema reconciliation re-executed SCHEMA_SQL in a scratch SQLite DB
(~85ms) per SessionDB(); the reference parse is now disk-memoized by
DDL hash (live-DB diffing still runs every startup).
- bundled-skills sync (~120-170ms rglob/hash) moved off the startup path
to a daemon thread; plugin discovery starts in the background and every
synchronous consumer joins via discover_plugins().
- hermes_cli.auth imported httpx eagerly (~30ms); now a lazy proxy that
test monkeypatching still reaches (setattr forwards to the real module).
- fast chat launch: unambiguous 'hermes'/'hermes chat' invocations skip
building all ~40 subcommand parsers (bails to full dispatch on anything
else, incl. container mode).
- -w path: git worktree add runs with checkout.workers=8 (0.6s→0.2s) and
overlaps HermesCLI construction; --skills preload runs in the background
and is folded in at agent init (finalize_preloaded_skills, same
fail-loud contract for fully-unknown skill lists); stale-worktree prune
moved off the banner path.
Warm results (PTY time-to-banner, 5-run): hermes 1.80s → 0.38-0.40s;
hermes -w -s hermes-agent-dev --yolo 2.82s → 0.57-0.69s.
_live_session_payload() falls back to _fallback_session_info() while a
session's agent is still None (lazy/deferred build). That fallback omitted
desktop_contract, so session.activate returned lazy metadata with no contract
field. Desktop feeds the value straight into reportBackendContract(), where a
missing field reads as contract 0 — a current backend is then falsely flagged
"Backend out of date" on every activate of a live lazy session.
The sibling session.create shape (_lazy_resume_info) was fixed the same way in
#36112; this closes the remaining session.activate gap by advertising
DESKTOP_BACKEND_CONTRACT in the fallback payload.
Adds test_session_activate_lazy_info_reports_desktop_contract pinning the
session.activate path against a lazy (agent=None) session.
read_file's .ipynb extraction previously dropped cell outputs entirely,
so a notebook's training logs, tracebacks, and printed results were
invisible to the model. Ported LobeHub's token-efficient conversion:
- stream text and error tracebacks are kept (ANSI-stripped, \r
progress-bar rewrites collapsed to the final frame)
- execute_result/display_data prefer text/plain over the HTML twin
- base64 images become sized placeholders ([image/png output — 3 KB,
omitted]); widget state and script-bearing HTML are omitted
- legacy nbformat v3 pyout/pyerr flat-field shapes handled
- per-cell output block capped at 20k chars