- build_environment_hints: split into _local_host_hints / _remote_backend_hint /
_embedder_environment_hint; backend probe split into _run_backend_probe +
_format_backend_probe with image-key / container-config dispatch tables
replacing the if/elif chain.
- Skills index: _SkillFilter (frozen dataclass) unifies the disabled+conditions
check that was copied 4x (snapshot, scan, project, external);
_collect_extra_skills dedupes the project/external scan loops;
_read_category_descriptions dedupes DESCRIPTION.md reading;
_label_visible_entries and _render_skills_index lift the org-labeling and
rendering regions out of _build_skills_system_prompt_inner; snapshot and scan
sources now feed one visibility pass.
- Context files: _read_context_file + _context_section unify the
read/strip/scan/section/truncate sequence across .hermes.md, AGENTS.md,
CLAUDE.md and .cursorrules loaders.
- Dead: _clear_backend_probe_cache (test-only helper; tests clear the dict
directly), unused org_id_of_path re-export.
- Comments/docstrings hand-compacted; every rule, invariant, ordering and
failure-mode rationale kept.
System prompt text verified byte-identical against origin/main over a 273-case
fixture corpus (env hints x backends/probe states, skills index x toolsets /
platforms / project / org / compact, context files x all loaders, full
AIAgent._build_system_prompt_parts x 9 configs). Tool schema byte-identical.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
* fix(desktop): MEDIA:-delivered non-media files route to the preview pipeline — PDFs/data files get the file card instead of a dead 'Open' anchor (extends #84951 to every extension)
* docs(prompt): desktop guidance aligned with any-file MEDIA: delivery — preview card truth, local-markdown-image block warning
* refactor(prompt): diet the memory/skills guidance block — schema-taught curricula removed, form rule + pruning contract kept (537 -> 223 tok in the combined block)
* polish: literal check-glyphs in source; memory capacity posture — save proactively, replace/consolidate when full
* refactor: single spine for memory/profile guidance — form rule + capacity posture written once, variants differ only in opening frame
* refactor: ONE memory-guidance builder — frame adapts to enabled stores, body written once, positive posture leads (maintainer direction)
* fix wording: memory is loaded per SESSION, not injected per turn (maintainer correction)
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.
- hermes-agent skill is now essential: cannot be disabled (config reads
strip it, hermes tools writes drop it), cannot be deleted by
skill_manage, is re-seeded past curator suppression, and is seeded
even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
file+terminal+vision+skills: read_file cannot read images and points
at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
tool isn't loaded.
All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
An inline ::preview widget could render and be clicked, but the click went
nowhere: the sandbox has no channel to the agent, so an interactive chart
was a dead end. Now the frame injects a second script beside the measurer
that gives the page one voice:
window.hermes.send('get-price eth')
<button data-hermes-send="get-price eth">ETH</button> (zero-script form)
The prompt rides postMessage up tagged with the mount token, then goes
through the composer's own send path (requestComposerSubmit -> prompt.submit)
flagged display_kind=hidden — the same row-typing auto-continue and internal
notifications already use. The agent wakes and takes a real turn; the
durable row persists (context, resume, DB audit); but NO bubble renders,
live or on reload. The user clicks ETH and the chart just changes — the
off-screen loop is click -> hidden turn -> agent rewrites the widget file ->
frame hot-swaps.
Trust boundary matches size reports and is tighter where it matters: mount
token required (frames can't forge each other's intents), string-only,
trimmed, capped at 500 chars, throttled to one intent per second per frame.
The gateway whitelists display_kind to "hidden" — the RPC can't mint
arbitrary row types — and the flag threads through both turn paths (inline
and compute-host isolation) so isolated sessions don't resurrect bubbles on
resume.
The desktop platform hint teaches the model to wire interactive widgets
with data-hermes-send and to answer clicks by updating the widget's file
rather than with prose; the SDK doc documents the contract.
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
The first inline frame was a full-width bordered box at a fixed height:
webpage-in-a-rectangle, not a widget. Now the frame disappears into the
message flow:
- Content-driven size. The injected measurer reports height (live) and
intrinsic width (adopted once, so %-width children can't feedback-loop the
frame toward zero). A sparkline shrink-wraps and sits flush left like an
inline image; a full-bleed page measures the whole viewport and stays
column-wide. The height attribute is now only a starting value.
- Theme bridge. A style prelude injects first with the app's resolved theme
tokens under stable names (--foreground, --muted-foreground, --accent,
--border, --card), the app font, zero body margin/padding, and a
transparent background — reference HTML written against those vars renders
native in any theme. Page styles override the prelude, so a page that
brings its own design keeps it.
- No chrome. Border, rounded box, and the rail-opener card under the frame
are gone; the fallback paths (non-HTML, remote gateway, unreadable file)
keep the classic card. The wheel gate went with the border — frames size
to content, so there is nothing to scroll inside, and widgets are fully
interactive.
- The desktop platform hint now teaches the default: an inline widget is
transparent, token-colored, flush left, no page chrome — only a standalone
page brings its own background. "Make me an inline sparkline" gets native
styling without the user spelling it out.
The first cut of the core ::preview consumer rendered the classic
preview-attachment card — a button into the right rail we already had, which
made the directive indistinguishable from an ordinary preview link. Now the
directive shows the thing itself: the workspace HTML file renders in a
sandboxed srcdoc iframe inline in the assistant message (opaque origin,
allow-scripts only — no reach into the app, its storage, or the bridge),
with an optional height attribute clamped to 120-1200px and the classic
card kept below as the rail escape hatch.
The frame waits for turn settle before reading the file (mid-stream it is
often mid-write), resolves relative paths against the session's own cwd,
and falls back to the plain card for non-HTML targets and remote gateways
(no local file door there).
The transcript becomes a contribution area (transcript.directives). A plugin
registers a named directive and the model addresses it by emitting
::name{key="value"} as its own paragraph; that leaf renders as the plugin's
component, wrapped in the contribution error boundary. Unclaimed or malformed
directives stay plain prose, so nothing changes for text that merely looks
like a directive (std::vector) or for users with the plugin disabled.
Core ships ::preview{file="..."} as the reference consumer (the existing
preview-attachment card), the desktop platform hint teaches the model the
syntax, and the SDK exports the area + types so runtime plugin.js files get
the surface through the normal plugins API.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):
- bounded serve flags corrected: the daemon accepts
--session-policy/--approve-session-policy, not the docs'
--capability-manifest names (which it rejects). Verified end-to-end:
a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
TTY) and its token is a legacy compatibility path disabled by default
on current drivers (per the live browser_prepare schema). Kept as a
passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
with cua-driver's trusted-launcher grant. config opt-in
computer_use.grant_existing_profile: true appends
--grant existing-profile to the standard-mode MCP spawn (MCP
initialize verified accepting the flag). Default false = attachment
keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
ladder: config grant > bounded manifest > YOLO; token = legacy.
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:
- hermes computer-use browser-approve: CLI passthrough that mints
cua-driver's five-minute single-use attachment token for one exact
(pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
browser_route), forwarded only for existing_profile and only as a
non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
private per-session embedded daemon launched with
--capability-manifest/--approve-capability-manifest; missing manifest
fails loudly. 'unrestricted' is deliberately NOT a config value —
it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
rungs and the isolated-profile-first default.
E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.
Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).
Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:
- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
reporter verified returns 200. Meaning, the skill_manage reference, and the
## Skill Safety Rule block are all preserved. The reword is empirically
validated rather than understood, so a comment records the bisect and warns
that any rewrite must be re-verified against an OAuth token, not an API key.
- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
longer asserts exhaustion as fact. It hedges the opening line, names the
content-filter alternative, and gives the operator a way to tell the two apart
(if the usage page still shows quota, suspect a content rejection). It also
points at `hermes auth reset anthropic`, because the credential exhaustion
latch replays the stored error for ~60 min without issuing a request — which
makes a real fix look like it did not work.
- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
not an API key, despite auth_type="api_key". It stays in api_key_env_vars
because that tuple doubles as the credential-discovery list; removing it would
stop Hermes finding a `claude setup-token` credential at all.
Docs updated to match the reworded prompt.
Fixes#82154
The messaging gateway multiplexes profiles over ONE shared launch-home
state.db, binding the profile per turn via the HERMES_HOME ContextVar
(copy_context into the worker thread). _agent_home derived the home from
db_path unconditionally, so on that lane the launch home stomped the
correctly-bound profile — deterministically inverting the leak #86313
fixed (found by @kshitijk4poor's post-merge probe; @helix4u flagged the
plugin-metadata half).
- _agent_home: bound override wins; session_db home is the unbound-thread
fallback
- _plugin_session_info: profile_name derives from _agent_home too
- full-prompt wiring regression (SOUL + skills + profile line on a bare
thread with the bot's DB) — reverting any call-site wire fails it;
multiplex, bare-thread, and plugin-metadata cases each pinned;
sabotage-verified both new tests fail against the merged behavior
- skills LRU cap 8 -> 32 (key is now per-profile x platform)
load_soul_md resolved the home ambiently, so a build thread that lost the
HERMES_HOME ContextVar read the launch profile's SOUL.md into another
profile's prompt — same class as the skills-index leak fixed in #86313.
load_soul_md and build_context_files_prompt now accept home_override, and
build_system_prompt_parts passes the agent's own home (from session_db)
at both SOUL call sites. Ambient behavior unchanged when no override.
A bot profile's system prompt could list the DEFAULT profile's ~80
skills and print 'Active Hermes profile: default' — while the live
skills_list() correctly showed the bot's real (often empty) set. The
agent plans against that index, so a false inventory makes it claim
capabilities it doesn't have, waste context tokens, and lose trust.
Root cause (confirmed empirically): the skills-prompt builder and the
active-profile line resolve the home through get_hermes_home(), which
reads a HERMES_HOME ContextVar. ContextVars do NOT propagate into
threading.Thread, so an agent build running on a thread that didn't
bind the profile's home falls back to the launch (default) home and
builds default's index. A bare no-override thread builds default's
full 7621-char block; the same thread with the fix builds empty.
Fix: resolve the agent's OWN home from its dedicated _session_db.db_path
(ground truth, ContextVar-independent) and pass it explicitly:
- build_skills_system_prompt(skills_dir_override=...) scopes the index,
the disk snapshot, and external-dir resolution to that home
- the active-profile line derives the profile name from the same home
Both fall back to ambient resolution when no db is present, so the CLI
and default-profile paths are unchanged.
Regression tests: an empty bot profile yields an empty skills block on
a bare thread even with ambient HERMES_HOME bound to a skills-rich
default; agent-home resolution from session_db.db_path. 3/3.
New desktop_ui tool: the agent proposes an MCP server (install/enable/
authorize + a one-line reason) and blocks on mcp.setup.request until the
renderer's consent card answers mcp.setup.respond with the outcome
(installed/enabled/authorized/declined/unanswered/error). Same lifecycle
as clarify: 10-min timeout, allow_expired late answers, tool lifecycle
events forced on so the card mounts even with tool progress off. Desktop
prompt hint steers the model to the tool instead of hand-editing config;
every other surface keeps the schema out and is pointed at hermes mcp
install.
Two follow-ups from live Windows sessions:
1. agent/prompt_builder.py: extend the Windows shell hint with the
native-binary path rule. Hermes disables MSYS path conversion for its
bash, so agents passing /c/Users/... or /tmp/... to NATIVE programs
(git -C, node, python, rg) hit 'cannot change to' / 'not found' while
the same path works in bash builtins — observed repeatedly in a live
session (git -C failures, git apply /tmp/x.patch failures). The hint
now says: forward-slash native form (C:/Users/x) for native tools,
$LOCALAPPDATA/Temp over /tmp for scratch files native tools read.
(/tmp is pure model habit from Linux training data — nothing
instructs it — so the hint is the right layer.)
2. tests: pin LF/CRLF preservation through write_file and patch_replace.
A live session saw a repo-LF file come back full-CRLF after an edit
(4699-line diff churn); not reproducible through current tool APIs,
so pin the correct behavior — LF files stay LF, CRLF files stay CRLF,
no mixed endings — to catch any regression on the Windows write path.
Sweep of open Windows issues affecting day-to-day agent operation
(explicitly excluding install/setup and locale classes):
- hermes_cli/_subprocess_compat.py: new split_command_line() — Windows-
safe command-line tokenizer (posix=False + quote stripping) so
backslash paths survive. POSIX behavior unchanged (plain shlex.split).
- hermes_cli/console_engine.py (#83934): console commands like
'sessions export C:\Users\me\out.jsonl' no longer silently mangle the
path into a relative filename in the cwd.
- agent/shell_hooks.py (#78293): hook commands with backslash paths now
spawn, resolve their script path, and pass hooks doctor instead of
reporting 'not executable'. All three shlex sites routed through the
shared splitter.
- agent/prompt_builder.py (#51755): system prompt now reports
Windows (11) on Windows 11 — platform.release() returns 10 for both;
distinguish via sys.getwindowsversion().build >= 22000.
- hermes_cli/commands.py (#42016): @ autocomplete no longer crashes the
prompt_toolkit event loop when rg emits a path on a different mount
(device paths \.\nul, other drive letters) — relpath ValueError is
skipped per-entry.
- tools/browser_use_cli.py (#83884): screenshot-path detection now
matches Windows drive-letter paths (C:\... and C:/...) in addition to
POSIX; Browser Use screenshots attach on Windows.
- tools/skills_hub.py + tools/skills_guard.py (#62310): the two 'MUST
stay symmetric' skill content hashes actually agree on Windows now.
Bundle keys are normalized to POSIX separators before hashing, and the
disk digest sorts by rel-posix STRING (case-sensitive) instead of Path
objects (case-insensitive on Windows). Fixes permanent false-positive
update_available for every installed skill.
Tests: tests/tools/test_windows_agent_loop_papercuts.py — 16 cases
covering each fix, including a disk-vs-bundle hash symmetry check built
with native Windows separators and a mixed-case filename.
* fix: warn agents off driving interactive console TUIs via pty on Windows
Driving 'gh auth login' (and other survey-style console TUIs) through a
pty background process on Windows silently hangs: these programs read
Win32 console key events via ReadConsoleInput, not the stdin byte
stream, so Enter keypresses submitted over process stdin never register.
The agent-visible symptom is a prompt frozen at 'Press Enter to open
browser...' while the user sees nothing, and a turn interrupt then kills
the process, invalidating any device code the user already entered on
github.com.
Two guidance fixes, both proven in a live session on Windows 10:
- agent/prompt_builder.py: extend _WINDOWS_BASH_SHELL_HINT to steer
agents toward non-interactive paths (flags, --with-token, config
files, curl-polled OAuth device flow) instead of answering console
prompts programmatically.
- skills/github/github-auth: document the pitfall and add the manual
OAuth device-flow procedure (curl against gh's public client_id,
poll for the token, finish with 'gh auth login --with-token'), which
succeeded first try after two interactive attempts hung.
* fix: send CRLF for Enter on Windows PTY submit; correct root cause in guidance
Review feedback (helix4u) was right on both counts:
1. Root cause correction. gh's 'Press Enter to open browser' prompt is
waitForEnter -> bufio.Scanner reading stdin, not a survey/console-API
prompt. The real bug is ours: submit_stdin appended a bare \n, and
through pywinpty/ConPTY a lone \n is not delivered as a line
terminator, so the child's blocking line read never returns. Verified
empirically against pywinpty 2.0.15 with a readline() child:
\n -> hang, \r -> line delivered, \r\n -> line delivered.
Fix: submit_stdin now appends \r\n for Windows PTY sessions (POSIX
PTYs and Popen pipes keep \n). Windows-only regression tests cover
the PTY and pipe branches.
2. Prompt hint rewritten: instead of claiming Windows console TUIs
cannot be driven, it now says to use process(submit) rather than raw
writes with bare \n, and to prefer non-interactive paths when a CLI
offers one.
3. Skill device flow rewritten as an executable script: parses the
device-code response, polls per the returned interval, handles
authorization_pending / slow_down (+5s per GitHub docs) /
expired_token / access_denied / unexpected responses, pipes the token
straight into gh without echoing it, and drops the undocumented
workflow scope (repo,read:org,gist is the documented minimum for
gh auth login --with-token). The pitfall note is narrowed to the
reproduced condition.
Adds the comment-based hotspot convention (no new primitives) across three
guidance surfaces:
- KANBAN_GUIDANCE worker lifecycle: new step 7 — when a file keeps colliding
with siblings or appears in other cards' recent comments, leave a
'hotspot: <path> — <reason>' kanban_comment and repeat it in completion
metadata so the orchestrator can decompose the file first.
- kanban.md (en + zh-Hans): 'Collision hotspots in parallel campaigns'
subsection — the convention, the orchestrator response (2+ flags on one
path => dedicated decomposition card before queuing more work touching
it), and the cross-link to merge-reconciler for conflicts that already
happened.
- merge-reconciler SKILL.md Pitfalls: repeated conflicts on the same file
across rounds are a hotspot signal — flag for decomposition rather than
serially reconciling.
Live-verified: guidance renders once via real import (6152 chars); hotspot
comment round-trips through add_comment -> list_comments -> worker context
on an isolated HERMES_KANBAN_DB; kanban tools, review-surfaces, and
merge-reconciler skill tests green (45 passed).
Design decisions belong to the orchestrator: decide naming schemes,
schemas, file formats, and API shapes before fanning out; never let two
subtree cards decide the same question; stamp every decision into each
dependent card body since workers cannot see sibling context. Mirrored
in the kanban docs (en + zh-Hans) with an exporter/importer worked
example, and bounded KANBAN_GUIDANCE size with an invariant test.
Add a non-terminal "review" status so a worker that finished implementation
can hand off for human review without abusing kanban_block. The old
kanban_block(reason="review-required: ...") convention routed the handoff
through the unblock-loop breaker, so a normal review -> changes -> review
cycle was falsely escalated to triage.
- kanban_db: request_review (running/ready -> review, non-block, emits
review_requested), reopen_review_task (review -> ready/todo, review_reopened),
complete_task accepts review -> done, and a review_dispatch gate (default off,
shared by the dispatcher loop and the gateway health probe).
- kanban_request_review worker tool + `request-review` / `reopen-review` CLI
verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS.
- Gateway notifier wakes the origin subscriber on review_requested and
block_loop_detected; the subscription survives until done/archived, so every
review cycle re-notifies.
- Dashboard PATCH + bulk route the review transitions (request_review /
reopen_review_task) and render the review column.
- goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator.
- Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests.
needs_input / failed are unchanged: they still route through kanban_block,
still count toward block_recurrences, and still escalate to triage.
The HUD-mode note tells the model that an unqualified "this" means the app
behind the strip. It says nothing about the app that was behind it a minute
ago, and the user drags the strip from app to app mid-thought: parked over
Spotify, "pause that and play X here" is one request spanning two apps, and
only the second half has a window under it.
Those earlier windows are already in context as read_window_below results, so
the note only has to say they still count. Without it the latest window reads
as the only one and half the request is silently dropped.
No new tool names, so the existing gating tests cover it unchanged.
In HUD mode Hermes is a strip over the app the user is actually working
in, so "what's under you?" or "look up the weather" is almost always
about that app — but the agent had no way to know it was floating, and
answered from its own browser and panes instead.
The desktop tags a HUD submit with `surface: 'hud'` and the gateway turns
that into a per-turn note pointing at read_window_below, and at carrying
the work out in the app underneath. It rides the model-bound message
beside the reaction and speech-interrupted notes rather than the system
prompt: one session can be driven from the app window on one turn and the
HUD on the next, and the system prompt has to stay byte-stable.
Every tool the note names is checked against the agent's own schema
first, so a session without computer_use or read_window_below is never
pointed at a tool it cannot call.
Commit b45a217e0 gated the TELEGRAM_RICH_MESSAGES_HINT extension behind
a config read at the top-level ``platforms.telegram.extra.rich_messages``
key, but the Telegram adapter reads the same setting from the canonical
``gateway.platforms.telegram.extra.rich_messages`` path. When users set
the setting in the canonical location (the only one documented), the
lookup returned None and the extension never fired — the model degraded
pipe tables to bullet lists, task lists to plain dashes, and never
produced <details> blocks or block math.
Fix: merge both ``gateway.platforms.telegram.extra`` and the top-level
``platforms.telegram.extra`` with the same precedence the adapter uses
(top-level leaf wins), so config-wizard writes and dashboard-setup keys
are visible alongside the canonical gateway location. Narrow the
except-guard to ImportError so real config-stack failures surface.
Port from nanocoai/nanoclaw#2748: Docker's built-in 64 MB /dev/shm silently
breaks shared-memory-hungry workloads inside the sandbox — Chromium/Playwright
renderers crash tabs, and PyTorch DataLoader workers die with 'bus error' /
'insufficient shared memory'. tmpfs is lazily allocated, so the higher ceiling
costs nothing until actually used, and usage still counts against the
container's --memory cgroup limit.
- tools/environments/docker.py: --shm-size 1g in resource args (not
cgroup-gated; tmpfs mount option). Skipped when docker_extra_args already
sets --shm-size, or when configured empty/'0' (Docker default).
- terminal.docker_shm_size config key (DEFAULT_CONFIG + all three
config->TERMINAL_DOCKER_SHM_SIZE env bridges: CLI, gateway, config.py map)
- tests: default emit, custom value, opt-out, extra_args precedence,
helper edge cases (sabotage-verified: default/custom tests fail without
the emit)