Commit Graph

76 Commits

Author SHA1 Message Date
Teknium a49a9d79b3 Port from lobehub/lobehub#19329: surface environment recreation in terminal tool results
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.

Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.

Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
2026-09-12 20:55:04 -07:00
kshitijk4poor 4fc64b6730 refactor(terminal): probe NOPASSWD only when a prompt would fire; drop the host-only probe
With BaseEnvironment supplying a backend-scoped probe to every production
caller, the module-level host-only `_sudo_nopasswd_works` (and its
TERMINAL_ENV gate) had no callers left; the `or` fallback and the outer
try/except around a callback that already fails closed were dead too.

Move the probe under `should_prompt_for_sudo`: headless callers (gateway,
cron, delegated children) reach `(command, None)` whether or not the probe
runs, so the extra backend round trip — an ssh exec on the SSH backend —
was pure waste on every headless sudo command.

test_subagent_sudo_prompt no longer needs to patch the host probe out:
bare `_transform_sudo_command(cmd)` calls have no probe by construction.
2026-09-11 11:19:53 +05:30
fangliquanflq 39abca492d fix(terminal): gate sudo probes by backend cancellation safety
Opt in only Local, Docker, SSH and Singularity: a timed-out `sudo -n true`
probe on those backends kills one process, while SDK adapters (Modal,
Daytona, Vercel) cancel by terminating the whole sandbox.

Net of the original PR's commits f7db10ef + 9965b1b9 (the intermediate
_ThreadedProcessHandle special-case was superseded by this gate).
2026-09-11 11:19:53 +05:30
fangliquanflq d38064e3c1 fix(terminal): probe remote passwordless sudo 2026-09-11 11:19:53 +05:30
Teknium 463292351f fix: mid-turn user message no longer waits behind a foreground terminal command
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.

redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.

Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
2026-09-05 13:51:06 -07:00
Tarek Belkahia 70a64db532 feat(gateway): deliver MEDIA files that live inside a remote terminal sandbox (#466)
`MEDIA:/path` only delivered when the file existed on the gateway host.
With the ssh / modal / daytona / singularity / vercel backends the agent's
artifact sits on another filesystem, so `validate_media_delivery_path`
rejected it and the attachment vanished with a "Skipping unsafe" log line —
the #1 gap for sandboxed deployments.

`BaseEnvironment.fetch_file` pulls a regular file out of any backend over
the exec channel (base64, marker-fenced, size-bounded INSIDE the sandbox so
/dev/zero cannot flood host memory — the same shape image_source already
uses). `gateway/media_fetch.py` runs only when a remote backend is active
and the host lookup failed: the sandbox path is screened against the SAME
denylist as host deliveries, again after `readlink -f`, then copied into
the document cache (an allowlisted root) and validated like any host file.
No new tool; the existing `MEDIA:` tag is the interface.

Salvaged from #68506 by @tokou (design and denylist mirroring); redone on
current main without the send_file tool, the per-backend transports and
the undeliverable-notice plumbing (the #66797 failure notice already covers
that surface).
2026-09-05 16:09:46 +05:30
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 98c140bc4b simplify(compat): code_execution_tool/environments.local — drop 36 re-exports, repoint 5 callers + 14 test files 2026-09-03 13:24:04 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 132d67645c refactor(tools): unify docker probe/start helpers, collapse egress arg parser, compact env module docstrings 2026-09-02 22:43:33 -07:00
Teknium 5b24a7a649 refactor(tools): compact environment base/docker/singularity modules, extract docker __init__ phases 2026-09-02 22:04:56 -07:00
Teknium 3bbec90f23 refactor(tools/environments): split local/base into output/wait/session-env siblings; extract docker_egress + remote_common; dedupe remote backends; compact process_registry 2026-09-02 14:45:15 -07:00
Teknium 53e3c14f1d fix: clamp invalid effective_timeout to the 120s wait default instead of unbounded (review follow-up for #94305) 2026-08-31 10:42:39 -07:00
HexLab98 85bc25c949 fix(terminal): bound env.execute wait so a wedged poll cannot disable every timer
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
2026-08-31 10:42:39 -07:00
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
Teknium 1dc552d5d1 refactor(terminal): honest schema, pager defaults in the env, unified notify arg (837 → 670 tok/call, −20%) (#95937)
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens

* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim

* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)

* fix(terminal): background param referenced the unadvertised legacy arg name

* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape

* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
2026-08-26 19:23:39 -07:00
Teknium 979b7d14fd fix(search): path-scoped grep pruning + execution-backend gating for macOS TCC exclusions
Two fixups the #75785 review required before landing:
- grep fallback no longer uses --exclude-dir for protected dirs: grep
  matches exclude-dir globs against BASENAMES anywhere in the tree, so
  --exclude-dir=Downloads silently skipped every nested directory named
  Downloads (a repo's own Downloads/ included). Protected-dir searches now
  route through find's path-scoped -prune (same traversal-prevention the
  find backend uses) feeding grep via -exec. Regression test proves a
  nested work/repo/Downloads/notes.txt is still found while ~/Downloads is
  not (live filesystem, real find+grep).
- exclusions gated on env.is_local (new BaseEnvironment flag, True on
  LocalEnvironment): sys.platform/Path.home() describe the controller, not
  the execution host — a macOS controller driving a Linux SSH/container
  backend must not prune the remote's unprotected Downloads. Environments
  without the flag default to local semantics (warning-carrying skip,
  never data loss).

Both sabotage-verified: restoring basename --exclude-dir fails 2 tests.
2026-08-26 06:26:02 -07:00
Teknium 410b1ec555 fix(tools): share the task-id path sanitizer across backends; cover singularity overlays
Hoist the sandbox-directory sanitizer into tools/environments/base.py as
sanitize_task_id_for_path() and route BOTH host-path consumers through it:
the docker persistent sandbox (get_sandbox_dir()/docker/<id>) and the
singularity persistent overlay (hermes-overlays/overlay-<id>). One helper,
one mapping, whole bug class fixed in one place instead of per-backend
copies (#92414, #92640, #93044).

docker.py keeps _sandbox_dir_name as an alias of the shared helper so the
sanitized mapping (safe ids verbatim, digest suffix on rewrite for
collision safety) is unchanged for existing sandboxes.

Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-08-23 21:12:32 -07:00
abundantbeing 095a1d078c fix(browser): preserve controller work across reconnects
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.

Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
2026-08-21 22:33:45 +05:30
ethernet 16a173a8d6 fix(tools): do not adopt a stale cwd after an interrupted command
The command wrapper prints the cwd marker after the command returns. A
killed or timed-out command emits no marker, so ``env.cwd`` still holds the
directory of the last command to FINISH. One local environment serves every
session, because ``_resolve_container_task_id`` collapses cwd-only overrides
to ``"default"``. That leftover directory is therefore routinely another
session's.

The post-command dual-write copied ``env.cwd`` into the interrupted session's
durable record. Every later command in that session then ran in the foreign
directory, and the cwd echo told the model it had moved there. A desktop chat
silently re-homed into a worktree that another chat had opened.

Report the observation instead of inferring it. The marker parse now sets
``result["cwd_observed"]``, and both the record write and the echo read that
flag. The local override clears the flag when it rolls back a path that does
not exist, because the restored value is also unobserved. When a command
reports no cwd, the session keeps the directory it already had.

This needs no second session to be wrong: a lone session that interrupts a
command re-adopts a stale value too. A second session only makes the wrong
directory belong to somebody else.

The same class of write exists in the file-tools rescue for a reaped
environment (#26211). That rescue copied the cached snapshot of the shared
``env.cwd`` into the session record. The rescue is now fill-only: it writes
the snapshot when the session has no record, and it never overwrites a
record that the session wrote for itself.

The tests drive ``terminal_tool`` itself through an interrupt, not a copy of
its gate. Review found that a revert of either call site passed the first
version of the tests. Each gate now has a test that fails when the gate is
removed (verified by mutation).

Two exact-dict assertions in the Vercel sandbox tests now assert the two
fields they care about, so a new result key does not fail them.
2026-08-13 20:30:26 -07:00
Teknium 005dfcbfcc fix(tools): symlink-safe exclusive creation for all spill/cache writers
Spill files (terminal overflow, hook context, web_extract full text,
subagent summaries) were written with plain open()/write_text into
predictable directories. A pre-planted symlink at any of those paths
redirected the write onto an arbitrary user-owned file, and raw
pre-redaction terminal/hook spills landed world-readable under the
default umask.

New tools/spill_safety.py helpers create files with
O_CREAT|O_EXCL|O_NOFOLLOW (a link-shaped path fails the write instead of
following it) and overwrite via lstat-checked unlink + exclusive
re-create, so even the redaction rewrite cannot be diverted. Private
tier (0o700 dir / 0o600 file) covers raw terminal and hook spills;
cache/web and cache/delegation keep umask perms because those dirs are
bind-mounted into remote backends that must read them.

Pattern borrowed from DeepSeek Harness dsh-spill-local (MIT):
private root + exclusive owner-only opens for spill artifacts.
2026-08-13 11:09:51 -07:00
Teknium e5bc6b2186 fix(attribution): correct AI_AGENT id to registry value and carry harness markers into all terminal backends
The Hugging Face agent-harness registry matches standard-var values
EXACTLY against the harness id. Our registry id is 'hermes-agent'
(huggingface.js agent-harnesses.ts), so AI_AGENT=hermes was counted as
'unknown' — fixed at both entry points.

Remote terminal backends (Docker/SSH/Modal/Daytona/Singularity/Vercel)
never inherit the Hermes process env, and the cross-session leak guard
deliberately strips HERMES_SESSION_* from subprocess envs in engaged
multi-session hosts — so hf/huggingface_hub traffic from those shells was
unattributable. _wrap_command now exports AI_AGENT/HERMES_AGENT inside
every wrapped command with ${VAR:-default} semantics (outer harness is
never clobbered), and the snapshot dump excludes both names so a baked
value can never shadow a later outer harness.

E2E: verified against real huggingface_hub 1.27.0 detect_agent() with a
cached registry — 'hermes-agent' detected via AI_AGENT and via
HERMES_SESSION_ID; old 'hermes' value reproduced the 'unknown' bug.
2026-08-10 11:07:22 -07:00
Theophilus Chinomona b0594118ab fix(environments): surface stdin write failures as stdin_error (#79178) 2026-08-08 12:31:19 -07:00
Theophilus Chinomona c5a1a5d7b0 fix(environments): surrogateescape-safe stdin piping, always close stdin (#79178) 2026-08-08 12:31:19 -07:00
Teknium 5c29566e8d feat(terminal): graceful degradation for remote backend connection failures
Connection-class infrastructure failures on remote terminal backends (SSH
host unreachable/timed out, Docker daemon down or missing, remote file
sync failing on a dead link) previously surfaced to the model as raised
RuntimeError tracebacks. The model got a stack blob with no guidance and
the failure was indistinguishable from a tool bug.

Now:

- New EnvironmentConnectionError(RuntimeError) in tools/environments/base.py
  carrying a reason + retry_hint. Subclassing RuntimeError keeps every
  existing catcher working.
- ssh.py classifies connect-refused, connect-timeout, scp, remote mkdir,
  bulk upload/download, and remote rm failures as connection errors.
- docker.py classifies all four _ensure_docker_available() failure paths
  (missing exe, non-executable exe, daemon timeout, `docker version`
  failure).
- terminal_tool catches EnvironmentConnectionError and returns a
  structured tool result the model can act on:
    {"status": "degraded", "reason": ..., "retry_hint": ..., "exit_code": -1}
  The failed backend is evicted from the environment cache so a later
  call retries from scratch — recovery is automatic once the backend is
  reachable again.
- Config gate terminal.degraded_mode: warn|fail (default warn) in
  config.yaml, bridged as TERMINAL_DEGRADED_MODE across all four bridge
  sites (cli.py env_mappings, gateway/run.py _terminal_env_map,
  TERMINAL_CONFIG_ENV_MAP, DEFAULT_CONFIG). "fail" preserves the
  historical error+traceback tool result.
- Command failures (nonzero exit, command-not-found) are NOT touched —
  only infrastructure failures classify as degraded.

Tests: tests/tools/test_terminal_degraded_mode.py (15 tests) covering
exception classification for ssh+docker, structured degraded results,
no-caching of degraded envs, recovery after the backend returns,
nonzero-exit results unaffected, fail-mode preservation, invalid-mode
fallback to warn, and the four-site config bridge invariant.

Inspired by: Claude Cowork degraded-backend behavior (idea-level,
docs-only evidence).
2026-08-07 09:07:55 -07:00
f1aggo_macair 911d380296 fix(tools): allocate snapshot temp paths with mktemp instead of $BASHPID
Extracted from #54314 (@flag0x369), re-derived onto current main: macOS
ships bash 3.2 as /bin/bash, which lacks $BASHPID entirely — the
variable expands to empty string, collapsing every concurrent writer's
'unique' temp path onto the same file (torn snapshot writes under
concurrency). mktemp allocates per-writer unique paths portably.
Live-verified: /bin/bash -c 'echo $BASHPID' prints empty on this box.
2026-08-03 13:47:29 +05:30
Teknium 80631c4aea feat(terminal): recoverable truncation — spill full output + report pre-truncation size
Truncated terminal output was information LOSS: the middle was gone
and the only recovery was re-running the command (data: 1,394
truncation markers in a 250k-call window, with re-runs and grep
retries chained behind the big ones).

Truncation is now deferred retrieval (opencode/goose/qwen-code
pattern, codex's original_token_count idea):

- tools/environments/base.py: _BoundedOutputCollector gains an
  optional spill tee — when foreground output overflows the capture
  window, the FULL stream is teed to
  ~/.hermes/cache/terminal-output/out-*.log (lazy file creation with
  backlog backfill, 5MB hard cap, 7-day opportunistic cleanup,
  disk errors never break execution). All three _wait_for_process
  returns attach {output_total_chars, full_output_path} via a shared
  finalizer.
- tools/terminal_tool.py: redacts the spill with the same
  redact_terminal_output pass as the visible output (no secret
  persists unmasked), then surfaces output_total_chars,
  full_output_path, and a truncation_note pointing at
  search_files/read_file instead of a re-run.

Non-truncated results are byte-identical; internal unbounded
consumers (file-ops cat reads, RPC reads) are untouched (spill only
arms with bounded_capture=true).
2026-08-02 15:52:14 -07:00
kshitij fb6446fc9e fix(cron): scope cron approval context per session
Replace the process-global HERMES_CRON_SESSION env var with a per-session
ContextVar so a cron tick in the gateway process cannot leak into unrelated
live gateway/API/TUI turns. The cron scheduler now sets the ContextVar
inside the job's try/finally scope and resets it on cleanup. Gateway, API
server, ACP adapter, and TUI gateway all pass cron_session='' to explicitly
mark their sessions as non-cron, masking any stale process env.

Co-authored-by: hinablue <hinablue@gmail.com>
Closes #37968
2026-08-03 00:25:20 +05:30
kshitijk4poor 8fd1a68106 refactor(cron): harden the run heartbeat (review follow-ups)
- heartbeat loop continues past a raising activity callback instead of
  silently stopping (matches delegate_task / touch_activity_if_due
  swallow-and-continue semantics) — one transient error must not drop
  watchdog protection for the rest of a long job
- hard 6h elapsed ceiling so a wedged job under HERMES_CRON_TIMEOUT=0
  (unlimited child watchdog) cannot mask the gateway watchdog forever
- public get_activity_callback() accessor in tools/environments/base.py
  instead of importing the private _get_activity_callback cross-module
- tests: deterministic heartbeat test (event-gated, no timing sleep),
  no-callback test now asserts the thread is truly never created, new
  exception-survival guard; dead started event removed
- fix comment: delegate_task heartbeat cadence is 30s, not 10s
2026-08-02 14:04:52 +05:30
Christopher fc61608a17 fix(security): isolate explicit Docker passthrough snapshots 2026-08-02 00:36:03 -07:00
Christopher 7138b9587a fix(security): scope passthrough env to routed profile 2026-08-02 00:36:03 -07:00
HexLab98 9677495004 fix(environments): exclude multiline session env from terminal snapshots
Line-based grep on export -p only drops the first declare -x line, so a
newline in HERMES_SESSION_CHAT_NAME/USER_NAME leaves shell payload in the
shared snapshot and runs it on the next source. Unset bridged vars in a
subshell before export -p instead (issue #71296).
2026-07-26 19:30:02 -07:00
teknium1 95a566b1e7 fix: brace-group the filtered export dump so $BASHPID expands in the parent shell
Follow-up for salvaged PR #69380. The snippet
'export -p | grep -vE ... > $tmp.$BASHPID || true' attaches the redirect to
the grep pipeline segment, so $BASHPID expands inside grep's pipeline
subshell — a DIFFERENT pid than the parent shell that expands the follow-up
'mv $tmp' operand. The dump landed in an orphaned temp file, mv failed
silently (2>/dev/null), and the shared snapshot never updated again:
exported env / venv activation stopped persisting between commands
(tests/tools/test_local_shell_init.py TestSnapshotEndToEnd caught it).

Wrap the pipeline in a brace group and attach the redirect to the group so
the expansion happens in the current shell, matching mv's expansion. Also
rename the shell-init probe var HERMES_SESSION_ENV_PROBE ->
HERMES_STICKY_ENV_PROBE: it matched the HERMES_SESSION_ prefix the salvaged
fix now intentionally strips from snapshots, and the snippet-shape test is
updated to pin the brace-group contract.
2026-07-24 23:09:02 -07:00
bo.fu f2e32ceead fix(terminal): stop shared bash snapshot leaking HERMES_SESSION_* across sessions
A single long-lived backend serves many sessions through one
_active_environments['default'] LocalEnvironment (the messaging gateway, TUI,
and desktop/web dashboard all collapse the terminal to 'default'). That
environment persists a bash session snapshot file and sources it before every
command. 'export -p' dumped the FIRST session's HERMES_SESSION_ID into the
snapshot, so every LATER session sourced that stale value and its
'echo $HERMES_SESSION_ID' reported a FOREIGN session's id — overriding the
correct per-command Popen env injected by _inject_session_context_env.

Confirmed on staging-fp: session A read its own id, session B (same backend)
read A's id. Reproduced locally with two threads sharing one LocalEnvironment;
the snapshot held 'declare -x HERMES_SESSION_ID=...' from the first session.

Fix: strip the per-session bridged vars (HERMES_SESSION_* / HERMES_UI_SESSION_ID
/ HERMES_CRON_AUTO_DELIVER_*, i.e. gateway.session_context._VAR_MAP prefixes)
from the snapshot at both dump sites in base.py. They are re-injected fresh on
every command, so a snapshot should carry only the user's own shell state, not
Hermes' per-turn session identity.

Complements the _set_session_context session_id fix: that ensures the ContextVar
carries the right id; this ensures the shared snapshot can't override it with a
neighbour's. Adds tests/tools/test_snapshot_session_id_leak.py (regex unit +
real two-session LocalEnvironment integration) and updates the export-shape
assertions in test_base_environment.py for the new 'export -p | grep' dump.
2026-07-24 23:09:02 -07:00
teknium1 d4b867cf9f fix(windows): sweep remaining unguarded text-mode subprocess sites codebase-wide
AST-driven pass over every subprocess.run/Popen/check_output/check_call/call
with text=True (or universal_newlines=True) and no explicit encoding=:
append encoding='utf-8', errors='replace' at the kwarg site. 136 call
sites across 28 files (cli.py, hermes_cli/main.py, tools_config.py,
environments, computer_use, gateway, scripts, skills helpers, agent/*).

Together with the salvaged #55339/#60741 commits this closes out issue
#53428's bug class; the salvaged #60751 linter rule in
check-windows-footguns.py now enforces it repo-wide (verified: 807 files
scanned, zero findings).
2026-07-24 11:45:57 -07:00
Teknium 49a8c3f836 fix(terminal): stop writing the cwd temp file entirely
Follow-up for salvaged PR #63255: with LocalEnvironment._update_cwd
delegating to the stdout marker parser, the cwd temp file has zero
readers left. Drop the 'pwd -P > file' writes from the bootstrap and
_wrap_command so every command stops paying a pointless file write
(and stops littering temp dirs with hermes-cwd-*.txt).
2026-07-16 01:28:24 -07:00
Teknium cab457d722 fix(terminal): make bounded capture opt-in for the foreground terminal path only
Review finding on the salvaged collector: _wait_for_process is the shared
drain for EVERY env.execute() consumer, not just the terminal tool. Applying
tool_output.max_bytes there silently truncated file-operation cat reads
(read_file_raw feeds the patch engine — read-modify-write on any file >50KB
would corrupt it), paginated read_file, code-execution RPC reads, and log
reads.

bounded_capture is now an explicit opt-in on execute()/_wait_for_process,
set only by the foreground terminal tool. Default preserves the historical
full-fidelity capture via an effectively-unbounded collector (single code
path). Modal transports accept the kwarg for signature parity.

New regression test: default execute() returns a 200KB payload complete and
untruncated. E2E: 20MB internal read intact; ShellFileOperations
read_file_raw round-trips byte-exact; terminal path still bounded at 50KB.
2026-07-14 22:00:28 -07:00
embwl0x 0a07609173 fix(terminal): bound foreground output capture 2026-07-14 22:00:28 -07:00
Brooklyn Nicholson c4622a1d5b fix(windows): survive broken Git Bash login shells
#63621 fixed path quoting, but Ainz's Git for Windows still dies on
`bash -l` itself (`Directory \drivers\etc`). Hermes then fell back to
bash -l *per command*, so every write_file/terminal call failed the same way.

After a failed login snapshot, probe non-login bash -c; if it works, skip
-l for the session. Also skip a stale HERMES_GIT_BASH_PATH that fails a
noprofile probe in favor of %LOCALAPPDATA%\hermes\git portable bash.
2026-07-13 16:14:18 -04:00
Brooklyn Nicholson f2fcf89c1f fix(windows): bash-safe snapshot paths after #63113
#63113 rewrote native drive paths in ShellFileOperations, but init_session
/_wrap_command still embedded C:/... hermes-snap paths from get_temp_dir.
MSYS arg-converts those during bash -l and surfaces Directory \drivers\etc
— including for relative write_file targets, since the wrapper is the fault.

Add _bash_safe_path, override BaseEnvironment._quote_shell_path on
LocalEnvironment (no base→local import), and normalize mixed /c/Users\...
paths in file ops.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-07-13 01:59:58 -04:00
Eugeniusz Gilewski a1e6ea7d71 fix(tools): keep shell snapshots owner-only
BaseEnvironment writes shell snapshots and cwd metadata through the process
umask. With a common 022 umask, snapshot files containing exported environment
state landed at mode 0644 even though they can include env-carried credentials
from the parent process.

Set umask 077 only around Hermes metadata writes: the initial snapshot
bootstrap and the post-command snapshot/cwd refresh. User commands still run
under the caller's original umask, while Hermes-owned snapshot and cwd files
are created owner-only.

This intentionally does not copy the source PR's global orphan sweep; deleting
all matching /tmp snapshot files could interfere with concurrent Hermes
processes. The security-critical local disclosure fix is the file mode clamp.

This is salvageable because the source report still identifies a concrete
credential-disclosure path, but the safe subset is smaller than the original
proposal: clamp only the Hermes-owned snapshot writes and leave process-wide
cleanup, user command umask, and concurrent sessions alone.

Salvages source PR: https://github.com/NousResearch/hermes-agent/pull/20056
Related issue: https://github.com/NousResearch/hermes-agent/issues/48441

Co-authored-by: Andrew Homeyer <andrew@hndl.app>
2026-07-07 05:22:42 -07:00
teknium1 5d613a5638 fix(terminal): route init_session bootstrap cd through Windows path conversion
The Windows _quote_cwd_for_cd override only reached _wrap_command; the
snapshot bootstrap cd in init_session still used a bare shlex.quote(),
so on Windows the bootstrap cd failed and pwd -P captured the login
shell's dir instead of terminal.cwd. Route it through _quote_cwd_for_cd
too, and add -- for hyphen-safety to match _wrap_command.
2026-07-01 05:35:34 -07:00
etherman-os 2a3dbcaf46 fix(terminal): prevent corrupted session snapshots during init
The init snapshot dumped functions with a line-based filter:

    declare -f | grep -vE '^_[^_]'

That strips a function's *header* line (e.g. `_foo () `) but leaves the
orphaned `{ ... }` body behind, corrupting the snapshot that is sourced
before every command. Sourcing the torn snapshot runs leftover body code
and breaks subsequent commands (intermittent exit 127).

- Filter private (`_`-prefixed) functions by NAME via `declare -F` and
  dump only the wanted whole definitions, so a body is never torn. Guard
  against an empty name list (bare `declare -f` dumps everything).
- Treat a non-zero bootstrap exit code as snapshot-init failure, so
  execution safely falls back to login-shell-per-command mode.
- Add a regression test asserting snapshot_ready stays false when
  bootstrap exits non-zero.

Preserves the atomic-write ($BASHPID temp + mv -f) machinery from #38249.
2026-06-30 15:51:17 -07:00
Teknium 9f17f16c66 fix(environments): use $BASHPID for atomic snapshot temp + harden failure path
The atomic mv approach (kyssta-exe's commit) narrows but does not close the
#38249 race: the temp name used $$ (parent shell PID), which is identical
across &-launched concurrent subshells. Two concurrent writers pick the same
temp file, clobber each other mid-write, and mv then publishes a torn snapshot
— a reader sourcing it absorbs declare-x/export fragments into PATH.

- Use $BASHPID (actual per-subshell PID) so concurrent writers never collide.
- Chain mv on export success (&&) and rm the temp on failure so a partial dump
  never replaces a good snapshot; apply the same to the init_session bootstrap.
- shlex-quote the static temp-path portion (Windows/spaces), $BASHPID outside.
- LocalEnvironment.cleanup sweeps orphaned snap.tmp.* temps.
- Regression tests: string-shape + a behavioral concurrent writers/readers test
  that proves the snapshot never tears (would still tear with $$).
2026-06-28 02:08:57 -07:00
kyssta-exe 6a2958a521 fix(environments): use atomic file replacement for snapshot writes
Fix race condition in terminal environment snapshots that could corrupt
PATH with declare -x entries. When concurrent terminal calls share the
same snapshot file, the non-atomic 'export -p > snapshot.sh' write could
be read mid-write by another process, causing partial/corrupted env vars
to be sourced and mixed into PATH.

The fix uses atomic file replacement:
- Write to a temp file: export -p > snapshot.sh.tmp.303651
- Atomically replace: mv -f snapshot.sh.tmp.303651 snapshot.sh

On POSIX, mv within the same filesystem is atomic, so source() will
either see the old complete snapshot or the new complete one, never a
partial/truncated file.

Fixes #38249
2026-06-28 02:08:57 -07:00
Gille e7d2f0b93c fix(windows): suppress console flashes and harden gateway restarts 2026-06-25 14:42:38 -07:00
kshitijk4poor 6f8975dcd8 fix(tools): don't compound-rewrite spawn_via_env background wrappers
Background tasks on non-local backends (SSH/Docker/Modal/Daytona/Singularity)
go through `ProcessRegistry.spawn_via_env`, which builds a hand-crafted,
shell-safe wrapper:

    mkdir -p T && ( nohup bash -lc CMD > LOG 2>&1; rc=$?; ... ) & echo $! > PID && cat PID

`BaseEnvironment.execute()` unconditionally ran `_rewrite_compound_background`
on every command, including this wrapper. The rewrite (meant to defuse the
`A && B &` subshell-wait trap for user commands) turns `( ... ) & echo $!` into
`{ ( ... ) & } echo $!` — note `} echo` with no separator, which is a bash
syntax error. The wrapper then never produces a PID, the redirected output file
is never created, and the agent sees an immediate exit code -1. This breaks
*every* background launch on a non-local backend (e.g. a simple
count-and-redirect script over SSH), not just edge cases.

Fix:
- Add `rewrite_compound_background: bool = True` to `BaseEnvironment.execute()`
  (and the `BaseModalExecutionEnvironment` override, which accepts and ignores
  it). Default preserves existing behavior; the user foreground terminal path
  still rewrites.
- `spawn_via_env` passes `rewrite_compound_background=False` so its already
  shell-safe wrapper is left intact.
- Treat a wrapper that produces no PID as a failed launch (mark the session
  exited with a real exit code instead of exposing a fake running session), and
  don't register/checkpoint a session that never started.

Verified empirically: with the rewrite skipped, the wrapper is valid bash,
launches the process, captures the PID, and writes the log/pid/exit files; the
old rewritten form fails `bash -n` with a syntax error.

Based on #33756 by @CharZhou (extracted from a multi-feature branch; the
unrelated image_gen / docker-media changes are not included here).

Co-authored-by: CharZhou <17255546+CharZhou@users.noreply.github.com>
2026-06-01 00:05:10 +05:30
Teknium 90b3c54de9 fix: drain thread no longer crashes on fd-less stdout streams (#34789)
* docs(code-execution): document HERMES_* env narrowing + passthrough workaround

The execute_code sandbox-child env scrub (108397726, #27303) deliberately
dropped the broad HERMES_ prefix passthrough, keeping only an operational
4-var allowlist (HERMES_HOME/PROFILE/CONFIG/ENV). A script that relied on a
non-secret HERMES_* var (HERMES_BASE_URL, HERMES_KANBAN_DB, HERMES_*_WEBHOOK,
or a plugin-defined one) now sees it unset in the child.

Document the behavior change and the two recovery routes (terminal.env_passthrough
in config.yaml, or required_environment_variables in skill frontmatter), plus
the debug log line that surfaces the drop for diagnosis.

* fix: drain thread no longer crashes on fd-less stdout streams

The _wait_for_process drain thread called proc.stdout.fileno()
unconditionally. ProcessHandle implementations whose stdout is not
backed by a real OS fd (iterator-style in-memory streams, mock procs)
raised 'list_iterator' object has no attribute 'fileno' (or
'fileno() returned a non-integer' from select.select), killing the
daemon thread and silently losing all process output.

Resolve the fd defensively at the top of _drain; when stdout has no
usable integer fileno, fall back to draining it as an iterable (the
legacy 'for line in proc.stdout' contract). The real subprocess /
os.pipe-backed select() fast path is unchanged.
2026-05-29 12:16:57 -07:00