`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).
POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.
Fixes#12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.
Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.
A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.
Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.
Fixes#9516
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.
Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.
Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
Ports #63762 forward onto current main per teknium1's review.
refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).
1. refresh_launchd_plist_if_needed(): return False after retry
exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
of the success message.
Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.
Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.
5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
`build_tool_preview()`'s generic-key fallback and the cute-message helpers
still truncated with a bare `text[:max_len - 3] + "..."`; for max_len 1-3 the
slice goes negative and returns almost the whole string (27 chars for
max_len=1). `_truncate_preview` already had the guard, so the two code paths
disagreed.
One truncation helper (`_tail_trunc`) with the guard, used everywhere; the
head-truncating `_cute_path` gets the same clamp.
Salvage of PR #48483 by @HeLLGURD (current-code fix); the earliest reports and
patches were #9464 (@LarHope), #9477 (@kagura-agent) and #9497.
Co-authored-by: LarHope <12761142+LarHope@users.noreply.github.com>
Fixes#9439
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (3ff18ffe14,
#10447: a server answering errors made the model hammer it 8x in 10s).
Route the recovered result through _record_call_outcome so the caller
sees the tool's answer and the counter still moves the right way.
`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.
Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.
Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.
Fixes#9814
`add_provider()` flipped `_has_external` and appended the provider before
calling `get_tool_schemas()`. When schema loading raised, the broken provider
stayed registered and the single-external slot was poisoned for the rest of
the process: every later provider was rejected as "already registered".
Materialize the schema list first; state changes only after it succeeds.
Exception propagation is unchanged.
Hand-port of PR #9997 by @zhouhe-xydt onto the current add_provider() (the
reserved-core-tool filter landed in between); one invariant test.
Fixes#9948
`_append_to_sqlite` caught and debug-logged its own exceptions, so the outer
handler in `mirror_to_session` never fired and every failed SQLite write was
reported as a successful mirror. Callers (cron in_channel seed, send_message)
had no way to know the transcript was never updated.
Let the write helper raise; the caller already warns and returns False.
Fixes#10130
Same bug class as the wake-gate call fixed in the picked commit: when a caller
invokes _build_job_prompt without a cached prerun_script, the sibling ran the
job's script inline via _run_job_script(script_path) and dropped the job's
configured workdir, so the script ran from the scripts-dir parent. Route it
through _resolve_job_workdir like the no_agent and wake-gate sites already do.
Sweep of every _run_job_script / _run_job_script_with_claim_heartbeat caller:
_run_no_agent_job and cron/monitor.py already pass workdir; the wake-gate site
is fixed by the salvaged commit; this was the last one.
The picked fix inlined the platform ternary inside terminalShellEnv(), which
reads process.platform and can only be exercised by booting Electron. Lift it
into a pure terminalLcCtype(env, platform) that takes the platform as data
(root AGENTS.md: never fake the host OS) and pin the contract with one vitest:
Linux reuses LANG, falls back to C.UTF-8, respects an explicit LC_CTYPE; macOS
keeps the bare "UTF-8" it accepts. Red on origin/main (helper absent), green here.
pyproject extras and the lazy installer must agree on the slack-sdk pin;
the salvaged bump only touched pyproject.toml + uv.lock, so the
`platform.slack` lazy-install spec would still have pulled 3.43.0.
slack-sdk 3.44.1 contains the upstream fix for the aiohttp SocketModeClient
zombie retry loop (slackapi/python-slack-sdk#1956, closing #1913): connect()
used 'while True:' and never checked self.closed, so once close() closed the
shared aiohttp ClientSession, any connect/reconnect task in flight spun
forever logging 'Failed to connect (error: Session is closed); Retrying...'
every ping_interval.
Observed in production on this repo's own Slack adapter: the gateway's
socket watchdog (plugins/platforms/slack/adapter.py) heals wedged sockets by
rebuilding the AsyncSocketModeHandler, but the orphaned connect() task from
the pre-heal client kept retrying against the dead session indefinitely —
22k+ error lines per process per day while Slack itself remained connected.
The adapter's teardown docstring already references slackapi#1913.
3.44.1 adds the self.closed exit; slack-bolt 1.30.0 declares
slack_sdk>=3.38.0,<4, so the bump is compatible.
tornado 6.5.7 is affected by GHSA-5w76-955r-9v8r (CVSS 8.7): parse_multipart_form_data
splits the body unbounded before the max_parts check, so a request with a very large
number of parts can exhaust memory. 6.5.8 caps the split at max_parts+1 so the flooding
part is never materialized.
uv lock --upgrade-package tornado on current main; only the tornado block changes
(13 insertions / 13 deletions in uv.lock), no other package or marker moves.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QghbdVnZSkJbRRDxCSs2zr
The version-less canonical Flash id is not matched by _is_deepseek_thinking_model, so agent.reasoning_effort and every auxiliary reasoning_effort were silently dropped on the OpenCode Go relay while the direct provider was fixed (8435a3ae00/aeecb110f8). Match it through a version-less id set, mirroring plugins/model-providers/deepseek. Verified against the Go relay: named levels are graded (low 906 / high 1371 / max >=2500 reasoning tokens on a multi-step prompt) and integer efforts are rejected, so the named level is the knob to send.
The salvaged comment restated the symptom at length; keep only the WHY.
Drop the `# pragma: no cover` markers (the repo does not gate on coverage).
One invariant test: with aiohttp.web_request lacking RequestKey, `web` stays
bound to the aiohttp module and RequestKey is None (red on origin/main).
On aiohttp < 3.14 the RequestKey import fails and the shared except clause
also resets the already-imported web module to None, so every admission
reply (non-streaming POST /v1/runs) raised AttributeError: 'NoneType'
object has no attribute 'json_response' and surfaced as HTTP 500.
Import the two names in separate try/except blocks; RequestKey already has
None-guards at its use sites.
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
Keep '' -> normal (the bug), None -> normal and priority -> fast (passthrough
sanity), both for a built agent and a pre-build pin, plus the 'status never
writes config' guard. Drop the redundant identity rows and raising=False
(both patched names exist on the module).
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.
Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
The two inline comments on the BaseException branch restated what the code
does; keep only the reason the branch exists (own process group means the
terminal's SIGINT never reaches the hook).
An unquoted redirect target breaks if the pytest tmp dir contains a
space: the hook then never writes its pid and the test waits out the
300 s hook instead of failing fast. The neighbouring tests in this file
are unquoted, so only the new test is changed here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_spawn puts the hook child in its own process group on POSIX, so the
terminal's SIGINT never reaches it and only _spawn's own cleanup can stop
it. That cleanup sat in `except Exception`, and KeyboardInterrupt is a
BaseException — Ctrl+C during a hook skipped kill_process_tree entirely
and left the hook (and anything it forked) running.
Catch BaseException, run the existing tree kill and drain, then re-raise
anything that is not an Exception subclass so KeyboardInterrupt and
SystemExit still reach the caller unchanged. Every other return path is
untouched.
Follow-up to #84901, which added the same cleanup for the timeout path.
The regression test mirrors the existing real-subprocess tests in
tests/agent/test_shell_hooks_tree_kill.py: it interrupts the process once
the hook is confirmed running, then asserts both that KeyboardInterrupt
propagates and that the hook pid is gone. Verified RED on the unfixed
tree (that test alone fails, the other five pass) and green with the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>