Commit Graph

33900 Commits

Author SHA1 Message Date
teknium1 07ead2249b chore(tests): explicit utf-8 encoding in test_claw (windows-footgun ratchet) 2026-09-12 08:24:52 -07:00
Drexuxux 45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1 d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1 fd15cd003e chore(tests): encoding="utf-8" on read_text/write_text in test_personality_single_owner.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:26 -07:00
teknium1 44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1 5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00
teknium1 6ea3b00e7f chore(tests): encoding="utf-8" on read_text/write_text in test_doctor.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:11 -07:00
skyc1e 7d25ae58a7 fix(doctor): image_gen reports "no provider configured", not "system dependency not met"
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.

A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.

Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.

Fixes #9516
2026-09-12 08:24:11 -07:00
teknium1 1dce69be4e chore: contributor mapping for ggascoigne 2026-09-12 08:23:58 -07:00
teknium1 81cec06c77 test(cli): banner hero line is not center-padded
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
2026-09-12 08:23:58 -07:00
Guy Gascoigne-Piggford e38d63b7f0 fix(cli): left-align banner hero art 2026-09-12 08:23:58 -07:00
Season 6d20c321d6 fix(gateway): enable linger on systemd user-service repair path (#12863)
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.

Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.

Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
2026-09-12 08:23:44 -07:00
ygd58 5be0c4921c fix(gateway): surface launchctl bootstrap failures in refresh_launchd_plist_if_needed
Ports #63762 forward onto current main per teknium1's review.

refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).

1. refresh_launchd_plist_if_needed(): return False after retry
   exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
   of the success message.

Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.

Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.

5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
2026-09-12 08:23:44 -07:00
HeLLGURD 9d09598111 fix(display): tool previews never exceed a tiny max_len (1-3)
`build_tool_preview()`'s generic-key fallback and the cute-message helpers
still truncated with a bare `text[:max_len - 3] + "..."`; for max_len 1-3 the
slice goes negative and returns almost the whole string (27 chars for
max_len=1). `_truncate_preview` already had the guard, so the two code paths
disagreed.

One truncation helper (`_tail_trunc`) with the guard, used everywhere; the
head-truncating `_cute_path` gets the same clamp.

Salvage of PR #48483 by @HeLLGURD (current-code fix); the earliest reports and
patches were #9464 (@LarHope), #9477 (@kagura-agent) and #9497.

Co-authored-by: LarHope <12761142+LarHope@users.noreply.github.com>
Fixes #9439
2026-09-12 08:23:42 -07:00
teknium1 dfa84e4a84 chore: map contributor email d.hermsmeier@mittwald.de -> Hermsi1337 2026-09-12 08:23:28 -07:00
teknium1 02b398bda5 fix(mcp): recovered application errors keep the breaker strike
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (3ff18ffe14,
#10447: a server answering errors made the model hammer it 8x in 10s).
Route the recovered result through _record_call_outcome so the caller
sees the tool's answer and the counter still moves the right way.
2026-09-12 08:23:28 -07:00
Dennis Hermsmeier 5dca47a651 fix(mcp): preserve application results after transport recovery 2026-09-12 08:23:28 -07:00
teknium1 b5b697e91d chore: contributor mapping for MonkeyLeeT 2026-09-12 08:23:13 -07:00
MonkeyLeeT 289e0d506f fix(insights): count tool usage per session before merging the two sources
`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.

Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.

Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.

Fixes #9814
2026-09-12 08:23:13 -07:00
teknium1 befa5735e9 chore(tests): encoding="utf-8" on read_text/write_text in test_memory_provider.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:22:57 -07:00
teknium1 f8f16b5e7f chore: contributor mapping for zhouhe-xydt 2026-09-12 08:22:57 -07:00
zhouhe-xydt 5084237f8e fix(memory): load provider schemas before mutating MemoryManager state
`add_provider()` flipped `_has_external` and appended the provider before
calling `get_tool_schemas()`. When schema loading raised, the broken provider
stayed registered and the single-external slot was poisoned for the rest of
the process: every later provider was rejected as "already registered".

Materialize the schema list first; state changes only after it succeeds.
Exception propagation is unchanged.

Hand-port of PR #9997 by @zhouhe-xydt onto the current add_provider() (the
reserved-core-tool filter landed in between); one invariant test.

Fixes #9948
2026-09-12 08:22:57 -07:00
teknium1 eca2cea5c6 chore(tests): encoding="utf-8" on read_text/write_text in test_mirror.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:22:42 -07:00
teknium1 6716f22396 fix(gateway): mirror_to_session reports False when the transcript write fails
`_append_to_sqlite` caught and debug-logged its own exceptions, so the outer
handler in `mirror_to_session` never fired and every failed SQLite write was
reported as a successful mirror. Callers (cron in_channel seed, send_message)
had no way to know the transcript was never updated.

Let the write helper raise; the caller already warns and returns False.

Fixes #10130
2026-09-12 08:22:42 -07:00
teknium1 66884c4397 fix(cron): honour job workdir on the inline _build_job_prompt script path too
Same bug class as the wake-gate call fixed in the picked commit: when a caller
invokes _build_job_prompt without a cached prerun_script, the sibling ran the
job's script inline via _run_job_script(script_path) and dropped the job's
configured workdir, so the script ran from the scripts-dir parent. Route it
through _resolve_job_workdir like the no_agent and wake-gate sites already do.

Sweep of every _run_job_script / _run_job_script_with_claim_heartbeat caller:
_run_no_agent_job and cron/monitor.py already pass workdir; the wake-gate site
is fixed by the salvaged commit; this was the last one.
2026-09-12 08:21:08 -07:00
teknium1 64edb8a9ce chore: map contributor email ycy1990@gmail.com -> metricano 2026-09-12 08:21:08 -07:00
Yong Chang Yi 62b0235d45 fix(cron): pass workdir to agent pre-run scripts 2026-09-12 08:21:08 -07:00
teknium1 5b0103ffdc refactor(desktop): extract terminalLcCtype so the Linux locale rule is testable
The picked fix inlined the platform ternary inside terminalShellEnv(), which
reads process.platform and can only be exercised by booting Electron. Lift it
into a pure terminalLcCtype(env, platform) that takes the platform as data
(root AGENTS.md: never fake the host OS) and pin the contract with one vitest:
Linux reuses LANG, falls back to C.UTF-8, respects an explicit LC_CTYPE; macOS
keeps the bare "UTF-8" it accepts. Red on origin/main (helper absent), green here.
2026-09-12 08:18:50 -07:00
NaviElLay 6f8d975fe7 fix(desktop): use valid LC_CTYPE on Linux terminals 2026-09-12 08:18:50 -07:00
teknium1 c7efbbdab2 chore(deps): mirror slack-sdk 3.44.1 pin in tools/lazy_deps.py
pyproject extras and the lazy installer must agree on the slack-sdk pin;
the salvaged bump only touched pyproject.toml + uv.lock, so the
`platform.slack` lazy-install spec would still have pulled 3.43.0.
2026-09-12 08:12:38 -07:00
teknium1 dc67f00e0b chore: map contributor email bill3wits@gmail.com -> bill3wits 2026-09-12 08:12:38 -07:00
arlai-mk a5437e592d fix(deps): bump slack-sdk to 3.44.1 — fixes aiohttp Socket Mode connect() retrying forever
slack-sdk 3.44.1 contains the upstream fix for the aiohttp SocketModeClient
zombie retry loop (slackapi/python-slack-sdk#1956, closing #1913): connect()
used 'while True:' and never checked self.closed, so once close() closed the
shared aiohttp ClientSession, any connect/reconnect task in flight spun
forever logging 'Failed to connect (error: Session is closed); Retrying...'
every ping_interval.

Observed in production on this repo's own Slack adapter: the gateway's
socket watchdog (plugins/platforms/slack/adapter.py) heals wedged sockets by
rebuilding the AsyncSocketModeHandler, but the orphaned connect() task from
the pre-heal client kept retrying against the dead session indefinitely —
22k+ error lines per process per day while Slack itself remained connected.
The adapter's teardown docstring already references slackapi#1913.

3.44.1 adds the self.closed exit; slack-bolt 1.30.0 declares
slack_sdk>=3.38.0,<4, so the bump is compatible.
2026-09-12 08:12:38 -07:00
Bill fc83fb42d3 deps: bump tornado 6.5.7 -> 6.5.8 (GHSA-5w76-955r-9v8r multipart max_parts DoS)
tornado 6.5.7 is affected by GHSA-5w76-955r-9v8r (CVSS 8.7): parse_multipart_form_data
splits the body unbounded before the max_parts check, so a request with a very large
number of parts can exhaust memory. 6.5.8 caps the split at max_parts+1 so the flooding
part is never materialized.

uv lock --upgrade-package tornado on current main; only the tornado block changes
(13 insertions / 13 deletions in uv.lock), no other package or marker moves.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QghbdVnZSkJbRRDxCSs2zr
2026-09-12 08:12:38 -07:00
teknium1 c7c86a7dc6 chore: map contributor email ada.a0uew@56.pt -> Arguably1341 2026-09-12 08:06:48 -07:00
secsson 456dd2b888 fix: preserve max reasoning for Daybreak alias 2026-09-12 08:06:48 -07:00
Ada 7c734c7838 fix(opencode-go): send reasoning for the canonical deepseek-flash id
The version-less canonical Flash id is not matched by _is_deepseek_thinking_model, so agent.reasoning_effort and every auxiliary reasoning_effort were silently dropped on the OpenCode Go relay while the direct provider was fixed (8435a3ae00/aeecb110f8). Match it through a version-less id set, mirroring plugins/model-providers/deepseek. Verified against the Go relay: named levels are graded (low 906 / high 1371 / max >=2500 reasoning tokens on a multi-step prompt) and integer efforts are rejected, so the named level is the knob to send.
2026-09-12 08:06:48 -07:00
teknium1 21bc1d3775 fix(api-server): trim RequestKey fallback comment, add import-isolation test
The salvaged comment restated the symptom at length; keep only the WHY.
Drop the `# pragma: no cover` markers (the repo does not gate on coverage).
One invariant test: with aiohttp.web_request lacking RequestKey, `web` stays
bound to the aiohttp module and RequestKey is None (red on origin/main).
2026-09-12 08:04:30 -07:00
teknium1 335c3f8843 chore: map contributor email hongzhennan@outlook.com -> Decent9967 2026-09-12 08:04:30 -07:00
hoozge 943e04ba2b fix(gateway): keep aiohttp web import alive when RequestKey is unavailable
On aiohttp < 3.14 the RequestKey import fails and the shared except clause
also resets the already-imported web module to None, so every admission
reply (non-streaming POST /v1/runs) raised AttributeError: 'NoneType'
object has no attribute 'json_response' and surfaced as HTTP 500.

Import the two names in separate try/except blocks; RequestKey already has
None-guards at its use sites.
2026-09-12 08:04:30 -07:00
teknium1 f34933605a test(identity): make the 0600 ledger test prove the mode under a permissive umask
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
2026-09-12 08:02:16 -07:00
codeshipsingh 5a8e651183 fix(identity): enforce 0600 on spawn ledger writes 2026-09-12 08:02:16 -07:00
teknium1 7b6fc1d234 test: trim fast-status regression test to the invariant cases
Keep '' -> normal (the bug), None -> normal and priority -> fast (passthrough
sanity), both for a built agent and a pre-build pin, plus the 'status never
writes config' guard. Drop the redundant identity rows and raising=False
(both patched names exist on the module).
2026-09-12 07:59:59 -07:00
Poxel2 d7ff5675e0 fix: normalize empty fast status to normal without changing tier 2026-09-12 07:59:59 -07:00
teknium1 28bf67b176 chore: map contributor emails for jinliyl and teddychenfeiyang-png 2026-09-12 07:57:13 -07:00
teddychenfeiyang-png a8ac7899b5 test(plugins): typo'd requires_hermes clause fails plugins validate admission
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.

Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
2026-09-12 07:57:13 -07:00
jinli.yl c4c508579d fix(plugins): validate requires_hermes from manifest module 2026-09-12 07:57:13 -07:00
teknium1 761fa5d497 refactor(hooks): fold the interrupt-path comments into one WHY line
The two inline comments on the BaseException branch restated what the code
does; keep only the reason the branch exists (own process group means the
terminal's SIGINT never reaches the hook).
2026-09-12 07:54:57 -07:00
yotamleo f8fbf7c385 test(hooks): quote the marker path in the interrupt test's hook script
An unquoted redirect target breaks if the pytest tmp dir contains a
space: the hook then never writes its pid and the test waits out the
300 s hook instead of failing fast. The neighbouring tests in this file
are unquoted, so only the new test is changed here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 07:54:57 -07:00
yotamleo bca94759d5 fix(hooks): reap the hook process tree when _spawn is interrupted
_spawn puts the hook child in its own process group on POSIX, so the
terminal's SIGINT never reaches it and only _spawn's own cleanup can stop
it. That cleanup sat in `except Exception`, and KeyboardInterrupt is a
BaseException — Ctrl+C during a hook skipped kill_process_tree entirely
and left the hook (and anything it forked) running.

Catch BaseException, run the existing tree kill and drain, then re-raise
anything that is not an Exception subclass so KeyboardInterrupt and
SystemExit still reach the caller unchanged. Every other return path is
untouched.

Follow-up to #84901, which added the same cleanup for the timeout path.

The regression test mirrors the existing real-subprocess tests in
tests/agent/test_shell_hooks_tree_kill.py: it interrupts the process once
the hook is confirmed running, then asserts both that KeyboardInterrupt
propagates and that the hook pid is gone. Verified RED on the unfixed
tree (that test alone fails, the other five pass) and green with the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 07:54:57 -07:00