When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).
- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
request-validation and overflow heuristics, returning format_error with
`retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
verdicts that legitimately reach that branch (policy block, TLS chain, MoA
shape/preset errors) now state `should_fallback=True` explicitly, so the gate
changes behaviour only for the new verdict. Local validation errors keep
their historical fallback.
Fixes#12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.
Co-authored-by: cuyua9 <2114364329@qq.com>
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.
Fixes#12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.
Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
Text and media batching advance the coalesced MessageEvent.message_id to
the newest message but left source.message_id at the first one, so after
a batch the reply anchor and the event id disagreed. Advance both together
at the two enqueue sites.
Same bug class as the Slack/Feishu picks: buzz, dingtalk, email, google_chat,
line, ntfy, photon, sms, teams, wecom and whatsapp already had the platform
message id on the MessageEvent but built the SessionSource without it, so
source.message_id consumers (reply anchor in run.py, /sethome synthetic-thread
check, relay _event_ids fallback, shutdown notice anchor) saw None. Only sites
where the id variable was already in scope are widened.
Invariant test for #9812 (red on origin/main: model_config was {"cwd"} only).
Adapted from the test in PR #9883 to the current save_session() flow, where an
empty-history session stays ephemeral until the first message.
`_persist()` prepared `session_meta` (cwd + provider/base_url/api_mode) but
the create path wrote only `{"cwd": ...}`; the snapshot only landed on a
later update. A restart before that update restored the session with
provider/base_url = None.
Use the same `session_meta` for create and update.
Cherry-pick of PR #9883 by @Ruzzgar (release.py mapping hunk dropped —
already mapped on main).
Fixes#9812
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
`tar xf node-*.tar.xz` shells out to the xz binary; minimal Debian, DietPi
and WSL images ship tar without it, so extraction died mid-way and the
installer then failed on a missing directory. Select .tar.xz only when
`xz` is on PATH, in both the installer and the runtime node bootstrap.
Same approach as the earlier #4229 (@JoshuaMart) and #39541 (@karnull);
#11278 (@vominh1919) attempted an apt-only install of xz-utils instead.
Refs #11197
`re.sub` with the home path as a template string parsed backslashes as
escapes (re.error dropped the whole injected config block); use a callable.
The early `~/…` return also skipped `expandvars`, leaving `~/$LEAF` half
resolved. Prefix-substitute and fall through to normal expansion instead.
Review finding on #109142: the JS bridge's DEFAULT_REPLY_PREFIX (the sender
actually used at runtime), scripts/install.sh, setup-hermes.sh and the site
favicon still carried the old glyph. Python and JS defaults now agree.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
The picked test bound an AF_UNIX socket at pytest's tmp_path, which overflows
the ~108-byte sun_path limit under scripts/run_tests.sh's deep temp root
("AF_UNIX path too long"). Bind by a relative name from inside the temp
HERMES_HOME instead; the walker still sees the same absolute entry.
Also list cache/ + runtime roots and non-regular entries in the `hermes backup`
"What's excluded" docs so the user-visible behaviour change is documented.
Review finding on #109143: `_handle_stream_error` returned True for the
stream_options rejection without checking whether `_call` had another
iteration. With HERMES_STREAM_RETRIES=0, or after the transient budget was
spent, the loop ended with neither a response nor an error set and the call
returned None instead of raising.
The compatibility retry now extends the loop by one attempt exactly once
(`_compat_retries`); the transient budget is untouched. Test pinned with
HERMES_STREAM_RETRIES=0 (red on the previous head).
Azure AI Foundry serverless (MaaS) endpoints validate the request body
strictly and reject `stream_options: {"include_usage": true}` with 422
`extra_forbidden`. Hermes sent the field on every streaming call, so the
agent was unusable against that endpoint family and the fallback chain
failed too.
When a 400/422 names `stream_options` as an extra/unsupported field and no
delta has been delivered yet, re-open the stream without the field and
remember the rejection on the agent (`_stream_options_unsupported`) so later
turns skip it up front. Streaming itself stays on — this is not the
"stream not supported" case. Usage accounting for such endpoints falls back
to the estimator, as it already does for native Gemini.
Salvage of PR #53271 by @DavidMetcalfe, reshaped onto the `_StreamingCall`
retry loop with per-agent state instead of a process-wide host set; one
end-to-end invariant test through `_interruptible_streaming_api_call`.
Fixes#9705
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
The _SKILL_INVALID_CHARS regex stripped all non-ASCII characters from
skill names, so skills with CJK, Cyrillic, or other Unicode names
(e.g. "小说拆条") produced an empty slug and were silently skipped.
Change the regex from [^a-z0-9-] to [\w-] so Unicode word characters
are preserved. Platform-specific sanitizers (Telegram, Discord) already
handle their own character restrictions downstream.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
`skip_context_files=True` kept AGENTS.md/CLAUDE.md out of the system prompt,
but `SubdirectoryHintTracker` was always on, so the first tool call touching
a directory with such a file spliced its full text onto the tool result.
Cron jobs without a workdir (which set skip_context_files) that relay exact
stdout then delivered `[Subdirectory context discovered: ...]` plus the file
body to Telegram/Discord.
The tracker now takes `enabled=` and agent init wires it to
`not skip_context_files`: one flag, both injection paths. Interactive
sessions and cron jobs with a workdir are unchanged.
Direction proposed in PR #9434 (@zhitiao), which gated on platform == "cron";
gating on the existing skip flag covers the same case without a platform
special-case.
Fixes#9441
Gateway sessions and batch workers append to the same default
trajectory_samples.jsonl / failed_trajectories.jsonl with a plain open("a") +
write(); concurrent writers interleaved mid-object and the file stopped
parsing (#12684). The append now holds an exclusive lock for write+flush:
flock on POSIX, a 1-byte msvcrt.locking range on Windows.
Tests: a foreign process holding the lock must block the append (red on
base); six processes appending oversized entries all land parseable.
Fixes#12684. Salvaged from #12685 by @shafdev; Windows arm and test trim ours.
`_render_text_element` ran every inbound post text element through
`_escape_markdown_text`, so a user's `` `print('hi')` `` reached the model as
`` \`print\('hi'\)\` ``. The agent echoed those backslashes in its reply and
Feishu's `md` renderer then showed literal `\*\*bold\*\*` instead of bold.
Post elements carry raw text plus separate style flags; the style wrappers
already rebuild the markdown, so the text itself is passed through untouched.
The outbound path never escaped. Mentions/emoji labels keep their escaping.
Salvage of PR #9991 by @nightq onto `plugins/platforms/feishu/adapter.py`.
Fixes#9816
Review findings on #109139 (the "aeskey half is already on main" claim was
wrong): `_decrypt_file_bytes` restored base64 padding but the key was never
percent-decoded, so `%2F`/`%3D` keys failed to decrypt and `_cache_media`
returned None. And `_store_media` received the CDN's `application/octet-stream`
as the image MIME, which the gateway image classifier rejects even with the
corrected `.png` name.
Decode the key once at the payload boundary; for images forward the response
MIME only when it is `image/*`, otherwise derive it from the resolved
extension. Covered by one end-to-end test through `_cache_media` (red on the
previous head).
WeCom's CDN returns `content-type: application/octet-stream` for images.
`mimetypes.guess_extension` maps that to ".bin", a truthy value that
short-circuited `_guess_extension()` before the magic-byte detector ran, so
every inbound image was cached as `img_*.bin`.
Treat ".bin" as "unknown" so the URL suffix / magic-byte fallback decides.
The second half of the report (URL-encoded, unpadded `aeskey`) is already
handled on main by `_decrypt_file_bytes`.
Salvage of PR #10187 by @nightq onto `plugins/platforms/wecom/media.py`.
Fixes#10085
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).
POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.
Fixes#12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.
Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.