Same site class as the e2e fixtures: tests-js/scripts/mock-server.ts writes
the config for `npm run dev:mock`. Entry-level providers.<name>.context_length
is honoured since #98387, so a 4096 window now fails agent init below
MINIMUM_CONTEXT_LENGTH exactly as it did for the Playwright fixtures.
The two get_hermes_dir symlink tests wrapped symlink_to() in a
try/except -> pytest.skip(). tests/conftest.py already has the
require_symlinks marker for exactly this (probes symlink support once,
skips at setup); using it keeps the test bodies straight-line and the
skip reason uniform with the sibling tests in this file.
Ladder rung 3 says a service-gated tool 'only appears when a prerequisite
is configured', while the SESSION-surface section below says check_fn must
not gate per-session surface. Both are correct but read as a contradiction
at the ladder; add the disambiguation inline with a pointer to the section.
asyncio.start_unix_server only exists where an AF_UNIX event loop does.
On native Windows the attribute is absent, so patch.object raised
AttributeError while arming the tripwire — the TCP-witness test could
only ever pass on POSIX, the platform it pretends not to be. create=True
arms the forbidden-call tripwire on every platform; mock removes the
created attribute on exit, so no cross-test leakage.
Before: AttributeError on native Windows. After: 4/4 pass on Windows
(verified) — and unchanged on POSIX, where the attribute exists.
Five tests were silently failing on the Windows CI job because they
either required a Linux/macOS runtime or created symlinks (which the
default Windows user cannot do without Developer Mode).
Mark them with the project's existing linux_only / require_symlinks
markers so CI on Windows skips them cleanly instead of reporting red:
* TestPosixNoOp, test_noop_on_posix, test_no_hermes_home_returns_native
-> @pytest.mark.linux_only (the asserts only make sense on POSIX
where '~/.hermes' is the natural HERMES_HOME fallback).
* test_atomic_replace_copy_fallback_preserves_symlink,
test_symlinked_target_survives_a_contended_rename
-> @pytest.mark.require_symlinks (matches the pattern already used
by sibling tests in the same file).
For the two symlink tests in test_hermes_constants.py that weren't
covered by require_symlinks, wrap the symlink_to() call in a
try/except(OSError, NotImplementedError) -> pytest.skip() so a host
without symlink privilege still skips instead of erroring out.
No production code touched. Verified locally on win32 Python 3.11.9:
88 passed, 25 skipped, 0 failed across the three files (down from
7 failed, 88 passed).
The `/browser connect` tip said Chrome 136+ "silently refuse[s] to open the
remote debugging port" on the default user-data-dir and that "there is no error
message". Chrome does surface it: on Chrome 152/macOS, launching the default
profile with --remote-debugging-port puts up an "Allow remote debugging?" dialog,
and the port opens once you press Allow. Port 9222 refusing connections is the
symptom while that consent is pending, not a permanent silent block.
Two details this cost time to rediscover, now written down:
- the consent is asked per launch, so it reappears on every browser restart
- the remote-debugging toggle in chrome://settings does not suppress it; that
toggle only makes the feature available
Also cross-references browser.use_real_profile for the case the tip leaves
unanswered — wanting existing logins in the agent's browser. A dedicated
user-data-dir avoids the dialog but starts signed out; the real-profile snapshot
gets both, since the snapshot copy is itself a non-default user-data-dir.
Verified locally: default profile + --remote-debugging-port=9222 shows the
dialog and leaves 9222 closed, while a --user-data-dir launch answers
/json/version in under a second with no prompt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_finish_turn` gated the history update on truthiness, so a turn that
legitimately returned an empty transcript left the previous history
in place and the client kept rendering stale state. Same falsey-vs-absent
class as the step-callback fix (#10845); gate on key presence.
Refs #10844
make_step_cb() used 'or' to fall back from 'result' to 'output' key:
result = tool_info.get('result') or tool_info.get('output')
This collapsed valid falsey results (empty string, 0, False) to None,
because Python's 'or' operator treats all falsey values as missing.
Use explicit key presence check instead:
result = tool_info.get('result') if 'result' in tool_info else tool_info.get('output')
This preserves the actual value when the 'result' key exists (even if
falsey), and only falls back to 'output' when 'result' is truly absent.
Adds regression tests for empty string, zero, and output-key fallback.
Fixes#10845.
The generic .filterBtn:hover rule outranks .filterActive and set the label to emphasis-1000 (near-black) on the emphasis-900 pill background, so hovering the selected filter made it unreadable in light mode and left the #91491 contrast fix incomplete. Add a .filterActive:hover rule keeping the emphasis-0 foreground. Dark mode is unaffected: [data-theme='dark'] .filterActive ties .filterBtn:hover on specificity and wins by source order.
The chat page's mobile side panel reads the same --component-sidebar-background
skin var as App.tsx and renders transparent under skins that do not define it,
so it gets the same var(--background-base) fallback. Trim the two salvaged CSS
comments to the WHY.
Active filter pills on /docs/user-stories used color: var(--ifm-background-color)
on a near-black background, but Infima's light theme resolves that variable to
transparent because custom.css only overrides it under [data-theme='dark'].
Default-active "All" pills therefore rendered as empty black blobs (#91491).
Switch to --ifm-color-emphasis-0 for theme-aware contrast. Apply the same fix to
skills page pick-button hover labels.
Co-authored-by: Cursor <cursoragent@cursor.com>
--component-header-background and --component-sidebar-background
resolve to an empty string in some themes. An unresolved CSS custom
property makes the whole background declaration invalid, so the
browser falls back to transparent -- the mobile hamburger menu then
renders with a see-through sidebar and page text bleeds through
behind the nav items, even though the click-to-close backdrop overlay
itself is fully opaque and working correctly.
Add --background-base as an explicit fallback in both var() calls so
the sidebar and mobile header always paint an opaque background, with
themes still free to override via the component tokens.
The job header was a single non-wrapping flex row carrying the title and up
to seven badges. On a phone the badges overflowed the card, collided with
the action buttons, and squeezed the title to zero width, so the scheduled
jobs showed no names at all.
flex-wrap lets the badges fall to the next line, and min-w-0 on the title
lets its existing truncate take effect instead of being ignored inside a
flex row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The final assertion read `aiohttp.web_request.RequestKey` directly, which
AttributeErrors on the very aiohttp versions (< 3.14) the guarded import
exists to support. Compare against getattr(..., None) so the test checks
the module rebinds to whatever the real aiohttp provides.
Follow-up to #109157.
`_append_entry` moved to `utils.atomic_json_write`, which dumps with
`ensure_ascii=False` through a utf-8 text handle. An argv token holding
surrogate-escaped bytes (a non-UTF-8 project path via os.fsdecode) makes
json.dump raise UnicodeEncodeError — a ValueError, so the `except OSError`
does not catch it and callers silently lose their registration.
Expose `ensure_ascii` on `atomic_json_write` (default unchanged) and pass
True at the ledger call site, restoring the previous json.dumps behaviour
while keeping mode=0o600. One round-trip test, red on base.
Follow-up to #109156.
`_start_gateway_shutdown_tail()` returned False on `should_exit_with_failure`
before `cron_stop.set()`, the cooperative thread waits, the planned-stop
watcher stop and MCP shutdown, so a failure exit leaked the cron ticker and
housekeeping daemon threads (and open MCP connections) for embedded/library
callers. The verdict is now resolved after the teardown, matching the
startup-abort path which already shuts MCP down first.
Fixes#12175. Salvaged from #55031 by @DavidMetcalfe, re-applied onto the
extracted shutdown tail with one thread-lifecycle invariant test.
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
Two CI regressions from the previous commit, both mine:
- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
`should_fallback` to True for the MoA adapter-shape and missing-preset
verdicts. #55933 made those deliberately NOT fall back (a fallback would
silently replace the MoA route with a single model); restore
`retryable=False` only. The gate in `settle_unrecovered_error` now honours
that for real: on main these verdicts never reached the fallback branch
because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
`_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
block instead.
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).
- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
request-validation and overflow heuristics, returning format_error with
`retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
verdicts that legitimately reach that branch (policy block, TLS chain, MoA
shape/preset errors) now state `should_fallback=True` explicitly, so the gate
changes behaviour only for the new verdict. Local validation errors keep
their historical fallback.
Fixes#12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.
Co-authored-by: cuyua9 <2114364329@qq.com>
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.
Fixes#12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.
Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
Text and media batching advance the coalesced MessageEvent.message_id to
the newest message but left source.message_id at the first one, so after
a batch the reply anchor and the event id disagreed. Advance both together
at the two enqueue sites.
Same bug class as the Slack/Feishu picks: buzz, dingtalk, email, google_chat,
line, ntfy, photon, sms, teams, wecom and whatsapp already had the platform
message id on the MessageEvent but built the SessionSource without it, so
source.message_id consumers (reply anchor in run.py, /sethome synthetic-thread
check, relay _event_ids fallback, shutdown notice anchor) saw None. Only sites
where the id variable was already in scope are widened.
Invariant test for #9812 (red on origin/main: model_config was {"cwd"} only).
Adapted from the test in PR #9883 to the current save_session() flow, where an
empty-history session stays ephemeral until the first message.
`_persist()` prepared `session_meta` (cwd + provider/base_url/api_mode) but
the create path wrote only `{"cwd": ...}`; the snapshot only landed on a
later update. A restart before that update restored the session with
provider/base_url = None.
Use the same `session_meta` for create and update.
Cherry-pick of PR #9883 by @Ruzzgar (release.py mapping hunk dropped —
already mapped on main).
Fixes#9812
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
`tar xf node-*.tar.xz` shells out to the xz binary; minimal Debian, DietPi
and WSL images ship tar without it, so extraction died mid-way and the
installer then failed on a missing directory. Select .tar.xz only when
`xz` is on PATH, in both the installer and the runtime node bootstrap.
Same approach as the earlier #4229 (@JoshuaMart) and #39541 (@karnull);
#11278 (@vominh1919) attempted an apt-only install of xz-utils instead.
Refs #11197
`re.sub` with the home path as a template string parsed backslashes as
escapes (re.error dropped the whole injected config block); use a callable.
The early `~/…` return also skipped `expandvars`, leaving `~/$LEAF` half
resolved. Prefix-substitute and fall through to normal expansion instead.
Review finding on #109142: the JS bridge's DEFAULT_REPLY_PREFIX (the sender
actually used at runtime), scripts/install.sh, setup-hermes.sh and the site
favicon still carried the old glyph. Python and JS defaults now agree.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
The picked test bound an AF_UNIX socket at pytest's tmp_path, which overflows
the ~108-byte sun_path limit under scripts/run_tests.sh's deep temp root
("AF_UNIX path too long"). Bind by a relative name from inside the temp
HERMES_HOME instead; the walker still sees the same absolute entry.
Also list cache/ + runtime roots and non-regular entries in the `hermes backup`
"What's excluded" docs so the user-visible behaviour change is documented.
Review finding on #109143: `_handle_stream_error` returned True for the
stream_options rejection without checking whether `_call` had another
iteration. With HERMES_STREAM_RETRIES=0, or after the transient budget was
spent, the loop ended with neither a response nor an error set and the call
returned None instead of raising.
The compatibility retry now extends the loop by one attempt exactly once
(`_compat_retries`); the transient budget is untouched. Test pinned with
HERMES_STREAM_RETRIES=0 (red on the previous head).