Use the ACP v1 session config contract advertised by session/new: locate the category=model option and apply the selected value through session/set_config_option. Retain session/set_model only as compatibility fallback for pre-configOptions agents. Reject unknown and policy-disabled values before prompting.
Verified against the installed Copilot ACP server: its model config option advertises the account-authorized choices, session/set_config_option returns the updated state, and live prompts route gpt-5.6-terra to Terra and claude-sonnet-5 to Sonnet 5.
Follow-up to the session/set_model wiring, caught in live use: picking an
org-policy-disabled model (claude-fable-5) produced a response claiming to
BE that model while Copilot actually served its default (Claude Sonnet 5).
Two causes:
1. The prompt preamble injected 'Hermes requested model hint: <id>', so
whatever model actually served the session parroted the requested name
back as its identity. Remove the line entirely — the model is applied
for real via session/set_model now, and identity must come from the
backend, not prompt suggestion.
2. session/new advertises policy-disabled ids alongside enabled ones
(_meta.copilotEnablement: 'disabled'); selecting one is accepted but
silently serves the default. Exclude disabled ids from the offered set
so the degrade-with-warning path handles them.
Verified live: requesting claude-fable-5 logs the does-not-offer warning
listing the 23 genuinely enabled models, serves the default, and the
response truthfully self-identifies as Claude Sonnet 5.
Selecting a model on the copilot-acp provider had no effect: the model id
never left Hermes. _create_chat_completion() dropped the model argument
before _run_prompt(), so the selection survived only as prompt text
('Hermes requested model hint: ...') and Copilot answered with its own
session default — a user picking gpt-5.6-terra visibly got Claude Sonnet 5.
Live-probing 'copilot --acp --stdio' shows the CLI validates but IGNORES
its --model spawn flag in ACP mode, while session/new advertises
models.availableModels and the ACP-native session/set_model call actually
switches the session. Wire that in: forward the model into _run_prompt,
and after session/new send session/set_model when the id is advertised
(or the server reports no list). Unknown ids degrade to the session
default with a warning instead of failing the turn; the provider-level
virtual slug 'copilot-acp' is never forwarded.
Verified live against the real CLI: requesting gpt-5.6-terra answers as
GPT-5.6 Terra and claude-sonnet-5 answers as Claude Sonnet 5.
Salvage hardening on top of #87308 (thanks @Dudeman456):
- Tri-state verdict: inconclusive probes (binary missing, --help
failed/timed out) return None and fall through to the normal spawn
path, preserving the established 'Could not start Copilot ACP
command' error instead of masking it. This also fixes the two
test_copilot_acp_client HOME-env regressions that went red on the
PR: their mocked-Popen path was intercepted by the new unmocked
subprocess.run probe.
- Cache definitive verdicts per binary path so CLIs that DO support
--acp pay the ~50ms --help cost once per process, not per prompt.
- Skip the probe entirely when custom ACP args don't include --acp.
- Fix the help-text regex: the old pattern never matched '[--acp]'
(leading '[' is neither start-of-string nor whitespace) and \b
after 'p' matched '--acpfoo'.
- Hermeticity: stub subprocess.run in the two HOME-env tests; add 6
probe-specific tests (fast-fail, fall-through, caching, skip).
The restart-routing, systemd-support, and subprocess-HOME tests asserted
branch behavior but left part of the real probe surface unmocked, so they
fail when the suite itself runs inside a container (self-hosted CI) or a
launchd-descended shell:
- /restart routing tests: the handler also consults the real /.dockerenv —
extract the inline probe to gateway.restart.is_container_restart_context()
(patchable seam, no behavior change) and pin it False; scrub ALL four
supervisor env markers (ambient XPC_SERVICE_NAME on macOS flipped one).
- supports_systemd_services tests: pin shutil.which('systemctl') and
is_container() so the test asserts the branch, not the host.
- copilot ACP real-HOME test: pin is_container() (auto mode prefers profile
home in containers) and scrub ambient HERMES_REAL_HOME/TERMINAL_HOME_MODE.
97 tests green on macOS dev box AND inside a docker CI runner container.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Several file-I/O call sites still use open() / Path.read_text() /
Path.write_text() without an explicit encoding, so they fall back to
the platform default. On Windows CN/JP/KR locales (GBK/CP932/CP949)
any non-ASCII byte in a config/state/user-content file raises
UnicodeDecodeError or UnicodeEncodeError and crashes the caller.
to the remaining hot paths:
- agent/copilot_acp_client.py: fs/read_text_file and fs/write_text_file
(Copilot's read_file / write_file tools,
directly reported in #18637 bug 2)
- agent/model_metadata.py: context-length YAML cache load + two
save sites (context probing is on the
call path of every model invocation)
- agent/nous_rate_guard.py: cross-session rate-limit JSON state
(read + atomic write via os.fdopen)
- cron/scheduler.py: user config.yaml read in run_job
- gateway/delivery.py: cron output writes for AI-generated
content, very likely non-ASCII
yaml.dump call sites also gain allow_unicode=True so the emitted
YAML preserves non-ASCII chars as-is instead of emitting \u escape
sequences.
Adds regression tests that monkeypatch builtins.open / Path.read_text
/ Path.write_text to simulate a GBK locale: each test raises
UnicodeDecodeError / UnicodeEncodeError unless the caller explicitly
passes encoding='utf-8'. Verified that the tests fail on main and
pass with this change, on Linux as well as on Windows.
Refs #18637
Return actionable errors when HERMES_WRITE_SAFE_ROOT blocks a path instead of
labeling every denial as a protected credential file. Wire the helper through
write_file, patch, delete/move, and the Copilot ACP shim; sync docs examples.
CI Tests workflow has been red on main for 40+ consecutive runs. This
commit recovers every failure visible in run 25130722163 (most recent
completed run prior to this PR).
Root causes, by group:
Test-mock drift after product landed (fix: update mocks)
- test_mcp_structured_content / test_mcp_dynamic_discovery (6 tests):
product added _rpc_lock (#02ae15222) and _schedule_tools_refresh
(#1350d12b0) without updating sibling test files. Install a real
asyncio.Lock inside the fake run-loop and patch at _schedule_tools_refresh.
- test_session.py: renamed normalize_whatsapp_identifier → canonical_
whatsapp_identifier upstream; keep a local alias so the legacy tests
keep working.
- test_run_progress_topics Slack DM test: PR #8006 made Slack default
tool_progress=off; explicitly set it to 'all' in the test fixture so
the progress-callback path still runs. Also read tool_progress_callback
at call time rather than freezing it in FakeAgent.__init__ — production
assigns it AFTER construction.
- test_tui_gateway_server session-create/close race: session.create now
defers _start_agent_build behind a 50ms timer — wait for the build
thread to enter _make_agent before closing, otherwise the orphan-
cleanup path never runs.
- test_protocol session.resume: product get_messages_as_conversation now
takes include_ancestors kwarg; accept **_kwargs in the test stub.
- test_copilot_acp_client redaction: redactor is OFF by default (snapshots
HERMES_REDACT_SECRETS at import); patch agent.redact._REDACT_ENABLED=True
for the duration of the test.
- test_minimax_provider: after #17171, dots in non-Anthropic model names
stay dots even with preserve_dots=False. Assert the new invariant
rather than the old 'broken for MiniMax' behavior.
- test_update_autostash: updater now scans `ps -A` for dashboard PIDs;
the test's catch-all subprocess.run stub needed stdout/stderr fields.
- test_accretion_caps: read_timestamps dict is populated lazily when
os.path.getmtime succeeds. Use .get("read_timestamps", {}) to tolerate
CI filesystems where the stat races file creation.
Change-detector tests (fix: rewrite as structural invariants)
- test_credential_sources_registry_has_expected_steps: was a frozen set
comparison that broke when minimax-oauth was added. Rewrite as an
invariant check (every step has description, no dupes, core steps
present) per AGENTS.md 'don't write change-detector tests'.
xdist ordering / test pollution (fix: reset state, use module-local patches)
- test_setup vercel: sibling test saved VERCEL_PROJECT_ID='project' to
os.environ via save_env_value() and never cleared it. monkeypatch.delenv
the VERCEL_* vars in the link-file test.
- test_clipboard TestIsWsl: GitHub Actions is on Azure VMs whose real
/proc/version often contains 'microsoft'. Patching builtins.open with
mock_open didn't reliably intercept hermes_constants.is_wsl's call in
xdist workers that had already cached _wsl_detected=True from an
earlier test. Patch hermes_constants.open directly and add
teardown_method to reset the cache after each test.
Pytest-asyncio cancellation hangs (fix: bound product await with timeout)
- test_session_split_brain_11016 (3 params) + test_gateway_shutdown
cancel-inflight: under pytest-asyncio 1.3.0, 'await task' and
'asyncio.gather(cancelled_tasks)' can stall for 30s when the cancelled
task's finally block awaits typing-task cleanup. Bound both with
asyncio.wait_for(..., timeout=5.0) and asyncio.shield — the stragglers
are released from adapter tracking and allowed to finish unwinding in
the background. This is also a legitimate hardening: a wedged finally
shouldn't stall the caller's dispatch or a gateway shutdown.
Orphan UI config (fix: merge tiny tab into messaging category)
- test_web_server test_no_single_field_categories: the telegram.reactions
config field lived in its own 'telegram' schema category with no
siblings. Fold it under 'discord' via _CATEGORY_MERGE so the dashboard
doesn't render an orphan single-field tab.
Local verification: 38/38 originally-failing tests pass; 4044/4044
gateway tests pass; 684/684 targeted subset (all 16 touched test files)
passes.
Pass an explicit HOME into Copilot ACP child processes so delegated ACP runs do not fail when the ambient environment is missing HOME.
Prefer the per-profile subprocess home when available, then fall back to HOME, expanduser('~'), pwd.getpwuid(...), and /home/openclaw. Add regression tests for both profile-home preference and clean HOME fallback.
Refs #11068.
file_safety now uses profile-aware get_hermes_home(), so the test
fixture must override HERMES_HOME too — otherwise it resolves to the
conftest's isolated tempdir and the hub-cache path doesn't match.