Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:
- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
try shutil.which('hermes') before the bare-name fallback, so
environments with a PATH but no venv sibling resolve exactly what an
interactive shell would. Platform test switched os.name -> sys.platform
('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
errors='replace' on both subprocess.run sites — without them the
child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
with which=None, and encoding-pin assertions in the deliver transport
test.
Refs #93590, #93597, #93601
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):
1. waiter_command embeds the reply path in generated python -c source
with !r. repr escapes each backslash, but the Windows execution layer
folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
and SyntaxErrors the whole waiter script. Raw-string literals keep
the folded single backslash a literal; POSIX paths have no
backslashes so the prefix is a no-op there, and \' inside a raw
literal still cannot terminate the string, keeping the #93091
injection defense intact.
2. local_delivery_command hardcoded "hermes", relying on PATH — absent
in service contexts (systemd units, desktop launchers, non-login SSH
shells), so delivery died with ENOENT. It now resolves the CLI next
to this gateway's own interpreter (venv bin/Scripts sibling,
hermes.exe on Windows) with a bare-name fallback. The #93091
per-profile turn-lock recognition in bot_mode_dm now matches the CLI
element by basename (split on both separators) so resolved absolute
paths still take the lock instead of silently bypassing it.
Fixes#93590
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:
- Desktop relay drain forwards bot_relay.deliver's error.data.reason
into bot_relay.reply (and prefers it for the attention badge over
free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
text, so the completion notification the sending agent receives is
machine-branchable.
Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.
bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
wrapping it in --run-delivery would make the child contend with its
parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
loaded CI runners; the wait-duration message assert matches ~Ns generally.
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:
- tools/bot_relay.py: expose the 6h stale sweep as
cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
into gateway housekeeping — previously it ran only when the Desktop
drained the outbox, so plaintext envelopes/replies queued while the
Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
waiter deliveries, which have no plaintext DM tempfile to reclaim.
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.
Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.