Commit Graph

25 Commits

Author SHA1 Message Date
teknium1 9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1 3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
Erosika 4f12985cd2 fix(gateway): refuse relay sender fields from a logged-in client and say what the author trusts
bot_relay.deliver accepted from_profile, from_handle and from_connection from any admitted JSON-RPC client. The handler now refuses them with error 4095 when the calling transport carries a browser login identity, since a logged-in browser never relays for another connection. The DeliveryAuthor docstring now says the author is trusted because an admitted client relays it, not because the sender is verified.
2026-09-10 10:27:07 -07:00
Erosika 3db4defcc1 fix(bot-mode): qualify a relayed author from the local connection too
delivery_turn_author kept the bare bot:<profile> id when the sender's connection was the Desktop's own "local", so a DM relayed from that machine collided with the recipient's profile of the same name. A relayed DM always crosses gateways, so the connection id is now part of the id whenever the Desktop sends one, and only the direct message_agent path in tools/bot_mode_dm.py stays bare. The relay.ts and session_auto_continue.py comments added earlier are cut to one line each.
2026-09-10 10:27:07 -07:00
Erosika d3c8bacbe3 fix(bot-mode): qualify relayed authors with the sender's connection id
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
2026-09-10 10:27:07 -07:00
Erosika 6881e4d3fc fix(bot-mode): carry the relay sender into a live Bot Chat turn
When the target Bot Chat is already open on this gateway, the relay handler delivers through
`prompt.submit` with `queued: true`, and that branch dropped the envelope's sender. The model
still saw the text prefix, but the turn reached the agent unattributed, the exact case the
subprocess branch fixes.

The relay handler now stamps the author on the submit as a `DeliveryAuthor`, an in-process object
a JSON client cannot build, so `prompt.submit` accepts it the way it accepts a hosted-room callback
and refuses a dict with error 4124. The busy queue keeps an authored envelope in its own slot, the
drain hands the author to the turn runner, and the runner passes it to an agent that declares the
keyword. A plain prompt after an authored dm carries no author.

Local deliveries to a desktop-owned Bot Chat take the live-owner mailbox instead. The admission
intent and the mailbox record now carry the author, a retry under the same id with a different
author is refused, and the owner gateway hands the author to the turn it runs. Isolated compute
turns still run unattributed, because the compute-host frame has no author field.
2026-09-10 10:27:07 -07:00
Erosika 16d15869ba feat(bot-mode): carry the sender through local and relay deliveries
a bot dm arrived as an ordinary user message. the only trace of the sender
was the "Message from" text prefix, which the model reads and nothing else
does. the recipient's memory provider saw its own configured user.

message_agent now passes the sender as {"id": "bot:<profile>", "name":
<handle>, "is_bot": true} to the delivery runner (--author <json>), which sets
HERMES_TURN_AUTHOR on the recipient one-shot only. the -Q turn reads it and
passes turn_author into run_conversation. the desktop relay forwards the
envelope's from_profile/from_handle to bot_relay.deliver, which sets the same
variable on its delivery turn. the runner drops any inherited author first so
a delivery without one stays unattributed. the text prefix is unchanged.
2026-09-10 10:27:07 -07:00
Erosika b1915cb02d fix(bot-mode): keep the relay waiter watching past the Desktop deliver deadline
The sender-side waiter gave up at 900s while the Desktop held bot_relay.deliver open for 1500s, so a turn finishing between minute 15 and minute 25 wrote a reply nobody read. REPLY_WAIT_SECONDS now rebuilds the Desktop budget from the same numbers and waits 60s past it. The two turn constants move into tools/bot_relay.py so the gateway handler and the waiter share one definition.
2026-09-10 10:18:27 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 4025afc45a refactor(tools): unify never-raise/liveness helpers and tighten supervisor frame/dialog paths 2026-09-02 22:47:50 -07:00
Teknium 88d1423e99 refactor(tools): compact bot-mode DM/relay/probe and browser supervisor modules 2026-09-02 22:26:43 -07:00
Teknium 6723628de9 refactor(tools/messaging): split send_message into senders/targets tables; dedupe discord/bot_mode/relay helpers; extract feishu_lark shared plumbing; compact graph client/auth 2026-09-02 14:45:42 -07:00
Teknium d4cec15b47 refactor(tools): first-wave simplification of tools/ (file ops split, lazy_deps, code_exec, approval, browser, delegate, mcp, skills, terminal, voice, media)
Behavior-neutral structural pass over tools/*: god-file extractions into
sibling modules (file_operations_common/lint/search, file_tools_paths/
read_tracking/write, code_execution_env/rpc, tool_search_catalog/names/
validation, tts_command_provider, ...), duplicate helper unification,
if/elif -> dispatch tables, dead-code removal, docstring compaction.
Tool schemas (get_tool_definitions) verified byte-identical to base.
2026-09-02 14:43:45 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
kshitij c460e87d10 fix(bot-mode): review follow-ups for the turn lock
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
  probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
  platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
  turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
  wrapping it in --run-delivery would make the child contend with its
  parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
  loaded CI runners; the wait-duration message assert matches ~Ns generally.
2026-08-24 02:39:06 +05:30
kshitijk4poor ac3f9a2dc4 feat(bot-mode): per-profile turn lock — concurrent deliveries queue instead of racing (#93091) 2026-08-24 02:08:53 +05:30
kshitij e00d6c1995 fix(desktop): don't push a live connection as absent when its profile fetch blips
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
2026-08-24 01:07:33 +05:30
kshitijk4poor b96369212c feat(bot-mode): envelope TTL + offline fast-fail for bot relay (#93091 item 2) 2026-08-24 01:05:07 +05:30
kshitijk4poor 64eb6bb7fc feat(bot-mode): typed failure-reason codes for bot turns and relay replies (#93091 item 1) 2026-08-24 00:57:15 +05:30
Teknium 764dba6953 fix(bot-relay): sweep stale relay artifacts + never leak the deliver tempfile
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:

- tools/bot_relay.py: expose the 6h stale sweep as
  cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
  into gateway housekeeping — previously it ran only when the Desktop
  drained the outbox, so plaintext envelopes/replies queued while the
  Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
  try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
  waiter deliveries, which have no plaintext DM tempfile to reclaim.
2026-08-23 03:57:43 -07:00
Teknium d3e087fd8c feat(bot-mode): bots on every Desktop connection can message each other
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.

Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
2026-08-23 02:16:11 -07:00