Commit Graph

1447 Commits

Author SHA1 Message Date
Teknium c4f96565d0 refactor(adapters/whatsapp_webhook): 6449->4465; whatsapp_cloud/common/plugin send + media unification, webhook route/filter dispatch tables 2026-09-02 14:06:38 -07:00
Teknium 40532d1b7f refactor(adapters/qq_signal): 8720->6042; qqbot keyboards/chunked_upload/onboard dedupe, signal send/receive helpers, bluebubbles + msgraph_webhook helper unification 2026-09-02 14:06:37 -07:00
Teknium 9837fd0bc4 refactor(adapters/yuanbao_weixin): 10420->6889; yuanbao_proto spec-table encode/decode (-46%), unify weixin HTTP/send/media/typing helpers, drop dead send_file/register_handler/encode_forward_msg_data, duplicate _on_delta 2026-09-02 14:06:33 -07:00
Teknium 581d97e545 refactor(adapters/api_server): 10589->8871; extract OpenAI-compatible routes into api_server_openai_routes.py, dedupe SSE/response builders, drop dead idempotency acknowledge path 2026-09-02 14:06:32 -07:00
Teknium d3db321828 refactor(adapters/base): base.py 7719->5348, helpers.py 942->640; extract _process_message_background helpers, unify media send/fallback/cache/notice paths, drop dead _resolve_extensionless_candidate + TextBatchAggregator, compact docs 2026-09-02 14:06:31 -07:00
Teknium 661fc669a0 refactor(adapters): add gateway/platforms/_shared.py (scoped-secret / profile-scope / port helpers) and lift shared text-batching into BasePlatformAdapter
Unifies 25 per-adapter copies of _get_scoped_secret/_get_wsecret/_resolve_qq_secret/_sig_secret/_yb_secret, 4+ _profile_scoped*, 3 _coerce_port, 5 _text_batch_key and 4 _enqueue_text_event.
2026-09-02 14:06:31 -07:00
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Hudson db639e1023 fix(gateway): bind every adapter to the runner at the _create_adapter boundary
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.

Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.

Salvaged from #70831 (Hudson). First reported in #68332.
2026-09-02 06:47:45 -07:00
Teknium 7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Teknium 74775df53f fix(gateway): route-stamp primary callback auth and carry is_bot through the adapter auth check
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.

- `_make_adapter_auth_check`: for the shared primary adapter under
  multiplex, mirror the inbound message path exactly — stamp the
  `profile_routes` match so the routed profile's pairing store is
  consulted, and authorize under the TRANSPORT home via
  `_is_user_authorized_for_source` (same split as
  `_make_default_profile_message_handler`, 2afed50863). A rejected route
  fails closed like the ingress gate. Retain the receiving adapter as
  `_transport_adapter_ref` so config.yaml policy reads stay on it.
  Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
  `thread_id` as keywords only when set, so legacy 3-positional callbacks
  keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
  the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
  honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
  introspection class — fall back to the injected `gateway_runner` and
  the adapter's owner profile.

Fixes #86296
Fixes #92840

Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Celio Monteiro d55d9d128a fix(gateway): route profile into topic prune, cooldowns, and docs
Address hermes-sweeper review on #76487:

- Prefer hermes_profile from send metadata when pruning stale topic
  bindings so profile_routes cannot delete the transport adapter's
  namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
  (profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
2026-09-02 05:59:24 -07:00
muhifni 1cd736ff63 fix(terminal): scope terminal config per turn under profile multiplexing
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.

Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:

- tools/terminal_scope.py: ContextVar holding the routed profile's
  COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
  TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
  resolves ONLY from it - an omitted key yields the defined default,
  never os.environ. Unreadable/malformed policy installs a refusal
  scope; terminal_tool / execute_code refuse instead of running under
  ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
  _profile_runtime_scope, tui_gateway session/build/turn scopes, cron
  per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
  (_get_env_config, _resolve_container_task_id shared key, orphan
  reaper lifetime, degraded mode), gateway/platforms/base.py docker
  media translation (volumes, shared key, persistence), runtime_cwd /
  agent_init / skill_utils / code_execution_tool / file_tools cwd
  anchors, prompt_builder / browser_tool / env_probe backend checks,
  gateway footer, @-refs and slash-command cwd. env_probe resolves the
  backend in the caller's context, since the probe worker thread does
  not inherit the ContextVar.

Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.

Fixes #68559
Fixes #94200
Fixes #101132
Fixes #95470

Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
2026-09-02 05:34:28 -07:00
Teknium 45b0d8cab5 feat(gateway): one gateway.trust_env key controls aiohttp proxy-env honoring at every adapter site (#48820 bug 3)
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.

- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
  resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
  auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
  teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
  bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.

Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
2026-09-02 04:13:02 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
JonthanaHanh 5360886f54 fix(gateway): exclude Ollama Cloud from GLM truncation detection; propagate partial flag (#72316)
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:

1. _is_ollama_glm_backend() matched "ollama" in base URL, which
   included Ollama Cloud (ollama.com). The hosted service correctly
   reports finish_reason and is not affected by the local Ollama
   stop-reason bug.  Exclude "ollama.com" before the substring check.

2. _handle_session_chat_stream() hardcoded "partial": False in the
   assistant.completed SSE event instead of reading result.get("partial").
   The WebUI could not detect truncation and rendered partial responses
   incorrectly (showing only the continuation instead of the full text).
   Read the partial flag from the agent result, matching the pattern
   used by other SSE paths in the same file.

Fixes #72316
2026-09-01 23:27:10 -07:00
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos 3739cf3b86 fix(api): resolve the declared conversation instead of minting a session per request
POST /v1/responses and POST /v1/runs parse and authenticate the client's
X-Hermes-Session-Key, pass it downstream for memory scoping, and then mint a
throwaway physical session id anyway whenever the client manages its own
history (no previous_response_id chain to carry one forward).

Every conversation-affinity hint Hermes sends is derived from that physical
id, so all four re-keyed on every single reply: prompt_cache_key on both
OpenAI-wire transports, the OpenRouter and Nous sticky session_id, and xAI's
x-grok-conv-id. The conversation never landed back on a warm prefix.

Fix the identity rather than the four consumers. The declared key resolves to
its live session through find_latest_gateway_session_for_peer -- the same
reset-fenced recovery every native gateway platform already uses -- and the
turn records the row it ended on through record_gateway_session_peer, which
AIAgent._ensure_db_session never did (it knows the key and writes the row
unkeyed, so the mapping the next reply needs did not exist).

Because the lookup is fenced on sessions.end_reason, the generation that must
rotate is already durable: session_reset (/new), session_switch, idle, daily,
suspended and resume_pending_expired all return None, so a new conversation
gets a new id and a cold affinity scope, and a retired generation can never be
resolved again. No counter, no new persisted field, and no new precedence rule
in the cache-scope resolver -- /branch, delegate and tool children keep the
isolation of #79161/#79017 byte for byte.

Precedence is unchanged where it already worked: an explicit body session_id
and the previous_response_id chain both still outrank the declared key, and a
request that declares nothing keeps its per-request id. Recording is opt-in
(bind_declared_conversation), so no other _run_agent caller's rows change.

Refs #96811

(cherry picked from commit e7c83dddf36784d1012bf483240ebc7f6b2ef9aa)
2026-09-01 02:14:35 -07:00
EmpireOperating 00394acfae fix(buzz): ingest verified native attachments 2026-08-31 10:06:34 -07:00
Mathias Gorf aaad054330 fix(buzz): gate authenticated inbound media on explicit authorization
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.

`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.

Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
2026-08-31 10:06:34 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00
kshitijk4poor c0ff25a1f8 fix(gateway): stop blocking the event loop — off-loop hot sites + ASYNC lint ratchet
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).

Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
  now runs via asyncio.to_thread — a webhook delivery no longer freezes
  every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
  5s worst case) become await asyncio.sleep() (re-lands #36163 at current
  line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
  the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
  move off the loop.

Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
  (blocking HTTP / Popen / subprocess.run / time.sleep in async def).
  These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
  sites (detached restart watchers; router sweep in flight via #84376;
  two platform adapters), each documented for burn-down. New files or
  new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).

Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
  subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
  liveness DURING delivery (ticker coroutine): 1 tick on the old
  blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.

Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
2026-08-28 07:50:51 -07:00
Teknium 31579f781e fix(cron): transient run prompt survives the relay-fronted gateway forward
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
2026-08-27 20:53:02 -07:00
Teknium 34393c32aa feat(plugins): wire plugin platform handlers into a2a, buzz, and qqbot adapters
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
2026-08-27 07:51:37 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
Tranquil-Flow 6cbb7b6115 fix(gateway): preserve exception type when error string is empty (#78183)
httpx timeout exceptions (ReadTimeout, WriteTimeout) stringify to "",
which defeats _is_timeout_error's first-line guard (if not error: return
False).  The base-layer plain-text fallback then re-sends an already-
delivered message — the user receives it twice.

Replace error=str(exc) with error=str(exc) or type(exc).__name__ at every
httpx-based adapter boundary so the existing matcher ("readtimeout",
"writetimeout") still fires.  ConnectTimeout intentionally stays
unmatched: if the connection never opened the message was not delivered,
so retry/fallback remains correct.

Applies to BlueBubbles (send + _create_chat_for_handle), WhatsApp Cloud
(text + interactive + media), QQ Bot (send chunk + keyboard + media), and
Yuanbao media handler — the same latent bug exists in every adapter that
stores error=str(exc) from an httpx call.
2026-08-27 17:21:26 +05:30
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
Teknium 82b32f32ef feat(terminal): wire shared-container key into profile-scoped resolver and MEDIA delivery
Follow-up on @fangliquanflq's opt-in (#84775): after the profile-scoping fix
(#94560) the container cache key is resolved in _resolve_container_task_id,
so the shared key must unify profiles there too — 'shared:<key>' for every
session of every opted-in profile AND for CLI/no-session runs. Delivery adds
the shared sandbox layout as the first translation candidate. Empty key
keeps strict per-profile isolation; SSH ignores the key entirely.
2026-08-25 04:00:27 -07:00
Teknium 15f7b7293c fix(terminal): persistent Docker containers are profile-scoped, not per-session
Commit a270c4ade's session-key fallback in _resolve_container_task_id was
added to stop cross-profile SSH environment reuse, but it wasn't backend-
gated: persistent Docker silently fragmented into one container per gateway
session, breaking the product contract (one long-lived container per profile,
shared by CLI and every session of that profile). #93950's vanishing MEDIA
attachments were downstream damage.

- persistent Docker (container_persistent: true) now keys to the profile:
  literal 'default' for the default profile (same container as CLI),
  'profile:<name>' for named profiles
- SSH and non-persistent Docker keep session scoping (the original leak fix
  and the #82731 isolation contract are untouched)
- gateway MEDIA translation follows the profile layout and keeps the legacy
  bug-window per-session sandboxes as fallback candidates, trying each until
  the file resolves — old sessions self-heal, no migration
- /root/.hermes credential-surface refusal preserved across all layouts
2026-08-25 02:30:38 -07:00
Michael Nguyen a15533b646 feat(gateway): warn when a Docker sandbox MEDIA path fails translation
De-silence the #93950 failure mode: when a container-absolute MEDIA path
under /workspace or /root cannot be resolved to a host sandbox file while
TERMINAL_ENV=docker, log the reason (no mounts / no prefix match / host
file missing) plus the delivering session key instead of only the generic
'Skipping unsafe MEDIA directive path' line upstream.
2026-08-24 23:50:23 -07:00
Michael Nguyen d4f31a8f36 fix(gateway): resolve session-scoped Docker sandboxes for MEDIA delivery (#93950)
Persistent Docker containers bind <sandboxes>/docker/<task>/{workspace,home}
where <task> is sanitize_task_id_for_path("session:<session_key>") — but the
gateway's synthetic mounts hardcoded the literal "default" sandbox
(_default_docker_workspace_host_root / _docker_persistent_home_host_root).
For any session-scoped deployment the longest-prefix match missed, the
container path fell through to a host-filesystem resolve that could not
exist, and every MEDIA attachment was silently dropped.

The post-handler delivery pipeline also runs after
_handle_message_with_agent cleared the turn's session contextvars, so even
a correct sandbox derivation consulting ambient state would collapse onto
"default". Thread the delivering session's key explicitly through
validate_media_delivery_path -> _translate_docker_container_media_path ->
the two host-root helpers (same pattern as the TTS fix for #57049/#36685).

Default-sandbox resolution and the /root/.hermes credential exclusion are
preserved; contexts without a key keep the historical behavior.
2026-08-24 23:50:23 -07:00
lkz-de cbc8d1804d fix(signal): chunk long cron deliveries instead of truncating 2026-08-24 20:03:24 -07:00
aniruddhaadak80 7befc1d2dd fix(gateway): route platform authorization reads through the profile secret scope
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):

- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
  WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
  GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
  GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
  hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
  _sig_secret helper; empty scoped group list previously meant "drop all
  groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
  WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
  while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
  startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
  though its sibling dm/group reads already used the scoped _getenv

All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.

Fixes #93522
2026-08-24 03:20:06 -07:00
chelsealong d7e4204e77 fix(gateway): scope multiplex-profile authorization reads (weixin/yuanbao/wecom)
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.

Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.

Fixes #93522.
2026-08-24 03:20:06 -07:00
joaomarcos f5a9ba9ee6 perf(bluebubbles): move attachment reads off the event loop 2026-08-23 20:00:07 -07:00
loulanyue 4e8419dadb fix(gateway): honor approval scope capabilities 2026-08-23 17:45:47 -05:00
kshitij 4865194772 fix(bot-mode): review follow-ups for recoverable-archive resurrection
- Clear the accidental end stamp on resurrection (at the lineage tip):
  a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
  archive auto-resurrect on the next lookup — the user could never retire
  the canonical chat. Test pins the resurrect -> deliberate-archive ->
  stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
  compressed lineage carries end_reason='compression', so tip-stamped
  accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
  dm resolution) filtered archived rows out via list_sessions_rich and
  still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
  hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
  into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
  by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
2026-08-24 03:27:11 +05:30
Teknium 6cb1085d3d fix(gateway): /p/<profile>/ on a non-multiplex gateway fails closed instead of serving the owner profile
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.

Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.

Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.

Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.

Fixes #91583 (defect 2). Repro and live validation by @kubaboski.
2026-08-23 03:57:47 -07:00
John Paul Soliva 265bdcac82 fix(api-server): a /p/<profile> prefix on a non-multiplexed gateway fails closed instead of misdelivering
The prefix is an address: the caller is naming WHICH agent the request is
for. With gateway.multiplex_profiles off, _resolve_request_profile ignored
the prefix entirely — "don't 404 a would-be valid route" — so a request
explicitly addressed to one agent was silently answered by a different one.
Observed live (Aug 2026): `hermes peer dm mini/researcher` was answered by
the mini's DEFAULT agent with no error on either side, because that host
runs one LaunchDaemon per profile and only the default daemon hosted an
api_server. A wrong-agent answer is strictly worse than an error: the
sender believes the addressee got the message.

With multiplexing off the process serves exactly one profile, so the prefix
is honored when it names that profile (peers address single-profile daemons
this way without knowing the host's topology — get_active_profile_name() is
the same identity the file already uses for model resolution) and rejected
otherwise through the existing _PROFILE_REJECTED path (404). A process that
cannot resolve its own identity rejects too: if it cannot prove who it is,
it must not answer as anyone.

Unprefixed requests are untouched, and multiplexed hosts are untouched —
the change is confined to the prefix-present, multiplexing-off branch that
previously discarded the caller's addressing.
2026-08-23 03:57:47 -07:00
Teknium 231e613d3d fix(peer): resolve hidden canonical Bot Chats in hermes peer dm
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.

- api_server: GET /api/sessions now accepts an exact-title lookup
  (?title=...) and honors include_hidden=1 ONLY alongside a title filter,
  so canonical hidden rows resolve without exposing a blanket hidden
  listing on the client surface. The title needle is pushed into SQL
  (search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
  peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
  naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
  older-peer visible fallback) + real-gateway E2E over a real state.db.

Root-cause analysis and regression recipe by @kubaboski in #91583.

Fixes #91583
2026-08-23 03:56:28 -07:00
kshitijk4poor 0b8a848754 perf(api): classify compaction rows once per message in run.completed transcript
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
2026-08-21 22:54:13 +05:30
kshitijk4poor 23a64a97ec fix(api): correct _handle_browser_control_frame return annotation
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
2026-08-21 22:33:45 +05:30
kshitijk4poor 847289864d fix(browser): make the artifact boundary compose end-to-end and scope stores per profile
Addresses both merge blockers from @andrexibiza's review of #85351:

1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
   artifact_scope_key hashed (principal, session, family), the HTTP routes
   store with an EMPTY session (API-key auth has no server session) while
   broker validation carries a session-bearing ControllerScope — every
   real upload->dispatch journey died with ArtifactScopeMismatch
   (reproduced before fixing). Canonical ownership is now
   principal/transport-family (documented in the scope-key docstring);
   ids stay unguessable server-minted 32-hex and downloads one-shot.
   New composition regression: HTTP-shape upload -> registered controller
   scope -> broker artifact dispatch, mutation-checked (re-adding session
   to the key makes it fail).

2. The 'profile-scoped' artifact store was first-profile-wins process
   state: one adapter-level singleton pinned profile B to profile A's
   physical root on multiplex listeners (same frozen-handle class as
   #88734). Stores are now cached by resolved profile, and the broker
   selects the store from the controller scope's profile_id (default-slot
   fallback preserves single-profile/test behaviour). New A/B multiplex
   regression proves distinct physical roots regardless of touch order.

Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
2026-08-21 22:33:45 +05:30
kshitijk4poor 13f209d4fd refactor(browser): dedupe auth-flow names and sentinel identity
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
  exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
  and BOTH are caught in the same WS auth flow this feature touches — two
  unrelated same-named exception types in one blast radius invited a wrong
  except clause.
- Import the 'server-internal' sentinel identity from its canonical
  definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
  re-declaring the strings; drift would have silently broken the
  internal-peer exclusion in _is_authenticated_identity.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
kshitijk4poor 45078eb941 fix(browser): offload broker lock acquisition off the event loop
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
abundantbeing 1977c3d2eb feat(browser): add scoped artifact endpoints, broker permission gates, and companion journal 2026-08-21 22:33:45 +05:30