_spawn_bridge, _env_enablement and register()'s platform_hint read
RAFT_PROFILE via raw os.environ. Under a multiplexed secondary profile
os.environ holds the DEFAULT profile's bridged value, so the bridge
subprocess / CLI hint pointed at another profile's Raft identity.
Resolve through get_secret() only when running inside a secondary
profile's scope (Buzz/SimpleX _profile_scoped() pattern, #98738); the
default profile keeps its unscoped os.environ read. No fallthrough to
os.environ after a scoped miss.
Salvaged from #100392 (tests trimmed to two).
Two sites dropped profile ownership on the tile-transcript refresh path:
- tui_gateway/server.py _sessions_sig statted only the launch home's
state.db, so a turn landing in a served sibling profile's store never
produced sessions.changed. The watcher now also probes every profile
home _profile_home() has resolved for this backend (empty set on
single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
transcript unscoped; it now passes the tile's ownerRoute scope, the
same way reconcileActiveTranscript already does for the main pane, and
keys the change signature by owner.
Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).
Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.
Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.
Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
`ProfileRoute.matches()` compared `chat_id` as an exact string, so a
WhatsApp route written as a phone number never matched the JID
(`…@s.whatsapp.net`) or LID (`…@lid`) the bridge actually delivers, and
the inbound fell through to the default profile. Allowlists and session
keys already canonicalize these via `gateway.whatsapp_identity`.
Exact compare still wins first; only whatsapp / whatsapp_cloud *user*
chats get the alias intersection fallback (applied to both `chat_id` and
`parent_chat_id`). Groups, broadcasts and every other platform stay exact.
Salvaged from #85081 (Kong); parent_chat_id fallback added on top.
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.
Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.
Salvaged from #70831 (Hudson). First reported in #68332.
Route every GOOGLE_CHAT_* / GOOGLE_APPLICATION_CREDENTIALS read through a
module-local `_get_scoped_secret` (scope-authoritative under multiplex,
os.environ fallback only for the unscoped default-profile constructor, so
startup/reconnect never hits UnscopedSecretError — #70652 class). Snapshot
Pub/Sub callback knobs on the instance while the scope is still installed,
and seed them into `extra` from `_env_enablement`.
When a scoped profile has no service-account setting, do NOT fall through
to google.auth.default(): ADC reads the process env directly and would
authenticate the profile as another profile's SA. Fail closed with an
explicit error (adapter and standalone send).
Also resolve the bot-id cache path at call time via get_hermes_home() so
profiles don't share one identity cache.
Fixes#73439.
Salvaged from #73445 (Jony) with the ADC guard from #57674 (Ray, first submitter).
Co-authored-by: Ray <rayjun0412@gmail.com>
lark_oapi.ws.client keeps the asyncio loop used by Client.start() in a
module-level global, and Hermes monkey-patched websockets.connect on the
shared module. Under multiplex every profile runs its own WS client on a
dedicated thread, so N threads overwrote each other's globals
(last-write-wins): clients scheduled tasks on a sibling's loop ('Future
attached to a different loop') or bound to the wrong loop and went deaf.
Install process-wide shims once: the module loop becomes a proxy that
forwards to the calling thread's registered loop, and websockets.connect
a dispatcher merging the calling thread's ping overrides. Add a per-adapter
supervisor: the executor future was only awaited by disconnect(), so a
dead WS thread left the profile silently deaf; now it is rebuilt with
capped exponential backoff while the adapter is meant to be connected.
Salvaged from PR #84165, trimmed: the legacy fallback path when the shim
cannot install was dropped (the shim only touches two attributes the
adapter already depended on); tests reduced to three. The loop-isolation
approach was first proposed by @zmlgit in #64247/#69904.
Co-authored-by: zmlgit <6995990+zmlgit@users.noreply.github.com>
Route FEISHU_ALLOW_BOTS, FEISHU_GROUP_POLICY, FEISHU_ALLOWED_USERS,
FEISHU_BOT_*, FEISHU_APP_ID and FEISHU_REQUIRE_MENTION through
_get_scoped_secret, and snapshot FEISHU_ALLOW_ALL_USERS /
GATEWAY_ALLOW_ALL_USERS into settings.allow_all_dm at construct time so
_admit (running on the lark_oapi WS thread, no secret scope) no longer
reads the default profile's os.environ. Feishu open_ids are app-scoped,
so a secondary bot's DMs were always dm_policy_rejected against the
default profile's allow-list.
Salvaged from PR #86908 trimmed to the adapter call-site migration; the
fail-closed _get_scoped_secret / _auth_env rewrites are already on main
(2912c36aa4, agent/secret_scope.py get_secret policy).
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.
Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.
Salvage of #69118 rebased onto the shared helper (which post-dates it).
Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.
Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.
Fixes#89161
Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
Slash dispatch already runs inside _profile_runtime_scope under multiplex,
but _save_gateway_config_key (/reasoning --global, /fast, show/hide),
/memory approval, /skills approval, /verbose and /footer built their write
path from the module constant gateway.run._hermes_home — the launch home —
so a routed profile's toggles landed in the default profile's config.yaml
while the reads (via _gateway_config_home()) saw the routed one.
Resolve the write path through _gateway_config_home() at all five sites so
reads and writes agree. Single-profile gateways never install the override
and keep resolving the launch home.
Fixes#87939Fixes#75684
Co-authored-by: Bao <nnqbao@gmail.com>
The multiplexed inbound handler wraps every message in _profile_runtime_scope,
which installs the routed profile's HERMES_HOME override and its secret scope
as contextvars. A bare loop.run_in_executor(None, fn) starts the worker with an
EMPTY context, so neither reaches the blocking work.
GatewaySlashCommandsMixin already knows this -- /compress goes through
_run_in_executor_with_context and the call site says why. Three siblings in the
same file still used the bare hop:
/insights SessionDB() with no explicit path resolves get_hermes_home() at
call time (_default_db_path), so the worker opened the DEFAULT
profile's state.db. Under multiplexing the command reported
another profile's conversations, session counts and sources to
this profile's user.
/debug collects that home's logs/config and uploads them to a public
paste, so it published the default profile's diagnostics from
another profile's chat.
/goal draft calls the auxiliary LLM, whose provider/credential resolution
reads the profile secret scope -- unscoped it falls back to
process-global os.environ, which under multiplexing may hold a
different profile's keys.
Route all three through _run_in_executor_with_context.
/reload-skills is deliberately left alone: tools.skills_tool binds SKILLS_DIR
at import time, so it does not follow the contextvar either way. Fixing that
needs the module-global retarget web_server._profile_scope performs under a
lock, which is a different change from context propagation.
Single-profile gateways never enter the scope, so their behaviour is unchanged.
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.
- cron/scheduler.py: factor the preflight's primary-config route loader into
`_primary_profile_routes_for_current_home()` (one owner for both halves,
so route semantics cannot drift) and add `SharedRouteAdapters`, a
read-only view over the primary adapter map that resolves an adapter for
a (platform, target) ONLY when an enabled primary route with a
chat_id/thread_id maps that exact target to the current profile —
using the same `ProfileRoute.matches` predicate as inbound routing.
`_deliver_result` resolves the transport per target from it; everything
else (unmatched chat, disabled route, route for another profile, no
primary adapter, guild-only route) is a miss and never uses the primary
bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
fallback: with no matching route the view is falsy and delivers nothing.
Fixes#101113
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:
1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
the caller already has a running loop — the desktop dashboard shape) ran
the standalone sender on a fresh thread with NO profile ContextVars: the
home override and secret scope were gone, so the sender resolved the
process default's home/token (or, fail-closed under multiplex, raised
UnscopedSecretError). Wrap the submit in copy_context().run like the
session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.
2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
whose own gateway (with live adapters) is running; winning the tick-lock
race meant the adapter-less desktop ticker delivered standalone. The
multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
the desktop wires it to `_check_gateway_running(home)` so such profiles
are neither ticked nor heartbeated by the dashboard while their gateway
is alive (re-evaluated every cycle, no restart needed).
Fixes#100489
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).
Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.
Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:
- startup recovery loop: a per-profile exception (e.g. an unreadable
executions.db raising sqlite3.DatabaseError) was uncaught and killed the
ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
exception escaped to the cycle-wide handler, skipping every remaining
profile that cycle and marking all of them failed.
Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).
Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).
Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.
Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.
Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.
Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.
Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).
Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.
Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.
Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.
Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.
bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.
On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.
Closes#100523
Supersedes #100544, #100542
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.
`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.
The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.
Supersedes #94147 (python half)
Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:
- `_request_owns_run` no longer admits run state that exists without an
owner stamp. Under gateway.multiplex_profiles every served profile holds
a valid key, so the "backward compatibility" branch made the boundary
allow-all whenever provenance was missing. Unstamped state now fails
closed; only an in-memory owner match or a durable idempotency record
under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
inside the request's profile scope, so its run is confined to the
creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
(`_release_run_owner_if_forgotten`) and runs at every retirement point
(task finally, SSE stream close, both sweep loops, chat-stream finally),
not only the terminal-status sweep — no stranded entries, no stateful id
ever left unowned.
Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).
Fixes#93689Fixes#90415
Supersedes #93747, #93704, #92822
Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
Salvaged from #100677. Fixes#100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.
- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
it off-thread every TTL window, so every surface on the machine reads
a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.
- Prefer the wired `set_authorization_check` callback (profile-bound
`_make_adapter_auth_check`) at both sites; keep `__self__` introspection
only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
(scoped miss under multiplex → "", never os.environ); drop the raw
`os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.
Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.
Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.
`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.
Extracted from #87240 (topic-table half handled separately via #76487).
Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.
- `_make_adapter_auth_check`: for the shared primary adapter under
multiplex, mirror the inbound message path exactly — stamp the
`profile_routes` match so the routed profile's pairing store is
consulted, and authorize under the TRANSPORT home via
`_is_user_authorized_for_source` (same split as
`_make_default_profile_message_handler`, 2afed50863). A rejected route
fails closed like the ingress gate. Retain the receiving adapter as
`_transport_adapter_ref` so config.yaml policy reads stay on it.
Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
`thread_id` as keywords only when set, so legacy 3-positional callbacks
keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
introspection class — fall back to the injected `gateway_runner` and
the adapter's owner profile.
Fixes#86296Fixes#92840
Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).
Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).
Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.
- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
_build_slash_event and _dispatch_thread_session built their SessionSource
without guild_id/parent_chat_id, while on_message passes both. Route
matching in build_source keys off exactly those fields, so under
gateway.multiplex_profiles a guild- or channel-routed profile never matched
a native slash command: /new, /reset, /model, /profile, /status ... all ran
against the default profile and reset the wrong session (#69178, #91633).
Pass guild_id (interaction.guild_id, falling back to channel.guild like the
message path) and the thread's parent channel id into build_source at both
sites. One test pins channel + thread routing parity with messages.
Fixes#69178Fixes#91633
Co-authored-by: Sora-bluesky <179361977+Sora-bluesky@users.noreply.github.com>
Co-authored-by: jondgilbert <42873618+jondgilbert@users.noreply.github.com>
Co-authored-by: tensorbit89-netizen <257030052+tensorbit89-netizen@users.noreply.github.com>
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.
Fixes#100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
_process_handoff caught a config load failure for a secondary profile,
logged a warning, and kept going with self.config, which is the primary
profile's config. The handoff then went out through the right bot to the
primary's home channel and the row was reported completed. That is the
exact wrong delivery the multi-profile handoff work exists to prevent,
and the same fail closed posture the no-live-adapters branch already
takes.
A load failure now logs an error and raises, which marks the row failed
so the CLI can report and retry it. The default profile path is
untouched, it never reloaded config.
Follow-ups on the salvaged #101081 guard:
- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
guard treated that as a lost generation and permanently halted the
handle, so the #94736 late-write self-heal reopen dropped transcript
tails (4 existing tests failed). close() now clears the recorded
sidecar generation, and _wal_generation_was_lost() re-adopts the
current sidecars after a clean /proc/self probe instead of relying on
a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
generation is recorded, the stat-based inode check alone detects an
unlink/replace. The fd probe only runs in the empty-identity state
(fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
gateway retry queue and run_agent flush divert transcripts to the
JSONL fallback exactly as they do for a replaced store, instead of
retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
open, halving the system-wide /proc scan; dropped the dead
include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
non-linux test now patches sys.platform (the real gate) instead of
_IS_WINDOWS.
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>