The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.
buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.
One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.
delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.
Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
The #82871 symptom — gateway default-denies every Buzz user because no
central env allowlist exists and the adapter's config.extra.allowed_users
was never consulted — is fixed by the plugin-platform extra.allowed_users
fallback salvaged from PR #98748 (registry-gated, normalize_user_id-aware).
That PR's tests cover the multiplex profile path; these pin the plain
single-profile gateway path from the issue's repro: npub-only and hex-only
config lists authorize, unlisted senders and empty lists stay default-deny.
Sabotage-verified: all four fail on origin/main's authz_mixin.
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
Regression for the cron/standalone path: credentials-file auth_tag must be
passed into _exec_buzz, and a bare BUZZ_PRIVATE_KEY must not invent a tag
from ambient credential files.
Reapplied from PR #83155 head (original commit carried a placeholder
local identity).
Co-authored-by: Alex P. Günsberg <alex@gunsberg.fi>
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
Summary:
The gateway's central allowlist check compared the inbound Buzz sender's
64-char hex pubkey against the raw BUZZ_ALLOWED_USERS entries. An operator
who listed only their npub saw every message rejected with
"Unauthorized user: <hex pubkey>" (gateway drops the message). npub
entries are now decoded to hex before the comparison, so npub and hex
forms of the same identity are equivalent.
Root Cause:
The Buzz adapter's own intake check already normalizes npub→hex via
_normalize_user_ref when building _allowed_pubkeys, but the gateway
applies BUZZ_ALLOWED_USERS centrally as well (authz_mixin._is_user_authorized
via the platform registry's allowed_users_env). That central path did a
raw string comparison of the allowlist entries against the hex user_id,
so an npub-only entry never matched.
Change:
- gateway/authz_mixin.py: add a pure-stdlib bech32 npub→hex decoder
(mirroring plugins/platforms/buzz/adapter.py) and normalize the buzz
allowlist set in _is_user_authorized: each npub1… entry is decoded and
its hex form added; hex entries pass through unchanged, so existing
hex-only allowlists keep working. Comparison stays fail-closed for
unrelated senders.
- tests/gateway/test_buzz_authz.py: new tests covering npub-only,
hex-only, mixed, uppercase-npub, and denial of unrelated users, plus
unit tests for the decoder/helper.
Verification:
- 43 passed (test_buzz_authz.py + test_buzz_adapter.py +
test_pairing_allowlist_bypass.py)
- 9 passed (test_multiplex_profile_authz.py + test_buzz_websocket.py)
Closes#78428
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).
The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.
Fixes#98738
Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
The four new #98738 authorization tests failed on CI shards that had
never looked Buzz up: plugin platforms have no static Platform member —
Platform._missing_ creates one on demand and caches it in _member_map_,
so attribute access (Platform.BUZZ) only works after an earlier value
lookup in the same process. Local full-suite runs happened to register
it first, which is why this only surfaced on a fresh shard.
Resolve the member once via Platform("buzz") at module import and use
that constant in the runner/source helpers.
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).
- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
profile scope extra is authoritative and env is not consulted (a missing
key fails closed to its default instead of borrowing the default
profile's value); single-profile and unscoped/default-profile reads keep
the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
home channel, poll interval, require_mention, transport, allowed users),
_resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
_standalone_send, and check_requirements (which now consults the
profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
_apply_yaml_config() skips the env bridge there, so the default profile's
env cannot fabricate Buzz for a profile that never configured it and a
secondary profile's YAML cannot be pinned into the process env
(first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
config.extra.allowed_users (gated on the registry entry declaring
allowed_users_env, with an optional normalize_user_id hook so Buzz npub
entries match hex-pubkey user ids) — under multiplex only the default
profile's list ever reached the env var, so listed secondary-profile
users were default-denied (#82871). Empty/absent lists change nothing;
default-deny is preserved.
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.
Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).
Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
The Discord adapter resolves username allowlist entries to numeric IDs at
connect and mirrors them into os.environ — but the gateway's per-turn .env
hot-reload (load_hermes_dotenv(override=True)) restores the raw usernames
from the file. From the second agent turn onward, _is_user_authorized
compared numeric user_ids against username strings and dropped every
message from the operator as 'Unauthorized user' while the adapter layer
still admitted them (bot reacted, never replied).
Fix: gateway authz unions the adapter's resolved numeric IDs
(DiscordAdapter.resolved_allowlist_user_ids()) into the env-derived
allowlist. Union only fires when an env allowlist is configured (never a
widening; fail-closed branch unchanged), is duck-typed + isinstance-guarded
against mock adapters, and filters non-numeric entries so unresolved
usernames and '*' can't leak through adapter memory.
Live repro: symptom fired on origin/main (authorized=False after reload),
passes with fix; stranger + empty-allowlist + raising-resolver negatives
hold. Sabotage run: incident test fails on unfixed authz_mixin.
Since the #98790 heartbeat guard, a never-ticked gateway prints the YELLOW
first-heartbeat notice (which also contains 'NOT fire'). The lock-first
test's real contract (#87033) is that an active runtime lock suppresses the
RED 'Gateway is not running' false alarm — assert that directly.
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
PID: 4165
Ticker heartbeat: 39s ago
4 active job(s)
Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.
Root causes:
1. executed ,
which enumerated the entire systemd fleet regardless of ,
violating the docstring "only PIDs belonging to the current profile".
2. checked when
(no heartbeat file) should trigger a warning — instead, it fell through
to the "✓ Gateway is running" green branch.
Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
explicit yellow warning: "ticker has not reported a heartbeat".
Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).
When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).
Addresses review feedback on #73363. The previous truthy
`profile_adapters.get(name)` check fell back to the shared (default-profile)
adapters whenever a secondary profile's adapter map was empty — which is the
normal state before that profile's bot connects (the map is created empty and
filled only on a successful connection). That reintroduced wrong-bot
cross-delivery for the secondary until it connected.
Thread the default profile identity and reserve the shared `adapters` set for
it alone; every other profile uses its own adapter map, or an empty set when
its bot has not connected yet (so it simply does not deliver that tick).
Add regression coverage for the default, connected-secondary, empty-secondary
and missing-secondary adapter-map cases.
The tui_gateway/run_agent writers now stamp profile_name explicitly, but
every OTHER creation path that passes no profile_name (cli.py /new,
hermes_cli/main.py --create-if-missing, foreign-session import, ACP
adapter, gateway branch/title paths, the #82616 peer self-heal INSERT,
and compression children of legacy NULL parents) still minted
profile_name = NULL rows. Rows minted NULL after the one-shot #94724
legacy-owner backfill ran stayed NULL forever: profile-keyed consumers
(desktop sidebar scope matching, @session:<profile>/<id> deep links, the
fail-closed owner ladder) treat NULL as unowned, so the sessions vanished
from the sidebar with their transcripts intact (#99222).
Fix the class at the choke point instead of chasing call sites: every
profile-tree state.db belongs to exactly one profile, so
SessionDB._insert_session_row (and the peer self-heal INSERT and the
compression-child publisher) derive the store's own profile from db_path
when the caller names none — <root>/state.db -> 'default',
<root>/profiles/<name>/state.db -> <name>. The same single-match contract
backfill_null_session_profiles and the web listing's row_profile stamp
already rely on. Explicit profile_name arguments always win; stores
outside the profile tree (tests, ad-hoc copies) keep NULL — never guess.
E2E-verified with real imports against a temp HERMES_HOME:
before (origin/main) a bare create_session on the default store persisted
NULL; after, 'default' / '<profile>' land in state.db for the default
store, a named-profile store, the peer self-heal insert, and a
compression child of a NULL parent, while explicit args and
outside-tree stores are unchanged.
Refs #99222
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).
NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.
Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.
Fixes#99222
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.
Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
config.get and config.set ignored the focused profile on a shared
app-global backend, so reads and persistent writes used the launch
config.yaml. Bind the existing @_profile_scoped decorator and write
_save_cfg through the request home override.
Fixes#95760
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
_env_line_defines_key() decides which .env lines the writers may rewrite or
drop. It matched on the `KEY=` prefix, but load_env() splits on the first
`=` and strips the name:
key, _, value = line.partition('=')
env_vars[key.strip()] = _parse_env_value(value)
so `OPENAI_API_KEY = sk-...` is a live assignment — the key resolves, the
provider works, and every UI shows it as set. The writers did not see it.
This is the same resurrection hole #40041 fixed for `export KEY=`, still
open for the whitespace form:
- DELETE /api/env 404s ("not found in .env") while the credential stays
active — a key the user revoked through the UI is never actually revoked
- PUT /api/env appends a SECOND line instead of replacing; a later delete
removes the appended line and the original value silently comes back
Rotate-then-delete on a spaced line therefore restores exactly the key the
user rotated away from.
Match load_env()'s parse instead of prefix-matching, so the writers accept
precisely what the reader accepts: skip blank/comment/no-'=' lines, strip an
`export ` prefix, then compare the stripped name. Commented-out lines stay
untouched and `KEY_EXTRA=`/`MY_KEY=` still do not match `KEY`.
Verified against the real dashboard endpoints on a temp HERMES_HOME: the
spaced line is now removed, rotation replaces it in place with no duplicate,
and a parity check asserts the writer matches a line iff load_env() does.
`_scrub_config_yaml_mirrors` reconciles the config.yaml copies of a credential
when it is rotated or removed through the dashboard. It walks `model`,
`auxiliary.<task>`, and `custom_providers.<name>` — but not the keyed
`providers` schema.
`providers` is not a niche section: `get_compatible_custom_providers` documents
it as "the newer keyed schema" (v12+), and it is exactly where the dashboard /
desktop write a custom endpoint's inline key —
`_write_custom_endpoint` sets `providers.<id>.api_key`. That value is a real
credential: the runtime resolver reads it (`runtime_provider` /
`hermes_cli.main` / `model_switch` all read `entry.get("api_key")` off a
`providers` entry), and an inline key outranks the env var.
So the section the scrub skips is the one the dashboard writes to, and both
callers break on it:
- save_provider_env_credential (rotation, #62269): a stale
`providers.<id>.api_key` is left at the OLD value and, being
higher-precedence than the freshly-rotated env var, shadows the rotation —
the "persistent 401 with a key the UI no longer shows" that #62269 fixed,
reintroduced for the newer schema.
- remove_provider_env_credential: its contract is to "remove a credential from
EVERY store it lives in", yet the `providers` copy survives, leaving the
secret in config.yaml after the user asked to delete it.
Walk `providers.<id>` too. The scrub stays value-matched, so an unrelated
endpoint's key is untouched. Only `api_key` is scrubbed here: in the keyed
`providers` schema `api` is the base_url alias, not a credential (unlike
model/auxiliary/custom_providers), so `_fix` takes an explicit field list and
this section passes `("api_key",)` — a provider's endpoint URL is never
rewritten even if it happened to equal the credential string.
tests/hermes_cli/test_credential_lifecycle.py: drive the real PUT/DELETE
/api/env endpoints against a `providers.<id>.api_key` mirror — rotation moves
it to the new key, delete clears it, and a `providers.<id>.api` base_url alias
is preserved. The two scrub tests fail on main (stale key survives); the
base_url guard passes on main as a control. 15 pass here; 345 pass across the
credential-lifecycle + web-server suites (the one failing honcho-merge test
fails identically on clean main).
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
comment pinning WHY it is the timeout ladder's first rung (transient
lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
test_compression_rotation_state.py / test_context_compressor.py; kept the
force-bypass test (only site pinning that behavior for split failures) and
the E2E test (now asserting the named constant).
The salvaged unit tests drive _record_compression_failure_cooldown directly;
this drives the real _compress_context split-failure path (archive boom on a
real SessionDB) and asserts the cooldown recording fires with the
session_split_failed error class.
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):
1. publish_compression_child gains require_lease_refresh: the lease is
extended inside the same transaction as the expiry check (same conn, no
TOCTOU), giving a worker whose refresher thread died from transient DB
failures one final chance to keep its completed work.
2. A failed compression split now records a 60s failure cooldown, so the
next turn cannot immediately re-trigger the same compression.
Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
Follow-ups on the salvaged #97797 transport:
- Blocker 2 from the #97681 exact-head review claimed near-expiry refresh
silently mints against the target's CURRENT policy. On this head the
handler DOES refuse drift (_require_unchanged_execution_policy -> 403
room_reauthorization_required), but nothing pinned the handler-level
behavior: removing the drift check still passed the entire grants suite
(the check was only unit-tested in isolation). New HTTP-level regression
test drives /v1/room-members/grants/refresh with a drifted-policy grant
and requires the 403; sabotage-verified (check removed -> test fails).
- cancel() conflict resolution: keeps our race-retry routing loop from
#99099 with this layer's peer-stop acknowledgement body inside it
(peer receipt -> settle completion -> local interrupt escalation).
- docs: NAT one-way-reachability note in bot-mode.md — Desktop is a viewer,
not a relay; put room authority on the host everyone can reach (field
finding from /bin/bash on #97681).
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.
Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.
Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
The stream consumer called prefers_fresh_final_streaming(text,
metadata=...) only, and no metadata producer stamps a platform key — so
RelayAdapter's hook always fell back to the PRIMARY descriptor's
platform (the scalar-vs-per-chat capability seam, third occurrence).
Two failure directions on multiplexed relays with
platforms.relay.extra.slack.unfurl_links/media: true (#97957):
- Slack primary fronting Telegram/Discord: every link-bearing streamed
final on the non-Slack chats finalized as a fresh send with no delete
op advertised -> the answer delivered TWICE (orphaned preview).
- Non-Slack primary fronting Slack: the hook returned False, leaving
the force-on unfurl feature dark on exactly the chats it shipped for.
Pass chat_id=self.chat_id from the consumer; the relay hook already
accepted it and resolves via _platform_by_chat + the per-platform
negotiated descriptor. Graduated TypeError fallback keeps the
single-platform hook signatures (Telegram, base class) and legacy test
doubles working unchanged.
Both regression tests verified RED against the unfixed consumer, GREEN
with the fix; single-platform relays are unaffected (#97957's own 30
tests unchanged-green).
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.
Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.
Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
The PKCE payload is a flat 'provider=...;state=...;verifier=...;next=...'
string. A raw ';' is a cookie-attribute terminator, so Python's
http.cookies emits the value in RFC 6265 quoted form with each ';'
escaped as the backslash-octal '\073'. Mainstream browsers echo that
form back verbatim and Python parsers decode it — the browser round
trip is fine. But '"' and '\' are outside the plain cookie-octet set,
and non-Python hops that re-serialize the Cookie header reject the
value and drop the cookie entirely: Go's net/http (Traefik middleware,
Authentik outposts, other gateways) refuses any cookie value
containing a backslash. The OIDC callback then 400s with "Missing
PKCE state cookie" even though the browser sent the cookie.
Field reproduction: support thread "Still unable to use Authentik for
signin with traefik" — devtools showed the browser sending the intact
quoted \073 cookie on /auth/callback while Hermes logged
missing_pkce_cookie behind a Traefik+Authentik chain.
Fix: URL-encode the whole payload in set_pkce_cookie (quote(payload,
safe='') — ';' becomes '%3B') so the wire value contains only
cookie-octets and no parser in the chain has anything to reject, and
decode through a single shared inverse, cookies.parse_pkce_payload(),
in BOTH readers: the OAuth /auth/callback and the native
password-login path (routes.login_submit), whose broker/provider
binding check would otherwise parse zero segments from the
newly-encoded value and silently disable itself.
Regression coverage: the wire-shape test pins the full cookie-octet
set (the '"'/'\' assertions are the ones a Go-parser hop fails
pre-fix), the round-trip tests drive the real /auth/login →
/auth/callback path, and the next= test pins the exact post-login
redirect byte shape. Native-flow broker assertions updated to decode
through parse_pkce_payload instead of substring-matching the raw wire
value.
Salvaged from #84065 (rebased onto current main, which gained the
SameSite=None PKCE attrs and the RFC 8252 native password flow since
the PR branched): kept main's _pkce_attrs cookie shape, extended the
fix to the login_submit reader the original PR predated, and reframed
the rationale — browsers do NOT truncate at the first ';' (there is
no literal ';' on the wire in the quoted form); the failing hop is a
strict middlebox cookie parser.
Closes#83832
Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
Query-file DM transports do not consume stdin. Use DEVNULL for both the initial attempt and policy-gated retry so Git Bash cannot pass an invalid pseudo-handle to Windows subprocess creation.