Commit Graph

14240 Commits

Author SHA1 Message Date
Reinhold b3af0b6863 fix(buzz): keep progress messages in thread 2026-08-31 07:30:44 -07:00
Han Ngo 34c10f83c3 feat(buzz): implement edit_message and delete_message so replies can stream
The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.

buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.

One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.

delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.

Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
2026-08-31 07:30:44 -07:00
686f6c61 e703717513 fix(cron): treat a live multiplexer as gateway-alive for satellite profiles
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
2026-08-31 07:28:39 -07:00
Teknium 5742f6987d test(buzz): pin the single-profile config-only allowlist path (#82871)
The #82871 symptom — gateway default-denies every Buzz user because no
central env allowlist exists and the adapter's config.extra.allowed_users
was never consulted — is fixed by the plugin-platform extra.allowed_users
fallback salvaged from PR #98748 (registry-gated, normalize_user_id-aware).
That PR's tests cover the multiplex profile path; these pin the plain
single-profile gateway path from the issue's repro: npub-only and hex-only
config lists authorize, unlisted senders and empty lists stay default-deny.
Sabotage-verified: all four fail on origin/main's authz_mixin.
2026-08-31 07:28:30 -07:00
Teknium 9113cf24a9 fix(buzz): reconcile scoped auth-tag resolution across salvaged fixes
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
2026-08-31 07:28:30 -07:00
Teknium 66c56d46fc test(buzz): cover standalone owner auth_tag injection
Regression for the cron/standalone path: credentials-file auth_tag must be
passed into _exec_buzz, and a bare BUZZ_PRIVATE_KEY must not invent a tag
from ambient credential files.

Reapplied from PR #83155 head (original commit carried a placeholder
local identity).

Co-authored-by: Alex P. Günsberg <alex@gunsberg.fi>
2026-08-31 07:28:30 -07:00
Alex P. Günsberg 29ee8230be fix(buzz): fail closed on multiplex credential discovery 2026-08-31 07:28:30 -07:00
Alex P. Günsberg 8c09c39530 fix(buzz): scope owner credentials per profile 2026-08-31 07:28:30 -07:00
Alex P. Günsberg 51ffeb059c fix(buzz): load owner auth tag from credentials 2026-08-31 07:28:30 -07:00
liuhao1024 a684d154bc fix(buzz): let the requirement gate see externally managed secrets
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
2026-08-31 07:28:30 -07:00
webtecnica c83121cf5f fix(buzz): normalize npub entries in BUZZ_ALLOWED_USERS to hex
Summary:
The gateway's central allowlist check compared the inbound Buzz sender's
64-char hex pubkey against the raw BUZZ_ALLOWED_USERS entries. An operator
who listed only their npub saw every message rejected with
"Unauthorized user: <hex pubkey>" (gateway drops the message). npub
entries are now decoded to hex before the comparison, so npub and hex
forms of the same identity are equivalent.

Root Cause:
The Buzz adapter's own intake check already normalizes npub→hex via
_normalize_user_ref when building _allowed_pubkeys, but the gateway
applies BUZZ_ALLOWED_USERS centrally as well (authz_mixin._is_user_authorized
via the platform registry's allowed_users_env). That central path did a
raw string comparison of the allowlist entries against the hex user_id,
so an npub-only entry never matched.

Change:
- gateway/authz_mixin.py: add a pure-stdlib bech32 npub→hex decoder
  (mirroring plugins/platforms/buzz/adapter.py) and normalize the buzz
  allowlist set in _is_user_authorized: each npub1… entry is decoded and
  its hex form added; hex entries pass through unchanged, so existing
  hex-only allowlists keep working. Comparison stays fail-closed for
  unrelated senders.
- tests/gateway/test_buzz_authz.py: new tests covering npub-only,
  hex-only, mixed, uppercase-npub, and denial of unrelated users, plus
  unit tests for the decoder/helper.

Verification:
- 43 passed (test_buzz_authz.py + test_buzz_adapter.py +
  test_pairing_allowlist_bypass.py)
- 9 passed (test_multiplex_profile_authz.py + test_buzz_websocket.py)

Closes #78428
2026-08-31 07:28:30 -07:00
Kosta Gorod a5e7355766 fix(buzz): resolve BUZZ_AUTH_TAG through the profile secret scope
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).

The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.

Fixes #98738

Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
2026-08-31 07:28:30 -07:00
liuhao1024 91d4c791be fix(tests): resolve Buzz Platform member by value, not attribute access
The four new #98738 authorization tests failed on CI shards that had
never looked Buzz up: plugin platforms have no static Platform member —
Platform._missing_ creates one on demand and caches it in _member_map_,
so attribute access (Platform.BUZZ) only works after an earlier value
lookup in the same process. Local full-suite runs happened to register
it first, which is why this only surfaced on a fresh shard.

Resolve the member once via Platform("buzz") at module import and use
that constant in the runner/source helpers.
2026-08-31 07:28:30 -07:00
liuhao1024 aaa5f27d0f fix(buzz): secondary multiplex profiles must not inherit the default profile's env
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).

- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
  profile scope extra is authoritative and env is not consulted (a missing
  key fails closed to its default instead of borrowing the default
  profile's value); single-profile and unscoped/default-profile reads keep
  the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
  home channel, poll interval, require_mention, transport, allowed users),
  _resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
  _standalone_send, and check_requirements (which now consults the
  profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
  _apply_yaml_config() skips the env bridge there, so the default profile's
  env cannot fabricate Buzz for a profile that never configured it and a
  secondary profile's YAML cannot be pinned into the process env
  (first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
  config.extra.allowed_users (gated on the registry entry declaring
  allowed_users_env, with an optional normalize_user_id hook so Buzz npub
  entries match hex-pubkey user ids) — under multiplex only the default
  profile's list ever reached the env var, so listed secondary-profile
  users were default-denied (#82871). Empty/absent lists change nothing;
  default-deny is preserved.
2026-08-31 07:28:30 -07:00
Wesley Simplicio 8edaa25746 fix(cron): isolate desktop profile persistence 2026-08-31 07:28:22 -07:00
Teknium 6fba07cb7e test(terminal): background spawn contract now includes owner_task_id 2026-08-31 07:28:18 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
Teknium 3a351a9665 feat(worktree): pushed open-PR lanes reclaim their disk; cron tick prunes worktrees
Two growth leaks closed:

1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
   trees, ~18GB on the reporting box): managed installs fetch with a
   single-branch refspec, so pushed PR branches never get refs/remotes/*
   entries and read as 'unpushed' forever. When a clean tree's branch head
   EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
   is redundant: reap the TREE, keep the BRANCH ref (shielded from the
   orphaned-branch pass). Anything diverged/unverifiable stays preserved.
   Applied to both the startup pruner and hermes worktree prune/list.

2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
   gateway-driven boxes accumulated trees for days. The scheduler tick now
   dispatches the same conservative pruner on a daemon thread, throttled
   to once per 6h, against the install checkout + job-workdir repos that
   have a .worktrees/ dir.
2026-08-31 07:27:50 -07:00
Teknium a0a63a1bc2 fix(gateway): username-based DISCORD_ALLOWED_USERS no longer locks out the operator after one turn
The Discord adapter resolves username allowlist entries to numeric IDs at
connect and mirrors them into os.environ — but the gateway's per-turn .env
hot-reload (load_hermes_dotenv(override=True)) restores the raw usernames
from the file. From the second agent turn onward, _is_user_authorized
compared numeric user_ids against username strings and dropped every
message from the operator as 'Unauthorized user' while the adapter layer
still admitted them (bot reacted, never replied).

Fix: gateway authz unions the adapter's resolved numeric IDs
(DiscordAdapter.resolved_allowlist_user_ids()) into the env-derived
allowlist. Union only fires when an env allowlist is configured (never a
widening; fail-closed branch unchanged), is duck-typed + isinstance-guarded
against mock adapters, and filters non-numeric entries so unresolved
usernames and '*' can't leak through adapter memory.

Live repro: symptom fired on origin/main (authorized=False after reload),
passes with fix; stranger + empty-allowlist + raising-resolver negatives
hold. Sabotage run: incident test fails on unfixed authz_mixin.
2026-08-31 06:28:25 -07:00
Teknium 9eb832aad7 test(cron): pin the lock-first contract on the red alarm, not the shared 'NOT fire' substring
Since the #98790 heartbeat guard, a never-ticked gateway prints the YELLOW
first-heartbeat notice (which also contains 'NOT fire'). The lock-first
test's real contract (#87033) is that an active runtime lock suppresses the
RED 'Gateway is not running' false alarm — assert that directly.
2026-08-31 06:02:32 -07:00
Teknium 8ad02486b6 test(cron): regression coverage for owning-profile home-target resolution (#94862 #97909 #99028) 2026-08-31 06:02:32 -07:00
Teknium 43ca78aa40 fix(cron): keep SessionDB kwargs-free in the context-preserving worker
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
2026-08-31 06:02:32 -07:00
kokhlo 82d7a13003 fix(cron): profile isolation — systemd filter + heartbeat guard
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
  PID: 4165
  Ticker heartbeat: 39s ago

  4 active job(s)
  Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.

Root causes:
1.  executed ,
   which enumerated the entire systemd fleet regardless of ,
   violating the docstring "only PIDs belonging to the current profile".
2.  checked  when
   (no heartbeat file) should trigger a warning — instead, it fell through
   to the "✓ Gateway is running" green branch.

Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
  when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
  explicit yellow warning: "ticker has not reported a heartbeat".

Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
2026-08-31 06:02:32 -07:00
liuhao1024 d5d0613778 fix(cron): profile_routes rescue for satellite-profile delivery preflight
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).

When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).
2026-08-31 06:02:32 -07:00
Kevin dfa5004a04 fix(cron): preserve profile context during session DB init 2026-08-31 06:02:32 -07:00
Muhammad Usama a07370abd0 fix(cron): reserve shared adapters for the default profile only
Addresses review feedback on #73363. The previous truthy
`profile_adapters.get(name)` check fell back to the shared (default-profile)
adapters whenever a secondary profile's adapter map was empty — which is the
normal state before that profile's bot connects (the map is created empty and
filled only on a successful connection). That reintroduced wrong-bot
cross-delivery for the secondary until it connected.

Thread the default profile identity and reserve the shared `adapters` set for
it alone; every other profile uses its own adapter map, or an empty set when
its bot has not connected yet (so it simply does not deliver that tick).

Add regression coverage for the default, connected-secondary, empty-secondary
and missing-secondary adapter-map cases.
2026-08-31 06:02:32 -07:00
ghosty93 07f6518b6d fix(cron): retain profile secret scope through delivery 2026-08-31 06:02:32 -07:00
Teknium 5cc3da6827 fix(state): SessionDB derives its own store's profile for unstamped session rows
The tui_gateway/run_agent writers now stamp profile_name explicitly, but
every OTHER creation path that passes no profile_name (cli.py /new,
hermes_cli/main.py --create-if-missing, foreign-session import, ACP
adapter, gateway branch/title paths, the #82616 peer self-heal INSERT,
and compression children of legacy NULL parents) still minted
profile_name = NULL rows. Rows minted NULL after the one-shot #94724
legacy-owner backfill ran stayed NULL forever: profile-keyed consumers
(desktop sidebar scope matching, @session:<profile>/<id> deep links, the
fail-closed owner ladder) treat NULL as unowned, so the sessions vanished
from the sidebar with their transcripts intact (#99222).

Fix the class at the choke point instead of chasing call sites: every
profile-tree state.db belongs to exactly one profile, so
SessionDB._insert_session_row (and the peer self-heal INSERT and the
compression-child publisher) derive the store's own profile from db_path
when the caller names none — <root>/state.db -> 'default',
<root>/profiles/<name>/state.db -> <name>. The same single-match contract
backfill_null_session_profiles and the web listing's row_profile stamp
already rely on. Explicit profile_name arguments always win; stores
outside the profile tree (tests, ad-hoc copies) keep NULL — never guess.

E2E-verified with real imports against a temp HERMES_HOME:
before (origin/main) a bare create_session on the default store persisted
NULL; after, 'default' / '<profile>' land in state.db for the default
store, a named-profile store, the peer self-heal insert, and a
compression child of a NULL parent, while explicit args and
outside-tree stores are unchanged.

Refs #99222
2026-08-31 05:56:18 -07:00
Teknium 6874b99d49 fix(sessions): stamp launch-profile name on new session rows instead of NULL
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).

NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.

Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.

Fixes #99222
2026-08-31 05:56:18 -07:00
ehz0ah 64b96bb5d2 fix(openviking): synchronize setup connection state 2026-08-31 17:20:45 +05:30
ehz0ah 823bcc887a feat(openviking): use user memory by default
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.

Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
2026-08-31 17:20:45 +05:30
teknium1 8fd144c502 fix(desktop): model assignment carries the credential pointer, not a resolved key (#88990, salvage #90484)
Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
2026-08-31 04:40:24 -07:00
StanleyStetson fa2dd28029 fix(tui_gateway): scope config.get/set RPC to params.profile
config.get and config.set ignored the focused profile on a shared
app-global backend, so reads and persistent writes used the launch
config.yaml. Bind the existing @_profile_scoped decorator and write
_save_cfg through the request home override.

Fixes #95760
2026-08-31 04:33:20 -07:00
Teknium a90be562f4 fix(web): stop mirroring env-backed provider keys into model.api_key (#88990)
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
2026-08-31 03:37:43 -07:00
Frowtek 1152d4d3ce fix(cli): recognize whitespace around '=' in .env save/remove
_env_line_defines_key() decides which .env lines the writers may rewrite or
drop. It matched on the `KEY=` prefix, but load_env() splits on the first
`=` and strips the name:

    key, _, value = line.partition('=')
    env_vars[key.strip()] = _parse_env_value(value)

so `OPENAI_API_KEY = sk-...` is a live assignment — the key resolves, the
provider works, and every UI shows it as set. The writers did not see it.

This is the same resurrection hole #40041 fixed for `export KEY=`, still
open for the whitespace form:

- DELETE /api/env 404s ("not found in .env") while the credential stays
  active — a key the user revoked through the UI is never actually revoked
- PUT /api/env appends a SECOND line instead of replacing; a later delete
  removes the appended line and the original value silently comes back

Rotate-then-delete on a spaced line therefore restores exactly the key the
user rotated away from.

Match load_env()'s parse instead of prefix-matching, so the writers accept
precisely what the reader accepts: skip blank/comment/no-'=' lines, strip an
`export ` prefix, then compare the stripped name. Commented-out lines stay
untouched and `KEY_EXTRA=`/`MY_KEY=` still do not match `KEY`.

Verified against the real dashboard endpoints on a temp HERMES_HOME: the
spaced line is now removed, rotation replaces it in place with no duplicate,
and a parity check asserts the writer matches a line iff load_env() does.
2026-08-31 03:37:43 -07:00
Drexuxux 22f9caf84e fix(credentials): scrub the keyed providers schema on rotate/remove
`_scrub_config_yaml_mirrors` reconciles the config.yaml copies of a credential
when it is rotated or removed through the dashboard. It walks `model`,
`auxiliary.<task>`, and `custom_providers.<name>` — but not the keyed
`providers` schema.

`providers` is not a niche section: `get_compatible_custom_providers` documents
it as "the newer keyed schema" (v12+), and it is exactly where the dashboard /
desktop write a custom endpoint's inline key —
`_write_custom_endpoint` sets `providers.<id>.api_key`. That value is a real
credential: the runtime resolver reads it (`runtime_provider` /
`hermes_cli.main` / `model_switch` all read `entry.get("api_key")` off a
`providers` entry), and an inline key outranks the env var.

So the section the scrub skips is the one the dashboard writes to, and both
callers break on it:

- save_provider_env_credential (rotation, #62269): a stale
  `providers.<id>.api_key` is left at the OLD value and, being
  higher-precedence than the freshly-rotated env var, shadows the rotation —
  the "persistent 401 with a key the UI no longer shows" that #62269 fixed,
  reintroduced for the newer schema.
- remove_provider_env_credential: its contract is to "remove a credential from
  EVERY store it lives in", yet the `providers` copy survives, leaving the
  secret in config.yaml after the user asked to delete it.

Walk `providers.<id>` too. The scrub stays value-matched, so an unrelated
endpoint's key is untouched. Only `api_key` is scrubbed here: in the keyed
`providers` schema `api` is the base_url alias, not a credential (unlike
model/auxiliary/custom_providers), so `_fix` takes an explicit field list and
this section passes `("api_key",)` — a provider's endpoint URL is never
rewritten even if it happened to equal the credential string.

tests/hermes_cli/test_credential_lifecycle.py: drive the real PUT/DELETE
/api/env endpoints against a `providers.<id>.api_key` mirror — rotation moves
it to the new key, delete clears it, and a `providers.<id>.api` base_url alias
is preserved. The two scrub tests fail on main (stale key survives); the
base_url guard passes on main as a control. 15 pass here; 345 pass across the
credential-lifecycle + web-server suites (the one failing honcho-merge test
fails identically on clean main).
2026-08-31 03:37:43 -07:00
Finn763 82733a3fdb fix(credential-pool): materialize pool entry on Desktop PUT /api/env save (#96058) 2026-08-31 03:37:43 -07:00
kshitijk4poor 3aee290899 refactor(compression): name the split-failure cooldown; drop duplicate tests
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
  comment pinning WHY it is the timeout ladder's first rung (transient
  lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
  condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
  test_compression_rotation_state.py / test_context_compressor.py; kept the
  force-bypass test (only site pinning that behavior for split failures) and
  the E2E test (now asserting the named constant).
2026-08-31 14:09:42 +05:30
kshitijk4poor 81ab11b821 test(compression): E2E-pin the split-failure cooldown through a real SessionDB
The salvaged unit tests drive _record_compression_failure_cooldown directly;
this drives the real _compress_context split-failure path (archive boom on a
real SessionDB) and asserts the cooldown recording fires with the
session_split_failed error class.
2026-08-31 14:09:42 +05:30
VVV 087cc49a26 fix(compression): refresh lease in-transaction before publish; arm cooldown on split failure
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):

1. publish_compression_child gains require_lease_refresh: the lease is
   extended inside the same transaction as the expiry check (same conn, no
   TOCTOU), giving a worker whose refresher thread died from transient DB
   failures one final chance to keep its completed work.

2. A failed compression split now records a 60s failure cooldown, so the
   next turn cannot immediately re-trigger the same compression.

Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
2026-08-31 14:09:42 +05:30
Teknium 1cf3639813 test(bot-mode): pin grant-refresh reauthorization on execution-policy drift
Follow-ups on the salvaged #97797 transport:

- Blocker 2 from the #97681 exact-head review claimed near-expiry refresh
  silently mints against the target's CURRENT policy. On this head the
  handler DOES refuse drift (_require_unchanged_execution_policy -> 403
  room_reauthorization_required), but nothing pinned the handler-level
  behavior: removing the drift check still passed the entire grants suite
  (the check was only unit-tested in isolation). New HTTP-level regression
  test drives /v1/room-members/grants/refresh with a drifted-policy grant
  and requires the 403; sabotage-verified (check removed -> test fails).
- cancel() conflict resolution: keeps our race-retry routing loop from
  #99099 with this layer's peer-stop acknowledgement body inside it
  (peer receipt -> settle completion -> local interrupt escalation).
- docs: NAT one-way-reachability note in bot-mode.md — Desktop is a viewer,
  not a relay; put room authority on the host everyone can reach (field
  finding from /bin/bash on #97681).
2026-08-31 01:04:11 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
itskaism 5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Ben Barclay 1f99a4b2f2 fix(relay): resolve fresh-final unfurl decision per chat, not per primary identity (#99206)
The stream consumer called prefers_fresh_final_streaming(text,
metadata=...) only, and no metadata producer stamps a platform key — so
RelayAdapter's hook always fell back to the PRIMARY descriptor's
platform (the scalar-vs-per-chat capability seam, third occurrence).
Two failure directions on multiplexed relays with
platforms.relay.extra.slack.unfurl_links/media: true (#97957):

- Slack primary fronting Telegram/Discord: every link-bearing streamed
  final on the non-Slack chats finalized as a fresh send with no delete
  op advertised -> the answer delivered TWICE (orphaned preview).
- Non-Slack primary fronting Slack: the hook returned False, leaving
  the force-on unfurl feature dark on exactly the chats it shipped for.

Pass chat_id=self.chat_id from the consumer; the relay hook already
accepted it and resolves via _platform_by_chat + the per-platform
negotiated descriptor. Graduated TypeError fallback keeps the
single-platform hook signatures (Telegram, base class) and legacy test
doubles working unchanged.

Both regression tests verified RED against the unfixed consumer, GREEN
with the fix; single-platform relays are unaffected (#97957's own 30
tests unchanged-green).
2026-08-31 16:28:21 +10:00
kshitijk4poor 6681f9ebc3 refactor(telegram): share exception-graph walk across classifiers
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.

Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
2026-08-31 11:35:24 +05:30
kshitijk4poor 777a501d09 test: settle generation verifier in shielded-close drain test 2026-08-31 11:34:16 +05:30
kshitijk4poor 8288129475 fix(telegram): bound polling drain with wall-clock deadline
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.

Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
2026-08-31 11:34:16 +05:30
Ben Barclay a65a517d04 fix(dashboard-auth): url-encode the PKCE cookie value so strict proxy hops stop dropping it (#99176)
The PKCE payload is a flat 'provider=...;state=...;verifier=...;next=...'
string. A raw ';' is a cookie-attribute terminator, so Python's
http.cookies emits the value in RFC 6265 quoted form with each ';'
escaped as the backslash-octal '\073'. Mainstream browsers echo that
form back verbatim and Python parsers decode it — the browser round
trip is fine. But '"' and '\' are outside the plain cookie-octet set,
and non-Python hops that re-serialize the Cookie header reject the
value and drop the cookie entirely: Go's net/http (Traefik middleware,
Authentik outposts, other gateways) refuses any cookie value
containing a backslash. The OIDC callback then 400s with "Missing
PKCE state cookie" even though the browser sent the cookie.

Field reproduction: support thread "Still unable to use Authentik for
signin with traefik" — devtools showed the browser sending the intact
quoted \073 cookie on /auth/callback while Hermes logged
missing_pkce_cookie behind a Traefik+Authentik chain.

Fix: URL-encode the whole payload in set_pkce_cookie (quote(payload,
safe='') — ';' becomes '%3B') so the wire value contains only
cookie-octets and no parser in the chain has anything to reject, and
decode through a single shared inverse, cookies.parse_pkce_payload(),
in BOTH readers: the OAuth /auth/callback and the native
password-login path (routes.login_submit), whose broker/provider
binding check would otherwise parse zero segments from the
newly-encoded value and silently disable itself.

Regression coverage: the wire-shape test pins the full cookie-octet
set (the '"'/'\' assertions are the ones a Go-parser hop fails
pre-fix), the round-trip tests drive the real /auth/login →
/auth/callback path, and the next= test pins the exact post-login
redirect byte shape. Native-flow broker assertions updated to decode
through parse_pkce_payload instead of substring-matching the raw wire
value.

Salvaged from #84065 (rebased onto current main, which gained the
SameSite=None PKCE attrs and the RFC 8252 native password flow since
the PR branched): kept main's _pkce_attrs cookie shape, extended the
fix to the login_submit reader the original PR predated, and reframed
the rationale — browsers do NOT truncate at the first ';' (there is
no literal ';' on the wire in the quoted form); the failing hop is a
strict middlebox cookie parser.

Closes #83832

Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
2026-08-31 15:57:21 +10:00
Lime-oss-hash f9908e2ed6 fix(bot-mode): avoid inherited stdin on Windows
Query-file DM transports do not consume stdin. Use DEVNULL for both the initial attempt and policy-gated retry so Git Bash cannot pass an invalid pseudo-handle to Windows subprocess creation.
2026-08-30 22:20:06 -07:00