Commit Graph

33736 Commits

Author SHA1 Message Date
Teknium 9c021a6bde fix(desktop): Bot Mode pet picker selects locally hatched pets via pet.thumb
Selecting a self-generated pet in the Bot avatar picker always failed with
"Could not load that pet — try another", and its tile showed only a name.
Locally hatched pets have no petdex manifest entry, so pet.gallery reports an
empty spritesheetUrl; the picker cropped frame 0 client-side from that URL and
bailed on the empty string. The same raw CDN fetch also lacked a User-Agent,
which the petdex CDN rejects with 403 (#90465), so manifest pets could fail on
the same path.

Route tile rendering and selection through the gateway's pet.thumb RPC, which
already backs the Settings pet picker: it crops frame 0 server-side from the
installed sheet on disk (or the host-validated CDN URL for uninstalled pets)
and returns a same-origin PNG data URI. The client cache is keyed by slug,
evicts failures so a blip never poisons a tile, and races a 15s deadline so a
hung RPC cannot park a pending promise forever.

Based on analysis from PR #90931, whose target (plugin.js) has since been
decomposed into pet.tsx.

Co-authored-by: m1k3s0 <44042865+m1k3s0@users.noreply.github.com>
2026-09-11 19:28:29 -07:00
hermes-seaeye[bot] 2f21d29f44 fmt(js): npm run fix on merge (#108749)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-12 02:24:53 +00:00
hermes-seaeye[bot] 794981e3d4 fmt(js): npm run fix on merge (#108745)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-12 02:19:48 +00:00
Teknium 20816c13cd feat(voice): pick the voice chat engine from the composer
Switching to GPT-Live meant Settings → Voice → Voice Chat Mode, which is not
where you are when you want to talk. The composer now offers the choice where
the voice button is:

- folded layout (HUD / narrow): a "Voice chat engine" radio group in the
  existing voice menu
- unfolded layout: a small chevron beside the start-voice button opening the
  same rows; the button tooltip names the engine that will mount

Rows are hidden until the backend reports a mode (older gateway = no switch
that would 4002); GPT-Live is disabled with the backend's reason when no
OpenAI key resolves. Selecting writes `voice.voice_chat_mode` through
`config.set` on the LIVE gateway (local, SSH, cloud alike) and re-reads the
resolved status; it applies to the next conversation and never touches a
running one.

Backend: `voice.voice_chat_mode` joins the `config.set` word setters
(chained|gpt-live).

Live: headed desktop, chained → menu → GPT-Live → start = RTCPeerConnection
connected; menu → chained → start = no peer connection, chained controls.
2026-09-11 19:14:24 -07:00
Teknium f923faa0b8 feat(voice): GPT-Live voice chat mode — a full-duplex voice frontend that delegates to Hermes (Desktop)
`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS
loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and
has no tools of its own. Every real request it hears becomes a normal Hermes
turn on the open chat — any model/provider the session selected, full toolset,
memory, approvals — and the voice paraphrases the reply aloud.

Backend
- tools/voice_live.py: mode/credential/persona resolution and the one server-side
  step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer,
  pinned to client delegation; the OpenAI key never reaches the renderer. Voice
  persona follows the vendor prompting guide (role, style, labelled delegation
  policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn
  model-input note (transcript in, speakable prose out).
- REST: GET /api/audio/voice-live/status (mode + readiness, non-secret),
  POST /api/audio/voice-live/session (SDP exchange). The offer is passed
  byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF".
- prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the
  note rides the model input via the existing _prepend_note seam, the persisted
  user row stays the user's words, the system prompt stays byte-stable.
- config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*.

Desktop
- lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript
  accumulation, session.commentary/thinking/instructions appends (500-token
  chunking), mute, graceful close waiting for session.closed.
- hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation;
  delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking
  appends; reply streamed back per sentence; spoken stop phrase ends the chat;
  a newer delegation interrupts an in-flight turn.
- use-composer-voice mounts both engines and latches one at conversation start
  from the backend-resolved status; gpt-live without a key falls back to chained
  with a notice. Settings → Voice gets the mode dropdown, voice picker, persona.

Live-verified on the worktree desktop build (headless Electron, CDP, synthetic
mic): "what is 17 times 23 and which model are you on" → delegation → Hermes
(Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through
OpenRouter"; follow-up "double that" resolved from the spoken context → 782;
"run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as
quiet context → spoken kernel version; "stop" closed the session
(reason=close_requested). Chained mode creates no RTCPeerConnection.
2026-09-11 19:14:24 -07:00
Teknium fafb27ee5c feat(plugins): on_room_member_activity hook projects Group Chat member runtime events to plugins
A hosted room member runs on a hidden room_plumbing session with no client
transport, so the tool.start/complete, approval.request, message.delta and
reasoning.delta frames its turn already emits bottom out at stdio and vanish.
Between turn.started and turn.settled in the durable room log a client sees a
black box, and community clients (Hermes Crew) cannot render tool cards,
approvals or live member status without inferring them from text.

One seam in write_json (plus the connector bypass in tool_progress) re-routes
those frames, stamped with the session's _hosted_room_task coordinates
(room_id, thread_id, member_id, turn_id, task_id, execution_generation), to a
new observer hook through the bounded per-consumer queues on_stream_* already
use, so plugin code never runs on the token path. Nothing is written to the
room log: deltas would exhaust a room's byte budget in minutes and checkpoint
replay must stay a pure function of the durable events. The task stamp gains
member_id (the driver already knows it; the proof did not carry it).

Group Chat keeps execution, scheduling and persistence; plugins own
presentation.
2026-09-11 19:04:58 -07:00
Teknium 8134d941be chore: map contributor email for ishangodawatta 2026-09-11 19:04:38 -07:00
Teknium 12d03eafed test(whatsapp): trim salvaged bridge tests to the quoted-media invariants (drop cache unit test) 2026-09-11 19:04:38 -07:00
Teknium 27481a6338 docs(whatsapp): quoted replies attach the quoted media (both adapters) 2026-09-11 19:04:38 -07:00
Teknium c8b56d43da fix(whatsapp): quoting the bot's own image/voice/document attaches the file
Salvage of #77660 covers quotes of INBOUND media via the bridge's download
cache. The reported case is the other half: a user quotes an image the bot
sent (a cron-delivered chart) and asks "what is this?" — the bridge cache
knows only inbound messages, and Meta's Cloud webhook ``context`` carries
only the quoted wamid, so the agent received a text-only turn and could not
see the image it had itself delivered.

Widen ``gateway/rich_sent_store`` (already the (chat_id, message_id) → text
index Telegram and the Cloud adapter use for quoted text) with
``record_media`` / ``lookup_media``: both WhatsApp adapters index the local
path + MIME of every media send at send time and every inbound media
receive, and on a quoted reply fold the resolved ``(path, mime)`` into the
event's own ``media_urls``/``media_types`` so the existing vision/audio
pipeline handles it like a direct attachment. ``lookup_media`` drops entries
whose file no longer exists. Baileys: the outbound index is the fallback
when the bridge cache misses; the cache-dir guard stays on bridge-supplied
paths only (our own sends are paths we chose). Cloud: link sends (public
URL, no local bytes) are not indexed.

One invariant test per adapter, red on origin/main.
2026-09-11 19:04:38 -07:00
ishangodawatta e948ea8347 fix(whatsapp): don't drop bare quote-replies with resolved media
The empty-message guard only checked the reply's own body/hasMedia,
so a caption-less quote of a cached image (no text, no media of its
own) was dropped even though extractBridgeEvent had already resolved
quotedMediaUrls for it.
2026-09-11 19:04:38 -07:00
ishangodawatta cd89e4b4a6 fix(whatsapp): resolve original media for quoted-media replies
Baileys' contextInfo.quotedMessage only ever carries a thumbnail-sized
stub for media, or nothing at all for an uncaptioned attachment — never
a way to fetch the original file. When a user replies to an earlier
photo/video/document/voice note with no caption on it (e.g. "did you
save this?" quoting an uncaptioned wedding invite image), the agent
saw no text and no media reference at all: it looked like the message
never had an attachment.

Add createQuotedMediaCache, a bounded in-memory cache (keyed by
chatId:messageId) of each inbound message's already-downloaded media
and text, populated as extractBridgeEvent processes every message.
When a later message quotes one of these, extractBridgeEvent resolves
quotedMediaUrls/quotedMediaType from the cache and falls back to a
human-readable quotedText ("sent an image", etc.) when the quote had
no caption to extract. The adapter folds resolved quoted media into
the event's own media_urls/media_types — reusing the existing
vision/audio pipeline and the existing _is_allowed_bridge_path path
validation — rather than adding a parallel reply-media code path.

Reimplements the same feature as #52875 (credit: dhruvkej9) against
current main, whose 11627fdcb refactor (native polls, locations, rich
inbound metadata) moved this code into bridge_helpers.js and made that
PR's diff no longer apply cleanly.
2026-09-11 19:04:38 -07:00
Teknium fa425c942e test: stub the safe.directory pre-read by default; map privacydied's email
noninteractive_git_env() now spawns `git config --get-all safe.directory` before
building the env. Eight tests fake subprocess.run/Popen with a fixed sequence of
expected git calls (update check, plugin pull, MCP install, bounded probe) and the
extra spawn tripped them in CI. An autouse fixture stubs the read to "no entries";
the two carve-out invariant tests opt back in with @pytest.mark.real_safe_directory
(and were confirmed to still exercise the real read: the ordering test would fail
against the stub).

contributors/emails: pry@privacydied.net -> privacydied (check-attribution).
2026-09-11 19:01:47 -07:00
Teknium 66adfaee6b refactor(git): trim safe.directory tests to two invariants, memoise the config read
Tests: fold the ambient GIT_CONFIG_KEY_n negative and the isolation-still-in-force
assertions into the ordering test, and drop the two positive-only cases it subsumes.
What remains pins the injected sequence to git's own `config -z --get-all` output and
asserts the real trust decision (named repo usable, unrelated cross-owner repo refused).

Cost: noninteractive_git_env() runs on every internal git call, including the banner
startup probe, and the two `git config` children added ~10 ms per call against ~0.2 ms
before. Memoise per process on the inputs that select the config files plus the
system/global candidates' mtimes, so an edit to ~/.gitconfig is picked up without a
restart (verified live: 128 -> 0 within one process after appending safe.directory).
2026-09-11 19:01:47 -07:00
privacydied 02200f0b65 fix(git): preserve safe.directory ordering and reset markers when carrying it
Review of #107748 found the carry path serialized the user's trust policy
incorrectly, and the mistake could WIDEN trust rather than merely reformat it.

safe.directory is an ordered multi-valued protected setting: an empty value
resets every entry seen so far. That is the documented mechanism for revoking a
system-wide `safe.directory=*` and then naming only the repositories you
actually trust. The previous helper broke all three properties that make it work:

  * read `--global` before `--system` (git's precedence is system, then global)
  * dropped empty entries via `if entry:`, deleting the reset marker
  * de-duplicated values, though it is a sequence and not a set

Reproduced with real git (2.55.0 here, 2.47.3 by the reviewer). Given
system `safe.directory=*`, global `safe.directory=` then `/trusted/only`:

  real git, user's own policy                     -> rc=128 (dubious ownership)
  previous head's sequence `/trusted/only`, `*`   -> rc=0   (trust widened)
  this head's sequence `*`, ``, `/trusted/only`   -> rc=128 (matches real git)

So a user who had deliberately revoked a machine-wide wildcard silently got it
back inside Hermes's internal git calls. That contradicted the "read-only and
non-widening" claim the original change rested on.

Fix: read scopes lowest-precedence first (system, then global) and replay every
value verbatim -- no de-duplication, no dropping of empty reset markers. Read
with `git config -z` so a value containing whitespace or a newline stays the one
entry git reads it as, instead of being split into several bogus trust entries
by splitlines()/strip(). The trailing field of a -z stream is always empty and
is discarded; interior empty fields are real resets and survive.

Two invariant tests, both proven red on the previous implementation:

  * the injected sequence equals `git config -z --get-all safe.directory` for
    the same two config files, asserting git's documented reset shape
  * an E2E trust decision under GIT_TEST_ASSUME_DIFFERENT_OWNER=1: the named
    repo stays usable and an unrelated cross-owner repo is still refused,
    proving the reset still revokes the wildcard

13 passed in tests/hermes_cli/test_noninteractive_git.py, 69 in
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py.
2026-09-11 19:01:47 -07:00
privacydied 01a3206e90 fix(git): carry user safe.directory past non-interactive config isolation
`noninteractive_git_env()` blanks GIT_CONFIG_GLOBAL/SYSTEM to /dev/null so a
user's config cannot hang Hermes's internal git plumbing with pagers, hooks or
credential prompts. Sound intent, but it also discards `safe.directory` — and
git honours that key ONLY from global/system config (it is rejected from
repo-level config by design, so a hostile repo cannot self-authorise).

Result: every internal git call fails on a repo whose st_uid != geteuid():

    $ hermes -w
    ✗ Failed to create worktree: fatal: detected dubious ownership in
      repository at '/mnt/nas/py/repo'

That hits NFS/CIFS mounts without idmapping, shared checkouts, and containers
with a remapped uid. The user's own `git config --global --add safe.directory`
is correctly set and their interactive git works — Hermes throws the setting
away before git reads it, so the error's own suggested remedy can never fix it.
There is no config or env escape hatch: the blanking is unconditional.

Reproducer (any repo where the checkout uid differs from the caller's):

    git rev-parse HEAD                                    # works
    GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null \
      GIT_CONFIG_NOSYSTEM=1 git rev-parse HEAD            # dubious ownership

Fix: read the user's real `safe.directory` values before the isolation is
applied, then re-inject them over the GIT_CONFIG_KEY_n channel, which survives
GIT_CONFIG_GLOBAL=/dev/null. Isolation is unchanged — global/system config stay
pointed at /dev/null and every hardening override still applies, since a later
key of the same name wins in git's config order.

Read-only and non-widening: only values already present in the user's own
config are carried, so this grants no trust they had not granted. Ambient
GIT_CONFIG_KEY_n injection is still stripped first, so a caller cannot launder
an attacker-controlled path in this way — covered by a regression test.

Tests: three cases in TestNoninteractiveGitEnv — entries carried past
isolation, no entries injected when the user configured none, and ambient
injection not trusted. All pin GIT_CONFIG_SYSTEM at an empty file, since a real
/etc/gitconfig on the test host can otherwise leak entries and mask the
assertions.

Verified on Arch Linux, git 2.x, repo on an NFSv4 mount (uid 1024 vs caller
1000): `git worktree add` under the patched env returns rc=0 where it
previously failed. 80 passed in tests/hermes_cli/test_noninteractive_git.py,
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py, including the pre-existing assertions
that the isolation stays in force.
2026-09-11 19:01:47 -07:00
Teknium 1021a03256 chore: map contributor email for jakobdylanc 2026-09-11 16:50:20 -07:00
Teknium 3a2cfb7871 test(models): trim routing-variant tests to invariants
One relation test per consumer file: a routed id resolves to its base's
metadata while a real :free SKU keeps its own window.
2026-09-11 16:50:20 -07:00
Teknium 30e4673776 refactor(models): import openrouter_variant_base from its defining module
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
2026-09-11 16:50:20 -07:00
jakobdylanc f28c52845a fix(models): resolve routed OpenRouter ids in models.dev catalog lookups
Extends the routing-variant fix to the sibling catalog paths, so the whole
bug class is covered rather than just context length (#97820).

_find_model_entry(), lookup_models_dev_context(), get_model_info(), and
get_model_capabilities() all keyed on the full suffixed id, so a routed id
such as z-ai/glm-5.3-flash:floor missed models.dev and fell through to the
generic "glm" family default (202,752 instead of 1,310,720).

The retry with the base id runs LAST — after exact and case-insensitive
matching — so a real catalog SKU always wins over its base.

:free and :batch are deliberately NOT stripped. They are real catalog SKUs,
not routing modifiers: OpenRouter's /models currently lists 18 :free and 65
:batch entries, and 13 of those carry a context window different from their
base (z-ai/glm-5.2:free is 256K vs the base's 1.05M). Stripping them would
report a window LARGER than the model has, so the request fails at the API
instead of merely compacting early — a worse failure than the under-report
this fixes. An absent SKU must also miss rather than inherit the base, so
model_overrides _default fill-gap semantics keep working.

The suffix set is shared from hermes_constants, so validation, context
resolution, and catalog lookup all agree on one definition.
2026-09-11 16:50:20 -07:00
jakobdylanc 284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Teknium bf51fee548 docs(multiplex): make the multiplexed-gateway page match what the code does
The "one gateway for all profiles" section had drifted from the runtime. Each
claim was re-verified at its defining symbol on current main and rewritten to
the behaviour users will actually see:

- named-profile guard: only `gateway run` refuses (exit 78 / EX_CONFIG,
  systemd RestartPreventExitStatus; launchd KeepAlive still retries);
  `start`/`install` do not refuse themselves and the message lands in the
  service log; `--force` is a `run`-only flag
  (hermes_cli/gateway.py::_guard_named_profile_under_multiplexer, _cmd_start)
- same (platform, token) in two profiles: the duplicate adapter is parked as
  fatal/duplicate_credential and the gateway keeps running — it was described
  as a fail-fast startup error (gateway/run_adapters.py::_refuse_duplicate_claim)
- status surfaces: one gateway_state.json under the default home with
  `<profile>:<platform>` entries + served_profiles; nothing is written under a
  secondary home (the page claimed a per-profile runtime_status.json), and
  `hermes status` does not list served profiles — `gateway list`,
  `-p X gateway status` and /api/status do (gateway/status.py::write_runtime_status,
  hermes_cli/status.py::_render_gateway, hermes_cli/gateway.py::_cmd_status)
- API_SERVER_KEY in a secondary .env auto-enables api_server and trips the
  port-binding skip; document the `enabled: false` pin
  (gateway/config_env.py::_api_server / _enable_from_env)
- allowlist: it is a start-time snapshot, and the Desktop backend's cron
  ticker enumerates every local profile regardless of it
  (hermes_cli/web_server.py::_start_desktop_cron_ticker)
- routed-profile cron via the shared bot: only when the profile has no live
  adapter of its own, and a route carrying guild_id never matches a cron
  target because delivery matches on chat_id/thread_id only
  (cron/scheduler_provider.py::tick_adapters_for,
  cron/scheduler_preflight.py::SharedRouteAdapters.get)
- add a "What is isolated per profile" table (credentials, authorization,
  endpoints, media denylist, MCP child env, outbound egress, session
  namespace, logs, terminal) describing behaviour, not PR numbers
- configuration.md: the ${VAR} scoping paragraph now says where it applies
  and links to the table
2026-09-11 15:51:11 -07:00
Teknium bdb5bf96f3 test(hooks): synthetic payload carries the new profile field like production 2026-09-11 15:44:00 -07:00
Teknium adf23550f5 fix(tools): profile-scoped checkpoint/snapshot paths, tool caches, TZ and schema paths under multiplex
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:

- tools/process_registry.py, tools/environments/{modal,singularity}.py:
  `_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
  call time (same seam as `tools/skills_tool._skills_dir`, so the existing
  `monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
  the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
  path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
  per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
  resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
  the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
  the `_X_resolved` flags and the lifecycle reset keep their shape.
  tools/file_tools.py drops its private `file_read_max_chars` memo and reads
  the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
  env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
  startup) is ignored in favour of the routed profile's config.yaml. Both
  sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
  the static schema text is profile-neutral and `dynamic_schema_overrides=`
  rebuilds the `display_hermes_home()` / create-dir hint per
  `get_definitions()`, so a routed profile's model sees its own paths.

Refs #95685.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
2026-09-11 15:44:00 -07:00
srojk34 e70db09f51 fix(security): re-resolve checkpoint/sticker-cache paths per call
tools/checkpoint_manager.py's CHECKPOINT_BASE and gateway/sticker_cache.py's
CACHE_PATH are resolved once at import time via get_hermes_home(), which is
a context-local ContextVar under the multiplexed gateway (multiple profiles
sharing one process). Freezing the path at import time pins every later
checkpoint/cache read-write to whichever profile's HERMES_HOME was active
when the module was first imported -- the same bug class already fixed for
cache dirs, skills_hub, rich_sent_store, and (this session) the OAuth/auth.json/
sessions.json paths.

CheckpointManager is "owned by AIAgent" per-instance, but its methods read
the frozen module constant directly instead of taking the store root from
the instance, so a profile's CheckpointManager can read/write code-edit
checkpoints into a different profile's store.

Add a per-call resolver for each path, following the established "respect
an existing test monkeypatch of the constant, otherwise re-resolve through
get_hermes_home()" pattern so the extensive existing test seams in
tests/tools/test_checkpoint_manager.py and tests/gateway/test_sticker_cache.py
keep working unmodified.

(cherry picked from commit 03ae075d969094cb584e6ab38d2a773d15ff875c)
(cherry picked from commit b850c4b18e2ae2158a97c6cb87bd2057918b8170)
2026-09-11 15:44:00 -07:00
Teknium 819517fbac fix(gateway): routed profiles get their own max_turns, fallback chain, hooks, aux auth and media policy
One multiplexed gateway process serves every profile, but several per-turn
reads still went through state frozen from the LAUNCH profile:

- `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the
  module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS,
  so every secondary ran with the default profile's turn budget. A routed turn
  (HERMES_HOME override) now resolves `agent.max_turns` from its own config.
- `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide
  slot, so secondaries fell back through the default's provider/model with their
  own keys. It now reads the active gateway home and keeps a last-known-good
  chain per home.
- `_load_prefill_messages` resolved relative paths against the launch home.
- `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a
  secondary's compression/title/vision calls authenticated to Nous with the
  default profile's token when it had no pool entry. Resolved per call via
  `hermes_cli.auth._auth_file_path()` (patched constant still wins in tests).
- `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was
  loaded outside any profile scope, so secondaries' `hooks/` never ran and the
  default profile's handlers received every profile's messages, responses and
  user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the
  runner holds one registry per served home, picked from the active scope at
  emit time and front-loaded under each secondary's startup scope.
- Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME
  and the default profile's secrets). They now get the routed HERMES_HOME via
  `build_subprocess_env`, scrubbed under multiplexing, and the stdin payload
  carries `profile` so one script can tell which profile fired it.
- Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`,
  `trust_recent_files*`) was bridged once into env at startup and read from env
  per delivery; under a HERMES_HOME override the validator now reads the routed
  profile's config. Single-profile runs keep the env-bridge contract.

Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro
(temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7,
fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values.
2026-09-11 15:44:00 -07:00
srojk34 0c74353c86 security(gateway): re-resolve hooks directory per call to fix profile isolation
gateway/hooks.py::HOOKS_DIR is resolved once at import time via
get_hermes_home(), which is a context-local ContextVar under the
multiplexed gateway (multiple profiles sharing one process, each owning
its own Gateway/HookRegistry instance). Freezing the path at import time
pins every later HookRegistry.discover_and_load() call to whichever
profile's HERMES_HOME was active when this module was first imported --
so a later-starting profile silently discovers and executes the FIRST
profile's hook handlers (arbitrary Python code, not just data) against
its own live event context, including session_id/message/response text.
Same bug class already fixed for cache dirs, skills_hub, rich_sent_store,
and the OAuth/auth.json/sessions.json/checkpoint/sticker-cache paths.

Add a per-call resolver, following the established "respect an existing
test monkeypatch of the constant, otherwise re-resolve through
get_hermes_home()" pattern so the existing test seam in
tests/gateway/test_hooks.py keeps working unmodified.

(cherry picked from commit 1e4f96af681de90b942ccb77e3f96ee03913ddc2)
2026-09-11 15:44:00 -07:00
Teknium 53e32d0581 fix(env_passthrough): tolerate an unresolvable home when keying the allowlist cache
_make_run_env runs with a stripped environ on Windows children; hermes_home_key()
raises RuntimeError there (no HOME/USERPROFILE). Fall back to an unkeyed slot
instead of failing the sandbox env build.
2026-09-11 15:29:15 -07:00
Teknium d198082172 chore(contributors): map see-k's commit email for release credit 2026-09-11 15:29:15 -07:00
Teknium 72cc96578a test: trim salvaged multiplex tests to the invariant pair per fix 2026-09-11 15:29:15 -07:00
Teknium 651878e1e4 docs(multiplex): per-profile config.yaml behaviour, env_passthrough and write guards 2026-09-11 15:29:15 -07:00
Teknium 388b881b33 fix(gateway,tools): per-profile Yuanbao home, env_passthrough allowlist and Slack ignored-channel guard under multiplex
- gateway/platforms/yuanbao.py::AutoSetHomeMiddleware: the first authorized DM
  to a SECONDARY Yuanbao bot wrote YUANBAO_HOME_CHANNEL into os.environ, making
  that tenant's chat the default profile's cron/notification home. The write
  now only happens unscoped; reads go through the scoped reader + config.
- tools/env_passthrough.py::_config_passthrough: one module slot froze the
  first profile's terminal.env_passthrough for every profile's sandbox children;
  keyed by hermes_home_key().
- gateway/run.py::_slack_ignored_channels_from_gateway_config: the runner-level
  fail-safe only had the DEFAULT profile's GatewayConfig, so a secondary Slack
  bot's traffic was judged by the default's ignored list. It now takes the
  source's routed adapter (whose extra is the secondary's own config) and reads
  the env fallback through the scoped gate reader.
2026-09-11 15:29:15 -07:00
Teknium 545e74d0ea fix(gateway): a secondary profile's config.yaml no longer poisons the process env under multiplex
Under gateway.multiplex_profiles every secondary profile's config loads inside
_profile_runtime_scope, yet every apply_yaml_config_fn hook (feishu, matrix,
whatsapp, slack, dingtalk, discord non-gate keys, telegram non-gate keys) and
gateway/config_loader.py::bridge_core_env_settings still wrote os.environ there.
First-writer-wins: the first secondary with a require_mention / allowlist /
allow_bots / reactions block made that policy the DEFAULT profile's (live:
TELEGRAM_REQUIRE_MENTION written from a secondary load), and secondaries read
the default's env for the same keys.

- gateway/platforms/_shared.py::yaml_env_setter: the one env-write shape for
  YAML->env bridges — env wins, skipped under an active secondary scope.
- Every hook now seeds its values into the profile's PlatformConfig.extra and
  uses yaml_env_setter; bridge_core_env_settings seeds telegram/signal
  require_mention into extra and skips the env write under scope.
- Readers that bypassed extra/scope now consult extra first (matrix flags +
  session_scope, slack reactions, telegram reactions/mention_patterns/_extra_bool,
  discord reactions/auto_thread/history_backfill/approval_mentions/allow_mentions,
  feishu allow_bots, dingtalk mention_patterns, signal require_mention).

Salvages the shape of PR #100604 (whatsapp, earliest report #80099), #100435
(discord) and #100448 (telegram) by @nftpoetrist on top of current main.

Fixes #80099

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-11 15:29:15 -07:00
PRATHAMESH75 7af5006b24 fix(file-safety): bind the write-guard resolver fallback to the active profile
The per-call home/config getters fell back to `_expand_tilde("~/.hermes...")`
when the primary resolver raised. `_expand_tilde` follows the subprocess-HOME
contract, which under host `auto` mode can be the real/default user home rather
than the active multiplex `HERMES_HOME`. So on the exception path the guards
recreated the very cross-profile authority bug the happy path fixed: beta's
`config.yaml` was compared against the default/root config (hard-block fails
open), and the protected-instruction exemption resolved against the wrong home.

Re-derive both fallbacks from the same `get_hermes_home()` key the happy path
uses (`Path(home)/config.yaml`, `realpath(home)`), and substitute no unrelated
home if the active security path cannot be established — a `None` fails closed
at the protected-instruction consumer (exemption skipped, gate runs).

Adds opposite-side regressions: forcing the primary config resolver to raise
keeps beta's own config refused (and does not spuriously protect alpha's under
beta's scope); forcing the primary home resolver to raise keeps beta's
instruction-file exemption resolved against beta.

Addresses the fallback-authority review on #107335 (thanks @andrexibiza).

(cherry picked from commit 119d88b46f745ba081f12adf3f6ebba457d95d68)
2026-09-11 15:29:15 -07:00
PRATHAMESH75 1271622e4b fix(file-safety): resolve HERMES_HOME/config per call so multiplex profiles don't poison the write guards (#107327)
In a multiplexed gateway (`gateway.multiplex_profiles: true`) each profile turn
scopes `HERMES_HOME` through a per-turn contextvar. But
`tools/file_tools_write_guards.py` memoised the resolved home and config path in
process-global module state, filled once by whichever profile ran first. Both
the protected agent-instruction approval gate (`_get_real_hermes_home` →
exemption for a profile's own home) and the `config.yaml` hard-block
(`_get_hermes_config_resolved`) therefore became order-dependent: a later
profile's own `workspace/AGENTS.md` was gated against a *sibling* profile's home,
and — worse — its own `config.yaml` stopped matching the block, so a
prompt-injected agent could rewrite the very file the block exists to protect
(reproduced end-to-end in #107327).

Resolve both values per call instead. `get_hermes_home()` / `get_config_path()`
are contextvar-scoped, so the getters now track the active profile; the guard
already pays a `realpath` per call, so the extra cost is negligible. The two
module slots are kept purely as a test-override surface (set the slot + its
`_loaded` flag to pin a value); production leaves them unset and resolves live,
which also removes the cross-test poisoning the process memo could cause.

Adds regression coverage: both getters track the active profile after a prior
profile's scope, and the `config.yaml` hard-block fires for beta's own config
even after an alpha turn ran first.

(cherry picked from commit 36b257391da497ac31e6c440727dc57ecd584e71)
2026-09-11 15:29:15 -07:00
infinitycrew39 606903badc fix(tui): bind launch-profile terminal scope once multiplexing is active
After any secondary profile home is served, launch-profile turns used to stay
unscoped and fall back to ambient os.environ. Bind the launch home's own
terminal policy in that case so a poisoned ambient bridge can never become
the launch turn's authority (#107422 residual of #68559).

(cherry picked from commit f81147c1e5d283837e5e27f4da79710da7025235)
2026-09-11 15:29:15 -07:00
infinitycrew39 a5c801c8dc fix(tools): never ambient-bridge TERMINAL_* under a profile home override
A multiplexed dashboard can call _ensure_terminal_env_bridged while a
secondary profile's HERMES_HOME override is active. The one-shot latch then
wrote that profile's docker policy into process-global os.environ and poisoned
later unscoped launch-profile tool calls (#107422).

Skip the ambient bridge whenever a context-local home override is set —
ambient env is launch-profile authority only; routed profiles must use
terminal_scope (same rule as env_loader._reapply_terminal_config_bridge).

(cherry picked from commit 2050efb24fdc9d54a282b24d0042b90f47486c5a)
2026-09-11 15:29:15 -07:00
Chike Okonta 954113839b fix(discord): apply YAML allow_bots to ingress policy
config.yaml discord.allow_bots was accepted but ignored because
_get_allow_bots() only read DISCORD_ALLOW_BOTS. Seed the key through
_apply_yaml_config like the other Discord gates, then resolve via
_gate_raw so env still wins over YAML.

(cherry picked from commit ddeb1bd39253404a3c0b43bc65372e5ada61bbbf)
2026-09-11 15:29:15 -07:00
Teknium d807a34c87 docs(multiplex): cron ticker allowlist, named multiplexer, guild-scoped route delivery 2026-09-11 15:28:37 -07:00
Teknium 1c08edccf0 fix(gateway): a named-profile multiplexer ticks its own cron store and owns the shared adapters
profiles_to_serve(multiplex=True) yields default + allowlist, so a gateway run
as `hermes -p <name>` with multiplex on never ticked its own profile's jobs
unless allowlisted (which would start a second adapter on the same token). The
ticker's home list now unions the active profile. The shared-adapter owner
passed to the ticker is the runner's launch profile instead of the literal
"default", so that profile's jobs reuse its live adapters rather than the
fail-closed empty map.

Co-authored-by: Paul Pincente <101599379+pincente@users.noreply.github.com>
Co-authored-by: r3x443 <325334945+r3x443@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
Teknium a6c5b7ada8 fix(serve): SSH-isolated idle-exit keeps the backend alive while a cron job runs
turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.

Addresses #107485
2026-09-11 15:28:37 -07:00
Teknium 0e57543908 fix(desktop): cron ticker follows the multiplexer allowlist and stands down for served satellites
The Desktop/serve backend ticked every installed profile (ignoring
gateway.multiplex_profile_allowlist) and gated only on the profile's OWN
gateway.pid. A satellite served by the default multiplexer has no pid file, so
both tickers raced its fires and the Desktop one won nondeterministically —
adapter-less standalone delivery, and the environment behind #107485.

Homes now come from profiles_to_serve with the default profile's allowlist
(the multiplexer's served set); the per-tick gate also consults
named_profile_served_by_running_multiplexer.

Addresses #107485, #94590
Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-11 15:28:37 -07:00
Teknium 444fa8166a fix(cron): routed-profile cron delivers through the shared bot for guild-scoped routes and profiles without a platforms block
SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.

_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.

Fixes #89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
tachyon-r 32c538851a fix(desktop): yield single-profile cron to its running gateway
(cherry picked from commit db92ff258f0138a251b2c4e14b32f2fa0b7b8d5b)
2026-09-11 15:28:37 -07:00
fangliquanflq c17629a0a2 fix(cron): scope restart-safe worker environment
(cherry picked from commit e57f719f942736104ee7ff79999f41b8f2a23d63)
2026-09-11 15:28:37 -07:00
Teknium 3efbd79a30 chore: map salvaged contributor emails (remi-td, nmediaie) 2026-09-11 15:28:00 -07:00
Teknium cdedfe9770 test(gateway): invariants for multiplex routing/authz (#104933, #103717)
Five behaviour contracts, each red on origin/main: shared-bot route does not
hijack a dedicated secondary bot; secondary busy follow-up authorized against
its own allowlist; mid-turn authorization reads the admitting transport's
allowlist; shared-bot satellite resolves the primary transport for restored
sources while a downed secondary stays fail-closed; completion pre-flight runs
in the target profile's scope.
2026-09-11 15:28:00 -07:00
Teknium c632437c3b fix(gateway): deliver secondary-profile async completions under their own profile scope
The supervised `_async_delegation_watcher` and startup-recovered process
watchers run under the ROOT scope, so a secondary profile's completion was
classified against the DEFAULT profile's state.db (row absent → "terminal" →
"permanently-gone session" warning, delivery dropped) and every durable-ledger
op (`claim`/`complete`/`release`) hit the default's ledger, stranding the real
row `pending` forever.

`_deliver_completion_notification` and `_deliver_async_delegation_group` now
run their whole pre-flight + claim + inject + settle sequence inside
`_completion_event_scope(evt)` — the runtime scope of the profile the event's
session belongs to. At multiplex startup `_restore_secondary_completion_ledgers`
replays every secondary's pending rows, which the process registry (launch
home only) never saw.

Builds on the classify-only half of #107247.
2026-09-11 15:28:00 -07:00
Teknium 77180acc2c fix(gateway): mid-turn authorization reads the admitting bot's allowlist; shared-bot satellites keep a transport
Inside a routed satellite's turn the ambient scope is the satellite's, whose
.env has no token or allowlist. Five sites still called `_is_user_authorized()`
directly there — `/topic`, the sibling-thread `/stop` grant, plugin message
injection, Discord voice transcripts and startup auto-resume — so the shared
bot's owner was refused ("not authorized to use /topic") and a satellite that
DID copy an allowlist widened who may drive those commands. They now go through
`_is_user_authorized_for_source`, and `_under_authorization_profile` derives the
transport home from the delivering adapter's owner when ingress did not stamp
one (restored/cached sources).

`_adapters_for_profile` (the resolver behind `_authorization_adapter`,
`_adapter_for_source` and now `_resolve_injection_adapter`) returns the primary
map for a shared-bot satellite: a served profile with the `{}` startup
placeholder, no reconnect pending, targeted by a default-bot route. Kanban and
cron already applied that rule; the gateway's own resolvers returned None, so
heartbeats, process completions, goal notices and delegation results for such
profiles were undeliverable after a restart. A secondary that owns a credential
(connected on any platform, or queued for reconnect) still fails closed.
2026-09-11 15:28:00 -07:00
Teknium 7ecee9e78f refactor(gateway): share the primary admit/scope step between message and busy handlers
`_make_default_profile_busy_session_handler` (salvaged from #105357) duplicated
the stamp-transport-home / stamp-route / resolve-home block of
`_make_default_profile_message_handler`. Both now call `_admit_primary_source`,
so the two ingress paths cannot drift apart again (#103717 was exactly that
drift).
2026-09-11 15:28:00 -07:00