One profile's broken cron store no longer takes the whole multiplex ticker
down with it:
- startup recovery loop: a per-profile exception (e.g. an unreadable
executions.db raising sqlite3.DatabaseError) was uncaught and killed the
ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
exception escaped to the cycle-wide handler, skipping every remaining
profile that cycle and marking all of them failed.
Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).
Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).
Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.
Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.
Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.
Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.
Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).
Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.
Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.
Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.
Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.
bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.
On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.
Closes#100523
Supersedes #100544, #100542
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.
`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.
The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.
Supersedes #94147 (python half)
Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:
- `_request_owns_run` no longer admits run state that exists without an
owner stamp. Under gateway.multiplex_profiles every served profile holds
a valid key, so the "backward compatibility" branch made the boundary
allow-all whenever provenance was missing. Unstamped state now fails
closed; only an in-memory owner match or a durable idempotency record
under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
inside the request's profile scope, so its run is confined to the
creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
(`_release_run_owner_if_forgotten`) and runs at every retirement point
(task finally, SSE stream close, both sweep loops, chat-stream finally),
not only the terminal-status sweep — no stranded entries, no stateful id
ever left unowned.
Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).
Fixes#93689Fixes#90415
Supersedes #93747, #93704, #92822
Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
Salvaged from #100677. Fixes#100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.
- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
it off-thread every TTL window, so every surface on the machine reads
a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.
- Prefer the wired `set_authorization_check` callback (profile-bound
`_make_adapter_auth_check`) at both sites; keep `__self__` introspection
only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
(scoped miss under multiplex → "", never os.environ); drop the raw
`os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.
Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.
Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.
`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.
Extracted from #87240 (topic-table half handled separately via #76487).
Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.
- `_make_adapter_auth_check`: for the shared primary adapter under
multiplex, mirror the inbound message path exactly — stamp the
`profile_routes` match so the routed profile's pairing store is
consulted, and authorize under the TRANSPORT home via
`_is_user_authorized_for_source` (same split as
`_make_default_profile_message_handler`, 2afed50863). A rejected route
fails closed like the ingress gate. Retain the receiving adapter as
`_transport_adapter_ref` so config.yaml policy reads stay on it.
Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
`thread_id` as keywords only when set, so legacy 3-positional callbacks
keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
introspection class — fall back to the injected `gateway_runner` and
the adapter's owner profile.
Fixes#86296Fixes#92840
Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).
Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).
Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.
- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
_build_slash_event and _dispatch_thread_session built their SessionSource
without guild_id/parent_chat_id, while on_message passes both. Route
matching in build_source keys off exactly those fields, so under
gateway.multiplex_profiles a guild- or channel-routed profile never matched
a native slash command: /new, /reset, /model, /profile, /status ... all ran
against the default profile and reset the wrong session (#69178, #91633).
Pass guild_id (interaction.guild_id, falling back to channel.guild like the
message path) and the thread's parent channel id into build_source at both
sites. One test pins channel + thread routing parity with messages.
Fixes#69178Fixes#91633
Co-authored-by: Sora-bluesky <179361977+Sora-bluesky@users.noreply.github.com>
Co-authored-by: jondgilbert <42873618+jondgilbert@users.noreply.github.com>
Co-authored-by: tensorbit89-netizen <257030052+tensorbit89-netizen@users.noreply.github.com>
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.
Fixes#100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
_process_handoff caught a config load failure for a secondary profile,
logged a warning, and kept going with self.config, which is the primary
profile's config. The handoff then went out through the right bot to the
primary's home channel and the row was reported completed. That is the
exact wrong delivery the multi-profile handoff work exists to prevent,
and the same fail closed posture the no-live-adapters branch already
takes.
A load failure now logs an error and raises, which marks the row failed
so the CLI can report and retry it. The default profile path is
untouched, it never reloaded config.
Follow-ups on the salvaged #101081 guard:
- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
guard treated that as a lost generation and permanently halted the
handle, so the #94736 late-write self-heal reopen dropped transcript
tails (4 existing tests failed). close() now clears the recorded
sidecar generation, and _wal_generation_was_lost() re-adopts the
current sidecars after a clean /proc/self probe instead of relying on
a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
generation is recorded, the stat-based inode check alone detects an
unlink/replace. The fd probe only runs in the empty-identity state
(fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
gateway retry queue and run_agent flush divert transcripts to the
JSONL fallback exactly as they do for a replaced store, instead of
retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
open, halving the system-wide /proc scan; dropped the dead
include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
non-linux test now patches sys.platform (the real gate) instead of
_IS_WINDOWS.
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).
Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.
Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.
Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
Follow-up to the salvaged #99560 commit: restore the original docstring
wording (user writes still always land everywhere else), name the
llm-outranks-derived hole the guard now closes, and annotate the two new
tests with what each pins.
Drop the no-field fallback vitest case and the get_compression_chain
unit test: the fallback is implied by the predicate (Boolean(undefined))
and the chain walk is covered by test_list_serves_full_lineage_ids_for_projected_rows
through the real list projection.
Compression rotates a conversation's tip id while tiles stay keyed by
whichever segment id they were opened with. focusOpenSession and
openSessionTile tested exact ids, so right after a rotation the same
chat read as 'not open' and opened again in a second tab — and a tile
keyed to a MIDDLE segment (the tip when it was opened) could no longer
prove it names the conversation at all, rendering as an untitled ghost.
The projected list row now carries the full chain
(SessionDB.get_compression_chain, served as _lineage_ids by
list_sessions_rich and the sidebar tree row), lineageAliases indexes
every segment, sessionMatchesStoredId accepts membership, and the tab
focus/open paths dedupe through the lineage instead of the exact id.
Older gateways omit the field and degrade to today's root/tip pairing.
Extends the TTS lease from #100912 (4e3feb8bbb) beyond built-in local
engines: when the configured tts.provider is user-declared, acquiring
the first lease and releasing the last one now reach it, so a
self-hosted TTS server can preload its model when read-aloud / voice
conversation turns on and unload when it turns off (Discord request).
- agent/tts_provider.py: TTSProvider gains concrete no-op warm() /
release() (not abstract — existing plugins are unaffected).
- tools/tts_tool.py: _signal_user_tts_provider() forwards the lease
hook; plugin providers get warm()/release(), command providers run
optional `warm_command` / `release_command` (config.yaml, under
tts.providers.<name>) through the existing _run_command_tts helper
on a daemon thread — best-effort, output discarded, failures at
debug. warm_tts_provider() and release_tts_provider() call it.
- tests/tools/test_tts_lifecycle_leases.py: fake plugin provider and
fake command provider observe warm/release through acquire/release
lease (both fail on main with action == "noop").
- docs: features/tts.md — lease section, command-provider optional
keys table, plugin optional hooks.
hermes_cli/container_boot.py resolved multiplexing from the
GATEWAY_MULTIPLEX_PROFILES env var only, while the gateway runtime
resolves env -> config.yaml -> default. A deployment enabling
multiplex_profiles via config.yaml alone therefore auto-started every
named profile's gateway slot at boot, which then crash-looped in the
double-bind guard against the multiplexing default gateway.
Resolve through load_gateway_config().multiplex_profiles (the shared
resolver, so env override precedence is preserved) and fall back to the
env var only when config loading fails.
Salvage of #85437 (test module trimmed to config-only + env-override).
Fixes#85413
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
Posts made with a user token (xoxp-) arrive with app_id and no
client_msg_id, so _event_declares_bot_sender dropped them as app traffic;
the only workaround was allow_bots: all. Adds
platforms.slack.extra.api_human_users (SLACK_API_HUMAN_USERS fallback), a
users-only allowlist consulted inside the predicate.
Salvaged from #100964 (users only: an app-id allowlist would also admit
the app's own xoxb bot posts, which share the user+app_id shape).
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).
Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).
Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
session-only rule: when neither model.default nor model.provider is
configured yet, persist. This preserves #86414's first-pick motivation
for CLI, gateway and Desktop alike. With a default configured, a plain
pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
lists --global.
- Docs: desktop.md picker note + slash-commands /model row.
Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).
Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).
Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.
Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.
Fixes#100661Closes#97766 (overflow-force idea; the bundled continuation changes were not taken)
Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):
- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
requests silently run at standard speed and bill standard rates
(usage.speed: 'standard'). Today's allowlist therefore shows 4.6
users a fast toggle that does nothing, while denying it to the two
models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
select fast inference via the model field and are explicitly
excluded from the param gate.
Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.
## How to test
scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q
113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
Follow-up to the #99641 salvage:
- One module-level _normalize_security() (ssl/tls/implicit -> tls, starttls,
plain/none -> plain; unknown -> WARNING + secure default) replaces the three
copies of the alias set; _connect_imap/_connect_smtp/_standalone_send all
compare against the canonical value. Unknown modes no longer raise.
- _tls_context(verify, host) is module-level and shared by all sites; when
verification is disabled for a non-loopback host it logs a WARNING.
- _esecret_bool: an unset/empty env var now yields the caller's default
(previously is_truthy_value('') returned False, silently disabling TLS
verification whenever EMAIL_*_TLS_VERIFY was unset).
- Documented surface is platforms.email.extra.{imap,smtp}_security and
{imap,smtp}_tls_verify in config.yaml; env vars remain an internal bridge
and are NOT added to plugin.yaml (optional_env feeds hermes setup prompts).
- Docs: Proton Mail Bridge / local relays recipe in user-guide/messaging/email.md.
- Tests: starttls builds IMAP4 then .starttls(); unknown mode falls back to
tls/starttls with verification still on.
ALIBABA_CODING_PLAN_CN_API_KEY is checked first for the China Coding Plan
endpoint (mirroring kimi-coding-cn), so the intl and CN rows no longer
light off the same key. Fixes#101122.