On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
`hermes import` published every zip member, including `state.db`, with
`_extract_member_atomically` — a rename that swaps the file's inode. Any
gateway, dashboard, or WebUI process holding the database open keeps its
descriptor on the now-unlinked inode: it goes on serving pre-import pages
and writing sessions no other process can see, while the sidecar WAL left
beside the new file describes the database that was just unlinked. Nothing
raises, so the import prints "Import complete" and the sessions are simply
absent from the database everyone opens next.
The live-safe path already exists: `/snapshot restore` has routed `.db`
files through `_safe_restore_db()` since #65942, writing snapshot pages
into the existing file so every open connection converges. `hermes import`
— the disaster-recovery path, reached by users who already lost something
once — never got that treatment.
Route `.db` members through it. A target that does not exist yet has no
holders and no inode worth preserving, so it keeps the ordinary atomic
publish. A refused or failed live-safe restore now raises, so the import
reports a skipped file instead of counting a silent success, and the
existing database is left untouched.
Importing an older backup over newer work stays allowed but no longer
silent: the summary reports the session/message counts the import replaced,
the same before/after evidence `restore_cron_jobs_if_emptied` uses for
`cron/jobs.json`.
Closes#100960
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ShTVU941HYygvypY9JMYE
(cherry picked from commit 8ff260a312341cb85bdf7cdd570de5ff888c6013)
OpenCode pins requests sharing an x-opencode-session value to one upstream
backend, which is what keeps its prompt cache warm across a conversation.
Hermes never sent it, so cache ratios on OpenCode traffic were poor.
- agent/opencode_affinity.py: single owner of the header — target detection
(built-in zen/go/free, custom opencode-* providers, any opencode.ai URL)
and the key (affinity scope → conversation root → session id, cron
timestamp stripped), same resolution as OpenRouter/xAI affinity hints.
- build_api_kwargs: merged once after the per-mode builder, so
chat_completions, codex_responses and anthropic_messages all carry it.
- auxiliary _build_call_kwargs: same key from the runtime-main session so
compression/title/vision calls stay on the conversation's backend; the
aux Codex and Anthropic adapters now forward extra_headers.
Closes#81584, #81832 (deepseek-v4-flash 400 without the header).
Salvage follow-up for PR #100269. The flag landed upstream in 0.3.x;
binaries before it (e.g. 0.2.8, verified locally) fatally reject the
flag with 'unknown argument', breaking every Browser Use launch.
Probe 'lightpanda help' once per process and omit the flag when the
binary predates it. Also soften the unverified concurrency claim in
the _http_cache_dir docstring and tell docs readers to stop sessions
before deleting the live sqlite cache.
Lightpanda's HTTP cache is opt-in (`--http-cache-dir`, off by default), and
the launcher never passed it, so every Browser Use navigation re-fetched
every asset.
Point all Hermes-spawned instances at one shared cache under
$HERMES_HOME/cache/browser-use/lightpanda/http-cache. Sharing it across
sessions keeps assets warm through session churn; Lightpanda stores it in
sqlite (WAL), so a write that loses a race degrades to a cache miss rather
than a failed load, and --http-cache-entry-limit (default 1000) bounds the
directory without Hermes managing eviction.
Measured over 25 navigations across 5 sites, median warm navigation drops
from 0.40s to 0.18s on news.ycombinator.com and 0.14s to 0.10s on
wikipedia; total navigation time 7.0s -> 5.9s.
The flag has existed since Lightpanda 0.3.x (April 2026), so this needs no
minimum-version bump.
Expose the read deadline as a top-level config.yaml key beside
context_file_max_chars (same load_config_readonly resolution shape), default
5s, documented in context-files.md. Narrow the reader thread's catch from
BaseException to Exception: control-flow exceptions can't originate inside
read_text on a worker thread, and re-raising one would bypass the sites'
except Exception / except (OSError, UnicodeDecodeError) handlers.
Twenty-three comments arrived as twenty-three flat blocks, so the agent made
twenty-three todos and ground through them one at a time. They now arrive
grouped by where they sit in the page, with a line telling the agent to work
the groups rather than the comments.
The renderer groups on structure, not meaning. Whether a comment is a UI nit
or a functional bug is a judgment only the model can make, and prose-matching
it here would be wrong constantly; which pins share a DOM subtree is something
the selector already answers. That split is also the one that makes parallel
work safe — grouping by theme instead ("all the spacing ones") cuts across the
same components and puts several workers in the same files, so the guidance
says to hand out whole groups and never to regroup by theme.
Grouping compares ancestor paths, so a heading and a paragraph in one card
stay together instead of becoming two singletons. Depth is derived rather than
tuned: descend the shared prefix until it stops being shared, then sub-split
any group still holding more than a third of the batch — without that pass a
normal page buries every section under `main`. Batches under four comments,
and batches that all land in one region, stay flat.
Grouping is advice in the prompt, never an action: the renderer does not spawn
or delegate anything. That stays the agent's call.
Comment mode shipped the crop and the note, so an agent got a picture of the
problem and had to grep for the element it showed. Each element comment now
also names its CSS selector, its markup, and the computed styles that decide
layout, which is what the agent needs to land in the right file.
The target line stays prose — it is what the user pointed at — and the DOM
detail rides labelled lines beneath it. Area pins have no element, so they
still get only the crop and the note.
Markup is redacted in the guest before it crosses to the host: password and
hidden input values, and any attribute reading as a key/token/secret, are
replaced with [redacted] on a clone, so a page's secrets never reach the
composer or the model. It is clipped to a 600-char budget so one comment
cannot paste a whole section.
AnnotateIdentity was a hand-copy of CompactIdentity that had already drifted;
it is now an alias, so the guest, the pin, and the packer cannot disagree
about the shape again.
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.
- resolve_nous_runtime_credentials(stale_access_token=): under the store
lock, skip the refresh POST when the on-disk token differs from the one
that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
pass the failed bearer through, and treat a lock TimeoutError as
'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
refresh/0 unrecovered.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Flip the state.db retention defaults per Teknium's decision on #54189:
- sessions.auto_prune: false -> true. A stock install now prunes ENDED
sessions inactive for retention_days at CLI/gateway/cron startup
(at most once per min_interval_hours). Open, pinned and mid-turn
sessions are never deleted; the only open rows touched are stale
automation sessions (#100903 sweep), which are closed, not deleted,
and aged a further full window before removal.
- sessions.retention_days stays 90 (already the default; verified).
- Auto-VACUUM is now additionally gated on the reclaimable fraction of
the file: PRAGMA freelist_count / page_count must exceed 25%
(AUTO_VACUUM_MIN_FREELIST_RATIO) on top of the existing
min_vacuum_interval_days throttle. Pruning a few small sessions on a
dense multi-GB DB no longer rewrites the whole file to reclaim a few MB.
Unknown ratio (pragma read failure) falls back to the time throttle.
Existing installs that explicitly set any sessions.* key keep their
values (load_config deep-merges DEFAULT_CONFIG under user YAML); only
unset keys pick up the new defaults. No _config_version bump needed.
cli-config.yaml.example documents the section commented-out so
installers that copy it verbatim never pin these as explicit settings.
Tests: ratio gate (below/above/at-threshold/unknown/override), real-DB
freelist ratio, default assertions, fresh-config startup hook reaches
the prune call, explicit opt-out respected, template-does-not-pin-keys.
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:
* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
agent.tools from live probes with no predecessor to preserve. Persist
the session's resolved tool-name order in a new `sessions.tool_names`
JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
alongside the system prompt and re-pinned on every published refresh
(so /reload-mcp and compaction naturally reset it; /new mints a new
row). On restore-for-existing-session the fresh definitions are folded
onto the saved order via the SAME `_merge_preserving_prefix` helper —
a probe-flipped tool is carried forward from the registry schema, a
deregistered one dropped, new tools appended at the tail.
* /reload-mcp (CLI, gateway, TUI RPC) now also calls
`reprobe_tool_availability()` — drops the check_fn verdict cache and the
get_tool_definitions memo — so a user can consciously pick up a
credential/daemon that appeared mid-session. Docs updated.
Under a multiplexed gateway every profile's outbound webhooks share one
delivery worker, so receivers could not tell which profile fired an
event. Add a top-level `profile` field to the payload, resolved at fire
time from the bound Hermes home via get_active_profile_name() ("default"
outside profiles). Documents the field in the wire-format section.
Reported by @vszgdcn8cj-ctrl.
Fixes#92674
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).
Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.
- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
`_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
retry purge and the integrity self-heal only ever clear the staging tree,
never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.
Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.
Closes#86443
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
`ProfileRoute.matches()` compared `chat_id` as an exact string, so a
WhatsApp route written as a phone number never matched the JID
(`…@s.whatsapp.net`) or LID (`…@lid`) the bridge actually delivers, and
the inbound fell through to the default profile. Allowlists and session
keys already canonicalize these via `gateway.whatsapp_identity`.
Exact compare still wins first; only whatsapp / whatsapp_cloud *user*
chats get the alias intersection fallback (applied to both `chat_id` and
`parent_chat_id`). Groups, broadcasts and every other platform stay exact.
Salvaged from #85081 (Kong); parent_chat_id fallback added on top.
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.
Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.
Salvaged from #70831 (Hudson). First reported in #68332.
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:
- `_request_owns_run` no longer admits run state that exists without an
owner stamp. Under gateway.multiplex_profiles every served profile holds
a valid key, so the "backward compatibility" branch made the boundary
allow-all whenever provenance was missing. Unstamped state now fails
closed; only an in-memory owner match or a durable idempotency record
under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
inside the request's profile scope, so its run is confined to the
creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
(`_release_run_owner_if_forgotten`) and runs at every retirement point
(task finally, SSE stream close, both sweep loops, chat-stream finally),
not only the terminal-status sweep — no stranded entries, no stateful id
ever left unowned.
Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).
Fixes#93689Fixes#90415
Supersedes #93747, #93704, #92822
Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.
- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
it off-thread every TTL window, so every surface on the machine reads
a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
Follow-up on @fortun8te's user-made roster sections:
- Sections start empty: no seeded General/Workforce/Clients. With no
sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
(the app's session-rename shape) instead of an inline caret; the row
menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
shows a faint outline while a drag is live, the hovered target lights
up, the source section refuses the drop, Escape cancels, and the moved
row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
the roster has no selection model. Per-bot saveBotMeta writes run in
sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
`onDoubleClick`), so user sections fold and look like the gateway
headings; ⋯ menu and right-click drive the same Rename / Move up /
Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
bucket, indented under a hairline rail (membership lives in the bot's
profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
saveBotMeta + reload, remainder = Unassigned, delete returns bots +
undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
- use-desktop-integrations: hold the restore latch until the config
record answers; when false, stay on the fresh chat (route and session
restore alike) while still remembering the open chat for next launch.
- wiring: read the shared config-record query; undefined while pending,
fetch failure falls back to the historical behavior (resume).
- appearance-settings: ToggleRow writing through the shared config
cache with rollback + notifyError on a failed save.
- ar/ru strings, docs line in user-guide/desktop.md, two hook tests.
Every bot's canonical chat is stored under the same title ("Bot Chat" — the
name the gateway resolves it by, and an invariant roster-actions.ts's stale-tile
probe and #90102 rely on), so the main tab strip captioned every open bot chat
identically and two bots' tabs were indistinguishable (#99152).
Fix at the presentation layer, leaving the stored title and tabTitle untouched:
- workspace-scope.ts gains `$workspaceOwnerLabels` + `workspaceOwnerTitle()`:
a bots-mode tab whose resolved title still equals its registered placeholder
reads its owner's label instead. Side threads / Sessions tabs are untouched.
- session-tile.tsx captions tiles through it (and the drag payload); the main
`workspace` tab (controller.tsx) does the same via `$botChatScopes`, the
bot-mode scope the main tab was last opened under (it has no tile).
- The hermes-bots roster publishes displayName() per owner key through the new
`host.setWorkspaceOwnerLabel` (feature-detected), so renames follow.
Supersedes #99177, which set tabTitle at open time — that reverts after mount
because tileTitle() prefers the stored row's title once the hidden row is
upserted, and breaks the `workspaceTabTitle === 'Bot Chat'` invariant.
Tests: one unit test on workspaceOwnerTitle() (bot chat → bot name; side
thread / sessions tab / unlabeled owner untouched) and one Electron e2e
(tab strip reads "Alpha", not "Bot Chat"); both fail on main, pass here.
Closes#99152
Supersedes #99177
Co-authored-by: twotnguyen <nguyenngoctinh011258@gmail.com>
Extends the TTS lease from #100912 (4e3feb8bbb) beyond built-in local
engines: when the configured tts.provider is user-declared, acquiring
the first lease and releasing the last one now reach it, so a
self-hosted TTS server can preload its model when read-aloud / voice
conversation turns on and unload when it turns off (Discord request).
- agent/tts_provider.py: TTSProvider gains concrete no-op warm() /
release() (not abstract — existing plugins are unaffected).
- tools/tts_tool.py: _signal_user_tts_provider() forwards the lease
hook; plugin providers get warm()/release(), command providers run
optional `warm_command` / `release_command` (config.yaml, under
tts.providers.<name>) through the existing _run_command_tts helper
on a daemon thread — best-effort, output discarded, failures at
debug. warm_tts_provider() and release_tts_provider() call it.
- tests/tools/test_tts_lifecycle_leases.py: fake plugin provider and
fake command provider observe warm/release through acquire/release
lease (both fail on main with action == "noop").
- docs: features/tts.md — lease section, command-provider optional
keys table, plugin optional hooks.
- display.bell_on_approval (default false): same BEL mechanism as
bell_on_complete, rings when a dangerous-command approval prompt
opens (_approval_callback / approval.request event). Complements
bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
(if without braces) that failed the CI JS & TS checks job.
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
Posts made with a user token (xoxp-) arrive with app_id and no
client_msg_id, so _event_declares_bot_sender dropped them as app traffic;
the only workaround was allow_bots: all. Adds
platforms.slack.extra.api_human_users (SLACK_API_HUMAN_USERS fallback), a
users-only allowlist consulted inside the predicate.
Salvaged from #100964 (users only: an app-id allowlist would also admit
the app's own xoxb bot posts, which share the user+app_id shape).
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).
Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).
Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
session-only rule: when neither model.default nor model.provider is
configured yet, persist. This preserves #86414's first-pick motivation
for CLI, gateway and Desktop alike. With a default configured, a plain
pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
lists --global.
- Docs: desktop.md picker note + slash-commands /model row.
Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).
Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).
Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.
Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.
Fixes#100661Closes#97766 (overflow-force idea; the bundled continuation changes were not taken)
Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
Follow-up to the #99641 salvage:
- One module-level _normalize_security() (ssl/tls/implicit -> tls, starttls,
plain/none -> plain; unknown -> WARNING + secure default) replaces the three
copies of the alias set; _connect_imap/_connect_smtp/_standalone_send all
compare against the canonical value. Unknown modes no longer raise.
- _tls_context(verify, host) is module-level and shared by all sites; when
verification is disabled for a non-loopback host it logs a WARNING.
- _esecret_bool: an unset/empty env var now yields the caller's default
(previously is_truthy_value('') returned False, silently disabling TLS
verification whenever EMAIL_*_TLS_VERIFY was unset).
- Documented surface is platforms.email.extra.{imap,smtp}_security and
{imap,smtp}_tls_verify in config.yaml; env vars remain an internal bridge
and are NOT added to plugin.yaml (optional_env feeds hermes setup prompts).
- Docs: Proton Mail Bridge / local relays recipe in user-guide/messaging/email.md.
- Tests: starttls builds IMAP4 then .starttls(); unknown mode falls back to
tls/starttls with verification still on.
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
ones it shares with its non-CN sibling, unless that CN provider is the
configured model.provider. With only the shared key: one row, not two;
DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
purge_stale_done_notify_subs only matched status='done', so a task the
circuit breaker parked in 'blocked' kept its notify-sub rows forever on
boards that never archive. Widen the predicate to done OR blocked while
keeping the existing age clause; backlog/ready cards are idle, not
abandoned, and stay exempt (test_gc_spares_reopened_task_even_when_old).
Watcher comment/log and docs updated to say done/blocked.
Closes#100955
Co-authored-by: itsflownium <itsflownium@users.noreply.github.com>
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.
- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.
Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.
- roster-actions: the open-tab shortcut may front only the canonical chat
(registry id or lineage tip, via a new onlyStoredIds allowlist on
focusWorkspaceOwnerSessionTile); anything else resolves the registry and
opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
none reach this window's stream), re-open it in place so the transcript
refreshes instead of waiting for an app restart (#99393 class).
Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.