Conflict in agent/error_classifier.py: both simp/agent-chat and
simp/agent-runtime simplified the same file. Kept simp/agent-chat's
rule-table version, which already subsumes agent-runtime's
billing/rate-limit/overflow verdict helpers (as _V_* verdicts + _first_match
rule tables) and its dead _THINKING_SIG_PATTERNS removal. Verified with a
60k-case differential fuzz: identical ClassifiedError output vs base and vs
agent-runtime's version.
- deadline: _result()/_abandon() replace 7 BoundedResult constructions and 2
cancel+callback sites; timers handled as a list; dead raise_if_timed_out removed.
- file_safety: retired classify_cross_profile_target (0 refs; get_cross_profile_warning
stub kept for external callers), _home_and_resolved/_mirror_warning shared by
the sandbox/container mirror guards; _find_sandbox_mirror_segments inlined.
- estop: _hermes_home/_canonical_root now the file_safety helpers;
_reset_log_state_for_tests inlined into its only test.
- subagent_lifecycle: _validate_request driven by _UNSUPPORTED_REQUEST_FIELDS table.
- turn_liveness: dead start()/_abort_message removed, _emit_warning shared.
- process_bootstrap: _enable_happy_eyeballs reused by the client variant.
- Comment/docstring compaction across the remaining leaf modules.
- shell_hooks is the shared home: _ToolMatcherMixin (matcher compile + matches_tool),
_payload_fields, _forget_home_registrations, _home_key, _utc_now_iso now serve
outbound_webhooks too (copies deleted; every log string byte-identical).
- shell_hooks: response parsing is a per-event dispatch table; _spawn diagnostic
dict + _evaluate_result shared by the live callback and run_once;
_locked_update_approvals POSIX/non-POSIX bodies merged via ExitStack.
- tool_guardrails: ToolCallGuardrailConfig thresholds from a _THRESHOLD_SOURCES
table (nested-wins-over-flat preserved); _int_at_least replaces
_positive_int/_non_negative_int; observe_identical_call (0 refs) folded into
observe_call; _halt helper for hard-stop decisions.
- tool_dispatch_helpers: _plan_tool_batch_segments split into _batch_admission +
close/extend helpers with the post-hoc normalization merged in.
- Comment/docstring compaction keeping every stated rule.
Drop dead read_error_body_or_default / LAYER_RUNTIME; inline single-use
_build_default_credential/_safe_close; compact incident narratives to their
invariants in the guards. Byte-cap and deadline semantics of
read_streaming_error_body verified identical.
PluginLlm's four public entry points share _gate/_finish/_host_kwargs; drop
dead classify_failure_scope/_REASON_SCOPES (and their tests) and unify the
three _norm_* helpers; should_skip_candidate routes through a scope predicate
table. Injected caller kwargs, audit dicts and log lines unchanged.
Drop dead call_converse_stream / classify_bedrock_error / is_context_overflow_error
(zero refs) and their tests; stop-reason mapping becomes a dict; extract
_cache_point/_assistant_blocks/_append_turn/_cached_client helpers; compact
incident narratives to their invariants. Converse wire output byte-identical.
- build_environment_hints: split into _local_host_hints / _remote_backend_hint /
_embedder_environment_hint; backend probe split into _run_backend_probe +
_format_backend_probe with image-key / container-config dispatch tables
replacing the if/elif chain.
- Skills index: _SkillFilter (frozen dataclass) unifies the disabled+conditions
check that was copied 4x (snapshot, scan, project, external);
_collect_extra_skills dedupes the project/external scan loops;
_read_category_descriptions dedupes DESCRIPTION.md reading;
_label_visible_entries and _render_skills_index lift the org-labeling and
rendering regions out of _build_skills_system_prompt_inner; snapshot and scan
sources now feed one visibility pass.
- Context files: _read_context_file + _context_section unify the
read/strip/scan/section/truncate sequence across .hermes.md, AGENTS.md,
CLAUDE.md and .cursorrules loaders.
- Dead: _clear_backend_probe_cache (test-only helper; tests clear the dict
directly), unused org_id_of_path re-export.
- Comments/docstrings hand-compacted; every rule, invariant, ordering and
failure-mode rationale kept.
System prompt text verified byte-identical against origin/main over a 273-case
fixture corpus (env hints x backends/probe states, skills index x toolsets /
platforms / project / org / compact, context files x all loaders, full
AIAgent._build_system_prompt_parts x 9 configs). Tool schema byte-identical.
init_agent (2711 LOC) becomes a ~280-line ordered orchestrator over
_resolve_api_mode / _finalize_routing / _init_* / _build_client /
_load_tools / _parse_compression_config -> CompressionSettings /
_resolve_context_length / _build_context_engine / ... phase helpers.
_build_client is further split per wire mode (_init_anthropic_client,
_init_moa_client, _init_bedrock_client, _init_openai_client with
_explicit_client_kwargs / _routed_client_kwargs). Statement order and
every side effect on the agent are preserved (AST body-parity checked
against origin/main).
Dedupe/dead code: drop _relay_moa_reference_event/_moa_reference_output_allowed
(zero callers; only their own test) and their test file; alias
_normalize_route_base_url; _parse_config_int replaces three copies of the
strict int parser; _cfg_flag replaces four inline truthy-set checks;
_client_kwargs_from_routed + _fallback_entries replace duplicated
routed-client/fallback-entry blocks; _warn_invalid_config_int unifies the
three log+stderr invalid-int warnings (byte-identical text);
_bedrock_region_from_url; _memory_provider_init_kwargs; the
host->default_headers if/elif chain becomes the _HOST_DEFAULT_HEADERS
dispatch table; callback params assigned from _CALLBACK_PARAMS.
Comments/docstrings hand-compacted to their rationale (invariants,
ordering, failure modes kept; issue numbers and narrative dropped).
test_pre_compress_checkpoint_contract source-check repointed at the
CompressionSettings field names.
Verified: tests/run_agent (2066 passed) + all agent_init-referencing tests
(1134 passed), get_tool_definitions() byte-identical vs origin/main, import
smokes for cli/run_agent/gateway.run/hermes_cli.main/agent.conversation_loop/
tui_gateway.server.
Submodule and worktree checkouts store .git as a file (gitdir: pointer);
the carry-over only looked at directory names, so that form was still lost
on rollback. Handle files with the same guard. Docstring now states the
deliberate limit: an excluded entry whose skill dir the target snapshot
lacks is dropped with staging (no orphan .git) and is not undoable via the
safety snapshot, which excludes these paths as well.
Excluding nested .git from snapshots has a side effect on rollback: the
staging move takes the whole live skill dir (including its .git) into
.rollback-staging-*, the extract restores the snapshot without it, and the
staging dir is then deleted — so a skill that is itself a git checkout lost
its .git on any rollback. Reproduced: main preserves it, the exclusion-only
branch did not.
After a successful extract, move excluded subtrees from the staged copy back
under their restored skill dir (mirroring how a top-level .git survives by
never being staged). Regression test included.
f50b5bb0fa taught the sync retry site to keep the cheap same-provider retry
when a Codex stream dies inside the 60s no-progress window (zero output),
skipping straight to fallback only on a stall or hard-ceiling timeout. The
async site never got that carve-out, so after widening the skip to vision
(#97572) an async vision call on a stillborn stream would have jumped to
fallback where the sync path retries. Both sites now apply the same rule.
Adds the async twin of the vision-skip test and a no-progress-still-retries
guard for the async site.
Issue #54465 established that a same-provider retry after a full-budget
timeout costs a second whole `timeout` window before the fallback chain is
reached, doubling the user-visible stall, and that compression must not pay
it because it sits on a critical path. The guard added for that is spelled
`task == "compression"`, so vision — which sits on the interactive path —
still retries.
The cost is the same and the stall is more visible: the turn holding the
image cannot answer, and because turns are serialised the following user
messages queue behind it. Two sequential full-budget timeouts on an
unhealthy vision provider is a long stall for something the fallback chain
could have served immediately.
Replaces the string comparison at both retry sites (sync `call_llm` and
`async_call_llm`) with `_TIMEOUT_NO_RETRY_TASKS = {"compression", "vision"}`,
so the two paths cannot drift again. Behaviour is unchanged for every other
task: fast blips (a streaming-close or a 5xx) still retry, and only
full-budget timeouts on those two tasks skip straight to fallback.
Tests: vision now falls straight through to fallback with the primary tried
exactly once, and a non-critical task still gets its one same-provider
retry, so the change stays scoped. Reverting the source change fails the
vision test and leaves the scoping test green.
Not the same as #51513, which fixes five separate defects in the vision
fallback chain (capability detection, sync/async client misuse, geo-block
and RemoteProtocolError classification, and chain iteration). This is about
what happens before that chain is reached.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On the default Nous config, call_llm acquires the auxiliary client via
_get_cached_client(resolved_model=None), so the cache key's model
element is "". On a 401, _refresh_nous_auxiliary_client rebuilt the
client but keyed the new entry on the resolved wire model (final_model,
e.g. "Hermes-4-405B"). The fresh client therefore landed under a
different key than the lookup, and the stale expired-credential client
under "" was never overwritten: every auxiliary call kept hitting the
dead client, 401ing and forcing a credential portal round-trip on each
request instead of self-healing after the first refresh.
The auto-provider dimensions had the same divergence: call_llm and
async_call_llm dropped task at both acquisition sites, and the async
path additionally dropped main_runtime at acquisition and at both of
its refresh sites, so the refreshed client shadowed the stale one under
a divergent (provider, task, model, runtime) key.
Pass the original lookup model (which may be None) into the refresh as
a separate lookup_model argument used only to build the cache key,
while the resolved model is still stored as the entry's usable model
and returned to the caller. Thread task into both acquisition sites and
main_runtime into the async acquisition and both async refresh sites,
so sync and async compute the same cache key on acquire and on refresh.
The stale client is now overwritten in place instead of lingering under
an orphaned key, preserving the per-model cache keying introduced in
on the default config.
Add end-to-end regression tests that drive the real call_llm and
async_call_llm through the real client cache; the existing 401 tests
patch _get_cached_client wholesale and so cannot observe
acquire/refresh key divergence.
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
``create_openai_client`` was a hardcoded if-ladder: copilot-acp builds an ACP
stdio shim, gemini builds a native client, everything else gets an
``openai.OpenAI``. There was no extension point, so a provider whose wire
protocol is not OpenAI-over-HTTP could only be added by editing this function —
which is exactly why an ACP provider cannot ship outside this tree today, even
though ``providers/__init__.py`` has discovered out-of-tree profiles from
``~/.hermes/plugins/model-providers/`` and pip entry points for a while.
``ProviderProfile.create_client(**client_kwargs)`` closes that gap. It returns
``None`` by default, so every provider that wants the standard client is
unaffected and the existing ladder still runs as the fallback. copilot-acp is
migrated onto it — its hardcoded branch is gone and its profile supplies the
client in three lines, which is the same three lines an external package writes.
Resolution goes by provider name first, then by ``base_url`` prefix, so a
runtime configured only by URL still reaches its profile — matching what the
replaced ``startswith("acp://copilot")`` branch did. A profile that raises is
logged and skipped: a third-party plugin can fail to provide a client, but it
cannot take the turn down.
Also replaces the two ``isinstance`` checks in ``agent/auxiliary_client.py``
that mean "this client is complete, do not wrap it" with capability flags the
client class declares — ``HERMES_SKIP_TRANSPORT_WRAP`` and
``HERMES_SKIP_ASYNC_WRAP``, mirroring ``SUPPORTS_HERMES_TOOL_CALLS`` in
``background_review.py``. Two in-tree consumers (the ACP shim and the Gemini
native client), an out-of-tree client is covered by the same declaration, and
the hot path no longer imports those modules just to type-test.
Co-Authored-By: Junie <junie@jetbrains.com>
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.
- resolve_nous_runtime_credentials(stale_access_token=): under the store
lock, skip the refresh POST when the on-disk token differs from the one
that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
pass the failed bearer through, and treat a lock TimeoutError as
'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
refresh/0 unrecovered.
Use the ACP v1 session config contract advertised by session/new: locate the category=model option and apply the selected value through session/set_config_option. Retain session/set_model only as compatibility fallback for pre-configOptions agents. Reject unknown and policy-disabled values before prompting.
Verified against the installed Copilot ACP server: its model config option advertises the account-authorized choices, session/set_config_option returns the updated state, and live prompts route gpt-5.6-terra to Terra and claude-sonnet-5 to Sonnet 5.
Follow-up to the session/set_model wiring, caught in live use: picking an
org-policy-disabled model (claude-fable-5) produced a response claiming to
BE that model while Copilot actually served its default (Claude Sonnet 5).
Two causes:
1. The prompt preamble injected 'Hermes requested model hint: <id>', so
whatever model actually served the session parroted the requested name
back as its identity. Remove the line entirely — the model is applied
for real via session/set_model now, and identity must come from the
backend, not prompt suggestion.
2. session/new advertises policy-disabled ids alongside enabled ones
(_meta.copilotEnablement: 'disabled'); selecting one is accepted but
silently serves the default. Exclude disabled ids from the offered set
so the degrade-with-warning path handles them.
Verified live: requesting claude-fable-5 logs the does-not-offer warning
listing the 23 genuinely enabled models, serves the default, and the
response truthfully self-identifies as Claude Sonnet 5.
Selecting a model on the copilot-acp provider had no effect: the model id
never left Hermes. _create_chat_completion() dropped the model argument
before _run_prompt(), so the selection survived only as prompt text
('Hermes requested model hint: ...') and Copilot answered with its own
session default — a user picking gpt-5.6-terra visibly got Claude Sonnet 5.
Live-probing 'copilot --acp --stdio' shows the CLI validates but IGNORES
its --model spawn flag in ACP mode, while session/new advertises
models.availableModels and the ACP-native session/set_model call actually
switches the session. Wire that in: forward the model into _run_prompt,
and after session/new send session/set_model when the id is advertised
(or the server reports no list). Unknown ids degrade to the session
default with a warning instead of failing the turn; the provider-level
virtual slug 'copilot-acp' is never forwarded.
Verified live against the real CLI: requesting gpt-5.6-terra answers as
GPT-5.6 Terra and claude-sonnet-5 answers as Claude Sonnet 5.
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.
- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
otherwise, same helper the runner uses) and wrap the script / render /
skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
`SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
listed default's skills; the #88023 home-keyed cache alone could not fix
that. Use the call-time `_skills_dir()` there and at the two other
SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
root, so a routed profile's absolute skill_dir was rejected by
`skill_view` ("must be a relative path within the skills directory").
Resolve against `_skills_dir()` — the root `skill_view` itself enforces.
Fixes#67277
Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Under a multiplexed gateway every profile's outbound webhooks share one
delivery worker, so receivers could not tell which profile fired an
event. Add a top-level `profile` field to the payload, resolved at fire
time from the bound Hermes home via get_active_profile_name() ("default"
outside profiles). Documents the field in the wire-format section.
Reported by @vszgdcn8cj-ctrl.
Fixes#92674
re_register_config_hooks() cleared the entire process-global idempotence
set on every force-reload, so a profile-local plugin force-reload dropped
another live profile's ledger key without touching its still-registered
callback — the next registration call for that profile then appended a
duplicate. Scope the clear to the reloading profile's own home, and give
outbound webhooks the same force-reload restoration shell hooks already
had, since unload() wipes both from the shared _hooks dict.
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).
Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).
Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.
Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.
Fixes#100661Closes#97766 (overflow-force idea; the bundled continuation changes were not taken)
Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):
- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
requests silently run at standard speed and bill standard rates
(usage.speed: 'standard'). Today's allowlist therefore shows 4.6
users a fast toggle that does nothing, while denying it to the two
models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
select fast inference via the model field and are explicitly
excluded from the param gate.
Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.
## How to test
scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q
113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.
Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.
Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).
Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.
Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.
- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
workers by exact delegation_id (heuristic shape/time grouping kept for
older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
returned by the dispatch and used for cache/delegation/live/<id>/.