Preserve profile ownership proven at the secondary socket boundary rather
than trusting arbitrary wire profile fields. Retire transient local owners
with the profile pool, while keeping durable and exact remote ownership first.
Salvage #103774 with two invariant tests and routing documentation. The
profile-only fallback was also identified in the earlier #103770; this
version retains the producer provenance and retirement boundary.
Live Desktop renderer with two isolated serve backends: Reject previously
failed after clearing durable bindings, leaving approval pending. With this
change, the same action sends deny to the owning backend and pending clears.
Existing-binding controls pass on both sides. No vendor inference used.
Fixes#103755
Salvaged-from: 8a3c545e255b66b6ca4bc4b99cd725c5f7a08632
Keep compact completion labels while rendering only result and job output bodies, including legacy and batch deliveries. Credit Gyarados4157's #101083 investigation; avoid rendering its full model instruction envelope.
Remove whole-document whitespace cleanup that erased hard breaks and fenced-code spacing across all five ingestion paths. Limit generated-image gap joining to the removed image itself. Keep soft newlines as soft breaks.
Investigated #97117 and the approaches in #97431 (@wooyongbin3-cpu) and #97175 (@Jackal991); both leave fenced-code whitespace and sibling ingestion paths exposed, so remove the destructive normalization instead.
Salvage #99095 (e7ea53074ab2b64a1530641659399d9f1bb4b435), completing one-time migration for both boolean values and using the existing storage helpers. Hydration and local toggles never edit backend configuration. Fixes#99076.
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Wire the existing consent-aware, profile-keyed hook registrar at agent
construction, where the correct session home is already bound. This covers
serve and TUI agent construction without a startup-only registration or a
new helper that swallows registration failures.
Live isolated serve/WebSocket probes reproduce the missing registration on
base for both write_file and terminal, then verify each consented profile
blocks its configured tool while the unapproved profile still runs normally.
Repeated alpha construction does not duplicate hook callbacks.
Slim implementation of the agent-build placement proposed in #57020;
thanks also to the profile-scoped analysis in #102691.
Co-authored-by: grimmjoww578 <willies578@gmail.com>
The real closed-WebSocket resident variant still wedges after the missing-timer repair. Re-enter existing transport cleanup before rearming orphan timers, preserving viewer transfer and the reconnect/delegation fence rather than deleting registry rows from a stale snapshot. Extend the same invariant to both dead-transport shapes. Live class investigation informed by #104710; no direct-vouch reclaim machinery imported.
Salvage #104704 (e9423d2d0bbe3e795c5eaccb86a913f1d95444ba, 5fa2b98fc833fc4e1ee7f1aaa7eb45cdd3cd7cde). Reuse guarded orphan teardown instead of deleting ownership fences. Real two-backend WebSocket probe reproduces the missing-timer wedge on base and proves reconnect/delegation protection and recovery. Add reusable probe and user documentation.
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>
Keep parsing, contracts, gates and persisted goal mutations in one dispatcher. Adapters retain authorization, rendering and scheduling; TUI drafting resolves the target session profile off the RPC reader. Document ACP as unsupported rather than implying a goal loop exists.
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.
`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.
GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.
Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.
Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
Folds the /heartbeat driver into the existing /loop poll slot (one cadence, one
try/except table) and closes the consumed-tick gap: a dispatch that raises OR is
refused by _admit_prompt_turn (returns False) releases the session claim and rewinds
the persisted fire via HeartbeatManager.abandon_fire(), so the tick stays due for
the next poll instead of advancing fire_count with no turn. abandon_fire refuses to
overwrite a pause/resume/clear that landed between claim and dispatch (mirrors
LoopManager.abandon_tick). Tests use the real SessionDB under a temp HERMES_HOME and
drive the poller loop itself; docs note the TUI/Desktop surface.
Rollback-on-failure idea credited to jerrygooch (#104011, 51df1ed4d04 / 283fa972169);
the driver placement is Halldrix's (#102118). Reported in #102056 / #103044 (Vksh07).
A background delegate_task call used to be ONE async unit: the runner joined on
every child and a single consolidated message re-entered the conversation when
the SLOWEST finished. Fifteen independent PR reviews therefore waited on the
fifteenth before the parent could act on the first.
Each task now carries an optional `group`. `_units_of` partitions the call's
children into units — one per distinct group, one per ungrouped task — and each
unit is dispatched to the async registry on its own, so its results re-enter the
conversation as soon as THAT unit is done. Tasks that must be compared or merged
share a group and still return together.
Capacity is unchanged: every unit of one call joins the first unit's pool slot
(`slot_key` in `async_delegation._dispatch`), so splitting never consumes more of
`delegation.max_concurrent_children` than the call did. Unit ids suffix the
call's id (`deleg_xxxx-1`, `-2`, …) so live transcripts stay under one dir; the
completion block names the group and notes that sibling units report separately;
`active_task_count` counts a unit's own tasks.
Portal serves anthropic/* on two routes. The native /v1/messages wire, which
Hermes has used since 02d5e23085, re-writes the previous turn's prompt cache
on 14-20% of consecutive calls in concurrent tool loops; the chat/completions
route does not. Measured 2026-09-06, 20 concurrent sessions x 6 tool calls on
Fable 5.1, same account, same hour, same first-party pin:
Nous /v1/messages 180 pairs, 25 stuck (13.9%); 3 earlier 40-session
runs 15.0 / 17.3 / 20.2%; unchanged by the pin
Nous /v1/chat/completions 320 pairs (2 runs), 0 stuck
OpenRouter direct, pinned 161 pairs, 0 stuck
"stuck" = cache_read on call k+1 equals cache_read on call k instead of
call k's prompt size: the last write was not visible, the turn (~28K) was
written twice. Every stuck pair had byte-identical system, tools and
per-message shas, so it is not client-side mutation. At $10-20/M for
cache writes that is 15-20% of a fan-out's write bill.
The cause is inside the portal's native route (every element of its
outgoing request reproduces clean from outside; NousResearch/api#227
carries the diagnostics). Until it is fixed, anthropic/* rides
chat/completions. nous.anthropic_wire: native opts back in. Cost of
chat: prior-turn thinking travels as OpenAI-style reasoning fields, and
cache_control scopes are translated by the portal's adapter.
Tests: new test_nous_anthropic_wire_default.py reads the knob through the
real config loader in a temp HERMES_HOME and through
resolve_runtime_provider (both settings); the existing native-wire
contract suite selects native via an autouse fixture so the wire keeps
working for the flip-back. Live: a real AIAgent on this branch,
provider=nous + anthropic/claude-fable-5.1, dispatches to
chat/completions with the session_id intact.
Two corrections on top of the mention-preservation pick:
- Stripping our own trigger handle from every group message was replaced by never stripping
it in groups, which broke the pre-existing contract that `@bot 2` answers a pending clarify
prompt and `@bot ok` approves (the intercepts match the exact stripped text). Strip only when
no other participant is named; when other bots/users are mentioned the text is kept verbatim
so `@research_bot , @ops_bot …` no longer arrives as `, @ops_bot …`.
- The addressing block carried a per-message fact ("explicitly mentions you: yes/no") inside
`channel_prompt`, which is part of the cached-agent signature: a mention turn followed by a
reply turn rebuilt the AIAgent (prompt-cache miss) every time the shape flipped. Keep only
the username line, byte-identical for the life of the session.
Tests rewritten to pin both contracts (sole-addressee stripping; identical channel_prompt and
agent signature across a mention turn and a reply turn).
Portal decides routing centrally per model and returns HTTP 400 on any
caller-supplied `provider` object (live-confirmed). Drop the five "OpenRouter or
Nous Portal" claims from the provider-routing page and say plainly that the
setting is ignored on Portal now that the profile no longer forwards it.
`provider_routing.models.<model-id>` now takes the same only/ignore/order/sort/
require_parameters/data_collection keys and overlays the flat provider_routing
values whenever the agent is on that model. Resolution lives in the one
chokepoint every request path already uses (_provider_preferences_for_agent),
so CLI, gateway, TUI/Desktop, cron, /model switches, fallback activation and
delegated children on another model all honour it with no per-surface plumbing.
Matching is spelling-tolerant, sharing _canonical_model_variants with
agent.reasoning_overrides.
The OpenRouter profile's speed-tier pin no longer overwrites an explicit user
`only` on the BASE gpt-6-astra slug: the pin exists to keep default routing off
flex/fast, and a user pin is the stronger intent (only: [openai] stays [openai]
instead of becoming [openai, azure, azure/us]). Tier slugs (-fast/-flex) keep
owning `only`.
Live A/B (config only: {gpt-6-astra: [openai], claude-fable-5.1: [anthropic]}):
main sent {"sort":"price"} for fable and OpenRouter served it from Azure; with
this change it sends {"only":["anthropic"],"sort":"price"} and Anthropic serves it.
Schema proposed in #24495 (samplesabotage) and #100711 (Artemonim); this is a
slim chokepoint implementation of that design.
Co-authored-by: samplesabotage <samplesabotage@users.noreply.github.com>
The default backend needs no Reddit account, login, cookie or API key; the
optional upgrade is a free 'script' app registration using the app-only
client_credentials grant, never a user login. Said in the skill's
Prerequisites (with a comparison table), Procedure, Pitfalls, the runtime
doctor/thread notes, and a delimited .env.example block.
Two zero-install research skills plus routing guidance, ported as ideas
(not code) from the Agent Reach skill's per-platform backend routing.
reddit-reading (skills/social-media): subreddit listings, site/subreddit
search, threads with comments, user pages. Live-verified from a datacentre
IP: www .json, api.reddit.com, old.reddit, r.jina.ai and the browser tool
all return 403 / an empty shell / a humanity check; the Atom .rss endpoints
are the only anonymous path and are throttled to ~1 request/min/IP. The
script waits out the x-ratelimit-reset window once and retries, and
switches to the OAuth API (scores, nested comments, ~100 req/min) when
REDDIT_CLIENT_ID/SECRET are present. `doctor` reports the active backend.
rss-feeds (skills/research): RSS 2.0 / RSS 1.0 / Atom / JSON Feed parsing
with UTC-normalised dates, --since/--limit, and feed discovery behind a
page URL (<link rel=alternate>, then well-known paths). Stdlib only; the
optional blogwatcher skill remains the stateful many-feed reader and now
points at this one for one-off reads.
grounded-citations gains a "Multi-Platform Sweeps" section routing
"what are people saying about X" tasks across web, Reddit, feeds, video,
code and X with per-platform attribution and coverage-gap reporting;
competitor-news-monitor references the new sources.
Tests: two invariant tests per skill (format normalisation + discovery;
anonymous 429 handling + OAuth routing/flattening), no network. The
reddit test caught a real bug: the feed footer's "/u/author" leaked into
post bodies.
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.
redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.
Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
Independent review: a YAML `true` coerced to int 1 and gave every child a
one-token compression trigger; "200k" silently disabled the default cap.
Values are now validated: an int >= 16000 is used, 0/false/null disable on
purpose, anything else warns and falls back to the 200K default so a typo
never costs money. Config comment and docs say it caps the compaction
trigger, not a hard request-size limit.
Test: true / "200k" / 5 -> default; 0 / false / None -> disabled; valid ints
pass.
A delegate_task child inherits the compression threshold as a RATIO of the
model window. On a 1M-window model at the run's configured 0.85 that is an
850K-token trigger: in the 1,393-agent refactor run 1,373 of 1,375 children
never compressed once, 62% of all API calls carried >150K of context, and the
calls above 200K carried ~$10.9k of the $19.3k bill (58% of it cache WRITES,
i.e. re-sending a 300-800K prefix on every call). A sawtooth replay of the
logged calls with a 200K cap / 65K floor cuts context spend by ~49% (~$7k).
Children are brief-driven and disposable; they re-read their brief and the
files they touch, so a large window buys them little. New
delegation.compression_threshold_tokens (default 200000) is applied to the
child's ContextCompressor right after construction as the lower of it and
any global compression.threshold_tokens; the parent's own trigger is
untouched. 0 disables the subagent-specific cap. The compressor applies
threshold_tokens_cap on first window resolution, so this is byte-equivalent
to the user having set compression.threshold_tokens for the child.
Live through the real spawn path (_build_child_agent, real imports, temp
HERMES_HOME, 1M-window model): main child trigger 500,000 / branch 200,000;
parent 500,000 on both.
Tests (3): default caps a 1M child at 200K; the cap is the lower of the
delegation and global values and never raises a small-window child's
trigger; 0 disables and an already-resolved trigger is re-clamped.
Docs: delegation.md, configuration.md.