Native Windows run 34096838164 reports false success for absent and corrupt executables, missing bundle files, missing chunks and missing or stale stamps. Reuse the existing build identity and PE validators, and check interpreter presence before waiting for Desktop. Preserve dependency recovery and exit-2 refusal behavior.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>
Keep parsing, contracts, gates and persisted goal mutations in one dispatcher. Adapters retain authorization, rendering and scheduling; TUI drafting resolves the target session profile off the RPC reader. Document ACP as unsupported rather than implying a goal loop exists.
The identity-guarded token pop appeared twice (the runner's finally and the
new worker-start except); a future edit to one copy would silently reintroduce
the sticky-token class. Hoist into a local closure called from both sites, and
extend the _run_hook_callback_bounded docstring with the new skip reason.
Behavior-preserving follow-up on #104651.
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).
- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
resumed turn while the durable transcript still matches, cleared with the row on
compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).
Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.
- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
The alias check re-implemented most of _is_same_auth_store (#101356) inline, but
that helper answers False on a samefile() OSError — on the strip path that would
fail OPEN. Resolve identity positively (same resolved path, or samefile says so)
and treat any error as "refuse"; three lines instead of the inline stat dance.
The `preserve_symlinks=False` save path (raw os.replace) bypassed atomic_replace's
EXDEV/Windows fallback to defend against a path swap no concurrent writer performs
on a directory this process just created — removed with its mock-injected test.
Tests kept: shared-store invariant [symlink|hardlink] and two fail-closed variants.
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
those files carry no windows_only marker so they never run on Windows, and the monkeypatched
predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
db.close() already releases a registry-shared handle, so the described NameError leak only
existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
teardown; replacement not published before the last close settles) plus the raising-close
settlement; comments say the WHY once.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.
`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.
GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.
Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.
Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
hermes_cli.foreign_sessions._walk owns discovery (glob per source, per-entry
OSError tolerance, symlink-escape and regular-file checks, newest first); the
CLI listing and the desktop browser's candidate rows both consume it instead of
carrying two scans of the same directories.
_first_user_line indexed splitlines()[0] on a whitespace-only user message
(image-only / tool-only turn) and raised IndexError, which aborted the whole
CLI picker and the desktop session.foreign.list RPC. partition("\n") yields
"" for blank text and the loop moves on to the first real user line.
Reported first in #92290.
- Sidebar: the entry joins SIDEBAR_NAV (same chrome, active state, data-tour
handle as the other rows) instead of a one-off Button below the rail;
i18n moves to sidebar.nav['session-import'].
- Reuse common.retry / common.refresh / common.back instead of duplicating them
under sessionImport in six locales.
- hermes_cli/foreign_session_browser.py -> foreign_sessions_browser.py so it
sorts as a sibling of the foreign_sessions module it extends.
One unreadable or rotated log must not hide the rest of the list; discovery
tolerates OSError from resolve()/stat() per entry and reuses the stat result.
Browse the backend host's foreign CLI session logs, preview a bounded read-only
transcript, and continue a copy in Hermes under the selected profile. Reuses the
hermes_cli.foreign_sessions parsers and the portability validator/writer;
imports are transactional and deduplicated on the recorded origin.
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.
Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
Folds the /heartbeat driver into the existing /loop poll slot (one cadence, one
try/except table) and closes the consumed-tick gap: a dispatch that raises OR is
refused by _admit_prompt_turn (returns False) releases the session claim and rewinds
the persisted fire via HeartbeatManager.abandon_fire(), so the tick stays due for
the next poll instead of advancing fire_count with no turn. abandon_fire refuses to
overwrite a pause/resume/clear that landed between claim and dispatch (mirrors
LoopManager.abandon_tick). Tests use the real SessionDB under a temp HERMES_HOME and
drive the poller loop itself; docs note the TUI/Desktop surface.
Rollback-on-failure idea credited to jerrygooch (#104011, 51df1ed4d04 / 283fa972169);
the driver placement is Halldrix's (#102118). Reported in #102056 / #103044 (Vksh07).
Portal serves anthropic/* on two routes. The native /v1/messages wire, which
Hermes has used since 02d5e23085, re-writes the previous turn's prompt cache
on 14-20% of consecutive calls in concurrent tool loops; the chat/completions
route does not. Measured 2026-09-06, 20 concurrent sessions x 6 tool calls on
Fable 5.1, same account, same hour, same first-party pin:
Nous /v1/messages 180 pairs, 25 stuck (13.9%); 3 earlier 40-session
runs 15.0 / 17.3 / 20.2%; unchanged by the pin
Nous /v1/chat/completions 320 pairs (2 runs), 0 stuck
OpenRouter direct, pinned 161 pairs, 0 stuck
"stuck" = cache_read on call k+1 equals cache_read on call k instead of
call k's prompt size: the last write was not visible, the turn (~28K) was
written twice. Every stuck pair had byte-identical system, tools and
per-message shas, so it is not client-side mutation. At $10-20/M for
cache writes that is 15-20% of a fan-out's write bill.
The cause is inside the portal's native route (every element of its
outgoing request reproduces clean from outside; NousResearch/api#227
carries the diagnostics). Until it is fixed, anthropic/* rides
chat/completions. nous.anthropic_wire: native opts back in. Cost of
chat: prior-turn thinking travels as OpenAI-style reasoning fields, and
cache_control scopes are translated by the portal's adapter.
Tests: new test_nous_anthropic_wire_default.py reads the knob through the
real config loader in a temp HERMES_HOME and through
resolve_runtime_provider (both settings); the existing native-wire
contract suite selects native via an autouse fixture so the wire keeps
working for the flip-back. Live: a real AIAgent on this branch,
provider=nous + anthropic/claude-fable-5.1, dispatches to
chat/completions with the session_id intact.
- Generalize Gemini 3 thinking config model prefix match to gemini-3*
- Add pricing snapshot entries for gemini-3.7-flash and gemini-3.8-flash
- Add model identifiers to Google, OpenRouter, and Vertex CLI lists
- Add test coverage for gemini-3.8-flash thinkingLevel configuration
Trim the salvaged suite to the behaviour contracts: a picker exception or Ctrl+C leaves
config.yaml's model exactly as snapshotted (parametrized), and an absent active_provider
stays absent through snapshot+restore. The mock-dispatch tests (asserting calls into our own
helpers) are dropped. evals/cli_fallback_add_picker_error.py drives the real
`hermes fallback add` under a Linux PTY into an ordinary picker OSError and checks the
persisted model: stranded on base, restored after.
Follow-up on the cherry-picked fix: _defer_update_notice no longer takes the Rich console it
can never safely use (the notice always lands after patch_stdout owns stdout), the unit test
asserts the contract at prompt_toolkit's boundary (an ANSI fragment reaches
print_formatted_text, with no raw ESC/markup in the visible text) instead of patching our own
cprint, and evals/cli_deferred_notice.py drives the real CLI under a Linux PTY with a FIFO-gated
update cache to A/B the late-notice cases (garbled on base, clean after).
The deferred update notice ("N commits behind") was rendered via
console.print which writes ANSI escapes directly to stdout. Under
patch_stdout (active in the interactive CLI), the StdoutProxy strips
ESC bytes, leaving visible [1;33m...[0m artifacts.
Fix: render Rich markup to an ANSI string via a captured Console, then
route through cprint (which uses prompt_toolkit ANSI parser) to
bypass the StdoutProxy mangling.
This is the same class of bug as #2262 (fixed by #2448 for agent._print_fn
and display.py), but the _defer_update_notice callsite in banner.py was
not covered by that fix.
Fixes#83969
Parametrized refusal test (bare/fenced/uppercase-fenced truncated objects) asserts
profile.yaml is byte-identical afterwards; one prose test keeps the lenient
fallback contract. Replaces five near-duplicate tests from #104075.
Campaign tracker: https://github.com/NousResearch/hermes-agent/issues/104154
hermes_cli/profile_describer.py's aux-LLM response parser is documented as
"lenient, never raises": when the reply doesn't parse as JSON via
_extract_json_blob, it falls back to treating the WHOLE raw reply as plain
prose and persists it (truncated to 280 chars) as profile.yaml's
description.
That fallback exists for models that ignore the JSON-output instruction and
just answer in prose -- reasonable. But it doesn't distinguish that case
from a reply that DID start out as the requested JSON object and was cut
off mid-object by the aux model or transport. _extract_json_blob requires a
matching closing brace, so a truncated object (no `}` at all, or one cut
off mid-string) returns None just like plain prose does, and the fallback
then persists the raw JSON fragment verbatim: literal leading `{`, `\n`
escapes, a sentence chopped mid-word (#104067).
Fix: when parsing fails, check whether the reply (after the existing
code-fence strip) still looks JSON-shaped (starts with `{`). If so, it's a
malformed/truncated structured reply, not prose -- refuse it and return
ok=False instead of persisting the fragment. Only fall through to the
raw-text-as-prose fallback when the reply never looked like JSON to begin
with, so genuinely prose-only aux replies keep working exactly as before.
Also logs one INFO line on the refusal path, per the issue's note that the
silent fallback made this undiagnosable.
Update (review feedback from @jerrygooch): profile_describer._FENCE_RE was
case-sensitive, so an uppercase ```JSON fence wasn't stripped -- stripping
left "JSON\n{..." which doesn't start with "{", so a truncated response
under that fence variant still fell through to the prose fallback and got
persisted, defeating the fix for that case. Added re.IGNORECASE to
_FENCE_RE, aligning it with kanban_specify._FENCE_RE (which already used
re.IGNORECASE), and added a regression test for the uppercase-fence case.
Added tests/hermes_cli/test_profile_describer.py coverage: a truncated
object cut off mid-sentence (the real-world shape from the issue), one cut
off right after the opening brace, the same truncation under a ```json
fence, the same truncation under an uppercase ```JSON fence, and a
regression guard that a genuine plain-prose reply still hits the existing
lenient fallback unchanged.
Fixes#104067
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LUtGTop5iGKi5vfWnKnwET
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).