Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).
Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.
`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.
Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.
Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
The #58262 assertion lived in test_scheduler.py against a harness that has
since moved; re-home it in the delivery-confirmation suite alongside the
positive-evidence tests, and widen it to the forum-topic route and the media
route so the marker cannot drift out of any lane.
The filter guards against bot-to-bot mirror loops of model chatter. Cron
output is an artifact: a job whose brief is legitimately terse ("...", a
single emoji from a script) has no loop partner, and dropping it while
returning {"success": True} is how a cron was logged as delivered with
nothing on the wire (#77763). Cron sends carry job_id in metadata; every
other caller keeps the filter unchanged.
A cron job fired, the scheduler logged "delivered to telegram:<chat> via
live adapter", and nothing reached Telegram (#77763). The log line was not
evidence of a send:
* the silence-narration filter returns {"success": True, "delivered": False}
(a successful *drop*), and the dict-normalization branch read only
"success", so a filtered message counted as delivered;
* an empty payload (no text, no media) skipped the send entirely and still
fell into the "delivered" branch;
* the log line named the chat but not the lane, so a wrong-thread delivery
and a phantom one are indistinguishable after the fact.
_confirm_adapter_delivery now inspects both result shapes: an explicit
`delivered: False` is a rejection even with a truthy `success`, and a
success with no message_id and no raw_response is accepted but logged as
UNVERIFIED. The empty-payload case fails closed into the existing
standalone/warn handling, and the delivered log carries thread= and
message_id=.
Failing closed on the live lane is only half the fix on a native target:
the standalone fallback sent the same empty payload, and the Telegram
adapter returns SendResult(success=True) for empty content without an API
call — a phantom live delivery became a phantom standalone one. Both
_send_to_platform call sites now sit behind one skip guard, so "empty
payload fails closed" holds on every lane (#77763).
The one-shot reasoning-off retry changes a request parameter that is part
of the provider cache key on config-sensitive providers (Anthropic renders
thinking/effort into the prompt; OpenAI lists reasoning.effort as
prefix-affecting), so that request is a deliberate single cache miss.
Pin the bound: the request AFTER it must carry the configured reasoning
again and the system prompt must be byte-identical across the whole retry
sequence. Sabotage-verified (sticky flag -> test fails on request 3).
Docstring on _consume_ephemeral_reasoning_off states the cost honestly.
Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
chat_completions path always re-emitted extra_body.reasoning with
enabled=True, so both reasoning_effort: none and the one-shot
length-continuation override went out as {enabled: true, effort: none}.
Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
turn start so a flag armed by an interrupted/errored turn can never
strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
/thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
reasoning.enabled=false) and a stale-flag turn-scope test.
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).
The length-continuation path handled that shape badly:
1. the empty response was appended as an interim assistant fragment,
poisoning the transcript until the pre-call sanitizer healed it
(observed 3+ healings per turn on the reporting user's session);
2. every continuation re-ran with thinking ON, re-deriving the whole
thinking budget against a growing context, so 4 attempts still produced
nothing and the turn died with 'Response remains truncated after 4
continuation attempts'.
Now:
- interim assistant fragments with no visible content are never appended
(whichever way they got empty);
- a thinking-only truncation sets a one-shot reasoning-off override that
build_api_kwargs consumes for the next request, so the continuation
writes the answer instead of re-thinking it;
- the ceiling exit clears a pending override and, when every fragment was
empty, returns an actionable final_response instead of an invisible None.
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):
- web dashboard CronPage: last_status was never rendered at all — a
delivery_failed job showed a green 'scheduled' badge and only a small red
'delivery: ...' line. New pure cronLastResult() helper maps the closed
literal set to tones (ok=success, delivery_failed/blocked_config=warning,
error/unknown=destructive) and the card now shows an amber
'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
literal; routineLastResult() spells out each one ('Ran, but delivery
failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
on this branch; no consumer compared == 'ok' for success apart from the
cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
detail field carries the reason.
Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.
Co-authored-by: webtecnica <webtecnica@gmail.com>
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.
Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.
Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.
Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.
Fixes#83993
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).
- windows.ps1: readiness handshake after BeginInvoke — one /progress
round-trip must succeed (≤15s) before the server is returned; on failure
tear the listener down and continue without UI. The URL now means
"serving", not "bound". Also fixes the browser opening to a page that never
loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
spawn the real PowerShell script; the Windows-only job now runs them only
when scripts/desktop-update/**, the Electron updater launcher, conftest,
pyproject, or those tests change (push/dispatch fail open). A PR that
never touched that surface cannot be failed by its process timing.
Widen the two salvaged fixes (#100490, #100493) to the whole class:
- match_runtime_outcomes: serve/dashboard rows never borrow gateway
bookkeeping at ANY site — not just the bare hermes-gateway unit name
(#100490) but also relaunched_profiles / externally_supervised_profiles
and the profile-substring unit match (hermes-gateway-work credited the
'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
units (exact names, scope prefix tolerated) or, when the caller passes
the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
feed the Phase-2 reconciliation, so a surviving unmanaged serve is
'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
remedy instead of 'hermes gateway restart', which cannot reach it.
Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
match_runtime_outcomes() treats any default-profile runtime as covered
once the bare "hermes-gateway" unit restarts, regardless of the
runtime's own kind. An sshd-spawned `serve --isolated` backend (no
systemd unit, supervisor "manual-serve") shares the default profile
and gets silently marked "restarted" even though its own PID was never
touched — so the #91277 Phase 2 unaccounted-runtime tripwire never
fires for it and `hermes update` reports success while it keeps
running pre-update code (#100479).
Restrict the "hermes-gateway" special case to kind == "gateway" so a
serve/dashboard runtime under the same profile falls through to
"unaccounted" instead of borrowing the gateway's outcome.
Regression for the single-connection/single-profile report: hasRegistryTopology()
is true on every modern Desktop, so the ambient escape hatch stays closed; the
approval.request event's own (connectionId, profile) stamp is what routes
approval.respond back to the primary socket.
recordSessionEventScope already captures the exact (connectionId, profile) a
runtime's inbound events proved, but knownOwnerForSession never consulted it:
with no tile/hint/row binding for the runtime id, approval.respond failed
owner resolution (SessionOwnerResolutionError) even though the event source
itself named the owner.
Add a structured owner twin of the scope ledger, written and cleared with it,
consumed as the LAST rung of knownOwnerForSession so durable stored identity
still outranks it and untagged/unknown runtimes keep failing closed.
Sibling site of the same loss class. The #72680 shutdown flush only
serialised the adapter slot (_pending_messages); the FIFO tail parked in
SessionState.conversation.queued_events was discarded with the process,
so every follow-up queued behind the head at restart time vanished the
same way the idle-orphan did. flush_overflow_to_file writes one payload
per overflow event in the slot-flush shape (plus seq for arrival order),
so the existing recover_pending_to_db startup replay inserts them with no
new reader. Wired into _stop_impl beside the slot flush.
Follow-up to the salvaged #99912 rescue. The original helper left the
rescued orphan IN the adapter slot while the caller also swapped it in as
the current turn, so the post-turn _dequeue_pending_event ran the same
follow-up a second time (live repro: TURNS=['Sent','C','C','D']). The
helper now pops the oldest orphan and returns it to run as this turn,
stages the NEXT orphan in the slot so the drain continues the chain in
arrival order, and the call site parks the incoming message behind the
chain via _enqueue_fifo (slot when free, overflow otherwise) instead of
always appending to overflow. The rescued event's own source drives the
turn so reply anchors point at the message actually being answered.
Tests: contract updated for the new return type; added the 2-orphan chain
case and the single-orphan-then-new-message slot case (both fail against
the original helper shape).
Review note on #99912: rescued = 1 followed by if rescued: is a constant
conditional — the log block runs unconditionally now that staging is
single-orphan by design.
When a follow-up is demoted to /queue during compression-in-flight,
it lands in SessionState.conversation.queued_events (overflow) with
the slot event in adapter._pending_messages. After the slot's turn
completes, _promote_queued_event should move the overflow head into
the slot for the recursive drain. When that drain never runs — the
#99882 shape: busy window ended through an exit that skipped the
promotion site — the overflow is silently orphaned: never dispatched,
never persisted, never logged. A 170-char Telegram follow-up vanished
without a trace; its re-send also vanished for the same reason.
Fix: _rescue_orphaned_overflow stages one orphan into the empty slot
on the next idle arrival, and the new message is enqueued behind it
so FIFO order (#28503) holds — oldest orphan runs as this turn, the
rest drain in order, the new message last. The helper is best-effort
(slot occupied or no overflow → no-op) and logs at WARNING when it
fires so a future drain regression is visible.
Tests (tests/gateway/test_fifo_overflow_rescue.py, 4 cases on the real
GatewayRunner FIFO):
- moves overflow head to empty slot
- no-op when slot occupied
- no-op when no overflow
- FIFO preserved: orphan-1, orphan-2, new-msg in exact arrival order
Existing queue suites pass unchanged (test_queue_consumption — 5 passed).
Fixes#99882
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:
- Edit -> re-run is progress. A successful mutating call (write_file/patch,
a green terminal/execute_code, browser actions, job/message/cron/memory/
skill mutations) marks progress for every failing signature still being
counted this turn; the next identical retry restarts its streak instead
of accumulating toward exact_failure_block_after. A pure replay never
mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
(terminal, execute_code, process pollers, browser_navigate, web_extract)
same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
task loops with a live parent/client and do real edit -> re-run work.
Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
this commit: COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.
Add regression coverage for repeated skill_view results under hard-stop guardrails.
BasePlatformAdapter._acquire_platform_lock emits `{scope}_lock` with
retryable=True on purpose (#54167): a MID-RUN reconnect must be able to
recover once the live holder exits or a stale record is cleared. The
startup router keyed solely off that flag, so a live foreign holder of the
bot token at zero-connected startup landed in `_failed_platforms` with
gateway_state=running — alive, deaf, and retry-storming the token every
backoff — instead of the exit-78 (EX_CONFIG / startup_failed) contract
that #51228 established for single-writer conflicts.
Minimal class fix, salvaged from #83183 (@alexgunsberg) against current
main:
- gateway/restart.py: `is_global_startup_conflict(error_code)` — matches
the `*_lock` / `lock_conflict` code families every adapter emits for
scoped-lock and identity conflicts. Code only, never message text.
- gateway/run.py primary startup routing: a lock-conflict failure is
routed as non-retryable (parked `fatal`, not queued). Nothing else
connected → exit 78; alongside a transient peer → NS-609 mixed mode,
gateway stays alive and only the peer retries.
- gateway/run.py `_schedule_secondary_profile_startup_reconnect`: the same
contract for multiplex secondaries — park `<profile>:<platform>` fatal
like `duplicate_credential` instead of scheduling a reconnect storm.
- Mid-run behavior is untouched: `_handle_adapter_fatal_error_impl` and
the reconnect watcher still treat `*_lock` as retryable (#54167).
Not carried over from #83183 (superseded on main or out of scope): the
`degraded` lifecycle write only fires on the all-retryable path and the
runner immediately overwrites it with `running` (so busy/drain already
see `running`); the secondary retry bridge landed separately in
96489f3c1b (#92064); Buzz/IRC/LINE lock-tuple unpack and the reconnect
ownership registry are separate class fixes.
Live repro (real GatewayRunner.start(), isolated HERMES_HOME + lock dir,
live holder subprocess owning the lock via production
acquire_scoped_lock): before — exit_code=None, gateway_state=running,
telegram `retrying`, queued in _failed_platforms; after — exit_code=78,
gateway_state=startup_failed, telegram `fatal`, _failed_platforms={}.
Co-authored-by: alexgunsberg <alex@gunsberg.fi>
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).
Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
Refs #86443
assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.
Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.
Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.
The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.
The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
All seven TestLoopTickWitness cases that need real UNIX-domain sockets
(socket.AF_UNIX socket nodes or asyncio.start_unix_server producers)
fail on native Windows, where neither primitive exists. Mark exactly
those cases with a shared skipif so a Windows run reports SKIPPED
instead of erroring, while the platform-independent witness-absent
contracts (mocked probes, file-only heartbeats) keep running there.
Split the legacy two-witness-contract test in two: its stale-file arm
is file-only and keeps running on Windows; its dead-listener-node arm
needs a real socket node and is skipped with the rest.
Follow-up to @zoser69's #78111 cherry-pick:
- lift the redirect into _redirect_platform_display_key() and apply it
BEFORE _validate_config_key / type coercion, so the unknown-key hint and
the string-vs-bool coercion both see the canonical path
- widen to the sibling surfaces: config get resolves the canonical key
(previously echoed the dead top-level value — the misleading half of the
report) and config unset removes the canonical leaf
- regression tests: get mirrors gateway resolve_display_setting, unset
removes the redirected leaf, note printed, helper touches ONLY
OVERRIDEABLE_KEYS (connection keys / 4-segment / already-canonical
paths untouched)
- docs: configuration.md per-platform section names the canonical CLI
path and the accepted shorthand
Problem A of #71047: 'hermes config set platforms.telegram.streaming false'
wrote to a key the gateway never reads. The connection config
(gateway/config.py) reads only token/extra/overrides from the top-level
platforms.<name> block, while per-platform display settings (streaming,
show_reasoning, tool_progress, ...) are resolved from
display.platforms.<name>.<setting> (gateway/display_config.py).
Redirect a platforms.<name>.<setting> key to
display.platforms.<name>.<setting> only when <setting> is a known per-platform
display setting (gateway.display_config.OVERRIDEABLE_KEYS), leaving real
connection keys (token, extra, channel_overrides, ...) untouched. The
gateway.display_config import is lazy/try-guarded to avoid a circular import
and to keep the CLI working where gateway is not importable.
Adds tests/hermes_cli/test_config_set_platforms_redirect.py covering the
redirect, connection-key non-redirect, and the no-stray-top-level-platforms
case.
Follow-up for salvaged #88472 (+ #100055 / #99995 / #88998 intent):
- add_contributor.py compares filenames with str.casefold(), the same key
scripts/check-case-collisions.py uses repo-wide, so non-ASCII folds
(ß ~ ss) are caught the way macOS/Windows fold them.
- The KNOWN_CASE_CONFLICTS allowlist is gone: the historical
agent@Agents-Mac-mini.local pair was removed on main (fcdae2cf0b), so the
repo-wide test asserts zero collisions.
- New tests: same-login different-spelling is still refused (the filename
pair is the problem, not the login) and the exact spelling stays
idempotent; casefold vs lower coverage.
Co-authored-by: alfred-amanda <288490622+alfred-amanda@users.noreply.github.com>
contributors/emails/ uses the email as the FILENAME, so two mappings differing
only in case are the same file on Windows and on default macOS. The tree has
such a pair today:
contributors/emails/agent@Agents-Mac-mini.local -> skip-agent
contributors/emails/agent@agents-Mac-mini.local -> momomojo
git writes one and then reports the other as modified in a FRESH clone, forever.
The repo cannot be checked out clean on those platforms, which breaks any tool
that gates on a clean tree -- our own Windows Desktop rebuild refuses with
"fresh clone is NOT clean" and never gets to build.
add_contributor() now refuses a mapping that case-collides with an existing one,
for the same reason it already refuses a conflicting login: the tool exists so a
typo cannot silently reassign commits, and a collision does exactly that on half
the platforms it lands on.
Two tests: the guard, and a directory-wide check that no NEW collision appears.
The existing pair is pinned in KNOWN_CASE_CONFLICTS rather than resolved here --
the two files name DIFFERENT logins, so picking one reassigns a contributor
commit history, and that is a maintainer call. Please resolve it; the pin keeps
the breakage visible and stops it spreading meanwhile.
Verified: pytest tests/scripts/test_contributor_map.py -- 9 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two guards for the reconnect-queue eviction fix:
- a complete Matrix password config (homeserver + user_id + password)
counts as credentialed and stays retryable;
- every incomplete variant is still dropped, which is what keeps the
#64674 empty-primary multiplex eviction intact.
The incomplete cases pin the "read extra, not the environment" property,
and they set a fully-populated MATRIX_* environment explicitly to do it.
That setenv is load-bearing: tests/conftest.py sandboxes HERMES_HOME to a
tempdir and scrubs MATRIX_* from the environment, so production .env is
never loaded under pytest. Without the explicit setenv these cases pass
against an os.getenv-reading implementation and guard nothing.
Verified both directions against throwaway worktrees, live checkout
untouched:
- pre-fix implementation: the positive case fails (1 failed, 5 passed);
- os.getenv-fallback implementation: 4 of the 5 incomplete cases fail.
The "blank" case cannot discriminate by construction -- whitespace is
truthy, so `extra.get(k) or os.getenv(...)` never consults the
environment -- it guards strip()-emptiness instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_platform_has_bot_credential() decides whether a failed platform may be
retried. It only inspected PlatformConfig.token / .api_key, but Matrix
supports password login (MATRIX_USER_ID + MATRIX_PASSWORD, no
MATRIX_ACCESS_TOKEN), and build_config() puts those on extra{} rather
than .token.
So a password-auth Matrix config read as credential-less, and the
reconnect watcher deleted it from the retry queue on the first transient
failure. A momentary DNS failure at boot therefore took Matrix down
permanently: the homeserver was healthy, but nothing ever retried and
recovery required a manual gateway restart. Observed live as a ~13h
outage after a boot-time "Temporary failure in name resolution".
Mirror the adapter's own gate (homeserver + user_id + password).
Read ONLY from extra, never os.getenv: build_config() already copies all
three env vars onto extra, and importing this module loads ~/.hermes/.env,
so an env fallback would report "has credential" for every Matrix config
on the host -- including the empty-primary multiplex case (#64674) that
this check exists to evict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the #100829 salvage. uv pip install writes no __pycache__ by
default (pip does), so --compile-bytecode covers the whole install including
transitive deps, which the per-spec warm never sees. Also skip *.dist-info /
*.egg-info roots in _installed_dist_roots — they own no importable code.
Live: fresh cpython-3.12.13 venv, real uv install of anthropic==0.87.0 via
_venv_pip_install: main -> 0 pyc, first import 0.468s; after -> 1212 pyc
(546 anthropic), first import 0.205s.
Refs #100461
A pip/uv install writes .py sources and no __pycache__ — and reinstalling
the same version still deletes the cache the previous copy had. Nothing in
Hermes compiles them, so the whole compile is paid by whoever imports the
package next. For a lazily installed backend that is the foreground of a
user request, with nothing printed while it runs.
Measured for anthropic==0.87.0 (541 modules) on cpython-3.12.13: the first
import after an install costs 2.2-2.7s against 0.7-1.0s warm, and 10.5s
under concurrent load. N per-profile daemons cold-starting together each
pay it in full, because none of them has written the cache yet.
Compile the freshly installed distributions in _venv_pip_install instead,
on the success path of both the uv and pip tiers. The caller is already
waiting on an installer there and can see why. Package directories are
resolved from each distribution's own file list, so specs whose import
name differs from their package name (python-telegram-bot -> telegram)
are covered. Best-effort: a compile failure never invalidates an install
that succeeded, and sys.dont_write_bytecode is honored.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MhAnkrFktFdmZwf64fUYLE