Salvage #101887 after native Actions 34097643131 reproduced WinError 32 using a ready Electron app with its cwd inside the live release. Reuse the existing install-scoped process cleanup before promotion and wait after forced termination, preserving rollback.
Co-authored-by: fangliquan <fangliquan@qq.com>
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:
1. delegation.independent_completions (new, default false). #104299 made every
ungrouped task its own completion message, so a 15-task call woke the
orchestrator up to 15 times; one chain received 132 notices and answered
130 of them with "already incorporated". A multi-task call now returns as
ONE consolidated message unless the flag is on; `group` is inert until then.
2. Queued units were killed before they started. Units of one call share a
pool slot but the executor was still sized by slots, so with 15 units live
a new unit queued behind a full pool; the stale monitor's clock ran from
dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
when its thread finally came up (13 such lanes in one session). The
executor now grows to the number of live units and the stall clock arms
when the runner actually starts.
3. The tool text said "do not wait or poll — just continue" without saying
that completions are delivered only BETWEEN turns. A model that never ends
its turn (one 203-minute turn, 717 API calls) never received 40 finished
results. Tool description, dispatch note and completion header now say to
finish independent work, give a one-line status, and end the turn.
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Slim adaptation of #83772 to the current schema and failure-hint table.
Generated helpers are module exports on every execution path, not globals.
Correct schema, recovery hints and CLI tip rather than injecting names or
changing the execution boundary. Two registry-driven invariants reproduce
both misleading instructions on main and execute the corrected guidance.
Additional tool fix discovered during campaign #104904.
Original diagnosis and correction: @yuzilongleif-collab (#83772).
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Preserve producer descriptions without identity wrappers and size the existing themed tooltip to the viewport. Replace skipped and structural tests with two behavioral invariants. Native Electron before/after hover, click and keyboard verification passed; campaign suite validation remains queued.
Slim adaptation of anombyte93/hermes-agent@d06d2a49c5; use the canonical board resolver instead of inferring the slug from a path. Live isolated CLI probe confirms current-file, env and explicit board banners; event delivery remains live.
Co-authored-by: Hayden (Atlas agents) <212644172+anombyte93@users.noreply.github.com>
check_for_skill_updates() fetched every lock-file entry remotely, even
when the entry's install directory no longer existed, and each fetch had
no wall-clock bound — a few dead sources turned a routine
`hermes skills update` into a multi-minute stall (#104291).
- Entries whose recorded install_path resolves but does not exist are
reported as "orphaned" and skipped without a remote fetch;
unresolvable paths keep the previous fetch behavior.
- Each fetch now runs under a daemon helper thread with a hard timeout
(default 30 s) and degrades to "unavailable" when abandoned.
- `hermes skills check` prints a removal hint for orphaned entries.
Fixes#104291
Retain partial-line output, UTF-8 decoding, failure output and cancellation cleanup. Based on streaming investigations by Artemonim (#101850) and lEWFkRAD (#104843); gateway tee adapted from fangliquanflq (#97402). Live Linux child/tee probe: withheld or dropped on base, visible in 0.02 seconds after. Campaign-locked tests and native Windows proof are pending.
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
The desktop app always sends profile=default on GET /api/messaging/platforms.
_is_current_profile() recognized only None/""/"current" as the dashboard's
own profile — NOT the string "default" — so a single-profile install (the
standard `hermes gateway setup` flow: token in .env, no platforms: section in
config.yaml) entered the profile-scoped branch of _config_profile_scope().
That branch derives platform enablement from config.yaml only and never calls
load_gateway_config()'s env-override pass (which enables the platform when the
token is in the environment). Result: a platform connected via .env reported
enabled=false, state="disabled" while it was actually running. The unscoped
GET (no profile param) correctly reported enabled=true, state="connected".
Fix: classify by resolved path, not by string. After _is_current_profile()
fails, _config_profile_scope() now resolves the requested profile dir and
compares it against get_process_hermes_home().resolve() — the same comparison
_is_other_profile() already uses. When they match (profile=default on a
default-home process), yield None (no override), taking the unscoped path that
calls load_gateway_config(). A named-profile process (`-p worker`) has a
different HERMES_HOME, so its profile=default resolves to a different directory
and still scopes correctly — cross-profile secret isolation is preserved.
The scoped branch's config.yaml-only enablement is DELIBERATE:
load_gateway_config()'s env pass reads os.environ and would leak the root
install's tokens into a genuinely different profile's state. This fix only
reclassifies requests that name the process's OWN home; it does not touch the
scoped branch or gateway/config_env.py.
Refs #104614
all() over an empty tuple evaluates True, so the scoped credential
fallback reported platforms with required_env == () (whatsapp, yuanbao,
api_server, webhook, a2a, msgraph_webhook, relay, whatsapp_cloud) as
enabled=True with no config entry and no credentials. Add the
bool(required) guard to the enabled computation (per review suggestion)
and to the configured field, which came from the same all()-over-empty
expression and reported configured=True for the same shape — the
unscoped branch reports enabled=False / configured=False there, so the
scoped branch now agrees.
Adds tests/hermes_cli/test_web_server_scoped_enablement.py covering the
empty-required_env shapes, the explicit-enabled precedence, and the
credentials-present path.
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
The scoped branch of _platform_enablement consulted only config.yaml's
platforms: section, but the `hermes gateway setup` wizard writes .env
credentials and never a platforms: entry. The desktop always sends
?profile=default (normalizeProfileKey maps the primary profile to
`default`), so the Settings - Messaging page showed a working bot as
"Disabled" while /api/status reported it connected (#104614).
Mirror _enable_from_env (gateway/config_env.py): env credentials alone
enable a platform, an explicit enabled: false still wins. Only the
profile's own .env (env_on_disk) is consulted, so the root install's
os.environ credentials still never leak into a profile's state.
Fixes#104614
Native Windows run 34096838164 reports false success for absent and corrupt executables, missing bundle files, missing chunks and missing or stale stamps. Reuse the existing build identity and PE validators, and check interpreter presence before waiting for Desktop. Preserve dependency recovery and exit-2 refusal behavior.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>
Keep parsing, contracts, gates and persisted goal mutations in one dispatcher. Adapters retain authorization, rendering and scheduling; TUI drafting resolves the target session profile off the RPC reader. Document ACP as unsupported rather than implying a goal loop exists.
The identity-guarded token pop appeared twice (the runner's finally and the
new worker-start except); a future edit to one copy would silently reintroduce
the sticky-token class. Hoist into a local closure called from both sites, and
extend the _run_hook_callback_bounded docstring with the new skip reason.
Behavior-preserving follow-up on #104651.
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).
- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
resumed turn while the durable transcript still matches, cleared with the row on
compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).
Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
Older desktop builds appended a frozen "## Messaging other agents" section (roster
included) to SOUL.md. Since the server started injecting the live section into Bot Chat
sessions, that copy did two wrong things: every CLI/TUI/messenger session paid ~600 tok
for a bot-only protocol, and in Bot Chat itself the probe went silent when SOUL carried
the heading, so bots saw the stale roster instead of the live one.
- load_soul_md strips the legacy section at read time (covers un-migrated profiles and
the ambient-home edge cases the same way the SOUL isolation fix does)
- bot_mode_probe drops the SOUL-carries-heading suppression; a SOUL-era stored Bot Chat
prompt now counts as legacy and is upgraded once (stamped, so it cannot loop)
- config migration v41 rewrites SOUL.md across the default + every profile once
The alias check re-implemented most of _is_same_auth_store (#101356) inline, but
that helper answers False on a samefile() OSError — on the strip path that would
fail OPEN. Resolve identity positively (same resolved path, or samefile says so)
and treat any error as "refuse"; three lines instead of the inline stat dance.
The `preserve_symlinks=False` save path (raw os.replace) bypassed atomic_replace's
EXDEV/Windows fallback to defend against a path swap no concurrent writer performs
on a directory this process just created — removed with its mock-injected test.
Tests kept: shared-store invariant [symlink|hardlink] and two fail-closed variants.