Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:
1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
but "acp" was absent from EXECUTION_SURFACES, so the contract's
closed-schema fallback folded every editor session into "other" --
the bucket meant for genuinely unclassifiable traffic.
2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
"platform" entirely, so every batch task run reported "unknown"
despite "batch" already being a first-class surface.
Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".
Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
self.platform = "batch" on the runner, and defaulted at the worker call
site so callers that build a config without it stay attributable
Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.
Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
M1 revert acp from EXECUTION_SURFACES -> 3 failed
M2 revert acp entrypoint mapping only -> 1 failed
M3 revert batch passthrough -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
HermesProviderMixin._handle_refresh_response overrides the SDK's handler (to
accept any 2xx and keep token bodies out of logs) but dropped the SDK's RFC 6749
section 6 carry-forward. An authorization server that does not rotate refresh
tokens (TinyFish, Google, Zoho, Asana, Futu) answers the refresh grant without a
refresh_token; we then stored the response verbatim, erasing the only refresh
token we had, so the next expiry had nothing to refresh with and forced a
browser re-auth roughly one TTL after every login.
Carry the prior refresh_token (and scope, per section 5.1) forward on the
OAuthToken before _store_tokens, so both the live provider and the on-disk
token file keep it. A rotating AS still wins: only None fields are filled.
Tests: two invariants on the real HermesMCPOAuthProvider + HermesTokenStorage
(omitted -> preserved in memory and on disk; provided -> rotated). The
carry-forward test is red on main.
Port Adolanium's focused-turn pose from Hermes-Bot-Mode#101 and
hermes-agent#88134 to the current typed Bot Mode implementation.
Match the busy signal's connection-qualified focused owner rather than
the gateway socket, retain worker activity, and ease transitions in
elapsed time on the existing shared face clock.
Includes owner-isolation and animated-pose invariants, both proven red
on origin/main, and native Electron before/after verification against
a real temporary Hermes backend with held loopback inference.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Add Move up/down controls for actual rooms without changing bot or folder
ordering. Preserve default pin/activity ordering until an explicit move,
retain hidden room slots, and persist Desktop-local order through room
updates, mirror merges, and hydration. No membership or routing writes.
Adapted narrowly from the group ordering idea in archived
NousResearch/Hermes-Bot-Mode#105 by @onuraycicek; rename already exists.
Co-authored-by: Onur Aycicek <onur.m.aycicek@gmail.com>
Lowering the session trigger must not replace the window-relative lean
selection budget with threshold times target_ratio. Invalidate the lean
cache through the existing property while preserving explicit legacy and
external-engine fallback behavior.
Narrow adaptation of the aux-sync diagnosis and invariants in #93576,
without adding a required recalibration method to context engines.
Related: #95681, #93576
Co-authored-by: Turgut Kural <58116817+TurgutKural@users.noreply.github.com>
The importer stays in the command palette. The labeled nav row was
clutter next to New session / Capabilities / Messaging.
Co-authored-by: Cursor <cursoragent@cursor.com>
hideOnly chrome pinned the Sessions/Bots strip on at any tab count, so
never was a silent no-op. The panes stay; ⌘⌥T brings the strip back.
Co-authored-by: Cursor <cursoragent@cursor.com>
Complete the PR #101452 salvage, preserving sprmn24 authorship. Credit @kokhlo PR #101591 for independently diagnosing the in-flight rename race; use scoped authoritative room bindings rather than persistent aliases. Keep mention attention independent and retire late work after disband.
$groupNeedsYou was written by two independent sources: an @user mention
(appendGroupChatEntry) and, until now, syncGroupClarify whenever a member
blocked on a clarify or approval. Nothing kept the two in sync, so every
path that consumed a clarify -- answering it, the server resolving it,
disbanding or renaming the room -- left a stale badge lit with nothing
behind it, because none of them cleared the copy syncGroupClarify had
written. A naive fix (writing false back into $groupNeedsYou on every
clarify cleanup path) would have created the opposite bug: clearing an
unrelated, still-unresolved @user mention.
$groupClarify is already the source of truth for pending clarify/
approval attention. syncGroupClarify no longer writes $groupNeedsYou; a
new groupHasPendingClarify(clarifies, group) derives the same signal
from a $groupClarify snapshot. It's intentionally pure -- the caller
(roster-pane) subscribes to $groupClarify itself via useValue and passes
the live snapshot in, so the subscription actually drives the
recalculation instead of existing only to force a re-render. A prompt
resolving, being answered, or its room being disbanded all self-correct
through the existing $groupClarify cleanup paths, with nothing left to
keep in sync.
Rename needed its own fix: the old clearGroupClarify(oldName) call on
rename dropped a clarify-only room's attention entirely, since clarify no
longer lives in $groupNeedsYou to be carried over by the existing
old->new key swap there. New renameGroupClarify(oldName, newName) re-keys
matching mirrors onto the new name in two passes -- unrelated entries
first, migrated entries last -- so a stale mirror already stranded at the
destination key (left behind by a known, separate in-flight-poll race)
can never clobber the just-migrated current prompt.
The roster now reads groupNeedsYou[group] || groupHasPendingClarify(...)
and subscribes to both stores so either one repaints the row.
Tests: multi-clarify sequencing, mention+clarify independence (resolving
one must not clear the other), server-side resolution cleanup, rename
migration, rename destination-collision (red-green verified against the
prior single-pass implementation), a GroupRow badge render test, and a
subscription harness using the real stores/useValue/helper end-to-end.
(cherry picked from commit e5038f12ccc1a4d1fb56ecc5f0e1a884f1e8b795)
_user_systemd_socket_ready() accepts systemd/private alone, which is enough for
systemctl --user but not for the systemd-run --user that restart-safe workers
need; systemd_user_bus_env() requires the bus socket. Replace the uid threading
through five helpers with one _wait_for_target_user_bus(uid) that polls
/run/user/<uid>/bus, and move the post-enable wait + restart hint out of
_ensure_linger_enabled into _ensure_system_service_linger so the activity probe
runs only when linger was actually just enabled. Kanban applies the bus env
unconditionally like the cron sibling. Refs #104893.
run_gateway() adopts the user bus once at boot; the generated system unit had
no ordering against user@<uid>.service, so after a reboot the two race and
the adoption can miss until the next gateway restart. Emit After=/Wants=
user@<uid>.service for the unit's User= (uid now returned by
_system_service_identity, which already resolved the account). Existing
system units are flagged outdated once and refreshed on the next
install/restart. Refs #104893.
A system-level gateway unit has no ordering against user@<uid>.service and
linger may be enabled after boot, so the bus can appear after the one-shot
adoption in run_gateway() ran. Derive XDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESS
fresh for the availability probe and every scoped spawn (cron worker, Kanban
worker, PTY/pipe terminal spawns, scope cleanup) so the 60s failure TTL can
actually recover. Refs #104893.
A system unit's `User=` has no login session on a headless host, so
user@<uid>.service never starts and `systemd-run --user --scope` — every
restart-safe cron/Kanban worker — has no bus to reach. `hermes gateway
install --system` runs as root and already knows the target user, so enable
linger for that user on fresh install, on an already-current unit, and on
repair.
- `get_systemd_linger_status()` / `_ensure_linger_enabled()` take the target
username; root is included (restart-safe workers cross `systemd-run --user`
regardless of who the gateway runs as).
- After `loginctl enable-linger` succeeds, wait for the TARGET uid's control
socket (`_wait_for_user_dbus_socket(uid=...)`) — logind starts the user
manager asynchronously and `--start-now` boots the gateway immediately; the
caller's own env (root's) says nothing about it and is never adopted.
- Messages and the manual-remediation hint are scope-aware (`sudo systemctl
restart`, not `systemctl --user`); when the repaired service is already
active, say that a restart is required — `systemctl start` on an active
unit is a no-op and the running gateway keeps its bus-less environment.
Refs #104893.
The docs batch (#105782-#105787, tracking #105788) fixed the website docs but
flagged two in-code strays it could not touch:
- hermes_cli/tips.py:121 still said delegate_task "spawns up to 3 concurrent
sub-agents"; the default has been 10 since the v33 config migration
(config_defaults.py:1256, config_migrations.py:613).
- hermes_cli/config_defaults.py:1778 said "remote backends run per-call",
contradicted by tools/code_kernel_remote.py (session kernels on remote
backends, failing open to per-call only when the backend cannot spawn one)
since #96991.
Comment-only changes: tips.py list entry updated, config_defaults.py comment
rewritten to match the actual kernel lifecycle. No behaviour change.
Validated: test_tips.py 6 passed; ruff, check-windows-footguns, and
check_compat_pointers all clean. Refs #105788 (crumb noted in the issue body).
Separate saved Cloud instances from the live window source, use the existing
registry activation path instead of repeating sign-in/apply, and persist the
chosen instance name while retaining registry identity and custom labels.
Naming metadata adapted from IAvecilla's contribution in PR #103224.
Co-authored-by: IAvecilla <ignacio.avecilla@lambdaclass.com>