Commit Graph

612 Commits

Author SHA1 Message Date
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
Teknium c0ce7473bd fix(serve): Windows conflict probe uses SO_EXCLUSIVEADDRUSE (SO_REUSEADDR binds over live listeners on WinSock) 2026-08-24 09:55:15 -07:00
Teknium de07bd5fab fix(serve): emit BACKEND_PORT_IN_USE sentinel + exit 75 on port bind conflict (#93608)
A held port made 'hermes serve' print only uvicorn's bare
'ERROR: [Errno 98/10048] error while attempting to bind on address'
and exit 1 — indistinguishable from a broken backend for the desktop
spawn and wrapping scripts.

- Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags)
  before uvicorn.Server; on conflict print machine-readable
  'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely
  holders, exit 75 (EX_TEMPFAIL — existing repo convention, see
  gateway/restart.py, kanban_db.py).
- Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind
  failure is re-checked and translated on both POSIX and Windows
  runner paths.
- --port 0 (ephemeral) short-circuits the probe: unchanged behavior.
- HERMES_BACKEND_READY contract untouched.
- Tests: real held-socket repro (sentinel + exit 75, sabotage-proven
  to fail as bare exit 1 without the fix), free-port boot regression,
  ephemeral-port regression, probe/classification units.
- Docs: port-conflict paragraph under 'hermes serve' in
  reference/cli-commands.md.
2026-08-24 09:55:15 -07:00
chelsealong 9c013eaaf8 fix(dashboard): follow scroll on implicit active-session resume (#93518)
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.
2026-08-24 03:21:49 -07:00
Teknium d9a48f656a fix(desktop): scheduled jobs on sleeping profiles keep firing
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
2026-08-24 03:14:30 -07:00
Teknium b03b8ac51d fix(dashboard): name the exact gate trigger in fail-closed refusals
When the bind is loopback and the only gate trigger is
dashboard.public_url, the startup refusal now says so explicitly and
gives both exits (configure a dashboard auth provider, or remove
dashboard.public_url if the proxy no longer exists). Prevents the
stale-public_url mystery-locked-dashboard upgrade trap.

Adds a truth-table regression suite for should_require_auth and the
fail-closed message shape.
2026-08-23 19:55:17 -07:00
e-macgregor d3df14a7e3 fix(dashboard): secure loopback public URL proxy mode 2026-08-23 19:55:17 -07:00
Teknium 65c58651b0 feat: review slot appears in every aux-model picker (desktop, dashboard, CLI)
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.

- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
  stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
  zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
  _AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
  of fallback-providers and the delegation /review section missed in
  #93339
2026-08-23 18:22:39 -07:00
Teknium fdd8d75ba0 fix(gateway): make ws keepalive and orphan-reap grace config-driven (#79635)
- New dashboard.ws_ping_interval / dashboard.ws_ping_timeout defaults
  (20.0/20.0) in DEFAULT_CONFIG; hermes_cli/web_server.py reads them for
  non-loopback binds. Loopback keeps ws_ping=None (event-loop stalls must
  never kill a healthy local connection).
- New dashboard.ws_orphan_reap_grace_s (20.0): tui_gateway/server.py's
  _WS_ORPHAN_REAP_GRACE_S now resolves from config via
  _resolve_ws_orphan_reap_grace(); the HERMES_TUI_WS_ORPHAN_REAP_GRACE_S
  env var is kept as an internal override for backward compat and wins
  when set.
- tests/test_ws_keepalive_config.py: real load_config against a temp
  HERMES_HOME yaml — defaults, propagation, deep-merge, env override,
  invalid-value fallback.
2026-08-23 17:43:39 -07:00
Gille b7cb321223 fix(dashboard): preserve placeholder cwd fallback 2026-08-23 15:21:48 -07:00
Gille 3a81721147 fix(dashboard): preserve exported terminal overrides 2026-08-23 15:21:48 -07:00
Gille 5a85c4a77b fix(dashboard): scope terminal config to selected profile 2026-08-23 15:21:48 -07:00
kshitijk4poor a07ada235c fix(test): deterministic delivery-spawn sentinel + fold bot_mode into agent config tab
Two CI failures: (1) the target_busy test's global subprocess.run patch
recorded unrelated gateway-init git calls (rev-parse/ls-remote) as the
delivery spawn — now local_delivery_command is monkeypatched to a
sentinel argv so only the real delivery path counts; (2) the new
bot_mode config section is single-field, tripping the dashboard
no-single-field-categories rule — folded into the agent tab via
_CATEGORY_MERGE (same as #93102's fix).
2026-08-24 02:08:53 +05:30
kshitijk4poor 3ac6306655 fix(web): fold single-field bot_mode config section into agent tab
The new bot_mode.envelope_ttl_seconds default created a one-field
dashboard category, tripping test_no_single_field_categories. Merge it
into the agent tab via _CATEGORY_MERGE like code_execution et al.
2026-08-24 01:05:08 +05:30
webtecnica f293e7206b fix(dashboard): detect stale code after hermes update and refuse model picker with clear 503 (#86207) 2026-08-23 06:36:57 -07:00
Teknium 10f0d2278b feat(desktop): client-direct voice — use the active profile's STT/TTS keys from the desktop, no audio relay
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.

Backend:
- tools/voice_client_config.py: single resolver; per-provider client
  wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
  elevenlabs-tts). Server-host-only providers (local whisper, edge,
  command/plugin) and missing credentials resolve to {mode: relay}.
  xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
  same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).

Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
  profile) with 60s TTL, provider-direct STT + TTS calls, sentence
  cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
  first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
  startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
  below it. Barge-in via the same stopVoicePlayback sequence bump.

Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.

Docs: voice-mode.md client-direct section ships in this PR.
2026-08-23 03:54:52 -07:00
Teknium 8804e78354 feat(dashboard): Desktop and dashboard read the update receipt instead of inferring success (#91277 Phase-1 bullet 3)
Builds on @mrsucesso's durable-marker recovery (previous commit):

- GET /api/hermes/update/receipt — the full durable receipt (steps,
  skips, gateway restart outcome, fleet matrix) + compact summary; the
  authoritative update-outcome record (written by every run since
  #91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
  and when BOTH the in-memory registries and the update.log marker are
  gone (dashboard restarted + log rotated — the #81193 state), a
  finished receipt reports the outcome: success→0, partial→1. A
  still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
  finished receipt whose run started at/after this apply is
  authoritative, replacing timeout-based failure inference across the
  update's restart gap ('Backend update failed' on successful updates,
  #81193; 'boot failed' during update restarts, #87359).

Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).
2026-08-23 00:50:41 -07:00
Mauricio Ruiz 091106092b fix(dashboard): recover update success after restart
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
2026-08-23 00:50:41 -07:00
Bruce Xu 50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
abundantbeing c9fd5223f6 feat(browser): enable extension controller actions 2026-08-21 22:33:45 +05:30
abundantbeing 5df1d0e113 feat(browser): add authenticated control broker 2026-08-21 22:33:45 +05:30
Teknium 95fa814269 feat(process): positive process identity — spawn tags, machine spawn ledger, Windows job-object self-attach
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:

- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
  spawn-ledger.json self-registration keyed on (pid, create_time) — PID
  reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
  and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
  the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
  themselves at startup and attach to the job; Desktop legacy
  HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
  works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
  _ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
  backends (purpose reapable + recorded spawner provably dead) in ANY update
  context. Ledger-unknown holders fall through to the existing rungs.

22 new tests, sabotage-verified.
2026-08-20 04:47:38 -07:00
Jack Lau 49cc3708e5 fix(dashboard): coalesce repeat gateway restarts for a short window
`_spawn_gateway_restart` already reuses an in-flight `hermes gateway
restart` child so a double-clicked button cannot start two racing
restarts. That guard evaporates exactly when it is needed most: the
child exits as soon as it has handed the restart to the supervisor (or
to the running gateway), long before the gateway is actually back, so a
stale cached dashboard frontend re-firing its own restart every few
seconds cleared the guard on every attempt and started a fresh restart
each time.

#89034 measured the result on an s6-supervised container: 77
`gateway-restart started` entries, 17 of them inside one minute. Each
one SIGHUPs a gateway that is still coming up, and killing it
mid-FTS5-write corrupted `state.db` ("database disk image is
malformed", 203x in agent.log) until the operator recreated the file by
hand.

Requests for the same profile within GATEWAY_RESTART_COOLDOWN_SECONDS of
the last spawn are now coalesced onto that spawn and logged, so a storm
produces one restart instead of one per request. The window is fixed
rather than health-gated on purpose: a gateway that never comes back
would leave a health-gated restart action permanently inert, which is a
worse failure than the flood it prevents. The cooldown state is kept
outside `_ACTION_PROCS` because completed action children are reaped out
of that table, and a guard that disappears when the child exits is the
bug being fixed.

Only the *frontend-flood* half of #89034 is addressed here. The s6
`finish` death-cap the report also asks for is a separate change to
`hermes_cli/service_manager.py` with a much larger blast radius, and is
left for a maintainer decision.
2026-08-20 02:58:31 +05:30
jonny 649c20629e Merge pull request #85429 from NousResearch/jb/yolo-settings-session-info-emit
fix(desktop): re-emit session.info when approvals config changes out of band
2026-08-19 09:29:42 +03:00
Brooklyn Nicholson 443c104e61 fix(desktop): treat profile=default as this process's own home
_is_other_profile only allowed empty/current, so a ?profile=default
save skipped the session.info broadcast on the process whose config
it just wrote. Compare the resolved target to the process HERMES_HOME.
2026-08-19 01:23:39 -05:00
ethernet d5a9c2ba6c feat(nix): home-manager module, shared with the NixOS module
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.

This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:

  - the options
  - the renderers for config.yaml, .env and the documents
  - the activation body
  - the command lines of the processes

nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.

`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.

The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.

The change also makes four corrections that apply to both modules:

- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
  had only the gateway. But Hermes Desktop and the web dashboard connect
  to a different process, so six of the configurations in public repos
  add a second unit by hand. serve and dashboard are one entry point with
  one flag of difference, and you can run only one of them. Thus the
  option is an enum. The NixOS module asserts against container mode with
  a backend, and does not make a unit that cannot start.

- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
  installs into the working directory, which is correct for AGENTS.md but
  wrong for SOUL.md and memories/. Hermes reads those files from
  HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
  made a workspace file that Hermes never loaded as the identity. The
  documentation said this in prose, but two directory diagrams showed the
  opposite. This change corrects both. A key in either option can now
  contain subdirectories.

- `documents` needs an explicit `workingDirectory`. The default of that
  option is bad on both modules. It is the home directory of the user on
  Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
  workspace files without a directory therefore gets a place that the
  user did not select. The place is also different on each module. The
  modules now refuse that combination.

  The test is on the priority of the option and not on its value. An
  option that nothing sets keeps the priority of its own default, and
  each definition from a user is stronger. Thus a directory with the same
  text as the default still counts as a selection, and so does a
  mkDefault. A comparison of values detects neither case.

- Each activation writes .env again from a base in the Nix store, and
  does not add to the file that exists. Thus a second activation cannot
  put the same secret in the file two times, and a removed
  environmentFile goes away. environmentFiles keeps the type `listOf
  str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
  the Nix store, which all users can read.

- HERMES_MANAGED and the .managed marker now hold the name of the system
  that manages the install. Thus a refusal says "managed by home-manager"
  and not "managed by NixOS", and `hermes update` gives the Nix guidance
  for both shapes. The CLI does not print a rebuild command for each
  system. It names the owner, and the user knows their own tool. A bare
  `true` and an empty marker still mean NixOS, so this does not change an
  existing install.

Verification. Six new checks, all built:

  nixos-module           evaluates the module with evalModules and the
                         NixOS module list. It asserts both units, one
                         HERMES_HOME, and that the module refuses
                         container mode with a backend.
  home-manager-module    evaluates the module with the
                         homeManagerConfiguration function of
                         home-manager. The process assertions run against
                         systemd units on Linux and launchd agents on
                         Darwin.
  module-option-parity   asserts that each shared option is on both
                         modules, and that the two exclusion lists name
                         only options that exist.
  env-file-assembly      runs the real .env script and checks the
                         contents, the mode, that a second run gives the
                         same bytes, and that a removed file goes away.
  workspace-files-need-a-directory
                         checks that the module refuses `documents`
                         without a directory, and accepts a directory
                         that has the same text as the default.
  service-argv           runs each command line that the modules build
                         through the real parser of the CLI, with one
                         sentinel flag added, and requires that argparse
                         refuses only the sentinel.

`nix flake check` passes, with 21 checks in total.

The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.

Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:

  - a lost --no-open
  - a backend that runs the gateway
  - an overwritten config.yaml
  - documents in the wrong directory
  - a different HERMES_HOME on the two processes
  - a lost HERMES_HOME export
  - a missing backend unit
  - a removed assertion
  - an .env file that grows at each activation
  - an install that reports NixOS
  - an empty .managed marker
  - an option on the NixOS module only
  - a stale entry in an exclusion list
  - a renamed subcommand
  - an unknown flag
  - the workspace-files assertion always passes
  - the assertion compares values instead of priorities
  - an off-by-one that lets an untouched default through
  - the assertion also fires for hermesHomeFiles
  - a mkDefault no longer counts as a selection
  - the Home Manager module stops wiring the assertion
  - the NixOS module stops wiring the assertion

The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.

These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:

  - test_git_probe_tree_kill.py (2 tests)
  - test_update_import_guard.py (1 test)
  - test_telegram_media_read_timeout.py (2 tests)
  - test_teams.py (a collection error)

Closes #9056

# Conflicts:
#	hermes_cli/main.py
#	hermes_cli/update_cmd.py
#	hermes_cli/web_server.py
2026-08-18 20:42:06 -04:00
Teknium 5ce09b3c1e perf(desktop): Bot Mode wakes paint-first — transcript paint completes the
wake instead of the full runtime boot (#89206 class)

zero trust's third bundle (on ae6578af, both prior fixes present) showed the
remaining failure: cold profile backends on slower Windows machines take
47-120s to fully boot, while the wake path's fixed budgets (20s hydration,
~15s resume retries) raced the whole boot and lost — "errors waking up BOTS"
while the backend came up healthy moments later.

Rather than raising timeouts, make the wake cheap:

- waitForFocusedSessionHydration: a history-bearing chat is hydrated when
  the persisted transcript is PAINTED on the right session. The REST
  prefetch delivers that seconds after the backend's HTTP is up; the full
  runtime resume (agent build, MCP discovery, 114-skill load) keeps warming
  in the background and binds the composer when it lands. Only an
  expected-empty chat still waits for the runtime (nothing to paint).
- On hydration timeout, log a [bot-wake] phase breakdown (activation ms,
  hydration ms, which conditions were unmet) to the renderer console so the
  next support bundle pinpoints the slow phase directly.
- web_server: flush the headless "listening" line — block-buffered on the
  Desktop's piped stdout, it surfaced minutes late and made boots look far
  slower than they were in support bundles (the 120s "gap" in this bundle
  was partly this artifact).

Sabotage-proven: restoring the runtime-gated wait fails the new paint-first
test by timing out — the exact field shape.
2026-08-18 15:52:39 -07:00
Teknium cc421cb697 fix: dashboard console skills commands no longer act on the wrong profile (#65828)
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.

Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.

Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.

Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root

Fixes #65828
2026-08-18 14:14:42 -07:00
Teknium a1682376ca feat(profiles): rename any agent — the default profile gets a display name (#45624)
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.

Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").

Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
2026-08-18 02:27:18 -07:00
Teknium c9ce66e25e fix(desktop): one roster row per bot even when a backend is registered under two addresses — /api/status install_id + roster collapse
Backend: /api/status now carries a stable random install_id persisted once
under the root HERMES_HOME, shared by every profile of the install.
Desktop: roster enumeration captures it per connection (TTL-cached probe),
buildAgentRoster collapses same-install rows with a deterministic canonical
pick (active > local > ssh > remote > cloud > earliest), the @name-device
handle rule runs after the collapse, and Settings → Gateways shows a
display-only 'Same backend as' hint. Backends without install_id bypass the
collapse (fully backward compatible).
2026-08-17 19:22:57 -07:00
adybag14-cyber d685ea4df5 docs(termux): document community pkg distribution 2026-08-17 15:47:17 -07:00
adybag14-cyber 26466d0594 feat(termux): support apt-managed Hermes installs 2026-08-17 15:47:17 -07:00
Teknium 0b13cafffa feat(cron): continuity toggle across dashboard, Bot Mode routines, and TUI cron RPC
Wire the continuity flag through every cron-creation surface, not just the
model tool:

- dashboard (web/): checkbox in the cron job editor; form state round-trips
  the stored reserved 'self' entry into the toggle and strips it from the
  context_from textarea; web_server dashboard validator skips 'self'
  (create precedes the job's existence)
- Bot Mode Routines tab (hermes-bots plugin): Continuity checkbox in the
  New Cronjob dialog, forwarded through cron.manage
- tui_gateway cron.manage RPC: optional continuity param on action=add

vitest cron-job suite 10/10 (4 new), tsc app project clean, py_compile clean.
2026-08-16 22:09:28 -07:00
Teknium 0e4552199a fix(status): keep fatal platform entries visible when gateway startup failed
Follow-up to #80451. /api/status cleared gateway_platforms whenever the
gateway process was down — correct for a clean stop (stale 'connected'
states are noise) but wrong for startup_failed, where the fatal entries
ARE the diagnosis: per-profile credential collisions and auth failures
(multiplex '<profile>:<platform>' keys) that the single exit_reason
string cannot express. #80451's writer-identity and freshness filters
already drop entries from other/older processes, so preserving
fatal-state entries here cannot leak another gateway's live state.

Live-validated shape: a real multiplex gateway (2 secondary profiles,
rejected tokens) persists telegram / alpha:telegram / beta:telegram
fatals in gateway_state.json; /api/status previously reported {} for
platforms while state was startup_failed.
2026-08-16 21:54:31 -07:00
Shannon Sands d66341ab28 fix(status): strict writer-identity ownership for aggregated platform entries (OOF-3)
The freshness window (updated_at >= live process create_time - 2s) had a
P1 boundary hole: a stale failure written by the PREVIOUS process
immediately before a fast restart landed inside the slack and was
aggregated; if that platform was then removed, the new process never
replaces the entry and NAS stays degraded indefinitely.

Replace clock heuristics with persisted writer identity:

- write_runtime_status now stamps every platform entry with the writing
  process's (writer_pid, writer_start_time) — the same PID-reuse
  fingerprint the liveness checks use, so a recycled PID never
  masquerades as the original writer.
- The aggregation ownership filter requires exact equality between an
  entry's stamp and the profile's validated live gateway process
  (get_runtime_status_running_pid + _get_process_start_time). No slack,
  no timestamps. Legacy entries without a stamp fail closed.
- Writer stamps are process recon (same class as the auth-gated
  gateway_pid) and are stripped from all /api/status projections, both
  active-profile and merged cross-profile entries.

Near-boundary regression test: prior-process entry stamped 100ms before
restart is excluded; recycled-pid-different-fingerprint excluded;
legacy no-stamp excluded; current-process entry kept.
2026-08-16 20:32:09 -07:00
Shannon Sands 57279cf2b9 fix(status): freshness-filter aggregated per-profile platform entries (OOF-3)
Gateway startup deliberately preserves plain platform entries in
gateway_state.json across restarts, and the active-profile endpoint
compensates by filtering against current configuration. The cross-profile
aggregation copied raw maps, so a fatal entry for a platform the operator
had since disabled/removed could keep NAS reporting the instance degraded
indefinitely.

The aggregation has no cheap per-profile config context (platform sets
depend on tokens in each profile's .env behind its secret scope), so use
freshness instead: an entry is aggregatable only when its updated_at is
at/after the live gateway process's create time (validated PID via
get_runtime_status_running_pid + psutil create_time; the record's own
start_time field is a PID-reuse fingerprint in clock ticks, not a
timestamp). Config changes require a restart to take effect, so
restart-anchored freshness is exactly the config filter's semantics.
Fail closed: unparseable timestamps or no live process exclude the entry
— a false 'degraded forever' is the worse failure mode.
2026-08-16 20:32:09 -07:00
Shannon Sands 1d46a9fe0f fix(status): aggregate independent per-profile gateway failures; harden key filter (OOF-3)
- /api/status now folds LIVE independent per-profile gateways' platform
  failures (gateway_mode == 'multiple', the OOF-3 deployment mode) into
  gateway_platforms under the validated <profile>:<platform> grammar, so
  NAS fleet health sees them without a schema change. ?profile= requests
  stay unmerged (single-profile view).
- Namespaced-key validation no longer fails open: colon-containing keys
  are grammar-checked even when configured-platform loading throws.
- Platform key segment now accepts hyphens, matching plugin platform IDs
  (plugins/platforms/<dir> names, e.g. foo-bar).
2026-08-16 20:32:09 -07:00
Shannon Sands f755ed5e90 fix(gateway): surface multiplex profile failures (OOF-3) 2026-08-16 20:32:09 -07:00
Shannon Sands 4822156923 fix(cron): stop retry storms when the gateway is deliberately stopped (OOF-266)
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).

Split the unreachable path by durable operator intent:

- desired_state == "stopped" (written only by the s6 lifecycle commands;
  the same intent signal container-boot reconciliation trusts) -> drop
  the fire with 200 + a structured log line, mirroring NAS's own
  instance_stopped drop. Jobs are not lost: the Chronos provider
  reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
  file without desired_state) -> keep the retryable 503, now stamped
  with Retry-After: 60 so a scheduler that honors it spaces retries
  past the wake/restart window instead of exhausting them inside it.
  The gateway's own pass-through 503s (draining) get the same hint.

The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
2026-08-16 19:50:07 -07:00
yu-xin-c c3d74603aa fix(desktop): avoid PowerShell parent marker boot gate 2026-08-16 02:14:35 -07:00
Teknium 1e6717baf2 fix(web_server): discover root user plugins under profile-scoped processes
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.

Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.

Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.

Fixes #87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
2026-08-16 02:07:16 -07:00
fangliquanflq 29dfbf2d6a fix(sessions): surface open sessions skipped by prune 2026-08-16 01:55:25 -07:00
Teknium 99500aca11 fix: guard breadcrumb writes for bare test fakes; fold session category into general
- cli.py breadcrumb call sites use getattr so object.__new__/SimpleNamespace
  test fakes without the mixin method don't AttributeError (pitfall 17)
- web_server: fold one-field 'session' category into 'general' per the
  single-field-category test contract
2026-08-15 18:15:18 -07:00
Benjamin Ang ad42ecfc06 fix(desktop): stream remote media without renderer credentials 2026-08-15 01:53:34 -07:00
Paul BlackSwan d5a865882a fix(desktop): save remote gateway files natively 2026-08-15 00:37:00 -07:00
Christopher 5bceb3e84b fix(dashboard): add idle back-off to PTY pump loop (#42627) 2026-08-15 00:32:53 -07:00
Lucas Oliveira f0cfe5a56f perf(dashboard): bound multi-profile sidebar polling 2026-08-15 00:32:53 -07:00
Teknium f45813ea77 fix(sessions): run state.db schema migration eagerly at backend startup and stop swallowing locked ALTERs
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).

Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):

1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
   writable open, typically the user's first NEW session. The dashboard
   backend now schedules one writable open of its own state.db from the
   lifespan (daemon thread, never blocks the ready-probe socket, never
   raises), so the store is brought current before the first session-
   list poll on every `hermes serve` / `hermes dashboard` / Desktop
   headless entrypoint.

2. _reconcile_columns caught sqlite3.OperationalError around every
   ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
   orphaned sibling backends made the ALTER fail silently — startup
   "succeeded" with a half-reconciled schema, and the open-time lock
   patience (#74478) never saw the error because it was swallowed
   inside first. Now: "duplicate column" races stay at DEBUG,
   locked/busy re-raises so _connect_and_init_with_lock_patience
   retries the whole idempotent init with jittered backoff, and any
   other failure (e.g. un-ADDable NOT NULL) logs at WARNING.

Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.

Fixes #79531
Fixes #80037

Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
2026-08-14 21:36:29 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00