Commit Graph

5936 Commits

Author SHA1 Message Date
Jeffrey Quesnelle 612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
Teknium 095f003377 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	hermes_cli/tools_config.py
2026-08-19 16:48:40 -07:00
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium 803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
Teknium d762ed9b3c feat: execution-discipline guidance now reaches all tool-capable models (config model.execution_guidance)
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.

Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.

The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
  file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
  format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
  broader retry before concluding)
- completion gated on verification (done = every named acceptance
  criterion verified, never a plausible subset)

The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.

Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.

Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).

Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
2026-08-19 16:16:37 -07:00
Teknium aa3c5e59d3 fix(tools): honor raw stt.provider: local; finish _reconfigure_provider provider-string migration
Two real gaps the CI-red sibling tests exposed:

- read_selection() treated EVERY raw stt.provider: local as the legacy
  DEFAULT_CONFIG seed and reported no-selection — but the seed never
  reached config.yaml (save_config strips schema defaults), so a
  picker- or hand-written local pick was silently discarded and the
  autodetect ladder could route an explicit local user to cloud STT.
  A raw 'local' is now a genuine selection; the merged-view ambiguity
  note replaces the over-broad shim (mirror comment updated in
  nous_subscription._selected_provider and _get_provider).

- _reconfigure_provider was half-migrated: the tts/stt/browser/web
  branches and the managed-category fallthrough still wrote
  use_gateway flags and vendor names for managed rows. They now write
  the single provider string ('nous' for managed rows) and pop the
  legacy key, matching _write_provider_config.
2026-08-19 16:10:01 -07:00
Teknium b10c5a8084 fix(cli): persist one provider string per picker row; mirror strict routing in status
Every hermes tools row now writes exactly one selection value per
category — managed 'Nous Subscription' rows write 'nous', BYOK rows the
vendor name (including the historically-unset BYOK-FAL image row) — and
use_gateway is no longer written; fresh picks drop any legacy key so the
read-time shim cannot override them. The non-managed clear now resolves
the category from the row's own markers, covering plugin-injected rows
the TOOL_CATEGORIES loop missed. Setup-flow writers (managed defaults,
gateway enablement) store 'nous', and the feature-state mirrors in
nous_subscription.py compute per-category selections with the same
legacy interpretation so hermes status matches runtime: a stored vendor
selection pins direct (managed availability no longer lights it up) and
an explicit non-camofox selection beats a stray CAMOFOX_URL.
2026-08-19 16:10:01 -07:00
Teknium 099258ef48 fix(voice): route TTS/STT OpenAI audio on the stored selection, not credentials
Both _resolve_openai_audio_client_config resolvers now switch on the
stored provider string: 'nous' (or legacy use_gateway: true) => managed
openai-audio gateway only, erroring by selection name when unentitled —
the STT twin previously never read the stored gateway intent at all, so
a direct OPENAI_API_KEY silently overrode the Nous Subscription pick;
stored vendor => direct credentials only with a selection-naming error
on missing keys (no silent managed fallback); never-configured keeps the
legacy ladder. DEFAULT_CONFIG stops seeding stt.provider: local, and the
seeded value on existing configs is treated as no-selection so autodetect
keeps working for that installed base.
2026-08-19 16:10:01 -07:00
Teknium f08d3e400f feat: hermes tools lets Exa/Parallel users pick the free keyless or paid keyed endpoint
Exa and Parallel now each render as two picker rows in hermes tools —
'Free (keyless)' and 'Paid (API key)'. Selection persists to
web.provider_tier.<name>:
- free: always the anonymous public endpoint, even with a key set
- paid: always the keyed SDK path; missing key errors instead of
  silently downgrading to the free tier (is_keyless_available also
  returns False so the auto-fallback walk can't route there)
- unset: auto (key present -> paid, else keyless)

Mechanism: get_setup_schema() gains a 'variants' list the picker
flattens into sibling rows sharing one web_backend; selection writes
the tier via both _write_provider_config sites; active-row detection
matches the tier (auto mirrors use_keyless). Routing goes through a
single use_keyless() chokepoint shared by search+extract in both
providers.

Live E2E: tier=free with a fake key present searched keyless OK (a
keyed call would have 401'd); tier=paid without key errored naming
PARALLEL_API_KEY; picker rows verified for both vendors x both tiers.
2026-08-19 15:36:20 -07:00
Teknium 96c2fd3c04 feat: web search/extract now work keyless on fresh installs via Parallel + Exa free tiers
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).

- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
  mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
  errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
  never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
  absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
  after every keyed/importable candidate; check_web_api_key() lights
  the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md

E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
2026-08-19 15:15:27 -07:00
Jack Lau 49cc3708e5 fix(dashboard): coalesce repeat gateway restarts for a short window
`_spawn_gateway_restart` already reuses an in-flight `hermes gateway
restart` child so a double-clicked button cannot start two racing
restarts. That guard evaporates exactly when it is needed most: the
child exits as soon as it has handed the restart to the supervisor (or
to the running gateway), long before the gateway is actually back, so a
stale cached dashboard frontend re-firing its own restart every few
seconds cleared the guard on every attempt and started a fresh restart
each time.

#89034 measured the result on an s6-supervised container: 77
`gateway-restart started` entries, 17 of them inside one minute. Each
one SIGHUPs a gateway that is still coming up, and killing it
mid-FTS5-write corrupted `state.db` ("database disk image is
malformed", 203x in agent.log) until the operator recreated the file by
hand.

Requests for the same profile within GATEWAY_RESTART_COOLDOWN_SECONDS of
the last spawn are now coalesced onto that spawn and logged, so a storm
produces one restart instead of one per request. The window is fixed
rather than health-gated on purpose: a gateway that never comes back
would leave a health-gated restart action permanently inert, which is a
worse failure than the flood it prevents. The cooldown state is kept
outside `_ACTION_PROCS` because completed action children are reaped out
of that table, and a guard that disappears when the child exits is the
bug being fixed.

Only the *frontend-flood* half of #89034 is addressed here. The s6
`finish` death-cap the report also asks for is a separate change to
`hermes_cli/service_manager.py` with a much larger blast radius, and is
left for a maintainer decision.
2026-08-20 02:58:31 +05:30
Brooklyn Nicholson ae6c973fc7 fix(update): run the Windows update hand-off unattended
The re-exec'd child inherits the console, so sys.stdin.isatty() still reported
a terminal and the update asked its local-changes question. By then the parent
shim had exited and the shell had taken the console back, so the prompt could
not be answered and the update sat there forever — worse than the lock it
replaced, because nothing recovers without closing the window.

Spawn the child with stdin closed. It then takes the same path the gateway and
Desktop updates take: honour updates.non_interactive_local_changes, which
stashes by default so nothing is lost, and keep going without asking.
2026-08-19 13:41:44 -05:00
Yingliang Zhang 73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30
Brooklyn Nicholson a9eb99d172 fix(update): stop deferring shim renames to next boot
MOVEFILE_DELAY_UNTIL_REBOOT was the quarantine's last resort, and it is worse
than doing nothing. It writes to HKLM, so a non-elevated update — every
Desktop-driven one, and most terminal ones — gets ERROR_ACCESS_DENIED and
reports nothing. When it does succeed it frees nothing for the install
running right now, and the queued operation outlives that update: at the next
boot it moves aside whatever sits at the shim path, including a shim a later
repair just wrote.

Drops the fallback and sweeps entries older versions queued, matching only
our own <shim> -> <shim>.old.<stamp> pairs so unrelated installers keep
theirs.

Salvaged from #88121 by @fangliquanflq.
2026-08-19 13:14:40 -05:00
Brooklyn Nicholson 867ab54e20 fix(update): re-run Windows updates off the console shim
`hermes update` launched as venv\Scripts\hermes.exe can never finish on
Windows. The launcher runs the interpreter with the shim as its script and
holds it open without FILE_SHARE_DELETE for the whole command, so the
quarantine rename is refused and uv fails to replace hermes.exe with
os error 32 — every time, with no Desktop, gateway or AV involved. The
concurrent-instance preflight cannot catch it because it excludes this
process and its ancestors by design.

Detect the shim from both the process ancestry and this process's own launch
paths (argv[0], __main__.__file__, the spec origin — the runpy/zipapp launch
puts <shim>\__main__.py there), intersected with the project venv's shims so
an unrelated hermes.exe never matches. When it matches, re-run the same
argv as `venv\Scripts\python.exe -m hermes_cli.main ...` and return, which
releases the shim before the child installs anything.

The hand-off sits ahead of the update lock so the child claims the marker
itself rather than adopting one the parent immediately releases, and any
failure falls through to the previous in-process behaviour with the manual
command printed.

Refs #88838, #89599, #86093
2026-08-19 13:13:09 -05:00
Brooklyn Nicholson 7a94b1fbf7 fix(update): resolve the project venv as venv or .venv
`uv venv` writes `.venv` while our installers write `venv`, and every venv
lookup in the update/repair paths hardcoded `venv`. On a `.venv` install
`_venv_scripts_dir()` returned None, so the Windows shim-lock preflight, the
quarantine, and the console-script verification all silently skipped
themselves — the update walked straight into the failure they exist to catch.

Adds `hermes_constants.project_venv_dir()` as the single resolver and routes
both `_venv_scripts_dir()` implementations plus the two VIRTUAL_ENV call
sites through it.

Refs #79542
2026-08-19 13:11:35 -05:00
Noah Lee cd3f9b64de feat(skills): add --yes/-y flag to hermes skills uninstall
`do_uninstall` already accepted a `skip_confirm` parameter and the
slash-command handler already passed `skip_confirm=True`, but the CLI
argparse path never exposed a flag to reach it. This adds `--yes`/`-y`
to `hermes skills uninstall`, matching the existing pattern on `install`
and `reset`.
2026-08-19 10:19:35 -07:00
Alex Fournier 31402f630b fix(relay): complete native plugin cutover
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 3fad83df31 fix(relay): guard native plugin ownership cutover
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Jeffrey Quesnelle b5455fdd16 Merge pull request #84645 from rroverin/feat/add-nemotron-lightning-35-model
feat(models): add Nemotron 3.5 Lightning 30B-A3B to NVIDIA NIM picker
2026-08-19 11:17:24 -04:00
Teknium e6ff4eacdb fix(tools-config): widen stale-model picker guard to legacy image and video pickers
Same bug class as the plugin image picker crash (#77238): an unguarded
`current_model = default_model` fallback that can index the catalog with a
key it doesn't contain when the provider's default drifts from its catalog.
Applies the `default if default in catalog else next(iter(catalog))` guard
to _configure_imagegen_model and _configure_videogen_model_for_plugin.
2026-08-19 01:19:37 -07:00
zhuermu c6b680c445 fix(image-gen): handle stale OpenRouter model defaults 2026-08-19 01:19:37 -07:00
Brooklyn Nicholson 6907c97847 fix(update): don't report success when the Desktop rebuild failed
hermes update treated a failed desktop pack as non-fatal and still printed
✓ Update complete!, so Windows users kept running an old Hermes.exe after a
"successful" update. Withhold the success banner, surface the stale app in
the summary, and write .update_exit_code=1 for gateway watchers.

Supersedes #88359, #87984.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-19 02:10:37 -05:00
jonny 649c20629e Merge pull request #85429 from NousResearch/jb/yolo-settings-session-info-emit
fix(desktop): re-emit session.info when approvals config changes out of band
2026-08-19 09:29:42 +03:00
Brooklyn Nicholson 443c104e61 fix(desktop): treat profile=default as this process's own home
_is_other_profile only allowed empty/current, so a ?profile=default
save skipped the session.info broadcast on the process whose config
it just wrote. Compare the resolved target to the process HERMES_HOME.
2026-08-19 01:23:39 -05:00
Teknium 4903993cdd test: enforce -q/--query-file exclusivity at parse time
The subprocess-based exclusivity test invoked hermes_cli.main in the CI
environment where startup exits 1 before the manual guard runs. Enforce
the conflict in argparse itself (mutually exclusive group, exit 2 at parse
time) and test the parser directly; the manual guard stays for programmatic
namespace fills.
2026-08-18 22:31:32 -07:00
Teknium 1190825652 fix(bot-mode): DM protocol no longer shell-interpolates message bodies
The Bot Mode teammate-DM protocol told agents to inline the message into a
double-quoted shell argument: quotes truncated the body and $(...)/backticks
executed on the sender's machine. The protocol now writes the message to a
temp file and delivers it via a new 'hermes chat --query-file' flag (or '-'
for stdin); 'hermes peer dm' already accepted stdin and the peer recipe now
uses it. No shell pass touches the body at any point.

Supersedes the tool-based approach in #89077 — same bug, fixed with a CLI
flag + protocol rewrite instead of a new model tool.

Co-authored-by: mehmetkr-31 <mehmetkr-31@users.noreply.github.com>
2026-08-18 22:31:32 -07:00
Teknium bdd0a79c6a fix(goals): /goal resume actually restarts work after budget exhaustion
After a standing goal auto-paused on turn-budget exhaustion, every
surface's /goal resume handler only flipped the persisted state back to
active (and reset turns_used) and rendered an acknowledgement — nothing
re-entered the conversation loop, so the goal sat idle until the user
sent another ordinary message.

Fix the whole class by scheduling the canonical
GoalManager.next_continuation_prompt() through each surface's existing
input path after a successful resume:

- Desktop/TUI (tui_gateway/methods_tools.py command.dispatch): return a
  sendable {type: "send"} dispatch with the continuation as the message,
  a "Continuing now" notice, and display "/goal resume" so the
  transcript shows the concise invocation instead of the model-facing
  scaffolding. No-goal keeps the exec response.
- Classic CLI (hermes_cli/cli_commands_mixin.py): put the continuation
  on _pending_input, same as the /goal <text> kickoff.
- Messaging gateway (gateway/slash_commands.py): enqueue a continuation
  MessageEvent through the adapter FIFO — the same path the post-turn
  judge uses — so queued real user messages preempt naturally and the
  pause/clear stale-continuation cleanup recognizes it.

Also correct the now-misleading gateway.goal.resumed copy ("Send any
message to continue…") across all 17 locale files.

Regression tests cover exact budget exhaustion → resume on the real CLI
handler, the real gateway handler (including the
_is_goal_continuation_event guard contract), and the TUI
command.dispatch boundary; verified each fails on the pre-fix code.

Fixes #75362
2026-08-18 19:51:52 -07:00
ethernet 1f234a1033 fix(nix): let the install-method stamp name a home-manager install
detect_install_method reads the stamp against an allowlist. The
allowlist held "nixos" but not "home-manager", and a stamp that names
home-manager gave "unknown". The managed path (step 3) returned the
correct name, so the gap was invisible: it appeared only for an install
that carries a stamp.

An install with the value "unknown" gets "hermes update" as its update
guidance. That command is the one command a managed install refuses, so
the user gets a dead end.

The test for this was also environment-dependent. It called the real
get_project_root(), and it passed here only because this worktree
carries no stamp. A checkout from the curl installer carries a "git"
stamp, and the assertion then failed for the contributor and not for
us. The test now detects against a temporary install tree.

The new test stamps each managed system and asserts the value that
comes back. With the allowlist reverted, the home-manager case fails
with "assert 'unknown' == 'home-manager'". The nixos case passes,
because that name was already in the allowlist.
2026-08-18 20:42:06 -04:00
ethernet d5a9c2ba6c feat(nix): home-manager module, shared with the NixOS module
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.

This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:

  - the options
  - the renderers for config.yaml, .env and the documents
  - the activation body
  - the command lines of the processes

nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.

`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.

The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.

The change also makes four corrections that apply to both modules:

- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
  had only the gateway. But Hermes Desktop and the web dashboard connect
  to a different process, so six of the configurations in public repos
  add a second unit by hand. serve and dashboard are one entry point with
  one flag of difference, and you can run only one of them. Thus the
  option is an enum. The NixOS module asserts against container mode with
  a backend, and does not make a unit that cannot start.

- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
  installs into the working directory, which is correct for AGENTS.md but
  wrong for SOUL.md and memories/. Hermes reads those files from
  HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
  made a workspace file that Hermes never loaded as the identity. The
  documentation said this in prose, but two directory diagrams showed the
  opposite. This change corrects both. A key in either option can now
  contain subdirectories.

- `documents` needs an explicit `workingDirectory`. The default of that
  option is bad on both modules. It is the home directory of the user on
  Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
  workspace files without a directory therefore gets a place that the
  user did not select. The place is also different on each module. The
  modules now refuse that combination.

  The test is on the priority of the option and not on its value. An
  option that nothing sets keeps the priority of its own default, and
  each definition from a user is stronger. Thus a directory with the same
  text as the default still counts as a selection, and so does a
  mkDefault. A comparison of values detects neither case.

- Each activation writes .env again from a base in the Nix store, and
  does not add to the file that exists. Thus a second activation cannot
  put the same secret in the file two times, and a removed
  environmentFile goes away. environmentFiles keeps the type `listOf
  str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
  the Nix store, which all users can read.

- HERMES_MANAGED and the .managed marker now hold the name of the system
  that manages the install. Thus a refusal says "managed by home-manager"
  and not "managed by NixOS", and `hermes update` gives the Nix guidance
  for both shapes. The CLI does not print a rebuild command for each
  system. It names the owner, and the user knows their own tool. A bare
  `true` and an empty marker still mean NixOS, so this does not change an
  existing install.

Verification. Six new checks, all built:

  nixos-module           evaluates the module with evalModules and the
                         NixOS module list. It asserts both units, one
                         HERMES_HOME, and that the module refuses
                         container mode with a backend.
  home-manager-module    evaluates the module with the
                         homeManagerConfiguration function of
                         home-manager. The process assertions run against
                         systemd units on Linux and launchd agents on
                         Darwin.
  module-option-parity   asserts that each shared option is on both
                         modules, and that the two exclusion lists name
                         only options that exist.
  env-file-assembly      runs the real .env script and checks the
                         contents, the mode, that a second run gives the
                         same bytes, and that a removed file goes away.
  workspace-files-need-a-directory
                         checks that the module refuses `documents`
                         without a directory, and accepts a directory
                         that has the same text as the default.
  service-argv           runs each command line that the modules build
                         through the real parser of the CLI, with one
                         sentinel flag added, and requires that argparse
                         refuses only the sentinel.

`nix flake check` passes, with 21 checks in total.

The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.

Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:

  - a lost --no-open
  - a backend that runs the gateway
  - an overwritten config.yaml
  - documents in the wrong directory
  - a different HERMES_HOME on the two processes
  - a lost HERMES_HOME export
  - a missing backend unit
  - a removed assertion
  - an .env file that grows at each activation
  - an install that reports NixOS
  - an empty .managed marker
  - an option on the NixOS module only
  - a stale entry in an exclusion list
  - a renamed subcommand
  - an unknown flag
  - the workspace-files assertion always passes
  - the assertion compares values instead of priorities
  - an off-by-one that lets an untouched default through
  - the assertion also fires for hermesHomeFiles
  - a mkDefault no longer counts as a selection
  - the Home Manager module stops wiring the assertion
  - the NixOS module stops wiring the assertion

The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.

These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:

  - test_git_probe_tree_kill.py (2 tests)
  - test_update_import_guard.py (1 test)
  - test_telegram_media_read_timeout.py (2 tests)
  - test_teams.py (a collection error)

Closes #9056

# Conflicts:
#	hermes_cli/main.py
#	hermes_cli/update_cmd.py
#	hermes_cli/web_server.py
2026-08-18 20:42:06 -04:00
ethernet 6e67841a9a refactor(goals): share one dropped-write warning across managers
Review fixups for #88965. The goal, loop, and heartbeat managers each
had a copy of the same WARNING text. The shared _warn_dropped_write
helper in goals.py keeps the three logs identical and greppable as one
bug class. The _warm_goals_session_db parameter is now label. The old
name ctx said context, but the value is a log label.
2026-08-18 16:08:29 -07:00
ethernet 46d8cf0be3 fix(gateway): /loop paths warm the SessionDB cache off-loop
The loops delegation (previous commit) moved /loop onto the shared
bootstrap windows, but the loop-class gateway callers still constructed
LoopManager on the loop thread with no warm-up — the same false-ack
class as /goal, one sibling over:

- The /loop command handler: a cold init past the window made
  save_loop discard the write while the reply claimed the loop was
  set. Reproduced with a 2s init: "↻ Loop set" with nothing persisted;
  with the warm-up the loop persists.
- _post_turn_loop_completion: a cold cache at the turn boundary
  stalled the loop for the init duration and could drop the
  tick-completion write.
- _loop_wakeup_watcher: the scan reads every persisted loop, so a cold
  cache ran the state.db init on the loop thread before the first
  read.

All three now warm the cache off-loop first (same helper, same reason
as the goal paths). save_loop also logs at WARNING when it drops a
write, matching save_goal: the reply has already told the user the
loop was set.

Review findings (PR #88965): the goal and heartbeat command paths
warmed off-loop, the loop-class paths did not.
2026-08-18 16:08:29 -07:00
ethernet 246477a809 fix(gateway): /goal no longer lies when state.db init is slow
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).

Two changes, one per caller shape:

- Async callers (_get_goal_manager_for_event,
  _get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
  the heartbeat poller) warm the SessionDB cache off-loop through the
  context-preserving executor before constructing the manager (shared
  _warm_goals_session_db helper). The loop never blocks and the first
  write lands at any init duration. A bare to_thread would lose the
  per-turn profile home override under multiplex; the executor hop
  keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
  cannot await, so the bootstrap windows stay: the call that starts the
  bootstrap waits a one-time init window (1.5s) instead of the short
  per-call window (0.25s), giving healthy cold inits room to land while
  a contended migration still degrades to None with only a bounded
  one-time stall. The bootstrap thread binds the caller's home as a
  contextvar override so a multiplexed worker cannot cache the default
  profile's DB under another profile's key.

save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).

Independent diagnosis + measurement by jackulau (#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
2026-08-18 16:08:29 -07:00
Teknium 5ce09b3c1e perf(desktop): Bot Mode wakes paint-first — transcript paint completes the
wake instead of the full runtime boot (#89206 class)

zero trust's third bundle (on ae6578af, both prior fixes present) showed the
remaining failure: cold profile backends on slower Windows machines take
47-120s to fully boot, while the wake path's fixed budgets (20s hydration,
~15s resume retries) raced the whole boot and lost — "errors waking up BOTS"
while the backend came up healthy moments later.

Rather than raising timeouts, make the wake cheap:

- waitForFocusedSessionHydration: a history-bearing chat is hydrated when
  the persisted transcript is PAINTED on the right session. The REST
  prefetch delivers that seconds after the backend's HTTP is up; the full
  runtime resume (agent build, MCP discovery, 114-skill load) keeps warming
  in the background and binds the composer when it lands. Only an
  expected-empty chat still waits for the runtime (nothing to paint).
- On hydration timeout, log a [bot-wake] phase breakdown (activation ms,
  hydration ms, which conditions were unmet) to the renderer console so the
  next support bundle pinpoints the slow phase directly.
- web_server: flush the headless "listening" line — block-buffered on the
  Desktop's piped stdout, it surfaced minutes late and made boots look far
  slower than they were in support bundles (the 120s "gap" in this bundle
  was partly this artifact).

Sabotage-proven: restoring the runtime-gated wait fails the new paint-first
test by timing out — the exact field shape.
2026-08-18 15:52:39 -07:00
Teknium 052fe7240a fix(providers): plugin-profile fallback skips endpoint-less placeholder profiles
CI slice 11 caught a regression from the salvaged get_provider() fallback:
the "custom" placeholder profile (aliases ollama/local/vllm, empty
base_url) now resolved as a bare ProviderDef before
resolve_provider_full() reached its custom_providers step, collapsing
keyed IDs like custom:local-127.0.0.1:11434 to an endpoint-less "custom"
(test_keyed_custom_provider_bare_custom_fallback_uses_stable_key).

Gate the fallback on the profile carrying a concrete base_url —
placeholder profiles completed by config.yaml keep flowing to the
custom-provider resolution path.
2026-08-18 14:27:36 -07:00
greyvito 7072fc4f87 fix(providers): resolve plugin-registered provider profiles in get_provider
Plugin-only providers (commandcode, tencent-tokenhub, ...) are absent from
models.dev and HERMES_OVERLAYS, so resolve_provider_full returned None and
/model switches failed with "Unknown provider ..." even though the picker
lists them (CANONICAL_PROVIDERS auto-extends from the same registry).

Fall back to providers.get_provider_profile() before giving up, mapping the
profile api_mode to the ProviderDef transport.
2026-08-18 14:27:36 -07:00
Teknium 72b7c6c8d1 fix(models): keep discovery sentinels out of the user-facing models mapping
PR #67934 marked auto-discovered catalogs by writing two sentinel keys
INSIDE the user-facing ``models`` mapping of custom provider entries:
``__discovered_model_catalog__`` (written by
_save_discovered_models_to_config) and ``__explicit_model_allowlist__``
(injected by _normalize_custom_provider_entry). Every consumer of that
mapping — pickers, selectors, gateway/agent readers, and the user's own
config.yaml — had to know to filter those keys, and any site that
didn't listed them as phantom model IDs (``__discovered_model_catalog__``
showing up as a selectable "model"). The v11→v12 config migration and
the ACP session-state test caught exactly that leak on main.

Replace the in-mapping sentinels with a single entry-level flag:

- ``models_discovered: true`` now sits next to ``models``/``base_url``
  on the provider entry; the models mapping stays a clean
  ``{model_id: metadata}`` dict with no reserved keys.
- _save_discovered_models_to_config writes the new shape and refreshes
  catalogs it previously discovered (entry-level flag or legacy
  sentinel) instead of treating them as user-curated metadata.
- _normalize_custom_provider_entry no longer injects
  ``__explicit_model_allowlist__``; a dict-shaped models mapping counts
  as an explicit allowlist exactly when the entry is NOT marked
  models_discovered.
- _models_config_is_allowlist takes the discovered flag as a parameter
  (new helper _entry_models_discovered resolves it, including the
  legacy in-mapping sentinel); all call sites updated
  (model_switch.py, model_setup_flows.py, acp_adapter/server.py).
- Backward compat, no config version bump: configs written by a
  pre-fix Hermes (sentinels inside models) still read correctly —
  ``__discovered_model_catalog__: true`` is treated as
  models_discovered, both sentinel keys are stripped from model
  listings, and the next discovery save migrates the entry to the
  clean shape. Covered by a new regression test.

Also restore ``except Exception:`` on the pre-existing guards this PR
had narrowed to specific exception tuples (the resolve_runtime_provider
fallback in switch_model, the picker discovery/cache guards in
list_authenticated_providers, _get_model_config_dict, and
_credential_fingerprint). Those guards were intentionally broad on
main — a failed resolution or probe must degrade to the fallback path,
never crash the model switch. Guards the PR introduced for its own new
probe code keep their authored tuples.

The ACP new_session payload also goes back to
probe_current_custom_provider=False, matching the contract main's
test_new_session_returns_authenticated_cross_provider_model_state pins
(session opens must not block on live-probing the current custom
endpoint).
2026-08-18 14:26:16 -07:00
Vadelma 015f9990c6 fix(models): persist named custom catalog identity
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 7fb6b28ec8 fix(models): complete selector parity safeguards
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 81e813507f fix(models): preserve empty catalog and custom model semantics
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 9d2ba6c655 fix(acp): isolate custom catalog identities
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 7e61bd38c6 fix(models): preserve discovery provenance and endpoint identity
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma f70e9abce9 fix(models): close final provider discovery edge cases
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma fa1bb88e3e feat(models): propagate native discovery across selectors
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Vadelma 638b72f5ef feat(models): add native Ollama catalog semantics
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 14:26:16 -07:00
Teknium cc421cb697 fix: dashboard console skills commands no longer act on the wrong profile (#65828)
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.

Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.

Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.

Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root

Fixes #65828
2026-08-18 14:14:42 -07:00
xxxigm 696fef6214 fix(skills): keep inspect/install from mixing same-named hub skills
ClawHub treated the last path segment as a slug, so a GitHub-style
id like owner/repo/skills/skillopt fetched a different author's
skillopt. Pair metadata and files from the same source so inspect
cannot show one registry's header and another's SKILL.md.
2026-08-19 01:23:17 +05:30