Commit Graph

1498 Commits

Author SHA1 Message Date
Teknium 1fa66f2577 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	website/docs/user-guide/configuration.md
2026-08-19 19:36:12 -07:00
Teknium 26da56fd53 docs: tool provider selection follows the hermes tools pick (post #90317) 2026-08-19 19:25:50 -07:00
Jeffrey Quesnelle 612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
Teknium 095f003377 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	hermes_cli/tools_config.py
2026-08-19 16:48:40 -07:00
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium 803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
Teknium 09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00
Teknium 4d87290d39 feat: keyless web traffic splits 50/50 between Exa and Parallel like opencode
Unpinned zero-credential installs now pick Exa or Parallel by the
parity of the per-process random session id (stable within a process,
even split fleet-wide) instead of always favoring Parallel. An explicit
hermes tools selection (web.backend / per-capability keys) bypasses the
split entirely; the runner-up vendor stays in the walk as fallback.

Live E2E: 6 fresh processes split 3/3 between vendors, each performed
a real keyless search via its picked endpoint; explicit pin verified.
2026-08-19 16:22:30 -07:00
Teknium d762ed9b3c feat: execution-discipline guidance now reaches all tool-capable models (config model.execution_guidance)
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.

Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.

The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
  file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
  format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
  broader retry before concluding)
- completion gated on verification (done = every named acceptance
  criterion verified, never a plausible subset)

The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.

Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.

Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).

Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
2026-08-19 16:16:37 -07:00
Teknium f08d3e400f feat: hermes tools lets Exa/Parallel users pick the free keyless or paid keyed endpoint
Exa and Parallel now each render as two picker rows in hermes tools —
'Free (keyless)' and 'Paid (API key)'. Selection persists to
web.provider_tier.<name>:
- free: always the anonymous public endpoint, even with a key set
- paid: always the keyed SDK path; missing key errors instead of
  silently downgrading to the free tier (is_keyless_available also
  returns False so the auto-fallback walk can't route there)
- unset: auto (key present -> paid, else keyless)

Mechanism: get_setup_schema() gains a 'variants' list the picker
flattens into sibling rows sharing one web_backend; selection writes
the tier via both _write_provider_config sites; active-row detection
matches the tier (auto mirrors use_keyless). Routing goes through a
single use_keyless() chokepoint shared by search+extract in both
providers.

Live E2E: tier=free with a fake key present searched keyless OK (a
keyed call would have 401'd); tier=paid without key errored naming
PARALLEL_API_KEY; picker rows verified for both vendors x both tiers.
2026-08-19 15:36:20 -07:00
Teknium 2d9dad0bae docs: correct Exa keyless rate-limit characterization
A 12-request sequential burst from the same IP that earlier saw the
free-tier rate-limit error went 12/12 OK — the limit is a transient
burst/load control, not a tight standing per-IP quota. Soften the docs
and setup-schema wording accordingly (opencode users hit Exa keyless
as their default path in practice without throttling).
2026-08-19 15:24:27 -07:00
Teknium 96c2fd3c04 feat: web search/extract now work keyless on fresh installs via Parallel + Exa free tiers
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).

- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
  mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
  errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
  never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
  absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
  after every keyed/importable candidate; check_web_api_key() lights
  the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md

E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
2026-08-19 15:15:27 -07:00
Teknium aba96d5251 feat(image-gen): route live-catalog models to the Image API; merge picker catalogs; docs
Follow-ups on top of the salvaged #82631 surface:

- _select_surface: an unknown model id found in the live /images/models
  catalog now ROUTES to the dedicated Image API instead of only logging a
  hint — without this, a model picked from the live picker that postdates
  the curated snapshot would fall onto chat-completions and fail. Curated
  defaults stay pinned to chat (no behaviour change for existing setups);
  offline probes still fall back to chat. _HINTED_MODELS removed.
- list_models (OpenRouter): union of the live GET /images/models catalog
  (43 models today) and the chat-completions image models, deduped,
  defaults first; curated metadata wins for known ids, API names for the
  rest. Nous Portal (no /images route) keeps its chat-only catalog.
  Offline fallback: static chain + curated Image API snapshot.
- Tests updated/added: unknown-id routing (flipped from the hint-only
  pinning test), non-catalog id stays on chat, merged-picker union/dedupe/
  order, Nous exclusion.
- Docs: image-generation.md gains the OpenRouter Image API section and an
  editing-support row.

Live-verified: picker lists 43 models; generation succeeded through the
dedicated API on google/gemini-3.1-flash-lite-image and on the previously
unreachable black-forest-labs/flux.2-klein-4b (config-selected, no kwarg).
2026-08-19 14:44:36 -07:00
Alex Fournier 31402f630b fix(relay): complete native plugin cutover
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Teknium d23a4875aa docs: cover hermes chat --query-file and the file-based Bot Mode DM transport
Follows PR #89762: cli-commands reference and CLI guide document the new
--query-file flag; bot-mode.md's DM and peer-dm recipes now show the
file/stdin transport instead of inlining message bodies into the shell.
2026-08-18 23:02:48 -07:00
David Dudok de Wit 9ed738fa0c fix(desktop): stabilize gateway settings interactions 2026-08-18 22:03:25 -07:00
David Dudok de Wit 72fc5ab7ca fix(desktop): use gateway terminology in connection UI 2026-08-18 22:03:25 -07:00
David Dudok de Wit 2228bc7d88 feat(desktop): remember the last used source 2026-08-18 22:03:25 -07:00
David Dudok de Wit ccd30f955c feat(desktop): distinguish sources from profiles 2026-08-18 22:03:25 -07:00
David Dudok de Wit 4ea516b8d4 docs(desktop): explain multi-source session scoping 2026-08-18 22:03:25 -07:00
rob-maron 0b588cb3a4 MCP CIMD auth 2026-08-18 20:03:18 -07:00
Teknium 0c5f195ee2 docs: document one-click plugin install links (hermes://plugin/install)
The deeplink-driven plugin install flow shipped in #89464 (salvage of
#82735 by @serefyarar) had no docs. Adds:

- user-guide/features/plugins.md: "One-click install links (Desktop)"
  section under Managing plugins — link forms (repo/enable/force), the
  confirm-first dialog contract (never auto-installs, same install-time
  security scanning as the CLI), hybrid-repo behavior, legacy
  plugin-agent/plugin-desktop routing, hermes-dev:// in dev builds, and
  the no-SDK anchor example. Cross-links the MCP "Add to Hermes link"
  equivalent.
- developer-guide/desktop-plugin-sdk.md: "Distributing with an install
  link" section so plugin authors find the link form next to the
  packaging docs.
2026-08-18 16:42:04 -07:00
Teknium a77ee88ce2 feat: Bot Mode avatars default to deterministic blob faces drawn from the agent's name
New agents now get a blobatar — a deterministic soft-body face generated
from the bot's name (same name, same face, forever) — as the default
shapes mode, with full manual control:

- Face follows the name live while typing in New Agent
- Randomize re-rolls the seed; Lock face pins the current one so a later
  rename can't change it (Unlock returns to name-following)
- Any of the six silhouettes (round/organic/boxy/nub/cloud/sun) can be
  pinned via frozen-per-major trait positions while the rest stays
  name-derived
- Classic geometric shapes remain one click away, and existing bots keep
  their stored looks untouched

Wiring: blobatar@0.2.0 (zero deps, ~3.7KB) exported through the plugin
SDK (blobatarSvg / Blobatar), feature-detected in plugin.js with a
legacy-shape fallback for older desktops. Blob shape strings are
'blobatar[:seed[:kind]]' inside the existing meta.shape field, so
persistence, cross-machine ui_meta sync, and the roster's PNG backfill
(data-bot-face tag preserved) all work unchanged.
2026-08-18 15:54:24 -07:00
Teknium c820a5d383 docs(teams-pipeline): document fetch --organizer-user-id
Follow-up to #89382: the operator runbook, bundled-skill docs page, and the
bundled SKILL.md now cover the organizer-scoped lookup flag and note that
/meet/ short URLs require it while webhook jobs derive the organizer
automatically.
2026-08-18 13:03:16 -07:00
Teknium d8e2386912 feat: Capabilities view configures the selected profile on its own gateway
A profile belongs to one gateway, but the Capabilities surface (Skills /
Tools / MCP) always read and wrote through the window's active backend —
scoping to a remote-owned profile silently edited the wrong machine.

- hermes.ts: capability REST helpers accept a ProfileScope
  (string | {connectionId, profile}); ambient path now also carries the
  active registry connection tag (same contract as the cron helpers,
  #87882); profileScopeKey namespaces cache keys per connection.
- SkillsView: scope selector lists (profile, device) rows from the union
  agent roster on multi-connection desktops; new fixedConnection prop
  pins the whole view to a registered connection (plugin door), with a
  probe-able SkillsView.supportsFixedConnection flag.
- MCP tab: live reload.mcp RPC withheld for cross-backend scopes (it
  rides the active gateway socket and would reload the wrong machine).
- Bot Mode: remote-target drafts now get the live Capabilities tab
  pinned to the target machine via fixedConnection, feature-detected so
  older desktops keep the staged checklists.
- Config-record/hub-action stores accept scopes; cache keys fold in the
  connection id so two gateways' same-named profiles never share rows.
2026-08-18 11:31:43 -07:00
Teknium a1682376ca feat(profiles): rename any agent — the default profile gets a display name (#45624)
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.

Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").

Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
2026-08-18 02:27:18 -07:00
Goktug Vatandas 9bea439189 feat(bot-mode): support multiple groups per bot 2026-08-18 00:26:31 -07:00
teknium1 8ce8ffd429 fix(update): 'hermes update' no longer claims success on a parked feature branch — switches back when safe, warns loudly when not
Live incident 2026-08-17: the source checkout was parked on a stale feature
branch (claude-code-inspired/local-terminal-memory-limit, days behind main),
left there by earlier tooling. 'hermes update' autostashed, refreshed lazy
backends, synced skills, and printed '✓ Code updated!' / '✓ Update complete!'
while the checkout stayed on the stale branch with none of main's new code.
Two sessions burned time on 'the fix is missing' confusion.

- Parked-branch guard: auto-switch back to the update target ONLY when the
  parked branch is clean and fully merged (git cherry origin/<target> shows
  nothing unmerged); the checkout then STAYS on the target instead of being
  re-parked. Otherwise: loud CODE UPDATE SKIPPED block naming the branch,
  behind-count, and resolution commands; exit 1; branch untouched.
- The up-to-date (commit_count == 0) path no longer switches back to a
  fully-merged parked branch either.
- Post-pull gate additionally refuses to print '✓ Code updated!' when HEAD
  ends up attached to a non-target branch.
- Summary lines now carry the actual branch + HEAD short-sha:
  '✓ Update complete! [main @ 30fcf9580]' — drift visible at a glance.
- New config toggle updates.auto_switch_parked_branch (default true).
- Real-git-fixture regression tests (init/clone/branch, no subprocess
  mocks): clean+merged auto-switch, dirty skip, unmerged skip, cherry-picked
  equivalence, config opt-out, unverifiable ref, on-main fast path,
  up-to-date no-repark, summary branch/sha assertions.
2026-08-17 23:53:05 -07:00
Teknium 6170f844c4 fix(desktop): remove profile scoping from the Gateways settings page
The unified Gateways settings page (from the recent settings merge) still
carried the legacy per-profile gateway-override machinery: an "Applies to"
profile-chip scope switcher, a scope state machine threaded through load/
save/test/sign-in paths, inherit-mode ModeCard variants, and an SSH
remote-profile mapping row.

The page is machine-level gateway management: it decides which gateway
backends this desktop can connect to, and profiles are discovered FROM the
connected gateways. It must not be profile-scoped.

- Delete the scope chips section, ScopeChip component, and the scope/setScope
  state; every scope-conditional collapses to its global (scope === null)
  branch. getConnectionConfig/save/apply/test/sign-in are all unscoped now.
- ModeCard local card always renders the local title/desc (inherit variants
  gone); SSH remote-profile mapping row removed.
- i18n: drop now-unused gateway keys (appliesTo, allProfiles,
  defaultConnection, profileConnection, inheritTitle, inheritDesc,
  sshRemoteProfileTitle, sshRemoteProfileDesc) from types.ts and en/zh/
  zh-hant/ja/ar in sync; rewrite the gateway intro in each locale to say
  connections are machine-level and profiles come from gateways.
- Tests: replace the scope-switching component tests with a machine-level
  assertion (loads getConnectionConfig(null), never a profile scope, no
  scope UI rendered).
- Docs: update desktop.md and multi-connection-desktop.md wording — gateway
  connections are machine-level; per-profile backend routing continues via
  the profile rail / session source surfaces, not the settings page.

The electron main-process per-profile override mechanism
(getConnectionConfig(profileName), route map) and the profile-rail connect
flows are intentionally untouched; only the settings page loses the
affordance.
2026-08-17 20:39:02 -07:00
Teknium d127b27303 docs(desktop): document the tabbed SESSIONS|BOTS sidebar, Bots-mode-only Cronjobs pane, per-bot Hide/Unhide, and host.paneVisibility (#88788, #88800) 2026-08-17 19:22:47 -07:00
Teknium bc76f62c20 feat(cron): configurable media-send timeout + non-empty failure reasons
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):

- Promote the media-send timeout to the standard resolution pattern:
  HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
  cron.media_send_timeout_seconds in config.yaml, then 300s default
  (mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
  (environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
  in delivery_errors (post-#88631 the reason reaches the run status, not
  just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
2026-08-17 17:51:08 -07:00
Teknium 24f7f9a9da docs: reflect the unified Gateways page, settings profile scope, plugins cleanup, Bot Mode group rows, and host.openWorkspace
Update the desktop docs for five just-merged desktop changes:

- Settings → Gateway + Settings → Connections are now one "Gateways" page:
  retitle every reference, describe the Add-connection flow's four kinds
  (Local / Hermes Cloud / Remote gateway / SSH) and the save-time duplicate
  rules (one local; URL-normalized dedupe across remote/cloud; user@host:port
  + remote profile for SSH), and describe the "Per-profile overrides"
  subsection that replaced the page-level Applies to chip row.
- Document the shared "Applies to" profile scope on the config-backed
  settings pages (Model, Workspace, Safety, Memory & Context, Voice, Chat,
  Advanced, Tools & Keys) and the Messaging overlay.
- Agent plugins section: bundled built-ins are hidden (user/git/project/
  pip/portable installs only), Example Plugin is gone, and the section has
  its own Applies to selector backed by plugins.manage's optional profile
  param.
- Bot Mode: group chats are standalone Discord-style roster rows and open
  in the main chat window (older builds fall back to the in-panel view).
- Desktop Plugin SDK: document the new host.openWorkspace(id, { render,
  title, minWidth, onClose }) door, its refresh/re-front semantics, and
  the feature-detection fallback pattern.

Also retitles the Settings → Gateway references in the web-dashboard guide.
No new pages; sidebars.ts unchanged. `npx docusaurus build` passes.
2026-08-17 17:23:25 -07:00
Teknium 3360590115 fix: address second-round SkillEvaluator review feedback
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):

- Incomplete-validator findings are now PRESERVED as partial evidence;
  only the validator's pass/fail verdict is excluded from the advisory
  verdict. A report with findings from an incomplete check no longer
  reads as clean.
- Clean-report wording is now "no findings from completed checks"
  whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
  config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
  discard-pinning test, plus the completed-checks wording case.
2026-08-17 17:04:40 -07:00
Teknium 2c2697b52e feat: widen Tier 1 advisory scan to license + security checks
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.

- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
  static (no measurable cost); security invokes NVIDIA SkillSpector in
  its keyless static-rules mode (~+1.2s per install). schema/quality
  stay excluded: hygiene signal ("author not specified" is
  high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
  or failing, the security check reports status="incomplete" and the
  adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
  note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
  This also absorbs a live upstream inconsistency: SkillEvaluator's
  anti-tamper cross-check on SkillSpector's risk score currently trips
  on moderate-finding skills (fail verdict with zero findings, e.g.
  github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
  NVIDIA separately; either way an evidence-free fail must not render
  as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
  not-run formatting.

E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
2026-08-17 17:04:40 -07:00
Teknium 183f18d530 feat: advisory NVIDIA SkillEvaluator Tier 1 scan on skill installs
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).

- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
  over the quarantined bundle, parses the JSON report, classifies
  secrets-class findings (private keys, tokens, credentialed connection
  strings) apart from advisory PII findings. Every failure mode
  (binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
  guard's policy decision and before the install confirmation. Findings
  are shown with file:line; secrets-class findings render red with a
  loud warning. Warn-and-continue by design — the built-in skills guard
  remains the only enforcement layer, because the upstream PII scanner
  has known false-positive classes (git@github.com, docs example
  emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
  endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
  optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
  config toggle.

Scanner install (optional):
  uv tool install --python 3.13 \
    "skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"

E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
2026-08-17 17:04:40 -07:00
Teknium 9aa1413781 feat(desktop): extend kanban native notifications to blocker/failure events
Builds on @nductien's completion-notify module (PR #87705):

- Notify on the gateway watcher's full terminal set — blocked, gave_up,
  crashed, timed_out, block_loop_detected — not just completed. A worker
  hitting a blocker while the user is away was the original community ask.
- Route all notification copy through the kanban plugin i18n bundles
  (en/ja/zh/zh-hant), with an English-bundle fallback when the translator
  isn't bound yet.
- Wire the ctx.os.notify door so events also fire a NATIVE OS notification
  while the user is away from the Hermes window (host.notify toast covers
  the foreground). OS-door failures are isolated from the toast path.
- Docs: Desktop notifications section in kanban.md, including the
  app-running coverage window.
2026-08-17 17:00:00 -07:00
Teknium 6229683b62 feat(cli): hermes peer — bot-to-bot DMs across machines and gateways
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.

- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
  `dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
  (list by title, create when missing), runs one synchronous agent turn via
  POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
  cross-machine twin of the local bot-messaging command, so the Bot Mode
  protocol composes over it unchanged. Named profiles route via the peer's
  /p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
  (`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
  ~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
  messaging protocol gains a cross-machine paragraph (peer roster +
  `hermes peer dm` pattern) so agents discover remote teammates on their
  own; peers join the capability fingerprint so registering/removing one
  refreshes eternal Bot Chat prompts on the next message (loud, one-time,
  user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
  reference (`hermes peer` section + summary row).

Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
2026-08-17 16:13:30 -07:00
Teknium 66221397a1 fix(bot-mode): always hide Bot Mode sessions from the global Sessions sidebar
Bot Mode's group chats spawned one per-member session per room, and those
"Group: ..." rows (plus canonical Bot Chats when the old eye-toggle pref was
off) flooded the global Sessions sidebar — a 6-bot room dumped six identical
rows into recents (reported with screenshot, Aug 17).

Plugin (apps/desktop/src/plugins/hermes-bots/plugin.js):
- session.create now passes hidden:true UNCONDITIONALLY for both canonical
  Bot Chats and group-room member sessions; the $hideBotChats pref, its eye
  toggle, and its storage hydrate are removed (Bot Mode sessions are plumbing
  or plugin-owned forever-chats, never scratch conversations).
- hideOwnedBotSessions(): idempotent reconciliation sweep over every owned
  session id (bot meta canonical chats + each room's member sessions) via
  session.set_hidden, run on plugin load and on each gateway reconnect, so
  rows born visible under the old pref get cleaned up.
- The Bots session browser and canonical-chat recovery scan pass
  include_hidden:true so they still see the rows they own.

Gateway (tui_gateway/methods_session.py):
- session.list honors an include_hidden param (default off — the resume
  picker and all global callers keep dropping hidden rows).
- session.set_hidden gains a durable fallback: when no LIVE runtime session
  matches, resolve the stored session id in the target profile's state.db
  (via resolve_session_id) and flip the flag there. The sweep holds stored
  ids for chats that aren't live; the old live-only lookup 4001'd them.

Validated E2E with real imports against a temp HERMES_HOME: born-hidden row
(hidden=1), profile-scoped session.list default vs include_hidden (0 vs 1),
and stored-id sweep on a non-live legacy row (hidden=1). Plugin suite
167/167; new RPC regression tests in tests/tui_gateway/test_session_hidden_rpc.py.
2026-08-17 16:02:42 -07:00
Teknium bf53915331 docs: Bot Mode guide — Create-on picker, cross-machine mentions + group chats
Community asks (Discord, Aug 17): unclear what happens across cloud vs
desktop, whether every bot replies in group chats, and how to persist
connections to multiple gateways.

Extends the Bot Mode user-guide page (landed on main today) with the new
cross-connection features:

- "Create on" picker: creating an agent on another registered machine, with
  the remote-target caveats (clone source, staged capability checklists,
  draft discard).
- Group chats: explicit "not every bot replies" explanation of the
  round-robin/pass model, and rooms spanning machines with device badges.
- @mentions across machines via the Connections registry (no gateway switch).
- Bots-across-machines section: persistent SSH inventory, last-known rows,
  and the stay-in-your-chat interaction model; cloud+desktop recipe.
- desktop.md Bot Mode section links to the full guide; multi-connection
  page's Bot Mode reference points at the docs page instead of the old
  standalone repo.
2026-08-17 14:32:52 -07:00
witcheer 851a30d0ab docs: add a Bot Mode user-guide page
Bot Mode ships built into the desktop app (default on) but only had a
short section in desktop.md. This adds a dedicated user-guide page
covering the Bots roster, creating and editing Bots, avatars, routines,
group chats, bot-to-bot messaging (agent.bot_mode_protocol), the
multi-connection roster, and CLI parity.
2026-08-17 14:08:38 -07:00
Teknium 6e22d26583 feat: project-skill quarantine + non-interactive trust inheritance
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.

Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).

Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.

Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
2026-08-17 14:06:16 -07:00
Teknium 481156139d feat: misfire catch-up for external cron providers
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).

fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:

- No-op for the built-in ticker (its tick loop already self-heals
  past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
  own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
  external retry is de-duplicated) and runs fire_claimed in a daemon
  thread, mirroring the webhook admission pattern, so housekeeping
  never blocks for the length of an agent run. Provider re-arm logic
  (Chronos NAS one-shots) runs exactly as for a normal fire.

Docs: cron.md section + cron.misfire_grace_minutes reference.
2026-08-17 11:42:25 -07:00
Teknium f891d702df feat: project-local skill discovery with per-repo trust gate
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.

Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.

- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
  get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
  project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
  [project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
  skill_view resolves cross-tier collisions in favor of the project tier
  (same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
  commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)

Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
2026-08-17 11:39:13 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
kshitij 97c4f9eeec test(delegate): assert the unproven-state contract, not its prose
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.

- The distinguishability test asserted the failure payload was byte-identical
  to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
  That freezes the AMBIGUITY as a required property: emitting `commits: None`
  for "unknown" would improve exactly what #88113 is about and fail the test.
  Now asserts what the parent actually depends on -- both keep the worktree,
  and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
  happy path, forbidding an always-present-but-False flag (a legitimately
  better JSON contract: stable key set for serializers). Now
  `assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
  of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
  was not even a cross-producer contract: delegate_tool's note said "state
  unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
  assert the note names the worktree AND branch -- the actionable part for a
  human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
  before any git call would keep it green while proving nothing). Now checks
  `call_count` and mirrors the branch-survival + note-names-path legs its
  sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
  parent reads, but delegate_tool's fallback -- the second producer -- was
  verified only by reading. It now AST-parses the real fallback dict literal
  and compares against live `finalize_subagent_worktree()` output, so the two
  producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
  commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
  raising, handled in delegate_tool), and the module docstring listed
  `inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
  `_break_git_index()` beside the file's other module-level helpers.

Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
2026-08-17 19:41:32 +05:30
kshitij 38ea711fd0 fix(delegate): tell the parent when a worktree was preserved un-inspected
The preserved worktree is invisible to the only consumer that can act on it.

Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":

  inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
  inspected OK, child produced nothing     -> {commits: 0, dirty: False, pruned: False}

The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.

Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
  naming the worktree/branch, warns, and returns the payload. Both unproven
  exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
  non-numeric rev-list stdout) produced the same unproven payload but logged at
  DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
  exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
  creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
  schema missing commits/dirty/pruned. It now emits the same flagged shape, and
  logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
  affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
  by restoring the unconditional prune and reintroducing this P1.

Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.

Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
  suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
  the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
  work intact on disk; proven-clean still prunes (pruned=true).
2026-08-17 19:41:32 +05:30
kshitij eb4bc1513f fix(telegram): log first-choice IPv4 stick as info, not warning
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
2026-08-17 17:13:27 +05:30