Commit Graph

34979 Commits

Author SHA1 Message Date
teknium1 7ee9fa52c0 test(stt): correct the de-flake provenance in the idle-timeout comment
andrexibiza's review is right: the base test already printed tick 0
before its first sleep, so this change never removed a
sleep-before-first-tick race. The actual de-flake is the larger idle
budget (0.1s -> 0.25s) plus a longer heartbeat sequence whose ~400ms
runtime still exceeds the idle window. Say exactly that so the causal
record is accurate.
2026-09-15 03:36:11 -07:00
andrexibiza a70ea873cf test(stt): stabilize idle-timeout progress test against spawn latency
The stderr-progress idle-timeout test used a 0.1s idle window with 0.04s
ticks — shorter than Windows process spawn, so the first chunk could never
arrive in time (deterministic failure on Windows, flake under Linux CI
load). Verified failing identically on pristine main before the change.

Fix: emit the first tick immediately, tick every 50ms for ~400ms total,
250ms idle window (5x tick period). The pass still depends on the progress
extension while tolerating real spawn/scheduling latency.

Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
2026-09-15 03:36:11 -07:00
teknium1 b00ebf6210 fix(skills): live-dashboard credits the human author and drops the product-name intro
Authoring standard 4 requires the human first in `author`; the skill was
drafted with Hermes so the tool was credited instead. The intro cited a
third-party product, which is allowed only in LICENSE/credit lines. Also
adds metadata.hermes.category and lowercases tags to match sibling
skills; docs page regenerated for this skill only.

Tests: the two prose tests asserted sentence literals (change-detectors);
they now assert structure — three Procedure phases, a "Done when" per step,
standard headings, and the tool wiring (cronjob/desktop_preview/[SILENT]).
2026-09-15 03:35:33 -07:00
teknium1 788601358d refactor(skills): live-dashboard becomes an optional skill reconfigured for the Desktop app
Why: the web dashboard is being deprecated in favour of the Electron Desktop
app, and a skill that ships a bespoke cron blueprint should not be bundled by
default. Reconfigure instead of just rebasing:

- Move skills/productivity/live-dashboard -> optional-skills/productivity/
  live-dashboard (install with `hermes skills install
  official/productivity/live-dashboard`); register it like every other
  optional skill: per-skill docs page under user-guide/skills/optional/,
  optional-skills-catalog row, sidebars entry.
- Desktop reality: add a "Show the dashboard" step — when `desktop_preview`
  is in the toolset (Desktop/GUI sessions) render index.html in the in-app
  preview pane after every build/tick and on request; otherwise report the
  absolute file path. Prerequisites section added (Enough1122 review).
- Drop the hard-wired cron/blueprint_catalog.py entry: the curated catalog
  is for bundled skills and would preload a skill that may not be
  installed. Use the skills-pipeline blueprint instead —
  `metadata.hermes.blueprint` on the SKILL.md registers a daily
  all-dashboards sweep as a /suggestions entry at install time (opt-in,
  never auto-scheduled), which is exactly the mechanism main provides for
  optional skills.
- Never hardcode ~/.hermes in prose the agent executes: refer to the Hermes
  home directory's dashboards/<slug>/ and write absolute paths into cron
  prompts (also answers the review's "tick prompt must name the state-file
  path" point).
- Tests follow the skill to optional-skills/, the catalog-blueprint tests
  are replaced by one parse_blueprint/blueprint_to_job_spec invariant and
  one desktop_preview-with-path-fallback invariant.
2026-09-15 03:35:33 -07:00
Teknium 2154458685 Inspired by Energy: one-sentence live dashboards — bundled skill + automation blueprint
Energy (getenergy.com) ships natural-language persistent dashboards:
describe what you want to see in one sentence and the agent builds a
self-updating status page fed by email threads, signed-in websites, and
files. This ports the concept onto Hermes's existing cron + connector
architecture:

- skills/productivity/live-dashboard: setup/tick split skill — pin the
  dashboard contract, verify one live read per source before scheduling,
  keep dashboard.json as source of truth with a self-contained HTML
  projection, stale-read discipline, deliver only on material change.
- cron/blueprint_catalog.py: live-dashboard automation blueprint
  (purpose/sources/time/recurrence/deliver slots) rendering to the
  dashboard form, /blueprint command, and hermes:// deep-link.
- tests/skills/test_live_dashboard_skill.py: skill standards + blueprint
  registration + real fill_blueprint E2E.
- docs: per-skill page, skills catalog row, sidebar entry.
2026-09-15 03:35:33 -07:00
teknium1 59c20a51aa test: drop product name from inbox-triage test docstring
Competitor product names are allowed only in LICENSE attribution and
salvage credit lines; the test docstring had an inspiration tag left over
from the port.
2026-09-15 03:34:55 -07:00
teknium1 0a208b6891 docs(skills): itemize inbox-triage voice calibration and bound its context cost
Why: the calibration step was a single ~120-word paragraph that models follow
less reliably than enumerable rules, and 20-50 full sent messages could
crowd inbox coverage out of context (Enough1122 review). Split sampling /
extraction / record / fallback into bullets, state that truncated excerpts
carry the style facts, add the Sent-folder-naming pitfall (`Sent`, `Sent
Messages`, `[Gmail]/Sent Mail`, localized) so a missing folder name does not
silently trigger the fallback, and compress the provenance parenthetical to
the operative fact. Docs page mirrors the SKILL.md body verbatim.
2026-09-15 03:34:55 -07:00
Teknium 7f2bf68308 Inspired by Energy: evidence-based voice calibration for inbox-triage reply drafts
Energy's cross-app reply agent analyzes ~100 of the user's past replies
before drafting, so drafts land in the user's actual voice instead of
generic-professional AI register. Port the mechanism into the
email-inbox-triage skill's drafting step:

- Step 4 now calibrates on a bounded sample (20-50) of the user's own
  sent replies — greeting/sign-off habits, length, formality, rhythm,
  per-audience differences, how the user pushes back — before drafting,
  with an explicit fallback when Sent is empty or inaccessible.
- New pitfall + verification item pinning the calibration discipline.
- Test locks the evidence-based calibration and fallback language.
2026-09-15 03:34:55 -07:00
teknium1 286e723db8 fix(skills): auto-load resolves under the agent's own home and skips internal forks
Why: build_auto_load_prompt read config via ambient load_config_readonly()
and looked skills up under the ambient SKILLS_DIR. Gateway bot threads lose
the HERMES_HOME ContextVar, so a bot profile's pinned skills came from the
launch profile — the docs promise profile scoping. _auto_load_parts now
passes home_override=_agent_home(agent) and build_auto_load_prompt binds it
for config, disabled-list and <home>/skills lookup, the same seam
_skills_prompt uses via skills_dir_override.

_auto_load_parts was unconditional and injected pinned SKILL.md bytes into
delegate children, curator/background_review forks and gateway hygiene
agents; it now mirrors _skills_prompt's gate (nothing without the skills
toolset) and returns [] when skip_context_files is set.

cli.py's HERMES_IGNORE_RULES check used == "1" while system_prompt used
is_truthy_value; both use is_truthy_value now.

Tests stay at 4: the build test asserts the home-scoped resolution, the
ignore-rules test also covers the subagent / no-skills-toolset gates.
2026-09-15 03:34:20 -07:00
Carl Taylor 075a256597 fix(skills): auto-load resolves once per agent and dedupes against -s
The system prompt must stay byte-stable for the life of a conversation:
`_auto_load_skills_result` is seeded in `_SESSION_STATE` and filled on
the FIRST prompt build only (HERMES_IGNORE_RULES captured then too), so
model switches, compression and static-prefix restoration reuse the
exact rendered bytes rather than re-reading config or skill files.

CLI: auto_load renders in the existing background `--skills` preload
thread (real session id for ${HERMES_SESSION_ID}), `-s` names dedupe
against the auto-loaded canonical names via
`build_preloaded_skills_prompt(excluded_loaded_names=)`, the activated
skills line shows auto_load first, and the lazily built agent is seeded
with the pre-resolved bytes. `--ignore-rules` skips auto-load with the
rest of the auto-injected context.

Re-implementation of #74060 by @ctaylor86 against current main.
2026-09-15 03:34:20 -07:00
ArcherQAQ 1976869c01 feat(skills): skills.auto_load pins skills into every new session's prompt
`skills.auto_load: [name, ...]` in config.yaml renders the listed skills
as fully loaded skill blocks in the system prompt of every new agent —
CLI, TUI, gateway, cron and API — the persistent counterpart of `-s`.
Missing or operator-disabled names are warned about and skipped; a
config typo never blocks session start.

Re-implementation of #26840 by @ArcherQAQ (via #74060) against current
main — the original patch targeted the pre-decomposition cli.py /
system_prompt.py god files; design and diagnosis preserved. The loader
reuses `_load_skill_blocks` (same disabled gate and Curator usage bump
as `-s`) instead of a parallel loop.
2026-09-15 03:34:20 -07:00
teknium1 bbe13059bc fix: gate kanban.default_assignee through the dispatch_profiles allowlist
_resolve_default_assignee imported the raw profile_exists instead of the
allowlist-gated predicate returned by _profile_exists_fn(). On a shared
board, a home with dispatch_profiles: [sage] and default_assignee: default
therefore passed the gate for "default" (every home has that profile), wrote
assignee=default plus an `assigned` event onto an unassigned card it may not
claim, and only then bucketed it skipped_nonspawnable. That is a row write on
a card another home owns. Routing through _profile_exists_fn() makes an
out-of-allowlist default_assignee resolve to None, so the card is left
untouched; the unimportable-profiles fallback is unchanged.

Review finding: default_assignee bypassed the dispatch_profiles gate and persisted assignee + assigned event onto unclaimable shared-board cards.
2026-09-15 03:33:53 -07:00
teknium1 923960fa4b fix(kanban): dispatch_profiles is config-only and fail-closed; trim tests
Follow-up to the salvaged #111004 commit, aligning it with the shape agreed
on #110995:

- Drop the HERMES_KANBAN_DISPATCH_PROFILES env bridge: non-secret behaviour
  lives in config.yaml only, like every other kanban.* key.
- Read the key via load_config_readonly() with the same fail-open config
  read as the sibling kanban.* readers (configured_max_in_progress).
- Fail closed when the key is set: the "none" sentinel is gone (an empty
  list already claims nothing), and an assignee that is not a valid
  profile id is never claimable instead of being lower-cased into the
  allowlist.
- Trim the regression file to two invariants (allowlist without `default`
  buckets the card as nonspawnable AND turns has_spawnable_ready off;
  unset key keeps upstream behaviour). Both drive the real dispatch tick
  against a real config.yaml + kanban.db; the first is red on origin/main.
- Docs: move the "Shared boards across homes" section out of the
  gateway-dispatcher paragraph, state that `default` collides by
  construction, add the config-reference row.
2026-09-15 03:33:53 -07:00
Kevin Rajan 2d46af3fa2 fix(kanban): per-home dispatch claim allowlist for shared boards
On a shared kanban.db, every home's profile_exists('default') is
unconditionally True, so any home's dispatcher could claim cards assigned
to 'default'. Wrap the _profile_exists_fn() predicate with an optional
per-home allowlist: kanban.dispatch_profiles (config.yaml, list or
comma-separated string) with a HERMES_KANBAN_DISPATCH_PROFILES env bridge.
Unset preserves upstream behavior; 'none' claims nothing. Foreign
assignees land in the existing skipped_nonspawnable bucket, and the gate
applies to the ready spawn path, _has_spawnable, and review dispatch alike.
Also documents the multi-home default collision in the kanban user guide.

Fixes #110995
2026-09-15 03:33:53 -07:00
teknium1 931387dd42 test(cron): two invariant ESTOP tests for the fire webhook and misfire backstop; docs
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.

Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
2026-09-15 03:33:13 -07:00
JulianCruzet 05e13a1fcc fix(cron): address Xipong's review nits on #110563
Two inline test-hardening suggestions + two follow-up coverage cases
(Xipong called them non-blocking but they're cheap and prove the
boundaries).

- tests/gateway/test_api_server_jobs.py::TestCronFireEstop::
  - test_fire_webhook_returns_503_when_estop_engaged: also patch
    `cron.scheduler_provider.resolve_cron_scheduler` and assert that
    `claim_fire`, `fire_claimed`, and `fire_due` are never reached.
    A 503 by itself does not prove admission never happened.
  - test_fire_webhook_401_when_verifier_crashes (new): crashing
    verifier → 401, ESTOP is never consulted. Proves auth runs before
    the ESTOP check so the sentinel state cannot leak to unauth callers.

- tests/cron/test_misfire_catchup.py::TestFireOverdueJobs::
  - test_estop_engaged_skips_backstop: spy on `provider.claim_fire`
    and `cron.scheduler_provider.threading.Thread`. Deterministic
    proof that no claim was attempted and no worker thread was
    constructed (the prior `wait_fired(timeout=0.5)` was timing-based).
  - test_estop_release_restores_backstop: spy on `provider.claim_fire`
    on the recovery path with an explicit `assert_called_once` so a
    regression that fires without claiming still fails the test.

All 58 tests across the affected files pass.
2026-09-15 03:33:13 -07:00
JulianCruzet 95fd5d6816 fix(cron): honor ESTOP in NAS fire webhook and misfire backstop
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:

- Door 2 (NAS fire webhook `_handle_cron_fire`,
  `gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
  the JWT/drain guards and `provider.claim_fire`. Added a
  `check_paused("cron-webhook")` guard inside the reservation block,
  returning 503 + Retry-After so the NAS retries after `hermes resume`
  rather than silently dropping the run.

- Door 3 (misfire backstop `fire_overdue_jobs`,
  `cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
  the function. Added an early-return `check_paused("cron-misfire")`
  guard. Self-healing — the next sweep after `hermes resume` catches
  everything up via the existing claim_fire path.

Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.

Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.

Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
2026-09-15 03:33:13 -07:00
kshitijk4poor fef0bc56b2 refactor(cli): one notification payload, one sink per branch in _ring_bell
Review fold: build "\a" + sequence once and dispatch it to exactly one sink (app loop
when the Application runs, write_tty otherwise) instead of two branches with two writes
and two try/excepts. terminal_notify.notify() had a single caller (that fallback), so
write_tty is the public no-app entry and notify() is gone. The loop-side write keeps the
never-raises contract on a dead tty (EIO/closed file) the same way _pet_flush_kitty_frame
does, and _ring_bell skips entirely once _terminal_io_broken is set.

pty A/B re-run on this stack: 400 rings vs 90 frames, 0 aborted, 0 painted, 400/400 delivered.
2026-09-15 14:52:26 +05:30
kshitijk4poor 6b69d0f6fd fix(cli): /bg completion rings through _ring_bell like every other bell site
The side-worker (/bg, /btw, /login) hand-rolled its bell with a bare sys.stdout
'\a' and so never emitted the OSC 9 / OSC 777 desktop notification the other
end-of-work sites get. Route it through _ring_bell, which also owns the
bell_on_complete gate and the app-loop serialization.
2026-09-15 14:52:26 +05:30
kshitijk4poor ac2359cd35 fix(cli): turn-end notifications no longer paint the pet's kitty frame as base64
Symptom (Ghostty, display.pet on, display.bell_on_complete on): at the end of a turn the
input line fills with several rows of base64 and a stale copy of the status bar + pet
stays above the response panel.

Root cause: _ring_bell runs on the agent thread and terminal_notify wrote the OSC 9 /
OSC 777 sequence through its own open("/dev/tty") (or sys.stdout). The prompt_toolkit
loop thread may at that moment be mid-write of a 12 KB kitty APC pet frame, which the tty
drains ~1 KB at a time. The second writer splices into the frame; the foreign ESC aborts
the APC and the terminal paints the remainder of the payload as text at the input cursor.
The wrapped garbage scrolls the screen, so the panel that follows is printed against a
stale cursor position and the old chrome survives above it.

Change: when the CLI's Application is running, _ring_bell hands "\a" + the notification
sequence to the app loop (_run_on_app_loop -> _write_terminal_sequence), serializing it
behind the renderer and the after_render frame writer. terminal_notify.notify() keeps the
/dev/tty path for callers without a running app; the sequence builder is split out as
notification_sequence().

Verification: pty A/B with the real Application + after_render frame writer, 400 rings
vs 91 frames — base: 3 leaks (4,324 base64 chars painted); fixed: 0 leaks, 400/400
notifications delivered, 0 aborted frames.
2026-09-15 14:52:26 +05:30
kshitijk4poor 2932195c22 fix(skills): give the provider cut one owner inside the parallel walker
The dashboard endpoint GET /api/skills/hub/search passes its user-supplied
`source` straight into parallel_search_sources and never applied the merged
provider cut, so ?source=nvidia returned a mixed set. That was the fourth
caller of the walker; the cut was copy-pasted at three of them and missing
at the fourth.

parallel_search_sources already computes the normalized provider filter, so
the cut now lives there — applied per source before results are counted and
merged. Every caller (CLI search via unified_search, CLI browse, TUI-gateway
browse, dashboard router) sees the same rule with no provider logic of its
own, source_counts stop reporting rows that are then dropped, and the three
duplicated call-site cuts are deleted. do_browse keeps its provider-specific
"No skills found for provider" message.

Also:
- HermesIndexSource.search now treats a whitespace-only provider_filter as
  "no filter", matching GitHubSource.search (the two adapters previously
  disagreed on the same keyword argument; unreachable through the walker,
  which pre-normalizes).
- The regression-test fixture seeds tap caches by github_provider_for label
  instead of case-sensitive repo literals, and serializes metas through
  _skill_meta_to_dict, so a DEFAULT_TAPS casing change can no longer silently
  unseed the fixture.

Validation: 121 targeted tests green; disabling the walker cut fails the
pre-existing test_unified_search_provider_filter_keeps_index_source with the
expected clawhub leak; 4/4 regression cases still red on unpatched main.
2026-09-15 14:15:59 +05:30
kshitijk4poor 69d181011b refactor(skills): skip wrong-provider taps and give the provider-filter idiom one owner
Follow-up polish on the provider-filter-before-limit fix:

- GitHubSource.search now skips taps whose repo maps to a different provider
  instead of enumerating every tap and filtering afterwards. A tap's repo fixes
  the provider of every result it yields (github_provider_for is the only source
  of extra.provider in this adapter), so the skip is lossless and avoids up to 23
  useless tap enumerations per provider-filtered search — real GitHub API calls
  against the 60/hr unauthenticated budget and the 30s overall timeout when the
  index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
  it replaces the four inline copies of the strip/lower/membership idiom in
  _select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
  the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
  pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
  provider cut as do_browse; it accepted a provider value but returned
  unfiltered results.

Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
2026-09-15 14:15:59 +05:30
Danylo Borodchuk e75b5a2d38 fix(skills): filter providers before limiting search results 2026-09-15 14:15:59 +05:30
kshitijk4poor a55c972e09 refactor(gemini): prefix rule only for no-id slot matching; share the provider-id guard
Slot arguments are always complete json.dumps output (Gemini re-sends full args), so the mid-stream JSON check could never fire; the id key needs no tool name; _new_call_id and the slot lookup now share _provider_call_id.
2026-09-15 13:27:24 +05:30
jmiguellucas 91665b5e79 fix(gemini): give each native tool call its own streaming slot
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.

Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.

Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
2026-09-15 13:27:24 +05:30
kshitijk4poor 5900f0f70a chore: map jmiguellucas and cdbartholomew contributor emails 2026-09-15 13:27:24 +05:30
kshitijk4poor 3ba5602b62 chore: map the earlier RFC 9207 iss relay submitters
#88612 (willfrombr, Aug 17, closed when the author retired their fork) and
#92765 (CoLaOnline, Aug 23) carried the same dashboard/gateway iss relay
before #105610; credit them alongside the commit this stack cherry-picks.

Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
Co-authored-by: Colin Lateano <222331716+CoLaOnline@users.noreply.github.com>
2026-09-15 13:00:12 +05:30
kshitijk4poor 8bdac1b17a fix(desktop): omit a null iss from the oauth.callback relay; hoist the loopback parse import
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
2026-09-15 13:00:12 +05:30
kshitijk4poor 4a164a8073 refactor(tui_gateway): parse the loopback redirect with the shared helper
The gateway loopback handler re-inlined tools.mcp_oauth._parse_redirect_query;
that copy is exactly how the gateway relay lost `iss` while the CLI path
kept it. Use the helper so the four callback keys have one owner, and
point the docstring at it instead of repeating the RFC 9207 rationale.
2026-09-15 13:00:12 +05:30
kshitijk4poor 1c243f86de fix(tui_gateway): relay RFC 9207 iss through the oauth.callback RPC
The oauth.callback handler parsed `iss` but never passed it to deliver_callback_flow, and McpOauthCallbackParams (extra="forbid") had no `iss` field, so the desktop renderer sending `iss: null` was rejected with 4000 "unknown key" — breaking every Desktop→remote-gateway MCP OAuth login. Add the field, forward it, and regenerate the OpenRPC/TS contract artifacts via scripts/gen_gateway_contracts.py.

Also update tests/hermes_cli/test_mcp_dashboard_oauth.py for the 3-tuple callback shape introduced by the cherry-picked commit (it was red on the stack).
2026-09-15 13:00:12 +05:30
OOOOOAO 1a6503a520 fix(mcp): thread RFC 9207 iss through every OAuth callback relay
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:

    OAuthFlowError: Authorization response missing iss parameter
    advertised by the authorization server

and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.

Forward `iss` on every producer, matching `_make_callback_handler()`:

- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
  `wait_for_callback()` returns `(code, state, iss)`. The bridge in
  tools/mcp_oauth.py already splats that tuple into
  `_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
  parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
  reads `iss` off the redirect (the renderer already spreads the whole
  callback object into the RPC, so it flows through unchanged).

Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.

Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.

Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case

Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
2026-09-15 13:00:12 +05:30
kshitijk4poor 9a1d3920ad chore: map OOOOOAO contributor email 2026-09-15 13:00:12 +05:30
brooklyn! d128ce2e25 fix(desktop): keep sudo commands visible before password entry 2026-09-15 02:22:22 -05:00
brooklyn! b79107c565 fix(gateway): include command context in sudo password requests 2026-09-15 02:22:22 -05:00
kshitijk4poor 30a299180c refactor(state): stdlib-only home for read_only_db_uri; reuse the doctor/repair helpers
hermes_state_common pulls in agent.* at import, so the URI builder moves to
hermes_state_holders (errno/os/sqlite3/pathlib only) where the gateway
readiness probe and backup can adopt it in a follow-up sweep. The doctor
structural-damage branch is one helper instead of two copies, the holder
scan goes through hermes_state_repair._live_writer_holds_db, the migration
hint uses _schema_not_built (the startswith("no such ") check also matched
"no such module: fts5"), and the hermes_state import is hoisted so an import
failure cannot mask itself as UnboundLocalError.
2026-09-15 12:51:27 +05:30
kshitijk4poor ddbd7a9437 fix(cli): only a missing column/table means the read-only store needs migration
Other OperationalErrors (locked, disk I/O) keep their real message.
2026-09-15 12:51:27 +05:30
kshitijk4poor 55d9a49c1d refactor(state): one read-only URI builder; probe a held store via snapshot only
read_only_db_uri() replaces four inline mode=ro URI sites (two of which
still used the raw f-string that truncates on ?/# in the home path:
state_db_has_structural_damage and collect_state_db_stats). The doctor
write probe now applies the live-holder gate in both modes: a quiet store
is probed in place as on main, a held store is probed through a read-only
snapshot, and a held store over 1 GB is skipped with an info line unless
--fix is given (the unconditional copy cost one full DB write per plain
doctor run). Connect/backup failures propagate to the existing
classification instead of being reported as FTS write-health failures.
Observational sessions commands print a migration hint instead of a raw
traceback when a read-only opener meets an older schema.

Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
2026-09-15 12:51:27 +05:30
kshitijk4poor e30c4ed50d test(cli): trim observational-store coverage to two invariants
- Keep (a) list/stats/pinned open SessionDB(read_only=True) — one parametrized test — and (b) a missing store prints empty results and is never created.
- Drop the insights read-only test (already covered on main), the status test, the mutating-action/live-writer/doctor-isolation tests, and the two doctor factory tests.
- Replace _EmptyObservationalStore + error-string sniffing with a plain '_default_db_path() does not exist' branch printing each action's empty output; the fake-store tests in test_sessions_pin keep working because the branch only runs when the open fails.
2026-09-15 12:51:27 +05:30
kshitijk4poor 4a268c9e4a fix(cli): keep doctor's read-only opens URI-safe and holder-gated
- _session_count: back to main's raw sqlite mode=ro COUNT(*) via as_uri() — routing it through SessionDB(read_only=True) both re-introduced the raw f-string URI ('?'/'#' in the home path truncate it) and queries columns (s.archived) an unmigrated store lacks, so doctor would report a healthy DB as broken.
- _write_health_reason: snapshot source URI built with as_uri() for the same reason; the --fix live probe (_db_opens_cleanly runs BEGIN IMMEDIATE) now falls back to the snapshot unless live_writer_holds_db proves the store quiet, matching _state_db_wal — hermes doctor --fix never becomes a second writer against a gateway's state.db (#103339).
- SessionDB._connect_read_only: same as_uri() form so every read-only opener is safe in a home containing '?' or '#'.
- test_sessions_export_output_dir: fixture accepts the read_only kwarg the PR introduced.
- Drop the two doctor tests that pinned the SessionDB factory kwargs; main's URI-reserved-chars test covers _session_count.

Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
2026-09-15 12:51:27 +05:30
KoNit-K b4e22481a8 fix(cli): open observational session stores read-only
## What does this PR do?

Makes observational CLI commands open `state.db` in read-only mode, so they can inspect a live Hermes installation without participating in writable WAL lifecycle handling.

### Symptom

Running `hermes status`, `hermes doctor` without `--fix`, `hermes sessions list`, `hermes sessions stats`, or `hermes insights` while a gateway owns the store could open another writable session handle. The live turn could then lose its WAL generation and stop.

### Impact

Users inspecting status or session history during an active turn could lose that in-flight turn and leave the gateway halted until recovery.

### Bug Cause

**Trigger:** observational CLI helpers constructed `SessionDB()` with its writable default.

**Causal chain:**

1. A live gateway holds the `state.db` WAL generation.
2. A nested observational CLI command opens a second writable handle.
3. Writable-handle close behavior can participate in WAL lifecycle work and retire the generation used by the live writer.

**Why it is wrong:** these commands only query state and should not have writer privileges.

**Working sibling / contrast:** repair and mutating session commands still use writable access intentionally.

**Ruled out:** no state schema, migration, or WAL checkpoint implementation changes are included.

### Fix

Routes status, non-fixing doctor state inspection, sessions list/stats, and both insights entrypoints through `SessionDB(read_only=True)`. Repair and mutating paths remain writable, and regression tests cover WAL preservation with a live writer.

## Related Issue

Fixes #110173

## Type of Change

- ✅ Bug fix (non-breaking change that fixes an issue)

## Changes Made

- `hermes_cli/status.py`, `hermes_cli/doctor_state.py`, and insights helpers — open observational state readers read-only.
- `hermes_cli/sessions_cmd.py` — make only `list` and `stats` read-only; retain writable access for mutations.
- `tests/hermes_cli/test_observational_sessiondb_modes.py` — verify access modes and a live writer's WAL remains usable.

## How to Test

- ✅ `scripts/run_tests.sh tests/hermes_cli/test_observational_sessiondb_modes.py tests/hermes_cli/test_cli_insights_command.py` — 9 passed.
- ✅ `scripts/run_tests.sh tests/hermes_cli/test_doctor.py tests/hermes_cli/test_doctor_structural_corruption.py tests/hermes_cli/test_sessions_error_exit_codes.py` — 75 passed; two sandbox-only failures came from blocked host process/symlink operations.
- ✅ A live `SessionDB` writer remains able to create and retrieve a session after `sessions stats` reads the store.

## Checklist

### Code

- ✅ I've read the Contributing Guide
- ✅ My commit messages follow Conventional Commits
- ✅ I searched for existing PRs to make sure this isn't a duplicate
- ✅ My PR contains only changes related to this fix
- ✅ I've run relevant tests locally (see How to Test)
- ✅ I've added tests for my changes
- ✅ I've tested on my platform: macOS

### Documentation & Housekeeping

- ✅ Documentation update: N/A
- ✅ `cli-config.yaml.example`: N/A
- ✅ `CONTRIBUTING.md` or `AGENTS.md`: N/A
- ✅ Cross-platform impact considered
- ✅ Tool descriptions/schemas: N/A
2026-09-15 12:51:27 +05:30
kshitijk4poor b82c79ac50 refactor(agent): share the shutdown-only socket primitives with the pool sweep
The stale-attempt socket shutdown re-implemented two blocks that already
live in agent_runtime_helpers: the settimeout(0)+shutdown(SHUT_RDWR)
body of force_close_tcp_sockets (now _shutdown_socket) and the
candidate->socket lookup of _iter_pool_sockets (now _socket_from_candidate).
The hand-unrolled _httpcore_stream unwrapping is dead since
_connection_candidates walks _stream/_httpcore_stream itself, so the helper
starts from the network_stream extension and the response stream only.

Also add ReadError to _TRANSIENT_TRANSPORT_ERRORS (the third classifier of
the same abort-induced read; the other two were already updated), drop the
incidental gettimeout() assertion from the shape test, and cut the E2E
from ~3 s to ~1.5 s (stale budget 1 s, serve_forever poll 50 ms).
2026-09-15 12:46:27 +05:30
kshitijk4poor e625602a67 test(agent): trim stale-kill unwedge tests to the two invariants
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.

Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
2026-09-15 12:46:27 +05:30
kshitijk4poor 13f9206e52 fix(agent): drop the stale-aligned read-timeout cap from _stream_timeouts
Capping ``read`` to ``stale`` silently overrode an explicit
HERMES_STREAM_READ_TIMEOUT and the local-endpoint ``read = base`` branch,
and let the stale-kill E2E pass via ReadTimeout alone rather than through
the socket-shutdown unwedge this change is about. The shutdown path is
sufficient on its own: the E2E still passes with the cap removed.
2026-09-15 12:46:27 +05:30
finn763 2355d593f3 fix(agent): stale-killed stream unwedges its reader and reconnects (#110769)
The stale-stream monitor aborted a wedged provider stream only via
force_close_tcp_sockets() -> shutdown(SHUT_RDWR). That is best-effort: a
parked body read is not unblocked on every platform (Windows keeps the
pending recv parked) and the sweep can miss the socket. The worker then
stayed blocked in the provider read, so the retry loop never retried; the
monitor re-killed every stale interval and the call only ended at the
byte-read timeout, far past the stale budget - the reported
"No response from provider for 180-240s ... Reconnecting" loop ending in
"The model server is not responding".

- _kill_stale_stream now also closes the killed attempt's own provider
  response (identity-guarded self._attempt_stream_response), which is what
  actually unblocks a parked reader; a racing retry's fresh response is
  never touched.
- an abort-induced httpx.ReadError counts as a transient connection error,
  so the aborted attempt reconnects instead of ending the turn.

Reproduced with a local SSE server: before, the worker stayed parked and no
second request was issued (recovery only at the byte-read timeout); after,
the kill unblocks the reader at the stale budget and the retry lands.
2026-09-15 12:46:27 +05:30
kshitijk4poor 1d8af13e4d refactor(cli): one special-file filter for every profile copytree
Move _non_exportable_entries next to its first caller, fold the .pyc/.pyo
suffixes into it (the clone-all closure kept its own copy), and route the
last un-ignored profile copytree (the skills/ copy in _bootstrap_profile_dir)
through it. Cut the three repeated "sockets abort copytree" comments down to
the helper docstring. Split the clone-all special-file case into its own
POSIX-marked test so the cron-jobs assertion keeps its Windows coverage, and
use monkeypatch.chdir in the socket-binding test helper.
2026-09-15 12:45:29 +05:30
kshitijk4poor fe6d3016cf fix(cli): skip special files when cloning a profile with --clone-all
Route the --clone-all copytree ignore through _non_exportable_entries so a
live source profile holding a gateway or agent-browser socket (or a FIFO)
no longer aborts the clone with [Errno 6] No such device or address.
.pyc/.pyo and the root exclude sets keep their existing handling.

hermes_cli/profile_distribution.py:_copy_dist_payload is left alone: it
copies from a freshly extracted distribution archive (a staged temp tree),
never from a live profile, so it cannot meet a socket.

Extends test_clone_all_does_not_copy_cron_jobs to cover the clone path.
2026-09-15 12:45:29 +05:30
Camil Blanaru 9a82c44e38 fix(cli): skip Unix sockets and other special files during profile export
Named-profile export staged its copy with an ignore callable that only
excluded credential files, so any Unix socket in the profile (e.g. a stale
agent-browser control socket under home/.agent-browser/) made
shutil.copytree collect "[Errno 6] No such device or address" and raise
shutil.Error, failing the entire export. The default-profile export
already excluded *.sock by suffix, but a socket without that suffix (or a
FIFO, or a device node) failed it the same way.

Extract the universal exclusions into _non_exportable_entries(), which
keeps the __pycache__/*.sock/*.tmp name rules and additionally drops any
entry that is not a regular file, directory, or symlink (os.lstat mode
check), and use it in both the default and named export branches.
2026-09-15 12:45:29 +05:30
kshitijk4poor e6f2f2b759 chore: map camilb contributor email 2026-09-15 12:45:29 +05:30
kshitijk4poor 54d3217578 refactor(gateway): name the machinery display kind once; drop a dead isinstance guard 2026-09-15 12:44:57 +05:30
kshitijk4poor d1637409e9 docs(gateway): silence-marker retraction no longer implies an empty send 2026-09-15 12:44:57 +05:30