Commit Graph

5936 Commits

Author SHA1 Message Date
Teknium 094dba5a2f refactor(web): extract gateway/update/action-status routes into web_routers/actions.py 2026-09-02 13:30:55 -07:00
Teknium 401efb0b2b refactor(cli): decompose _cmd_update_impl into phase helpers, unify duplicated update helpers, compact comments
hermes_cli/update_cmd.py 11217 -> 10204 LOC; _cmd_update_impl 2797 -> 521 LOC.

Decomposition (behavior-neutral, AST free-name verified) — _cmd_update_impl is
now a thin orchestrator calling, in order:
  _clear_windows_venv_holders_or_exit -> _prepare_checkout_for_update
  (-> _CheckoutPlan) -> _repair_current_checkout | _pull_updates ->
  _sync_python_dependencies_after_pull -> _run_post_update_maintenance ->
  _restart_gateway_fleet_after_update (-> _GatewayRestartOutcome;
  _restart_systemd_gateway_units) -> _resume_windows_gateways_and_merge_outcome
  -> _verify_fleet_after_update. _resolve_manage_cmd hoisted to module level.

Unified helpers: _git_run (49 captured git subprocess.run sites),
_systemctl / _systemctl_reset_and_restart (17 sites + 2 pairs),
_sweep_bytecode_after_update (3x), _print_bundled_skills_sync_report (ZIP+git),
_ensure_venv_pip (ZIP+git), _self_and_non_gateway_ancestor_pids (2x),
_record_update_step (4x), _write_gateway_update_exit_code for the 3 inline
".update_exit_code" writes. Dropped duplicate nested copies of
_wait_for_service_active/_service_restart_sec/_print_items, dead
upstream_exists, a dead if/pass branch and unused imports.

Comments/docstrings hand-compacted (236 blocks, AST-identical with docstrings
normalized); rationale/invariant sentences kept.

Tests: source-inspection guards repointed to the helper that now owns the code
(test_update_self_lock, test_update_fleet_check_fail_closed,
test_update_apply_shallow_count); new regression test drives the real Windows
resume/merge helper.
2026-09-02 13:30:54 -07:00
Teknium 9de17448dd refactor(web): extract audio/TTS routes into web_routers/audio.py 2026-09-02 13:30:53 -07:00
Teknium feeb1074a3 refactor(web): extract config/env/custom-endpoint routes into web_routers/config_env.py 2026-09-02 13:30:53 -07:00
Teknium 1f4ab19eae refactor(web): extract model assignment routes into web_routers/models.py 2026-09-02 13:30:53 -07:00
Teknium 11bc6bbe11 refactor(web): extract memory-provider setup routes into web_routers/memory_providers.py 2026-09-02 13:30:51 -07:00
Teknium 1f77a129da refactor(web): extract messaging onboarding/platform routes into web_routers/messaging.py 2026-09-02 13:30:50 -07:00
Teknium 0cca195418 refactor(web): extract pairing/webhooks/credential-pool/memory/ops routes into web_routers/ops.py 2026-09-02 13:30:50 -07:00
Teknium 3b1ecfc0a1 refactor(cli): decompose list_authenticated_providers and switch_model
model_switch.py (4405 -> 3712):
- list_authenticated_providers (1367 lines) is a thin driver over _PickerBuild
  state and one builder per section (_lap_builtin_rows, _lap_overlay_rows,
  _lap_canonical_rows, _lap_user_provider_rows, _lap_bare_custom_row,
  _lap_custom_provider_rows). Credential checks shared with
  _collect_authed_provider_slugs are single-copy helpers
  (_iter_builtin_candidates, _auth_store_has_provider, _raw_pool_usable,
  _pool_usable, _overlay_has_env_creds, _has_aws_sdk_creds_for_listing);
  live/curated lookup (_live_or_curated_ids, _aws_live_or_curated_ids,
  _nous_picker_model_ids), endpoint discovery (_discover_endpoint_models) and
  config-entry parsing are deduped across sections 1/2/2b/3/3b/4.
- switch_model: _switch_fail, _runtime_creds, _entry_configured_key,
  _ollama_configured_base, _unknown_provider_message, _aggregator_alias_error,
  _aggregator_catalog_match, _config_declares_model,
  _apply_direct_alias_endpoint replace inline blocks; four copies of the
  resolve_runtime_provider unpack collapse to one.
- _model_sort_key: one _flush() for the four repeated version-flush blocks
  (identical on 400k random ids).
Row dicts, ordering, error strings and lazy-import patch points unchanged;
picker rows and the interactive hermes model screen byte-identical vs base.
2026-09-02 13:30:39 -07:00
Teknium 3d5f9269c7 refactor(web): extract files/fs/media routes into web_routers/files.py 2026-09-02 13:30:05 -07:00
Teknium f0f7f55f76 refactor(web): dashboard_auth/web_routers/kanban dedupe + _common helpers (resume — verified partial work) 2026-09-02 13:30:05 -07:00
Teknium 9dab23c5a6 refactor(cli): simplify hermes_cli/commands.py registry, gateway collectors and completer
- Drop dead code: _telegram_effective_priority, _prioritize_telegram_menu_commands,
  discord_skill_commands, _TG_NAME_LIMIT/_clamp_telegram_names compat aliases,
  the empty _SLACK_PRIORITY_ALIASES pinning pass (and tests that only pinned them).
- Unify the Telegram and Discord skill collectors behind _iter_gateway_skills
  (one eligibility/root-matching implementation) and _truncate_desc.
- Table-drive Telegram menu priority modes (_TELEGRAM_PRIORITY_TIERS) and the
  completer's per-command dynamic completions (_DYNAMIC_COMPLETIONS).
- Unify path and @file:/@folder: directory-listing completions (_dir_completions),
  command-completion construction (_short_desc), /tools candidate rows, the
  Slack canonical/alias passes and the derived COMMANDS/COMMANDS_BY_CATEGORY loops.
- Compact docstrings/comments, keeping every rule, invariant and rationale.
Behavior-neutral: registry-derived outputs, menus, manifests and completions
byte-identical against origin/main on a fixture sweep.
2026-09-02 13:29:45 -07:00
Teknium 150ee2a35d refactor(cli): simplify CLI mixins (billing, agent-setup, commands) — dedupe helpers, dispatch tables, god-method splits
hermes_cli/cli_billing_mixin.py (1566 -> 1173): shared _modal_choice / _step_up_remote_spending /
_print_org_line / _print_portal_line / _print_total_spendable / _try_usage_model /
_billing_patch_auto_top_up; copy tables for charge-failed / charge-error / upgrade-status.

hermes_cli/cli_agent_setup_mixin.py (962 -> 787): _current_runtime + _route_signature (3 sites),
_compression_descendant (2 sites), resume event/skin tables, _load_resumed_history_late lifted
out of _init_agent; fixes latent undefined 'logger' in _resume_history_limit_error.

hermes_cli/cli_commands_mixin.py (4175 -> 3634): _end_current_session + _sync_agent_to_session
(/resume + /branch), _print_side_result_panel (/bg + /btw), _queue_prompt_turn (/learn /plan /init),
_toggle_target + _command_arg (/footer /timestamps /focus /voice /wake), _print_diff_body,
_summarize_paths, _split_scope_flags + _scope_outcome (/reasoning /fast); dispatch tables for /cron
flags + subcommands, /busy, /fast, /voice; /browser, /cron, /goal god-methods split into helpers.
Docstrings/comments compacted by hand; rationale kept.
2026-09-02 13:29:44 -07:00
Teknium ab2fb71b23 refactor(cli): drop 18 update_cmd lazy re-exports with zero references outside update_cmd itself 2026-09-02 13:29:44 -07:00
Teknium 5c125d3bf6 refactor(cli): lift cmd_gui's install+build+stage/swap region into _build_desktop_app; drop 2 unused imports
Zero-back-ref region (inputs desktop_dir/source_mode/npm/env, single output
packaged_executable) becomes a helper returning the swapped-in executable.
Removes the unused functools/_add_accept_hooks_flag imports left by the
parser-builder migration.
2026-09-02 13:29:44 -07:00
Teknium 64f301dfc0 refactor(cli): dedupe the 3x --oneshot dispatch block and chat-arg defaults into helpers
_run_oneshot_from_args replaces three identical confirm+run_and_exit blocks
(main, fast chat, Termux fast cli); _default_to_chat reuses the existing
_set_chat_arg_defaults instead of a second attr table. Also restores the
'bare hermes profile' note as _profile_status's docstring.
2026-09-02 13:29:44 -07:00
Teknium ed8aa5e3e4 refactor(cli): move session browse/status/relative-time helpers from main.py into sessions_cmd.py
_relative_time, _session_status_tag, _annotate_session_statuses,
_session_browse_picker and _size_delta_label (410 LOC) replace the
call-time delegating wrappers in sessions_cmd; main.py re-exports them via
_LAZY_COMMAND_EXPORTS so hermes_cli.main.<name> imports/patches still work.
2026-09-02 13:29:44 -07:00
Teknium 277d51bbc9 refactor(cli): route select_provider_and_model flows through a _PROVIDER_MODEL_FLOWS table
The 20-way if/elif on selected_provider becomes a dict of uniform
flow(config, current_model, args) lambdas plus a _GENERIC_API_KEY_PROVIDERS
frozenset; custom-slug / remove-custom / api-key fallthroughs keep their
order. Lambdas resolve _model_flow_* by name at call time so existing
hermes_cli.main monkeypatches still intercept.
2026-09-02 13:29:44 -07:00
Teknium 853646f143 refactor(cli): move cmd_profile to hermes_cli/profile_cmd.py as a PROFILE_ACTIONS dispatch table
The 604-line 14-way if/elif on profile_action becomes one _profile_<action>
handler each, with per-handler imports derived from the branch's free names.
main.py re-exports cmd_profile and _render_distribution_plan so existing
imports/monkeypatches resolve unchanged.
2026-09-02 13:29:44 -07:00
Teknium 41b658d2b4 refactor(cli): unify duplicate reasoning-effort config helpers into hermes_cli.setup 2026-09-02 13:29:44 -07:00
Teknium 060039d169 refactor(cli): split main() into _build_cli_parser/_parse_cli_args/_default_to_chat helpers
main() is now a ~110-line orchestrator: startup prologue, parser build,
container routing, bpo-9338-safe parse, --version/--yolo/--oneshot, chat
default, dispatch. The two identical plugin add_parser blocks collapse into
_attach_plugin_cli_command; the two default-to-chat attr loops merge. Parser
tree, --help output and set_defaults are unchanged (399 parsers byte-diffed).
2026-09-02 13:29:44 -07:00
Teknium b81ab901f3 refactor(cli): extract the last 16 inline parser groups from main() into hermes_cli/subcommands/
moa, fallback, worktree, browser, secrets, egress, migrate, whatsapp-cloud,
checkpoints, bundles, curator, pets, journey, computer-use, sessions and
completion each become a build_<group>_parser() builder. Closure handlers
that only closed over their own parser moved verbatim; sessions/completion
take the handler by injection. --help/usage/defaults byte-identical for all
399 parsers in the tree (in-process dump before/after).
2026-09-02 13:29:44 -07:00
Teknium 0ccf6714be refactor(cli): drop zero-ref startup-fast wrappers and _build_provider_choices from main.py 2026-09-02 13:29:43 -07:00
Teknium 3d3d2c89f5 refactor(cli): table-drive setup.py wizard sections; drop dead helpers
setup.py (3933 -> 3651):
- Dead code (zero references repo-wide, refs.py-verified): _model_config_dict,
  _get/_set_credential_pool_strategy, _supports_same_provider_pool_setup,
  _current/_set_reasoning_effort, _prompt_container_resources, _setup_telegram_auto.
- setup_terminal_backend: (slug, label) table + _TERMINAL_BACKEND_SETUP dispatch
  to one _setup_backend_* per backend; _ensure_sdk / _prompt_secret_env dedupe
  the modal/daytona install+token prompts.
- _setup_tts_provider: label/choice/api-key/local-engine tables, three step
  functions replace the 9-branch elif; _pip_install_tts_package dedupes the
  neutts/kittentts installer tail.
- _print_setup_summary: TTS/STT rows from _TTS/_STT_SUMMARY_ROWS via
  _voice_provider_status; banner and hint lines from small tables.
- setup_agent_settings: session-reset mode table + _prompt_int_setting.
- setup_gateway: _HOME_CHANNEL_CHECKS table.
Printed output and .env/config writes verified byte-identical against base
over exhaustive input matrices (summary 630, tts 3840, agent 3375 combos).
2026-09-02 13:29:29 -07:00
Teknium 5212a3077d refactor(cli): dedupe model_setup_flows boilerplate
model_setup_flows.py (3313 -> 2848):
- _load_config_model_section, _begin/_commit_model_config, _ensure_flow_api_key,
  _pick_model_or_prompt, _run_login, _models_dev_merged, _copilot_model_list,
  _show_curated replace ~15 copies of config-save / api-key / picker boilerplate.
- _gemini_tier_ok and _api_key_provider_model_list lift the two inline blocks
  out of _model_flow_api_key_provider; five-way provider branch -> early returns.
- Comments compacted, keeping every rationale (Bedrock geo routing, key_env
  hygiene, discover_models semantics, Nous free/paid partition, etc.).

Also drops two tests that only asserted the existence of setup.py helpers
removed in the next commit.
2026-09-02 13:29:29 -07:00
kshitijk4poor a3d33fe22f fix(copilot): single-flight the token exchange and close abandoned responses
Follow-up to the off-loop move: once the credential-pool handlers run on
worker threads, the dashboard's periodic /api/credentials/pool polls can
overlap, and during a DNS outage each poll would have started its own
exchange and abandoned its own hung resolver thread.

- Per-fingerprint threading.Lock around the exchange: concurrent callers
  wait on the one in-flight attempt, then hit the positive or negative
  cache (bounded worker count, no duplicate network calls).
- _urlopen_bounded: when the hard cap fires and the abandoned worker later
  succeeds, close the HTTPResponse instead of leaking the socket.
- Tests (none shipped with the original PR): hard cap + late-close,
  single-flight success and failure paths, and the pool endpoint running
  off-loop / keeping the loop responsive under a 200 ms blocking read.
2026-09-03 01:57:09 +05:30
peetteerr ff0afff0e4 fix(server): move blocking credential-pool calls off the event loop
Network off (unplugged) froze the backend 17 minutes: the async
/api/credentials/pool endpoints (GET/POST/DELETE) called load_pool()
synchronously on the event-loop thread -> Copilot token exchange ->
blocking urlopen -> getaddrinfo stuck in C for 1016s, immune to
urlopen(timeout=10). WS dropped (1006), sessions detached, even log
writes stalled.

Fix 1 (web_server.py): move the three endpoint bodies into
asyncio.to_thread, matching the file's existing _run pattern.

Fix 2 (copilot_auth.py): _urlopen_bounded() runs the request in a
daemon thread with a wall-clock hard cap (timeout+5s) so DNS hangs
can no longer block any caller indefinitely.

Measured: simulated DNS hang now raises TimeoutError after 6.0s
instead of freezing the loop.
2026-09-03 01:57:09 +05:30
kshitijk4poor 2b4e70ec07 refactor(cli): use the router's run_in_threadpool alias; offload profile model write
The module already binds run_in_threadpool (used by list_profiles_endpoint)
and every sibling router uses the same starlette helper; the nine new
loop.run_in_executor(None, _run) sites now go through that alias so the
file has one offload idiom. Behaviour-identical (both hand the callable to
a worker thread).

Also sweeps the one endpoint the PR left synchronous:
update_profile_model_endpoint's _write_profile_model reads and rewrites
the profile's config.yaml on the event loop.
2026-09-03 01:56:47 +05:30
briandevans 34a8e1dc50 fix(cli): run profile document I/O off the dashboard event loop
The remaining in-scope handlers in this router read and write profile
documents inline on the ASGI event loop:

- GET  /api/profiles/{name}/soul            reads SOUL.md
- PUT  /api/profiles/{name}/soul            atomic_write_text(SOUL.md)
- PUT  /api/profiles/{name}/description     write_profile_meta(profile.yaml)
- GET  /api/profiles/{name}/desktop-overlay reads desktop.json

The persona save is the sharpest of the four: atomic_write_text() writes a
temp file, fsyncs it and replaces the original, so the loop is parked for
however long the filesystem takes to durably commit — unbounded on a slow
or contended disk, and paid on every Save in the editor.

Each handler keeps its existing status-code mapping. The reads probe and
load in a single executor hop rather than two, which also avoids widening
the gap between the existence check and the read.

Both readers return a _MISSING sentinel rather than None for an absent
file. desktop.json may legitimately contain the document `null`; collapsing
that onto None would newly report an existing-but-empty overlay as absent.
The same distinction is what the SOUL.md durability tests rely on, where
"file missing" and "file empty" must not both read as never-set.

_resolve_profile_dir() stays on the loop in all four, as it does in the
rest of this sweep: it is a name check plus one stat, and it owns the
400/404 responses.
2026-09-03 01:56:47 +05:30
briandevans 63d42cd0e2 fix(cli): run profile rename and active-profile state off the event loop
Three more handlers in this router did filesystem work inline on the ASGI
event loop:

- PATCH /api/profiles/{name} calls rename_profile(), which stops a running
  gateway through the same 10-second _stop_gateway_process() poll that
  delete uses, then renames the profile directory, rewrites the Honcho
  host blocks and regenerates the wrapper script.
- GET /api/profiles/active reads the active_profile state file and
  resolves HERMES_HOME against the profiles root. The sidebar polls it.
- POST /api/profiles/active stats the target profile, creates the state
  directory and writes active_profile via a temp file plus replace.

Rename carries the same worst case as delete and belongs off the loop for
the same reason. The two active-profile handlers are individually cheap,
but they are the routes the dashboard polls, so they are the ones most
likely to be queued behind something slower — and leaving them inline is
what made the router inconsistent with list_profiles_endpoint, which
already offloads a plain directory listing eight lines above.

The two reads in GET share one executor hop rather than taking one each.
2026-09-03 01:56:47 +05:30
briandevans 4da5689b80 fix(cli): run auto-describe LLM round-trip off the dashboard event loop
POST /api/profiles/{name}/describe-auto called
profile_describer.describe_profile() inline. That function is a plain def;
it reaches agent.auxiliary_client.call_llm(), also a plain def, which makes
a synchronous provider request with a 60-second ceiling.

Held on the ASGI event loop that is six times the 10-second WebSocket
ready-probe threshold web_server.py records as the point where the desktop
app gives up (GH-73083). A single describe-auto on a slow or unreachable
auxiliary provider therefore takes the whole dashboard offline for up to a
minute, including the /api/ws and /api/pty sockets the desktop app and the
Chat tab run on.

Move the import and call into the default executor. _resolve_profile_dir()
deliberately stays on the loop ahead of the hop: it is a name validation
plus a single stat, and it owns the 400/404 responses that the handler's
`except Exception` would otherwise turn into a 500.
2026-09-03 01:56:47 +05:30
briandevans a2504a0a59 fix(cli): run profile deletion off the dashboard event loop
DELETE /api/profiles/{name} called profiles.delete_profile() inline on the
ASGI event loop. When the target profile has a gateway running, that call
stops it via _stop_gateway_process(), which polls the PID every 500 ms for
up to 10 s before escalating to a force kill, and then removes the profile
tree.

For the whole of that window the dashboard process serves nothing else.
web_server.py's own notes record what that costs: a stall of this length
"caus[ed] the Desktop's 10-second WebSocket ready-probe to time out
(GH-73083)", and both the desktop app and the dashboard's Chat tab drive the
agent over those WebSockets. Deleting a profile whose gateway is up is a
routine action that reliably reaches the full ten seconds — the handler's
own output announces "Gateway is running - it will be stopped".

Move the call into the default executor via run_in_executor, matching
list_profiles_endpoint, export_profile_endpoint and import_profile_endpoint
in this same module. The exception-to-status mapping is unchanged:
FileNotFoundError/ValueError are raised inside the worker and re-raised by
the await, so they still map to 404/400.
2026-09-03 01:56:47 +05:30
briandevans 8f6df4a82a fix(proxy): resolve the 401/429 retry credential off the event loop
`handle_proxy` already offloads the two credential reads on the happy path,
but the rotation inside the `upstream_resp.status in {401, 429}` branch still
called `adapter.get_retry_credential` inline on the event loop.

That is the most expensive of the three blocking methods on the
`UpstreamAdapter` contract, not the cheapest:

  * `NousPortalAdapter.get_retry_credential` routes into
    `_get_credential(force_refresh=True)`, so the token-refresh POST that
    `get_credential` performs only near expiry is unconditional here — and it
    runs under the same `_auth_store_lock()`, a cross-process advisory lock
    with a 15s timeout.
  * `XAIGrokAdapter.get_retry_credential` loads the key pool off disk and
    calls `try_refresh_current` / `mark_exhausted_and_rotate` under its lock.

So every upstream 401 or 429 froze the proxy's single event loop — and with
it every other in-flight streaming completion — for the whole rotation. A 429
is exactly when the proxy is busiest, which is the worst moment to stall.

Wrap it in `asyncio.to_thread`, matching the two sites above. The error
contract is unchanged: `to_thread` re-raises the worker's exception in the
awaiting frame, so the existing `except Exception -> retry_cred = None` still
swallows a failed rotation and streams the upstream's own rejection back.
2026-09-03 01:56:32 +05:30
briandevans 985b034a51 fix(proxy): keep the health endpoint off the auth-store lock
`handle_health` called `adapter.is_authenticated()` inline from an
`async def`. `UpstreamAdapter.is_authenticated` is documented as
"Should be cheap — no network calls. Used by `proxy start` for a clear
up-front error before binding a port." (`adapters/base.py`), and that is
true of the `proxy start` preflight, which runs in a plain synchronous
CLI function. It is not true on the event loop:
`NousPortalAdapter.is_authenticated` goes through `_read_state()`, which
takes `_auth_store_lock()` — the same cross-process lock with a 15s
timeout as credential resolution — and `XAIGrokAdapter` reads its key
pool off disk.

`/health` is precisely what a supervisor, systemd unit, container
healthcheck or load balancer polls, on a fixed interval, so it is the
endpoint least able to afford a lock wait; and a wait here freezes every
concurrent proxied stream, not just the healthcheck.

Offload it with `asyncio.to_thread`. The response body is byte-identical;
only the scheduling changes.
2026-09-03 01:56:32 +05:30
briandevans 5b63c6174b fix(proxy): resolve upstream credentials off the event loop
`create_app` registers `handle_proxy` as an `async def`, and it called
`adapter.get_credential()` directly on the aiohttp event loop.

`UpstreamAdapter` is a synchronous contract (`adapters/base.py` — every
method is a plain `def`), and the shipped adapters implement it with
blocking I/O. `NousPortalAdapter.get_credential` takes
`_auth_store_lock()` — a cross-process advisory lock with
`AUTH_LOCK_TIMEOUT_SECONDS = 15.0` (`hermes_cli/auth.py:110`) — reads
`auth.json` off disk, and may issue a token-refresh POST; on a terminal
`AuthError` it takes that lock a second time to persist the quarantined
state. `XAIGrokAdapter.get_credential` reads its key pool off disk.

The proxy is one process with one loop, and `handle_proxy` streams with
`sock_read=300`, so long-lived completions are the normal case. Blocking
inside credential resolution therefore freezes *every* concurrent
in-flight stream mid-token for the duration — a concurrent `hermes auth`
command holding the auth-store lock is enough to do it. Every proxied
request goes through this path.

Dispatch through `asyncio.to_thread` instead. This is a pure scheduling
change: `to_thread` re-raises the worker's exception in the awaiting
frame, so the existing `except Exception` -> 401 `upstream_auth_failed`
mapping is unchanged, and the adapters' own `self._lock` still serialises
concurrent resolutions exactly as before. Fixing it at the handler also
leaves the synchronous `UpstreamAdapter` ABC untouched, so it covers
every adapter without conflicting with in-flight work that subclasses it.
2026-09-03 01:56:32 +05:30
Teknium 10088c569e fix(update): hermes update no longer hangs on a GitHub username prompt
GitHub answers anonymous fetches with HTTP 401 during outages (and for
renamed/private repos). git then prompts `Username for 'https://github.com':`
on the inherited terminal and `hermes update` sits there — users read it as
Hermes demanding a GitHub login.

Every network git call in the updater (fetch/pull/push, apply + --check +
fork sync) now runs with GIT_TERMINAL_PROMPT=0 / stdin=DEVNULL, so the 401
fails fast into the fetch-failure classifier, which now reports it as a
GitHub-side rejection (likely outage) rather than blaming the user's
credentials. Credential helpers/askpass are left configured so private-fork
origins still authenticate.

Live repro: PTY-attached update --check against a 401 origin hung 15s+ on
the prompt before; exits rc=1 in 0.2s with the diagnosis after.

Same class as #73751 (@Frowtek, pre-main.py decomposition); passive banner
half salvaged from #101421 (@RobbertC5).
2026-09-02 12:34:26 -07:00
RobbertC5 6e775907d7 fix(cli): keep passive update checks noninteractive 2026-09-02 12:34:26 -07:00
Teknium f6234d00c5 fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session
directory automatically — the coding-workspace snapshot, gateway
project-tree build, /diff, @diff|@staged context refs, goal-gate
fingerprint, and -w startup worktree add — before any prompt, tool call,
approval, or trust gate. Those probes ran the system git without
stripping the repository's own config, so a repo delivered as files with
its .git directory intact (a shared zip, sync folder, or USB stick;
git clone never transfers .git/config) could set an execution-sink git
setting and get arbitrary host code execution as the user with nothing
on screen.

- core.fsmonitor / core.hooksPath / pager / editor / credential helper:
  neutralized by routing every automatic probe through
  noninteractive_git_env(), which pins those keys to inert values via
  GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the
  reported sink, coding_context._git + tui_gateway.git_probe) now defaults
  to that env; worktree-add, working_diff, web_git, context_references,
  goals, and subagent_worktree route through it too.
- Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker
  names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't
  enumerate them. Added harden_git_argv(), which inserts
  --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/
  log/blame) only — status et al reject the flags. Both flags required
  (verified empirically; each alone leaves the other live).

Builds on the noninteractive_git_env config-scrubbing from the
gemini-cli #28792 port. Real-git E2E regression suite arms a malicious
repo and asserts every automatic path neutralizes fsmonitor, hooks,
external-diff, and textconv; a baseline test proves the repo is armed.
2026-09-02 10:33:43 -07:00
Teknium 413b6ba3dd Port from google-gemini/gemini-cli#28792: harden internal git env
(cherry picked from commit 09bb9c3b2f20f22bdb1688b6efcb78b24f1f19d6)
2026-09-02 10:33:43 -07:00
Teknium 1b49ad9be9 fix(dashboard): fold the one-field nous config section into the agent settings tab 2026-09-02 10:22:30 -07:00
olopez25 c6c8c74c70 Move the keepalive interval to config.yaml and tighten the schedule guard
Addresses review feedback on #84928.

The tick interval was exposed as HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS.
AGENTS.md reserves .env for credentials and puts behavioural thresholds in
config.yaml, so the knob moves to `nous.keepalive_interval_seconds`,
following the existing `vertex:` section's precedent for non-secret
provider settings. The env var is dropped rather than bridged: it was never
released, so nothing depends on it. Adding a key to a new section is handled
by the deep-merge, so no _config_version bump is required.

test_keepalive_interval_fits_inside_the_token_lifetime asserted
`900 < 899 * 4 - 120`, which is true for any realistic interval and could
never fail. It also tested the wrong value: the configured constant is only
a ceiling, while the schedule that ships is the derived tick. Replaced with
an assertion over the derived tick for each observed lifetime, which does
fail if the derivation constants regress -- verified against both
TICKS_PER_LIFETIME=1 and MIN_INTERVAL_SECONDS=5000.

Also adds coverage for an unreadable config.yaml, which must fall back to
the module default rather than take the keepalive thread down.

pytest tests/hermes_cli/test_nous_auth_keepalive.py -> 9 passed

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 10:22:30 -07:00
olopez25 fd05b97913 Derive Nous keepalive tick from the issued credential lifetime
Follow-up to the interval fix: ticking faster narrowed the gap but did not
close it, and the hardcoded interval was wrong for short-lifetime accounts.

Two independent problems, both needed:

1. Lifetime is not a constant. Installs have been observed issuing ~3594s
   and ~899s (see #35752, which reports expires_in: 899 while this account
   reports 3594). Any hardcoded interval is right for one and wrong for the
   other. The tick now derives from the lifetime the server actually issued,
   capped by the configured interval and floored at 60s. The access token and
   the invoke agent key carry separate lifetimes, so the shorter one governs.

2. Ticking faster alone never closes the gap. The refresh only fires once a
   credential is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of expiry,
   so a tick spaced wider than that window steps straight over it. With a 900s
   tick against a 3594s lifetime the last tick before expiry still saw 894s
   remaining, declined to refresh, and the credential died 6s before the next
   one. The keepalive now asks "will this outlive my next tick?" (tick + skew)
   rather than reusing the request path's bare skew.

Measured against both observed lifetimes:

  lifetime 3594s: was never refreshed proactively; now refreshes 900s early
  lifetime  899s: was never refreshed proactively; now refreshes 227s early

The lifetime is re-read every pass rather than cached, since it can change
when the account, plan, or server-side policy does -- exactly the case the
keepalive exists to cover.

min_access_ttl_seconds is threaded through as an optional parameter, so the
request path keeps its existing 120s behaviour and only the keepalive widens
the window.
2026-09-02 10:22:30 -07:00
olopez25 a301d19ba4 Fix reactive Nous 401s by ticking keepalive inside the token lifetime
Nous Portal access tokens carry a one-hour lifetime, and the keepalive only
refreshes once a token is within ACCESS_TOKEN_REFRESH_SKEW_SECONDS (120s) of
expiry. The tick interval was 6 hours, so it could only land inside that
2-minute window by coincidence. In practice every hour rolled over untouched
and the next inference call paid a 401 plus a re-auth round trip.

Observed on a production install: 71 "refreshed Nous runtime credentials
after 401, retrying" events across current logs.

Drop the default interval to 15 minutes, which gives four ticks per token
lifetime and leaves ample margin under the TTL-minus-skew ceiling of 3480s.
Add HERMES_NOUS_KEEPALIVE_INTERVAL_SECONDS so the interval is tunable without
a source edit, matching the existing HERMES_NOUS_TIMEOUT_SECONDS convention.
Zero still disables the thread.

Callers pass no interval, so the default is what actually shipped; the
signature now resolves at call time rather than binding the constant at
import.
2026-09-02 10:22:30 -07:00
Alexander Prendota 1bd0b8ed41 feat(providers): let an external-process provider ship out of tree
An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:

* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
  ``providers/`` covered api-key providers only, so an external-process profile
  never entered it and ``hermes -m <that provider>`` died with "Unknown
  provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
  literal ``"copilot-acp"``, so anything else silently fell through to the
  OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
  (``copilot``), the argv (``--acp --stdio``), the env var names and the
  placeholder api_key — so a third-party provider would have been handed
  another vendor's CLI.

Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.

The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.

Co-Authored-By: Junie <junie@jetbrains.com>
2026-09-02 09:57:39 -07:00
Teknium 0ee98eda52 feat(models): add google/gemini-3.8-flash to nous + openrouter catalogs
Slots above gemini-3.7-flash (kept) in OPENROUTER_MODELS and
_PROVIDER_MODELS["nous"]; openrouter plugin fallback_models bumped
3.7 -> 3.8; model-catalog.json regenerated.

Verified live with test completions on both Nous Portal and OpenRouter
(model echo + billed). Same 1,048,576 window / 65,536 output / pricing
as 3.7-flash, so provider-agnostic metadata resolves via the existing
gemini entries and both routes bill live (official_models_api) — no
pricing snapshot needed.

Scoped to the two named providers: vertex/gemini/kilocode/gmi curated
lists, setup.py samples, and aux defaults untouched.
2026-09-02 09:46:15 -07:00
Teknium 3312947e14 fix(auth): concurrent Nous 401 recovery adopts a peer's refresh instead of re-rotating the shared grant
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.

- resolve_nous_runtime_credentials(stale_access_token=): under the store
  lock, skip the refresh POST when the on-disk token differs from the one
  that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
  pass the failed bearer through, and treat a lock TimeoutError as
  'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
  refresh/0 unrecovered.
2026-09-02 09:33:05 -07:00
unsupportedpastels 5487222658 fix(models): live Copilot catalog for CLI-login users; unbreak picker-flag-empty catalogs
Three gaps between the copilot-acp picker row and what the user's
subscription actually serves (reported: picker showed the stale curated
list while the Copilot CLI offered Sonnet 5 / Opus 5 / GPT-5.6):

1. _resolve_copilot_catalog_api_key() never looked at the Copilot CLI's
   own token store (~/.copilot/config.json copilotTokens). A user whose
   only credential is 'copilot login' got no catalog key, the live fetch
   401'd, and copilot-acp silently fell back to the stale curated list.
   Add it as resolution source 3, JSONC-tolerant, with each candidate
   validated and exchanged like pool entries.

2. The existing credential-pool branch unpacked exchange_copilot_token()
   into two names, but it returns (api_token, expires_at, base_url) —
   the ValueError was swallowed by the enclosing except, disabling that
   entire resolution path. Latent since the base_url return was added.

3. GitHub now returns model_picker_enabled: false for EVERY model on
   some accounts/token types, so honoring the flag rejected the whole
   live catalog. Treat the flag as a display hint: when it empties the
   result, refilter without it (chat/endpoint checks still exclude
   embeddings and non-chat rows).

Verified live: catalog resolves 44 models for a copilot-login-only
account, matching the CLI's own picker (claude-sonnet-5, claude-opus-5,
gpt-5.6-sol/terra, gemini, kimi).
2026-09-02 20:51:07 +05:30
unsupportedpastels 6b2d32d3f6 fix(picker): keep signed-in copilot-acp visible in explicit-only desktop pickers
Two follow-up gaps found by actually running 'copilot login' end-to-end:

1. The CLI (without an OS keychain) stores its token in
   ~/.copilot/config.json under copilotTokens — a JSONC file with
   //-comment header lines. Add it as an auth-evidence source in
   _external_process_auth_evidence(), parsed comment-tolerantly and
   counting only a non-empty copilotTokens map (config.json exists after
   first launch even when logged out).

2. The desktop chat picker requests explicit_only rows, and
   _filter_explicit_provider_rows() dropped copilot-acp because a CLI
   login leaves no trace in active_provider, model.provider, or env vars
   — exactly the Anthropic-OAuth carve-out case. Keep external_process
   rows when their CLI credentials are verified (auth_verified), while
   still dropping ambient executable-on-PATH-only rows so the filter's
   narrower contract holds.

Net effect: after 'copilot login', copilot-acp appears in the desktop
picker and the Accounts card reads signed in; a machine with only the
binary installed keeps today's hidden-until-configured behavior.
2026-09-02 20:51:07 +05:30
unsupportedpastels 323168a289 fix(dashboard): correct copilot-acp sign-in command, honest status card
The Accounts-tab card told users to run 'copilot /login', which is not a
valid invocation — slash-commands only exist inside an interactive session.
Use 'copilot login', the CLI's device-code login subcommand.

The card's status_fn also hardcoded logged_in: False with a static label.
Wire it to get_external_process_provider_status(): claim logged_in only on
positive credential evidence (auth_verified), show which executable Hermes
resolved when merely configured, and say so when the CLI is missing from
PATH entirely.

The rendered cli_command now substitutes the executable the user actually
configured (HERMES_COPILOT_ACP_COMMAND / COPILOT_CLI_PATH) so a custom
binary path gets a copy-pasteable command that matches what Hermes spawns.
2026-09-02 20:51:07 +05:30
unsupportedpastels fd439ac1b8 fix(auth): dispatch external-process providers by auth_type, add positive auth evidence
get_auth_status() special-cased the literal slug 'copilot-acp'; any other
external_process provider (the pending kiro/devin/junie ACP backends) fell
through to {'logged_in': False}. Dispatch on
PROVIDER_REGISTRY[target].auth_type == 'external_process' instead so the
whole class gets a real status.

get_external_process_provider_status() equated 'logged_in' with 'the
executable resolves', which says nothing about whether the Copilot CLI is
actually signed in. Add auth_verified/auth_source: positive-only evidence
from supported env tokens (validated via copilot_auth, classic ghp_* PATs
excluded) or known on-disk GitHub Copilot credential stores. No evidence
means unknown — never presented as signed out, because the CLI may keep its
session in an OS keychain. Deliberately subprocess-free to avoid re-creating
the gh-auth-token cold-start stall (#60800).
2026-09-02 20:51:07 +05:30