Commit Graph

684 Commits

Author SHA1 Message Date
Teknium beb56b6a00 refactor(web): re-export list_credential_pool from web_server for off-loop tests 2026-09-02 13:49:24 -07:00
Teknium 5c7096e7cf refactor(web): delete dead web_server helpers (_demo, _ws_request_reason, _TERMINAL_BACKEND_NAMES, _mcp_oauth_callback_url_from_base) 2026-09-02 13:33:23 -07:00
Teknium 4fd970c62f refactor(web): absorb single-router web_server helpers into their routers (75 defs); drop dead imports 2026-09-02 13:33:22 -07:00
Teknium c90af45f57 refactor(web): unify the three post-config gateway auto-restart helpers 2026-09-02 13:32:52 -07:00
Teknium f2c506845b refactor(web): unify device-code poller lifecycle into _oauth_poller decorator 2026-09-02 13:32:49 -07:00
Teknium ba4f4b8161 refactor(web): collapse try/except-HTTPException/log/raise boilerplate into _common.http_failure; trim unused fastapi imports 2026-09-02 13:32:46 -07:00
Teknium ba91166b1c refactor(web): hoist stdlib imports in extracted routers; drop unused web_server imports 2026-09-02 13:32:28 -07:00
Teknium e4d8a78967 refactor(web): move /api/logs into web_routers/status.py (logs_router) 2026-09-02 13:32:08 -07:00
Teknium de109d4097 refactor(web): extract chat-tab WebSocket routes into web_routers/chat_ws.py 2026-09-02 13:32:06 -07:00
Teknium 4bae44d654 refactor(web): extract dashboard theme/font/plugin routes into web_routers/dashboard_ui.py 2026-09-02 13:32:05 -07:00
Teknium 94309ad40a refactor(web): extract raw-config + analytics routes into web_routers/analytics.py 2026-09-02 13:32:02 -07:00
Teknium 067028eb22 refactor(web): extract OAuth provider routes into web_routers/oauth.py 2026-09-02 13:32:01 -07:00
Teknium d3b768232e refactor(web): extract status/health/curator/learning/portal routes into web_routers/status.py 2026-09-02 13:30:56 -07:00
Teknium 094dba5a2f refactor(web): extract gateway/update/action-status routes into web_routers/actions.py 2026-09-02 13:30:55 -07:00
Teknium 9de17448dd refactor(web): extract audio/TTS routes into web_routers/audio.py 2026-09-02 13:30:53 -07:00
Teknium feeb1074a3 refactor(web): extract config/env/custom-endpoint routes into web_routers/config_env.py 2026-09-02 13:30:53 -07:00
Teknium 1f4ab19eae refactor(web): extract model assignment routes into web_routers/models.py 2026-09-02 13:30:53 -07:00
Teknium 11bc6bbe11 refactor(web): extract memory-provider setup routes into web_routers/memory_providers.py 2026-09-02 13:30:51 -07:00
Teknium 1f77a129da refactor(web): extract messaging onboarding/platform routes into web_routers/messaging.py 2026-09-02 13:30:50 -07:00
Teknium 0cca195418 refactor(web): extract pairing/webhooks/credential-pool/memory/ops routes into web_routers/ops.py 2026-09-02 13:30:50 -07:00
Teknium 3d5f9269c7 refactor(web): extract files/fs/media routes into web_routers/files.py 2026-09-02 13:30:05 -07:00
peetteerr ff0afff0e4 fix(server): move blocking credential-pool calls off the event loop
Network off (unplugged) froze the backend 17 minutes: the async
/api/credentials/pool endpoints (GET/POST/DELETE) called load_pool()
synchronously on the event-loop thread -> Copilot token exchange ->
blocking urlopen -> getaddrinfo stuck in C for 1016s, immune to
urlopen(timeout=10). WS dropped (1006), sessions detached, even log
writes stalled.

Fix 1 (web_server.py): move the three endpoint bodies into
asyncio.to_thread, matching the file's existing _run pattern.

Fix 2 (copilot_auth.py): _urlopen_bounded() runs the request in a
daemon thread with a wall-clock hard cap (timeout+5s) so DNS hangs
can no longer block any caller indefinitely.

Measured: simulated DNS hang now raises TimeoutError after 6.0s
instead of freezing the loop.
2026-09-03 01:57:09 +05:30
Teknium 1b49ad9be9 fix(dashboard): fold the one-field nous config section into the agent settings tab 2026-09-02 10:22:30 -07:00
unsupportedpastels 323168a289 fix(dashboard): correct copilot-acp sign-in command, honest status card
The Accounts-tab card told users to run 'copilot /login', which is not a
valid invocation — slash-commands only exist inside an interactive session.
Use 'copilot login', the CLI's device-code login subcommand.

The card's status_fn also hardcoded logged_in: False with a static label.
Wire it to get_external_process_provider_status(): claim logged_in only on
positive credential evidence (auth_verified), show which executable Hermes
resolved when merely configured, and say so when the CLI is missing from
PATH entirely.

The rendered cli_command now substitutes the executable the user actually
configured (HERMES_COPILOT_ACP_COMMAND / COPILOT_CLI_PATH) so a custom
binary path gets a copy-pasteable command that matches what Hermes spawns.
2026-09-02 20:51:07 +05:30
kshitijk4poor 95f62ca3bf fix(cron): validate failure_deliver at preflight and dashboard update lanes
Follow-up to the failure_deliver salvage (#100375):

- _preflight_check_delivery also checks the failure lane, so a typo'd
  failure_deliver platform blocks at config-validation time instead of
  surfacing only when a failure occurs — exactly when the notice must
  not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
  deliver (text normalization; empty clears the optional override
  instead of coalescing), closing the one update path that could write
  an unnormalized value into jobs.json.

4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
2026-09-02 20:16:14 +05:30
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 4e3feb8bbb feat(tts): speech toggles warm up and unload local TTS engines (#100881)
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.

- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
  acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
  registry; piper/kittentts loaders extracted so warm-up and synthesis
  share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
  failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
  intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.

Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
2026-09-01 21:43:59 -07:00
Jeffrey Quesnelle c56f8cdd48 Merge pull request #100667 from NousResearch/feat/local-models-squash
feat: local models — managed llama.cpp runtime with one-click desktop  setup
2026-09-01 17:53:28 -04:00
Pedro Fontana b3576a29c3 Merge pull request #97354 from NousResearch/fix/nous-org-model-policy
fix(nous): honour the org model policy in the model pickers
2026-09-01 18:20:18 -03:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
David Metcalfe 6545812c86 fix(dashboard): use config-only scope for /api/model/options to prevent lock-contention freeze
_profile_scope holds _SKILLS_PROFILE_LOCK (threading.RLock) across the
entire context-manager yield.  When get_model_options' worker thread
blocks on fetch_models_dev → requests.get() (up to 15s on a models.dev
cache miss), the lock stays held for the full duration.  Concurrent
requests to /api/config (get_config also enters _profile_scope) then
block the main event-loop thread on the RLock, freezing the server.

Switch to _config_profile_scope which uses only the contextvar-based
HERMES_HOME override (thread-safe, no lock) — sufficient for the config
reads + credential checks that build_model_options_payload needs, and
already used by other await-safe endpoints.

Refs #58576
2026-09-01 12:07:00 -07:00
xxxigm c99a3919a8 fix(dashboard): drop systemd-only advice from the model-picker code-skew 503 (#97046)
Desktop-owned serve and macOS hosts have no hermes-dashboard unit, so the
restart hint now follows HERMES_SERVE_HEADLESS instead of hardcoding systemctl.
2026-08-31 20:43:36 -05:00
Mariano Nicolini 6e20ec4101 fix(nous): apply the org policy before the free/paid tier split
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
2026-08-31 15:40:49 -03:00
Agi-Asi 54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
chelsealong 6d407ca1a4 fix(desktop): stop model_context_length edits from being dropped or wiped
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.

Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
2026-08-31 10:07:17 -07:00
Alvin T. Veroy e17fd0a708 fix(state): decode errors now reach the heal path and fail loud in TUI (residual #98924 surfaces)
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:

- web_server._open_session_db_at_path: the one-writable-open heal only
  caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
  to decode SQLite's own error message over corrupt file bytes) bypassed
  it, so the heal documented for malformed schema never fired (#98924
  Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
  vtables whose probe raised UnicodeDecodeError, the same too-narrow
  catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
  could not open, so prompt.submit streamed the turn while persisting
  nothing (#98924 Failure 2). It now returns False and prompt.submit
  fails the RPC with code 5072 so desktop maps it to a toast, mirroring
  the disk-full/5070 convention. session.create stays silent per its
  pinned degraded-mode contract.
2026-08-31 09:56:43 -07:00
Kshitij Kapoor c26762d60e fix(desktop): keep the orphan sweep unconditional; guard the probe fix with a mutation-checked test
Follow-up on the salvaged #87158: the reaper-side fix (probe with
cleanup_stale=False so the sweep never deletes its own exclusion
evidence) fully protects a healthy standalone gateway, so the
web_server.py skip-sweep gate is dropped — a stale-but-present
registration must not veto the #77276 orphan reap that motivated the
sweep in the first place.

New regression test drives the exact failure: a registration that
fails liveness validation only surfaces its PID when probed
non-destructively; with the destructive default the standalone gateway
would be hard-killed (TerminateProcess, no drain). Mutation-checked:
reverting the probe to get_running_pid() fails the test.
2026-08-31 21:09:43 +05:30
butbutbutbutbutbut 1ec4b569f2 fix(desktop): don't reap the healthy standalone gateway on Windows desktop startup 2026-08-31 21:09:43 +05:30
Wesley Simplicio 8edaa25746 fix(cron): isolate desktop profile persistence 2026-08-31 07:28:22 -07:00
teknium1 8fd144c502 fix(desktop): model assignment carries the credential pointer, not a resolved key (#88990, salvage #90484)
Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
2026-08-31 04:40:24 -07:00
Teknium a90be562f4 fix(web): stop mirroring env-backed provider keys into model.api_key (#88990)
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
2026-08-31 03:37:43 -07:00
David Dudok de Wit 93c7089f70 feat(bot-mode): run same-gateway Group Chats without Desktop 2026-08-30 22:19:06 -07:00
David Dudok de Wit cbc67b939f feat(bot-mode): add durable Group Chat authority and replay 2026-08-30 19:46:14 -07:00
joaomarcos 0099f250c2 fix(auth): close Anthropic OAuth review gaps 2026-08-29 18:34:35 -07:00
joaomarcos e1a210652a fix(auth): harden claude_code refresh lock and remove dashboard Anthropic OAuth
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).

Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
2026-08-29 18:34:35 -07:00