Commit Graph

79 Commits

Author SHA1 Message Date
Teknium 7c0d8f2c6f refactor(web): second-pass compaction of local_models/profiles/web_git (helpers, dict literals, docstrings) 2026-09-02 13:35:48 -07:00
Teknium de626c5a36 refactor(web): unify profiles router error mapping, target enumeration and row tagging 2026-09-02 13:34:56 -07:00
Teknium 8c5af5733a refactor(web): collapse local_models job/download/server boilerplate into shared helpers 2026-09-02 13:33:24 -07:00
Teknium 4fd970c62f refactor(web): absorb single-router web_server helpers into their routers (75 defs); drop dead imports 2026-09-02 13:33:22 -07:00
Teknium c90af45f57 refactor(web): unify the three post-config gateway auto-restart helpers 2026-09-02 13:32:52 -07:00
Teknium ba4f4b8161 refactor(web): collapse try/except-HTTPException/log/raise boilerplate into _common.http_failure; trim unused fastapi imports 2026-09-02 13:32:46 -07:00
Teknium 8ac12223bb refactor(web): late() proxies replace per-handler lazy web_server imports in extracted routers 2026-09-02 13:32:45 -07:00
Teknium ba91166b1c refactor(web): hoist stdlib imports in extracted routers; drop unused web_server imports 2026-09-02 13:32:28 -07:00
Teknium e4d8a78967 refactor(web): move /api/logs into web_routers/status.py (logs_router) 2026-09-02 13:32:08 -07:00
Teknium de109d4097 refactor(web): extract chat-tab WebSocket routes into web_routers/chat_ws.py 2026-09-02 13:32:06 -07:00
Teknium 4bae44d654 refactor(web): extract dashboard theme/font/plugin routes into web_routers/dashboard_ui.py 2026-09-02 13:32:05 -07:00
Teknium 94309ad40a refactor(web): extract raw-config + analytics routes into web_routers/analytics.py 2026-09-02 13:32:02 -07:00
Teknium 067028eb22 refactor(web): extract OAuth provider routes into web_routers/oauth.py 2026-09-02 13:32:01 -07:00
Teknium d3b768232e refactor(web): extract status/health/curator/learning/portal routes into web_routers/status.py 2026-09-02 13:30:56 -07:00
Teknium 094dba5a2f refactor(web): extract gateway/update/action-status routes into web_routers/actions.py 2026-09-02 13:30:55 -07:00
Teknium 9de17448dd refactor(web): extract audio/TTS routes into web_routers/audio.py 2026-09-02 13:30:53 -07:00
Teknium feeb1074a3 refactor(web): extract config/env/custom-endpoint routes into web_routers/config_env.py 2026-09-02 13:30:53 -07:00
Teknium 1f4ab19eae refactor(web): extract model assignment routes into web_routers/models.py 2026-09-02 13:30:53 -07:00
Teknium 11bc6bbe11 refactor(web): extract memory-provider setup routes into web_routers/memory_providers.py 2026-09-02 13:30:51 -07:00
Teknium 1f77a129da refactor(web): extract messaging onboarding/platform routes into web_routers/messaging.py 2026-09-02 13:30:50 -07:00
Teknium 0cca195418 refactor(web): extract pairing/webhooks/credential-pool/memory/ops routes into web_routers/ops.py 2026-09-02 13:30:50 -07:00
Teknium 3d5f9269c7 refactor(web): extract files/fs/media routes into web_routers/files.py 2026-09-02 13:30:05 -07:00
Teknium f0f7f55f76 refactor(web): dashboard_auth/web_routers/kanban dedupe + _common helpers (resume — verified partial work) 2026-09-02 13:30:05 -07:00
kshitijk4poor 2b4e70ec07 refactor(cli): use the router's run_in_threadpool alias; offload profile model write
The module already binds run_in_threadpool (used by list_profiles_endpoint)
and every sibling router uses the same starlette helper; the nine new
loop.run_in_executor(None, _run) sites now go through that alias so the
file has one offload idiom. Behaviour-identical (both hand the callable to
a worker thread).

Also sweeps the one endpoint the PR left synchronous:
update_profile_model_endpoint's _write_profile_model reads and rewrites
the profile's config.yaml on the event loop.
2026-09-03 01:56:47 +05:30
briandevans 34a8e1dc50 fix(cli): run profile document I/O off the dashboard event loop
The remaining in-scope handlers in this router read and write profile
documents inline on the ASGI event loop:

- GET  /api/profiles/{name}/soul            reads SOUL.md
- PUT  /api/profiles/{name}/soul            atomic_write_text(SOUL.md)
- PUT  /api/profiles/{name}/description     write_profile_meta(profile.yaml)
- GET  /api/profiles/{name}/desktop-overlay reads desktop.json

The persona save is the sharpest of the four: atomic_write_text() writes a
temp file, fsyncs it and replaces the original, so the loop is parked for
however long the filesystem takes to durably commit — unbounded on a slow
or contended disk, and paid on every Save in the editor.

Each handler keeps its existing status-code mapping. The reads probe and
load in a single executor hop rather than two, which also avoids widening
the gap between the existence check and the read.

Both readers return a _MISSING sentinel rather than None for an absent
file. desktop.json may legitimately contain the document `null`; collapsing
that onto None would newly report an existing-but-empty overlay as absent.
The same distinction is what the SOUL.md durability tests rely on, where
"file missing" and "file empty" must not both read as never-set.

_resolve_profile_dir() stays on the loop in all four, as it does in the
rest of this sweep: it is a name check plus one stat, and it owns the
400/404 responses.
2026-09-03 01:56:47 +05:30
briandevans 63d42cd0e2 fix(cli): run profile rename and active-profile state off the event loop
Three more handlers in this router did filesystem work inline on the ASGI
event loop:

- PATCH /api/profiles/{name} calls rename_profile(), which stops a running
  gateway through the same 10-second _stop_gateway_process() poll that
  delete uses, then renames the profile directory, rewrites the Honcho
  host blocks and regenerates the wrapper script.
- GET /api/profiles/active reads the active_profile state file and
  resolves HERMES_HOME against the profiles root. The sidebar polls it.
- POST /api/profiles/active stats the target profile, creates the state
  directory and writes active_profile via a temp file plus replace.

Rename carries the same worst case as delete and belongs off the loop for
the same reason. The two active-profile handlers are individually cheap,
but they are the routes the dashboard polls, so they are the ones most
likely to be queued behind something slower — and leaving them inline is
what made the router inconsistent with list_profiles_endpoint, which
already offloads a plain directory listing eight lines above.

The two reads in GET share one executor hop rather than taking one each.
2026-09-03 01:56:47 +05:30
briandevans 4da5689b80 fix(cli): run auto-describe LLM round-trip off the dashboard event loop
POST /api/profiles/{name}/describe-auto called
profile_describer.describe_profile() inline. That function is a plain def;
it reaches agent.auxiliary_client.call_llm(), also a plain def, which makes
a synchronous provider request with a 60-second ceiling.

Held on the ASGI event loop that is six times the 10-second WebSocket
ready-probe threshold web_server.py records as the point where the desktop
app gives up (GH-73083). A single describe-auto on a slow or unreachable
auxiliary provider therefore takes the whole dashboard offline for up to a
minute, including the /api/ws and /api/pty sockets the desktop app and the
Chat tab run on.

Move the import and call into the default executor. _resolve_profile_dir()
deliberately stays on the loop ahead of the hop: it is a name validation
plus a single stat, and it owns the 400/404 responses that the handler's
`except Exception` would otherwise turn into a 500.
2026-09-03 01:56:47 +05:30
briandevans a2504a0a59 fix(cli): run profile deletion off the dashboard event loop
DELETE /api/profiles/{name} called profiles.delete_profile() inline on the
ASGI event loop. When the target profile has a gateway running, that call
stops it via _stop_gateway_process(), which polls the PID every 500 ms for
up to 10 s before escalating to a force kill, and then removes the profile
tree.

For the whole of that window the dashboard process serves nothing else.
web_server.py's own notes record what that costs: a stall of this length
"caus[ed] the Desktop's 10-second WebSocket ready-probe to time out
(GH-73083)", and both the desktop app and the dashboard's Chat tab drive the
agent over those WebSockets. Deleting a profile whose gateway is up is a
routine action that reliably reaches the full ten seconds — the handler's
own output announces "Gateway is running - it will be stopped".

Move the call into the default executor via run_in_executor, matching
list_profiles_endpoint, export_profile_endpoint and import_profile_endpoint
in this same module. The exception-to-status mapping is unchanged:
FileNotFoundError/ValueError are raised inside the worker and re-raised by
the await, so they still map to 404/400.
2026-09-03 01:56:47 +05:30
Brooklyn Nicholson 3d81650c2f fix(desktop): a failed sidebar scan keeps the rows it could not re-read
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.

Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.

The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.

Part of #73847
Part of #88528

Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
2026-09-01 22:42:19 -05:00
Brooklyn Nicholson 3a0e7df799 fix(state): a busy session store reads as busy, not as damaged or empty
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.

Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.

Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.

On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.

Fixes #100436

Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
2026-09-01 22:42:19 -05:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium a1c25d393a feat(desktop): built-in optional-skills catalog in Capabilities → Skills with one-click install
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.

- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
  scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
  install-state spinners off $hubActions, and an OfficialSkillDetail pane
  (hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
2026-09-01 01:39:59 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Nio Thomas ba7743b076 fix(state-db): report corruption instead of "session not found", detect it early
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.

1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
   errors via the existing is_malformed_db_error() and raises 503 at all five
   call sites. delete_session_endpoint was the worst — an unresolvable id
   counted as idempotent success, so DELETE reported it had removed a session
   that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
   quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
   verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
   for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
   before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
   not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
   mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
   _await_gateway_exit() that also re-checks after the final sleep (a PID
   exiting in the last interval must not be SIGKILLed — PID-reuse hazard).

NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.

Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
2026-08-31 09:56:54 -07:00
joaomarcos 0bfe715e4d fix(sessions): don't let the empty-session sweep delete an archived transcript
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:

  * `replace_messages(..., archive_dropped=True)` — the rewind / edit /
    regenerate mode added in #82756 so a taken-back turn stays recoverable.
  * `archive_and_compact` — in-place compaction, which archives the
    pre-compaction transcript under the same session id (#38763).

A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.

Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.

Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.

Fixes #95868

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
2026-08-27 19:54:32 +05:30
Teknium 01f7ce5b76 feat(sessions): one-shot single-match owner backfill for legacy NULL-profile rows (#94724)
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.

Refs #94724
2026-08-27 02:17:56 -07:00
fangliquanflq cb54576b1a fix(gateway): isolate control routes from default executor 2026-08-26 21:38:20 -07:00
unsupportedpastels 08cf4fea5d fix(mcp): restrict catalog environment writes 2026-08-26 09:54:31 -07:00
fangliquanflq 66186dc58f fix(desktop): keep bot reconciliation off inactive backends 2026-08-26 00:21:29 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Teknium a1682376ca feat(profiles): rename any agent — the default profile gets a display name (#45624)
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.

Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").

Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
2026-08-18 02:27:18 -07:00
Teknium 2c2697b52e feat: widen Tier 1 advisory scan to license + security checks
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.

- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
  static (no measurable cost); security invokes NVIDIA SkillSpector in
  its keyless static-rules mode (~+1.2s per install). schema/quality
  stay excluded: hygiene signal ("author not specified" is
  high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
  or failing, the security check reports status="incomplete" and the
  adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
  note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
  This also absorbs a live upstream inconsistency: SkillEvaluator's
  anti-tamper cross-check on SkillSpector's risk score currently trips
  on moderate-finding skills (fail verdict with zero findings, e.g.
  github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
  NVIDIA separately; either way an evidence-free fail must not render
  as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
  not-run formatting.

E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
2026-08-17 17:04:40 -07:00
Teknium 183f18d530 feat: advisory NVIDIA SkillEvaluator Tier 1 scan on skill installs
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).

- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
  over the quarantined bundle, parses the JSON report, classifies
  secrets-class findings (private keys, tokens, credentialed connection
  strings) apart from advisory PII findings. Every failure mode
  (binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
  guard's policy decision and before the install confirmation. Findings
  are shown with file:line; secrets-class findings render red with a
  loud warning. Warn-and-continue by design — the built-in skills guard
  remains the only enforcement layer, because the upstream PII scanner
  has known false-positive classes (git@github.com, docs example
  emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
  endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
  optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
  config toggle.

Scanner install (optional):
  uv tool install --python 3.13 \
    "skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"

E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
2026-08-17 17:04:40 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
Shannon Sands 4822156923 fix(cron): stop retry storms when the gateway is deliberately stopped (OOF-266)
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).

Split the unreachable path by durable operator intent:

- desired_state == "stopped" (written only by the s6 lifecycle commands;
  the same intent signal container-boot reconciliation trusts) -> drop
  the fire with 200 + a structured log line, mirroring NAS's own
  instance_stopped drop. Jobs are not lost: the Chronos provider
  reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
  file without desired_state) -> keep the retryable 503, now stamped
  with Retry-After: 60 so a scheduler that honors it spaces retries
  past the wake/restart window instead of exhausting them inside it.
  The gateway's own pass-through 503s (draining) get the same hint.

The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
2026-08-16 19:50:07 -07:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
Teknium e22fa90769 feat(desktop): MCP fleet cost/usage overlay with schema token estimates and 30-day usage
Each configured server row on the MCP Capabilities page now shows what it
costs and whether it earns its keep:

- ~per-call token estimate of the server's tool schemas, summed over ENABLED
  tools only (ceil(schema_chars/4) via the existing include/exclude filter)
- 30-day usage count from getUsageAnalytics(30), cached per scope profile
  like the Toolsets tab's toolCallsCache, mapped to servers via the
  mcp__<server>__<tool> registry-name convention (tools/mcp_tool.py)
- a subtle muted "unused" pill on enabled, probed-ok servers with nonzero
  schema cost and zero 30-day uses — never a dialog

Backend: the /api/mcp/servers/{name}/test probe now fills an additive
per-tool `schema_chars` (length of the SAME converted registry schema the
agent registers). Older backends omit it → renderer shows counts only;
older renderers ignore the extra key. Display-only: nothing changes what
schemas are sent to models, no config knobs.

i18n keys (costTokens/usage30d/unusedPill) added to types/en/zh/zh-hant/ja
(ar inherits en via defineLocale overrides). Pure math lives in
lib/mcp-cost.ts with unit tests; Python wire shape pinned in
tests/hermes_cli/test_web_server_profile_unification.py.
2026-08-16 02:24:32 -07:00
Teknium fc8ebff6d8 feat(mcp): unify the desktop MCP suggestion directory into the catalog
The desktop app carried its own hardcoded list of 17 vendor MCP endpoints
(apps/desktop/src/lib/mcp-directory.ts) powering the composer suggestion
pills — a second PR-reviewed vendor list, overlapping and drifting from the
Nous-approved MCP catalog (optional-mcps/).

This makes the catalog the single source of truth:

- manifest schema: optional `suggest:` block (keywords + hosts), parsed,
  validated, and normalized in mcp_catalog.py
- 15 new URL-only hosted-remote catalog entries (atlassian, sentry, datadog,
  notion, stripe, vercel, supabase, netlify, hugging_face, asana, intercom,
  airtable, webflow, paypal, square); figma + linear manifests gain suggest
  blocks
- GET /api/mcp/catalog now serves the suggest metadata
- desktop suggestion provider builds its match index from the catalog;
  the static directory remains only as a compatibility rung for older
  backends without suggest metadata
- setup card source line prefers the catalog entry's transport URL

GitHub stays out of the catalog on purpose: its hosted MCP rejects generic
DCR and the bundled github/* skills (gh CLI) are the stronger integration.
New desktop `github` suggestion provider offers the github-auth skill
instead — gated on a new cached GET /api/git/gh-auth probe so already-
authenticated users never see the pill.
2026-08-15 02:03:24 -07:00