The module already binds run_in_threadpool (used by list_profiles_endpoint)
and every sibling router uses the same starlette helper; the nine new
loop.run_in_executor(None, _run) sites now go through that alias so the
file has one offload idiom. Behaviour-identical (both hand the callable to
a worker thread).
Also sweeps the one endpoint the PR left synchronous:
update_profile_model_endpoint's _write_profile_model reads and rewrites
the profile's config.yaml on the event loop.
The remaining in-scope handlers in this router read and write profile
documents inline on the ASGI event loop:
- GET /api/profiles/{name}/soul reads SOUL.md
- PUT /api/profiles/{name}/soul atomic_write_text(SOUL.md)
- PUT /api/profiles/{name}/description write_profile_meta(profile.yaml)
- GET /api/profiles/{name}/desktop-overlay reads desktop.json
The persona save is the sharpest of the four: atomic_write_text() writes a
temp file, fsyncs it and replaces the original, so the loop is parked for
however long the filesystem takes to durably commit — unbounded on a slow
or contended disk, and paid on every Save in the editor.
Each handler keeps its existing status-code mapping. The reads probe and
load in a single executor hop rather than two, which also avoids widening
the gap between the existence check and the read.
Both readers return a _MISSING sentinel rather than None for an absent
file. desktop.json may legitimately contain the document `null`; collapsing
that onto None would newly report an existing-but-empty overlay as absent.
The same distinction is what the SOUL.md durability tests rely on, where
"file missing" and "file empty" must not both read as never-set.
_resolve_profile_dir() stays on the loop in all four, as it does in the
rest of this sweep: it is a name check plus one stat, and it owns the
400/404 responses.
Three more handlers in this router did filesystem work inline on the ASGI
event loop:
- PATCH /api/profiles/{name} calls rename_profile(), which stops a running
gateway through the same 10-second _stop_gateway_process() poll that
delete uses, then renames the profile directory, rewrites the Honcho
host blocks and regenerates the wrapper script.
- GET /api/profiles/active reads the active_profile state file and
resolves HERMES_HOME against the profiles root. The sidebar polls it.
- POST /api/profiles/active stats the target profile, creates the state
directory and writes active_profile via a temp file plus replace.
Rename carries the same worst case as delete and belongs off the loop for
the same reason. The two active-profile handlers are individually cheap,
but they are the routes the dashboard polls, so they are the ones most
likely to be queued behind something slower — and leaving them inline is
what made the router inconsistent with list_profiles_endpoint, which
already offloads a plain directory listing eight lines above.
The two reads in GET share one executor hop rather than taking one each.
POST /api/profiles/{name}/describe-auto called
profile_describer.describe_profile() inline. That function is a plain def;
it reaches agent.auxiliary_client.call_llm(), also a plain def, which makes
a synchronous provider request with a 60-second ceiling.
Held on the ASGI event loop that is six times the 10-second WebSocket
ready-probe threshold web_server.py records as the point where the desktop
app gives up (GH-73083). A single describe-auto on a slow or unreachable
auxiliary provider therefore takes the whole dashboard offline for up to a
minute, including the /api/ws and /api/pty sockets the desktop app and the
Chat tab run on.
Move the import and call into the default executor. _resolve_profile_dir()
deliberately stays on the loop ahead of the hop: it is a name validation
plus a single stat, and it owns the 400/404 responses that the handler's
`except Exception` would otherwise turn into a 500.
DELETE /api/profiles/{name} called profiles.delete_profile() inline on the
ASGI event loop. When the target profile has a gateway running, that call
stops it via _stop_gateway_process(), which polls the PID every 500 ms for
up to 10 s before escalating to a force kill, and then removes the profile
tree.
For the whole of that window the dashboard process serves nothing else.
web_server.py's own notes record what that costs: a stall of this length
"caus[ed] the Desktop's 10-second WebSocket ready-probe to time out
(GH-73083)", and both the desktop app and the dashboard's Chat tab drive the
agent over those WebSockets. Deleting a profile whose gateway is up is a
routine action that reliably reaches the full ten seconds — the handler's
own output announces "Gateway is running - it will be stopped".
Move the call into the default executor via run_in_executor, matching
list_profiles_endpoint, export_profile_endpoint and import_profile_endpoint
in this same module. The exception-to-status mapping is unchanged:
FileNotFoundError/ValueError are raised inside the worker and re-raised by
the await, so they still map to 404/400.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
`projects.tree` answers for the backend's own profile, so the grouped
sidebar had nothing to draw once the user asked to see every profile.
Run the same authoritative builder once per profile against that
profile's state.db and merge the results by folder, so one checkout is
one group no matter how many profiles work in it, and the owning profile
rides on each session row where the badge and filter can read it.
Group totals are summed in SQL rather than over the loaded page — a
number that shrank as you scrolled would be worse than no number.
Scope the batched sidebar slices while we're here: cron and messaging
came back cross-profile unconditionally, which is why a concrete profile
showed another profile's Telegram threads and cronjobs.
Closes#65710Closes#42651Closes#70629
A session row can say whether its work is open, merged or closed, and link
to it. The join is the session's own repo + branch, asked of GitHub in one
batched GraphQL request per repo (branch aliases, not a `gh pr list` page
that a busy repo crowds ours out of), through the remote-aware git facade so
a desktop on a remote gateway asks the backend's `gh`.
Two ways a session's branch can't answer, both covered:
- It ran on trunk. Fork PRs share our branch namespace, so asking about
`main` badges a stranger's PR onto it — trunk is never asked about, and
cross-repository PRs are dropped server-side either way.
- It worked in a worktree, so the branch it recorded at start isn't where
the PR came from. Creating a PR from the review pane binds the session to
the branch it actually used, and for sessions that predate that, the PR is
recovered from the transcript: `gh pr create` prints a bare PR url and
nothing else, so a tool result whose whole output is one is a claim rather
than a mention. Scanned read-only across profiles, once per session ever.
After `hermes update`, the desktop sidebar showed "No sessions yet" until
the user's first message. #72424 added sessions.last_activity_at, which
list_sessions_rich now selects — but column adds only land through
_reconcile_columns() in the writable _init_schema, and read-only opens
skip that by design. Every sidebar read path opens state.db read-only, so
each poll raised "no such column: s.last_activity_at" until the first
prompt's lazy session-row persist forced a writable open and reconciled.
A heal for exactly this class already existed (_open_session_db_for_profile
probes the read-only handle and does a one-time writable reopen on
staleness), but its probe was a hand-written four-column list that never
learned last_activity_at — it went stale three days after shipping. And the
batched sidebar route (/api/profiles/sessions/sidebar) bypassed the helper
entirely, swallowing per-profile failures into an errors array the desktop
never surfaces, so the incident produced an empty sidebar with clean logs.
The fix removes the maintenance burden instead of paying it once more:
- hermes_state_schema.schema_read_probe_statements() derives one
`SELECT <every declared column> FROM <table> LIMIT 0` per table from
SCHEMA_SQL via the existing _parse_schema_columns() — the same source of
truth the writable reconciler diffs against, so any future ADD COLUMN is
probed with no list to update. Column references are table-qualified:
an unqualified double-quoted identifier that fails to resolve silently
degrades to a string literal (SQLite's double-quoted-string misfeature)
and would make the probe pass on exactly the store it exists to catch.
- web_server splits the heal into a path-level _open_session_db_at_path
(semantics unchanged) so the cross-profile session routes can share it;
both profiles.py loops and _count_status_active_sessions (the remaining
raw read-only sibling) now open through it. The heal stays a helper
rather than a SessionDB classmethod on purpose: escalation-to-writable
must remain an explicit caller decision — update_cmd.py opens read-only
mid-update and must never write.
- Exhaustion guard: if the writable heal SUCCEEDS and the re-probe still
fails (a schema problem ADD COLUMN cannot express), the store is marked
exhausted — warn once, skip the probe, serve reads probe-less — instead
of re-running the full writable init on every poll against a possibly
live DB. A FAILED writable open (transient lock) is deliberately not
recorded, so the next poll retries the heal.
- The per-profile swallow sites in profiles.py now also log a deduplicated
warning, so a persistent read failure is loud in errors.log even though
the response errors array stays invisible to the sidebar.
Tests: probe/SCHEMA_SQL coverage invariants (tests/test_schema_read_probe.py),
last_activity_at added to the /api/sessions heal parametrize, a sidebar-route
heal test reproducing the shipped symptom (errors == [] and the session
returned against a store missing the column), and an exhaustion test pinning
exactly one writable open. The sidebar and last_activity_at tests fail on
main.
Follow-up to the salvaged #79323 commits. The three hand-rolled
stat -> atomic_write_text -> chmod blocks (xai migration, uninstaller
shell-rc rewrite, dashboard SOUL.md editor) collapse into an opt-in
preserve_mode=True kwarg on utils.atomic_write_text, plus create_mode=
on both atomic_write_text and atomic_yaml_write for first-create paths
(SOUL.md first save, write_manifest's allowlist create path).
Beyond deduplication this closes two gaps the hand-rolled copies had:
- Owner preservation: the old in-place writes kept the inode, so file
ownership survived root-run rewrites for free. atomic_write_text
swaps in a new inode owned by the writing user, and the hand-rolled
blocks restored only the mode -- a root-run 'hermes migrate xai' or
sudo uninstall on a user-owned Docker/NAS volume would flip
config.yaml / ~/.zshrc ownership to root. preserve_mode now routes
through the same _preserve_file_owner/_restore_file_owner helpers
atomic_yaml_write and atomic_json_write already use.
- chmod-after-replace window: the mode is applied to the temp fd via
fchmod BEFORE the replace (mirroring atomic_json_write's mode= param),
so the target never transits through mkstemp's 0600.
Also removes write_manifest's caller-side existed/chmod block (and its
small TOCTOU) in favor of atomic_yaml_write(create_mode=0o644), and
corrects the SOUL.md mode comment (the default profile's runtime seeder
does run it through _secure_file; named profiles do not).
preserve_mode defaults to False so the existing callers (memory store,
skill manager, cron, agent importer) keep their current semantics.
New tests in tests/test_atomic_write_text_metadata.py cover mode
preservation, owner restore through symlinks, fchmod-before-replace,
create_mode on both writers, and no-behavior-change without opt-in;
all mutation-checked.
Both files were routed through the shared atomic writers earlier in this
branch. tempfile.mkstemp creates the temp file 0600 and the atomic swap
carries that mode onto the target, so the *create* paths silently tightened
two files that previously landed at the umask default:
- web_routers/profiles.py: the dashboard persona editor's first-ever Save has
no prior SOUL.md to copy permissions from, so the existing guard skipped the
chmod entirely -- contradicting the comment directly below it, which states
profile SOUL.md is created 0644 and is not run through _secure_file.
- profile_distribution.py: atomic_yaml_write only restores a mode it captured
from a file that already existed. _materialize() calls write_manifest() with
no manifest on disk whenever a distribution declares an explicit
distribution_owned allowlist that omits distribution.yaml, so the staged
copy is never placed in the profile.
Both are fixed with a local chmod at the two sites this branch regressed;
utils.py's public mode semantics are left alone. profiles.py now also
distinguishes "no file yet" (FileNotFoundError -> 0644) from "stat failed for
some other reason" (-> leave the mode alone rather than guess at it).
uninstall.py and xai_retirement.py have no create path and are unchanged: the
former captures prior_mode unconditionally after a successful read_text(), and
the latter runs require_readable_config_before_write() first.
`utils.atomic_write_text`'s docstring states the invariant: it exists "so that
every destructive file rewrite in the codebase shares one implementation."
Four full-file rewrites of *existing user-authored files* still bypass it and
use a bare truncating `open(path, "w")` / `Path.write_text()`, which truncates
the target before the new content is produced. A crash, SIGINT, or ENOSPC
mid-write therefore leaves the file empty or half-written.
In all four cases the read half degrades silently to a default rather than
erroring, so the damage is invisible and the next write cements it:
* `xai_retirement.apply_migration()` rewrites the user's config.yaml. Merged
commit beaa1a08e added a readability guard here and noted the writer "lives
outside the atomic_yaml_write path, so the chokepoint didn't cover it"; this
closes the durability half it left open. `--no-backup` is a documented flag,
so on that path the truncated file is the only copy that exists, and the
loader returns early on `doc is None` — the next run reports nothing to
migrate rather than surfacing the damage.
* `uninstall.remove_path_from_shell_configs()` rewrites the user's shell rc
(~/.bashrc, ~/.zshrc, ...). Hermes does not own these files and this function
takes no backup; the enclosing `except Exception` downgrades a partial write
to a warning, so the next login just starts a bare shell.
* `web_routers.profiles.update_profile_soul()` replaces SOUL.md from the
dashboard editor. The paired GET reports an unreadable file as
`{"content": "", "exists": False}`, so an interrupted save presents as "your
persona was never set" and the editor's next Save persists the empty document.
* `profile_distribution.write_manifest()` rewrites distribution.yaml on every
install/update. `read_manifest` treats an unparseable manifest as "not a
distribution", silently dropping update tracking and env_requires.
The xAI migration keeps its ruamel round-trip dumper (comments, key order and
quoting must survive) and now serializes to a string before handing the bytes
to the shared writer. `write_manifest` moves to `atomic_yaml_write`, whose
SafeDumper output the manifest already round-trips through, retiring the local
`_dump_yaml` helper.
`atomic_write_text` recreates the target from a 0600 temp file, so each of its
call sites re-applies the file's previous permission bits: `_secure_file`
deliberately leaves config.yaml alone under managed (NixOS 0640) and container
installs, shell rc files are normally 0644, and profile SOUL.md is created 0644
and never secured. `atomic_yaml_write` already preserves mode and owner itself.
Routing through `atomic_replace` also keeps a symlinked config.yaml or ~/.zshrc
(dotfiles repo, managed deployment) pointing at the real file.
Tests: one regression test per site fails on clean main (the interrupted write
completes there and destroys the file) and passes here; the remaining cases are
behaviour guards covering symlink survival, permission preservation, comment
round-tripping, and the existing happy paths.
export_profile() accepts extra_files (root-relative filename -> text) so a
caller can stage companion files into the archive; the desktop uses it for
desktop.json, its appearance/interface overlay, now part of the default
profile's export allow-list.
New routes wrapping the existing hermes profile export/import machinery:
- POST /api/profiles/{name}/export (extra_files + optional output path)
- POST /api/profiles/import (returns the bundled desktop overlay)
- GET /api/profiles/{name}/desktop-overlay
Paths cross the API, not bytes - the desktop's native dialogs and its
local/pooled backends share a filesystem.
le=100 would 422 real desktop callers: sessions-settings fetches
archived at limit=200, the command palette lists at 200, and the
electron remote-merge over-fetches limit+offset (exceeds 100 at
offset>=81, and its .catch(()=>null) silently drops remote sessions).
Clamp must sit above real client maxima. New test pins limit=200 w/
offset.
Re-derivation of aydnOktay's twin clamp PRs onto current main (the
session-list endpoints moved into web_routers/; the analytics endpoints
gained asyncio.to_thread wrappers since the originals):
- limit le=100 on /api/sessions, /api/sessions/search and the
/api/profiles/sessions fan-out (one unbounded request could drag every
session row + correlated-subquery preview work out of SQLite, times
every profile's state.db on the fan-out).
- days ge=1 le=365 on /api/analytics/usage + /api/analytics/models
(huge or non-positive values force full-history InsightsEngine work or
inverted windows; the UI only offers 7/30/90 presets).
FastAPI Query bounds reject at the validation layer (422). 8 new tests;
both clamp classes mutation-checked (clamp removed -> its tests fail).
Ports the negative limit/offset fix onto the current router modules
(hermes_cli/web_routers/sessions.py, profiles.py) since the handlers
moved out of web_server.py in 011ec4513e after this PR was opened.
Per review feedback: only add Query(..., ge=0) — no le=500. The
messages route already clamps oversized requests with min(limit, 500)
and must keep that behavior (succeed + cap) rather than reject them;
the two session-list routes never had a public 500 cap and shouldn't
gain a new rejecting one as a side effect of this fix.