Commit Graph

14240 Commits

Author SHA1 Message Date
zengzheqing 37ae0e1d39 fix(curator): remove terminal from the consolidation fork (issue #96962)
The curator LLM fork was steered by its own prompt to re-home skill
support files with terminal `mkdir -p ... && mv ...`. A terminal move
writes the same bytes with NO ledger entry, so the archive that follows
snapshots an already-stripped package (files: 1) and `hermes curator
rollback` restores a hollow skill — SKILL.md back, references/ gone.

Remove the capability rather than guard it: the fork's enabled_toolsets
drops "terminal", so terminal and process disappear together and there
is no shell to parse, no process stdin to feed, no remote-backend
divergence — a heuristic command guard over a Turing-complete input
space can guarantee none of that. Every mutation the pass needs has a
ledgered skill_manage action (write_file / remove_file / delete), and
the prompt now steers exactly those. Reading works through skill_view.

Tests pin both halves: the call-site kwarg (["skills"] only), the
resolved surface (no execution/write tools), and the prompt steering
(no mkdir -p / mv shapes).
2026-08-31 10:08:13 -07:00
zengzheqing 98e2f110cf fix(curator): restore complete skill packages on ledger rollback (#96962)
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.

The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.

Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
2026-08-31 10:08:13 -07:00
Teknium 6fe933e709 fix(cli): launch-context-independent Linux desktop-entry Exec (salvaged from #94874)
Rewrites resolve_exec_command so the generated .desktop Exec no longer
depends on how the installer happened to be launched: fixes the bare
repo-script form whose shebang escapes the venv, and the symlinked-venv
form that .resolve() dereferenced into the base interpreter store.

Salvaged squashed from PR #94874 (24 commits) after the original branch
was found to carry stray __pycache__/.gitignore payload.

Co-authored-by: Gökhan <gkhn.yldrmlr@gmail.com>
2026-08-31 10:07:51 -07:00
Agi-Asi 54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
chelsealong 6d407ca1a4 fix(desktop): stop model_context_length edits from being dropped or wiped
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.

Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
2026-08-31 10:07:17 -07:00
Teknium c8329384a9 test(send_message): drop duplicate buzz UUID target tests
Dispatch cluster (#99431) landed equivalent coverage first; the media
branch's copies shadowed them and tripped
test_no_shadowed_test_definitions.
2026-08-31 10:06:34 -07:00
Teknium c816957a43 fix(buzz): reconcile media pipeline with landed dispatch + threading contracts
Post-rebase composition over #99431/#99429/#99427: file-attachment sends
route through _run_message_send so the mention-recovery ladder covers
media captions; _send_file_attachment/_send_local_file honor the
resolved thread-root anchor and reply_to_mode opt-out; send() records
event_meta on the verified receipt id (#75826); test fakes gain the
auth_tag kwarg and accepted-receipt shape.
2026-08-31 10:06:34 -07:00
EmpireOperating a55d66b79a fix(buzz): redact media paths before bounding errors 2026-08-31 10:06:34 -07:00
EmpireOperating fcd34e57a2 fix(buzz): complete media-only delivery reporting 2026-08-31 10:06:34 -07:00
EmpireOperating 9c25704257 fix(buzz): verify live media delivery receipts 2026-08-31 10:06:34 -07:00
EmpireOperating 37c943997b fix(buzz): support media in standalone sends 2026-08-31 10:06:34 -07:00
phil baker fafc3ddce5 fix: deliver Buzz media as native attachments 2026-08-31 10:06:34 -07:00
Riyaaz ebda1fbf8e fix(buzz): deliver local images through native upload 2026-08-31 10:06:34 -07:00
EmpireOperating f857f13b30 test(buzz): isolate authorization cases from CLI lookup 2026-08-31 10:06:34 -07:00
EmpireOperating 366506205a fix(gateway): require boolean authorization decisions 2026-08-31 10:06:34 -07:00
EmpireOperating d15cbbcbf3 fix(buzz): gate inbound attachment side effects 2026-08-31 10:06:34 -07:00
EmpireOperating 00394acfae fix(buzz): ingest verified native attachments 2026-08-31 10:06:34 -07:00
Mathias Gorf aaad054330 fix(buzz): gate authenticated inbound media on explicit authorization
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.

`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.

Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
2026-08-31 10:06:34 -07:00
joelbrilliant 55136adcc4 fix(buzz): preserve inbound media captions 2026-08-31 10:06:34 -07:00
joelbrilliant bce94cc1b9 fix(buzz): localize inbound relay media 2026-08-31 10:06:34 -07:00
fangliquanflq f10a231efa fix(scripts): clarify Windows update retry marker semantics 2026-08-31 10:05:13 -07:00
fangliquanflq fd24ac94e0 fix(update): resume deferred Windows desktop updates 2026-08-31 10:05:13 -07:00
fangliquanflq 915ec169a0 fix(update): verify failed restore cleanup 2026-08-31 10:04:56 -07:00
fangliquanflq 53cf38c3b8 fix(update): fail closed on incomplete restore checks 2026-08-31 10:04:56 -07:00
fangliquanflq 3e4dc5f2c3 fix(update): preserve unknown restore cleanup state 2026-08-31 10:04:56 -07:00
fangliquanflq 8236b51878 fix(update): authenticate import health markers 2026-08-31 10:04:56 -07:00
fangliquanflq 6c608e2f59 fix(update): reject terminated import probes 2026-08-31 10:04:56 -07:00
fangliquanflq 68f9681acb fix(update): capture terminating restored imports 2026-08-31 10:04:56 -07:00
fangliquanflq f6c3942957 fix(update): compare every restored module failure 2026-08-31 10:04:56 -07:00
fangliquanflq 7b958b3575 fix(update): detect restored import-time failures 2026-08-31 10:04:56 -07:00
fangliquanflq 716b10314a fix(update): reject unsafe stash restores 2026-08-31 10:04:56 -07:00
Teknium e49df40682 fix(update): refuse to mutate a venv containing foreign-owned files (#83529)
A venv ever touched by sudo pip / sudo hermes contains root-owned files
(classically site-packages/*.dist-info/INSTALLER). A later normal-user
'hermes update' pulls code fine, then 'uv pip install -e .' dies with
'Permission denied (os error 13)' mid-mutation — venv/bin/hermes already
deleted, CLI bricked.

Add a bounded, pure-stat ownership preflight (_venv_foreign_owned_paths)
that runs after the code pull and immediately before the dependency
install. If foreign-owned paths are found it refuses up front, names the
offending paths + owner uid, prints the exact recovery command
(sudo chown -R $(id -un): <root>), and confirms the venv is untouched.
Windows (no os.geteuid) and root skip entirely. Never raises, capped at
~2000 stat calls, no subprocess use (update tests mock subprocess.run).

Same refuse-before-mutate philosophy as the contended-venv gate (#87331).

Fixes #83529
Diagnosis and documented recovery by @eabase.
2026-08-31 10:04:40 -07:00
walker 1f713dce38 test(packaging): exempt every exact pin from exclude-newer — release-day brick class
Each release exact-pins at least one dependency to a version published
days before the release (v0.20.6: snowballstemmer==3.1.1, psutil==7.2.2).
For two weeks after release the relative exclude-newer cutoff filters
those versions out, so any venv that predates the release cannot resolve
the new pins at all ('no version of snowballstemmer==3.1.1' — observed
2026-08-29 updating three production installs v0.20.0 -> v0.20.6, one
Termux and two Linux servers; the Termux host additionally bricked on
psutil==7.2.2 sdist resolution, and cryptography's isolated build
environment resolved maturin/setuptools-rust under the same cutoff).

Same zero-float-protection logic as the setuptools/pillow/mcp/
firecrawl-anydoc exemptions: the pin bump WAS the review, so the cutoff
adds nothing for an exact pin and can only brick. Extend
exclude-newer-package to every exact-pinned package in
[project].dependencies / optional-dependencies (table moved to
one-key-per-line — 97 entries), plus maturin and setuptools-rust for
wheel-less sdist builds of the exempted cryptography pin.

test_exact_pinned_deps_exempt_from_exclude_newer enforces the invariant
going forward: adding a name==version pin without a matching
exclude-newer-package entry fails CI.
2026-08-31 10:03:34 -07:00
Olympusbuildz a300a7fa79 fix(packaging): include wheel in PEP 517 build-system requires
Windows installer editable builds fail in uv's isolated sandbox with
ModuleNotFoundError: wheel.cli because build-system.requires only listed
setuptools. setuptools.build_meta and our setup.py bdist_wheel guard both
import wheel during the build.

Also whitelist wheel in tool.uv.exclude-newer-package so the existing
build-system exclude-newer brick guard stays green.

Fixes #96488

Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
Co-authored-by: Olympusbuildz <Olympus.roots@outlook.com>
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
2026-08-31 10:03:34 -07:00
Teknium f57802c7ad fix(state): bound the rewind-tail walk at the watermark and rewind concurrent-tail originals too 2026-08-31 10:02:22 -07:00
BrunoBza 9d9d9194d4 fix(state): archive carried-forward compaction tail as rewind rows (#86366)
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.

Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.

Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
  (_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
  message);
- micro-compaction splices [prefix, marker, suffix] - everything except
  the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
  keeping the historical archive-everything behavior.

Fixes #86366
2026-08-31 10:02:22 -07:00
xxxigm 59c5bd0913 test(install): pin the workspace selection away from apps/desktop
Runs the installer's real node_deps_workspace_args against fabricated
checkout layouts by sourcing install.sh in --manifest mode, which defines
its functions without performing an install.

The load-bearing assertion is the invariant that no checkout shape lets
apps/desktop resolve, including the empty-argument case that would silently
hand npm the whole workspace glob back.

Also cover classify_changes recovering the PR file list when compare
returns nothing, so fail-open does not demand ci-reviewed for a CLI-only
install change.
2026-08-31 10:00:55 -07:00
liuhao1024 cb89872c28 fix(install): clean up a broken managed Node and guard the termux probe
AI-review follow-up on #87467:
- On probe failure, remove the extracted ~/.hermes/node tree and the
  node/npm/npx bin links so later installer steps and retry runs start
  clean instead of resolving node to a binary that cannot start.
- The termux pkg branch had the same silent-success class: an empty
  version probe logged success and set HAS_NODE=true. Degrade with the
  binary's own error instead.
2026-08-31 10:00:46 -07:00
liuhao1024 8f8351aff3 fix(install): report a managed Node that cannot start, and preinstall libatomic1
install_node's post-install probe was
installed_ver=$(node --version 2>/dev/null) under set -e: when the
downloaded Node exists but cannot start (Node 26 linux-x64 builds link
libatomic.so.1, missing on minimal Debian/Ubuntu), the assignment
aborted the whole installer at exit 127 with the loader's explanation
discarded — installs died mid-sentence with no output at all (#87460).

- Probe now captures stderr and degrades with log_error carrying the
  loader message plus the libatomic1 hint instead of aborting.
- Debian/Ubuntu installs preinstall libatomic1 (best-effort, mirroring
  the existing apt idiom) so the common case just works.
- Termux branch's same-shaped probe gets a || true guard.
Fixes #87460
2026-08-31 10:00:46 -07:00
liuhao1024 0aa6b44917 fix(install): defer the partial clone's checkout so the throttle fallback engages
Review feedback on this PR: without --no-checkout, the blob fetch runs
inside git clone's own checkout step, so when the repo-scoped 429 hits
that fetch the whole clone exits non-zero, the else branch removes the
directory, and the fallback degrades to one more failed clone under
exactly the condition it exists for.

- Clone with --no-checkout (commits+trees only — small, passes the
  throttle); the blobs are then fetched by a separate 'git reset --hard
  HEAD' the retry can actually wrap. Verified on a local file://
  filtering remote: the no-checkout clone materializes nothing and the
  reset alone produces the full working tree.
- Fail closed: both reset attempts failing now removes the checkout and
  reports 'Failed to clone repository' instead of the previous '|| true'
  + unconditional clone_ok=true handing the installer a half-materialized
  tree printed as a success.
- The reset runs under a subshell cd so a failed materialization never
  leaves the shell in a deleted cwd, and the direct-retry loop bound now
  derives from $max_attempts (seq) instead of a hardcoded 1 2 3 4 that
  could drift from the reported attempt count.
2026-08-31 10:00:38 -07:00
liuhao1024 11afd07f16 fix(install): retry the HTTPS clone and degrade past repo-scoped 429s
GitHub throttles packfile generation for this repository with
repo-scoped HTTP 429s that are not client IP rate limits: an
anonymous clone of a small repo succeeds and the API quota is
untouched, but the single big pack behind --depth 1 dies
mid-transfer with 'RPC failed; HTTP 429 / expected packfile'. The
fresh-install clone path had no retry and no fallback, so a clean
machine exited 1 at the download stage and left a half-populated
install directory (same throttle as the update path in #89287).

Retry the HTTPS clone with linear backoff, removing the partial
clone between attempts; when every direct attempt fails, degrade
to a blobless partial clone and materialize the working tree with
a hard reset — many small packs instead of one big one, which is
what gets past the throttle. SSH-first ordering, the existing
installation update branch, and the commit-pin flow are unchanged.
2026-08-31 10:00:38 -07:00
Teknium f2f7a3bf15 feat(cron): doctor flags overdue next_run_at as silent non-firing
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
2026-08-31 10:00:29 -07:00
joe102084 b028fe632e feat: add cron doctor health check 2026-08-31 10:00:29 -07:00
Justin Adkins d957e0e403 fix(cron): bound local fire-fence waits 2026-08-31 09:59:47 -07:00
Andrew Bagrin 5cf4ffc6ae test(cron): cover execution identity and sync docs 2026-08-31 09:59:39 -07:00
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
Teknium 89bcad4d0c test: create real profile-home dirs for the multiplex yield test
The multiplex loop on current main filters profile homes through
_existing_profile_homes (#47368); literal non-existent /tmp paths are
skipped, so the salvaged test's homes must exist on disk.
2026-08-31 09:59:07 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
Teknium 6cc4f3913d fix(timezone): make profile-keyed tz cache atomic and add cron persistence regression
Follow-up on the cherry-picked #92489 base: replace the four separate
process-global cache slots with one lock-guarded identity->(name, zone)
mapping so racing profile-scoped threads can never publish a mixed
identity/value pair (the P1 interleaving flagged in the #92489 review),
keep each profile's resolved zone hot across multiplex switches, and pin
the #97905 symptom with a real-store regression test: a foreign-process
tick (desktop multiplex ticker pattern) must persist next_run_at with the
job-owning profile's UTC offset.

Fixes #97905. Refs #88220, #92489.
2026-08-31 09:58:40 -07:00