Reconciles #84113 (authenticated same-relay URL localization) with #78051
(native imeta ingestion): _dispatch_message now merges caller-provided
verified imeta attachments with text-localized relay media instead of
clobbering them, dedupes paths, and downgrades mixed-source media to
DOCUMENT semantics so audio members are not routed through STT.
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.
`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.
Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
Renders the real CommandCenterView + ConfirmDialog: trash click alone must
not call onDeleteSession, delete fires only after explicit confirm, and
cancel closes without deleting. All three fail against the unguarded
pre-fix Command Center (verified by A/B against origin/main).
The Command Center -> Sessions delete button fired instantly on click,
hard-deleting the session (row + messages + request_dump files) with no
confirm and no undo. e6708af1f confirmed the sidebar rows, tab menus and
chat header, but missed the Command Center's independent entry point in
command-center/index.tsx.
Gate the row's delete button behind the same ConfirmDialog used by the
sidebar path, reusing t.sidebar.row copy and t.common.delete, so every
delete entry point is confirmed as e6708af1f intended.
A venv ever touched by sudo pip / sudo hermes contains root-owned files
(classically site-packages/*.dist-info/INSTALLER). A later normal-user
'hermes update' pulls code fine, then 'uv pip install -e .' dies with
'Permission denied (os error 13)' mid-mutation — venv/bin/hermes already
deleted, CLI bricked.
Add a bounded, pure-stat ownership preflight (_venv_foreign_owned_paths)
that runs after the code pull and immediately before the dependency
install. If foreign-owned paths are found it refuses up front, names the
offending paths + owner uid, prints the exact recovery command
(sudo chown -R $(id -un): <root>), and confirms the venv is untouched.
Windows (no os.geteuid) and root skip entirely. Never raises, capped at
~2000 stat calls, no subprocess use (update tests mock subprocess.run).
Same refuse-before-mutate philosophy as the contended-venv gate (#87331).
Fixes#83529
Diagnosis and documented recovery by @eabase.
uv refuses --locked/--check syncs when pyproject exclusion options differ from
the lockfile options block. Regenerated lock is metadata-only: the
[options.exclude-newer-package] table plus marker refinements; zero resolved
version or hash changes (verified: git diff has no version/sha256 lines).
Each release exact-pins at least one dependency to a version published
days before the release (v0.20.6: snowballstemmer==3.1.1, psutil==7.2.2).
For two weeks after release the relative exclude-newer cutoff filters
those versions out, so any venv that predates the release cannot resolve
the new pins at all ('no version of snowballstemmer==3.1.1' — observed
2026-08-29 updating three production installs v0.20.0 -> v0.20.6, one
Termux and two Linux servers; the Termux host additionally bricked on
psutil==7.2.2 sdist resolution, and cryptography's isolated build
environment resolved maturin/setuptools-rust under the same cutoff).
Same zero-float-protection logic as the setuptools/pillow/mcp/
firecrawl-anydoc exemptions: the pin bump WAS the review, so the cutoff
adds nothing for an exact pin and can only brick. Extend
exclude-newer-package to every exact-pinned package in
[project].dependencies / optional-dependencies (table moved to
one-key-per-line — 97 entries), plus maturin and setuptools-rust for
wheel-less sdist builds of the exempted cryptography pin.
test_exact_pinned_deps_exempt_from_exclude_newer enforces the invariant
going forward: adding a name==version pin without a matching
exclude-newer-package entry fails CI.
Windows installer editable builds fail in uv's isolated sandbox with
ModuleNotFoundError: wheel.cli because build-system.requires only listed
setuptools. setuptools.build_meta and our setup.py bdist_wheel guard both
import wheel during the build.
Also whitelist wheel in tool.uv.exclude-newer-package so the existing
build-system exclude-newer brick guard stays green.
Fixes#96488
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
Co-authored-by: Olympusbuildz <Olympus.roots@outlook.com>
Signed-off-by: Olympusbuildz <Olympus.roots@outlook.com>
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.
Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.
Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
(_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
message);
- micro-compaction splices [prefix, marker, suffix] - everything except
the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
keeping the historical archive-everything behavior.
Fixes#86366
Runs the installer's real node_deps_workspace_args against fabricated
checkout layouts by sourcing install.sh in --manifest mode, which defines
its functions without performing an install.
The load-bearing assertion is the invariant that no checkout shape lets
apps/desktop resolve, including the empty-argument case that would silently
hand npm the whole workspace glob back.
Also cover classify_changes recovering the PR file list when compare
returns nothing, so fail-open does not demand ci-reviewed for a CLI-only
install change.
The browser-tools step ran a bare `npm install` at the repo root, which
resolves the root package.json's `apps/*` workspace glob. That materializes
apps/desktop and with it node-pty, which ships no Linux prebuild and falls
back to `node-gyp rebuild` — so the installer needs make/gcc on a machine
that will never launch Electron or a PTY addon. Since #85297 made a failed
npm install fatal, a host without a C toolchain (a stock CentOS/RHEL box,
for instance) cannot complete a CLI-only install at all; it just reports
"npm install failed or timed out".
Name the workspaces the install actually needs instead. ui-tui and web are
selected when present, with --include-workspace-root so the root's shared
ESLint devDependencies are not pruned by the scoped install — the same
closure `hermes update` already installs. A checkout with neither workspace
falls back to a root-only install, since npm fails hard on a workspace it
cannot find. Desktop dependencies keep coming from install_desktop(), which
is only reachable via --include-desktop.
Against a pristine tree the unscoped install reifies 1362 packages including
node-pty 1.1.0; the scoped one reifies 582 with no native desktop addon.
A fork force-push can 404 the compare API used by detect-changes, which
fail-opens with ci_review=true and blocks the PR on a ci-reviewed label
the install change does not need. Recover the file list from the pull
request files endpoint before that fail-open.
AI-review follow-up on #87467:
- On probe failure, remove the extracted ~/.hermes/node tree and the
node/npm/npx bin links so later installer steps and retry runs start
clean instead of resolving node to a binary that cannot start.
- The termux pkg branch had the same silent-success class: an empty
version probe logged success and set HAS_NODE=true. Degrade with the
binary's own error instead.
install_node's post-install probe was
installed_ver=$(node --version 2>/dev/null) under set -e: when the
downloaded Node exists but cannot start (Node 26 linux-x64 builds link
libatomic.so.1, missing on minimal Debian/Ubuntu), the assignment
aborted the whole installer at exit 127 with the loader's explanation
discarded — installs died mid-sentence with no output at all (#87460).
- Probe now captures stderr and degrades with log_error carrying the
loader message plus the libatomic1 hint instead of aborting.
- Debian/Ubuntu installs preinstall libatomic1 (best-effort, mirroring
the existing apt idiom) so the common case just works.
- Termux branch's same-shaped probe gets a || true guard.
Fixes#87460
Review feedback on this PR: without --no-checkout, the blob fetch runs
inside git clone's own checkout step, so when the repo-scoped 429 hits
that fetch the whole clone exits non-zero, the else branch removes the
directory, and the fallback degrades to one more failed clone under
exactly the condition it exists for.
- Clone with --no-checkout (commits+trees only — small, passes the
throttle); the blobs are then fetched by a separate 'git reset --hard
HEAD' the retry can actually wrap. Verified on a local file://
filtering remote: the no-checkout clone materializes nothing and the
reset alone produces the full working tree.
- Fail closed: both reset attempts failing now removes the checkout and
reports 'Failed to clone repository' instead of the previous '|| true'
+ unconditional clone_ok=true handing the installer a half-materialized
tree printed as a success.
- The reset runs under a subshell cd so a failed materialization never
leaves the shell in a deleted cwd, and the direct-retry loop bound now
derives from $max_attempts (seq) instead of a hardcoded 1 2 3 4 that
could drift from the reported attempt count.
GitHub throttles packfile generation for this repository with
repo-scoped HTTP 429s that are not client IP rate limits: an
anonymous clone of a small repo succeeds and the API quota is
untouched, but the single big pack behind --depth 1 dies
mid-transfer with 'RPC failed; HTTP 429 / expected packfile'. The
fresh-install clone path had no retry and no fallback, so a clean
machine exited 1 at the download stage and left a half-populated
install directory (same throttle as the update path in #89287).
Retry the HTTPS clone with linear backoff, removing the partial
clone between attempts; when every direct attempt fails, degrade
to a blobless partial clone and materialize the working tree with
a hard reset — many small packs instead of one big one, which is
what gets past the throttle. SSH-first ordering, the existing
installation update branch, and the commit-pin flow are unchanged.
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
The multiplex loop on current main filters profile homes through
_existing_profile_homes (#47368); literal non-existent /tmp paths are
skipped, so the salvaged test's homes must exist on disk.
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.
tick() now checks, BEFORE acquiring the tick lock:
skew detected (boot fingerprint != disk revision)
AND this process does not own the gateway runtime lock
AND that lock is held (a fresh gateway is alive)
-> raise CronTickYielded, skipping the tick entirely
Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
proceed; yielding is a certainty claim, never a guess
The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.
Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.
gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
Follow-up on the cherry-picked #92489 base: replace the four separate
process-global cache slots with one lock-guarded identity->(name, zone)
mapping so racing profile-scoped threads can never publish a mixed
identity/value pair (the P1 interleaving flagged in the #92489 review),
keep each profile's resolved zone hot across multiplex switches, and pin
the #97905 symptom with a real-store regression test: a foreign-process
tick (desktop multiplex ticker pattern) must persist next_run_at with the
job-owning profile's UTC offset.
Fixes#97905. Refs #88220, #92489.