Commit Graph

790 Commits

Author SHA1 Message Date
Agi-Asi 54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
Teknium ce94f2b8f4 chore: contributor email mappings for Buzz media salvage 2026-08-31 10:06:34 -07:00
kshitijk4poor 3e5534d713 chore: map itkingtao@126.com -> walker83 (attribution for #97955) 2026-08-31 10:03:34 -07:00
Teknium 1617ff0254 chore: contributor mapping for Clarion1631 2026-08-31 09:59:47 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
Teknium 963c3619d2 chore: contributor email mapping for salvage #99362 2026-08-31 09:56:54 -07:00
Teknium 3d4712c744 chore: contributor email mapping for salvage cluster 2026-08-31 09:56:43 -07:00
Teknium b907b7eb85 fix(buzz): compose #97502's membership-rejection matching into the per-subscription CLOSED handler
- Widen the permanent-rejection match to the exact relay phrasings seen in
  production (#97502): 'not a channel member' and 'auth-required', alongside
  'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
  dms-list path and the channels-list fallback) now skips channels in
  _restricted_channels, so a restricted channel dropped at runtime cannot be
  silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
  per-subscription drop + restricted set is PR #76850 by @xozai.
2026-08-31 09:49:30 -07:00
Alessandro Boni d36827b086 fix(buzz): resume watched channels from a durable cursor across restarts (#90464)
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".

Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.

Details worth noting:

- The file records the identity and relay it was written for. A cursor from a
  different bot or relay is ignored rather than trusted — the channel ids
  would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
  seeding instead of failing the connect. Writes go through
  `utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
  cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
  newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
  without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
  rewrite the file every poll interval. Both inbound transports are covered:
  the poll sweep and the WebSocket event path share the same check.

Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.

Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
2026-08-31 09:49:30 -07:00
Teknium 92ac4ca4f0 chore: contributor email mappings for salvaged Buzz dispatch cluster 2026-08-31 09:05:41 -07:00
kshitij 4c9f130bf9 Merge pull request #99455 from kshitijk4poor/chore/author-map-lightpanda-prs
chore: contributor mapping for adria.arrufat@gmail.com (PR #99312)
2026-08-31 21:28:16 +05:30
Teknium ca9952cbb3 fix(plugins): doctor temp home can no longer be stranded by a failed staging copy
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.
2026-08-31 08:41:15 -07:00
kshitijk4poor 7e5cf09a23 chore: map adria.arrufat@gmail.com -> arrufat (PR #99312) 2026-08-31 20:39:45 +05:30
Teknium 26f178e5fa fix(terminal): gate the BUZZ_* terminal carve-out on actual Buzz agent context
Compose the two salvaged approaches (#78065 + #78511):

- Keep #78065's terminal-only scrub-path exemption (first-party prefix
  predicate in _make_run_env / _sanitize_subprocess_env, plain env values
  never scope-resolved, snapshot exclusion for cross-profile isolation,
  every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
  an import-time blocklist discard: the blocklist is shared by every
  scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
  execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
  process env (Buzz Desktop buzz-acp harness, #76243) OR the live
  session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
  concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
  session on a host that also runs a Buzz gateway does NOT get the
  signing key in its terminal children (maintainer triage note on
  #76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
  is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
  strips) and positive test (buzz session platform enables carve-out);
  docs updated accordingly.

Closes #78026, closes #76243.
2026-08-31 07:35:47 -07:00
Teknium 972f0314de feat(buzz): compose thread-topology cluster — reply_in_thread opt-out, NIP-10 root anchoring on all send paths, _PLATFORM_DEFAULTS tier
Compose/fix-up on top of the cherry-picked cluster commits:

- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
  field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
  users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
  through _apply_yaml_config and honored by send(), send_image(), and
  _standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
  _progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
  so the synthetic-thread fallback and the progress reply anchor are both
  suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
  _extract_thread_root (marked root > reply > legacy positional e-tag)
  instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
  edit_message now implemented, accumulate-style progress works, but without
  the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
  update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
2026-08-31 07:30:44 -07:00
Teknium e0b9f38177 chore: map contributor alex@gunsberg.fi -> alexgunsberg 2026-08-31 07:28:30 -07:00
Teknium b2995159f5 chore: map salvage contributors (cargo8, andthexi, kokhlo) 2026-08-31 06:02:32 -07:00
Paula Rossi 3efb514b9e fix(desktop): reuse primary WS for named profiles on attached shared remote (#96493)
Bot Mode group turns on a Windows Desktop attached to a remote gateway
dialed a registry secondary per member profile. That second WebSocket
accept/closed in ~30ms and never ran session.create. When the window
primary is already that one-host-many-profiles source (sharedRemote),
request/retain/open/ensure now reuse the primary socket and scope RPCs
with profile=. Isolated SSH/pooled backends are unchanged.
2026-08-31 05:58:52 -07:00
Teknium 37c856d337 chore: map contributor Bergmann89 for attribution CI 2026-08-31 05:56:18 -07:00
Teknium 38b7d0f4cf chore: contributor email mappings for salvage cluster 2026-08-31 03:37:43 -07:00
kshitijk4poor 19dd5bcee6 chore(contributors): map vaibhavdahiya28@gmail.com -> vsd2807
Attribution mapping for the #98137 salvage (author commit rides in the
same PR); must precede the authored commit so every commit passes
check-attribution.
2026-08-31 14:09:42 +05:30
Teknium 6cfadafb76 chore: map contributor email for @sambai-dev 2026-08-30 22:20:06 -07:00
Teknium b215e8d5f9 chore: map contributor email for itskaism 2026-08-30 22:19:10 -07:00
Teknium 1b2aaae1f8 chore: map steveonjava contributor email (PR #94036/#97292 salvage) 2026-08-30 05:16:10 -07:00
Teknium 0af3d62e7b chore: map ijnotion@pm.me -> james47kjv (PR #98008 salvage) 2026-08-30 05:16:02 -07:00
Teknium 11c8c05dc3 chore: map contributor email for fifeli 2026-08-29 20:39:09 -07:00
Teknium 9107b891c9 chore(attribution): map fabiantax@hotmail.com -> fabiantax (PR #90953 salvage) 2026-08-29 19:13:23 -07:00
Teknium ac5186ed82 chore(release): map adamfortuna1324@gmail.com to 0xAdamFortuna 2026-08-29 19:13:12 -07:00
Teknium 5a59ba82dd docs(providers): note extra_body survives gateway turns and /model switches; map gitabtion attribution 2026-08-29 19:13:00 -07:00
Teknium 1fa3edcb62 chore(attribution): map bsbofmusic noreply email 2026-08-29 19:12:51 -07:00
Teknium 2215fb0e35 fix(providers): mirror new Qwen Cloud models onto alibaba-cn
Follow-up to the #87808 salvage: the domestic alibaba-cn picker list
gets the same five additions (same DashScope catalog, per models.dev).
2026-08-29 19:12:19 -07:00
Brin Shadewater 0ebce1de81 chore: map contributor email
Agent: codex
2026-08-29 19:10:06 -07:00
Hermes e298dcce49 chore: map contributor email for injaneity 2026-08-29 18:35:17 -07:00
Teknium f8546c2eac fix(browser): real-profile follow-ups — reap launched Chrome, headless display-less Linux, register real_profile_pin default + docs
- _terminate_real_profile_chrome(): directly-launched real browsers are ours
  to reap (agent-browser only attaches); wired into the atexit emergency
  cleanup and both launch-failure paths so orphaned Chrome processes can't
  accumulate.
- Display-less Linux gate: append --headless=new (shares the profile's normal
  cookie store, unlike legacy headless) so the direct-launch path doesn't
  regress servers without DISPLAY/WAYLAND_DISPLAY.
- Register browser.real_profile_pin in config_defaults.py and document the
  new launch model + pin in website/docs/user-guide/features/browser.md.
- Drop unused tempfile import from the cherry-picked commit.
2026-08-29 18:35:12 -07:00
Cheri Wen 4bc7e624d6 feat(cli): show prompt cache hit rate in status bar
Add a ◎ XX% indicator to the CLI status bar showing the prompt cache
hit rate (cache_read / prompt_tokens). This helps users monitor how
effectively their provider's prompt caching is working.

Features:
- Color-coded: green (≥70%), yellow (40-70%), red (<40%)
- Adaptive precision: integer on narrow terminals, one decimal on wide
- Only shown when cache data is available (provider supports it)
- Compatible with OpenAI, Anthropic, DeepSeek, xiaomi, and other
  providers that return prompt_tokens_details.cached_tokens

Tests: 6 new test cases, 43/43 passing
2026-08-29 18:34:51 -07:00
salch-cred 1d13fe70b9 fix(cron): accept bare duration units like 'hour' in schedule parsing 2026-08-29 18:33:19 -07:00
Gille cd1c3211ec chore: map contributor email 2026-08-29 18:10:03 -07:00
kshitij b4b7727ea0 Merge pull request #97912 from kshitijk4poor/chore/author-map-provider-salvages
chore: map provider-salvage contributor emails (Neel49, amrrs)
2026-08-29 20:14:05 +05:30
kshitijk4poor 5e8ff7ce4b chore: map abdi.moya@gmail.com -> AxDSan, sam@odio.email -> srosro (attribution for #97875) 2026-08-29 20:03:48 +05:30
kshitijk4poor dfcdf13bcd chore: map provider-salvage contributor emails to GitHub logins
Neel49 (PR #93548 Ramp Router), amrrs (PR #28253 Nebius Token Factory).
simonweng@tencent.com (PR #96939) is already mapped.
2026-08-29 19:02:45 +05:30
teknium1 260adeb075 chore(skills/grill-me): contributor mapping + docs catalog/sidebar regen 2026-08-29 03:12:20 -07:00
kshitijk4poor 264ae5b2c8 chore: map jack@powries.com -> powriej (attribution for #97582) 2026-08-29 12:23:41 +05:30
Amandeep Khurana a7e7de6407 fix(desktop): make gateway file saves failure-atomic so a failed download never destroys an existing file
`pumpStreamToFile` opened the user-chosen destination with
`fs.createWriteStream`, which truncates the target the instant it opens,
and its error path then unlinked that same path. When a user picked an
existing file in the Save dialog (and confirmed the overwrite) and the
gateway dropped mid-stream, the original was gone: truncated first,
deleted second, with nothing written in its place. The data-URL
compatibility fallback (`saveGatewayFileViaDataUrl`) had the same class
of bug via `fs.promises.writeFile`, which truncates before the write
completes.

Both paths now go through one failure-atomic primitive. Bytes land in a
short, randomly named sibling temp file (`.hermes-download-<hex>.part`,
same directory so the final step is a same-volume rename), created with
`flags: 'wx'`, and are renamed onto the destination only after the whole
body has been written and the descriptor released. The destination is
never opened before that point, so a failed download leaves whatever was
there untouched.

- Ownership-gated cleanup: the temp file is unlinked only after the
  stream's 'open' event proved THIS operation created it. An exclusive
  create that fails before open (EEXIST collision, EACCES, missing
  parent) never removes a file that belongs to someone else.
- `WriteStream.close(cb)` rather than `end(cb)` before renaming: `end`'s
  callback fires on 'finish' while the fd may still be open, and Windows
  refuses to rename a file with an open handle. Falls back to `end` for
  stream shapes without `close`.
- The failure path waits for 'close' (bounded by a 2s grace period)
  before unlinking, for the same reason: `destroy()` releases the fd
  asynchronously and an unlink racing the open handle would leak the
  `.part` file on Windows.
- A rename failure (destination locked, permissions) removes the owned
  temp file and rejects; nothing is left behind.
- Fixed-length temp name so a long user-chosen filename cannot push it
  past the filesystem's name limit.
- `fsPumpDeps()` is the single production deps factory (`'wx'` create,
  `fs.promises.rename`, `fs.promises.unlink`); `writeBufferToFile()`
  routes the data-URL fallback through the same pump. `PumpDeps` gains
  `rename` and a `tempPathFor` test seam.

Tests. Fakes: temp-then-rename on success, close-before-rename ordering,
the regression itself (destination neither opened nor unlinked when the
response fails mid-stream), write-error cleanup, close-before-unlink
ordering, rename-failure cleanup, pre-open EEXIST leaves the colliding
file alone, `writeBufferToFile` success and post-open write failure, the
temp-name length bound, and the `main.ts` wiring. Real filesystem
(`gateway-file-download.fs.test.ts`, exact production deps in a scratch
dir): completed download replaces the destination with no temp left;
mid-stream failure leaves the pre-existing destination byte-for-byte
with no `.part`; failure into a fresh name leaves nothing; seeded temp
path survives a pre-open EEXIST with no rename; rename failure (directory
at the destination) cleans the owned temp; data-URL fallback success and
missing-directory failure.

Adds the contributor email mapping required by the attribution check.

Fixes #96597

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015u8q2pHVPZmxpSrgkt94jC
2026-08-29 11:33:05 +05:30
kshitijk4poor 9c137e6163 chore: map eric.maddox@outlook.com -> ericmaddox (attribution for #97618) 2026-08-29 11:33:05 +05:30
Mohamed HAMMANE 6da0ae1cf5 fix(install): preserve project config for locked uv sync (#82446) 2026-08-28 12:36:15 -07:00
Teknium c5b44e0756 chore: map contributor emails for TiberiuD and fedebyes 2026-08-28 07:51:23 -07:00
Teknium 3f3ae6850a chore: map brianbaldock contributor email 2026-08-28 07:51:16 -07:00
Teknium 95cf7dc9e8 feat: session temp root moves off tmpfs /tmp to ~/.hermes/cache/terminal by default; auto-pruned after 72h
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
2026-08-28 07:50:33 -07:00
Teknium 9f2ab334c8 fix(desktop): route the MCP health sweep and command palette through getServers
Two sibling readers of raw config.mcp_servers duplicated their own (weaker)
shape guard: mcp-health.ts guarded the map but still passed null entries to
isUrlServer (crash on .url read), and the command palette re-implemented the
map check inline. Both now go through getServers(), the single choke point
that drops malformed entries, so a null entry can't crash the sweep and the
palette lists exactly the servers the MCP tab shows.

Also records the contributor email mapping for the salvaged commits.

Follow-up to the cherry-picked #94338.
2026-08-28 06:33:40 -07:00
Teknium 911a41ec97 chore: map contributor email for hakanbaysal 2026-08-28 05:17:26 -07:00