Commit Graph

758 Commits

Author SHA1 Message Date
Teknium 7d8b450a6a chore: map contributor email for overtoneblue 2026-09-01 08:30:45 -07:00
Teknium fcdae2cf0b chore: drop case-colliding contributor email file (#88168)
contributors/emails/agent@Agents-Mac-mini.local collides on
case-insensitive filesystems (macOS/Windows checkouts) with its
lowercase sibling, breaking clean checkouts. The identity is a
local-machine artifact, not a real contributor address.

Fixes #88168
Fixes #100047
Fixes #99966
Fixes #99821
2026-09-01 08:12:20 -07:00
lesseradmin 64a7c0aa8a chore: map contributor emails
Map a.dltorreruiz@gmail.com to GitHub user 4dlt for PR #98135 attribution.
2026-09-01 02:32:39 -07:00
kshitijk4poor 78372d22d1 chore: map yongliabc888@gmail.com to yongli-abc for PR #87629 salvage 2026-09-01 14:56:30 +05:30
Teknium c532a9db02 chore: add contributor mapping for vjpixel (PR #93439 salvage) 2026-09-01 02:00:05 -07:00
Sergey0515 f85edd3d94 fix(desktop): stop plugin-launched gateways before the Windows release gate (#70337)
releaseBackendLock tree-kills only the Desktop's own backends
(backendConnectionState + backendPool). A messaging gateway launched by
the gateway-launcher desktop plugin via /api/gateway/start lives outside
those structures; on Windows its launcher (venv\Scripts\python.exe)
keeps the venv mandatory-locked, so the 15s release gate aborts the
hand-off BEFORE the venv-blocker scan — the scanner's pausable-gateway
exemption and the CLI updater's pause machinery never get their chance.

Delegate to `hermes gateway stop --all` (launcher + worker discovery
across every profile; gateway.pid records only the uv WORKER and
taskkill /T from the worker never reaches its parent). Per review, adds
the drain-semantics counterpart: every applyUpdates abort path
(lock-held, venv-blocked, probe-failed, updater-spawn-failed) restores
gateways via `gateway start --all`, so a failed update no longer
strands every profile's gateway stopped. Pure DI'd module + tests.

Salvaged from PR #76057 (issue #70337; overlap credit: #70477 by
@JonthanaHanh targeted the same symptom earlier via ZIP-dir preservation).
2026-09-01 01:30:09 -07:00
xy952666680 6b18224df5 fix(desktop): reap detached hindsight venv daemon before update handoff (Windows)
The pre-handoff teardown tree-kills only the backends the Desktop owns
(backendConnectionState + backendPool). The memory plugin's hindsight
daemon is spawned DETACHED off venv\Scripts\pythonw.exe, so it survives
the teardown, keeps venv files mapped, and either dead-ends the
venv-blocker scan with no in-app remedy or (pre-#74805 shim-only gate)
raced the updater into a half-updated venv.

Add a narrowly-scoped reap: kill only processes whose exe lives under
venv\Scripts (ordinal case-insensitive prefix — no PowerShell -like
wildcard hazards) AND whose cmdline references hindsight_api.main.
External holders (user terminals, unrelated scripts) are never killed —
scanVenvBlockers still reports them and the hand-off aborts, per existing
design. Selection logic is a pure DI'd module with unit tests.

Salvaged from PR #75477 (scoped per review: the narrow daemon kill; the
PR's generic every-exe kill was rejected as over-broad, its updateInFlight
half was superseded by #75778/#73822, and its generic lock-probe half by
the #74805 release gate + #99724 scanner classification).
2026-09-01 01:30:09 -07:00
kshitijk4poor 1eaa2b0a9b chore: map everest.kill1@gmail.com -> anhtahaylove in contributor emails
PR #98424 merged with commits authored as everest.kill1@gmail.com
(anhtahaylove) but the email was never added to contributors/emails/,
so the next release's contributor_audit would fail on the unmapped
address. One-file mapping, same mechanism as every other entry.
2026-09-01 03:48:40 +05:30
caya8205-2 5ffaed6e45 fix(gateway): resolve session storage from the key's profile, not ambient scope
#88734 made SessionStore._db follow the ambient HERMES_HOME so a multiplexed
profile's rows reach its own state.db. That is correct for the inbound message
path, which installs the scope via _profile_runtime_scope. Nothing else does.

_session_expiry_watcher (gateway/run.py) walks the single process-wide
_entries dict — every profile's keys — and finalizes expired sessions with no
scope installed, so _db resolved the ROOT store for rows that live under
profiles/<name>/state.db. The scoped inbound path and the unscoped background
path then maintained two copies of the same logical session whose end_reason
drifted apart independently. Once they disagreed, the #54878 stale-routing
guard read one copy while the routing index pointed at the other, and a live
conversation was dropped and recreated — silently, since that branch only sets
was_auto_reset when a reset policy also fired.

Field evidence from a live two-profile install: session 20260814_234313 was
end_reason=None in the root store but agent_close in the profile store, while
20260822_225807 was inverted. Both directions, which rules out a single
mis-scoped writer.

The owning profile is already encoded in the session key, so derive the store
from it: _profile_home_for_key / _db_for_key, plus _db_for_session_id for the
entry points addressed by session id. 40 self._db uses across 14 methods now
resolve that way. No signature changed and no existing test was modified.

_profile_home_for_key returns None when multiplexing is off, when the key
carries the legacy agent:main namespace, or when the profile has no live
directory, so single-profile installs resolve exactly where they always did.
The explicit-path branch still goes through SessionDB.__init__ ->
_ensure_test_isolation, keeping the live-DB guard over per-profile paths.

Part of #66887. The routing-index half — _routing_scope() and the sessions.json
mirror still pinned to one frozen sessions_dir while the handle moves — is left
for a follow-up rather than mixed in here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
Xuezhao Lan ca177221b1 chore: map contributor email for PR 68556 2026-08-31 14:40:09 -07:00
Dev 532ee1a199 fix(gateway): persist gateway routing identity on lazily-created session rows
When the default/global state.db is corrupt at gateway startup,
SessionStore degrades (_db=None) and record_gateway_session_peer never
self-heals. Under multiplexed profile routes the AIAgent's lazy
_ensure_db_session was then the ONLY durable write for the session, and
it created the row identity-less (session_key/chat_id/chat_type/
thread_id/user_id/origin_json all NULL) — unrecoverable by
find_latest_gateway_session_for_peer, so Telegram chats forgot prior
turns.

Persist the routing identity the agent already carries into
create_session; plain CLI sessions keep the old identity-less shape.

Salvaged leg 1 of #88804; the transcript-recovery leg is covered by the
scope-aware session DB resolution already on main.
2026-08-31 12:32:33 -07:00
Teknium 29112bef09 chore: release v0.21.0 (2026.8.31) 2026-08-31 12:29:27 -07:00
Teknium 7acc65a399 test(windows): live E2E for the git trampoline self-heal on the wine2e lane
Real windows-latest coverage for the #88136 salvage: probes drive the
actual _git_is_trampoline/_locate_real_git/_ensure_non_trampoline_git
helpers against the runner's genuine Git-for-Windows install plus a real
fork-bomb-guard trampoline stand-in. Wired into the on-demand
windows-venv-e2e lane (wine2e/** pushes only).
2026-08-31 12:21:46 -07:00
Teknium 069d9d3529 chore: map komzpa@gmail.com -> Komzpa in contributor email registry 2026-08-31 12:19:29 -07:00
Teknium f680a4dc7d fix(gateway): require FTS provenance before transcript rebuild-and-retry
Widen #96038's fail-closed classifier to the gateway transcript retry
path: SessionStore._is_fts_corruption_error no longer treats a generic
'database disk image is malformed' as FTS-only damage. It now delegates
to SessionDB._is_fts_write_corruption_error (SQLITE_CORRUPT_VTAB result
code or explicit fts5 corrupt-structure text) and only keeps the
messages_fts-named cases. Structural corruption falls through to the
bounded retry/backoff path instead of rebuilding FTS and retrying writes
against a damaged database.

Sibling site spotted in PR #98090 by @fangliquanflq.
2026-08-31 11:42:23 -07:00
semao0 39540a03fc fix(install): never adopt a pre-release Node.js build
install_node() picks the newest tarball out of
nodejs.org/dist/latest-v${NODE_VERSION}.x/ and installs it without ever asking
whether the binary inside is usable. That index currently serves
node-v26.8.0-<os>-<arch>.tar.xz -- a final-looking filename -- whose binary
reports v26.8.0-alpha.0.0.0. Node publishes the headers tarball named by
process.release.headersUrl only for final releases, so node-gyp cannot compile
against that build and every native module fails to install.

Probe the extracted tree before it replaces anything on disk, and fall back to
an older release line when the probe rejects it, instead of leaving the install
with an unbuildable runtime. Mirror the guard in node-bootstrap.sh, and let
_managed_node_tree_outdated() treat a pre-release tree as outdated so an
already-broken install heals itself -- the existing heal only fires below the
target major, and a pre-release sits above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GrEaXSjvFoBXKxAHTjnUbS
2026-08-31 10:53:18 -07:00
Teknium 8bce17432f chore: add contributor email mapping for humdrum00001010 2026-08-31 10:08:31 -07:00
Teknium 82c95a2a01 chore: map contributor emails for gokhanyildirimlar and mottledMantis 2026-08-31 10:07:51 -07:00
Agi-Asi 54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
Teknium ce94f2b8f4 chore: contributor email mappings for Buzz media salvage 2026-08-31 10:06:34 -07:00
kshitijk4poor 3e5534d713 chore: map itkingtao@126.com -> walker83 (attribution for #97955) 2026-08-31 10:03:34 -07:00
Teknium 1617ff0254 chore: contributor mapping for Clarion1631 2026-08-31 09:59:47 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
Teknium 963c3619d2 chore: contributor email mapping for salvage #99362 2026-08-31 09:56:54 -07:00
Teknium 3d4712c744 chore: contributor email mapping for salvage cluster 2026-08-31 09:56:43 -07:00
Teknium b907b7eb85 fix(buzz): compose #97502's membership-rejection matching into the per-subscription CLOSED handler
- Widen the permanent-rejection match to the exact relay phrasings seen in
  production (#97502): 'not a channel member' and 'auth-required', alongside
  'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
  dms-list path and the channels-list fallback) now skips channels in
  _restricted_channels, so a restricted channel dropped at runtime cannot be
  silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
  per-subscription drop + restricted set is PR #76850 by @xozai.
2026-08-31 09:49:30 -07:00
Alessandro Boni d36827b086 fix(buzz): resume watched channels from a durable cursor across restarts (#90464)
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".

Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.

Details worth noting:

- The file records the identity and relay it was written for. A cursor from a
  different bot or relay is ignored rather than trusted — the channel ids
  would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
  seeding instead of failing the connect. Writes go through
  `utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
  cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
  newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
  without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
  rewrite the file every poll interval. Both inbound transports are covered:
  the poll sweep and the WebSocket event path share the same check.

Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.

Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
2026-08-31 09:49:30 -07:00
Teknium 92ac4ca4f0 chore: contributor email mappings for salvaged Buzz dispatch cluster 2026-08-31 09:05:41 -07:00
kshitij 4c9f130bf9 Merge pull request #99455 from kshitijk4poor/chore/author-map-lightpanda-prs
chore: contributor mapping for adria.arrufat@gmail.com (PR #99312)
2026-08-31 21:28:16 +05:30
Teknium ca9952cbb3 fix(plugins): doctor temp home can no longer be stranded by a failed staging copy
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.
2026-08-31 08:41:15 -07:00
kshitijk4poor 7e5cf09a23 chore: map adria.arrufat@gmail.com -> arrufat (PR #99312) 2026-08-31 20:39:45 +05:30
Teknium 26f178e5fa fix(terminal): gate the BUZZ_* terminal carve-out on actual Buzz agent context
Compose the two salvaged approaches (#78065 + #78511):

- Keep #78065's terminal-only scrub-path exemption (first-party prefix
  predicate in _make_run_env / _sanitize_subprocess_env, plain env values
  never scope-resolved, snapshot exclusion for cross-profile isolation,
  every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
  an import-time blocklist discard: the blocklist is shared by every
  scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
  execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
  process env (Buzz Desktop buzz-acp harness, #76243) OR the live
  session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
  concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
  session on a host that also runs a Buzz gateway does NOT get the
  signing key in its terminal children (maintainer triage note on
  #76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
  is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
  strips) and positive test (buzz session platform enables carve-out);
  docs updated accordingly.

Closes #78026, closes #76243.
2026-08-31 07:35:47 -07:00
Teknium 972f0314de feat(buzz): compose thread-topology cluster — reply_in_thread opt-out, NIP-10 root anchoring on all send paths, _PLATFORM_DEFAULTS tier
Compose/fix-up on top of the cherry-picked cluster commits:

- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
  field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
  users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
  through _apply_yaml_config and honored by send(), send_image(), and
  _standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
  _progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
  so the synthetic-thread fallback and the progress reply anchor are both
  suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
  _extract_thread_root (marked root > reply > legacy positional e-tag)
  instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
  edit_message now implemented, accumulate-style progress works, but without
  the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
  update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
2026-08-31 07:30:44 -07:00
Teknium e0b9f38177 chore: map contributor alex@gunsberg.fi -> alexgunsberg 2026-08-31 07:28:30 -07:00
Teknium b2995159f5 chore: map salvage contributors (cargo8, andthexi, kokhlo) 2026-08-31 06:02:32 -07:00
Paula Rossi 3efb514b9e fix(desktop): reuse primary WS for named profiles on attached shared remote (#96493)
Bot Mode group turns on a Windows Desktop attached to a remote gateway
dialed a registry secondary per member profile. That second WebSocket
accept/closed in ~30ms and never ran session.create. When the window
primary is already that one-host-many-profiles source (sharedRemote),
request/retain/open/ensure now reuse the primary socket and scope RPCs
with profile=. Isolated SSH/pooled backends are unchanged.
2026-08-31 05:58:52 -07:00
Teknium 37c856d337 chore: map contributor Bergmann89 for attribution CI 2026-08-31 05:56:18 -07:00
Teknium 38b7d0f4cf chore: contributor email mappings for salvage cluster 2026-08-31 03:37:43 -07:00
kshitijk4poor 19dd5bcee6 chore(contributors): map vaibhavdahiya28@gmail.com -> vsd2807
Attribution mapping for the #98137 salvage (author commit rides in the
same PR); must precede the authored commit so every commit passes
check-attribution.
2026-08-31 14:09:42 +05:30
Teknium 6cfadafb76 chore: map contributor email for @sambai-dev 2026-08-30 22:20:06 -07:00
Teknium b215e8d5f9 chore: map contributor email for itskaism 2026-08-30 22:19:10 -07:00
Teknium 1b2aaae1f8 chore: map steveonjava contributor email (PR #94036/#97292 salvage) 2026-08-30 05:16:10 -07:00
Teknium 0af3d62e7b chore: map ijnotion@pm.me -> james47kjv (PR #98008 salvage) 2026-08-30 05:16:02 -07:00
Teknium 11c8c05dc3 chore: map contributor email for fifeli 2026-08-29 20:39:09 -07:00
Teknium 9107b891c9 chore(attribution): map fabiantax@hotmail.com -> fabiantax (PR #90953 salvage) 2026-08-29 19:13:23 -07:00
Teknium ac5186ed82 chore(release): map adamfortuna1324@gmail.com to 0xAdamFortuna 2026-08-29 19:13:12 -07:00
Teknium 5a59ba82dd docs(providers): note extra_body survives gateway turns and /model switches; map gitabtion attribution 2026-08-29 19:13:00 -07:00
Teknium 1fa3edcb62 chore(attribution): map bsbofmusic noreply email 2026-08-29 19:12:51 -07:00
Teknium 2215fb0e35 fix(providers): mirror new Qwen Cloud models onto alibaba-cn
Follow-up to the #87808 salvage: the domestic alibaba-cn picker list
gets the same five additions (same DashScope catalog, per models.dev).
2026-08-29 19:12:19 -07:00
Brin Shadewater 0ebce1de81 chore: map contributor email
Agent: codex
2026-08-29 19:10:06 -07:00