Commit Graph

22261 Commits

Author SHA1 Message Date
kshitij 285eaaddc0 chore: add contributor email mapping for seze@andrew.cmu.edu
Maps to GitHub user Shedrackeze002 (PR #85616).
2026-08-14 13:15:56 +05:30
Shedrack Eze ce20857f5a fix: exclude launchd-managed gateway from orphan reaper on macOS
`_reap_unsupervised_gateway_orphans()` short-circuits on Linux hosts
with systemd via `supports_systemd_services()`, but returns `False` on
macOS — there is no systemd.  This means the orphan reaper runs
unconditionally on macOS and treats the launchd-managed gateway as an
unsupervised orphan, SIGTERM-ing it.

When Hermes Desktop opens, `hermes serve` calls this function during
startup (web_server.py line ~251, gated on `HERMES_DESKTOP == "1"`).
The launchd gateway is killed, launchd restarts it via `KeepAlive: true`,
and the user sees a spurious gateway restart every time they reopen the
Desktop app.

Fix: exclude PIDs managed by launchd (`_get_service_pids()` already
returns launchd-managed PIDs on macOS) from the orphan scan, the same
way systemd PIDs are excluded on Linux.

Tested on macOS 26.5 (Tahoe) with Hermes Desktop 0.16+ and a
launchd-managed gateway (`ai.hermes.gateway` plist with `KeepAlive`).
Before the fix, quitting and reopening Hermes Desktop restarted the
gateway every time. After the fix, the gateway stays running across
Desktop quit/reopen cycles.
2026-08-14 13:15:56 +05:30
Shedrack Eze 8c0164cef9 test: add regression tests for launchd PID exclusion on macOS
Two tests in TestReapUnsupervisedGatewayOrphansMacOS:
- test_macos_excludes_launchd_pid_from_kill: verifies a launchd-managed
  PID is not SIGTERM'd while a real orphan is
- test_macos_no_orphans_when_only_launchd_gateway_running: verifies the
  reaper returns False when the only gateway PID is launchd-managed

Both tests patch is_macos() to True and supports_systemd_services() to
False to simulate the macOS code path where the short-circuit does not
fire.
2026-08-14 13:15:56 +05:30
Teknium 17d6a7d426 fix(desktop): avatar misses expire after 30s instead of caching forever (#85908)
resolveAgentAvatar cached null permanently (window lifetime), so a bot
whose avatar reached the asset store moments after its first notice
rendered — freshly created bots, art backfills in flight — kept the 🤖
glyph until an app restart even though profiles.get_asset had the pfp
(user report: brand-new bots' notices never picked up their faces).
Hits stay cached for the window; misses re-probe after 30s. Same
dedupe/inflight behavior otherwise.
2026-08-14 00:42:02 -07:00
Teknium 6c40300f93 feat(gateway): roster previews show the latest message, not the first (#85905)
profiles.list's last_session.preview reused list_sessions_rich's
shared preview, which is the session's FIRST user message — right for
session lists (recognition), wrong for a messaging-style roster where
the line under each agent should track the conversation ('Hey, tell
me about yourself!' forever, per user report). Override with the
newest active user/assistant text (same query shape and lock
discipline as SessionDB.latest_message_row_id); best-effort, falls
back to the first-message preview on any failure.

E2E: live profiles.list now shows each bot's latest exchange.
2026-08-14 00:33:51 -07:00
Teknium 1c971769ec feat(gateway): concise background process notifications by default
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
2026-08-14 00:26:11 -07:00
brooklyn! 529ee80ac0 Recover from a half-replaced desktop bundle instead of requiring a reinstall (#85887)
* fix(desktop): load the intact renderer bundle when an update tears one copy

index.html and the hashed chunks it names are one generation. A packaged app
ships that bundle twice (inside app.asar and, via asarUnpack, beside it in
app.asar.unpacked), so an update that replaces the app while its files are
locked can leave the two copies from different generations. resolveRendererIndex
took the first index.html that existed, so it could pick the torn one and the
window died on its first lazy import with "Failed to fetch dynamically imported
module" -- with no way out, because every relaunch reloaded the same copy.

Check each candidate's declared modules and prefer a complete generation; when
both are torn, log which files are missing and how to repair instead of leaving
the crash unexplained.

* fix(cli): rebuild the desktop app when its renderer bundle is half-replaced

The content stamp hashes the SOURCE tree, which an interrupted update leaves
intact, so `hermes desktop` reported "up to date" and skipped the rebuild that
would repair a torn bundle -- the app relaunched into the same crash and
reinstalling looked like the only option.

Treat a bundle whose index.html names missing chunks as stale regardless of the
stamp, and say so on the way into the rebuild.
2026-08-14 02:21:20 -05:00
hermes-seaeye[bot] bee9ef375b fmt(js): npm run fix on merge (#85898)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-14 07:19:47 +00:00
Teknium 0842655382 feat(desktop): sender-side delivery notices — 'Messaged X' / 'Message from X' (#85888)
The sending bot's chat showed inter-agent deliveries as raw terminal
tool rows (the hermes -p … chat … command + output transcript) —
plumbing, not conversation. When a terminal call matches the delivery
convention (-p <agent> chat … -q "Message from …", the shape from
#85855), it now renders as compact centered notices instead:
'Messaging <agent>…' while running, 'Messaged <agent>' on completion,
and — when the quiet run returns the recipient's reply — a 'Message
from <agent>' notice with the text behind a 'show message' expander.
Avatar resolution reuses the #85855 helper (now exported); glyph
fallback everywhere it can't resolve. Failed commands keep the real
terminal row (debuggable). Ordinary terminal calls untouched.

Sender + receiver now speak one visual language: the exchange is a
pair of timeline events on both ends (Grok-bots parity), with #85884
collapsing the receiving side's reply.

agent-delivery tests 6/6; thread suite 20 files / 117 tests green.
2026-08-14 00:10:07 -07:00
Teknium bd22451f0d feat(desktop): collapse inter-agent exchange replies (Grok-bots parity) (#85884)
The recipient's reply to an inter-agent delivery rendered as a full
assistant message, so the receiving bot's chat read like a normal
human conversation. The exchange is an EVENT in that bot's timeline,
not conversation content: when the immediately preceding user message
is an inter-agent delivery (AGENT_MESSAGE_RE, shipped in #85855), the
reply now renders as a compact centered 'Replied to <sender>' notice
with the full text behind a 'show reply' expander — mirroring the
delivery notice above it. Never collapses while streaming (progress
stays visible); ordinary assistant messages untouched.

Thread suite 19 files / 111 tests green.
2026-08-14 00:02:26 -07:00
Brooklyn Nicholson e19f00d770 docs(desktop): document HUD mode — long-press to move, resize, snap, exit
The desktop docs had no HUD mode section at all. Add one covering the
key interaction (long-press the composer to drag the bar), plus resize,
snap-to-pointer, and exit — all sourced from the actual implementation
in composer-drag.ts, resize-handle.ts, and keybinds/actions.ts.
2026-08-14 01:55:23 -05:00
Teknium e8debfbdc3 chore: map contributor email for attribution audit 2026-08-13 23:43:54 -07:00
Teknium fbec3c78fd fix(compression): clear preflight block on provider-confirmed rearm
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.

Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
2026-08-13 23:43:54 -07:00
Henry Mascot 2b9da1f252 fix(compression): re-arm same-turn attempt budget 2026-08-13 23:43:54 -07:00
Teknium 5e8d25d7e7 fix(cli): --in accepts Git Bash / MSYS-style paths on Windows (#85865)
* fix(cli): --in accepts Git Bash / MSYS paths on Windows

Under Git Bash, 'hermes chat --in ~' reaches the CLI as /c/Users/<user>
(the shell expands ~ to an MSYS POSIX path; MSYS2 argument conversion
is disabled for native executables), and the isdir check failed with
'--in directory not found: /c/Users/...'. Route the value through the
existing _msys_to_windows_path translator (MSYS + Cygwin + WSL drive
spellings; no-op elsewhere) before expanduser/abspath.

Hit live: Bot Mode's agent-messaging protocol delivers with --in ~, so
every bot-to-bot send from a Git Bash-driven agent failed on Windows.

Tests pin both the translation cases and (source-level) the call site
actually using it.

* test: assert the MSYS translation, not platform abspath

The prior assertion ran os.path.abspath on the translated Windows path,
which on the Linux CI runner (posixpath) treats 'C:\Users\alice' as
relative and prepends the runner cwd. Pin the translation output and
ntpath absoluteness instead — same contract, platform-independent.
2026-08-13 23:41:49 -07:00
Brooklyn Nicholson ad9e8c9b57 feat(desktop): expose data-attributes on sidebar sessions area for custom skinning
Add data-sessions-mode ('flat' | 'projects' | 'project' | 'archived' |
'search') and data-sessions-project (entered project id) to the sidebar
sessions wrapper so custom UIs can target project mode without relying
on internal class names.
2026-08-14 01:36:50 -05:00
Teknium 2a6eef77b7 chore: enforce LF line endings for all source files via .gitattributes (#85866)
Windows contributors' tools default to CRLF; without repo-level
normalization an edit becomes a whole-file phantom diff (583-line
'change' observed today from one 2-line edit), string-match patch
tooling breaks on invisible \r, and review is polluted. Extend the
existing LF rules (shell/Docker) to every source/text extension:
normalize at check-in AND check out as LF so working trees match the
index on all platforms. *.ps1 stays CRLF (PowerShell 5.1 tooling).

git add --renormalize: exactly one tracked file had mixed endings
(tests/tools/test_windows_agent_loop_papercuts.py) — normalized here,
so no phantom diffs land on anyone's next commit.
2026-08-13 23:36:25 -07:00
hermes-seaeye[bot] efad7c9161 fmt(js): npm run fix on merge (#85867)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-14 06:33:35 +00:00
Teknium 7a9634568c feat(desktop): agent-to-agent messages render as attributed cards, not user bubbles (#85855)
* feat(desktop): render agent-to-agent messages as attributed cards, not user bubbles

Bot-to-bot deliveries arrive on the user role (alternation requires
it) but are not the human speaking — they rendered as if the user
typed them. Detect the delivery prefix ('Message from 🤖 <sender>: …',
emoji-less, and the legacy bracket form) and render an attributed
inter-agent card: left-aligned, robot + sender header, 'agent message'
label, body through the same minimal markdown pipeline. Anchored regex
cannot fire mid-prose. Same pattern as ProcessNotificationNote.
Presentational only; content/roles/caching untouched.

* reshape: inter-agent card -> Grok-style compact timeline notice

Per maintainer screenshot: the delivery renders as a subtle centered
'🤖 Message from <sender>' notice (ProcessNotificationNote's shape),
with the delivered text behind a 'show message' expander instead of a
full-width card. The recipient's reply remains a normal assistant
message below it.

* feat: sender avatar on the inter-agent notice

The delivery prefix may carry the sender's profile handle —
'Message from 🤖 <Display Name> (@<handle>): …'. The notice resolves
it through profiles.list (has_avatar) + profiles.get_asset and renders
the sender's actual avatar in place of the 🤖 glyph, with module-level
memoization (one resolution per sender per window) and inflight
de-dup. Glyph fallback covers handle-less prefixes, older gateways
without profiles.*, avatarless profiles, and failures. 'hermes'
resolves to the primary profile by convention. Tests 6/6.

* lint: sort the $gateway import (perfectionist/sort-imports)
2026-08-13 23:19:04 -07:00
Gille 380e4da3d1 fix(compression): rearm budget from verified usage 2026-08-13 23:08:19 -07:00
ai-ag2026 72c828ca2c fix(agent): refund per-turn compression budget after real progress
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.

Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.

Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.

Rückbau: Commit revertieren; kein Zustand, keine Migration.

(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
2026-08-13 23:08:19 -07:00
Teknium 9504edbaea fix(desktop): agent mentions insert as plain @name, not an @simple: chip (#85841)
Picking a colon-less completion row (agent profile mentions from
complete.path, e.g. '@mr-tester') fell through serialize()'s typed-
reference branch: classify() marks colon-less entries type 'simple'
with insertId = text, so the serializer minted '@simple:`@mr-tester`'
— which rendered as a weird chip in the composer AND broke downstream
@mention routing (the backtick-quoted form no longer parses as a bare
mention). Colon-less rawText now inserts verbatim, matching what @diff
and @staged already did via the empty-insertId path.

Test: mention rows serialize to plain @name and commit to the editor
without an @simple: wrapper (directive-label 5/5).
2026-08-13 22:24:20 -07:00
unsupportedpastels ad8365d533 fix(desktop): preserve multi-pane plugins when closing panes
Closing a pane contributed by a plugin used to disable the entire plugin,
unloading every one of its contributions. For a plugin that owns several
independent panes (e.g. Bot Mode's Cronjobs pane alongside its Bots roster
and composer middleware), closing one pane silently killed the rest.

Now: closing one pane of a multi-pane plugin dismisses only that pane; the
plugin stays enabled and its other panes/commands/middleware keep working.
Reset layout restores dismissed contributed panes. A single-pane plugin
keeps the existing symmetric behavior (Close disables the plugin, with
Settings -> Plugins as the recovery path).

Adds regression coverage for both cases.
2026-08-13 22:18:42 -07:00
Teknium 86379c519a fix(desktop): add missing update mock to ComposerActionsScope test double
main's check:lint (tsc) is red: 071d27d1c3 added a required update()
member to ComposerActionsScope, but the use-composer-actions.test.ts
scope double was never extended. TS2741 at line 294.
2026-08-13 22:11:07 -07:00
Teknium 44bdcf3a30 docs(dashboard): document memory/disk pressure banner and /api/status resource blocks (NS-656)
Covers the advisory memory and disk blocks added to GET /api/status in
#84965 (thresholds, staleness handling, fail-safe degradation) and the
dashboard resource-pressure banner (trigger precedence, boot-scoped
dismissals).
2026-08-13 22:10:12 -07:00
Jaaneek 9c15f0191c fix(models): refresh xAI picker via models.dev; pin grok-4.6
Stop freezing the xAI/xAI-OAuth catalog at import so /model and setup
pick up new Grok IDs after the models.dev cache refreshes. Put xai and
xai-oauth on the shared picker-time models.dev merge path and pin
grok-4.6 as the default headline model.
2026-08-13 22:09:50 -07:00
teknium1 5970ac73ab chore(skills/box): Hermes-compliance polish + docs registration
- author field to 'Chris Kim (iskysun96), Hermes Agent' convention
- 'Use this skill for' -> 'When to Use' section heading
- related_skills: google-workspace
- register auto-gen docs page, catalog row, sidebar entry
- contributors/emails mapping for iskysun96
2026-08-13 22:09:44 -07:00
chriskim e450b09dd1 feat(skills): add bundled Box productivity skill
Box cloud content management via the official @box/cli through the
terminal tool: files, folders, sharing, search, metadata, Box AI,
Hubs, bulk operations, webhooks, and a REST fallback via box request.
OAuth-only auth; SKILL.md routes to ten scoped reference files.

Salvaged from PR #52107 by @iskysun96.
2026-08-13 22:09:44 -07:00
Teknium d4b0039940 chore: map contributor email for @silence-de 2026-08-13 22:09:37 -07:00
joaomarcos 44463ef802 test(usage): codex cache_write_tokens + Qwen/Kimi flat cached_tokens coverage (#70543)
Regression tests salvaged from PR #70522 by @JoaoMarcos44. The Qwen flat
cached_tokens behavior is provided by the shared top-level fallback from
PR #66105 (@mehmetkr-31); the codex cache_write_tokens read landed in the
previous commit.
2026-08-13 22:09:37 -07:00
joaomarcos e331b834fe fix(usage): preserve OpenAI-wire cache writes at canonical accounting boundary (#85706)
Salvaged from PR #85702 by @JoaoMarcos44, composed onto the mapping-safe
_usage_get reads (PR #74591 by @RelaxJonh) and the flat cached_tokens /
Anthropic-name fallbacks (PRs #66105, #52571):

- cache-write precedence in the chat_completions branch:
  details.cache_write_tokens > details.cache_creation_input_tokens >
  usage.cache_creation_input_tokens > usage.cache_write_tokens
- codex_responses branch reads details.cache_write_tokens (GPT-5.6+
  documented name) with cache_creation_tokens fallback (from PR #70522)
- _usage_count(): clamp malformed negative counters to 0
- all reads in every branch are mapping-safe via _usage_get
2026-08-13 22:09:37 -07:00
RelaxJonh 11de734331 fix(usage): support dict-shaped usage objects in normalize_usage (#74314)
When the Responses API returns usage as a plain dict (e.g. from a
middleware or proxy that deserialises JSON to dict instead of a typed
SDK object), normalize_usage() used getattr() exclusively, which
silently returned 0 for every field on a dict.

Add _usage_get() helper that reads via .get() for dicts and getattr()
for attribute-style objects. All accessor sites in normalize_usage()
now use this helper, so token counts and cost are correct regardless
of the usage object's type.

Regression tests: two new tests feed the same payload as both a dict
and a SimpleNamespace through the codex_responses and
chat_completions branches, asserting identical output and non-zero
values.
2026-08-13 22:09:37 -07:00
mehmetkr-31 5b82e69693 fix(agent): read Kimi/Moonshot top-level cached_tokens in normalize_usage (#65722)
Kimi/Moonshot's native API (api.moonshot.cn / .ai) reports context-cache hits
as a top-level ``usage.cached_tokens``. The chat-completions branch of
normalize_usage() walks a fallback chain of
prompt_tokens_details.cached_tokens -> cache_read_input_tokens ->
prompt_cache_hit_tokens; none of those names match, so direct Kimi sessions
normalized to cache_read_tokens=0. The hits were invisible in accounting and
the cached prefix was billed at the full input rate.

Appended as the last link in that chain, so it only fills a genuine zero and
cannot override a provider that reports the nested OpenAI shape or DeepSeek's
prompt_cache_hit_tokens.

Rebuilt on current main rather than rebased — the branch was ~3400 commits
behind. The DeepSeek half of the original branch is dropped: 03c0b00f4
(#65678) landed prompt_cache_hit_tokens on main, so this is Kimi-only as the
review asked. The scripts/release.py addition to the frozen LEGACY_AUTHOR_MAP
is dropped too; contributors/emails/mehmet.kar@std.yildiz.edu.tr already
exists on main.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:09:37 -07:00
Chengxi Zhou 69bc3159e8 fix(agent): fallback to Anthropic-style token fields in normalize_usage
Local OpenAI-compatible servers like mlx_vlm.server emit
input_tokens/output_tokens in chat completion responses instead of
prompt_tokens/completion_tokens. The OpenAI Python client preserves
these as extra attributes, but normalize_usage() only looked at
OpenAI-style field names, causing input_tokens to always be 0.

This made the context progress bar stay at 0% forever and prevented
auto-compression from triggering.

Fix: add Anthropic-style fallback (input_tokens/output_tokens) in the
default else-branch of normalize_usage, with OpenAI-style names taking
priority.

Fixes #14686
2026-08-13 22:09:37 -07:00
Teknium eca85e81d6 feat(complete): offer agent profiles as @mention completions (#85799)
Typing '@t' in the composer only ever completed paths/directives —
agent profiles were invisible even though multi-agent UIs (Bot Mode)
route @<profile> as a handoff. complete.path now surfaces matching
profiles for bare-word @queries, ranked above file hits (there are at
most a handful), plus the full list on a bare '@'. The primary profile
is also offered as '@hermes' when no real profile claims that name.
@kind: directive queries are untouched.

Verified live: '@t' -> @turqoise first, '@her' -> @hermes,
'@file:tur' unaffected.
2026-08-13 20:43:41 -07:00
ethernet 16a173a8d6 fix(tools): do not adopt a stale cwd after an interrupted command
The command wrapper prints the cwd marker after the command returns. A
killed or timed-out command emits no marker, so ``env.cwd`` still holds the
directory of the last command to FINISH. One local environment serves every
session, because ``_resolve_container_task_id`` collapses cwd-only overrides
to ``"default"``. That leftover directory is therefore routinely another
session's.

The post-command dual-write copied ``env.cwd`` into the interrupted session's
durable record. Every later command in that session then ran in the foreign
directory, and the cwd echo told the model it had moved there. A desktop chat
silently re-homed into a worktree that another chat had opened.

Report the observation instead of inferring it. The marker parse now sets
``result["cwd_observed"]``, and both the record write and the echo read that
flag. The local override clears the flag when it rolls back a path that does
not exist, because the restored value is also unobserved. When a command
reports no cwd, the session keeps the directory it already had.

This needs no second session to be wrong: a lone session that interrupts a
command re-adopts a stale value too. A second session only makes the wrong
directory belong to somebody else.

The same class of write exists in the file-tools rescue for a reaped
environment (#26211). That rescue copied the cached snapshot of the shared
``env.cwd`` into the session record. The rescue is now fill-only: it writes
the snapshot when the session has no record, and it never overwrites a
record that the session wrote for itself.

The tests drive ``terminal_tool`` itself through an interrupt, not a copy of
its gate. Review found that a revert of either call site passed the first
version of the tests. Each gate now has a test that fails when the gate is
removed (verified by mutation).

Two exact-dict assertions in the Vercel sandbox tests now assert the two
fields they care about, so a new result key does not fail them.
2026-08-13 20:30:26 -07:00
Shannon Sands 6977d21fa7 feat(gateway): disk-usage telemetry + dashboard disk-pressure banner (NS-656)
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).

- gateway/disk_status.py (new): collect_disk_status() samples
  shutil.disk_usage(HERMES_HOME) and classifies pressure
  (critical: <256 MB free or >=95% used; elevated: <512 MB free).
  Never raises — degrades to pressure="unknown" with null telemetry,
  same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
  not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
  resource banner with worst-first triggers (disk critical > memory
  critical > OOM restart > disk elevated > memory elevated) and
  cascading dismissals — hiding the top trigger surfaces the next one
  instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
  with English fallback per existing pattern).

Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
2026-08-13 20:30:12 -07:00
Shannon Sands f5a26b1575 fix(web): scope every memory-banner dismissal to boot_id (NS-656)
Review edge case: OOM dismissal was boot-keyed but live critical/elevated
dismissals were keyed only by severity. Dismiss critical, gateway reboots,
next poll still critical with no observed "ok" in between — the new
incident stayed hidden, and masked the OOM notice too, since critical
takes precedence.

Every dismissal key now embeds boot_id, so a gateway restart invalidates
prior dismissals of any kind. Within a boot, semantics are unchanged:
severity-scoped masking, escalation re-opens, confirmed "ok" recovery
clears live dismissals, "unknown" clears nothing. Missing boot_id
(pre-NS-656 image) degrades to a shared per-severity bucket as before.
2026-08-13 20:30:12 -07:00
Shannon Sands ba5dc00bb2 fix(memory-status): review follow-ups — incident-keyed dismissal, honest copy, single mobile offset (NS-656)
Addresses the human review findings on the memory-pressure feature:

* [P2] Dismissal hid later incidents of the same kind. The gateway now
  publishes `boot_id` (the lifecycle sentinel's started_at — changes on
  every gateway life) in the /api/status memory block, and the dashboard
  keys OOM-restart dismissal on it: acknowledging one restart no longer
  mutes the NEXT one (the OOM-loop case this banner exists for). Live
  pressure dismissals now also reset once pressure is demonstrably back
  to "ok" — "unknown" (stale heartbeat) is absence of evidence and
  clears nothing. Dismissal storage moved to a JSON list; old bare-string
  entries fail JSON.parse and degrade to a clean reset.

* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
  heartbeat), not proof the OOM killer acted — banner copy now says
  "restarted unexpectedly, most likely because it ran out of memory"
  instead of stating OOM as fact.

* [P3] Mobile header clearance was applied per-banner (mt-14 on both
  MemoryPressureBanner and ProfileScopeBanner) AND on the content
  (pt-14), double/triple-stacking 56px gaps when banners were visible.
  Replaced with a single h-14 spacer above the banner stack.
2026-08-13 20:30:12 -07:00
Shannon Sands 1745cf3b40 fix(web): rename memory-pressure interface to avoid declaration merge with existing MemoryStatus
api.ts already declares MemoryStatus for the /api/memory providers
endpoint; the NS-656 pressure block reused the name, and TypeScript
declaration merging fused the two shapes — 'tsc -b' in the Docker image
build failed on every test fixture. Renamed to MemoryPressureStatus.
2026-08-13 20:30:12 -07:00
Shannon Sands e11d1ddc7f feat(status): surface memory pressure and suspected-OOM restarts to users (NS-656)
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.

This is the read-side fix:

* New gateway/memory_status.py distills the existing 30s loop heartbeat
  (gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
  sentinel into a compact `memory` block: pressure ok/elevated/critical/
  unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
  Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
  future-dated heartbeats degrade pressure to "unknown" so a dead
  gateway's final gasp can't render a live "critical" banner forever.
  Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
  level would make a later unclean death "suspected OOM", warn at that
  level while the process is still alive.

* lifecycle_ledger.record_startup now carries prior_unclean_exit /
  prior_suspected_oom onto the reclaimed sentinel — previously the
  verdict survived only in append-only diag prose. Flags age out on the
  next sentinel rewrite (scoped to the life after the crash).

* /api/status serves the block (profile-aware, executor-offloaded,
  fail-safe to pressure=unknown). Deliberately NOT folded into
  components/overall: memory pressure is advisory, and flipping overall
  to "degraded" on it would page NAS's availability sweep for a
  condition the eviction valve is already handling. Public-safety:
  coarse numbers/enums/booleans only — same disclosure class as the
  existing nous_session_valid field, added for the same NAS-sweep
  audience.

* Dashboard: new MemoryPressureBanner (app-shell, next to
  ProfileScopeBanner) with worst-first trigger precedence
  (critical > suspected-OOM restart > elevated), per-trigger
  session-scoped dismissal, and escalation re-opening past a dismissal.
  i18n keys optional with English fallbacks, matching the
  managingProfileBanner convention.

Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.

NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.

Refs NS-656; context: NS-608, NS-657, OOF-77.
2026-08-13 20:30:12 -07:00
Ben Barclay f52feed1ef fix(azure-foundry): scope Responses reasoning suppression to post-tool turns (#84320)
Azure Foundry's OpenAI-compatible Responses surface rejects the post-tool
follow-up payload with HTTP 400 `invalid_payload` when a replayed encrypted
`reasoning` item is sent alongside `function_call` / `function_call_output`.
The initial function-call request and ordinary multi-turn continuity are both
accepted, so the failure only appears after the first tool executes.

Detect the Foundry endpoint in `ResponsesApiTransport.build_kwargs` and drop
only the encrypted reasoning replay on that follow-up turn, leaving
function_call / function_call_output continuity intact.

Salvage of #59981, rebuilt on current main. Same root cause and fix direction
as the original, which was correct; this version resolves three defects:

- No `chat_completion_helpers.py` change. main already forwards `provider`
  and `base_url` to the Responses transport, so the original's re-added
  arguments produced `SyntaxError: keyword argument repeated: provider` on
  merge. Dropping the hunk removed the syntax error and the conflict.

- Host matching uses `utils.base_url_host_matches`, not a substring test.
  `".services.ai.azure.com" in base_url` also matches URLs carrying the
  domain in a path or query segment, which would silently disable reasoning
  replay on an unrelated provider.

- The post-tool predicate tests the trailing messages, not the whole history.
  Scanning for any tool call plus any tool result made it sticky: one tool
  call early in a conversation suppressed reasoning on every later turn.

- Tool calls pair on `call_id` as well as `id`. Responses histories carry the
  function call id in `call_id` while `id` holds the response item id
  (`fc_...`). Identity is resolved via the converter's own
  `_split_responses_tool_id`, covering composite `"call_x|fc_y"` ids and bare
  `fc_` ids on both sides of the pairing.

Tests: 27 cases across the transport and the live `build_api_kwargs` bridge,
including six parametrized tool-call id shapes, non-Foundry host lookalikes,
the sticky-history guard, parallel tool results, and an unpaired tool result.
Each guard was confirmed to catch its defect by reverting the fix.

Verified with `scripts/run_tests.sh tests/agent/ tests/run_agent/`:
532 files, 5602 tests passed, 0 failed.

Not verified against a live Azure Foundry endpoint — no credentials. The
original HTTP 400 reproduction and post-fix Foundry Project / Azure Container
Apps harness runs are @AshuJoshi's, from #59981. This change is verified at
the payload-construction layer only.

Closes #59981.

Co-authored-by: Ashu Joshi <AshuJoshi@users.noreply.github.com>
2026-08-14 12:41:26 +10:00
Gille edb33be511 fix(desktop): persist dropped image bytes before attach 2026-08-13 21:24:10 -05:00
Teknium 486f4ace20 fix: voice dictation broken in profiles created via profiles.create (missing stt/tts config) (#85755)
* fix: mirror voice config (stt/tts/voice) into profiles created via profiles.create

Desktop dictation is profile-scoped: /api/audio/transcribe resolves the
stt section inside the TARGET profile's home. Profiles created through
profiles.create got only a model section, so dictation and TTS silently
fell back to defaults (local whisper, often not installed) — 'voice
dictation doesn't work in bot mode but is fine in regular mode'.

Mirror the launch profile's stt/tts/voice sections (key-wise, never
overwriting sections the clone already has) under the same
mirror_credentials flag that gates .env/auth mirroring, and report it
as mirrored.voice in the receipt.

* guard: route voice-config mirror through canonical loaders

read_user_config_raw (write-back round-trip; load_config would merge
DEFAULT_CONFIG and no-op the mirror) + save_config under the target
profile's HERMES_HOME override — same mechanism as _write_profile_model.
Satisfies test_config_read_guard.
2026-08-13 19:01:11 -07:00
Gille 2707183fed fix(desktop): stop offering unsupported GitHub MCP OAuth 2026-08-13 20:55:45 -05:00
Brooklyn Nicholson e50db44b12 docs(hyperframes): document preview teardown to stop leaked Chrome workers
npx hyperframes preview starts a long-lived next-server that keeps
chrome-headless-shell render workers resident. On GPU-less hosts (WSL,
containers, CI) each idle worker falls back to software WebGL (swiftshader)
and busy-spins a CPU core; a preview left open stacks these up until the
host is wedged. The skill never said preview was long-lived, so leaking
was the default outcome.

Add a Cleanup section + pitfall to SKILL.md and a Runaway CPU
troubleshooting entry with diagnose/fix/avoid steps.
2026-08-13 20:30:13 -05:00
Brooklyn Nicholson 423f92e607 fix(desktop): connect pills reload tools into the session that clicked them
Both connect providers captured `sessionId` when the suggestion was built
and ignored the one the pill hands `invoke`. An offer that outlived a
session switch therefore aimed its `reload.mcp` at the session the draft was
sampled in, so the chat the user actually clicked from resumed without the
tools the pill just said were ready.

Prefer the invoking pill's session; the captured one stays as the fallback.
2026-08-13 18:31:29 -05:00
Brooklyn Nicholson 685a5c95ad fix(desktop): a withdrawn suggestion pill drops its phase and cancels its work
Phase lived in a `Record<key, phase>` that only ever grew, keyed by
`provider:id` — keys that repeat constantly, since a provider withdraws and
re-offers the same suggestion whenever the draft loses and regains its
trigger. A leftover `done` then painted a genuine new offer as "Added
GitHub" and swallowed clicks, because only `idle` invokes.

The same map outlived a session switch. One composer stays mounted across
it, so connecting GitHub in one chat left the next chat's real offer inert.

Withdrawal also stranded in-flight work. The pill is the only cancel
affordance — clicking a working pill sets the flag the provider polls — so
once it left the strip an OAuth flow could poll forever, hold the server's
in-progress slot against a retry, and resolve into a config write with no UI
left to narrate or roll it back.

Phase now lives and dies with the pill: withdrawn keys drop their phase and
flip their cancel flag, unmount cancels everything in flight, and the strip
remounts per session.
2026-08-13 18:31:29 -05:00
Brooklyn Nicholson 0280cf09c4 fix(desktop): suggestion pills paint the current offer, not the first one seen
The bus's change gate compared offers by `provider:id` alone. Providers
rebuild their suggestion objects on every draft sample, so that key is equal
constantly and the write bailed out — pinning the FIRST object for the life
of the offer.

Two consequences, both user-visible. The pill keeps painting a stale reason
("you mentioned linear" after the user pasted a linear.app link), and it
keeps calling a stale `invoke` closure — work built for a draft that no
longer exists.

Compare the fields the pill actually renders instead. The reference-identity
bail-out survives for the common case (same draft, same match, no re-render),
which is what the gate was there for.
2026-08-13 18:31:29 -05:00
Brooklyn Nicholson 8c8d55bd07 feat(desktop): grow the MCP suggestion directory to 18 official hosted remotes
Vercel, Supabase, Netlify, Hugging Face, Asana, Intercom, Airtable,
Webflow, PayPal, and Square join the directory. Every entry is a
vendor-operated remote with its docs page linked, same URL-only rule
as the founding eight. Trigger notes where words are ambiguous:
'square' the English word never fires (squareup only), and
vercel.app/netlify.app deploy-preview hosts are deliberately absent
(a pasted preview link is about the site, not the platform). Brand
glyphs wired for all newcomers.
2026-08-13 17:41:47 -05:00