`_reap_unsupervised_gateway_orphans()` short-circuits on Linux hosts
with systemd via `supports_systemd_services()`, but returns `False` on
macOS — there is no systemd. This means the orphan reaper runs
unconditionally on macOS and treats the launchd-managed gateway as an
unsupervised orphan, SIGTERM-ing it.
When Hermes Desktop opens, `hermes serve` calls this function during
startup (web_server.py line ~251, gated on `HERMES_DESKTOP == "1"`).
The launchd gateway is killed, launchd restarts it via `KeepAlive: true`,
and the user sees a spurious gateway restart every time they reopen the
Desktop app.
Fix: exclude PIDs managed by launchd (`_get_service_pids()` already
returns launchd-managed PIDs on macOS) from the orphan scan, the same
way systemd PIDs are excluded on Linux.
Tested on macOS 26.5 (Tahoe) with Hermes Desktop 0.16+ and a
launchd-managed gateway (`ai.hermes.gateway` plist with `KeepAlive`).
Before the fix, quitting and reopening Hermes Desktop restarted the
gateway every time. After the fix, the gateway stays running across
Desktop quit/reopen cycles.
Two tests in TestReapUnsupervisedGatewayOrphansMacOS:
- test_macos_excludes_launchd_pid_from_kill: verifies a launchd-managed
PID is not SIGTERM'd while a real orphan is
- test_macos_no_orphans_when_only_launchd_gateway_running: verifies the
reaper returns False when the only gateway PID is launchd-managed
Both tests patch is_macos() to True and supports_systemd_services() to
False to simulate the macOS code path where the short-circuit does not
fire.
resolveAgentAvatar cached null permanently (window lifetime), so a bot
whose avatar reached the asset store moments after its first notice
rendered — freshly created bots, art backfills in flight — kept the 🤖
glyph until an app restart even though profiles.get_asset had the pfp
(user report: brand-new bots' notices never picked up their faces).
Hits stay cached for the window; misses re-probe after 30s. Same
dedupe/inflight behavior otherwise.
profiles.list's last_session.preview reused list_sessions_rich's
shared preview, which is the session's FIRST user message — right for
session lists (recognition), wrong for a messaging-style roster where
the line under each agent should track the conversation ('Hey, tell
me about yourself!' forever, per user report). Override with the
newest active user/assistant text (same query shape and lock
discipline as SessionDB.latest_message_row_id); best-effort, falls
back to the first-message preview on any failure.
E2E: live profiles.list now shows each bot's latest exchange.
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
* fix(desktop): load the intact renderer bundle when an update tears one copy
index.html and the hashed chunks it names are one generation. A packaged app
ships that bundle twice (inside app.asar and, via asarUnpack, beside it in
app.asar.unpacked), so an update that replaces the app while its files are
locked can leave the two copies from different generations. resolveRendererIndex
took the first index.html that existed, so it could pick the torn one and the
window died on its first lazy import with "Failed to fetch dynamically imported
module" -- with no way out, because every relaunch reloaded the same copy.
Check each candidate's declared modules and prefer a complete generation; when
both are torn, log which files are missing and how to repair instead of leaving
the crash unexplained.
* fix(cli): rebuild the desktop app when its renderer bundle is half-replaced
The content stamp hashes the SOURCE tree, which an interrupted update leaves
intact, so `hermes desktop` reported "up to date" and skipped the rebuild that
would repair a torn bundle -- the app relaunched into the same crash and
reinstalling looked like the only option.
Treat a bundle whose index.html names missing chunks as stale regardless of the
stamp, and say so on the way into the rebuild.
The sending bot's chat showed inter-agent deliveries as raw terminal
tool rows (the hermes -p … chat … command + output transcript) —
plumbing, not conversation. When a terminal call matches the delivery
convention (-p <agent> chat … -q "Message from …", the shape from
#85855), it now renders as compact centered notices instead:
'Messaging <agent>…' while running, 'Messaged <agent>' on completion,
and — when the quiet run returns the recipient's reply — a 'Message
from <agent>' notice with the text behind a 'show message' expander.
Avatar resolution reuses the #85855 helper (now exported); glyph
fallback everywhere it can't resolve. Failed commands keep the real
terminal row (debuggable). Ordinary terminal calls untouched.
Sender + receiver now speak one visual language: the exchange is a
pair of timeline events on both ends (Grok-bots parity), with #85884
collapsing the receiving side's reply.
agent-delivery tests 6/6; thread suite 20 files / 117 tests green.
The recipient's reply to an inter-agent delivery rendered as a full
assistant message, so the receiving bot's chat read like a normal
human conversation. The exchange is an EVENT in that bot's timeline,
not conversation content: when the immediately preceding user message
is an inter-agent delivery (AGENT_MESSAGE_RE, shipped in #85855), the
reply now renders as a compact centered 'Replied to <sender>' notice
with the full text behind a 'show reply' expander — mirroring the
delivery notice above it. Never collapses while streaming (progress
stays visible); ordinary assistant messages untouched.
Thread suite 19 files / 111 tests green.
The desktop docs had no HUD mode section at all. Add one covering the
key interaction (long-press the composer to drag the bar), plus resize,
snap-to-pointer, and exit — all sourced from the actual implementation
in composer-drag.ts, resize-handle.ts, and keybinds/actions.ts.
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.
Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
* fix(cli): --in accepts Git Bash / MSYS paths on Windows
Under Git Bash, 'hermes chat --in ~' reaches the CLI as /c/Users/<user>
(the shell expands ~ to an MSYS POSIX path; MSYS2 argument conversion
is disabled for native executables), and the isdir check failed with
'--in directory not found: /c/Users/...'. Route the value through the
existing _msys_to_windows_path translator (MSYS + Cygwin + WSL drive
spellings; no-op elsewhere) before expanduser/abspath.
Hit live: Bot Mode's agent-messaging protocol delivers with --in ~, so
every bot-to-bot send from a Git Bash-driven agent failed on Windows.
Tests pin both the translation cases and (source-level) the call site
actually using it.
* test: assert the MSYS translation, not platform abspath
The prior assertion ran os.path.abspath on the translated Windows path,
which on the Linux CI runner (posixpath) treats 'C:\Users\alice' as
relative and prepends the runner cwd. Pin the translation output and
ntpath absoluteness instead — same contract, platform-independent.
Add data-sessions-mode ('flat' | 'projects' | 'project' | 'archived' |
'search') and data-sessions-project (entered project id) to the sidebar
sessions wrapper so custom UIs can target project mode without relying
on internal class names.
Windows contributors' tools default to CRLF; without repo-level
normalization an edit becomes a whole-file phantom diff (583-line
'change' observed today from one 2-line edit), string-match patch
tooling breaks on invisible \r, and review is polluted. Extend the
existing LF rules (shell/Docker) to every source/text extension:
normalize at check-in AND check out as LF so working trees match the
index on all platforms. *.ps1 stays CRLF (PowerShell 5.1 tooling).
git add --renormalize: exactly one tracked file had mixed endings
(tests/tools/test_windows_agent_loop_papercuts.py) — normalized here,
so no phantom diffs land on anyone's next commit.
* feat(desktop): render agent-to-agent messages as attributed cards, not user bubbles
Bot-to-bot deliveries arrive on the user role (alternation requires
it) but are not the human speaking — they rendered as if the user
typed them. Detect the delivery prefix ('Message from 🤖 <sender>: …',
emoji-less, and the legacy bracket form) and render an attributed
inter-agent card: left-aligned, robot + sender header, 'agent message'
label, body through the same minimal markdown pipeline. Anchored regex
cannot fire mid-prose. Same pattern as ProcessNotificationNote.
Presentational only; content/roles/caching untouched.
* reshape: inter-agent card -> Grok-style compact timeline notice
Per maintainer screenshot: the delivery renders as a subtle centered
'🤖 Message from <sender>' notice (ProcessNotificationNote's shape),
with the delivered text behind a 'show message' expander instead of a
full-width card. The recipient's reply remains a normal assistant
message below it.
* feat: sender avatar on the inter-agent notice
The delivery prefix may carry the sender's profile handle —
'Message from 🤖 <Display Name> (@<handle>): …'. The notice resolves
it through profiles.list (has_avatar) + profiles.get_asset and renders
the sender's actual avatar in place of the 🤖 glyph, with module-level
memoization (one resolution per sender per window) and inflight
de-dup. Glyph fallback covers handle-less prefixes, older gateways
without profiles.*, avatarless profiles, and failures. 'hermes'
resolves to the primary profile by convention. Tests 6/6.
* lint: sort the $gateway import (perfectionist/sort-imports)
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.
Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.
Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.
Rückbau: Commit revertieren; kein Zustand, keine Migration.
(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
Picking a colon-less completion row (agent profile mentions from
complete.path, e.g. '@mr-tester') fell through serialize()'s typed-
reference branch: classify() marks colon-less entries type 'simple'
with insertId = text, so the serializer minted '@simple:`@mr-tester`'
— which rendered as a weird chip in the composer AND broke downstream
@mention routing (the backtick-quoted form no longer parses as a bare
mention). Colon-less rawText now inserts verbatim, matching what @diff
and @staged already did via the empty-insertId path.
Test: mention rows serialize to plain @name and commit to the editor
without an @simple: wrapper (directive-label 5/5).
Closing a pane contributed by a plugin used to disable the entire plugin,
unloading every one of its contributions. For a plugin that owns several
independent panes (e.g. Bot Mode's Cronjobs pane alongside its Bots roster
and composer middleware), closing one pane silently killed the rest.
Now: closing one pane of a multi-pane plugin dismisses only that pane; the
plugin stays enabled and its other panes/commands/middleware keep working.
Reset layout restores dismissed contributed panes. A single-pane plugin
keeps the existing symmetric behavior (Close disables the plugin, with
Settings -> Plugins as the recovery path).
Adds regression coverage for both cases.
main's check:lint (tsc) is red: 071d27d1c3 added a required update()
member to ComposerActionsScope, but the use-composer-actions.test.ts
scope double was never extended. TS2741 at line 294.
Covers the advisory memory and disk blocks added to GET /api/status in
#84965 (thresholds, staleness handling, fail-safe degradation) and the
dashboard resource-pressure banner (trigger precedence, boot-scoped
dismissals).
Stop freezing the xAI/xAI-OAuth catalog at import so /model and setup
pick up new Grok IDs after the models.dev cache refreshes. Put xai and
xai-oauth on the shared picker-time models.dev merge path and pin
grok-4.6 as the default headline model.
Box cloud content management via the official @box/cli through the
terminal tool: files, folders, sharing, search, metadata, Box AI,
Hubs, bulk operations, webhooks, and a REST fallback via box request.
OAuth-only auth; SKILL.md routes to ten scoped reference files.
Salvaged from PR #52107 by @iskysun96.
Regression tests salvaged from PR #70522 by @JoaoMarcos44. The Qwen flat
cached_tokens behavior is provided by the shared top-level fallback from
PR #66105 (@mehmetkr-31); the codex cache_write_tokens read landed in the
previous commit.
Salvaged from PR #85702 by @JoaoMarcos44, composed onto the mapping-safe
_usage_get reads (PR #74591 by @RelaxJonh) and the flat cached_tokens /
Anthropic-name fallbacks (PRs #66105, #52571):
- cache-write precedence in the chat_completions branch:
details.cache_write_tokens > details.cache_creation_input_tokens >
usage.cache_creation_input_tokens > usage.cache_write_tokens
- codex_responses branch reads details.cache_write_tokens (GPT-5.6+
documented name) with cache_creation_tokens fallback (from PR #70522)
- _usage_count(): clamp malformed negative counters to 0
- all reads in every branch are mapping-safe via _usage_get
When the Responses API returns usage as a plain dict (e.g. from a
middleware or proxy that deserialises JSON to dict instead of a typed
SDK object), normalize_usage() used getattr() exclusively, which
silently returned 0 for every field on a dict.
Add _usage_get() helper that reads via .get() for dicts and getattr()
for attribute-style objects. All accessor sites in normalize_usage()
now use this helper, so token counts and cost are correct regardless
of the usage object's type.
Regression tests: two new tests feed the same payload as both a dict
and a SimpleNamespace through the codex_responses and
chat_completions branches, asserting identical output and non-zero
values.
Kimi/Moonshot's native API (api.moonshot.cn / .ai) reports context-cache hits
as a top-level ``usage.cached_tokens``. The chat-completions branch of
normalize_usage() walks a fallback chain of
prompt_tokens_details.cached_tokens -> cache_read_input_tokens ->
prompt_cache_hit_tokens; none of those names match, so direct Kimi sessions
normalized to cache_read_tokens=0. The hits were invisible in accounting and
the cached prefix was billed at the full input rate.
Appended as the last link in that chain, so it only fills a genuine zero and
cannot override a provider that reports the nested OpenAI shape or DeepSeek's
prompt_cache_hit_tokens.
Rebuilt on current main rather than rebased — the branch was ~3400 commits
behind. The DeepSeek half of the original branch is dropped: 03c0b00f4
(#65678) landed prompt_cache_hit_tokens on main, so this is Kimi-only as the
review asked. The scripts/release.py addition to the frozen LEGACY_AUTHOR_MAP
is dropped too; contributors/emails/mehmet.kar@std.yildiz.edu.tr already
exists on main.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Local OpenAI-compatible servers like mlx_vlm.server emit
input_tokens/output_tokens in chat completion responses instead of
prompt_tokens/completion_tokens. The OpenAI Python client preserves
these as extra attributes, but normalize_usage() only looked at
OpenAI-style field names, causing input_tokens to always be 0.
This made the context progress bar stay at 0% forever and prevented
auto-compression from triggering.
Fix: add Anthropic-style fallback (input_tokens/output_tokens) in the
default else-branch of normalize_usage, with OpenAI-style names taking
priority.
Fixes#14686
Typing '@t' in the composer only ever completed paths/directives —
agent profiles were invisible even though multi-agent UIs (Bot Mode)
route @<profile> as a handoff. complete.path now surfaces matching
profiles for bare-word @queries, ranked above file hits (there are at
most a handful), plus the full list on a bare '@'. The primary profile
is also offered as '@hermes' when no real profile claims that name.
@kind: directive queries are untouched.
Verified live: '@t' -> @turqoise first, '@her' -> @hermes,
'@file:tur' unaffected.
The command wrapper prints the cwd marker after the command returns. A
killed or timed-out command emits no marker, so ``env.cwd`` still holds the
directory of the last command to FINISH. One local environment serves every
session, because ``_resolve_container_task_id`` collapses cwd-only overrides
to ``"default"``. That leftover directory is therefore routinely another
session's.
The post-command dual-write copied ``env.cwd`` into the interrupted session's
durable record. Every later command in that session then ran in the foreign
directory, and the cwd echo told the model it had moved there. A desktop chat
silently re-homed into a worktree that another chat had opened.
Report the observation instead of inferring it. The marker parse now sets
``result["cwd_observed"]``, and both the record write and the echo read that
flag. The local override clears the flag when it rolls back a path that does
not exist, because the restored value is also unobserved. When a command
reports no cwd, the session keeps the directory it already had.
This needs no second session to be wrong: a lone session that interrupts a
command re-adopts a stale value too. A second session only makes the wrong
directory belong to somebody else.
The same class of write exists in the file-tools rescue for a reaped
environment (#26211). That rescue copied the cached snapshot of the shared
``env.cwd`` into the session record. The rescue is now fill-only: it writes
the snapshot when the session has no record, and it never overwrites a
record that the session wrote for itself.
The tests drive ``terminal_tool`` itself through an interrupt, not a copy of
its gate. Review found that a revert of either call site passed the first
version of the tests. Each gate now has a test that fails when the gate is
removed (verified by mutation).
Two exact-dict assertions in the Vercel sandbox tests now assert the two
fields they care about, so a new result key does not fail them.
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Review edge case: OOM dismissal was boot-keyed but live critical/elevated
dismissals were keyed only by severity. Dismiss critical, gateway reboots,
next poll still critical with no observed "ok" in between — the new
incident stayed hidden, and masked the OOM notice too, since critical
takes precedence.
Every dismissal key now embeds boot_id, so a gateway restart invalidates
prior dismissals of any kind. Within a boot, semantics are unchanged:
severity-scoped masking, escalation re-opens, confirmed "ok" recovery
clears live dismissals, "unknown" clears nothing. Missing boot_id
(pre-NS-656 image) degrades to a shared per-severity bucket as before.
Addresses the human review findings on the memory-pressure feature:
* [P2] Dismissal hid later incidents of the same kind. The gateway now
publishes `boot_id` (the lifecycle sentinel's started_at — changes on
every gateway life) in the /api/status memory block, and the dashboard
keys OOM-restart dismissal on it: acknowledging one restart no longer
mutes the NEXT one (the OOM-loop case this banner exists for). Live
pressure dismissals now also reset once pressure is demonstrably back
to "ok" — "unknown" (stale heartbeat) is absence of evidence and
clears nothing. Dismissal storage moved to a JSON list; old bare-string
entries fail JSON.parse and degrade to a clean reset.
* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
heartbeat), not proof the OOM killer acted — banner copy now says
"restarted unexpectedly, most likely because it ran out of memory"
instead of stating OOM as fact.
* [P3] Mobile header clearance was applied per-banner (mt-14 on both
MemoryPressureBanner and ProfileScopeBanner) AND on the content
(pt-14), double/triple-stacking 56px gaps when banners were visible.
Replaced with a single h-14 spacer above the banner stack.
api.ts already declares MemoryStatus for the /api/memory providers
endpoint; the NS-656 pressure block reused the name, and TypeScript
declaration merging fused the two shapes — 'tsc -b' in the Docker image
build failed on every test fixture. Renamed to MemoryPressureStatus.
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
Azure Foundry's OpenAI-compatible Responses surface rejects the post-tool
follow-up payload with HTTP 400 `invalid_payload` when a replayed encrypted
`reasoning` item is sent alongside `function_call` / `function_call_output`.
The initial function-call request and ordinary multi-turn continuity are both
accepted, so the failure only appears after the first tool executes.
Detect the Foundry endpoint in `ResponsesApiTransport.build_kwargs` and drop
only the encrypted reasoning replay on that follow-up turn, leaving
function_call / function_call_output continuity intact.
Salvage of #59981, rebuilt on current main. Same root cause and fix direction
as the original, which was correct; this version resolves three defects:
- No `chat_completion_helpers.py` change. main already forwards `provider`
and `base_url` to the Responses transport, so the original's re-added
arguments produced `SyntaxError: keyword argument repeated: provider` on
merge. Dropping the hunk removed the syntax error and the conflict.
- Host matching uses `utils.base_url_host_matches`, not a substring test.
`".services.ai.azure.com" in base_url` also matches URLs carrying the
domain in a path or query segment, which would silently disable reasoning
replay on an unrelated provider.
- The post-tool predicate tests the trailing messages, not the whole history.
Scanning for any tool call plus any tool result made it sticky: one tool
call early in a conversation suppressed reasoning on every later turn.
- Tool calls pair on `call_id` as well as `id`. Responses histories carry the
function call id in `call_id` while `id` holds the response item id
(`fc_...`). Identity is resolved via the converter's own
`_split_responses_tool_id`, covering composite `"call_x|fc_y"` ids and bare
`fc_` ids on both sides of the pairing.
Tests: 27 cases across the transport and the live `build_api_kwargs` bridge,
including six parametrized tool-call id shapes, non-Foundry host lookalikes,
the sticky-history guard, parallel tool results, and an unpaired tool result.
Each guard was confirmed to catch its defect by reverting the fix.
Verified with `scripts/run_tests.sh tests/agent/ tests/run_agent/`:
532 files, 5602 tests passed, 0 failed.
Not verified against a live Azure Foundry endpoint — no credentials. The
original HTTP 400 reproduction and post-fix Foundry Project / Azure Container
Apps harness runs are @AshuJoshi's, from #59981. This change is verified at
the payload-construction layer only.
Closes#59981.
Co-authored-by: Ashu Joshi <AshuJoshi@users.noreply.github.com>
* fix: mirror voice config (stt/tts/voice) into profiles created via profiles.create
Desktop dictation is profile-scoped: /api/audio/transcribe resolves the
stt section inside the TARGET profile's home. Profiles created through
profiles.create got only a model section, so dictation and TTS silently
fell back to defaults (local whisper, often not installed) — 'voice
dictation doesn't work in bot mode but is fine in regular mode'.
Mirror the launch profile's stt/tts/voice sections (key-wise, never
overwriting sections the clone already has) under the same
mirror_credentials flag that gates .env/auth mirroring, and report it
as mirrored.voice in the receipt.
* guard: route voice-config mirror through canonical loaders
read_user_config_raw (write-back round-trip; load_config would merge
DEFAULT_CONFIG and no-op the mirror) + save_config under the target
profile's HERMES_HOME override — same mechanism as _write_profile_model.
Satisfies test_config_read_guard.
npx hyperframes preview starts a long-lived next-server that keeps
chrome-headless-shell render workers resident. On GPU-less hosts (WSL,
containers, CI) each idle worker falls back to software WebGL (swiftshader)
and busy-spins a CPU core; a preview left open stacks these up until the
host is wedged. The skill never said preview was long-lived, so leaking
was the default outcome.
Add a Cleanup section + pitfall to SKILL.md and a Runaway CPU
troubleshooting entry with diagnose/fix/avoid steps.
Both connect providers captured `sessionId` when the suggestion was built
and ignored the one the pill hands `invoke`. An offer that outlived a
session switch therefore aimed its `reload.mcp` at the session the draft was
sampled in, so the chat the user actually clicked from resumed without the
tools the pill just said were ready.
Prefer the invoking pill's session; the captured one stays as the fallback.
Phase lived in a `Record<key, phase>` that only ever grew, keyed by
`provider:id` — keys that repeat constantly, since a provider withdraws and
re-offers the same suggestion whenever the draft loses and regains its
trigger. A leftover `done` then painted a genuine new offer as "Added
GitHub" and swallowed clicks, because only `idle` invokes.
The same map outlived a session switch. One composer stays mounted across
it, so connecting GitHub in one chat left the next chat's real offer inert.
Withdrawal also stranded in-flight work. The pill is the only cancel
affordance — clicking a working pill sets the flag the provider polls — so
once it left the strip an OAuth flow could poll forever, hold the server's
in-progress slot against a retry, and resolve into a config write with no UI
left to narrate or roll it back.
Phase now lives and dies with the pill: withdrawn keys drop their phase and
flip their cancel flag, unmount cancels everything in flight, and the strip
remounts per session.
The bus's change gate compared offers by `provider:id` alone. Providers
rebuild their suggestion objects on every draft sample, so that key is equal
constantly and the write bailed out — pinning the FIRST object for the life
of the offer.
Two consequences, both user-visible. The pill keeps painting a stale reason
("you mentioned linear" after the user pasted a linear.app link), and it
keeps calling a stale `invoke` closure — work built for a draft that no
longer exists.
Compare the fields the pill actually renders instead. The reference-identity
bail-out survives for the common case (same draft, same match, no re-render),
which is what the gate was there for.
Vercel, Supabase, Netlify, Hugging Face, Asana, Intercom, Airtable,
Webflow, PayPal, and Square join the directory. Every entry is a
vendor-operated remote with its docs page linked, same URL-only rule
as the founding eight. Trigger notes where words are ambiguous:
'square' the English word never fires (squareup only), and
vercel.app/netlify.app deploy-preview hosts are deliberately absent
(a pasted preview link is about the site, not the platform). Brand
glyphs wired for all newcomers.