Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.
utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.
Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
They are teaching examples in hermes-example-plugins, not products users install; docs still link them there. Delist only (not a security pull), so removed.yaml stays empty.
create_main used O_WRONLY|O_CREAT on the main file and closed the fd, which drops this process's POSIX locks whenever state.db already exists (the gateway's own async_delegation import path). O_EXCL restricts the descriptor to a brand-new inode; existing files take the chmod(2) path.
Refs #109786#109687
POSIX fcntl locks are owned per (process, inode): closing any descriptor
for state.db releases every lock the process holds on that inode,
including the locks of an already-open SQLite connection. The
owner-only hardening cycle opened the live database and its -wal/-shm
read-only, fchmod'ed, and closed, so any process that already held a
connection (gateway, desktop hermes serve, dashboard share one) dropped
its live locks on every SessionDB init. A sibling process then took the
shared-memory DMS exclusively at its own close, checkpointed, and
unlinked the sidecars while long-lived holders kept the deleted inodes
open, tripping the deleted-WAL generation guard.
chmod(2) on the path never opens the file, so it cannot disturb locks.
The descriptor path remains only for first-time main-db creation, where
no locks can exist yet.
Widens the new hook to the sibling interrupt surface: the TUI/desktop
session.interrupt path stops a live turn exactly like the gateway's
/stop, so plugins holding per-turn external resources get the same
signal there (platform='tui'). Gated on a genuinely running turn;
dispatch failures are swallowed so a plugin can never break the
interrupt. Docs updated to describe both surfaces.
Inspired by ChatGPT Work / Codex CLI 0.150.0 'Interrupt' hooks
(hooks that run when an active top-level turn is interrupted).
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.
_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.
Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.
Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
meta/muse-image/text-to-image + paired meta/muse-image/edit, the FAL
listing of Meta's Muse Image model (launched on the Meta Model API in
Aug 2026 at $0.01/image).
- aspect_ratio size family (16:9 / 1:1 / 9:16 from the vendor's
21:9..9:21 enum); always sent on t2i for deterministic framing,
deliberately omitted on edits so Muse follows the input image.
- No seed in the vendor schema (Grok Imagine 2.0 precedent) - the
supports whitelist filters it.
- Edit takes 1-10 reference image_urls (max_reference_images=10).
Schema verified against FAL's OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403), same as prior catalog
additions.
The Matrix docs described six agent-exposed matrix_* tools and three MATRIX_TOOLS_ALLOW_* env gates that were never implemented — the tool names and gates appear nowhere in code. Docs now describe actual behavior: no Matrix-specific agent tools; reactions/redactions are internal to approval prompts and pickers; MATRIX_ALLOWED_ROOMS scopes responses. Fixes#100535.
Slack deprecates the Assistant messaging experience (assistant_view) in
February 2027: assistant.threads.setStatus/setTitle are replaced by
agents.sessions.setStatus/rename. slack-sdk 3.44.0 (Aug 27 2026) ships
the typed methods with drop-in-compatible signatures.
- adapter: capability probe on the AsyncWebClient CLASS (never instance —
mock auto-attributes lie), cached; status set/clear + thread title route
through agents.sessions.* when available, legacy otherwise
- pins: slack-sdk 3.43.0 -> 3.44.0 (pyproject messaging+slack extras,
lazy_deps, uv.lock)
- tests: autouse fixture pins the probe to legacy under the mocked SDK;
5 new tests cover both routing paths for typing, clear, and title
- docs: slack.md scope table + status-line notes mention both methods
"0 3 * * *" fires at a fixed UTC minute; when CI runs in the half hour before
it (observed 02:30:56Z), the 30-minute rung lands after the natural occurrence
and plan_retry correctly yields to the schedule, clearing the state the test
asserts on. "every 24h" always has its natural fire a full day out, so every
rung is strictly earlier regardless of when the test runs.
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."
A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.
Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.
- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
Port from earendil-works/pi#9300 fix (acaa253cc): a plugin registering a
tool whose schema["parameters"] is not a dict (a list, string, etc.)
previously registered fine and the malformed schema was serialized into
every provider request, 400-ing turns far from the offending plugin.
Live probe on main confirmed the bad schema flows into _fn_def() and the
OpenAI wire unchanged.
Fail at registry.register() with the tool name in the error instead. The
plugin loader already catches registration exceptions and marks the
plugin errored, so a broken plugin degrades gracefully rather than
breaking every session. Schemas that omit "parameters" stay valid
(no-argument tools); MCP tools are unaffected (their schemas pass
through _normalize_mcp_input_schema first, which always returns a dict).
Programs run via terminal(background=true, pty=true) can block forever when
they probe their terminal — device-status (ESC[5n), window-size (ESC[18t),
cursor-position (ESC[6n), or DEC private-mode (ESC[?N$p) queries — because
nothing on the PTY master side answers, and the raw query bytes leak into
captured output.
- tools/pty_query_responder.py: incremental byte scanner that strips the
handled queries from PTY output (chunk splits included) and produces
bounded replies; everything else passes through untouched.
- tools/process_registry.py: wire the responder into _pty_reader_loop
(POSIX only — ConPTY answers its own queries); flush partial escape
tails at end-of-stream.
- tests mirror the codex fixtures plus a live-PTY E2E where a subprocess
blocks on ESC[6n until answered.
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
Browser login against openrouter.ai/auth (S256 PKCE, bind-first OS-assigned loopback port,
POST /api/v1/auth/keys code exchange) that stores the minted key as a plain api_key pool entry
with source manual:openrouter_pkce, so it rotates and resolves exactly like a pasted key.
Salvaged from #102639 (nyx573) onto the facade+siblings layout: the flow lives in the new
auth_openrouter sibling and rides the shared loopback helpers instead of appending to the
auth.py facade; auth_commands gains a table row rather than a provider branch. OpenRouter echoes
no `state`, so the CSRF nonce rides in the callback path (a redirect that guesses the port but not
the nonce is a 404 and never reaches the exchange); remote/SSH sessions use OpenRouter's
documented headless paste-the-code mode instead of an unreachable loopback listener. OpenRouter
keeps its API-key default when --type is omitted so the documented `--api-key` form is unchanged.
Move the S256 verifier/challenge pair, the loopback callback handler, the bind-first
listener and the serve-until-redirect loop out of auth_spotify into auth_device_flow so a
second loopback PKCE provider does not copy 80 lines of HTTP-server plumbing. Spotify's
behaviour and error codes are unchanged; only its private copies are deleted.
Carve the direct-child list reconciliation from Indigo Karasu's earliest
PR #60506 (2a96ae2cbf806ccdc3e9b584911774f32622f421), corroborated by
fangliquanflq's narrow #81385 (50ffd243d92627e4a03a3ee8427ad0f5e090c3ab).
Run the existing helper after task/session filtering and reuse the
idempotent owner-stamped completion path. Do not import cross-session
disclosure, bare-PID healing, forget RPCs, or reader rewrites.
A real-child regression fails before this change and passes after it:
the direct child exits while its descendant keeps writing to stdout;
listing reports exit without consuming the result or waiting for EOF.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
The multiplex-scoped rebase replaced the flat _permanent_baseline set with
_permanent_baseline_by_home (keyed by profile home, "" = unscoped); the
fixture must reset and seed that map.
Review follow-up. Two of the three items taken as written; the third declined
with a reason.
1. Taken. The reconcile semantics mean `patterns` may only ADD -- an entry left
out of it is not removed, because the on-disk list wins for anything this
process did not approve itself. Every caller in the tree is additive today,
so nothing breaks, but the signature does not say so. Stated in the
docstring, and pinned by
`test_a_caller_that_passes_a_smaller_set_does_not_remove` so a future
`allowlist remove` finds out here instead of in production.
NOT taken: the `reconcile: bool = True` opt-out. There is no caller that
wants it, and AGENTS.md:98-101 names exactly this -- "Speculative
infrastructure. Hooks, callbacks, or extension points with no concrete
consumer." The removal path is editing config.yaml, which the docstring now
says.
2. Taken. website/docs/user-guide/security.md, next to the existing
`hermes config edit` tip, which is where an operator reads about removing a
pattern: the list is read at startup, a pattern removed while a session is
running stays approved in that session until the next write or a restart,
and if it was removed for safety reasons, restart.
3. Taken. `test_save_failure_is_logged_not_raised` asserted non-raising but
never asserted the log its name promises. Now asserts
"Could not save allowlist" via caplog.
scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
=== Summary: 1 files, 9 tests passed, 0 failed (100% complete) in 0.4s
`load_permanent_allowlist()` runs exactly once, at module import
(tools/approval.py, the call at the bottom of the module), and
`load_permanent()` only unions into `_permanent_approved` (:2866-2869) --
nothing ever removes. `save_permanent_allowlist()` then wrote that in-memory
set straight back over `config["command_allowlist"]`, at eight call sites.
`command_allowlist` is a file the operator edits, and deleting a line from it
is the documented way to withdraw a standing approval. Any hand edit made
while a Hermes process is live was undone by that process's next `[a]lways`,
in both directions at once.
Reproduced on this tree with a temp HERMES_HOME:
BEFORE (tools/approval.py at fcbd107)
on disk before this process starts : ['git status', 'ls *']
operator edits config.yaml by hand : ['ls *', 'npm test']
(revoked 'git status', added 'npm test')
after ONE [a]lways : ['docker *', 'git status', 'ls *']
is_approved still honours revoked? : True
AFTER
after ONE [a]lways : ['docker *', 'ls *', 'npm test']
is_approved still honours revoked? : False
`npm test` was silently deleted from the operator's own config file, and
`git status` -- a standing approval they had just withdrawn -- was written
back and kept auto-approving. Neither prints anything.
The same shape loses writes between two live Hermes processes: whichever
saves second overwrites the other's entry.
The fix reconciles at write time. The file is re-read and the result is what
is on disk now, plus what this process approved since its own baseline, where
the baseline is what `command_allowlist` held the last time this process
synchronised with the file. That difference is what separates "the operator
granted this here" from "this was on disk at import and may since have been
revoked". Revoked entries are also dropped from `_permanent_approved` so
`is_approved()` stops honouring them for the rest of the process.
It does NOT make a revocation take effect the instant the file changes --
nothing re-reads the file on the approval hot path, and adding a stat there is
a separate change with its own cost. It makes the next write stop undoing the
operator's edit.
`_lock` is `threading.Lock` and not reentrant; all eight call sites were
checked and none holds it across the call, so the added critical section
cannot deadlock. The failure path still logs and returns rather than raising,
as before.
Searched open and merged PRs and issues for `command_allowlist revoke`,
`permanent allowlist reload`, `approval allowlist clobber` and
`save_permanent_allowlist` -- nothing covers this.
Tests: tests/tools/test_permanent_allowlist_reconcile.py, 8 cases -- both
halves of the bug, the two-process race, idempotence, the unedited round trip,
and the existing contract that a config write failure is logged rather than
raised.
scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
=== Summary: 1 files, 8 tests passed, 0 failed (100% complete) in 0.5s
No regression across the 29 test files in tests/ that touch the allowlist or
the approval module: 25 failed before and after, byte-identical failure set
(pre-existing missing-dependency failures in my local venv).
Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
test_prompt_stash_cli.py stubs prompt_toolkit with a bare module, so the
module-level 'from prompt_toolkit.enums import EditingMode' crashed every
test importing cli. Import it with the same ImportError fallback as
CursorShape and pass editing_mode via extra_kw only when available.
- Relocate _handle_vim_command and _vim_mode_label onto
cli_status_bar_mixin.py (the mixin owning status-bar commands and
rendering) — the original targeted pre-decomposition cli.py; slash
dispatch picks the handler up by naming convention
- Wire the vim label into the current status-bar fragment builder
- Trim tests to 3 invariant tests per the salvage bar
- Document /vim in reference/slash-commands.md
Live E2E (tmux + PTY, temp HERMES_HOME): /vim toggles on, Esc/i flip
NORMAL/INSERT in the status bar, /vim off restores emacs bindings,
display.vim_mode persisted to config.yaml.
Adds 11 tests for the vim mode surface:
- /vim status reports without persisting
- bare /vim toggles; on|off set explicitly; each persists to
display.vim_mode
- editing_mode is applied to a running Application in both directions
- a missing Application is not fatal (toggle before the TUI starts)
- invalid arguments leave state untouched and print usage
- _vim_mode_label() maps vi input modes to NORMAL/INSERT/REPLACE and
stays empty when disabled or before the app exists
- the config default is off and /vim is registered with subcommands
Implements vim mode for the classic CLI input composer:
- Pass editing_mode=EditingMode.VI to Application when display.vim_mode
is set, else EditingMode.EMACS (prompt_toolkit's own default), so the
change is inert for users who have not opted in
- Add _handle_vim_command() to toggle at runtime and persist the choice
to display.vim_mode; the live Application is updated in place so the
toggle takes effect without a restart
- Add _vim_mode_label() and surface NORMAL/INSERT/REPLACE in the status
bar while vim mode is active, as requested in the issue
Credit to #4325 by @SHL0MS, which first identified the EditingMode
wiring and the /vim toggle; that PR has gone stale against main. This
revives the approach, adds the missing config default and the status
indicator, and covers it with tests.
Closes#4254.
Introduces the opt-in surface for vi editing in the input composer:
- display.vim_mode config default (False, so existing users are
unaffected and prompt_toolkit keeps its standard emacs bindings)
- /vim command registered with on|off|status subcommands, matching
the established /battery and /timestamps pattern
Part of #4254.
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.
Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
Port from cline/cline#13970: models that send patch calls with an empty
old_string got back 'old_string cannot be empty' — an error that names the
problem but not the recovery, so the next call was byte-identical and the
run burned turns until loop detection killed it (upstream repro: Kimi K3
looping on old_text: null).
The rejection now states the recovery: set old_string to the exact text the
replacement should replace, read the file first if unsure, use write_file
for new files/full rewrites, and do not re-send the call unchanged. The
whitespace-only rejection gets the same treatment. No behavior change for
valid calls.
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).
LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.
- `_status_429` now honors a structured billing code first (decisive
signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
(429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
limit" was already there; providers use both orders). "terminal billing
limit" text is deliberately NOT matched: substring rules cannot negate
the "non-terminal billing limit" wording — the structured code covers it.
The streamed download path now routes through the SSRF-safe client, which
(correctly) refuses the test fixture's 127.0.0.1 registry. Set
HERMES_ALLOW_PRIVATE_URLS for the fixture's lifetime and reset the module
cache on both sides so the guard still fail-closes everywhere else.
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.
Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.
Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.
Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.
Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:
- Snapshot status/pid/claim INSIDE the archive txn so the kill only
happens when this caller wins the archive transition; a losing
concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
hold the SQLite write lock. Safe because archived is terminal — no
dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
-> no signal, no event), live E2E: worker survived archive on main,
terminated (<0.3s, clean SIGTERM) with the fix.
Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
archive_task() was a pure DB operation — it cleared worker_pid,
claim_lock, and status from the tasks row but never sent SIGTERM
to the actual OS process. A running worker stayed alive until it
next called kanban_complete/kanban_block and discovered it was
archived, burning API quota and compute resources.
Fix: snapshot pid+claim_lock before the write_txn clears them,
then call _terminate_reclaimed_worker() — same function reclaim_task
uses — which sends SIGTERM, waits 5s, then SIGKILL if still alive.
The termination metadata is included in the 'archived' event so
operators can see what happened.
Order matches reclaim_task: terminate first, then DB update.
For non-running / non-local tasks, _terminate_reclaimed_worker
returns immediately as a no-op.
Closes#33774 reprise: the scratch-workspace side was fixed in
fc8afd500, but the orphaned-process side was never addressed.
Port of https://github.com/achimala/dream-loop (MIT, 400+ stars in 48h).
An autonomous loop for building visually impressive 3D scenes/games:
generate photorealistic concept art, build (three.js/WebGL/Blender),
screenshot the live build, judge screenshot-vs-concept on a 5-tier score
ladder, iterate to convergence with explicit exit criteria.
Prose-only port rebound to Hermes-native tools: image_generate for
concept art, vision_analyze for judging (side-by-side composite
workaround documented), browser_exec capture_screenshot for live builds,
delegate_task for parallel asset work. Upstream ladder, failure modes,
time-budget and exit rules preserved. optional-skills/ placement.
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.
_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.
Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
A community deep-dive found that running one never-ending gateway session for weeks means memory injection, session_search, and pre-reset distillation almost never fire, while token cost grows. Sessions doc stated conversations never expire but never explained why users should still create boundaries. Adds a Session hygiene subsection with the mechanism and a practical rule.
Three CI-load flakes from the same class — a fixed per-test timeout billed
for one-time module-transform/env-init cost:
- apps/desktop messaging/index.test.tsx: `await import('./index')` ran inside
renderMessaging(), so the FIRST test paid the whole MessagingView transform.
On loaded runners that alone blew the 15s testTimeout and cascade-failed all
subsequent tests in the file (unmounted DOM). Red on main runs 34599517793,
34600757569, 34601269252 (green file takes 15.7s on a green main run —
already over the first test's budget when billed there). Import moved to
module scope, where vitest bills it to collection.
- apps/desktop skills/index.test.tsx: same pattern, 9 call sites; the file ran
18.6s on a green main run. Deduplicated to one module-scope import (the
existing 60s describe-timeout stays for the legitimately slow tests).
- web SessionsPage.test.tsx: the web vitest project still ran on vitest's 5s
default while its per-row routing test legitimately takes 3.6-4.6s on GREEN
runs; run 34600757569 tipped it to 5079ms. Gave web/vitest.config.ts the
same 15s testTimeout the desktop project already carries, with the same
rationale comment.
Validation: both desktop files 5x consecutive green + green pinned to 1 CPU
core (worst-case contention); SessionsPage 3x green; full desktop ui project
(801 files / 7622 tests) green; tsc + eslint clean on touched files.
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.
Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).
Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."
Fixes#57155
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Ports the system-atlas agent skill: one data.mjs file renders both an
interactive isometric HTML map (progressive-disclosure chapters, moving
data packets, question tracking by Q-ID) and a generated text twin
(SYSTEM.md) for the repo.
Why: architecture discussions that outgrow a single diagram — the skill
encodes hard-earned rules (max 3 structures per chapter, shapes+labels,
docs-policy ask before committing). Upstream 410 stars since Aug 20,
VoltAgent-listed; complements architecture-diagram (static SVG) with an
explorable, stateful artifact.
- optional-skills/creative/system-atlas/: SKILL.md (89 lines),
references/ (design language, process lessons), assets/ (build.mjs,
template.html, data.example.mjs) vendored verbatim; LICENSE.txt (MIT,
Harshyt Goel)
- Live smoke: node --check both .mjs OK; node build.mjs produced
atlas.html (41.5 KB) + SYSTEM.md from the example data, zero deps
- docs: own catalog row + generated page + sidebar entry only
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).
wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.
Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.
Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
Google Docs can hold a tree of tabs, each with its own body and its own
1-based character index space; the legacy top-level `body` only carries
the first tab. `docs get` was silently dropping every other tab's
content, and `docs append` computed its insert index from the first
tab's body and sent the write with no tabId — so on a tabbed doc the
append could land at a wrong offset in the wrong tab.
- Requests now pass includeTabsContent=true (both gws and SDK paths).
- `docs get` returns a `tabs` array (preorder flattening of the
tabs/childTabs tree, nested tabs included); single-tab docs keep the
`body` field so existing callers work, and `--tab <tabId>` reads one
tab. Legacy no-tabs responses are unchanged.
- `docs append` targets exactly one tab: the insert location carries
the tabId, the end index is computed inside that tab's own body, an
unknown `--tab` errors instead of falling back to the first tab, and
a multi-tab doc without `--tab` errors with the tab list rather than
guessing. Tabs are never merged — index spaces are independent.
Adapted from cloudflare/cloudflare-os#450 (gatekeeper-google), which
fixed the same provider behavior: reads must traverse Document.tabs and
every write Location must carry the immutable tabId.
Two pre-existing bare read_text/write_text in the touched test file
gained encoding="utf-8" (windows-footgun sweep rule).
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.
- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
after the read loop; surface EOF-without-[DONE] (warn + error result
when nothing was received); extract _consume_sse_line so line parsing
and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
proven red on origin/main.