Optional skill: scroll-as-timeline landing pages on a deterministic
CSS/JS engine, with interview → page grammar → signature move workflow
and screenshot-based scroll verification. Engine and scripts vendored
verbatim; asset generation re-anchored on image_generate with the
upstream kie.ai flow kept as an optional path.
- The Kanban plugin cannot import `@/lib/ime` (plugin fence), so its
`./ime-enter` twin duplicated the predicate. Export `isSubmitEnter` from
`@hermes/plugin-sdk` and drop the twin so one helper owns the policy.
- `BoardNameField` in kanban/board-switcher.tsx (main's refactor of the
create/rename board dialogs) submitted on composition Enter; guarded.
- Telegram allowed-ID input and the clarify-card textarea submitted on the
post-compositionend keyCode-229 Enter; both now use `isSubmitEnter`.
Widens the salvaged Kanban fix (PR #94611 by @huklaa) to the whole bug
class, following cloudflare-os PR #291 which centralized one
isImeComposing predicate and swept every Enter-submit site.
- New shared helper src/lib/ime.ts (isImeComposing / isSubmitEnter):
handles nativeEvent.isComposing (React), event.isComposing (DOM), and
the legacy Chromium/Safari keyCode 229 commit-Enter that arrives after
compositionend.
- Sweeps 21 previously unguarded Enter-submit sites: dialogs (project,
worktree, profile-remote-override, pet rename), settings fields
(credential keys, model API key, quick-entry shortcut, attachment
size, combobox), quick entry, pet overlay composer, pet generate +
hatch, review ship bar, file rename, preview browser address bar,
model catalog menu, MCP setup + approval strips, Kanban drawer.
- Upgrades two partial guards (onboarding, session-actions-menu) that
checked isComposing but missed keyCode 229.
- find-in-page: Enter step no longer fires mid-composition (typed CJK
search queries jumped the viewport on every conversion commit).
Validation: tsc clean, eslint clean, 135 tests green across ime/find-bar/
composer/kanban suites; sabotage run (guard removed) fails the new test.
Pasting exactly one http(s) link while composer text is selected now
turns the selection into a markdown link ([selected text](url)) instead
of replacing it — the behavior every rich text editor ships.
- resolveExactLinkPaste(): recognizes a clipboard payload that is exactly
one supported link (bare or <...>-wrapped, host required, no prose or
trailing punctuation) and returns the href.
- selectionLinkLabel(): the selected composer text eligible for linking;
rejects collapsed selections, selections spanning ref chips or line
breaks, and whitespace-only selections.
- markdownLinkFor(): builds the markdown link, escaping square brackets.
- handlePaste wires the three together ahead of the @url: chip path;
any non-qualifying paste falls through to existing behavior.
Adapted from Buzz's TipTap link-mark approach to Hermes' contenteditable
composer: we emit a markdown link (the composer's native rich construct)
rather than a ProseMirror mark.
The prompt template is built from _PROMPT_GOOD_EXAMPLES, so asserting the
template contains them can never fail independently of the code it checks.
The behavioural invariants (echo rejected, greeting allowed, near-miss
passes) stay.
Small title models parroting a prompt example back verbatim produced
sessions named "Fix login button on mobile" with no relation to the
conversation. The example lines in _TITLE_PROMPT_TEMPLATE now render
from _PROMPT_GOOD_EXAMPLES so the guard set and prompt cannot drift,
and generate_title rejects exact (case-insensitive, wrapper-stripped)
echoes so the instant derived title survives instead. 'Friendly
greeting' stays allowed — it is prescribed output for bare greetings.
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.
- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
and it now honors the base_url argument instead of hardcoding
api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
incl. single-page and stuck-cursor termination; updated the two URL-pinning
pool-discovery tests for the ?limit=1000 contract
Live anonymous probes (2026-09-13, HermesAgent UA, 4 rounds): nemotron-3.5-lightning-free
returned zero bytes for >90s on every attempt and nemotron-3-ultra-free took ~40s, while
mimo-v2.5-free answered 200 in 2-4s each time. An aux model (titles, compression, memory)
that hangs is worse than a delisted one, so the default moves to the model that works.
The hy3-free fixture swap left two identical nemotron-3.5-lightning-free
assertions; use mimo-v2.5-free so the line still covers a second live
free-tier slug.
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).
- hermes_cli/models_catalog_static.py: remove both slugs from the
opencode-free offline floor and the opencode-zen discovery floor;
document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
invariant to cover both.
The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.
Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.
Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.
Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
Drop the no-op and still-running change-detectors (4 invariant tests remain:
pipe closed on finish, PTY closed on finish, poll still serves buffered
output, prune releases handles). The accretion-caps fake session now
carries process/_pty like the real dataclass, since prune reads them.
_release_finished_handles read the dataclass fields through getattr
fallbacks and swallowed every exception; use the real attributes, suppress
only OSError/ValueError on the pipe close (an stdin flush can hit EPIPE),
and rely on ptyprocess/pywinpty close() idempotence for the master fd.
The call-site comment claimed the reader had always drained the pipe;
on kill_process/_reconcile_local_exit the reader may still be reading,
so state what actually happens (its next read raises, the loop exits).
Widen #75162: _prune_if_needed() drops finished sessions (TTL expiry and
oldest-finished eviction at MAX_PROCESSES) — release their Popen/PTY
handles there too, covering sessions inserted into _finished without
passing through _move_to_finished(). The release helper is idempotent,
so double-close on the normal path is a no-op. Adds two tests: prune
releases handles of dropped sessions, and a still-running session's
pipe stays open.
Finished sessions retained their subprocess.Popen pipe objects (and PTY
masters) until the finished-process TTL (FINISHED_TTL_SECONDS, default 30
minutes) elapsed. Under heavy background churn — deployments, archivers,
watchers — finished-but-unpruned sessions accumulated one open pipe FD
each, exhausting the gateway process's file descriptor budget and
surfacing as a 'file descriptor limit' error on new background spawns.
The registry never rejects spawns (it prunes oldest-finished at
MAX_PROCESSES), so the real defect was the retained-handle leak, not a
registry-cap rejection. The fix closes each finished session's Popen
stdout/stderr/stdin streams and PTY master in _move_to_finished(), right
after the reader loop drains EOF. poll()/wait()/read_log() serve output
from the buffered output_buffer — never from the pipe — so the release is
lossless.
Tests: 4 new cases in TestFinishedHandleRelease — Popen pipes closed,
PTY closed, no-handle sessions safe, and poll() still serves buffered
output after the release. All 4 fail on main (reproduction) and pass
with the fix.
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).
A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.
Docs: relevance-floor bullet added to tool-search.md implementation
details.
Adds fal-ai/kling-image/v3/text-to-image ($0.028/img, native 2K default,
8 aspect ratios) with its image-to-image edit endpoint. The i2i schema
takes a SINGULAR `image_url` string instead of the usual `image_urls`
list, so the catalog gains an `edit_image_param` knob that
_build_fal_edit_payload honors (first source image only); the
edit-contract test now validates whichever image key the entry declares.
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
The ACP schema (agent-client-protocol 0.9.0, ContentChunk.messageId) says
"Both clients and agents MUST use UUID format for message IDs". The
ported allocator emitted hermes-assistant-N strings, which a strict
client may reject or fail to group. A fresh uuid4 per message keeps the
grouping semantics and can never collide with an earlier turn's id, so
the counter/prefix state is gone.
Tests trimmed to three invariants: chunks share one UUID until the None
flush sentinel (empty deltas ignored), thought + text share an id, and
the no-allocator shape stays unchanged.
ACP clients group streamed agent_message_chunk / agent_thought_chunk
updates into one assistant reply by messageId, and use a new id to
start the next reply (root-reply replacement semantics). Hermes' ACP
adapter sent every chunk without a messageId, so clients that replace
'the current assistant message' per chunk collapsed separate
autonomous turns into one bubble.
- AssistantMessageIdAllocator (per ACP session, monotonic across
turns): a contiguous run of reasoning + text deltas shares one
hermes-assistant-N id; the None flush sentinel Hermes core emits
before tool execution / at end of stream closes it.
- make_message_cb / make_thinking_cb stamp update.message_id when an
allocator is provided; legacy no-allocator shape unchanged.
- Unstreamed final responses open their own id; plugin-transformed
responses reuse the streamed message's id (replacement).
- Tests: grouping until flush, thought+text sharing, monotonic ids,
empty-string vs None sentinel, legacy shape.
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.
- _render_slack_table_block(): projects rows as 'cell | cell' lines,
collecting text leaves from raw_text/rich_text cell subtrees; capped
at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
(thread history + attachment-nested blocks), the live inbound
attachment loop, and _extract_additional_text_from_slack_blocks
(top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
allowlist drops 'rows', so it only emitted an empty husk.
Parametrize the reject case over send_video/send_document/send_voice
(each asserts no file upload, base fallback never runs, error names the
file, size and limit), fold the all-oversized batch case into the batch
test, and drop the duplicated fixture setup.
send_voice built its own discord.File from the audio bytes and never ran
the size preflight, so an oversized audio attachment still burned the
doomed 413 round-trip that #50846 is about. Route it through the same
_reject_oversized_upload helper as _send_file_attachment.
The limit constant and _discord_upload_limit_bytes lived on the adapter
facade while every consumer is in adapter_media.py; the facade+sibling
layout puts topic code in the sibling, so they move there.
Discord raised the default file upload limit from 10 MiB to 20 MiB for
users, bots, webhooks and interaction responses (developer changelog,
Sep 3 2026). The 25 MiB constant here predates the preflight salvage and
never matched the platform; more importantly discord.py 2.7.1 still
reports 10 MiB via guild.filesize_limit for unboosted guilds, so the
guild-aware path under-reported the cap and rejected 10-20 MiB files
Discord now accepts. Floor the guild value at the platform default so a
stale library constant can only widen, never shrink, the preflight.
The salvaged preflight (#67040) covers _send_file_attachment, but
send_multiple_images opened local files straight into discord.File with
no size check — one oversized image 413'd the whole chunk and dumped its
siblings into the per-image fallback. Preflight each local file against
the same boost-aware limit, skip oversized ones with a user-visible
notice, and still deliver the rest of the chunk.
Sibling-site widening for lobehub-scout salvage of PR #67040 (#50846).
The adapter facade is ~6.6k lines; new behaviour belongs in a topical sibling per the
facade+siblings layout. expand_link_entities() now lives in telegram_entities.py and
reuses the encode/decode UTF-16 slicing the adapter already uses for entity spans.
Also: skip inlining when the anchor text already is the URL (no 'url (url)' duplication),
trim the test file to the invariants and point it at the sibling.
Image.open() was never closed; convert()/resize() return new images so the
source file object lingered until GC (a real leak on Windows where the open
handle blocks later deletion of the original). Use the context manager.
- Keep main's media_write_timeout=60s (PR's HERMES_* env var dropped per
.env-is-secrets-only policy; main already fixed the timeout half).
- Replace stdlib imghdr (removed in Python 3.13) with a magic-byte sniff.
- Exclude GIFs: JPEG conversion flattens animations to one frame.
- Fix transparent-PNG handling: RGBA hit the len(getbands())==4 branch
before the white-background composite, rendering transparency black.
- Clean up temp JPEGs after send (both single and media-group paths);
the docstring promised caller cleanup that neither call site did.
- Write temp files via tempfile default dir instead of an undefined
DEFAULT_OUTPUT_DIR (NameError at runtime in the original PR).
- Add real-Pillow regression tests incl. a sabotage-verified
white-background test.
Behind an HTTP proxy (e.g. tgapi.indevs.in) the PTB
media_write_timeout (~20s) is exceeded by raw PNGs > 1-2MB, causing
TimedOut errors on both send_photo and the send_document fallback.
Add TelegramAdapter._compress_image_to_jpeg() which converts large
PNG/raster images (>1MB) to progressive JPEG at 85% quality, with
optional resize above 1600px. Applied in send_image_file() and in the
media-group path of send_multiple_images(). Compression is a no-op for
JPEGs, small files, and non-raster formats, and falls back gracefully
if Pillow is unavailable.
Co-authored-by: user
Passive update checks no longer run git fetch (338bf9ea9a); the thread
still shells out to rev-parse/remote get-url, which is what the pytest
no-op guards against. Also explain why the predicate checks sys.modules.
The prefetch_update_check daemon thread (started at tui_gateway.server
import time) shells out to git via the shared subprocess singleton at an
arbitrary point after import. Tests that patch subprocess.run/Popen
process-wide can capture that stray spawn in call_args, flaking their
assertions: on 2026-08-28 CI, test_slash_worker_popen_uses_utf8_replace
saw encoding=None from the thread's un-encoded 'git fetch' (red on main,
run 33175879563) and test_deliver_validates_profile_and_runs_transport
captured argv ['rev-parse', 'FETCH_HEAD'] from the shallow-checkout
banner path (FLAKY frame, run 33183215857).
Fix the class at the source: _skip_background_prefetch() makes both
prefetch_update_check and prefetch_banner_data no-ops under pytest
(nothing in tests needs a live update check; the done event is set so
get_update_result callers don't burn their timeout). Tests exercising
the prefetch itself monkeypatch the predicate. Sabotage-verified
regression tests pin both no-ops.
The pinned-argv test read the command line back; replace it with the
invariant: a logged-in user on a gh that rejects --json authenticated
(2.98+) is still reported as authenticated. Red on the old code, green
on the fix. Exercises hermes_cli.doctor_state._gh_authenticated, where
production now reads it (doctor.py is a facade).
gh CLI 2.98+ removed the 'authenticated' field from 'gh auth status --json'
(only 'hosts' remains), causing the command to exit 1 even when the user
is authenticated. The doctor then falsely reports 'No GITHUB_TOKEN' despite
the user being logged in via 'gh auth login'.
Since the code only checks the return code (it never parses stdout), the
'--json' flag is unnecessary. 'gh auth status' without it works across all
gh versions and returns exit 0 when authenticated.
Fixes: gh auth status --json authenticated → gh auth status
Add a regression test asserting the GitHub-auth check invokes plain
`gh auth status` (exit-code based) and never `--json authenticated`.
The existing test only loosely matched cmd[:2], so it passed both before
and after the fix and wouldn't catch a re-addition of the removed flag.
Addresses review feedback on #95162 (point 2).
omo's omo-agent-toolkit worktree-sweep (their PR #7151) added three
capabilities our hermes worktree command lacked:
- --json on list and prune: machine-readable audit/result payloads so
scripts and agents can consume verdicts without scraping table output.
- --older-than DAYS: an age floor that only ever RESTRICTS reaping
(young-but-reapable trees are kept); it never widens eligibility, so
the existing safety invariants are untouched.
- External-tree visibility: linked worktrees registered outside
.worktrees/ are now reported read-only in the audit (branch, locked,
missing) instead of being invisible, and registrations whose
directory has vanished are dropped via git worktree prune (metadata
only, no files touched) during prune.
Not ported: omo's ancestor-of-default-branch merge test (our git cherry
patch-equivalence is strictly stronger under rebase/squash merges), and
their hardcoded external-root exclusion list (we exclude by location:
everything outside .worktrees/ is hands-off).
Tests: 9 new contracts in tests/hermes_cli/test_worktree_gc.py (age gate
restrict-only, external trees never reaped, stale-registration prune
dry-run/real, JSON shapes, negative --older-than rejected). Live E2E on
a scratch repo verified all three flags end to end.
Tests collapse 13 change-detectors into 7 invariants (silent below threshold /
without context, fires above, same-model re-select silent, config override and
0-disables, registry threading, agent context derivation). The legacy 5-arg
guard test goes with the TypeError fallback it covered: that fallback was
defence for a case nobody has (every in-tree guard and test double is
*args-tolerant) and would re-run a guard whose real TypeError it masked, so the
rebased port passes the context positionally like every other argument.
Docs no longer claim the confirm fires on the Telegram/Discord pickers or the
dashboard: those surfaces call combined_selection_warning() without a live
agent, so the context-cache guard is (correctly) silent there.
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).
- hermes_cli/model_selection_guards.py: new context_cache guard +
SelectionContext carrier + selection_context_for_agent() helper;
registry threads live-session facts to guards (6-arg signature with a
TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
Twelve credential-driven env branches were routed through _enable_from_env on
main (867e4158f0, #48820), which honors the loader's `_enabled_explicit`
marker. The flag-driven WhatsApp step was the one survivor: `_whatsapp` still
set `wa_cfg.enabled = True` on WHATSAPP_ENABLED=true regardless of an explicit
YAML disable. The dashboard's disable action writes only
`platforms.whatsapp.enabled: false` and leaves the env flag on disk, so the
Baileys bridge reconnected to real contacts on the next full restart
(reported live on #73289).
Route the truthy branch through _enable_from_env like every other platform;
WHATSAPP_ENABLED=false still forces a disable. Register the flag in
_ENV_ENABLE_CREDENTIALS so the one-time explicit-disable WARNING can name it.
Remaining scope of #96557 (the other ~20 sites) landed on main in 867e4158f0
and the config_env.py extraction; this is the delta. Fix direction from
@CryptoDombili in #73303.
Co-authored-by: Professor Dombili <Cryptodombili@gmail.com>
A scheduled job whose prompt carries its own cadence phrasing ('Each
Monday, review...') can convince the agent to create ANOTHER cron job
at execution time instead of just doing the work — each run spawning a
sibling job. Hermes policy-denies the cronjob toolset in cron context
by default, but cron.allow_agent_scheduling: true re-enables it and
opens exactly this loop.
_build_job_prompt now extends the always-injected cron hint with a
RECURSION clause: this is a run of an existing job; never create or
update a cron job from schedule language in the task prompt; treat
cadence phrasing as context for this run.
Adapted from paradigmxyz/centaur#1479 (same failure mode in their
scheduled-task workflow runner).
main now passes persist_user_platform_id (and friends) into run_conversation; the fakes
only need the message, so swallow the rest with **_kwargs instead of pinning today's signature.