Commit Graph

33187 Commits

Author SHA1 Message Date
chelsealong 21d5df4100 fix(session-state): keep the canonical Bot Chat hidden when pinned
set_session_pinned's unconditional hidden-clear (0be7f931) also
unhides the canonical Bot Chat, which the desktop contract requires
to stay hidden and reachable only through the bot row. Exposing it
breaks the sidebar and disables the rename guard that protects its
identity title.

Skip the unhide when the pinned row is hidden and carries the exact
canonical title; ordinary hidden sessions are unaffected.
2026-09-09 10:52:21 -07:00
chelsealong c59aa2d041 fix(session-state): clear hidden flag when a session is pinned
A bot-mode session is created with hidden=true. Pinning it only wrote
pinned=1 and left hidden=1, so the row was filtered out of both the
default listing (`s.hidden = 0`) and the pinned back-fill (which reuses
the same WHERE), making a pinned session vanish from the sidebar
entirely. set_session_pinned now clears hidden across the session's
compression lineage whenever pinned is set to true.

Fixes #106171
2026-09-09 10:52:21 -07:00
teknium1 87cc4de430 fix(dashboard): place the deactivation ref reset after the callbacks that also write those refs
PR #106445's effect was inserted above `reconnectPty` / `startFreshPty` /
`startFreshDashboardChat`. React Compiler's `react-hooks/immutability`
rule then treats `ptyInputLineRef` and `mobileReplacementInputUntilRef`
as effect-owned values and flags the six pre-existing assignments in
those callbacks as errors ("This value cannot be modified"), which is
what turned the `JS & TS checks` job red. Declaring the effect after the
callbacks keeps the rule quiet; eslint on the file is back to main's
0 errors / 3 warnings.

Adds one pure-function vitest on `normalizePtyMobileInput` proving the
symptom: a stale tracked line ("hel") makes an in-window re-emission of
"hello" come out as 3×DEL + "hello", which over a composer already
holding "hello" leaves "hehello"; with the tracker reset the same input
passes through untouched.
2026-09-09 10:48:49 -07:00
Konstantin Khlopkov cbd1018ea6 fix(dashboard): clear input refs on tab deactivation (#106403)
(cherry picked from commit cb6fd96ef0521ed887bb136e12dadf2937053c3c)
2026-09-09 10:48:49 -07:00
KeyArgo 070128f056 test(agent): provider-neutral name for 429 server-overload regression test 2026-09-09 10:47:49 -07:00
KeyArgo 43ecad8fe6 fix(agent): classify Novita 'server overload' 429 as overloaded
Novita returns HTTP 429 with message 'server overload, please try again
later' and error type 'server_overload' when its server is genuinely busy.
Neither phrase was in _OVERLOADED_PATTERNS, so the 429 fell through to
_V_RATE_LIMIT and set should_fallback=True + should_rotate_credential=True —
rotating the credential / falling back early instead of retrying the same
key. Add 'server overload' and 'server_overload' to the overload tuple so
this reaches the existing _V_OVERLOADED verdict (retryable, no rotation).
Closes #106205
2026-09-09 10:47:49 -07:00
Teknium fb85b57ee9 docs(matrix): note LaTeX math rendering in the Matrix behaviour table 2026-09-09 10:43:58 -07:00
Teknium 1d651b3bb2 feat(matrix): render LaTeX in the standalone (cron) sender too
_standalone_send builds formatted_body through its own markdown call and
never went through _markdown_to_html, so cron-delivered equations still
arrived as raw dollars. Tokenize/expand around that conversion as well.
2026-09-09 10:43:58 -07:00
Teknium e22e04692a test(matrix): trim LaTeX salvage to two invariant tests, tidy _markdown_to_html
Keep the two tests that fail without the fix: inline + display math reach
formatted_body as data-mx-maths markup with the TeX HTML-escaped while the
plain body keeps the raw TeX; unpaired dollars and text colliding with the
sentinel format pass through unchanged (no IndexError). The 17 unit tests
from #106452 are dropped per the salvage bar (≤ 2 invariant tests).

Also drop the dead `_tex_store = []` pre-assignment and the `_fb`/`html`
temporaries in `_markdown_to_html` — pure tidy, no behaviour change.
2026-09-09 10:43:58 -07:00
romanovzky 199f4d5b96 feat(matrix): render LaTeX math via Element data-mx-maths markup
Element (feature_latex_maths) typesets <div|span data-mx-maths="TEX">
elements at display time, but the outbound HTML sanitizer allowlists tags
and attributes, so data-mx-maths markup sent by the gateway never reaches
Element intact - messages containing $...$ render as raw dollars.

Convert $...$ (inline) and $$...$$ (display) to opaque sentinel tokens
before Markdown conversion and expand them to data-mx-maths markup after
sanitization. Tokens are printable text with no HTML/Markdown meaning, so
neither the converter nor the sanitizer touches the TeX. Unpaired dollars
(prices, literals) are untouched, and adversarial text colliding with the
sentinel format passes through verbatim (index-checked expansion).

(cherry picked from commit eb73aafe0e08502c54e63dbd03e99796c3b7fee3)
2026-09-09 10:43:58 -07:00
Teknium 1a500a43a7 fix(gateway): keep launchd_stop's bootout quiet too; prove the fd-2 invariant with a fake launchctl
Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.

Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
2026-09-09 10:42:32 -07:00
Davy ebcc87ad4f fix(gateway): silence expected launchctl bootout/kickstart noise on macOS
Best-effort bootout calls (unloaded-job recovery, stale-EIO retry,
plist refresh, uninstall) and the handled kickstart -k in
launchd_restart inherited the terminal's stderr, so an expected
unloaded job printed raw launchctl errors around the CLI's own lines:

  Could not find service "ai.hermes.gateway" in domain for user gui: 501
  ↻ launchd job was unloaded; reloading
  Boot-out failed: 3: No such process

Capture them with _CAPTURE_TEXT instead. The kickstart error stays
available as e.stderr for the update_cmd failure diagnostic, and the
post-bootstrap kickstart intentionally stays loud (its failure feeds
the domain-unsupported fallback). Same precedent as the reload
helper, which already runs bootout with 2>/dev/null.

[salvage: picked hermes_cli/gateway.py only; the two capture_output kwarg-assertion tests are
replaced by fd-level invariant tests with a fake launchctl in the follow-up commit]
2026-09-09 10:42:32 -07:00
teknium1 79d73e8f63 test(telegram): trim bots_require_mention coverage to the two invariants
Keep the loop-breaker test (bot quote-reply processed on main, dropped with the
flag, explicit @mention still passes) and the human quote-reply test (unaffected
with the flag on). Drop the defaults-off snapshot, the /cmd@botname variant, the
plain-chatter duplicate and the own-echo helper test: they re-cover paths
(_message_mentions_bot, _is_own_message) already pinned elsewhere in this file.
2026-09-09 10:37:37 -07:00
Konstantin Khlopkov 1a981d8e83 fix(telegram): bots_require_mention gates bot quote-replies behind an explicit @mention (closes #106430)
(cherry picked from commit a67fa8c4b21e31a44019d850aedc5be9015a546d)
2026-09-09 10:37:37 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
teknium1 113304199e refactor(desktop): gate the WSLg D3D12 selection inside the helper and test the real launch env
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.

Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.

Follow-up to Xipong's fix for #106117 (salvaged from #106118).
2026-09-09 10:28:52 -07:00
Xipong 51fa04d49d fix(desktop): select installed WSL D3D12 driver before Electron exec 2026-09-09 10:28:52 -07:00
teknium1 a785db3672 fix(cli): inline the $HERMES_HOME/npmrc lookup, trim tests to two, document it
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.

Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.

Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.

Refs #106373
2026-09-09 10:26:06 -07:00
Konstantin Khlopkov ce4a33a9f7 fix(cli): keep $HERMES_HOME/npmrc npm config across updates (#106373)
(cherry picked from commit 339368f2215ffcfd04acf295c81806bf1e850770)
2026-09-09 10:26:06 -07:00
teknium1 d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
KoNit-K ef903617ea docs(providers): note the OpenSSL 3.5 PQ-group middlebox failure on Codex device login
Docs hunk carried from PR #106389 (hermes_cli changes superseded by #106394's
smaller equivalent); the reporter's exact openssl.cnf snippet follows in a
maintainer commit.

Refs #106384
(cherry picked from commit abb323d86ced829859477ed72d59f409c8d8b895, docs hunk only)
2026-09-09 10:14:58 -07:00
liuhao1024 d8eb177c93 fix(cli): keep SSL detail and add middlebox hint on Codex device-login transport errors
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.

- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
  (device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
  workaround hint; KeyboardInterrupt handling is unchanged

(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
2026-09-09 10:14:58 -07:00
Teknium 42d28d64f0 test(hermes_state): trim compacted-paging tests to the invariants that fail on base
Six new tests collapsed into three that each go red without the indexed
projection: bounded VM steps for a page read and for an append (the
symptom), composite user-handoff identity + first-row position + internal
columns never leaking + signed-zero timestamp identity (the parity
contract), and a legacy store read-only then lazily migrated (the
persistence contract). The parametrized symmetry/visibility-toggle test
exercised the same key through raw SQL edits no production path performs
and passed on base, so it was a change-detector for the trigger shape.
2026-09-09 10:05:59 -07:00
Xipong 1c6683e8e0 fix: index compacted display identity writes 2026-09-09 10:05:59 -07:00
Xipong 49e6d661a0 fix: bound compacted display history paging 2026-09-09 10:05:59 -07:00
Eva 2536772301 fix: preserve summary boundaries when restoring model replay 2026-09-09 09:59:13 -07:00
teknium1 5de30f36f1 refactor(gateway): reuse recorded_gateway_home_conflicts for the scoped PID home check
The salvaged fix added `_pid_record_matches_home`, a near-copy of
`recorded_gateway_home_conflicts(record, expected_home=...)` (already the
scoped-home predicate used by the #89315 stop guard). Reuse it instead of
carrying a second helper with the same semantics (legacy records without
`hermes_home` are accepted by both).

Also trims the salvaged tests to the two invariants that prove the symptom
(scoped probe reports a live foreign profile's PID and leaves its
identity files intact; a dead-PID scoped record is still cleaned). The
third-home-claim variant exercised the same `saw_live_pid` branch.
2026-09-09 09:55:52 -07:00
liuhao1024 c2a3aff3f8 fix(gateway): keep unscoped poison-file cleanup after scoped PID scoping (#106406)
The live-record unlink protection must only apply to scoped reads
(polling another profile must not delete its identity files). The
unscoped path keeps main's behavior: a live record owned by another
home inside this home's gateway.pid is unlinked on refusal, per the
#89315 cross-profile stop contract
(test_cross_profile_kill_refusal.py).

(cherry picked from commit 1c7343f9c0175867b9649ab8c54cf924d04e2266)
2026-09-09 09:55:52 -07:00
liuhao1024 a23319d0f7 fix(gateway): scoped PID queries validate against the probed home (#106406)
get_running_pid(pid_path) validated records against the serve process's
HERMES_HOME, so a dashboard scoped status poll (?profile=) rejected a live
foreign profile's record, force-unlinked its gateway.pid/gateway.lock, and
reported running=false. Validate scoped queries against the probed home
(pid_path.parent, like the expected_home sibling) and never cleanup-unlink
a live record's identity files; dead-PID records still clean up.

(cherry picked from commit 472ca1b8d87df5246d83864e9b1492d06e99e8e5)
2026-09-09 09:55:52 -07:00
Teknium 69c999f8a7 fix(agent): keep the tripping correction and explain the restart-bound exit
When the redirect cap trips, the correction that cancelled the final attempt is
still sitting in _pending_redirect; finalize_turn's clear_interrupt() would drop
it silently. Drain it into the steer slot so it rides result["pending_steer"] and
becomes the next user turn on every surface that already honours that key.

Both new exit reasons get a turn-completion explanation so the user sees why the
turn stopped instead of an empty reply.
2026-09-09 09:51:31 -07:00
Teknium b24a781b6f test(agent): trim the restart-bound tests to two invariants
Collapse the four class-based tests into two parametrized invariants over both
refunding restart flags and move them to tests/agent/ (the phase modules live in
agent/): a single restart still refunds-and-continues; a re-armed restart breaks
after max_retries refunds. The stub grows the redirect seam the follow-up commit
uses so the queued-correction contract is covered by the same test.
2026-09-09 09:51:31 -07:00
yoyodine-industries e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
teknium1 91adf584a4 fix(gateway): every send_multiple_images override returns the aggregate SendResult
#106167 widened the base contract to SendResult but left six native-batch
overrides (Discord, Email, Matrix, Mattermost, Slack, Telegram) returning
None, which (a) meant a media-only reply on those platforms still reported
FAILURE because _record_delivery(None) records nothing, and (b) produced new
`ty` invalid-method-override diagnostics against the widened base (#106192).

Make the contract honest instead of annotating it Optional: each override
now rolls its batches (and any per-image fallback) into one SendResult, so
the turn-outcome accounting works on every platform, not just Signal and
the base loop. The "legacy overrides return None" comment in
_send_image_batch goes away with the legacy.

ty on the 8 touched files: origin/main 241 diagnostics / 20 override,
this branch 241 / 20 — byte-identical diagnostic set; the intermediate
`-> SendResult` head without this commit had 246 / 25.

Refs #106192
2026-09-09 09:45:54 -07:00
Halldrix 78de23053f fix(gateway): report media-only turns as SUCCESS when attachments deliver
Thread record_delivery through _deliver_attachments, _deliver_media_attachments
and _send_image_batch so attachment sends feed the turn outcome tracker.
send_multiple_images (base default and Signal override) now returns SendResult
(success when at least one image/batch was accepted); _send_attachment_batch
returns bool. Legacy native-batch overrides that still return None record
nothing, keeping their previous behavior until migrated.

Fixes #106153
2026-09-09 09:45:54 -07:00
Teknium 00bcef9b5f test(kanban): install the reviewer profile the review-surface fixture hands off to
kanban_request_review now rejects reviewers that are not installed profiles
(#106163); the cross-surface lifecycle test used a bare "reviewer" name with no
profile behind it, which is exactly the phantom the guard exists to catch.
2026-09-09 09:45:13 -07:00
teknium1 e2763baf1c refactor(kanban): route the reviewer guard through _check and tighten its tests
Use the module's `_check`/`_Reject` idiom instead of an inline
`return tool_error(...)` so every kanban_request_review validation
failure renders through the same path, and drop the unreachable
`or "none"` (list_profile_names() always contains "default").

Tests: compare the task's (status, assignee, run) tuple and the event
log before/after instead of the unordered 6-assert block, use the
context-managed kanban_db_connect.connect (the kb.connect alias is a
plugin-compat pointer — scripts/check_compat_pointers.py flagged it),
and reference #106163 in the invariant's docstring.

Salvage note vs #106214 (@gaoanze888): that PR guards the same condition
inside hermes_cli/kanban_db.py::request_review, but the DB primitive is
also the chokepoint for `hermes kanban request-review` and the
dashboard's drag-to-review, both operator surfaces where a non-profile
assignee (external/human review lane) is a documented board shape
(website/docs/user-guide/features/kanban-worker-lanes.md) — and it forced
five unrelated test fixtures to monkeypatch profile_exists to True. The
model-facing tool wrapper is the layer where a typo'd string is a bug,
so the guard lives there.
2026-09-09 09:45:13 -07:00
auroracapital 1d89286b36 fix(kanban): reject phantom worker reviewers
kanban_request_review(reviewer=<name>) reassigned the task to whatever
string the model supplied. A non-profile value (e.g. the literal
"reviewer") parked the card in `review` on an assignee the dispatcher
can never spawn, with no error to the worker — the chain stalled
silently (#106163). Validate the explicit reviewer against installed
profiles before touching the board and return a tool error listing the
installed profiles so the model can self-correct.

Salvage of #97429: the kanban_diagnostics `review_reopened` hunk and its
tests were dropped (main already replaced that loop with
`_latest_event_ts`; the tests targeted the PR's pre-refactor base).
Re-authored from the placeholder identity `regen <regen@local>` to the
PR author's GitHub noreply address (misconfigured local git, not malice).
2026-09-09 09:45:13 -07:00
Teknium e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Teknium 3026f4a993 test(compression): trim batch clarify coverage to two invariants
Drop the sentinel-only batch test: a batch sentinel is already rejected by the
shared _is_clarify_non_response_sentinel list check that the existing sentinel
tests pin, so the case adds no new contract. Also add the contributor email
mapping for the cherry-picked commit so release CI can attribute it.
2026-09-09 09:21:55 -07:00
gaoanze888 8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
Teknium bca7cd0eb0 fix(state): projected compression tip inherits the root's title when the tip is untitled
`hermes peer dm` resolves the target's canonical Bot Chat with
GET /api/sessions?title=Bot%20Chat&include_hidden=1. list_sessions_rich
admits the hidden root via the chain search, then _project_compression_tips
overwrites every surfaced field — title included — with the live tip's. The
title is carried root->tip by the agent AFTER publish_compression_child's
transaction; a rotation cut off in between (crash, closed app, the tip's
title write failing) leaves "Bot Chat" on the ended root and NULL on the tip,
so the projected row carries title=None, the handler's exact-title filter
drops it, the peer POSTs a duplicate and the UNIQUE(title) guard answers
400 "Title already in use" (#106165).

Fix at the projection: fall back to the root's title only when the tip has
none (a titled tip keeps winning). Same COALESCE in the bounded recent-
sessions lister, the other place that projects a lineage onto its tip. This
replaces the handler-level fallback in PR #106365 (a second lookup path
bolted onto _handle_list_sessions with try/except: pass) with a 5-line fix
at the one place the title is lost, so every list consumer sees the name.

Salvage of #106365 by @finn763.
2026-09-09 09:21:46 -07:00
finn763 0c604d7e80 test(gateway): peer dm e2e against a hidden Bot Chat that rotated through compression (#106165)
Regression test from PR #106365 (one of its two tests kept: the full peer-dm
e2e; the listing-shape test asserted the same row and was dropped). The
handler-level fallback that shipped with the test is replaced by a SessionDB
fix in the follow-up commit, so this commit carries only the test.
2026-09-09 09:21:46 -07:00
Teknium 8a50413ba2 refactor(state): drop inline comments duplicating the cooldown-rollback docstring
The docstring already states WHY a vanished session row is tolerated (#106271); the two inline paragraphs restated it per branch.
2026-09-09 09:21:39 -07:00
1052326311 17fe7bc75f test(compression): cover cooldown rollback on a concurrently deleted session (#106271)
Regression tests carried from #106277 (its fix hunk is redundant with #106276's, picked before this).

Salvaged from #106277.
2026-09-09 09:21:39 -07:00
liuhao1024 c0e14092e5 fix(state): also tolerate session deletion between cooldown rollback and read-back
Review follow-up for #106276: the compensating UPDATE can succeed and the
session row can still be retired before the verification read-back runs.
That window raised the same RuntimeError as a genuine restore mismatch.
Accept the absent session (warn + return) exactly like the pre-update
deletion window, keep strict verification for a surviving row, and add a
regression test that deletes the session after the UPDATE commits.

[salvage: test hunk dropped from this pick, see previous commit]
2026-09-09 09:21:39 -07:00
liuhao1024 ccc5cba3d2 fix(state): tolerate a vanished session row in compression cooldown rollback
The exact-row rollback in restore_compression_failure_cooldown_row raised
RuntimeError when its UPDATE hit rowcount == 0, so a session row retired or
expired mid-attempt (e.g. by the maintenance sweep) crashed the turn
dispatcher during the compression summary failure path (#106271). A missing
session row means the cooldown died with it: nothing is left to restore, so
treat rowcount == 0 as a tolerated no-op (warn + return, skipping read-back
verification), mirroring the existing early-return for snapshots taken while
the session already did not exist. Write and verification failures still
propagate.

[salvage: contributor test file dropped from this pick; regression tests are carried from #106277 and trimmed to two invariants]
2026-09-09 09:21:39 -07:00
teknium1 02005cfe20 fix(kanban): promote refuses undone parents instead of a false --force success
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.

The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.

- drop `--force` from `promote` (parser, CLI handler, `promote_task`
  kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
  the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
  fake `ready`; the flag no longer parses

Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
2026-09-09 09:21:29 -07:00
Teknium e8bac35a40 refactor(whatsapp): bridge.js reuses the helpers' getMessageContent
bridge.js carried its own copy of getMessageContent with the single-layer
envelope list, alongside an unused getContextInfo. With envelope peeling now
living in bridge_helpers.js, the copy would drift (a nested envelope around
a pollUpdateMessage is peeled by the helpers but not by the copy), so import
the shared one instead of keeping a second unwrapping list.
2026-09-09 09:21:12 -07:00
Konstantin Khlopkov 7ecd4d14a5 fix(whatsapp): peel nested envelopes and restore immediate payload return
A quoted message can carry its own envelope stack (ephemeral
wrapping viewOnce wrapping the payload), and a reply itself can be
enveloped: both layers now unwrap iteratively (bounded) before quote
extraction. Peeled top-level envelopes return their inner message
immediately, as before — the template/buttons/list branches only
apply to unenveloped messages.
2026-09-09 09:21:12 -07:00
Konstantin Khlopkov 5de769956f fix(whatsapp): unwrap quoted-message envelopes before extracting quote text 2026-09-09 09:21:12 -07:00