Commit Graph

35699 Commits

Author SHA1 Message Date
teknium1 e1114bdcf9 refactor(mcp): one cycle-safe exception walker for every connect-error scan
The salvaged fix gave `_find_missing` and `_flatten_messages` each their own
visited-set loop, next to the one `_is_session_expired_error` already had —
three copies of the same idiom in one module. Collapse them into
`_iter_exception_nodes` (pre-order, left-to-right, each node once, bounded by
`_EXC_TRAVERSAL_MAX_NODES`) and read all three scans off that list. Acyclic
output is byte-identical: the missing-executable search keeps its depth-first
order and a message-less leaf still renders as its class name.

Tests move from the issue-numbered file into `tests/tools/test_mcp_tool_errors.py`
(mirror of the source module): a two-node cycle renders the real messages, and a
missing stdio binary wrapped deeper than the recursion limit with the chain
looping back to the top is still reported as the missing executable. Both are
red on origin/main (RecursionError).

Co-authored-by: Stephan Mongstad <stephan@users.noreply.github.com>
2026-09-15 19:02:39 -07:00
KoNit-K 030d4caa0e fix(mcp): bound nested connection error traversal 2026-09-15 19:02:39 -07:00
teknium1 2588c908e7 fix: recognise fences opened inside list items and blockquotes
Review finding on #112198: _mask_prose_link_destinations matched
_FENCE_LINE against the raw line, so a fence behind a CommonMark
container prefix (`- ```sh`, `1. ```sh`, `> ```sh`, nested) was not
seen and its body was scored as prose with link destinations masked.
Strip the container prefix before fence matching (open and close).
Bundled-skill rescan vs origin/main: 208 skills, 1447 findings on
both, no new/gone findings, no verdict changes.
2026-09-15 19:01:39 -07:00
teknium1 33292affd6 fix(skills): fenced blocks close only on a matching fence; temp-root rm covers //.. and ..;
Follow-up to the two cherry-picked contributor commits.

The picked fence tracker never checked for a closing fence once a block was
open (the closer test sat inside the not-in-code branch), so every prose link
after any code block was scanned verbatim again and the #111254 documentation
link exemption was lost; a fence line carrying an info string was also accepted
as a closer, which handed the scanner back to prose mode mid-block. Rewrite the
loop around CommonMark fence semantics: a block opens on 3+ backticks/tildes
indented at most 3 spaces (backtick info strings may not contain a backtick)
and closes only on a fence with the same marker, at least as long, and nothing
after it; tab- or 4-space-indented lines are code; an unclosed fence stays
code to EOF. plugin_guard inherits the behaviour through scan_file.

The temp-root exemption in destructive_root_rm now also refuses a parent
segment reached through an empty path segment or followed by a shell
separator, which the first cut let through.

Tests trimmed to one invariant per fix: the fence test covers the six code
shapes plus the prose-link-after-fence control that the picked version broke;
the rm test gains the two residual shapes.

Part of #111334
Fixes #112129
Fixes #111335
2026-09-15 19:01:39 -07:00
JulianCruzet 725713d575 fix(skills): track CommonMark fence state in prose-link masking exemption
replace the boolean fence toggle in _mask_prose_link_destinations with
proper (marker_char, opener_length) tracking so a mismatched-markdown-fence
body or an indented code block cannot re-enable prose-masking over live
command lines. closes an exploitable bypass in the community-source
install path; plugin_guard inherits the fix through scan_file.

also tighten is_indented_code to treat any tab indent (single or double)
as code, per CommonMark §4.4.
2026-09-15 19:01:39 -07:00
KoNit-K 4537869dc8 fix(skills): detect temp-root traversal deletes 2026-09-15 19:01:39 -07:00
Zheqing Zeng 8b58562614 docs(curator): align the consolidation prompt with the delete guard
_curator_consolidation_delete_guard refuses background deletes whose
absorbed_into is missing OR empty (fail-closed, #29912): pruning with no
forwarding target belongs to the deterministic staleness pass. The prompt
still instructed the model to pass absorbed_into="" for exactly that case —
a dead-end instruction that burned tool iterations on guaranteed refusals.
Tell the model the rule the guard actually enforces.
2026-09-15 19:01:11 -07:00
Zheqing Zeng 9b003f201f fix(curator): seed shared read-marks store in the LLM consolidation fork
The read-before-write guard requires a skill_view mark from the SAME review
run before any skill_manage write. mark_background_review_skill_read
auto-creates a store when the ContextVar is unset, but tool workers run on
copied contexts, so marks recorded in one worker stayed invisible to the
others: every patch was refused with "current SKILL.md content has not been
loaded in this review turn" even after fresh full reads, and the
consolidation pass burned its iterations retrying a dead-end write.

The background-review fork already seeds a shared store before
run_conversation (agent/background_review.py); do the same in the curator
fork so every copied worker context shares one store.
2026-09-15 19:01:11 -07:00
KoNit-K 551fe883d9 test: device login through a later issuer-bound authorization server
Extend the real-wire device fixture with a `multi_issuer` mode whose
protected-resource metadata lists an issuer-mismatching server before the
valid one, and run the production CLI login through it. Red on main
(`Authorization server metadata issuer mismatch`), green with the scan.

Ported from PR #112068.
2026-09-15 19:00:42 -07:00
teknium1 e133f3f607 fix: device OAuth login scans every advertised authorization server
`hermes mcp login <server> --flow device` took `authorization_servers[0]`
from the protected-resource metadata and failed when that entry was a
browser-only or issuer-inconsistent server, even though a later entry was
the issuer-bound device_code server meant for headless clients (Higgsfield
advertises exactly this shape: a PKCE server first, the device server second).

Discovery now tries each advertised server in order and binds to the first
whose metadata issuer matches its advertised URL and that offers device
authorization. Issuer validation (RFC 8414 / SEP-2468) is unchanged per
server; a single-server resource raises exactly the error it raised before,
and a multi-server resource with no usable entry reports every attempt.

The browser path (`tools/mcp_oauth_manager.py` pre-flight) is deliberately
left on the SDK's own first-entry selection: the SDK's 401-branch discovery
re-selects `authorization_servers[0]` itself, so a divergent pre-flight pick
would only desynchronise the cached metadata from what the SDK authorizes against.
2026-09-15 19:00:42 -07:00
teknium1 60f436b5f6 fix: redact '/'- and '~'-led secrets whose segments cannot be a path
Review finding (agent/redact.py::_should_redact_assignment): a '/'-prefixed
secret with a second '/' (`AWS_SECRET_ACCESS_KEY=/wJalrXUtnFEMIK7MDENG/bPx…`)
still parsed as a multi-segment path and leaked under a strong key; the same
held for a '~'-led value. Apply the opaque bar per segment (16+ chars, no
'.', mixed case and digits) to every '/' or '~' value instead of only to
single-segment ones, so `/home/u/.docker`, `~/.ssh/id_rsa` and
`S.gpg-agent.ssh`-style paths stay readable while base64 secrets mask.
2026-09-15 18:59:55 -07:00
teknium1 34067a7b6d fix: keep $(cmd) substitutions readable under strong-key assignments
Review finding (agent/redact.py::_PATH_OR_VAR_VALUE_RE): anchoring the
path/var exemption regressed `export SSH_AUTH_SOCK=$(gpgconf --list-dirs
agent-ssh-socket)` vs main — the `$(gpgconf` token no longer parsed as a
reference and was masked. Accept a leading `$(` as a reference atom in the
grammar and pin the gpg-agent line in
test_real_path_and_var_references_stay_readable.
2026-09-15 18:59:55 -07:00
teknium1 93a269c026 fix(redact): keep $VAR interpolations inside rc path values readable
The anchored path/variable grammar from the salvaged fix only allowed one
leading $VAR; a strong-key rc line such as SSH_AUTH_SOCK=/run/user/$UID/ssh or
SSH_AUTH_SOCK=$XDG_RUNTIME_DIR/agent.$USER.sock no longer parsed as a reference
and was masked, undoing the readability contract from 979576d938 for exactly
the lines it was written for.

Allow $VAR / ${VAR...} anywhere in the value (and ':' list separators). Crypt
digests still fall through to the credential checks: their '$' fields start
with a digit or carry '=' / ',', which the grammar rejects.
2026-09-15 18:59:55 -07:00
beardthelion 93a0ec705a fix(redact): anchor the path/var exemption so leading-/ and $ secrets still mask
_PATH_OR_VAR_VALUE_RE was an unanchored character class, so re.match made it a
first-character test: any assignment value beginning with '$', '/', or '~'
returned early from _should_redact_assignment, ahead of the strong-key and
opaque-credential checks. AWS secret access keys (~1 in 64 begin with '/') and
argon2/bcrypt digests (always '$'-prefixed) leaked verbatim through
redact_sensitive_text.

Anchor the pattern on both ends so the exemption only fires on a complete
$VAR/${VAR}/~/path//abs/path reference, and require a single-segment absolute
path — indistinguishable by shape from a high-entropy secret — to clear the
opaque-credential bar first. $VAR and ~/ references stay exempt
unconditionally, preserving the rc-readability contract that motivated the
exemption (SSH_AUTH_SOCK=$HOME/.ssh/agent.sock,
DOCKER_AUTH_CONFIG=/home/u/.docker).
2026-09-15 18:59:55 -07:00
teknium1 7c5296ce1c fix(buzz): a clean relay close backs off and publishes retrying like any other disconnect
A relay that accepted, authenticated and subscribed and then closed cleanly
made the read loop return without raising, so _websocket_loop reconnected
immediately with no backoff and never flipped health to "retrying". The read
loop now raises ConnectionError on StopAsyncIteration so the clean close takes
the same backoff + degraded path as an idle or send-side disconnect.
2026-09-15 18:59:25 -07:00
teknium1 f529986abf fix(buzz): a dead socket ends the WebSocket connection from either side and publishes retrying
Follow-up to the cherry-picked watchdog from #112052 (@KoNit-K), finishing the
class the reporter of #112049 laid out:

- `_websocket_loop` runs the read loop and the discovery sweep as sibling
  tasks and ends the connection when EITHER finishes. The discovery sweep
  re-raises `ConnectionClosed` instead of logging it and retrying next tick:
  a send that sees the socket closed is proof the read the loop is parked on
  will never return. That is exactly the traceback the reporter watched for
  22-86 h while inbound stayed silent.
- Health is invalidated while reconnecting: the first disconnect publishes
  `retrying` (`_mark_degraded`) and a successful re-subscribe publishes
  `connected` again. Before, `connect()` wrote "connected" once and nothing
  ever changed it, so `/health/detailed` claimed delivery during the silence.
- The teardown awaits both tasks with `gather(return_exceptions=True)` instead
  of a bare `except (CancelledError, Exception): pass`, which could swallow a
  `disconnect()` cancellation landing mid-teardown.
- Slims the salvaged read-loop hunk: the extra "receive task remained parked"
  warning and the in-loop `_mark_degraded()` are dropped; the reconnect log
  line and the loop-level health flip cover both.

Docs: the Buzz page still described inbound as poll-only and the WebSocket
transport as a future optimization; it now describes the watchdog and the
`retrying` health state.
2026-09-15 18:59:25 -07:00
KoNit-K 8f1e0ec990 fix(buzz): recover from uninterruptible websocket reads 2026-09-15 18:59:25 -07:00
teknium1 8ebc3d420f fix(feishu): warn once when the empty-allowlist default drops group messages
Under multiplex a secondary profile reads FEISHU_GROUP_POLICY from its own
secret scope only (deliberate 0.21.3 isolation), so a profile whose .env
carries no FEISHU_* policy keys falls back to `allowlist` with an empty
FEISHU_ALLOWED_USERS and every human group message is rejected while DMs
keep working. That deny was logged only at DEBUG, making it look like the
events never arrived (#111420).

Keep the scoped read as is — no environ fallthrough. Instead, the first
group drop caused by the untouched allowlist default logs once at WARNING
naming the chat and the keys to set (FEISHU_GROUP_POLICY /
FEISHU_ALLOWED_USERS in the profile's own .env, or group_rules in its
config.yaml). Operator-configured denies (populated allowlist, per-chat
rule, non-allowlist policy) and later drops stay at DEBUG. The predicate
lives in the topical sibling feishu_admission_diagnostics.py.

Docs: the Group Message Policy section now states the per-profile read and
where to put the keys under a multiplexed gateway.

Co-authored-by: bear0328 <bear0328@users.noreply.github.com>
Co-authored-by: NanPan <111261006+poijygfdyy@users.noreply.github.com>
2026-09-15 18:58:59 -07:00
teknium1 92d861c1f8 test(gateway): pin mid-turn completion receipt and live-loop arming
- agent-notify watcher sends the concise receipt only while the launching turn is still
  busy on its adapter; idle session stays receipt-free (the agent reports).
- GatewayRunner.arm_process_watcher schedules the watcher on a live loop only; while the
  gateway is not serving it returns False so the caller keeps the pending fallback.

Both red on origin/main (AttributeError: arm_process_watcher; send awaited 0 times).
2026-09-15 18:58:30 -07:00
teknium1 704f0b1c91 fix(gateway): background completions reach the chat while the launching turn is still running
`terminal(background=true, notify_on_complete=true)` appended its watcher descriptor to
`process_registry.pending_watchers`, which only the post-turn hooks drain. A process that
finished while the turn that launched it was still running (an agent sleep-polling for
hours) had no watcher task at all: the completion_queue entry sat inert, nothing was
injected, and the chat stayed mute until that turn ended (#112033).

- `_register_completion_watcher` arms the watcher on the live gateway loop at registration
  (`GatewayRunner.arm_process_watcher`, via the existing `_gateway_runner_ref` /
  `_gateway_loop` seam that send_message and cron already use); `pending_watchers` stays
  the fallback while the gateway is not serving (checkpoint recovery at startup, shutdown).
- The agent-notify branch of `_run_process_watcher` keeps its design (the agent's next turn
  is the user-facing report) but, when the launching turn is still active at process exit,
  the injection only queues a follow-up — so the concise receipt is sent to the chat right
  away instead of never. The busy check is taken before injection because the injected turn
  itself installs the adapter's session guard.

Live probe (real process, real GatewayRunner loop, fake telegram adapter, busy session):
before — pending_watchers=1 after exit, 0 watcher tasks, 0 injections, 0 receipts;
after — pending_watchers=0, watcher task armed at launch, 1 injection, 1 concise receipt.
Control (idle session): 1 injection, 0 receipts, unchanged.

Slimmer redo of #112038 by @KoNit-K: same two gaps closed, without a second scheduler
registry / loop attribute on ProcessRegistry and GatewayRunner.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:58:30 -07:00
teknium1 9660195388 fix(gateway): /status keeps the default runtime endpoint and context pin when no /model switch owns the route
A route without its own base_url (persisted route / SessionDB row / plain config) no longer
replaces the _resolve_runtime_agent_kwargs read in _resolve_gateway_model_context, so the
custom endpoint is still probed and the matching model.context_length pin survives; only a
/model switch carrying its own endpoint bypasses the default runtime read.
2026-09-15 18:58:02 -07:00
teknium1 de132792f6 fix(gateway): /status resolves the switched-to model's context window; validator names the missing slug
Two diagnostics diverged after a session-only `/model` switch (#111436).

/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).

Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).

Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:58:02 -07:00
teknium1 fcd4778e1b fix(gateway): trim planned-restart notice replay to a single write-once marker pass
Slim the salvaged #112111 mechanism (issue #112109) while keeping its
behaviour: the boot pass and the reconnect hook both run
_replay_pending_planned_restart_notification, which sends to every home
channel still owed an online notice, records delivered targets in
.restart_pending.json and unlinks the marker only when every owed target
(configured home with gateway_restart_notification=true) has been reached.

Dropped from the contributor diff:
- the per-target on_delivered checkpoint callback and pending_targets
  field: delivered targets are written once after the pass. Residual is a
  benign duplicate notice only if the process dies mid send-loop.
- getattr-based lazy lock -> class attribute default, same idiom as
  run_profile_reconcile._reconcile_lock.
- _clear_planned_restart_notification in gateway/run.py: no production
  caller remained; the roundtrip test unlinks the path directly.
- tests trimmed to two invariants: offline-at-boot is replayed once on
  reconnect (with live-at-boot control), and partial delivery is persisted
  so a fresh process does not re-notify and an opted-out home never keeps
  the marker alive.

Live probe (temp HERMES_HOME, Discord home, adapter absent at boot then
reconnected): base consumed the marker with 0 sends; fixed head retains it
and sends the online notice exactly once on reconnect, then clears it.
2026-09-15 18:57:39 -07:00
Steven Tartakovsky 2ae1630bb7 fix(gateway): persist and replay pending planned-restart notices 2026-09-15 18:57:39 -07:00
teknium1 490ee1607a fix(tui): heartbeat ownership follows the gateway's live routing index, not the row source
_notif_gateway_owns_heartbeat decided by the immutable sessions.source column,
so a heartbeat on an ARCHIVED gateway-sourced row (Telegram /reset, idle/daily
auto-reset, compression rotation) was skipped by the Desktop poller and never
registered by the gateway either — restore_heartbeat_watches only claims a key
whose current session_id is that row. The tick belonged to nobody and stayed
due forever, where origin/main's Desktop fired it.

Ownership now uses the same predicate the gateway does: a gateway_routing entry
whose current session_id is this session, with an origin and not suspended
(SessionDB.gateway_routing_entry_for_session; both the session's profile store
and the launch store are consulted so multiplexed and per-profile gateways are
covered). No entry is fail-open, as on main. The check runs after the cheap
is_active/is_due gate so idle sessions never touch the DB per poll.

Probe (SessionStore telegram -> force_new, heartbeat on the archived sid):
before: desktop fired False / gateway watches [] / still due True;
after: desktop fired True / still due False; the current gateway sid is still
left to the gateway (desktop fired False) — the hijack fix stays intact.
2026-09-15 18:57:09 -07:00
teknium1 037771a692 docs: heartbeat and loop ticks stay with the surface that registered them
Users pointed at the Desktop app as a workaround-breaker ("don't keep the shared session open"); state
the ownership rule so the behaviour is discoverable: a chat-registered heartbeat or /loop is fired by the
gateway and replies into the chat even while a TUI / Desktop viewer has the same session open.
2026-09-15 18:57:09 -07:00
teknium1 8a86c56ddb fix(tui): session-owner poller leaves gateway-routed /loop ticks to the gateway
Sibling of the heartbeat fix: a /loop set from a messaging chat carries the gateway's
pinned ``route`` (platform + chat_id), and ``gateway/run_goals.py::_loop_wakeup_fire_one``
already defers route-less CLI/TUI loops to their own schedulers. The TUI/Desktop poller
never returned the favor, so a Desktop viewer of the same session could fire the wakeup
on its own surface and the reply never reached the chat. Mirror the rule: a routed loop
is skipped before the session is claimed, so the tick stays due for the gateway scanner.

Live probe (temp HERMES_HOME, qqbot-routed loop, Desktop viewer of the same session):
before -> loop.ticks=1 awaiting=True (consumed on Desktop); after -> ticks=0, still due.
Desktop-owned control loop still fires.
2026-09-15 18:57:09 -07:00
KoNit-K cf1e727067 fix(gateway): preserve routed heartbeat ownership 2026-09-15 18:57:09 -07:00
teknium1 9a41a8be66 fix(gateway): key Telegram forum handoffs on the group slot the adapter replies on
A /handoff into a Telegram forum supergroup created a topic and bound the CLI session under
telegram🧵<chat>:<topic>, while the Telegram adapter keys every topic reply
telegram:group:<chat>:<topic> — the same restart-orphaning shape as the Slack case. Non-private
Telegram homes now use chat_type group; private-chat DM topics are unchanged.
2026-09-15 18:56:45 -07:00
teknium1 79bf0be53b fix(slack): resolve a cold channel's workspace from the sole authenticated team
scope_id_for_chat only consulted the channel→team map, which is empty right after boot (and
after a reconnect) until an inbound event from that channel arrives. A /handoff into a Slack home
without a stored scope_id (SLACK_HOME_CHANNEL env homes, or config homes never re-set via
/sethome) therefore built a key without the team while every thread reply carries it — the
handed-off thread was still orphaned across a restart (#111896).

When the map has no entry and the channel is not known to be shared across workspaces, fall back
to the single authenticated workspace (filled by auth.test at connect); multi-workspace installs
keep returning None.
2026-09-15 18:56:45 -07:00
teknium1 1f3f45e87b fix(gateway): key Slack handoffs the way Slack thread replies are keyed
`/handoff slack` bound the CLI session under `slack🧵<channel>:<ts>` while every inbound
reply in that thread resolves (via the Slack adapter's source shape) to
`slack:dm|group:<team>:<channel>:<ts>`. After a gateway restart the reply key found no
binding and the gateway opened a fresh empty session, orphaning the handed-off one (#111896).

The handoff destination now mirrors the adapter: chat_type `dm` for a D… home channel, else
`group`, plus the workspace scope_id (home channel provenance, falling back to the adapter's
channel→team map). Channel handoffs were equally affected (`thread` vs `group`), so the fix
covers both, not only DMs. Discord and Telegram destinations are unchanged.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:56:45 -07:00
teknium1 fb975fb098 fix: slim the Slack edit-failure log to one line via the existing payload helper
Follow-up to the salvaged #111938 commit: `_slack_response_payload` already normalizes a
SlackResponse/dict body, so the new `_slack_api_error_code` helper and the two-branch
logger.error were redundant. One log line now always carries `api_error=<code|none>` so an
HTTP 200 + ok=false failure (e.g. message_not_found) is readable without exc_info.

Tests trimmed to one invariant per fix (session key on the response-ready line; API error
code on the edit failure); the `session=unknown` fallback test was a change-detector.
2026-09-15 18:56:20 -07:00
KoNit-K 6214769189 fix(gateway): improve response and Slack error logs 2026-09-15 18:56:20 -07:00
teknium1 6062ad5aee fix(platforms): QR fallback tip names Hermes' own uv when it is not on PATH
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
2026-09-15 18:55:47 -07:00
teknium1 c2b93ae5ac fix(platforms): QR fallback tip uses uv against the running interpreter
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.

Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:55:47 -07:00
Kevin Rajan 505b36b7af fix(platforms): make QR fallback install tip target the active interpreter
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.

Fixes #111695
2026-09-15 18:55:47 -07:00
teknium1 3250020b34 fix(telegram): slim the scheduled typing re-arm and trim its tests to two invariants
Fold the three re-arm helpers (_typing_retrigger_state, _clear_typing_retrigger,
_send_typing_quietly) into _retrigger_typing itself; the semantic change from #111886
is unchanged: schedule sendChatAction as a tracked background task instead of awaiting
it on the send path, one in-flight re-arm per chat, at most one per
typing_retrigger_min_interval_seconds (2s default, matching _keep_typing), and honour
typing_indicator: false, which previously only gated the refresh loop.

Tests: keep the two invariants that are red on origin/main — an intermediate send()
returns while sendChatAction is stalled, and a 20-chunk stream costs one sendChatAction
per chat — and drop the eight change-detector variants.

Co-authored-by: aurel282 <aurelien.gek@gmail.com>
2026-09-15 18:55:19 -07:00
Aurélien Gekiere 7c10c249ce fix(telegram): schedule the post-send typing re-arm instead of awaiting it
`_retrigger_typing` awaited `sendChatAction` inline on the send path, and
streaming re-arms after *every* intermediate send. `sendChatAction` is a
fire-and-forget UI hint whose result nobody reads, but awaiting it ran its
TLS round-trip on the same event loop as the `getUpdates` long-polls.

With several agents streaming concurrently the loop stayed pinned, the
long-polls were never serviced, and they decayed into CLOSE-WAIT while the
adapter still reported `connected` — a gateway that is deaf but healthy, which
`Restart=always` cannot recover because the process never exits.

py-spy put 20 of 20 MainThread samples in `send_typing` -> `send_chat_action`
-> `start_tls`. The handshakes are what cost: with `max_keepalive_connections=4`,
a re-arm per chunk churns the pool so most calls pay a fresh TLS handshake on
the loop thread.

Three changes, all in the re-arm path:

- Schedule the re-arm as a tracked task rather than awaiting it, so a
  round-trip never delays a send or a poll. It joins `_background_tasks`, so
  shutdown cancels it and it cannot outlive the adapter.
- One in-flight re-arm per chat, and at most one per
  `typing_retrigger_min_interval_seconds` (default 2s, `extra` knob; 0 restores
  a call per send). Telegram's bubble lasts ~5s and `_keep_typing` already
  refreshes every 2s, so the re-arm only has to cover the gap left by a landed
  message.
- Honour `typing_indicator: false`. Only `_keep_typing` consulted it, so the
  documented workaround still paid for a `sendChatAction` on every
  intermediate send.

Simulating 200 streamed chunks with a 10ms loop-blocking handshake:
200 `sendChatAction` calls and 2037ms of send-path time before, 1 call and
13ms after.

Fixes #111727

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 18:55:19 -07:00
Aurélien Gekiere 0b265ef4fc chore: map aurel282 contributor email 2026-09-15 18:55:19 -07:00
teknium1 8017dfa4a8 fix(web): drop dead subprocess import from git router 2026-09-15 18:54:51 -07:00
teknium1 b8052c8d4a test(web): trim the gh auth probe tests to two invariants
Keep: overlapping cache misses share one probe (never two `gh` at once);
a refresh issued mid-probe after a login flip gets the fresh answer and
still never overlaps. Dropped: the bounded-helper call-shape detector (the
descendant cleanup is proven by the live wrapper probe, not by asserting
the helper's name), and the two tests of pre-existing behaviour (fresh
cache short-circuit, missing `gh`).
2026-09-15 18:54:51 -07:00
teknium1 ac829e8dee fix(web): gh auth refresh waits out a probe that started before it was asked for
With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.

A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.

Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed:   refresh=true right after login -> authenticated True,  cached True;
         5 concurrent requests -> peak concurrent probes 1

Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
2026-09-15 18:54:51 -07:00
KoNit-K 460768056d fix(web): bound and deduplicate gh auth probes 2026-09-15 18:54:51 -07:00
teknium1 aa03d38612 test(cron): desktop ticker scope test reads the live profile enumerator
_start_desktop_cron_ticker now hands the built-in scheduler a callable that
re-enumerates profiles every tick (so a deleted profile stops being ticked
without an app restart). The sibling scope test in tests/cron still compared
the kwarg to a list snapshot and went red in CI; it now asserts the callable
resolves to the same homes, which is the contract the scheduler consumes.
2026-09-15 18:54:28 -07:00
teknium1 8550f9084e docs(profiles): Desktop cron ticker follows profile create/delete live
Also satisfy padding-line-between-statements on the two salvaged hunks.
2026-09-15 18:54:28 -07:00
teknium1 1220491468 chore(desktop): keep the slot-storm fix to the retry backoff
Drop the extra assertLocalProfileCanStart call added ahead of the slot
queue: the deleted-profile fence already runs at the spawn boundary below
and is not the mechanism behind the #111338 retry storm.
2026-09-15 18:54:28 -07:00
KoNit-K 68e4833134 fix(desktop): bound background profile hydration retries 2026-09-15 18:54:28 -07:00
teknium1 05051691c6 test(desktop): trim the archived-view reload coverage to two invariants
Keep the two discriminating cases (a sessions.changed tick reloads the archived
set while the view is open; a failed refresh retains the last good rows) and
drop the closed-view control, which passes on the unfixed base as well.
2026-09-15 18:54:00 -07:00
KoNit-K dbda53b8be fix(desktop): refresh archived sessions on external changes
Co-authored-by: DavidMetcalfe <80915+DavidMetcalfe@users.noreply.github.com>
2026-09-15 18:54:00 -07:00
teknium1 a40d90e9be test(desktop): trim the voice-live toast tests to two invariants
Seven per-case tests collapsed to two: one per behaviour class (our own
close reasons → copy, server reasons verbatim; mic DOMException → recorder
copy, non-mic failures untouched). Source is byte-identical to the
contributor commit. Dropped: the `closed`/blank-reason case, the
`close_requested` filter (pre-existing behaviour), and the per-name
DOMException matrix — the mapping table lives in `micError`, asserting one
name proves the live path routes through it.
2026-09-15 18:53:32 -07:00