Commit Graph

2309 Commits

Author SHA1 Message Date
KoNit-K 8f1e0ec990 fix(buzz): recover from uninterruptible websocket reads 2026-09-15 18:59:25 -07:00
teknium1 8ebc3d420f fix(feishu): warn once when the empty-allowlist default drops group messages
Under multiplex a secondary profile reads FEISHU_GROUP_POLICY from its own
secret scope only (deliberate 0.21.3 isolation), so a profile whose .env
carries no FEISHU_* policy keys falls back to `allowlist` with an empty
FEISHU_ALLOWED_USERS and every human group message is rejected while DMs
keep working. That deny was logged only at DEBUG, making it look like the
events never arrived (#111420).

Keep the scoped read as is — no environ fallthrough. Instead, the first
group drop caused by the untouched allowlist default logs once at WARNING
naming the chat and the keys to set (FEISHU_GROUP_POLICY /
FEISHU_ALLOWED_USERS in the profile's own .env, or group_rules in its
config.yaml). Operator-configured denies (populated allowlist, per-chat
rule, non-allowlist policy) and later drops stay at DEBUG. The predicate
lives in the topical sibling feishu_admission_diagnostics.py.

Docs: the Group Message Policy section now states the per-profile read and
where to put the keys under a multiplexed gateway.

Co-authored-by: bear0328 <bear0328@users.noreply.github.com>
Co-authored-by: NanPan <111261006+poijygfdyy@users.noreply.github.com>
2026-09-15 18:58:59 -07:00
teknium1 79bf0be53b fix(slack): resolve a cold channel's workspace from the sole authenticated team
scope_id_for_chat only consulted the channel→team map, which is empty right after boot (and
after a reconnect) until an inbound event from that channel arrives. A /handoff into a Slack home
without a stored scope_id (SLACK_HOME_CHANNEL env homes, or config homes never re-set via
/sethome) therefore built a key without the team while every thread reply carries it — the
handed-off thread was still orphaned across a restart (#111896).

When the map has no entry and the channel is not known to be shared across workspaces, fall back
to the single authenticated workspace (filled by auth.test at connect); multi-workspace installs
keep returning None.
2026-09-15 18:56:45 -07:00
teknium1 fb975fb098 fix: slim the Slack edit-failure log to one line via the existing payload helper
Follow-up to the salvaged #111938 commit: `_slack_response_payload` already normalizes a
SlackResponse/dict body, so the new `_slack_api_error_code` helper and the two-branch
logger.error were redundant. One log line now always carries `api_error=<code|none>` so an
HTTP 200 + ok=false failure (e.g. message_not_found) is readable without exc_info.

Tests trimmed to one invariant per fix (session key on the response-ready line; API error
code on the edit failure); the `session=unknown` fallback test was a change-detector.
2026-09-15 18:56:20 -07:00
KoNit-K 6214769189 fix(gateway): improve response and Slack error logs 2026-09-15 18:56:20 -07:00
teknium1 6062ad5aee fix(platforms): QR fallback tip names Hermes' own uv when it is not on PATH
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
2026-09-15 18:55:47 -07:00
teknium1 c2b93ae5ac fix(platforms): QR fallback tip uses uv against the running interpreter
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.

Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:55:47 -07:00
Kevin Rajan 505b36b7af fix(platforms): make QR fallback install tip target the active interpreter
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.

Fixes #111695
2026-09-15 18:55:47 -07:00
teknium1 3250020b34 fix(telegram): slim the scheduled typing re-arm and trim its tests to two invariants
Fold the three re-arm helpers (_typing_retrigger_state, _clear_typing_retrigger,
_send_typing_quietly) into _retrigger_typing itself; the semantic change from #111886
is unchanged: schedule sendChatAction as a tracked background task instead of awaiting
it on the send path, one in-flight re-arm per chat, at most one per
typing_retrigger_min_interval_seconds (2s default, matching _keep_typing), and honour
typing_indicator: false, which previously only gated the refresh loop.

Tests: keep the two invariants that are red on origin/main — an intermediate send()
returns while sendChatAction is stalled, and a 20-chunk stream costs one sendChatAction
per chat — and drop the eight change-detector variants.

Co-authored-by: aurel282 <aurelien.gek@gmail.com>
2026-09-15 18:55:19 -07:00
Aurélien Gekiere 7c10c249ce fix(telegram): schedule the post-send typing re-arm instead of awaiting it
`_retrigger_typing` awaited `sendChatAction` inline on the send path, and
streaming re-arms after *every* intermediate send. `sendChatAction` is a
fire-and-forget UI hint whose result nobody reads, but awaiting it ran its
TLS round-trip on the same event loop as the `getUpdates` long-polls.

With several agents streaming concurrently the loop stayed pinned, the
long-polls were never serviced, and they decayed into CLOSE-WAIT while the
adapter still reported `connected` — a gateway that is deaf but healthy, which
`Restart=always` cannot recover because the process never exits.

py-spy put 20 of 20 MainThread samples in `send_typing` -> `send_chat_action`
-> `start_tls`. The handshakes are what cost: with `max_keepalive_connections=4`,
a re-arm per chunk churns the pool so most calls pay a fresh TLS handshake on
the loop thread.

Three changes, all in the re-arm path:

- Schedule the re-arm as a tracked task rather than awaiting it, so a
  round-trip never delays a send or a poll. It joins `_background_tasks`, so
  shutdown cancels it and it cannot outlive the adapter.
- One in-flight re-arm per chat, and at most one per
  `typing_retrigger_min_interval_seconds` (default 2s, `extra` knob; 0 restores
  a call per send). Telegram's bubble lasts ~5s and `_keep_typing` already
  refreshes every 2s, so the re-arm only has to cover the gap left by a landed
  message.
- Honour `typing_indicator: false`. Only `_keep_typing` consulted it, so the
  documented workaround still paid for a `sendChatAction` on every
  intermediate send.

Simulating 200 streamed chunks with a 10ms loop-blocking handshake:
200 `sendChatAction` calls and 2037ms of send-path time before, 1 call and
13ms after.

Fixes #111727

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 18:55:19 -07:00
teknium1 e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 4b00842536 fix(web): keyless MCP honours SSE CR terminators and declared charsets
Review follow-up on the keyless Exa UTF-8 fix:

- `_parse_mcp_body` split SSE frames on `\n` only, so a body using bare CR
  line terminators (permitted by the SSE spec, and parsed fine before via
  `splitlines()`) failed with "Unrecognized MCP response shape". Split on
  the SSE terminators CRLF / CR / LF instead.
- `mcp_call` decoded the body as UTF-8 unconditionally, ignoring a charset
  the server did declare. Decode with the declared charset when the
  Content-Type carries one and fall back to UTF-8 only when it is absent.
- The HTTP >= 400 branch and `_keenable_request` still surfaced
  `response.text`, so a non-ASCII error body from a charset-less text/*
  response reached the user as ISO-8859-1 mojibake. Both now go through the
  same decode helper as the success path.
2026-09-15 18:43:28 -07:00
teknium1 28cdd72815 fix(web): keyless Exa decodes its SSE body as UTF-8 so CJK searches work
Exa's MCP endpoint answers `text/event-stream` without a charset, so
`requests` decoded `.text` as ISO-8859-1: every non-ASCII character came back
as mojibake, and for CJK results the UTF-8 continuation byte 0x85 became
U+0085, which `str.splitlines()` treats as a line break — the `data:` JSON
line was cut in two and the perfectly valid result surfaced as "Unrecognized
MCP response shape", after which the ring fell through to keyless Firecrawl
(403). Decode the bytes as UTF-8 (JSON-RPC and SSE are UTF-8 by spec) and
split SSE frames on newlines only. A parsed envelope that really carries no
text is now reported as "no text content" instead of an unrecognized shape.
2026-09-15 18:43:28 -07:00
KoNit-K ffd02f37e8 fix(web): cache extracts by returned URL 2026-09-15 18:42:38 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1 f7ea39481a fix(kanban): dashboard estimate calls declare a relay-affinity key too
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.

Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
2026-09-15 18:33:43 -07:00
teknium1 b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
KoNit-K 4fe3d288eb fix(providers): send DeepInfra reasoning effort 2026-09-15 18:22:40 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 9c311038fa fix(discord): slash registration and /skill refresh scan the catalog off the event loop
connect() called _register_slash_commands inline on every (re)connect, and
/reload-skills called refresh_skill_group inline; both run
discord_skill_commands_by_category, the same per-skill path-resolution walk
the Telegram menu paid in #110707. On a 1.5k-skill install that holds the
loop past the liveness watchdog. Registration now hops through
asyncio.to_thread from connect(); refresh_skill_group is a coroutine that
hops the rescan the same way (the reload handler already awaits an awaitable
result). Contextvar-scoped profile overrides propagate through to_thread.

One invariant test: the loop keeps ticking while the scan blocks, from both
sites. Red on the PR head, green here.

Review finding: Discord _register_slash_commands/refresh_skill_group ran the skill catalog disk scan synchronously on the event loop.
2026-09-15 06:35:04 -07:00
teknium1 ab782a7685 fix(gateway): skill-slash fallthrough and the Telegram inline picker run off the event loop
The gateway's idle-command path resolved skill slash commands inline on the
loop: a cold skill scan, skill file loads and the unavailable-skill rglob over
every skills dir. On a 1.5k-skill install that held the loop ~2 minutes, the
loop-liveness watchdog fired and the gateway exited mid-session (#111091).
_hm_skill_slash_rewrite now runs through _run_in_executor_with_context so the
profile contextvars the scan is scoped to survive the hop. Known commands still
short-circuit before any I/O (previous commit).

Same class in the Telegram inline picker: build_inline_results rebuilds the
command/skill catalog per keystroke via _collect_gateway_skill_entries, the
same path-resolution pass #110707 traced in the command menu.

One invariant test: a known command triggers no scan; an unknown command's scan
runs while the loop keeps ticking. Red on origin/main and on the reorder-only
tree, green here.
2026-09-15 06:35:04 -07:00
teknium1 660d4b8d87 fix(telegram): hop only the menu build off the loop; one invariant test covers both menu sites
Follow-up to the salvaged #110716 commit. Drops the module-level
_build_telegram_command_menu wrapper (telegram_menu_max_commands is a config
read and stays on the loop; asyncio.to_thread takes kwargs directly), and
applies the same hop to _ensure_forum_commands, which rebuilds the menu on the
inbound-message path for every forum chat until registration succeeds — the
second live fire site named in #110707.

Replaces the contributor's test with one that proves the invariant for both
sites: the loop keeps ticking while telegram_menu_commands blocks. Red on
origin/main, green here.
2026-09-15 06:35:04 -07:00
KoNit-K 0a5846d37f fix(telegram): build command menu off event loop 2026-09-15 06:35:04 -07:00
teknium1 89ef145254 fix(kanban): dashboard link reports the gate; docs; trim to two invariant tests
The dashboard's POST /links is the fourth writer of link_tasks (CLI, tool,
dashboard, plus the graph builder); return the same ``gated`` flag so every
surface that can create the deadlock can see it. Document the
``dependency_wait`` payload the link path emits and the delegation rule the
reporter derived (never link a support card under the card it unblocks).

Drops the CLI output test (a change-detector on prose); the two DB-level
invariants (event emitted on demotion / none for a done parent) stay.
2026-09-15 06:25:42 -07:00
teknium1 060e2a4286 fix: exempt the Matrix quote block from the mention strip only on real replies
The reply-pill fix split the body on a leading "> " so the strip would not
rewrite the "> <@bot:srv>" pill. That keyed the exemption on the body shape,
not on the relation: a hand-typed blockquote in a plain (non-reply) message
that mentions the bot inside the quote reached the agent with the raw
"@hermes:example.org" text, which main used to strip. Split around the quote
only when m.relates_to carries m.in_reply_to; otherwise strip the whole body.

Review finding: quote-block exemption keyed on body.startswith('> ') instead of m.in_reply_to.
2026-09-15 06:19:19 -07:00
KeyArgo 0813a2276c fix(gateway): keep the Matrix reply fallback pill out of the mention strip
The Matrix adapter strips the bot's mention from the inbound body before
_extract_reply_context parses the inline reply fallback, and that strip is a
blind whole-body replace. A reply to the bot names the bot in the fallback
pill ("> <@bot:server> quoted"), which is exactly what makes the message
count as a mention under the default MATRIX_REQUIRE_MENTION=true -- so the
strip runs on every reply-to-the-bot and rewrites the pill to "> <>", after
which the pill regex no longer matches. reply_to_author_id is lost,
reply_to_text becomes the mangled "<> quoted" remnant, and the prompt
renders "[Replying to: "<> ..."]" instead of "[Replying to your previous
message: ...]".

Split the body into (quote block, reply text) and strip the mention from the
reply text only. The mention gate still sees the raw body, so a reply to the
bot keeps waking the bot; the visible reply text and the quote-block strip
are unchanged.

Fixes #111233
2026-09-15 06:19:19 -07:00
liuhao1024 a019976430 test(slack): address review nits on the subtype allowlist
Drop the unreachable handle_message assertion in the drop tests
(_prefilter_inbound never calls it) and state the deliberate
file_comment drop decision in the allowlist comment.
2026-09-15 04:40:49 -07:00
teo-nex 90aa5e3511 fix(slack): preserve allowed bot posts and canvas mentions 2026-09-15 04:40:49 -07:00
liuhao1024 cac288a0c3 fix(slack): gate inbound turns on a conversational-subtype allowlist
Housekeeping subtypes (channel_join/leave/topic/name/purpose,
convert_to_private/public, pins, deletions) are not a person speaking,
yet _prefilter_inbound only rejected message_changed/message_deleted,
so each of them started a full agent turn in free-response channels.
Replace the denylist with an allowlist: a message passes when subtype
is absent, file_share, thread_broadcast or me_message; everything else
is dropped. Fixes #110778.
2026-09-15 04:40:49 -07:00
teknium1 6a22abe5ba fix: retire clarify cards on an explicit no-answer signal, not the '[' prefix
`_clarify_callback_sync` decided "no answer arrived" by testing whether the
response text starts with '[' (the shape of the timeout / undeliverable
sentinels). A real answer can start with '[' too — a "[A] staging" choice
label picked by number, or "[urgent] ..." free text after Other — so the
clarify resolved and the agent got the answer, yet the Slack card was
rewritten to "This prompt expired" and typing was never re-armed.
`_clarify_send_then_wait` now returns `(response, answered)` and the runner
branches on that flag only.

The Slack click handler popped the retire entry as soon as Other was
clicked, but Other is not terminal: the clarify stays pending for typed
text, so a later timeout or /new reset found nothing to retire and the card
stayed stuck on "Awaiting typed answer". The entry is now popped only on a
terminal outcome (a choice click, or Other on an already-dead entry).

A typed answer to a native card (numeric pick, or text after Other) never
reaches the click handler, so the card kept its buttons forever; the
TEXT_RESOLVED intercept now retires it with the answer.

Review finding: '[' prefix mistaken for the timeout sentinel; Other click dropped the retire entry; typed answers never rewrote the card.
2026-09-15 04:40:04 -07:00
teknium1 71cac9426d fix(gateway): retire native clarify cards on timeout, reset and prose cancel
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:

- TurnRunner._clarify_callback_sync: when the bounded wait returns a
  sentinel (timeout, /new or run-end clear_session), schedule the retire
  with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
  the prose is routed as a follow-up (#111019). Lookup is on the adapter
  class so MagicMock doubles cannot fabricate the method; no platform ==
  SLACK special-case.

The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.

Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
2026-09-15 04:40:04 -07:00
KoNit-K a41187e187 fix(gateway): retire Slack clarify cards on prose cancellation 2026-09-15 04:40:04 -07:00
teknium1 1847ad2883 fix(telegram): trim the liveness heartbeat, fix the polling error-callback log lines
- Drop the 900s "inbound liveness" INFO line from the salvage: a periodic
  log heartbeat is a feature with a separate scoping call (#111211 item 3);
  the existing stall watchdog already escalates when getUpdates stops.
- The polling error_callback interpolated the raw exception and left the
  redaction call inside the format string, so its lines read
  "Telegram network _redact_telegram_error_text(error), scheduling
  reconnect: ..." and leaked unredacted text. Redact for real.
- Tests: keep two invariants (recovered wording after network errors;
  clean bootstrap stays "confirmed healthy") and the transport test.
2026-09-15 04:39:16 -07:00
liuhao1024 ff307aea5e fix(telegram): name empty-string transport errors in adapter log lines
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.

Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
2026-09-15 04:39:16 -07:00
fangliquan e1e943a9df fix(telegram): expose polling transport recovery 2026-09-15 04:39:16 -07:00
teknium1 2298dc8122 fix: propagate caller contextvars into the Feishu adapter-owned executor
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.

Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
2026-09-15 04:38:26 -07:00
teknium1 b447554ce5 fix(feishu): thread-reply lookup uses the adapter-owned pool too
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
2026-09-15 04:38:26 -07:00
JackJin 5fbef868a9 fix(gateway): keep Feishu websocket off default executor 2026-09-15 04:38:26 -07:00
liuhao1024 fa2c72746c fix(feishu): run the inbound dedup flush on the adapter-owned pool
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.

#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.
2026-09-15 04:38:26 -07:00
fangliquan 992b517fdd fix(photon): preserve Unicode NDJSON separators 2026-09-15 04:34:13 -07:00
wang2 fa12d7556c fix(telegram): allow opt-in CJK rich messages 2026-09-15 04:24:53 -07:00
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1 624777617c fix(discord): filter obfuscated channels on the explicit-id backfill path too
Review finding on #90154: only the wildcard ("*") branch of
_iter_missed_message_backfill_candidates skipped obfuscated channels. A
channel the bot previously talked in that later lost VIEW_CHANNEL still
entered the explicit-id branch and produced a failing history read on
every startup. Apply the predicate once, after both branches build the
candidate list, and import it at module level like the other helpers.

(During the rebase the predicate was also tightened, resolved in the original
commit: is_discord_channel_obfuscated catches AttributeError (the precise
expected failure) instead of a bare Exception so a genuine attribute bug
is not hidden behind the name fallback, and its docstring states the
deliberate bias of the sentinel-name fallback (a visible channel literally
named ___hidden___ is also skipped on discord.py builds without the flag).)
2026-09-15 03:50:10 -07:00
Teknium 1f34742552 fix(discord): skip obfuscated channels in directory and backfill enumeration
Discord's Channel Obfuscation change (announced Aug 12 2026, HTTP
enforcement Nov 16 2026) dispatches channels the bot lacks VIEW_CHANNEL
on with name "___hidden___", flag 1 << 17 (CHANNEL_OBFUSCATED), and
nulled fields. Without filtering, the channel directory lists phantom
"___hidden___" entries the agent can never post to, and wildcard
missed-message backfill wastes history reads on channels that always
403.

Adds is_discord_channel_obfuscated() to gateway/platforms/helpers.py
(checks the flag bit plus the sentinel name for discord.py builds that
don't expose the new flag) and applies it at both enumeration sites.
2026-09-15 03:50:10 -07:00
teknium1 7abe9502ee fix(kanban): worker liveness and kills require the spawn-time start fingerprint
tasks.worker_pid outlives a reboot; afterwards the number can belong to any
process. _pid_alive answered from bare existence, so reclaim_stale_claims kept
extending the claim of a "live" stranger (stuck running task), and
enforce_max_runtime / _terminate_reclaimed_worker SIGTERM'd then SIGKILL'd it.

_set_worker_pid now records gateway.status.get_process_start_time(pid) as
tasks.worker_started_at (additive column, NULL on legacy rows). _worker_alive
(pid, started_at) is the liveness check every reader uses (reclaim, defer,
reconcile, crash sweep, max-runtime, archive, reopen invalidation); a live pid
whose fingerprint disagrees is a recycled PID: treated as dead, never
signalled (termination reports pid_recycled). Legacy rows without a
fingerprint keep the existence answer until their next spawn.

Row hermes_cli/kanban_db_dispatch.py:458 (lane4_high) confirmed by tracing:
the SIGKILL at :461 was gated only on _pid_alive.
2026-09-15 03:47:15 -07:00
teknium1 8b181940b4 fix(multiplex): background threads and teardown paths carry the turn's profile scope
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:

- agent/title_generator.py: the auto-title thread read
  auxiliary.title_generation (model, language, provider key) from the default
  profile's config and billed the default's key for a secondary's session.
  Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
  (ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
  session.close pool RPC, superseded_by_resume or compute_host flush - none
  carries a scope, yet on_session_end / commit_memory_session / agent.close ->
  shutdown_memory_provider read the provider's config + credentials at call
  time. Under multiplex they failed closed (tail never committed, #110622
  class); on the Desktop backend a secondary's transcript went to the launch
  profile's memory tenant. _finalize_session and _teardown_session now bind
  _session_profile_runtime_scope(session) around those blocks, which covers
  every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
  SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
  context onto the loop task, so _dispatch_message and everything under it
  (attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
  _load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
  the launch profile. connect() captures its scope; _on_pubsub_message and
  _submit_on_loop run under a per-callback copy of it.

spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
2026-09-15 03:47:15 -07:00
teknium1 d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
Sudhir Patil ec11359b44 fix(discord): clear fatal status on successful reconnect
DiscordAdapter.connect() set self._running = True directly instead of
calling self._mark_connected(), unlike every other platform adapter
(Telegram, WeCom, Matrix, Feishu, Google Chat, IRC, LINE, Mattermost,
ntfy, photon, raft, simplex, a2a, buzz, dingtalk, whatsapp).

_mark_connected() clears _fatal_error_code/_fatal_error_message/
_fatal_error_retryable and rewrites the runtime status file as
"connected". Bypassing it meant a transient connect failure (e.g. a
one-off DNS blip: "Cannot connect to host discord.com:443 ssl:default
[Temporary failure in name resolution]") left the platform reported as
permanently fatal in gateway_state.json / the dashboard, even after the
adapter successfully reconnected and was actively serving messages for
hours.

Reproduced under gateway.multiplex_profiles: true with a secondary
profile's Discord bot (sarathi:discord) — the bot reconnected
repeatedly ("Connected as ..." logged many times over 18+ hours) while
the dashboard kept showing the original fatal error the entire time.

Fixes #102554.

Added a regression test asserting connect() clears a previously
recorded fatal error.
2026-09-15 03:44:36 -07:00
teknium1 fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
Teknium 5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00