httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.
Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.
Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
_connect_websocket now submits the SDK thread to _get_sdk_executor() instead
of the loop default executor; the SimpleNamespace adapter stub in the
profile-scope test predates that contract and lacked the method.
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
The regression test swaps in a torn-down default executor but never
restores the pristine one, leaking the poisoned executor to anything
sharing the loop. Save/restore via the private slot because
set_default_executor() rejects None, which is exactly the pristine
(lazy-creation) state. Also apply ruff format to the file.
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.
#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.
The bridge wrote GATEWAY_ALLOW_ALL_USERS into os.environ only when unset and
never cleared it. In-process restart paths (gateway restart watcher, dashboard
profile actions) copy os.environ into the child, so a config.yaml grant became
a sticky env var: flipping allow_all_users to false and restarting left the
gateway OPEN. The bridge now tracks its own write (module flag), overwrites or
clears it on reload, exports only a truthy grant (presence-based readers such
as the Telegram intake prefilter treated "false" as configured auth), and the
two restart env builders drop the bridge-owned value so the child re-derives
the posture from its own config.yaml. Under multiplex_profiles the DEFAULT
profile's events are authorized inside its secret scope, where gate readers
never fall to os.environ; the bridged grant is now seeded into that profile's
scope mapping only (secondaries never inherit it).
Review finding: bridged GATEWAY_ALLOW_ALL_USERS survives restart and overrides a flipped config.yaml; inert for the default profile under multiplex; "false" exported as configured auth.
`gateway.allow_all_users: true` (and the top-level spelling) in config.yaml
was a silent no-op: `_TOPLEVEL_BRIDGE` never forwarded it, GatewayConfig has
no field, and every allow-all reader (authz mixin default-deny branch,
startup access check, own-policy adapters, Discord/Matrix/Email plugin
gates) consults the GATEWAY_ALLOW_ALL_USERS env var only.
Bridge the YAML key into that env var in `bridge_core_env_settings`, the
one seam every reader already shares, instead of threading a new config
attribute through ten readers:
- first-writer-wins: an explicit env var beats YAML (matches every other
{PLATFORM}_* gate);
- skipped inside a multiplexed secondary profile's scope (#80099 class):
the secondary's config.yaml must not become the default profile's policy,
and the isolation test now asserts GATEWAY_ALLOW_ALL_USERS stays unset;
- a startup warning names config.yaml as the grant source, because the key
was inert until now and a forgotten `true` flips the posture to open.
Tests: both spellings authorize a stranger through `_is_user_authorized`;
`false`, absent key, and env=false over YAML=true all stay denied.
Docs: security guide, env-var reference, gateway internals.
Fixes#110690
Re-applied from #23955 (37de4cbd64) in current main's shape: the top-level
spelling joins _EXTRA_KNOWN_ROOT_KEYS (main derives known roots from
DEFAULT_CONFIG plus this set) instead of a DEFAULT_CONFIG root entry.
Drop the change-detector asserting the default attribute value; the two remaining tests
pin the invariants (a delta-fed consumer keeps the WeCom ack-race warning, an interim-only
consumer does not).
The cron/webhook lane compared the stripped response and its first/last
line against the marker set without stripping edge punctuation, while
the interactive predicate did. A Chinese lane naturally renders the
sentinel as 【静默】 or 静默。, which chat suppressed but cron still
delivered — the very lane #110935 was filed against. Route the three
autonomous candidates through is_intentional_silence_response so both
lanes share one canonical form; the bracketed-prefix rule and the
embedded-in-prose rejection are unchanged.
Review finding: is_autonomous_silence_response never de-punctuated, so 【静默】 / 静默。 passed cron/webhook while is_intentional_silence_response suppressed them.
Follow-up to the salvaged #110940 commit:
- `cron/scheduler_prompt.py::_CRON_HINT` tells the model the sentinel is a literal
ASCII control token that must never be translated or rephrased — the prompt-side
half of the fix, so a lane answering in any other language is steered to the
canonical token instead of relying on the filter knowing that language.
- Docs: the supported-token list ('Intentional Silence Tokens') gains the zh-Hans
forms in the English page and the zh-Hans mirror gets the section it lacked.
- Tests trimmed to two invariants (translated forms match in every shape the English
ones do; prose that mentions the word is still delivered), proven red on origin/main.
A model lane that does not think in English answers the cron instruction
("respond with exactly [SILENT]") by translating the sentinel rather than
dropping it. LIVE_GATEWAY_SILENT_MARKERS is English-only, so the whole control
token is delivered to the user as content. Carry the zh-Hans forms (the only
non-English locale this project documents) and derive the autonomous lane's
bracketed-prefix rule from the set so the two rules cannot drift.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011oAAQ4T859Y8fWNXRoiNEB
The archived-holder carve-out in _set_session_title lets the next Bot open
claim the "Bot Chat" title from an archived canonical row. That is correct
for a deliberate sidebar archive, but archive_stale_sessions (the opt-in
sessions.auto_archive sweep) could also archive an idle hidden Bot Chat with
end_reason NULL, which unarchive_recoverable_session refuses; the next Bot
click then stripped the old row's title, retiring the bot's whole history
on an idle timer with no way back.
Skip the hidden canonical Bot Chat in the sweep SELECT using the same
hidden + exact-title predicate set_session_pinned already uses to protect
it, so only an explicit archive retires a Bot Chat. Reword the bot-mode doc
so it no longer promises the retired (still hidden) chat is reachable from
the archive view.
Review finding: auto_archive sweep could irreversibly retire the canonical Bot Chat; docs over-promised archive-view reachability.
The API server's _resolve_media_to_data_urls ran MEDIA_TAG_CLEANUP_RE over the
raw final_response, bypassing the sentinel seam this branch added in
gateway/platforms/base.py. A provider that leaked its end-of-sequence token
glued to the last MEDIA path (MEDIA:/x.png<|eos|>) therefore still produced no
inline image and returned the raw tag to the HTTP client (non-stream and SSE
chat completions) - the exact symptom fixed for the chat platforms. The scan
now stops at _terminal_sentinel_start and drops the control token only when a
tag actually resolved, so unrelated responses stay byte-identical.
_terminal_sentinel_start also recognises a run of repeated exact tokens
(MEDIA:/x.png<|eos|><|eos|>), so a doubled leak no longer hides the attachment
either; mid-text or non-exact sentinels are still not a media boundary.
Review finding: api_server sibling MEDIA surface bypassed _mask_media_scan_text; doubled terminal sentinel still leaked
A provider that leaks its end-of-sequence control token glued to the last
MEDIA path (MEDIA:/x.png<|eos|>) made BasePlatformAdapter.extract_media
return no media: neither MEDIA_TAG_CLEANUP_RE nor MEDIA_EXTENSIONLESS_TAG_RE
accepts '<' as a path terminator (deliberately - widening the delimiter set
reopens the glue classes #68773/#88038), so the file was silently dropped
and the raw tag leaked into the chat.
Recognise the one exact token at the shared normalization seam instead:
_mask_media_scan_text blanks a terminal <|eos|> (offset-preserving) so the
glued tag ends on whitespace for every scan - extraction, display strip and
stream cleanup - across the known-extension, unknown-extension and
extension-less branches; _deliverable_tag_spans deletes the sentinel with
the tag it terminated. Not the exact token, not terminal, or inside a
protected code/JSON span: byte-identical to before.
Fixes#111046
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
#111809 accepted any *.nousresearch.com https host once the operator's Portal override
pointed at a non-production Portal. That let a network-provenance value — the Portal's
refresh response — pick any Nous-owned host as the bearer recipient, including hosts that
are not inference gateways. Owning the DNS suffix is not the same as being an authorized
recipient, and the validator's threat model (an injected refresh response) is exactly the
case a suffix rule fails to bound.
The recipient is now the operator's own NOUS_INFERENCE_BASE_URL: a non-production host
returned by the Portal is accepted exactly when it equals that override's host, otherwise
the strict production set stands. The Portal override grants nothing by itself. What the
operator gains over plain use of the override is that the Portal's value is then persisted
and used for the pricing scope and proxy instead of being healed to production, and the
per-turn "refusing inference URL host" warning stops. No environment is named in code.
Raised on #111809 review. Tests: recipient match accepted, unrelated Nous host refused, no
override refused, Portal override alone grants nothing, the match follows the profile scope
under multiplexing; the widening cases are red on main. Docs row for NOUS_INFERENCE_BASE_URL.
Co-authored-by: Ben Barclay <ben@nousresearch.com>
get_secret() already reads os.environ in a single-profile process; it raises
UnscopedSecretError only when multi-profile hosting is active and the call has no profile
scope. _nous_inference_env_override and _nous_portal_env_override caught that and read the
ambient env anyway, which is the launch profile's value — so a routed call that lost its
scope could POST a secondary's refresh token to the launch profile's Portal, or send its
inference bearer to the launch profile's host. Since #111620 the launch profile itself
carries a frozen-env scope while hosting, so the exception never means "single-profile
CLI"; it means the caller has no authority over that value. Both readers now go through
_scoped_operator_override, which returns None in that state — the override is absent, the
stored/default routing applies, nothing raises (pricing, runtime_provider and the guest
path rely on these never raising).
Raised on #111809 review. One parametrized invariant test over both readers: environ when
unscoped single-profile, the scope's value under hosting, None on a scoped miss, None with
no scope; the no-scope leg is red on main.
Main routes sends, finalize-edits and rich drafts through one _rich_content_ok
predicate, so the draft-specific opt-in test duplicated the send test's
invariant. Keep the send (BMP + astral CJK) and finalize-edit cases.
The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.
Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.
Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
Drop the wait-timeout call-shape assertion; the invariants are that the
session records the real exit code and that a failed reap does not publish a
false completion.
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.
Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.
- LICENSE: header naming the upstream repo, the pinned upstream commit
(b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
shape mono-color and pr-lens use — and the adaptation-notes blockquote now
links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
with website/scripts/generate-skill-docs.py (scoped to this skill).
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.
Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
(square aspect, main-prompt constraints mode — no negative_prompt
parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
proposal-round skip for pre-authorized batches, dimensions-reporting
rule when the backend returns only a URL, limbless-subject note
Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.
Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).
Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
The Memory Wiki from #89940 (salvage of #31244 by @Araja119) ships as a
standalone plugin in NousResearch/hermes-memory-wiki instead of landing in
core; this entry lists it in the catalog at the repo's reviewed commit.
The pin is the repo's first commit (younger than the two-week maturity
window) and needs the repo owner's explicit waiver to merge.
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.
Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.
Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
A detached or service-managed gateway can have a blocked stderr. The watchdog exit escort then hard-exits after ten seconds before the file-based traceback is reached, leaving only the metadata record. Write the durable dump first and pin the ordering with a regression test.
Construction-time maybe_auto_archive / maybe_auto_prune_and_vacuum ran
synchronously with no report_startup_progress lease. A multi-minute
VACUUM of a large state.db accrues near-zero CPU, so the startup
watchdog misread it as a parked deadlock and killed the attempt with
exit 75, live-locking gateway restarts. Renew the lease per long step
(prune, orphan sweep, VACUUM, archive) since leases clamp at 900 s,
plus one lease around the gateway maintenance block.
Fixes#111092
carry_unadmitted_user_message appends the interrupted turn's user row to the
early result's in-memory history only; it is never persisted because that turn
never owned the lease. When the follow-up turn also has to wait for the lease
(the common case: the other process that made the first turn wait is usually
still busy), admit_durable_turn_lease reloads conversation_history from the DB
after admission and replaced the caller's list wholesale, so the tagged row was
dropped from both the model input and state.db. Re-append the tagged rows that
have no _row_id after the reload so the follow-up turn sees and flushes them.
Review finding: waited-reload branch of admit_durable_turn_lease discarded the
_persist_after_admission_interrupt row carried from the aborted turn.
Move the pre-admission carry-forward out of run_conversation (facade) into
agent/turn_facade_lease.py::carry_unadmitted_user_message next to the early
result it repairs, and drop the extra turn-start flush in build_turn_context:
the follow-up turn's normal turn-start persist already writes the marked row
because _db_flush_collect no longer stamps it as durable (live probe: exactly
one A row in state.db after two flushes). Trim to two invariant tests
(carry-forward with metadata; flushed exactly once); the hard-stop negative is
covered by the E2E probe in the PR body.
The dynamic-shell-word rules fired on any `-del*`/`-exec*` substring after a
`find` anywhere in the segment, so quoted predicate arguments the shell never
expands were flagged: `find . -name 'log-del*'`, `find . -name 'pre-exec*.sh'`,
`find src -path '*-exec[0-9]*'`, and `echo find . -{delete,print}`.
Three changes close that:
- `find` must be the command word (_CMDPOS anchored, same as mkfs/rm/dd) and
the dynamic word must start a whitespace-delimited token (`(?<!\S)`), for
both the find rule and the rg/sort/ag/man program-option rule.
- Both rules now scan the quote-masked variant (_QUOTE_MASKED_DANGEROUS_
DESCRIPTIONS, same _mask_quoted_prose used by the positionless hardline
rules) so glob characters inside quotes are data, not expansion.
- _iter_shell_command_starts no longer treats the `{` inside a brace-expansion
word (`-{delete,print}`) as a brace-group opener; it split the word across a
marked start so `echo x; find . -{delete,print}` matched nothing. A brace
group opener is `{` as its own word (after whitespace or a separator).
The quoted-name cases join the inert parametrized test; the separator case
joins the dangerous one. Still two parametrized functions.
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.
The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
search_tool resolved its root via _resolve_path_for_task BEFORE the
NT-namespace check saw it (review finding): on Windows the resolve is the
SMB-auth trigger, on POSIX the task-base join hides the prefix from the
resolved-path denylist. The raw-string guard now runs first there too,
and the guard rides the existing top-level agent.file_safety import
instead of two function-local imports.
The tool-layer chokepoint test now covers all four entries and proves
none of them touched Path/_resolve_path_for_task/realpath before the
refusal; the 12 form-matrix tests collapse into one blocked/allowed
invariant over both classifiers.
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:
- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
raw-string check (never resolves — resolving IS the leak trigger).
Wired as the first check in get_read_block_error() and the write
denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
_check_sensitive_path (covers write_file_tool + patch_tool), before
the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md