Commit Graph

25247 Commits

Author SHA1 Message Date
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Teknium 0268c0b8c0 chore: map 2ndNatureAI attribution 2026-08-25 02:31:34 -07:00
Teknium ce9b9a6351 feat(computer_use): guide models from full-screen grabs to interactive lanes
Full-screen captures carry no element tree, so the CaptureResult now has a
'note' field surfaced in the tool summary telling the model to call
capture(app='<AppName>') or capture(app='desktop') when it needs to act on
what it sees. Schema description updated to distinguish app='screen'
(composited full-screen image) from app='desktop' (shell surface with
clickable elements); docs + regression tests (14, sabotage-verified) added.
2026-08-25 02:31:34 -07:00
2ndNatureAI aeac982223 fix(computer_use): route explicit screen capture to get_desktop_state
'Screenshot my screen' previously resolved the 'screen' sentinel to the OS
shell window (Progman/WorkerW) via list_windows — capturing the wallpaper +
icons layer, never the windows actually displayed. cua-driver's
get_desktop_state does a real composited full-screen grab; the
screen/fullscreen/all sentinels now route there directly, bypassing window
enumeration (which also keeps screenshots working when Windows UIA
enumeration hangs — trycua/cua#2110/#2113).

app='desktop' keeps the shell-window lane so desktop icons/taskbar stay
clickable.

Salvaged from PR #60081 by @2ndNatureAI (surgical reapply; original branch
predates the capture-routing refactor).
2026-08-25 02:31:34 -07:00
Teknium 15f7b7293c fix(terminal): persistent Docker containers are profile-scoped, not per-session
Commit a270c4ade's session-key fallback in _resolve_container_task_id was
added to stop cross-profile SSH environment reuse, but it wasn't backend-
gated: persistent Docker silently fragmented into one container per gateway
session, breaking the product contract (one long-lived container per profile,
shared by CLI and every session of that profile). #93950's vanishing MEDIA
attachments were downstream damage.

- persistent Docker (container_persistent: true) now keys to the profile:
  literal 'default' for the default profile (same container as CLI),
  'profile:<name>' for named profiles
- SSH and non-persistent Docker keep session scoping (the original leak fix
  and the #82731 isolation contract are untouched)
- gateway MEDIA translation follows the profile layout and keeps the legacy
  bug-window per-session sandboxes as fallback candidates, trying each until
  the file resolves — old sessions self-heal, no migration
- /root/.hermes credential-surface refusal preserved across all layouts
2026-08-25 02:30:38 -07:00
kshitij 4c1f53be10 Merge pull request #94568 from kshitijk4poor/fix/85125-2e-approval-outcome-parity-v2
fix(approval): machine-readable outcome parity on the gateway tails + sudo human-wait exclusion (#85125 2e)
2026-08-25 13:31:49 +05:30
kshitij dc3716d451 Merge pull request #94536 from kshitijk4poor/salvage/94439-computer-use-media-path
fix(gateway): widen computer-use media path repair to background and cron delivery
2026-08-25 13:31:05 +05:30
kshitijk4poor c8c3f4c448 fix(approval): machine-readable outcome parity on the gateway tails + sudo human-wait exclusion (#85125 2e) 2026-08-25 13:27:20 +05:30
kshitijk4poor d634b37047 refactor(gateway): retire private repair alias per replay_cleanup precedent
Phase 2c on the full final diff flagged the re-export shim as
contradicting the adjacent house pattern (agent.replay_cleanup import,
which documents retiring private aliases once tests migrate). Migrate
all six tests to the canonical gateway.media_repair seam, import the
canonical name in run.py, and drop the dead 'and result' guard at the
background-task call site.
2026-08-25 13:25:48 +05:30
kshitijk4poor e5032945cb chore: map contributor email for salvage attribution
Map macd@google.com (Mark McDonald, @markmcd) so the contributor
attribution check passes for the salvage of PR #94522.
2026-08-25 13:14:30 +05:30
Mark McDonald eaf6545ab4 docs(gemini): update to use latest gemini models 2026-08-25 13:14:30 +05:30
kshitijk4poor 105999a0c9 refactor(gateway): unify computer-use repair call sites after review
- Make repair_explicit_computer_use_media_paths fail-open internally
  (cosmetic repair must never abort delivery); drop the cron-only
  try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
  no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
  gateway.media_repair seam; pre-existing tests keep pinning the
  gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
2026-08-25 13:06:43 +05:30
kshitijk4poor bb0d5503c2 fix(gateway): widen computer-use media path repair to sibling surfaces
Follow-up to the salvaged fix from PR #94439:

- Extract the repair into gateway/media_repair.py (shared module) and
  re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
  background tasks (_run_background_task_inner) and cron job delivery
  (cron/scheduler.py) — both call agent.run_conversation directly and
  never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
  content first instead of regex-scanning the raw string, which yielded a
  doubled-backslash path artifact and rewrote the response to a path the
  model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
  gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
  the compression-fallback last-user slice (incl. no-user fail-closed).
2026-08-25 13:06:43 +05:30
Teknium 760d3af0d7 chore: map contributor email for salvage attribution 2026-08-25 00:31:14 -07:00
Guilherme Artiles 8ad20d065a test(curator): assert the instruction through the delivered prompt, not the source
The first version of this test used inspect.getsource() on skill_manager_tool
and regex-parsed the guarded action literals. AGENTS.md bans source-text tests,
and the ban is right here: that test would pass against a guard wired to the
wrong call site and fail on a pure rename, neither of which is the thing worth
guarding.

Replaced with a behavioral assertion in the shape of the neighbouring
dry-run-banner test: stub _run_llm_review, run run_curator_review, and assert
the prompt that actually reached the model names skill_view and all four
guarded actions (edit, patch, write_file, remove_file).

It still discriminates: with the prompt block removed, action=edit and
action=remove_file no longer appear anywhere in the assembled prompt (the
toolset list only mentions patch, create, write_file and delete), so the test
fails. Runtime guard behavior stays where it belongs, in
tests/tools/test_skill_manager_tool.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:31:14 -07:00
Guilherme Artiles 610e2e02fb fix(curator): tell the background reviewer to read before it writes
`_background_review_read_before_write_guard` refuses a background-review
`skill_manage` write when the target file was not loaded via `skill_view` in
the same review turn (patch, edit, write_file over an existing file,
remove_file).

`CURATOR_REVIEW_PROMPT` never says so. It lists `skill_view` only under "read
the current landscape", so the reviewer goes straight to the write and every
mutation is refused. The failure is silent from the outside: the curator run
completes, writes nothing, and reads like a pass that simply found nothing to
consolidate. On our deployment that was 32 of 32 attempted writes rejected over
48h before anyone read the logs.

This adds the missing instruction to the toolset block, plus a test that fails
if a future guarded action is added to `skill_manager_tool` without being named
in the prompt — the guard and the prompt have to drift together or not at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:31:14 -07:00
Teknium a70d2ffce5 fix(background-review): teach review prompts the enforced read-before-write handshake
The skill_manage guard (added in #55906) refuses any patch/edit of an
existing SKILL.md, or overwrite/removal of an existing support file,
unless the exact target was loaded via skill_view during the review.
Neither _SKILL_REVIEW_PROMPT nor _COMBINED_REVIEW_PROMPT ever mentioned
this, so models routinely issued the write without the pre-read, got
refused, and burned review iterations (#62397).

Both prompts now carry a Read-before-write section scoped to the
guard's actual contract: existing targets only, exact-path pre-read for
support files, transcript quotes don't count, new skills/new support
files exempt, and a bounded one-view-one-retry recovery instead of a
loop. Direction follows #60331 by @kkwills13 with the scope corrections
requested in review (existing-target-only wording, no delete claim,
bounded retry, contract tests for both prompt variants).

Fixes #62397.
2026-08-25 00:18:35 -07:00
kshitijk4poor b0cf2597c2 fix: follow-up for salvaged PR #93985 — cache key, snapshot, dead code
- Key _user_space_cache on _conn_snapshot instead of client object identity,
  so _new_client() results from the same connection share the cached user
  (previously every on_memory_write triggered an uncached /api/v1/system/status
  probe with a 30s default timeout)
- Thread a short timeout (0.05s) through the write-path identity probe
- Harden _tool_remember to snapshot the client before URI construction + POST,
  matching the pattern already established in on_memory_write
- Remove dead instance method _user_scoped_uri (zero callers; all call sites
  use the module-level function directly)

Co-authored-by: ehz0ah <haozhe4547@gmail.com>
2026-08-25 12:27:21 +05:30
ehz0ah 4387e03960 fix(memory): keep OpenViking identity operations consistent 2026-08-25 12:27:21 +05:30
ehz0ah 5ff03cb0c4 fix(memory): scope OpenViking user cache to connection 2026-08-25 12:27:21 +05:30
liuhao1024 7cd43cdf52 fix(memory): emit explicit-uid OpenViking URIs resolved from system status
Review follow-up to the viking://~ migration: the ~ home alias only
expands for USER/ADMIN roles. The DEFAULT dev auth mode (no
server.auth_mode, no root_api_key) resolves every request as ROOT,
which bypasses current-user expansion — the canonical parser rejects
viking://~ with 400 'Home alias URI is not canonical' (verified on a
live 0.4.16 server). A deployment upgrading to 0.4.16 with an
untouched ov.conf is in dev mode, so the ~ spelling would break
exactly the way the old uid-less one will.

Mirror the upstream first-party plugin pattern instead: resolve the
user space client-side from /api/v1/system/status (result.user,
'default' fallback) and emit explicit-uid
viking://user/<user>/memories/... URIs, which are canonical under
every auth mode (dev/ROOT, trusted/USER, api-key) and every server
version. viking://~/... input typed by the user keeps passing through
untouched. (#91995)
2026-08-25 12:27:21 +05:30
liuhao1024 fc4c2f456b fix(memory): migrate OpenViking URIs to the viking://~ home alias
Upstream OpenViking removed the uid-less viking://user/<segment>
shorthand (#4196, merged 2026-08-21): reserved segments like memories
and peers no longer expand to the caller's space and the server
rejects them with HTTP 400 (NamespaceShapeError). First-party clients
were migrated to viking://~ in the same change; the Hermes plugin was
not (#91995).

Migrate every URI the plugin constructs — the profile/preferences/
entities session-start reads, the _build_memory_uri memory-mirroring
write path, and the tool-schema example — to viking://~/... README
uid-less references updated to match; canonical user-scoped forms
(viking://user/default/...) are unchanged. The ~ alias requires
OpenViking server >= 0.4.16 (#4167).
2026-08-25 12:27:21 +05:30
Jony 335c60ecdd fix(skills): preserve review marks across contexts 2026-08-24 23:55:39 -07:00
Michael Nguyen 34041faea8 test(gateway): accept session_key kwarg in media resend dedup stubs
Same follow-through as the other filter-static stub updates: the three
lambdas patching filter_local_delivery_paths rejected the new keyword
and failed CI (tests/gateway/test_73771_media_resend_dedup.py).
2026-08-24 23:50:23 -07:00
Michael Nguyen a15533b646 feat(gateway): warn when a Docker sandbox MEDIA path fails translation
De-silence the #93950 failure mode: when a container-absolute MEDIA path
under /workspace or /root cannot be resolved to a host sandbox file while
TERMINAL_ENV=docker, log the reason (no mounts / no prefix match / host
file missing) plus the delivering session key instead of only the generic
'Skipping unsafe MEDIA directive path' line upstream.
2026-08-24 23:50:23 -07:00
Michael Nguyen d4f31a8f36 fix(gateway): resolve session-scoped Docker sandboxes for MEDIA delivery (#93950)
Persistent Docker containers bind <sandboxes>/docker/<task>/{workspace,home}
where <task> is sanitize_task_id_for_path("session:<session_key>") — but the
gateway's synthetic mounts hardcoded the literal "default" sandbox
(_default_docker_workspace_host_root / _docker_persistent_home_host_root).
For any session-scoped deployment the longest-prefix match missed, the
container path fell through to a host-filesystem resolve that could not
exist, and every MEDIA attachment was silently dropped.

The post-handler delivery pipeline also runs after
_handle_message_with_agent cleared the turn's session contextvars, so even
a correct sandbox derivation consulting ambient state would collapse onto
"default". Thread the delivering session's key explicitly through
validate_media_delivery_path -> _translate_docker_container_media_path ->
the two host-root helpers (same pattern as the TTS fix for #57049/#36685).

Default-sandbox resolution and the /root/.hermes credential exclusion are
preserved; contexts without a key keep the historical behavior.
2026-08-24 23:50:23 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
Gille 1fac440865 fix(gateway): recover explicit computer-use media paths 2026-08-24 23:39:05 -07:00
xxxigm fcda325cd4 test(telegram): cover send() waiting for reconnect after a network blip
Pin immediate replacement, mid-wait restore, timeout retryable=True,
and permanent-fatal fail-closed without waiting.
2026-08-25 12:07:36 +05:30
xxxigm 6c1bfff65c fix(telegram): wait for reconnect before failing send as Not connected
A short Telegram drop used to fail the final reply immediately. The
answer then sat in the delivery ledger until the next gateway boot.
Wait up to 15s for the bot (or a replacement adapter) so a brief
blip delivers now, matching QQBot.
2026-08-25 12:07:36 +05:30
xxxigm 27640c5844 test(teams): cover connect when the namespace exists but App is unbound
Pin the NoneType crash: TEAMS_SDK_AVAILABLE true plus a failed bind must
return False without calling App(). Also pin plugin import when the
microsoft_teams parent namespace is missing.
2026-08-25 12:07:33 +05:30
xxxigm f84f94b494 fix(teams): do not call App() when the SDK was never bound
find_spec("microsoft_teams") can be true from sibling namespace packages
while App is still None, so a failed lazy-install crashed connect with
'NoneType' object is not callable instead of a missing-SDK error.
Probe microsoft_teams.apps via the parent first — a dotted find_spec
raises ModuleNotFoundError on 3.11 when the namespace is absent.
2026-08-25 12:07:33 +05:30
Leegenux 6ce7ab8bfb feat(browser): make snapshot threshold configurable 2026-08-24 21:51:44 -07:00
Teknium 9cce872505 docs: note httpcore pin dependency in _enable_happy_eyeballs 2026-08-24 21:46:05 -07:00
Nathan Shan d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
pierrenode bf8b28f27a fix(tools): route browser snapshot storage through the symlink-safe writer
Today's spill/cache-writer hardening (tools/spill_safety.py,
write_text_exclusive/ensure_spill_dir with O_CREAT|O_EXCL|O_NOFOLLOW)
migrated tools/web_tools.py::_store_full_text() — which writes to the same
cache/web directory with the same content-hash filename scheme — but left
its near-identical sibling, tools/browser_tool.py::_store_full_snapshot(),
on the pre-fix plain open()/write_text() pattern. A pre-planted symlink at
the content-hash path redirected the write onto an arbitrary user-owned
file, same as the sites that commit fixed.

Reproduced live: with a symlink planted at the exact
browser-snapshot-<digest>.txt path (predictable from the snapshot content
hash), the pre-fix write followed the link and overwrote the link's
target with the (secret-redacted but otherwise user/page-controlled)
snapshot content.

Fix mirrors _store_full_text's exact usage: ensure_spill_dir(private=False)
+ write_text_exclusive(private=False, overwrite=True) — not private since
cache/web is bind-mounted into remote backends whose container UID must
read it; overwrite=True because re-snapshotting the same page state
legitimately reuses the same content-hash name (the overwrite path
lstat-unlinks the link itself, never following it to write through).

Added a regression test planting a symlink at the exact digest path and
asserting the link's target is untouched (only the link itself gets
safely replaced by a real file). Mutation-verified: with the fix stashed,
the pre-fix code wrote the snapshot content into the symlink's target
file, reproducing the vulnerability exactly.
2026-08-24 21:45:56 -07:00
Teknium c1b295d003 test(desktop): stop the syntax-diff mock factory from leaking unhandled rejections (#94415)
The diff-lines error-boundary test mocked './syntax-diff' with a factory
that THREW. vitest hoists the factory and registers its module promise in
the mocker registry; a throwing factory leaves rejected promises there,
and under CI load one escapes as "Vitest caught 1 unhandled error during
the test run" attributed to whichever sibling file the worker is running
(user-message-edit.test.tsx in run 32803716726) — an intermittent js-tests
red on green code.

Rework: the factory now resolves to a component that throws the fetch
error during render — the same way React surfaces a rejected lazy payload
— so no rejected promise ever sits in the registry.

Guard proof (sabotage A/B): with the local syntax-diff ErrorBoundary
removed from diff-lines.tsx, the reworked test still fails (workspace
fallback renders), so the #93479 regression pin is intact. Full ui
project: 596 files / 5729 tests green, no unhandled errors in 3 full-run
repetitions.
2026-08-24 21:45:44 -07:00
Teknium 8172be0e8e fix(memory): log when a configured provider's tools are gated off by toolset config
Follow-up to the cherry-picked gate-parity fix: the silent 'return 0'
in inject_memory_provider_tools made #81014 undiagnosable — a
configured provider looked half-on with no hint which config key
suppressed its tools. Now an INFO line names the withheld providers
and the gating keys.
2026-08-24 21:45:30 -07:00
Chen Jin b45b028573 fix(agent): gate memory provider system_prompt_block on toolset config (#81014)
The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provider_tools_enabled()` via platform_toolsets or
disabled_toolsets. Result: the agent received instructions to call
`mnemosyne_remember`, `mnemosyne_recall`, etc., that did not exist in
its tool surface.

Centralize the gating into `memory_provider_tools_exposed(agent)`, use
it from both `inject_memory_provider_tools` and the system prompt
assembly path, and add regression tests covering:
* memory toolset enabled -> both tools and prompt block exposed,
* memory in disabled_toolsets -> neither exposed,
* memory not in enabled_toolsets and not built-in -> neither exposed,
* the built-in "memory" tool present as an opt-in -> both exposed,
* parity between `inject_memory_provider_tools` and
  `memory_provider_tools_exposed`.
2026-08-24 21:45:30 -07:00
Teknium bc47fcd3f9 fix(tests): e2e group-restart test no longer flakes on cold SessionDB init
The /goal post-turn hook constructs a real SessionDB on an executor
thread at the turn boundary. On a cold or loaded CI runner that
state.db init can exceed send_and_capture's 2s poll window, so the
send lands after the assertion and the test reports the bare
'Expected mock to have been called once. Called 0 times.' (#92130).
Mock _run_post_turn_hooks in the e2e runner — these tests exercise
gateway command dispatch, not goal hooks.

Also scrub TELEGRAM_GROUP_ALLOWED_CHATS / *_GROUP_ALLOWED_USERS / QQ
allowlist env vars in the hermetic conftest: a developer shell with
those set flips _get_unauthorized_dm_behavior to 'ignore' and fails
the pairing e2e test locally.
2026-08-24 21:21:34 -07:00
Teknium 64a6f42cb3 test: opted-out profile seeding now exercises the essential-only sync subprocess 2026-08-24 20:25:10 -07:00
Teknium 3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
hermes-seaeye[bot] 7c97343950 fmt(js): npm run fix on merge (#94410)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-25 03:17:26 +00:00
Teknium beb7941236 fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219)
The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.

- server: events_since() now returns bare event objects (the frame's
  params), the exact shape the live dispatch path consumes; ring stores
  params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
  seq-gated afterward — no double dispatch of deltas, no gap-skip from
  a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
  reset them while clients kept high watermarks (replay forever empty,
  truncated=false). New replay_epoch advertised in gateway.ready and
  echoed by session.events.since; the client clears watermarks on epoch
  change.
- methods_session no longer reaches into event_replay privates
  (is_truncated() accessor).

Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.
2026-08-24 20:11:30 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
Teknium 48f69e51d3 fix(signal): chunk long standalone sends and cover both delivery paths (salvage #57929 + #67279)
Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.

- tools/send_message_tool.py: register Signal's 8000-char limit in
  _MAX_LENGTHS (imported from the adapter module so the two paths can't
  drift) so hermes send / cron standalone / MCP sends split via the
  shared truncate_message() pass instead of signal-cli rejecting them.
  Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
  adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).

Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.
2026-08-24 20:03:24 -07:00
lkz-de cbc8d1804d fix(signal): chunk long cron deliveries instead of truncating 2026-08-24 20:03:24 -07:00
Brooklyn Nicholson f8b52e4d80 test(desktop): cover UI scale across recordless hash routes
Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.

Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>
2026-08-24 22:01:47 -05:00
Clark Vines b637ee0fc6 fix(desktop): keep UI scale across in-page route navigation
Desktop is a HashRouter over one file:// document, so every route is a
distinct URL to Chromium's per-URL zoom store. A route the user never
zoomed on has no record at all and resolves to the host default (100%) —
that is every fresh session and every never-visited settings tab.

In-page navigation fires neither did-finish-load nor any window event,
so nothing re-asserted the persisted level. The window dropped to 100%
while the Appearance control kept reading the chosen scale, because the
renderer only learns of zoom changes through 'hermes:zoom:changed',
which never fired. Touching the setting sent a fresh apply, which is
why it appeared to fix itself.

Re-assert the persisted level on main-frame did-navigate-in-page.
Verified on real Electron 40.10.2 / Chromium 144 (win32): a recordless
hash route reports 100% at the event, so the existing drift-guard sees
the drop and re-applies, and still no-ops when the route's record
already matches.

Fixes #48658
Fixes #38854
Fixes #79863

Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
2026-08-24 22:01:47 -05:00