Merge latest origin/main into skill metrics

Signed-off-by: Alex Fournier <afournier@nvidia.com>
This commit is contained in:
Alex Fournier
2026-08-04 15:06:36 -07:00
91 changed files with 5603 additions and 366 deletions
+122
View File
@@ -0,0 +1,122 @@
name: Install & Update E2E (reusable)
# Runs ONE update route against ONE starting commit, in the dev sandbox, with a
# real install (uv, a managed Python, Node, the venv) behind it.
#
# Reusable so callers can fan out over the combinations that matter -- update
# from the tip vs. from an older release, `hermes update` vs. re-running the
# installer -- without duplicating the runner setup. Each leg is independent:
# its own sandbox, its own install, nothing rewound or shared.
#
# Call it:
#
# jobs:
# tip:
# uses: ./.github/workflows/install-e2e-run.yml
# with:
# route: update
# install-ref: refs/heads/main
on:
workflow_call:
inputs:
route:
description: 'Update path to exercise: update (hermes update) or installer (re-run install.sh).'
required: true
type: string
install-ref:
description: 'What to install before updating: a branch, a tag (v2026.7.7), or a SHA reachable from main.'
required: false
type: string
default: refs/heads/main
runner:
description: 'Runner label.'
required: false
type: string
default: ubuntu-latest
timeout-minutes:
description: 'Job timeout. A cold run installs real toolchains twice.'
required: false
type: number
default: 45
permissions:
contents: read
jobs:
e2e:
name: ${{ inputs.route }} from ${{ inputs.install-ref }}
runs-on: ${{ inputs.runner }}
timeout-minutes: ${{ inputs.timeout-minutes }}
steps:
# Full history: the sandbox fetches the starting commit and the test
# compares against this commit, so a shallow clone is not enough.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
# bubblewrap + slirp4netns are what the sandbox is built on; util-linux
# supplies the `unshare` that builds the multi-uid userns for the
# user-level (non-root) install.
- name: Install sandbox dependencies
run: |
set -euo pipefail
sudo apt-get update -qq
sudo apt-get install -y -qq bubblewrap slirp4netns uidmap util-linux
# Ubuntu 24.04 restricts unprivileged user namespaces through AppArmor,
# which is exactly what bwrap needs. Report the state before touching it
# so a future runner-image change is visible in the log rather than
# silently altering what this job proves.
- name: Permit unprivileged user namespaces
run: |
set -euo pipefail
echo "--- kernel userns settings (before)"
sysctl kernel.unprivileged_userns_clone 2>/dev/null || echo " (sysctl absent)"
sysctl kernel.apparmor_restrict_unprivileged_userns 2>/dev/null || echo " (sysctl absent)"
if sysctl -n kernel.apparmor_restrict_unprivileged_userns >/dev/null 2>&1; then
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
fi
echo "--- subuid/subgid for $(id -un)"
grep "^$(id -un):" /etc/subuid /etc/subgid || echo " (none — sandbox will say so)"
- name: Run install + update E2E
run: |
set -euo pipefail
tests/install/install-update-e2e.sh \
--route '${{ inputs.route }}' \
--install-ref '${{ inputs.install-ref }}'
env:
# Outside the workspace on purpose: the script creates this directory
# up front, and an untracked dir inside the repo makes the worktree
# dirty -- which dev-sandbox reacts to by snapshotting the working
# copy into a fresh fake-main commit on every invocation, moving the
# update target mid-run.
HERMES_E2E_LOG_DIR: ${{ runner.temp }}/e2e-logs
# Artifact names cannot contain '/', and install-ref may be a full ref
# like refs/heads/main. GitHub Actions expressions have no string-replace
# function, so build the safe name here. Runs even on failure -- that is
# exactly when the logs are wanted.
- name: Build artifact name
if: always()
id: artifact
run: |
set -euo pipefail
safe_ref='${{ inputs.install-ref }}'
safe_ref="${safe_ref//\//-}"
echo "name=install-e2e-${{ inputs.route }}-${safe_ref}" >> "$GITHUB_OUTPUT"
# The installer's own transcripts say far more than the assertion that
# tripped when a real install breaks.
- name: Upload installer logs
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
# Unique per leg: a matrix over releases runs this workflow several
# times per route, and same-named artifacts collide.
name: ${{ steps.artifact.outputs.name }}-${{ github.sha }}
path: ${{ runner.temp }}/e2e-logs
retention-days: 14
if-no-files-found: ignore
+110
View File
@@ -0,0 +1,110 @@
name: Install & Update E2E
# Can a user on a released version get to this commit?
#
# For each release we sample, a leg installs that release through the real
# `curl | install.sh` one-liner (uv, a managed Python, Node, the venv) inside
# scripts/dev-sandbox.sh, then applies one update route and requires the
# checkout to land on this commit with a working `hermes`.
#
# The starting versions are chosen at runtime from the repo's release tags
# (scripts/sandbox/pick-release-tags.sh): newest, oldest, and a spread between.
# A hardcoded list would stop covering the newest release the day after it
# ships, and would pin an "oldest" that nobody still runs.
#
# Triggers:
# * every 12 hours, so upstream drift (a new uv, a Node bump, a PyPI change)
# surfaces on a schedule rather than in someone's review cycle;
# * when a release tag is created -- the moment the set of versions users can
# update FROM changes, and the moment a broken updater would strand them;
# * manually, where you can pick the route and how many releases to sample.
#
# Deliberately NOT on pull_request: a leg takes ~11 minutes of real toolchain
# installation, and the matrix multiplies that. Updating is release-shaped work,
# so it is gated on releases and the clock instead.
on:
workflow_dispatch:
inputs:
route:
description: 'Which update route to exercise.'
required: false
type: choice
default: both
options: [both, update, installer]
tag-count:
description: 'How many release tags to sample (newest, oldest, and a spread between).'
required: false
type: string
default: '5'
schedule:
# Every 12 hours, off the hour to avoid the top-of-hour runner crunch.
- cron: '20 7,19 * * *'
push:
tags:
# Release tags only: the repo also carries backup/* and one-off tags.
- 'v[0-9]+.[0-9]+.[0-9]+'
- 'v[0-9]+.[0-9]+.[0-9]+.[0-9]+'
permissions:
contents: read
concurrency:
group: install-e2e-${{ github.ref }}
cancel-in-progress: true
jobs:
# Which released versions do we test updating FROM? Resolved once and shared
# by both route matrices, so the two routes cover the same set.
pick-releases:
name: Pick release tags
runs-on: ubuntu-latest
timeout-minutes: 5
outputs:
tags: ${{ steps.pick.outputs.tags }}
steps:
# This job only reads tag names and runs one script, so take the cheap
# checkout: no blobs (filter), no other files (sparse), but DO fetch tags
# -- they are the whole input, and the default shallow checkout has none.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
filter: blob:none
fetch-tags: true
sparse-checkout: scripts/sandbox/pick-release-tags.sh
sparse-checkout-cone-mode: false
- id: pick
run: |
set -euo pipefail
tags="$(scripts/sandbox/pick-release-tags.sh --count '${{ inputs.tag-count || 5 }}')"
echo "Testing updates from: $tags"
echo "tags=$tags" >> "$GITHUB_OUTPUT"
# `hermes update` -- the route most users take.
update:
if: github.event_name != 'workflow_dispatch' || inputs.route != 'installer'
needs: pick-releases
strategy:
# One release breaking is worth knowing about even if another already
# failed, so let every leg report.
fail-fast: false
matrix:
install-ref: ${{ fromJSON(needs.pick-releases.outputs.tags) }}
uses: ./.github/workflows/install-e2e-run.yml
with:
route: update
install-ref: ${{ matrix.install-ref }}
# Re-running the curl one-liner over an existing checkout: autostash + pull
# rather than the updater's own git handling.
installer:
if: github.event_name != 'workflow_dispatch' || inputs.route != 'update'
needs: pick-releases
strategy:
fail-fast: false
max-parallel: 3
matrix:
install-ref: ${{ fromJSON(needs.pick-releases.outputs.tags) }}
uses: ./.github/workflows/install-e2e-run.yml
with:
route: installer
install-ref: ${{ matrix.install-ref }}
+3
View File
@@ -145,6 +145,9 @@ docs/superpowers/*
# Persistent dev sandbox dir (scripts/dev-sandbox.sh --persistent)
.hermes-sandbox/
# Sandbox dirs used by the install/update E2E (tests/install/). The suffix is
# the route name, so each route gets its own tree and two can run at once.
.hermes-sandbox-e2e*/
# Interrupted-update breadcrumb + recovery lock written next to the shared venv
# by `hermes update` / launch-time self-heal. Runtime state, never a code change
+7
View File
@@ -1020,6 +1020,13 @@ def recover_with_credential_pool(
}
if _credential_id:
kwargs["credential_id"] = _credential_id
# Hand the pool the classified semantics, not just the status. A
# billing 403 (OpenRouter "key limit exceeded", xAI spending limit)
# and an edge-throttle 403 are the same number but need opposite
# cooldowns — the pool can only tell them apart if we say which.
# ``effective_reason`` is resolved below; this closure runs after.
if effective_reason is not None:
kwargs["failure_reason"] = effective_reason.value
return pool.mark_exhausted_and_rotate(**kwargs)
effective_reason = classified_reason
+30
View File
@@ -1291,6 +1291,14 @@ def run_conversation(
agent._last_compression_attempt_recorded = False
agent._last_compression_attempt_in_place = None
# Adopt any ~/.hermes/.env credential/base-url edits made since the last
# turn — a Settings save updates .env but not this worker's client, which
# was built at agent init (#67821). No-op when .env is unchanged.
try:
agent._try_refresh_env_client_credentials()
except Exception:
logger.debug("per-turn env credential refresh failed", exc_info=True)
# ── Per-turn setup (the prologue) ──
# All once-per-turn setup — stdio guarding, retry-counter resets, user
# message sanitization, todo/nudge hydration, system-prompt restore-or-
@@ -4234,6 +4242,28 @@ def run_conversation(
f" Check which providers support tools: https://openrouter.ai/models/{_model}"
)
# Actionable hint for a bare 404 on a provider whose catalogue
# uses ``vendor/model`` ids. A model id that lost its prefix
# (e.g. ``nemotron-…`` instead of ``nvidia/nemotron-…``) gets
# a content-free "404 page not found" from the provider that
# never names the model, so it reads like an outage or an auth
# failure. Name the real cause and the exact id to use (#78796).
if getattr(api_error, "status_code", None) == 404:
try:
from hermes_cli.model_normalize import suggest_prefixed_model_id
_suggestion = suggest_prefixed_model_id(_provider, _model)
except Exception:
_suggestion = None
if _suggestion:
agent._buffer_vprint(
f" 💡 Model '{_model}' is not a valid id for provider {_provider} — "
f"it is missing its vendor prefix."
)
agent._buffer_vprint(
f" Did you mean '{_suggestion}'? Re-pick it with `hermes model`."
)
# Check for interrupt before deciding to retry
if agent._interrupt_requested:
# Preserve a pending redirect (mid-stream correction): the
+133 -32
View File
@@ -124,6 +124,17 @@ SUPPORTED_POOL_STRATEGIES = {
EXHAUSTED_TTL_401_SECONDS = 5 * 60 # 5 minutes
EXHAUSTED_TTL_429_SECONDS = 60 * 60 # 1 hour
EXHAUSTED_TTL_DEFAULT_SECONDS = 60 * 60 # 1 hour
# When a pool has no other credential to rotate to (the offending key is the
# sole non-DEAD entry), a 1-hour bench means an hour of hard failures with
# nothing to fall back to. Throttles (429/403/5xx) are transient and reset in
# seconds, so a sole credential cools down briefly instead — same rationale as
# the short 401 cooldown above. Provider-supplied reset_at still overrides.
EXHAUSTED_TTL_SOLE_CREDENTIAL_SECONDS = 60 # 1 minute
# ``FailoverReason.billing`` as a bare string. The pool stores classified
# failure semantics as plain text (it persists to JSON and must not import
# the classifier), so the value is duplicated here rather than referenced.
FAILURE_REASON_BILLING = "billing"
# Throttle window for the "no available entries" INFO line. Credential
# selection runs on a hot path (every model call, plus auxiliary tasks like
@@ -150,6 +161,13 @@ _EXTRA_KEYS = frozenset({
"token_type", "scope", "client_id", "portal_base_url", "obtained_at",
"expires_in", "agent_key_id", "agent_key_expires_in", "agent_key_reused",
"agent_key_obtained_at", "tls", "secret_source", "secret_fingerprint",
# Classified failure semantics for the last exhaustion, as decided by
# agent/error_classifier.py. The raw HTTP status is not enough to size a
# cooldown: providers return 403 for both an edge throttle (transient,
# seconds) and a spending/key limit (billing, needs a real fix). Persisted
# with the entry so a restart doesn't downgrade a billing bench back to a
# 60s transient cooldown.
"failure_reason",
})
@@ -289,13 +307,39 @@ def _is_manual_source(source: str) -> bool:
return normalized == SOURCE_MANUAL or normalized.startswith(f"{SOURCE_MANUAL}:")
def _exhausted_ttl(error_code: Optional[int]) -> int:
"""Return cooldown seconds based on the HTTP status that caused exhaustion."""
def _exhausted_ttl(
error_code: Optional[int],
*,
sole_credential: bool = False,
failure_reason: Optional[str] = None,
) -> int:
"""Return cooldown seconds based on the HTTP status that caused exhaustion.
When *sole_credential* is True the pool has no other entry to rotate to, so
a long bench just blocks the only key. Transient throttles (429 and the
catch-all default, which covers 403/5xx/unknown) are capped to a brief
cooldown so the sole key can recover — mirroring the short 401 path. 401
keeps its own (already short) TTL.
*failure_reason* is the classified semantics from
``agent/error_classifier.py``. The raw status alone can't size the
cooldown: an OpenRouter ``key limit exceeded`` and an xAI spending-limit
block both arrive as **403** but classify as ``billing``, and a 60s retry
on a spent account just re-fails every minute. Billing keeps the full
bench regardless of status; 402 does too, since it is billing by
definition even when nothing classified it.
"""
if error_code == 401:
return EXHAUSTED_TTL_401_SECONDS
if error_code == 429:
return EXHAUSTED_TTL_429_SECONDS
return EXHAUSTED_TTL_DEFAULT_SECONDS
base = EXHAUSTED_TTL_429_SECONDS if error_code == 429 else EXHAUSTED_TTL_DEFAULT_SECONDS
# Sole credential: shorten only TRANSIENT throttles (429 rate-limit, 403
# edge-throttle, 5xx server, or unknown). Billing exhaustion — whether
# classified as such or self-evident from a 402 — is a genuine depletion
# where a quick retry can't help, so it keeps the full bench.
is_billing = error_code == 402 or failure_reason == FAILURE_REASON_BILLING
if sole_credential and not is_billing:
return min(base, EXHAUSTED_TTL_SOLE_CREDENTIAL_SECONDS)
return base
def _parse_absolute_timestamp(value: Any) -> Optional[float]:
@@ -376,14 +420,18 @@ def _normalize_error_context(error_context: Optional[Dict[str, Any]]) -> Dict[st
return normalized
def _exhausted_until(entry: PooledCredential) -> Optional[float]:
def _exhausted_until(entry: PooledCredential, *, sole_credential: bool = False) -> Optional[float]:
if entry.last_status != STATUS_EXHAUSTED:
return None
reset_at = _parse_absolute_timestamp(getattr(entry, "last_error_reset_at", None))
if reset_at is not None:
return reset_at
if entry.last_status_at:
return entry.last_status_at + _exhausted_ttl(entry.last_error_code)
return entry.last_status_at + _exhausted_ttl(
entry.last_error_code,
sole_credential=sole_credential,
failure_reason=getattr(entry, "failure_reason", None),
)
return None
@@ -645,11 +693,18 @@ class CredentialPool:
available, _pending = self._available_entries()
if available:
return None
# Mirror _available_entries: if the pool has no other credential
# to rotate to, the sole entry's transient throttle cools down in
# seconds — next_available_at must report that shorter window too,
# or the fallback restore gate waits an hour for a 60s cooldown.
sole_credential = sum(
1 for e in self._entries if e.last_status != STATUS_DEAD
) <= 1
candidates: List[float] = []
for entry in self._entries:
if entry.last_status != STATUS_EXHAUSTED:
continue
until = _exhausted_until(entry)
until = _exhausted_until(entry, sole_credential=sole_credential)
if until is not None:
candidates.append(until)
return min(candidates) if candidates else None
@@ -742,6 +797,7 @@ class CredentialPool:
error_context: Optional[Dict[str, Any]] = None,
*,
persist: bool = True,
failure_reason: Optional[str] = None,
) -> PooledCredential:
normalized_error = _normalize_error_context(error_context)
# Permanent OAuth failures (token_invalidated, token_revoked, etc.)
@@ -755,6 +811,15 @@ class CredentialPool:
terminal_status = STATUS_DEAD
else:
terminal_status = STATUS_EXHAUSTED
# Carry the classifier's verdict onto the entry so the cooldown can be
# sized by what actually failed, not just the HTTP status (a billing
# 403 must not get the sole-credential transient cooldown). Absent a
# classification, clear any stale verdict from a previous failure.
updated_extra = dict(entry.extra)
if failure_reason:
updated_extra["failure_reason"] = failure_reason
else:
updated_extra.pop("failure_reason", None)
updated = replace(
entry,
last_status=terminal_status,
@@ -763,6 +828,7 @@ class CredentialPool:
last_error_reason=normalized_error.get("reason"),
last_error_message=normalized_error.get("message"),
last_error_reset_at=normalized_error.get("reset_at"),
extra=updated_extra,
)
self._replace_entry(entry, updated)
if persist:
@@ -1759,6 +1825,12 @@ class CredentialPool:
# that can block for 20+ seconds. We collect them under self._lock
# and refresh outside the lock to avoid stalling all pool consumers.
pending_refresh: List[tuple] = [] # (entry, sync_entry_fn)
# DEAD entries never re-enter rotation, so if at most one non-DEAD entry
# exists there is nothing to rotate to: an exhausted sole credential
# should cool down briefly rather than bench the only key for an hour.
sole_credential = sum(
1 for e in self._entries if e.last_status != STATUS_DEAD
) <= 1
for entry in self._entries:
# Borrowed credentials persist as metadata-only references and are
# hydrated from their live source on load. A stale duplicate row
@@ -1839,7 +1911,7 @@ class CredentialPool:
# the re-auth case for OAuth singletons.
continue
if entry.last_status == STATUS_EXHAUSTED:
exhausted_until = _exhausted_until(entry)
exhausted_until = _exhausted_until(entry, sole_credential=sole_credential)
if exhausted_until is not None and now < exhausted_until:
# Codex quota windows can reopen EARLY: the user redeems a
# banked rate-limit reset (Codex CLI / ChatGPT UI), upgrades
@@ -1963,6 +2035,7 @@ class CredentialPool:
error_context: Optional[Dict[str, Any]] = None,
api_key_hint: Optional[str] = None,
credential_id: Optional[str] = None,
failure_reason: Optional[str] = None,
) -> Optional[PooledCredential]:
with self._lock:
entry = None
@@ -2041,7 +2114,9 @@ class CredentialPool:
if entry is None:
return None
_label = entry.label or entry.id[:8]
self._mark_exhausted(entry, status_code, error_context)
self._mark_exhausted(
entry, status_code, error_context, failure_reason=failure_reason
)
# A 402/429/401 is an API-key–level failure: the account is out of
# balance, rate-limited, or its key is rejected. The same key can
# back more than one pool entry (e.g. an explicit pool entry plus a
@@ -2061,7 +2136,11 @@ class CredentialPool:
continue
if sibling.runtime_api_key == failed_runtime_key:
self._mark_exhausted(
sibling, status_code, error_context, persist=False
sibling,
status_code,
error_context,
persist=False,
failure_reason=failure_reason,
)
siblings_marked = True
if siblings_marked:
@@ -2300,6 +2379,11 @@ def _upsert_entry(entries: List[PooledCredential], provider: str, source: str, p
field_updates = {}
extra_updates = {}
_field_names = {f.name for f in fields(existing)}
token_changed = (
"access_token" in payload
and payload["access_token"] is not None
and payload["access_token"] != existing.access_token
)
for key, value in payload.items():
if key in {"id", "priority"} or value is None:
continue
@@ -2311,6 +2395,15 @@ def _upsert_entry(entries: List[PooledCredential], provider: str, source: str, p
elif key in _EXTRA_KEYS:
if existing.extra.get(key) != value:
extra_updates[key] = value
# When the credential token itself changes (key rotation), clear any
# exhaustion/error state — the old status is stale for the new key.
if token_changed and existing.last_status is not None:
field_updates["last_status"] = None
field_updates["last_status_at"] = None
field_updates["last_error_code"] = None
field_updates["last_error_reason"] = None
field_updates["last_error_message"] = None
field_updates["last_error_reset_at"] = None
if field_updates or extra_updates:
if extra_updates:
field_updates["extra"] = {**existing.extra, **extra_updates}
@@ -2728,6 +2821,31 @@ def _seed_from_singletons(provider: str, entries: List[PooledCredential]) -> Tup
return changed, active_sources
# Prefer ~/.hermes/.env over os.environ — the user's config file is the
# authoritative source for Hermes credentials. Stale env vars from parent
# processes (Codex CLI, test scripts, etc.) should not override deliberate
# changes to the .env file. load_env() memoizes on the .env mtime, so
# per-call reads (pool seeding, per-turn credential refresh) cost a stat()
# when the file is unchanged.
def get_env_prefer_dotenv(key: str) -> str:
env_file = load_env()
raw = env_file.get(key, "").strip()
scoped_value = (_get_secret(key, "") or "").strip()
# If .env contains an unresolved op:// reference, prefer the
# already-resolved value supplied by the active secret scope (or by
# os.environ in legacy single-profile mode), set by
# load_hermes_dotenv() -> apply_onepassword_secrets()). The raw
# "op://Vault/Item/field" string would otherwise win and every
# provider auth attempt would receive a URL instead of a key. This
# happens during a partial migration, or when the user wrote op://
# references straight into .env rather than the secrets.onepassword
# config block. For every non-op:// value the original
# .env-takes-precedence behaviour is preserved unchanged.
if raw.startswith("op://") and scoped_value:
return scoped_value
return raw or scoped_value
def _seed_from_env(provider: str, entries: List[PooledCredential]) -> Tuple[bool, Set[str]]:
changed = False
active_sources: Set[str] = set()
@@ -2746,27 +2864,10 @@ def _seed_from_env(provider: str, entries: List[PooledCredential]) -> Tuple[bool
if provider == "copilot":
return False, active_sources
# Prefer ~/.hermes/.env over os.environ — the user's config file is the
# authoritative source for Hermes credentials. Stale env vars from parent
# processes (Codex CLI, test scripts, etc.) should not override deliberate
# changes to the .env file.
def _get_env_prefer_dotenv(key: str) -> str:
env_file = load_env()
raw = env_file.get(key, "").strip()
scoped_value = (_get_secret(key, "") or "").strip()
# If .env contains an unresolved op:// reference, prefer the
# already-resolved value supplied by the active secret scope (or by
# os.environ in legacy single-profile mode), set by
# load_hermes_dotenv() -> apply_onepassword_secrets()). The raw
# "op://Vault/Item/field" string would otherwise win and every
# provider auth attempt would receive a URL instead of a key. This
# happens during a partial migration, or when the user wrote op://
# references straight into .env rather than the secrets.onepassword
# config block. For every non-op:// value the original
# .env-takes-precedence behaviour is preserved unchanged.
if raw.startswith("op://") and scoped_value:
return scoped_value
return raw or scoped_value
# The .env-preferring resolution lives at module level
# (``get_env_prefer_dotenv``) so the pool seeder and the per-turn
# credential refresh share one implementation.
_get_env_prefer_dotenv = get_env_prefer_dotenv
# Honour user suppression — `hermes auth remove <provider> <N>` for an
# env-seeded credential marks the env:<VAR> source as suppressed so it
+38
View File
@@ -339,6 +339,32 @@ _MODEL_NOT_FOUND_PATTERNS = [
"no endpoints found that support tool use",
]
def _model_id_missing_known_prefix(model: str, provider: str) -> bool:
"""True when a bare model id is only known to the provider as ``vendor/id``.
Some providers answer a malformed model id with a naked 404 that names
nothing — NVIDIA NIM returns ``404 page not found`` for a bare
``nemotron-3-ultra-550b-a55b``, indistinguishable from a bad endpoint
path. Consulting the curated catalogue tells the two apart: if the id
carries no ``/`` but the catalogue has exactly one entry ending in
``/<id>``, the prefix was dropped and the failure is deterministic.
Never guesses — an id absent from the catalogue (a local NIM container,
a proxied model) returns False so genuine endpoint problems keep their
retryable ``unknown`` classification.
"""
name = (model or "").strip()
if not name or "/" in name:
return False
try:
from hermes_cli.model_normalize import suggest_prefixed_model_id
return bool(suggest_prefixed_model_id((provider or "").strip(), name))
except Exception:
return False
# Malformed-message-array 400s. Deterministic request-shape rejections that
# describe the *transcript* being invalid, not a parameter. The canonical
# case: a stream dies mid-response and Hermes persists a content-less
@@ -1061,6 +1087,18 @@ def _classify_by_status(
retryable=False,
should_fallback=True,
)
# A bare id that the provider's catalogue only knows in prefixed form
# is a malformed model id, not a routing glitch — NVIDIA NIM answers
# one with a naked ``404 page not found`` that names nothing, so the
# generic branch below burns three retries and reports what looks
# like an outage (#78796). Deterministic: don't retry, and let the
# model_not_found surface carry the real cause.
if _model_id_missing_known_prefix(model, provider):
return result_fn(
FailoverReason.model_not_found,
retryable=False,
should_fallback=True,
)
# Generic 404 with no "model not found" signal — could be a wrong
# endpoint path (common with local llama.cpp / Ollama / vLLM when
# the URL is slightly misconfigured), a proxy routing glitch, or
+17 -1
View File
@@ -10333,7 +10333,7 @@ ipcMain.handle('hermes:notify', (_event, payload) => {
// kind+session can arrive here twice. Collapse it at this single choke point.
// Return true (not false): a notification for the event IS being shown by the
// first caller, so the settings "send test" success probe stays honest.
if (isDuplicateNotification(`${payload?.kind ?? ''}:${payload?.sessionId ?? ''}`)) {
if (isDuplicateNotification(`${payload?.kind ?? ''}:${payload?.sessionId ?? payload?.tag ?? ''}`)) {
return true
}
@@ -10510,6 +10510,22 @@ ipcMain.handle('hermes:writeClipboard', (_event, text) => {
return true
})
// Native save-location picker (profile export etc.) — the write itself happens
// elsewhere (the backend, for profile archives); this only picks the path.
ipcMain.handle('hermes:selectSavePath', async (_event, options: any = {}) => {
const result = await dialog.showSaveDialog(mainWindow, {
title: options?.title || 'Save',
defaultPath: options?.defaultPath ? String(options.defaultPath) : undefined,
filters: Array.isArray(options?.filters) ? options.filters : undefined
})
if (result.canceled || !result.filePath) {
return null
}
return result.filePath
})
// Paired reader for the GUI terminal's paste chord: the renderer's
// navigator.clipboard.readText() throws "Document is not focused" whenever a
// portaled overlay has focus, and there's no way to route a read through the
+1
View File
@@ -115,6 +115,7 @@ contextBridge.exposeInMainWorld('hermesDesktop', {
},
readFileText: filePath => ipcRenderer.invoke('hermes:readFileText', filePath),
selectPaths: options => ipcRenderer.invoke('hermes:selectPaths', options),
selectSavePath: options => ipcRenderer.invoke('hermes:selectSavePath', options),
writeClipboard: text => ipcRenderer.invoke('hermes:writeClipboard', text),
readClipboard: () => ipcRenderer.invoke('hermes:readClipboard'),
saveImageFromUrl: url => ipcRenderer.invoke('hermes:saveImageFromUrl', url),
+34 -3
View File
@@ -8,6 +8,7 @@ import { useLocation } from 'react-router'
import type { SubmitTextOptions } from '@/app/session/hooks/use-prompt-actions/utils'
import { Thread } from '@/components/assistant-ui/thread'
import { TranscriptWindowProvider } from '@/components/assistant-ui/thread/transcript-window'
import { Backdrop } from '@/components/Backdrop'
import { COMPOSER_HEART_CONFIG, HeartField } from '@/components/chat/vibe-hearts'
import { usePaneVisible } from '@/components/pane-shell/pane-visibility'
@@ -63,6 +64,7 @@ import { ScrollToBottomButton } from './scroll-to-bottom-button'
import { useSessionView } from './session-view'
import { SessionActionsMenu } from './sidebar/session-actions-menu'
import { threadLoadingState } from './thread-loading'
import { selectTranscriptWindow } from './transcript-window'
interface ChatViewProps extends Omit<React.ComponentProps<'div'>, 'onSubmit'> {
gateway: HermesGateway | null
@@ -220,9 +222,34 @@ function ChatRuntimeBoundary({
onThreadMessagesChange,
suppressMessages
}: ChatRuntimeBoundaryProps) {
const storeMessages = useMessagesWhileVisible(useSessionView().$messages)
const view = useSessionView()
const runtimeId = useStore(view.$runtimeId)
const storeMessages = useMessagesWhileVisible(view.$messages)
const messages = suppressMessages ? NO_MESSAGES : storeMessages
const runtimeMessageRepository = useRuntimeMessageRepository(messages)
const [windowPages, setWindowPages] = useState(1)
const [windowSessionKey, setWindowSessionKey] = useState(runtimeId)
// Reset the window on session swap during RENDER, so a large expand from the
// previous chat can't leak into the next one's first paint (#55191).
if (windowSessionKey !== runtimeId) {
setWindowSessionKey(runtimeId)
setWindowPages(1)
}
const { messages: windowedMessages, windowed } = useMemo(
() => selectTranscriptWindow(messages, windowPages),
[messages, windowPages]
)
const runtimeMessageRepository = useRuntimeMessageRepository(windowedMessages)
const expandWindow = useCallback(() => setWindowPages(pages => pages + 1), [])
const transcriptWindow = useMemo(
() => ({ olderAvailable: windowed, expandWindow }),
[expandWindow, windowed]
)
const runtime = useIncrementalExternalStoreRuntime<ThreadMessage>({
messageRepository: runtimeMessageRepository,
@@ -237,7 +264,11 @@ function ChatRuntimeBoundary({
onReload
})
return <AssistantRuntimeProvider runtime={runtime}>{children}</AssistantRuntimeProvider>
return (
<TranscriptWindowProvider value={transcriptWindow}>
<AssistantRuntimeProvider runtime={runtime}>{children}</AssistantRuntimeProvider>
</TranscriptWindowProvider>
)
}
// Memoized: the tile caller (session-tile.tsx) and the contrib surface re-render
@@ -51,4 +51,32 @@ describe('useRuntimeMessageRepository', () => {
expect(feedToRepository(result.current).map(item => item.id)).toEqual(['user-1', 'assistant-stream-1', 'user-2'])
})
it('anchors a branch group to its fork point, and a windowed cut keeps it', () => {
// Branch groups record their fork parent the first time they are seen. A
// window that started mid-group would anchor the survivors to whatever
// preceded them instead — selectTranscriptWindow aligns the cut so the
// whole group arrives together (#55191).
const branch = (id: string): ChatMessage => ({
...text(id, 'assistant', 'branch'),
branchGroupId: 'group-1'
})
const messages = [text('user-1', 'user', 'hi'), branch('a-1'), branch('a-2'), text('user-2', 'user', 'more')]
const { result } = renderHook(() => useRuntimeMessageRepository(messages))
const parents = new Map(result.current.messages.map(item => [item.message.id, item.parentId]))
expect(parents.get('a-1')).toBe('user-1')
expect(parents.get('a-2')).toBe('user-1')
// The same group fed as a window that begins AT the group start keeps the
// fork intact (parent becomes null: the group is now the transcript root).
const { result: windowed } = renderHook(() => useRuntimeMessageRepository(messages.slice(1)))
const windowedParents = new Map(windowed.current.messages.map(item => [item.message.id, item.parentId]))
expect(windowedParents.get('a-1')).toBe(windowedParents.get('a-2'))
})
})
@@ -59,6 +59,7 @@ import {
setShowAllProfiles,
sortByProfileOrder
} from '@/store/profile'
import { runExportProfileFlow, runImportProfileFlow } from '@/store/profile-share'
import type { ProfileInfo } from '@/types/hermes'
import { CreateProfileDialog } from '../../profiles/create-profile-dialog'
@@ -264,6 +265,7 @@ export function ProfileRail() {
profiles={named}
/>
<AddProfileButton label={p.newProfile} onClick={() => setCreateOpen(true)} />
<ImportProfileButton label={p.importProfile} />
</div>
) : (
<div
@@ -302,6 +304,7 @@ export function ProfileRail() {
)}
<AddProfileButton label={p.newProfile} onClick={() => setCreateOpen(true)} />
<ImportProfileButton label={p.importProfile} />
</div>
)}
@@ -435,6 +438,24 @@ function AddProfileButton({ label, onClick }: { label: string; onClick: () => vo
)
}
// Import-archive door beside the "+": adopt a shared profile bundle (theme,
// skills, layout) as a new profile. Same chrome as AddProfileButton; the whole
// flow (picker → import → apply overlay → switch) lives in the store.
function ImportProfileButton({ label }: { label: string }) {
return (
<Tip label={label}>
<button
aria-label={label}
className="grid size-5 shrink-0 place-items-center rounded-[3px] text-(--ui-text-tertiary) opacity-55 transition hover:bg-(--ui-control-hover-background) hover:text-foreground hover:opacity-100"
onClick={() => void runImportProfileFlow()}
type="button"
>
<Codicon name="cloud-download" size="0.75rem" />
</button>
</Tip>
)
}
// The condensed rail: every named profile in one compact select. The trigger
// shows the active profile (tinted initial + name); on default/all scope it
// falls back to the placeholder since the left toggle pill carries that state.
@@ -692,6 +713,10 @@ function ProfileSquare({
<Codicon name="edit" size="0.875rem" />
<span>{p.editSoul}</span>
</ContextMenuItem>
<ContextMenuItem onSelect={() => void runExportProfileFlow(label)}>
<Codicon name="package" size="0.875rem" />
<span>{p.exportProfile}</span>
</ContextMenuItem>
<ContextMenuItem
className="text-destructive focus:text-destructive"
onSelect={onDelete}
@@ -22,7 +22,14 @@ vi.mock('@/i18n', () => ({
t: {
common: { cancel: 'Cancel', close: 'Close', delete: 'Delete', save: 'Save' },
sidebar: {
projects: { menuAppearance: 'Appearance', noColor: 'No color' },
projects: {
menuAppearance: 'Appearance',
moveFailed: 'Could not move session',
moveNoProjects: 'No other projects',
movedTo: (name: string) => `Moved to ${name}`,
moveToProject: 'Move to project',
noColor: 'No color'
},
row: {
archive: 'Archive',
branchFrom: 'Branch from here',
@@ -50,6 +57,12 @@ vi.mock('@/lib/profile-color', () => ({ PROFILE_SWATCHES: [] }))
vi.mock('@/lib/session-export', () => ({ exportSession: vi.fn() }))
vi.mock('@/store/gateway', () => ({ activeGateway: vi.fn(() => null) }))
vi.mock('@/store/notifications', () => ({ notify: vi.fn(), notifyError: vi.fn() }))
vi.mock('@/store/projects', () => ({
$projectTree: atom<unknown[]>([]),
moveSessionToProject: vi.fn(),
projectIdForCwd: vi.fn(() => null),
projectRootCwd: vi.fn(() => '')
}))
vi.mock('@/store/session', () => ({
$activeSessionId: atom<null | string>(null),
$selectedStoredSessionId: atom<null | string>(null),
@@ -30,6 +30,7 @@ import { PROFILE_SWATCHES } from '@/lib/profile-color'
import { exportSession } from '@/lib/session-export'
import { activeGateway } from '@/store/gateway'
import { notify, notifyError } from '@/store/notifications'
import { $projectTree, moveSessionToProject, projectIdForCwd, projectRootCwd } from '@/store/projects'
import {
$activeSessionId,
$selectedStoredSessionId,
@@ -133,6 +134,44 @@ function SessionColorSwatches({ sessionId }: { sessionId: string }) {
)
}
// The project list inside the session menu's "Move to project" submenu. Its own
// component so only an OPEN submenu subscribes to the stores (same reasoning as
// SessionColorSwatches). Re-homes the session's workspace at the target
// project's root — the fix for a chat created in the wrong folder. The current
// owner and folderless projects (the Home bucket) are excluded: there is
// nothing to move into.
function MoveToProjectItems({ kit, sessionId, profile }: { kit: MenuKit; sessionId: string; profile?: string }) {
const { t } = useI18n()
const p = t.sidebar.projects
const tree = useStore($projectTree)
const session = useStore($sessions).find(s => sessionMatchesStoredId(s, sessionId))
const cwd = session?.cwd?.trim() || ''
const currentProjectId = cwd ? projectIdForCwd(cwd) : null
const targets = tree.filter(node => node.id !== currentProjectId && !node.isNoProject && projectRootCwd(node))
if (targets.length === 0) {
return <kit.Item disabled>{p.moveNoProjects}</kit.Item>
}
return (
<>
{targets.map(node => (
<kit.Item
key={node.id}
onSelect={() => {
triggerHaptic('selection')
moveSessionToProject(sessionId, node.id, profile)
.then(() => notify({ durationMs: 2_000, kind: 'success', message: p.movedTo(node.label) }))
.catch(err => notifyError(err, p.moveFailed))
}}
>
{node.label}
</kit.Item>
))}
</>
)
}
function useSessionActions({
sessionId,
title,
@@ -355,6 +394,15 @@ function useSessionActions({
/>
<kit.Separator />
{workItems.map(item => renderActionItem(kit, item))}
<kit.Sub>
<kit.SubTrigger disabled={!sessionId}>
<Codicon name="folder" size="0.875rem" />
<span>{t.sidebar.projects.moveToProject}</span>
</kit.SubTrigger>
<kit.SubContent>
<MoveToProjectItems kit={kit} profile={profile} sessionId={sessionId} />
</kit.SubContent>
</kit.Sub>
{tabItems.length > 0 && (
<>
<kit.Separator />
@@ -0,0 +1,144 @@
import { describe, expect, it } from 'vitest'
import type { ChatMessage } from '@/lib/chat-messages'
import { RENDER_WEIGHT_CHARS } from '@/lib/render-weight'
import {
alignToBranchGroup,
selectTranscriptWindow,
TRANSCRIPT_WINDOW_BUDGET,
TRANSCRIPT_WINDOW_MIN_MESSAGES
} from './transcript-window'
const message = (id: string, chars: number, branchGroupId?: string): ChatMessage => ({
id,
parts: [{ type: 'text', text: 'x'.repeat(chars) }],
role: id.startsWith('u') ? 'user' : 'assistant',
...(branchGroupId ? { branchGroupId } : {})
})
/** Messages of `chars` each, newest last. */
const transcript = (count: number, chars: number): ChatMessage[] =>
Array.from({ length: count }, (_, i) => message(`m-${i}`, chars))
describe('selectTranscriptWindow', () => {
it('does not window a transcript that fits the budget', () => {
const messages = transcript(50, 100)
const window = selectTranscriptWindow(messages)
expect(window.windowed).toBe(false)
// Reference identity preserved — a fresh array would re-render the runtime.
expect(window.messages).toBe(messages)
})
it('windows a HEAVY-but-SHORT transcript that a message-count cap would miss', () => {
// 40 messages, each a big tool result. Well under any sane count cap, but
// this is the shape that exhausts the renderer heap (#55191).
const messages = transcript(40, RENDER_WEIGHT_CHARS * 400)
const window = selectTranscriptWindow(messages)
expect(window.windowed).toBe(true)
expect(window.messages.length).toBeLessThan(messages.length)
expect(window.messages.at(-1)).toBe(messages.at(-1))
})
it('keeps far MORE messages when they are light than when they are heavy', () => {
// The contract is weight, not count: a message-count cap would treat these
// two identically. 500 tiny messages are cheaper than 500 tool results, so
// many more of them survive the same budget.
const light = selectTranscriptWindow(transcript(500, 20))
const heavy = selectTranscriptWindow(transcript(500, RENDER_WEIGHT_CHARS * 40))
expect(light.messages.length).toBeGreaterThan(heavy.messages.length * 10)
})
it('leaves a long transcript whole when the whole thing is cheap', () => {
const messages = transcript(600, 20)
const window = selectTranscriptWindow(messages)
expect(window.windowed).toBe(false)
expect(window.messages).toBe(messages)
})
it('keeps a floor of messages when single turns are enormous', () => {
const messages = transcript(80, RENDER_WEIGHT_CHARS * TRANSCRIPT_WINDOW_BUDGET)
const window = selectTranscriptWindow(messages)
expect(window.messages.length).toBeGreaterThanOrEqual(TRANSCRIPT_WINDOW_MIN_MESSAGES)
})
it('grows by one budget page per expand and eventually covers everything', () => {
const messages = transcript(400, RENDER_WEIGHT_CHARS * 40)
const first = selectTranscriptWindow(messages, 1)
const second = selectTranscriptWindow(messages, 2)
expect(first.windowed).toBe(true)
expect(second.messages.length).toBeGreaterThan(first.messages.length)
let pages = 1
let window = selectTranscriptWindow(messages, pages)
while (window.windowed && pages < 100) {
window = selectTranscriptWindow(messages, ++pages)
}
// Paging terminates at the full transcript — never a dead end.
expect(window.windowed).toBe(false)
expect(window.messages).toHaveLength(messages.length)
})
it('never cuts inside a branch group, so branches keep their fork point', () => {
const heavy = RENDER_WEIGHT_CHARS * 200
// A branch group sits right where a weight-only cut would land.
const messages: ChatMessage[] = [
...transcript(20, heavy),
message('a-branch-1', heavy, 'group-1'),
message('a-branch-2', heavy, 'group-1'),
message('a-branch-3', heavy, 'group-1'),
...transcript(20, heavy).map(m => ({ ...m, id: `tail-${m.id}` }))
]
for (let pages = 1; pages <= 6; pages++) {
const kept = selectTranscriptWindow(messages, pages).messages
const groupMembers = kept.filter(m => m.branchGroupId === 'group-1')
// Either the whole group survives or none of it does — never a partial
// group, which would re-parent the surviving branches.
expect([0, 3]).toContain(groupMembers.length)
}
})
it('handles an empty transcript', () => {
const messages: ChatMessage[] = []
expect(selectTranscriptWindow(messages)).toEqual({ messages, windowed: false })
})
})
describe('alignToBranchGroup', () => {
const messages = [
message('u-1', 10),
message('a-1', 10, 'g'),
message('a-2', 10, 'g'),
message('u-2', 10)
]
it('widens a cut that lands mid-group back to the group start', () => {
expect(alignToBranchGroup(messages, 2)).toBe(1)
})
it('leaves a cut on a non-branch message alone', () => {
expect(alignToBranchGroup(messages, 3)).toBe(3)
})
it('clamps out-of-range indices', () => {
expect(alignToBranchGroup(messages, -5)).toBe(0)
expect(alignToBranchGroup(messages, 99)).toBe(messages.length)
})
})
@@ -0,0 +1,111 @@
import type { ChatMessage } from '@/lib/chat-messages'
import { messageRenderWeight } from '@/lib/render-weight'
/**
* Bound what reaches assistant-ui at all.
*
* Rendering the full transcript of an oversized session rebuilds an unbounded
* runtime repository on every store update and exhausts the renderer's V8 heap
* (#55191). The DOM budget in `thread/list.tsx` already bounds what PAINTS, but
* every message still gets normalized into the repository first — so a session
* only has to be heavy, not visible, to crash the window.
*
* The window spends the same currency as the DOM budget: render weight, not
* message count. A count cap gets this wrong in both directions — measured on a
* real 1,175-session store, a 400-message cap would disable itself on 37
* sessions that are heavy but short (one was 133 messages / 1.05MB) while
* engaging on 92 long-but-light sessions that were never at risk.
*/
/**
* One window page, in render-weight units.
*
* Four DOM pages (the `RENDER_BUDGET` of 300 in `thread/list.tsx`). "Show
* earlier" spends the DOM budget first, so the user pages through the
* already-materialized window three times before this asks the store for more
* — and the reported crash shape (~231K tokens ≈ 2,260 units) is windowed
* rather than handed to the repository whole.
*/
export const TRANSCRIPT_WINDOW_BUDGET = 1200
/**
* Floor on messages kept regardless of weight. A transcript of enormous turns
* must still render the turn the user is having; without this a single
* multi-megabyte tool result could window everything after it away.
*/
export const TRANSCRIPT_WINDOW_MIN_MESSAGES = 30
export interface TranscriptWindow {
/** The tail assistant-ui is allowed to materialize. */
messages: ChatMessage[]
/** Store holds older messages than this window. */
windowed: boolean
}
/**
* Widen a cut backwards so it never lands inside an assistant branch group.
*
* `useRuntimeMessageRepository` records a group's fork point the first time it
* sees the group (`branchParentByGroup`). A cut through the middle of a group
* therefore anchors the surviving branches to whatever message happens to
* precede them in the window — silently re-parenting a branch. Include the
* whole group or none of it.
*/
export function alignToBranchGroup(messages: readonly ChatMessage[], start: number): number {
if (start <= 0 || start >= messages.length) {
return Math.max(0, Math.min(start, messages.length))
}
const group = messages[start].branchGroupId
if (!group) {
return start
}
let aligned = start
while (aligned > 0 && messages[aligned - 1].branchGroupId === group) {
aligned--
}
return aligned
}
/**
* Select the tail of the transcript that fits one window, grown by `pages`.
*
* Walks newest-first accumulating weight until the budget is met, keeps at
* least MIN messages, then aligns the cut off a branch-group boundary.
*/
export function selectTranscriptWindow(
messages: readonly ChatMessage[],
pages = 1
): TranscriptWindow {
const budget = TRANSCRIPT_WINDOW_BUDGET * Math.max(1, Math.floor(pages))
if (messages.length === 0) {
return { messages: messages as ChatMessage[], windowed: false }
}
let start = messages.length
let weight = 0
for (let i = messages.length - 1; i >= 0; i--) {
weight += messageRenderWeight(messages[i].parts)
start = i
if (weight >= budget && messages.length - i >= TRANSCRIPT_WINDOW_MIN_MESSAGES) {
break
}
}
start = alignToBranchGroup(messages, start)
if (start <= 0) {
// Preserve reference identity when the whole transcript fits: handing React
// a fresh array of the same messages re-renders the runtime for nothing.
return { messages: messages as ChatMessage[], windowed: false }
}
return { messages: messages.slice(start), windowed: true }
}
+27 -1
View File
@@ -39,7 +39,7 @@ import { useContributions } from '@/contrib/react/use-contributions'
import { registry } from '@/contrib/registry'
import { discoverRuntimePlugins } from '@/contrib/runtime-loader'
import { sessionTitle as storedSessionTitle } from '@/lib/chat-runtime'
import { FileText, LayoutDashboard, PanelBottom, Terminal, Zap } from '@/lib/icons'
import { Download, FileText, LayoutDashboard, PanelBottom, Terminal, Upload, Zap } from '@/lib/icons'
import { type KeybindContribution, KEYBINDS_AREA } from '@/lib/keybinds/actions'
import { setYoloEnabled } from '@/lib/yolo-session'
import { pruneComposerPopoutZones } from '@/store/composer-popout'
@@ -56,6 +56,7 @@ import {
SIDEBAR_MAX_WIDTH
} from '@/store/layout'
import { $previewOpenRequest, $previewTabs, closeRightRail } from '@/store/preview'
import { runExportProfileFlow, runImportProfileFlow } from '@/store/profile-share'
import { $reviewOpen, closeReview, openReview, REVIEW_PANE_ID } from '@/store/review'
import { $currentCwd, $selectedStoredSessionId, $sessions, $yoloActive, sessionMatchesStoredId } from '@/store/session'
import { watchSessionPins } from '@/store/session-pin-sync'
@@ -314,6 +315,31 @@ registry.registerMany([
keywords: ['keybinds', 'shortcuts', 'hotkeys', 'keyboard'],
run: () => window.dispatchEvent(new CustomEvent('hermes:open-keybinds'))
} satisfies PaletteContribution
},
// Profile sharing: bundle the active profile (config, skills, theme, layout)
// into a portable archive, or adopt someone else's. Both open native dialogs,
// so the palette closing on select is correct.
{
id: 'profile.export',
area: PALETTE_AREA,
data: {
id: 'profile.export',
label: 'Export profile…',
icon: Upload,
keywords: ['profile', 'export', 'share', 'bundle', 'theme', 'settings', 'backup'],
run: () => void runExportProfileFlow()
} satisfies PaletteContribution
},
{
id: 'profile.import',
area: PALETTE_AREA,
data: {
id: 'profile.import',
label: 'Import profile…',
icon: Download,
keywords: ['profile', 'import', 'share', 'bundle', 'archive', 'restore'],
run: () => void runImportProfileFlow()
} satisfies PaletteContribution
}
])
+37 -6
View File
@@ -42,7 +42,7 @@ import {
} from '@/store/profile'
import { openFolderAsProject, requestNewWorktree } from '@/store/projects'
import { toggleReview } from '@/store/review'
import { setModelPickerOpen } from '@/store/session'
import { $selectedStoredSessionId, setModelPickerOpen } from '@/store/session'
import { reopenLastClosedTile } from '@/store/session-states'
import {
$switcherOpen,
@@ -62,11 +62,13 @@ import { useTheme } from '@/themes/context'
import { requestComposerFocus, requestModelMenuToggle, requestVoiceToggle } from '../chat/composer/focus'
import { openSession } from '../open-session'
import {
$workspaceIsPage,
AGENTS_ROUTE,
ARTIFACTS_ROUTE,
CRON_ROUTE,
MESSAGING_ROUTE,
navigateToWorkspacePage,
NEW_CHAT_ROUTE,
PROFILES_ROUTE,
sessionRoute,
SETTINGS_ROUTE,
@@ -99,11 +101,29 @@ export function useKeybinds(deps: KeybindRuntimeDeps): void {
const profileSwitchHandlers: HandlerMap = {}
// A tab key that lands on the WORKSPACE tab while a full page (skills /
// messaging / artifacts / a plugin route) covers it must also route back to
// the chat: the workspace pane is already the zone's active tab behind the
// page, so fronting it alone changes nothing on screen and the key reads
// dead. Mirrors `openSession`'s full-page rule — only a route change puts
// the chat back.
const leavePageForWorkspaceChat = (paneId: null | string) => {
if (paneId === 'workspace' && $workspaceIsPage.get()) {
const selected = $selectedStoredSessionId.get()
navigate(selected ? sessionRoute(selected) : NEW_CHAT_ROUTE)
}
}
for (let slot = 1; slot <= PROFILE_SLOT_COUNT; slot += 1) {
// ⌘1…⌘9 switch the FOCUSED zone's tab when it's a real tab strip; only a
// single-pane (or unfocused) layout falls through to the profile switch.
profileSwitchHandlers[`profile.switch.${slot}`] = () => {
if (!activateTreeTabSlot(slot)) {
const pane = activateTreeTabSlot(slot)
if (pane) {
leavePageForWorkspaceChat(pane)
} else {
switchProfileToSlot(slot)
}
}
@@ -132,6 +152,19 @@ export function useKeybinds(deps: KeybindRuntimeDeps): void {
goToSession(openOrAdvanceSwitcher(direction))
}
// ⌃Tab cycles the focused session/main tab strip; only a non-tabbed focus
// falls through to the recent-session switcher. Landing on the workspace
// under a full page routes back to the chat (same as ⌘1).
const cycleTab = (direction: 1 | -1) => {
const pane = cycleTreeTabInFocusedZone(direction)
if (pane) {
leavePageForWorkspaceChat(pane)
} else {
stepSession(direction)
}
}
const showFiles = () => {
setFileBrowserOpen(true)
setTerminalTakeover(false)
@@ -170,10 +203,8 @@ export function useKeybinds(deps: KeybindRuntimeDeps): void {
},
'session.newTab': () => deps.openNewSessionTab(),
'session.newWindow': () => void openNewWindow(),
// ⌃Tab cycles the focused session/main tab strip; only a non-tabbed focus
// falls through to the recent-session switcher.
'session.next': () => void (cycleTreeTabInFocusedZone(1) || stepSession(1)),
'session.prev': () => void (cycleTreeTabInFocusedZone(-1) || stepSession(-1)),
'session.next': () => cycleTab(1),
'session.prev': () => cycleTab(-1),
...sessionSlotHandlers,
'session.focusSearch': requestSessionSearchFocus,
'session.togglePin': deps.toggleSelectedPin,
@@ -1,5 +1,7 @@
import { describe, expect, it } from 'vitest'
import { messageRenderWeight, RENDER_WEIGHT_CHARS } from '@/lib/render-weight'
import {
buildGroups,
firstVisibleGroupIndex,
@@ -7,8 +9,6 @@ import {
LIVE_TAIL_PARTS,
liveTailStart,
type MessageGroup,
messageRenderWeight,
RENDER_WEIGHT_CHARS,
resolveThreadScrollTarget
} from './list'
@@ -16,6 +16,7 @@ import {
import { type GetTargetScrollTop, useStickToBottom } from 'use-stick-to-bottom'
import { useI18n } from '@/i18n'
import { messageRenderWeight } from '@/lib/render-weight'
import { cn } from '@/lib/utils'
import {
onScrollToBottomRequest,
@@ -28,6 +29,8 @@ import { isSecondaryWindow } from '@/store/windows'
import { MessageRenderBoundary } from '../message-render-boundary'
import { resolveShowEarlierAction, useTranscriptWindow } from './transcript-window'
type ThreadMessageComponents = ComponentProps<typeof ThreadPrimitive.MessageByIndex>['components']
export type MessageGroup = { id: string; weight: number } & (
@@ -47,8 +50,6 @@ export type MessageGroup = { id: string; weight: number } & (
// a virtualizer — pure rendering, never touches scrollTop, so it can't fight
// use-stick-to-bottom (the single scroll owner).
const RENDER_BUDGET = 300
export const RENDER_WEIGHT_CHARS = 512
const MAX_MEASURED_MESSAGE_CHARS = RENDER_BUDGET * RENDER_WEIGHT_CHARS
// On session switch, paint a small budget first (enough for the bottom turn(s)
// the user actually sees after scroll-to-bottom), then bump to the full budget
// in a requestAnimationFrame — defers the heavy markdown+syntax-highlight render
@@ -76,70 +77,6 @@ export const resolveThreadScrollTarget: GetTargetScrollTop = (targetScrollTop, {
return remaining >= 0 && remaining <= SCROLL_TARGET_EPSILON_PX ? currentScrollTop : targetScrollTop
}
const contentWeightCache = new WeakMap<object, number>()
const NON_RENDERED_CONTENT_FIELDS = new Set(['id', 'role', 'toolCallId', 'toolName', 'type'])
/**
* Estimate the synchronous renderer cost of one assistant-ui message.
*
* The traversal is capped once a single message has enough text to consume a
* complete render page. Going further cannot affect which whole turn crosses
* the budget, and avoiding an unbounded walk matters for deeply nested tool
* payloads. A WeakMap keeps settled history O(message count) on later store
* updates; assistant-ui publishes a new content array when a streaming message
* changes, so the live tail still receives a fresh weight.
*/
export function messageRenderWeight(content: unknown): number {
if (!Array.isArray(content)) {
return 1
}
const cached = contentWeightCache.get(content)
if (cached !== undefined) {
return cached
}
const seen = new WeakSet<object>()
const pending: unknown[] = [...content]
let characters = 0
while (pending.length > 0 && characters < MAX_MEASURED_MESSAGE_CHARS) {
const value = pending.pop()
if (typeof value === 'string') {
characters += Math.min(value.length, MAX_MEASURED_MESSAGE_CHARS - characters)
continue
}
if (!value || typeof value !== 'object' || seen.has(value)) {
continue
}
seen.add(value)
if (Array.isArray(value)) {
for (const nested of value) {
pending.push(nested)
}
continue
}
for (const [key, nested] of Object.entries(value)) {
if (!NON_RENDERED_CONTENT_FIELDS.has(key)) {
pending.push(nested)
}
}
}
const weight = Math.max(1, content.length) + Math.ceil(characters / RENDER_WEIGHT_CHARS)
contentWeightCache.set(content, weight)
return weight
}
interface ThreadMessageListProps {
clampToComposer: boolean
components: ThreadMessageComponents
@@ -315,6 +252,8 @@ const ThreadMessageListInner: FC<ThreadMessageListProps> = ({
targetScrollTop: resolveThreadScrollTarget
})
const { olderAvailable, expandWindow } = useTranscriptWindow()
const [renderBudget, setRenderBudget] = useState(FIRST_PAINT_BUDGET)
// Cut the budget during RENDER, not in the post-commit layout effect. An
@@ -540,11 +479,26 @@ const ThreadMessageListInner: FC<ThreadMessageListProps> = ({
// Prepend an older page while preserving the on-screen position. The user is
// scrolled up (reading history) so the stick-to-bottom lock is escaped and
// won't fight this manual restore.
// won't fight this manual restore. Spend the already-materialized DOM page
// first; only when that is exhausted pull more messages out of the session
// store (#55191).
const showEarlier = useCallback(() => {
const action = resolveShowEarlierAction(hiddenCount, olderAvailable)
if (!action) {
return
}
anchorBeforePrepend()
setRenderBudget(budget => budget + RENDER_BUDGET)
}, [anchorBeforePrepend])
if (action === 'dom') {
setRenderBudget(budget => budget + RENDER_BUDGET)
return
}
expandWindow()
}, [anchorBeforePrepend, expandWindow, hiddenCount, olderAvailable])
useLayoutEffect(() => {
const el = scrollRef.current
@@ -553,7 +507,8 @@ const ThreadMessageListInner: FC<ThreadMessageListProps> = ({
el.scrollTop = el.scrollHeight - restoreFromBottomRef.current
restoreFromBottomRef.current = null
}
}, [scrollRef, renderBudget])
// renderBudget covers DOM pages; groups.length covers store-window expands.
}, [scrollRef, renderBudget, groups.length])
// The row array is memoized on the inputs the rows actually read. This
// component re-renders on every isAtBottom flip — and use-stick-to-bottom
@@ -647,7 +602,7 @@ const ThreadMessageListInner: FC<ThreadMessageListProps> = ({
data-slot="aui_thread-content"
ref={contentRef as React.RefCallback<HTMLDivElement>}
>
{hiddenCount > 0 && (
{(hiddenCount > 0 || olderAvailable) && (
<button
className="mx-auto mb-(--conversation-turn-gap) rounded-full border border-border/65 bg-(--composer-fill) px-3 py-1 text-xs text-muted-foreground hover:text-foreground"
onClick={showEarlier}
@@ -0,0 +1,18 @@
import { describe, expect, it } from 'vitest'
import { resolveShowEarlierAction } from './transcript-window'
describe('resolveShowEarlierAction', () => {
it('spends the already-materialized DOM page first', () => {
expect(resolveShowEarlierAction(3, true)).toBe('dom')
expect(resolveShowEarlierAction(3, false)).toBe('dom')
})
it('expands the store window once the DOM page is exhausted', () => {
expect(resolveShowEarlierAction(0, true)).toBe('window')
})
it('is a no-op when neither DOM nor store has older content', () => {
expect(resolveShowEarlierAction(0, false)).toBe(null)
})
})
@@ -0,0 +1,43 @@
import { createContext, type ReactNode, useContext } from 'react'
export interface TranscriptWindowValue {
/** Store holds older messages the runtime window has not materialized. */
olderAvailable: boolean
/** Pull one more page of older messages out of the session store. */
expandWindow: () => void
}
const TranscriptWindowContext = createContext<TranscriptWindowValue>({
olderAvailable: false,
expandWindow: () => {}
})
export function TranscriptWindowProvider({
children,
value
}: {
children: ReactNode
value: TranscriptWindowValue
}) {
return <TranscriptWindowContext.Provider value={value}>{children}</TranscriptWindowContext.Provider>
}
export function useTranscriptWindow(): TranscriptWindowValue {
return useContext(TranscriptWindowContext)
}
/**
* "Show earlier" pages the DOM budget first and only then asks the store for
* more messages — the DOM page is already-materialized content, so spending it
* first keeps the click cheap and the store window as small as it can be.
*/
export function resolveShowEarlierAction(
hiddenCount: number,
olderAvailable: boolean
): 'dom' | 'window' | null {
if (hiddenCount > 0) {
return 'dom'
}
return olderAvailable ? 'window' : null
}
@@ -48,14 +48,14 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteActiveTreeGroup('grp-main')
tree.noteHoveredTreeGroup('grp-side')
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-side')).toBe('session-tile:c')
// The focused zone is untouched — the pointer won the target, not both.
expect(activeOf('grp-main')).toBe('workspace')
// Same key, pointer moved: the other zone's slot 2.
tree.noteHoveredTreeGroup('grp-main')
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-main')).toBe('session-tile:a')
})
@@ -66,7 +66,7 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteHoveredTreeGroup('grp-main')
tree.noteHoveredTreeGroup(null)
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-side')).toBe('session-tile:c')
expect(activeOf('grp-main')).toBe('workspace')
})
@@ -77,7 +77,7 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteActiveTreeGroup('grp-main')
tree.noteHoveredTreeGroup('grp-side')
expect(tree.cycleTreeTabInFocusedZone(1)).toBe(true)
expect(tree.cycleTreeTabInFocusedZone(1)).toBeTruthy()
expect(activeOf('grp-side')).toBe('session-tile:c')
expect(activeOf('grp-main')).toBe('workspace')
@@ -98,9 +98,9 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteActiveTreeGroup('grp-side')
tree.noteHoveredTreeGroup(null)
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-side')).toBe('session-tile:c')
expect(tree.cycleTreeTabInFocusedZone(1)).toBe(true)
expect(tree.cycleTreeTabInFocusedZone(1)).toBeTruthy()
expect(activeOf('grp-side')).toBe('session-tile:b')
})
@@ -112,7 +112,7 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteActiveTreeGroup(null)
tree.noteHoveredTreeGroup(null)
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-main')).toBe('session-tile:a')
expect(activeOf('grp-side')).toBe('session-tile:b')
})
@@ -138,7 +138,7 @@ describe('hovered zone retargets the tab verbs', () => {
tree.noteActiveTreeGroup(null)
tree.noteHoveredTreeGroup('grp-files')
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBeTruthy()
expect(activeOf('grp-main')).toBe('session-tile:a')
// ⌘W must not close the file tree from a rung that can't serve it.
expect(activeOf('grp-files')).toBe('files')
@@ -317,6 +317,66 @@ export function movePane(
return shapeSignature(next) === shapeSignature(root) ? root : next
}
/**
* Move a SELECTION of panes together (multi-tab drag), preserving their strip
* order. The lead pane lands exactly like a single `movePane` (center joins at
* `before`, an edge opens the split); the rest stack in behind it. `activeId`
* (the pressed tab) fronts in the landing group. Same no-op guard as
* `movePane`: a drop that rebuilds the visible arrangement returns `root`.
*/
export function movePanes(
root: LayoutNode,
paneIds: readonly string[],
target: { groupId: string; pos: DropPosition; before?: null | string },
activeId: string = paneIds[0] ?? ''
): LayoutNode {
if (paneIds.length <= 1) {
return paneIds.length === 1 ? movePane(root, paneIds[0], target) : root
}
let without: LayoutNode | null = root
for (const id of paneIds) {
without = without && removePane(without, id)
}
// The selection was the whole tree, or removal dissolved the target zone
// (the selection was its only occupancy) — nowhere left to land.
if (!without || !findGroup(without, target.groupId)) {
return root
}
// The lead insert decides geometry; the rest stack into the lead's group at
// the same slot (each lands before `before`, so the block keeps its order).
// Only the lead activates — `insertAtGroup(activate)` would otherwise front
// each follower in turn.
const lead = paneIds[0]
let next: LayoutNode | null = insertAtGroup(without, target.groupId, lead, target.pos, target.before)
for (let i = 1; next && i < paneIds.length; i++) {
const leadGroup = findGroupOfPane(next, lead)
if (!leadGroup) {
return root
}
const before = target.pos === 'center' ? (target.before ?? null) : null
next = insertAtGroup(next, leadGroup.id, paneIds[i], 'center', before, false)
}
if (!next) {
return root
}
const landed = findGroupOfPane(next, lead)
if (landed && landed.panes.includes(activeId)) {
next = setActivePane(next, landed.id, activeId)
}
return shapeSignature(next) === shapeSignature(root) ? root : next
}
/** Group ids of every leaf under a node, in tree order. */
export function groupLeafIds(node: LayoutNode): string[] {
return node.type === 'group' ? [node.id] : node.children.flatMap(groupLeafIds)
@@ -347,26 +407,32 @@ function findCover(node: LayoutNode, set: Set<string>): LayoutNode | null {
}
/**
* FancyZones span: merge the highlighted zones into ONE group holding
* `paneId`, absorbing any panes that lived in those zones as tabs. Only works
* when the highlighted set forms a rectangular subtree (it always does for a
* combined zone range on a guillotine tree); returns null otherwise so the
* caller can fall back to a single-zone drop.
* FancyZones span: merge the highlighted zones into ONE group holding the
* dragged pane block (one pane, or a multi-tab selection in strip order),
* absorbing any panes that lived in those zones as tabs. Only works when the
* highlighted set forms a rectangular subtree (it always does for a combined
* zone range on a guillotine tree); returns null otherwise so the caller can
* fall back to a single-zone drop.
*/
export function mergeZonesWithPane(root: LayoutNode, groupIds: string[], paneId: string): LayoutNode | null {
export function mergeZonesWithPane(
root: LayoutNode,
groupIds: string[],
paneId: string | readonly string[]
): LayoutNode | null {
const paneIds = typeof paneId === 'string' ? [paneId] : [...paneId]
const set = new Set(groupIds)
if (set.size <= 1 || !findCover(root, set)) {
return null
}
// Panes from the merged zones (tree order), minus the dragged one.
// Panes from the merged zones (tree order), minus the dragged block.
const panesInSet: string[] = []
const collect = (n: LayoutNode) => {
if (n.type === 'group') {
if (set.has(n.id)) {
panesInSet.push(...n.panes.filter(p => p !== paneId))
panesInSet.push(...n.panes.filter(p => !paneIds.includes(p)))
}
} else {
n.children.forEach(collect)
@@ -375,16 +441,19 @@ export function mergeZonesWithPane(root: LayoutNode, groupIds: string[], paneId:
collect(root)
// If the dragged pane lives OUTSIDE the merged set, pull it from its origin
// Any dragged pane living OUTSIDE the merged set is pulled from its origin
// first (leaving that origin an empty zone). Inside the set it's absorbed.
const origin = findGroupOfPane(root, paneId)
let working = root
if (origin && !set.has(origin.id)) {
working = removePane(root, paneId) ?? root
for (const id of paneIds) {
const origin = findGroupOfPane(working, id)
if (origin && !set.has(origin.id)) {
working = removePane(working, id) ?? working
}
}
const merged = group([paneId, ...panesInSet])
const merged = group([...paneIds, ...panesInSet])
const replace = (n: LayoutNode): LayoutNode => {
if (sameSet(groupLeafIds(n), set)) {
@@ -409,16 +478,23 @@ export function setActivePane(root: LayoutNode, groupId: string, paneId: string)
return mapGroups(root, g => (g.id === groupId && g.panes.includes(paneId) ? { ...g, active: paneId } : g))
}
/** Reorder a pane within its group's tab stack (browser-tab drag semantics). */
export function reorderPaneInGroup(root: LayoutNode, groupId: string, paneId: string, toIndex: number): LayoutNode {
/** Reorder a block of panes within a group as one unit (browser-tab drag
* semantics; a single-tab drag is a one-id block): the block lands at
* `toIndex` among the remaining tabs, keeping its own order. */
export function reorderPanesInGroup(
root: LayoutNode,
groupId: string,
paneIds: readonly string[],
toIndex: number
): LayoutNode {
return mapGroups(root, g => {
if (g.id !== groupId || !g.panes.includes(paneId)) {
if (g.id !== groupId || !paneIds.every(p => g.panes.includes(p))) {
return g
}
const without = g.panes.filter(p => p !== paneId)
const without = g.panes.filter(p => !paneIds.includes(p))
const index = Math.max(0, Math.min(without.length, toIndex))
const panes = [...without.slice(0, index), paneId, ...without.slice(index)]
const panes = [...without.slice(0, index), ...paneIds, ...without.slice(index)]
return { ...g, panes }
})
@@ -0,0 +1,147 @@
import { describe, expect, it } from 'vitest'
import { findGroup, findGroupOfPane, group, mergeZonesWithPane, movePanes, reorderPanesInGroup, split } from './model'
import { $tabSelection, clearTabSelection, selectionFor, selectTabRange, toggleTabSelected } from './tab-selection'
describe('movePanes (multi-tab drag)', () => {
it('stacks the whole block into the target group at the divider slot, in strip order', () => {
const tree = split('row', [
group(['a', 'b', 'c'], { active: 'a', id: 'left' }),
group(['x', 'y'], { active: 'x', id: 'right' })
])
const next = movePanes(tree, ['a', 'c'], { before: 'y', groupId: 'right', pos: 'center' }, 'c')
const right = findGroup(next, 'right')
expect(right).toMatchObject({ panes: ['x', 'a', 'c', 'y'], active: 'c' })
expect(findGroup(next, 'left')).toMatchObject({ panes: ['b'] })
})
it('an edge drop opens ONE split holding the block as tabs, pressed tab fronted', () => {
const tree = split('row', [
group(['a', 'b', 'c'], { active: 'a', id: 'left' }),
group(['x'], { active: 'x', id: 'right' })
])
const next = movePanes(tree, ['b', 'c'], { groupId: 'right', pos: 'bottom' }, 'b')
const landed = findGroupOfPane(next, 'b')
expect(landed).toMatchObject({ panes: ['b', 'c'], active: 'b' })
// One new zone, not one per pane: b and c share a group.
expect(findGroupOfPane(next, 'c')).toBe(landed)
expect(findGroup(next, 'left')).toMatchObject({ panes: ['a'] })
})
it('dragging a whole zone into a sibling dissolves the source zone', () => {
const tree = split('row', [
group(['a', 'b'], { active: 'a', id: 'left' }),
group(['x'], { active: 'x', id: 'right' })
])
const next = movePanes(tree, ['a', 'b'], { groupId: 'right', pos: 'center' }, 'a')
expect(next).toMatchObject({ type: 'group', panes: ['x', 'a', 'b'], active: 'a' })
})
it('is a no-op when removal dissolves the target zone itself', () => {
const tree = split('row', [
group(['a', 'b'], { active: 'a', id: 'left' }),
group(['x'], { active: 'x', id: 'right' })
])
// Dropping right's only pane (as part of a block) "into right" — the
// target vanishes with the removal, so nothing moves.
expect(movePanes(tree, ['x', 'a'], { groupId: 'right', pos: 'center' }, 'x')).toBe(tree)
})
it('falls back to single-pane semantics for a one-id block', () => {
const tree = split('row', [
group(['a', 'b'], { active: 'a', id: 'left' }),
group(['x'], { active: 'x', id: 'right' })
])
const next = movePanes(tree, ['b'], { groupId: 'right', pos: 'center' })
expect(findGroup(next, 'right')).toMatchObject({ panes: ['x', 'b'], active: 'b' })
})
})
describe('reorderPanesInGroup (block reorder)', () => {
it('moves a selection as one unit, preserving its internal order', () => {
const tree = group(['a', 'b', 'c', 'd'], { active: 'a', id: 'g' })
// [a, c] to the end: index 2 among the remaining [b, d].
expect(reorderPanesInGroup(tree, 'g', ['a', 'c'], 2)).toMatchObject({ panes: ['b', 'd', 'a', 'c'] })
})
it('leaves the group alone when any id is missing (stale selection)', () => {
const tree = group(['a', 'b'], { active: 'a', id: 'g' })
expect(reorderPanesInGroup(tree, 'g', ['a', 'ghost'], 0)).toBe(tree)
})
})
describe('mergeZonesWithPane with a multi-tab block', () => {
it('merges the span into one group led by the block in strip order', () => {
const tree = split('row', [
group(['a', 'b'], { active: 'a', id: 'left' }),
split('column', [group(['x'], { active: 'x', id: 'mid' }), group(['y'], { active: 'y', id: 'right' })])
])
const next = mergeZonesWithPane(tree, ['mid', 'right'], ['a', 'b'])
expect(next).toMatchObject({ type: 'group', panes: ['a', 'b', 'x', 'y'] })
})
it('returns null for a non-rectangular span (caller falls back to a single-zone drop)', () => {
const tree = split('row', [
group(['a', 'b'], { active: 'a', id: 'left' }),
group(['x'], { active: 'x', id: 'mid' }),
group(['y'], { active: 'y', id: 'right' })
])
expect(mergeZonesWithPane(tree, ['mid', 'right'], ['a', 'b'])).toBeNull()
})
})
describe('tab selection (Chrome grammar)', () => {
it('⌥-click seeds with the active tab, toggles, and dissolves at ≤1', () => {
clearTabSelection()
toggleTabSelected('g', 'c', 'a')
expect([...$tabSelection.get()!.ids].sort()).toEqual(['a', 'c'])
toggleTabSelected('g', 'c', 'a')
expect($tabSelection.get()).toBeNull()
})
it('shift-click ranges from the anchor and re-ranges on the next shift-click', () => {
clearTabSelection()
const order = ['a', 'b', 'c', 'd']
selectTabRange('g', order, 'c', 'a')
expect(selectionFor('g', order, 'b')).toEqual(['a', 'b', 'c'])
// Anchor holds at a (Chrome): re-ranging to d replaces, not extends.
selectTabRange('g', order, 'd', 'a')
expect(selectionFor('g', order, 'd')).toEqual(['a', 'b', 'c', 'd'])
})
it('selectionFor answers null for an unselected pressed tab and drops stale ids', () => {
clearTabSelection()
toggleTabSelected('g', 'b', 'a')
toggleTabSelected('g', 'c', 'a')
// Pressed tab outside the selection = single-tab drag.
expect(selectionFor('g', ['a', 'b', 'c', 'd'], 'd')).toBeNull()
// 'a' closed since: it silently falls out, strip order preserved.
expect(selectionFor('g', ['b', 'c', 'd'], 'b')).toEqual(['b', 'c'])
// Another zone never sees it.
expect(selectionFor('other', ['b', 'c'], 'b')).toBeNull()
clearTabSelection()
})
})
@@ -35,7 +35,8 @@ import { ESCAPE_PRIORITY, pushEscapeLayer } from '@/lib/escape-layers'
import { reorderCommitHaptic, reorderStepHaptic } from '@/lib/reorder'
import type { DropPosition } from '../model'
import { $dropHint, $treeDragging, type DropHint, mergeTreeZones, moveTreePane, reorderTreePane } from '../store'
import { $dropHint, $treeDragging, type DropHint, mergeTreeZones, moveTreePanes, reorderTreePanes } from '../store'
import { clearTabSelection } from '../tab-selection'
import { type EngineZone, HighlightedZones, primaryZone, type ZoneRect } from '../zones-engine'
const DRAG_THRESHOLD_PX = 4
@@ -96,10 +97,18 @@ const stripSlots = (strip: HTMLElement): StripSlot[] =>
})
/** Insertion slot from the pointer x against the OTHER tabs' midpoints:
* stack BEFORE the returned pane id (`null` = append). */
export function slotBefore(slots: StripSlot[], x: number, excludePaneId = ''): { before: null | string } {
* stack BEFORE the returned pane id (`null` = append). `exclude` is the
* dragged tab — or the whole selection on a multi-tab drag, so the block
* can't target a slot inside itself. */
export function slotBefore(
slots: StripSlot[],
x: number,
exclude: readonly string[] | string = ''
): { before: null | string } {
const excluded = typeof exclude === 'string' ? [exclude] : exclude
for (const slot of slots) {
if (slot.id === excludePaneId) {
if (excluded.includes(slot.id)) {
continue
}
@@ -422,7 +431,11 @@ export function startPaneDrag(
onTap?: () => void,
reorder?: ReorderContext,
double?: DoubleTapContext,
ghostLabel?: string
ghostLabel?: string,
/** Multi-tab selection riding this drag (strip order, includes `paneId`).
* The whole block moves/reorders together; `paneId` stays the pressed tab
* (it fronts at the destination). */
selection?: readonly string[]
) {
if (e.button !== 0) {
return
@@ -431,17 +444,28 @@ export function startPaneDrag(
e.preventDefault()
e.stopPropagation()
// The moving block: the selection when the pressed tab rides one, else just
// the pressed tab. Order is strip order (selectionFor guarantees it).
const moving: readonly string[] = selection && selection.length > 1 ? selection : [paneId]
const highlighted = new HighlightedZones()
let zones: EngineZone[] = []
let strips: StripSnapshot[] = []
let mode: 'reorder' | 'zone' | null = null
let dimmed: HTMLElement | null = null
let dimmed: HTMLElement[] = []
const markSource = () => {
// The dragged tab dims for the drag's life — the divider says where it
// GOES, the dim says what MOVES. No live shuffle (placement-on-release).
dimmed ??= reorder?.strip.querySelector<HTMLElement>(`[data-tree-tab="${CSS.escape(paneId)}"]`) ?? null
dimmed?.style.setProperty('opacity', '0.45')
// Every dragged tab dims for the drag's life — the divider says where they
// GO, the dim says what MOVES. No live shuffle (placement-on-release).
if (dimmed.length === 0 && reorder) {
dimmed = moving
.map(id => reorder.strip.querySelector<HTMLElement>(`[data-tree-tab="${CSS.escape(id)}"]`))
.filter((el): el is HTMLElement => el !== null)
}
for (const el of dimmed) {
el.style.setProperty('opacity', '0.45')
}
}
const enterZoneMode = () => {
@@ -494,7 +518,7 @@ export function startPaneDrag(
groupId: reorder!.groupId,
groupIds: [reorder!.groupId],
pos: 'center',
stack: slotBefore(reorderStrip().slots, x, paneId)
stack: slotBefore(reorderStrip().slots, x, moving)
}
}
@@ -525,7 +549,7 @@ export function startPaneDrag(
const strip =
groupIds.length === 1 && groupId ? strips.find(s => s.groupId === groupId && rectContains(s.rect, x, y)) : null
const stack = strip ? slotBefore(strip.slots, x, paneId) : undefined
const stack = strip ? slotBefore(strip.slots, x, moving) : undefined
const pos: DropPosition = stack
? 'center'
@@ -537,17 +561,26 @@ export function startPaneDrag(
},
onCommit(hint) {
// A multi-tab selection is spent by a LANDED drop (reorder or zone) —
// a deny-area release keeps it, so a missed drop can just be retried.
const spendSelection = () => {
if (moving.length > 1) {
clearTabSelection()
}
}
if (mode === 'reorder' && reorder && hint?.stack !== undefined) {
// Slot -> index among the OTHER tabs (reorderPaneInGroup inserts there).
// Slot -> index among the OTHER tabs (the block re-inserts there).
const others = [...reorder.strip.querySelectorAll<HTMLElement>('[data-tree-tab]')]
.map(el => el.dataset.treeTab)
.filter((id): id is string => Boolean(id) && id !== paneId)
.filter((id): id is string => Boolean(id) && !moving.includes(id!))
const toIndex = hint.stack.before ? others.indexOf(hint.stack.before) : others.length
if (toIndex >= 0) {
reorderTreePane(reorder.groupId, paneId, toIndex)
reorderTreePanes(reorder.groupId, moving, toIndex)
reorderCommitHaptic()
spendSelection()
}
}
@@ -559,18 +592,28 @@ export function startPaneDrag(
const targets = hint?.groupIds ?? []
if (targets.length > 1) {
// Shift-span: merge the highlighted zones, dropping the pane across them.
mergeTreeZones([...targets], paneId, hint?.groupId ?? null)
// Shift-span: merge the highlighted zones, dropping the block across them.
mergeTreeZones([...targets], moving, hint?.groupId ?? null)
spendSelection()
} else if (hint?.groupId) {
// strip = stack at the divider slot; center = join the stack;
// an edge = split the zone and land there.
moveTreePane(paneId, { groupId: hint.groupId, pos: hint.pos ?? 'center', before: hint.stack?.before })
// an edge = split the zone and land there. The whole selection
// rides — the pressed tab fronts at the destination.
moveTreePanes(
moving,
{ groupId: hint.groupId, pos: hint.pos ?? 'center', before: hint.stack?.before },
paneId
)
spendSelection()
}
}
},
onEnd() {
dimmed?.style.removeProperty('opacity')
for (const el of dimmed) {
el.style.removeProperty('opacity')
}
highlighted.reset()
}
})
@@ -50,6 +50,14 @@ import {
setTreeGroupMinimized,
treeTabCloseTargets
} from '../store'
import {
$tabSelection,
clearTabSelection,
isToggleSelectClick,
selectionFor,
selectTabRange,
toggleTabSelected
} from '../tab-selection'
import { type DoubleTapContext, startPaneDrag } from './drag-session'
import { forceLoneHeaderForPanes } from './lone-header'
@@ -186,6 +194,9 @@ export function TreeGroup({
const narrow = useStore($narrowViewport)
const newSessionTabAction = useStore($newSessionTabAction)
const panesWithCloser = useStore($panesWithCloser)
// Multi-tab selection (⌥/Ctrl-click, Shift-click) — null for every zone but
// the one holding it, so this subscription is quiet during normal use.
const tabSelection = useStore($tabSelection)
// Reload epochs: only an explicit tab-menu Reload writes here, so this
// subscription costs nothing on a normal render.
const paneEpochs = useStore($treePaneEpochs)
@@ -435,6 +446,7 @@ export function TreeGroup({
const chrome = paneChrome(paneFor(paneId))
const closeable = closeableTab(paneId)
const title = paneFor(paneId)?.title ?? paneId
const isSelected = tabSelection?.groupId === node.id && tabSelection.ids.has(paneId)
const tab = (
<PaneTab
@@ -444,11 +456,36 @@ export function TreeGroup({
key={paneId}
onClose={closeable ? () => closeTab(paneId) : undefined}
onPointerDown={e => {
// Chrome's tab-selection grammar, ahead of activate/drag:
// Shift-click ranges from the anchor, ⌥-click (Ctrl-click
// off-Mac) toggles. Neither activates nor starts a drag —
// the press IS the selection edit. ⌘-click stays close
// (PaneTab claims it first) and ⌃-click stays the macOS
// context menu.
if (e.button === 0 && e.shiftKey) {
e.preventDefault()
e.stopPropagation()
selectTabRange(node.id, shown, paneId, activeId)
return
}
if (isToggleSelectClick(e)) {
e.preventDefault()
e.stopPropagation()
toggleTabSelected(node.id, paneId, activeId)
return
}
// Tabs ACTIVATE (restoring a collapsed group). Minimize
// lives on the chevron / single-pane label — overloading
// the active tab made double-click a minimize/restore/hide
// lottery.
// lottery. A plain click also collapses any multi-tab
// selection back to the one tab (Chrome semantics).
const onTap = () => {
clearTabSelection()
if (node.minimized) {
restoreTreePane(paneId)
}
@@ -465,6 +502,26 @@ export function TreeGroup({
e.stopPropagation()
}
// Dragging a SELECTED tab carries the whole selection as
// one block through the generic pane move — a multi-tab
// drag outranks the pane's own tab drag (the session drop
// language is single-session).
const dragSelection = selectionFor(node.id, shown, paneId)
if (dragSelection) {
startPaneDrag(
paneId,
e,
onTap,
stripRef.current ? { groupId: node.id, strip: stripRef.current } : undefined,
hideHeaderDoubleTap,
t.zones.tabCount(dragSelection.length),
dragSelection
)
return
}
// A pane may own its tab drag (a session tab speaks the
// session drop language — link/stack/split); `false` defers
// to the generic pane move (the workspace tab on a fresh
@@ -481,6 +538,7 @@ export function TreeGroup({
}
}}
role="tab"
selected={isSelected}
style={{ cursor: 'grab' }}
>
{chrome.tabLead ? (
@@ -28,9 +28,10 @@ import {
mergeZonesWithPane as mergeZonesWithPaneOp,
mirrorTreeHorizontal,
movePane as movePaneOp,
movePanes as movePanesOp,
normalize,
removePane,
reorderPaneInGroup as reorderPaneInGroupOp,
reorderPanesInGroup as reorderPanesInGroupOp,
setActivePane as setActivePaneOp,
setGroupHeaderHidden as setGroupHeaderHiddenOp,
setGroupMinimized,
@@ -549,26 +550,29 @@ function shownPanesInGroup(group: { panes: readonly string[] }): string[] {
/** ⌘1…⌘9: activate the Nth *visible* tab of the target zone — the first of
* hovered / focused / workspace that is a real tab strip (≥2 shown panes).
* Pointing at the sidebar (or nothing) therefore still switches main's tabs
* instead of dead-ending. Returns false so the caller falls back to its
* instead of dead-ending. Returns the activated pane id — the caller needs to
* know when the slot landed on the workspace tab (a full page covering it
* must also route back to the chat) — or null so it falls back to its
* default (profile switch) when no zone qualifies. */
export function activateTreeTabSlot(slot: number): boolean {
export function activateTreeTabSlot(slot: number): null | string {
const group = tabTargetGroup(candidate => shownPanesInGroup(candidate).length >= 2)
const panes = group ? shownPanesInGroup(group) : []
if (!group || slot < 1 || slot > panes.length) {
return false
return null
}
activateTreePane(group.id, panes[slot - 1])
return true
return panes[slot - 1]
}
/** ⌃Tab / ⌃⇧Tab: cycle the target zone's *visible* tabs (wrapping) — the first
* of hovered / focused / workspace that is a chat strip with ≥2 shown tabs.
* Returns false so the caller falls back to the recent-session switcher when
* no zone qualifies. */
export function cycleTreeTabInFocusedZone(direction: 1 | -1): boolean {
* Returns the activated pane id (see `activateTreeTabSlot` — landing on the
* workspace under a full page must route back to the chat), or null so the
* caller falls back to the recent-session switcher when no zone qualifies. */
export function cycleTreeTabInFocusedZone(direction: 1 | -1): null | string {
const group = tabTargetGroup(candidate => {
const shown = shownPanesInGroup(candidate)
@@ -576,7 +580,7 @@ export function cycleTreeTabInFocusedZone(direction: 1 | -1): boolean {
})
if (!group) {
return false
return null
}
const panes = shownPanesInGroup(group)
@@ -595,7 +599,7 @@ export function cycleTreeTabInFocusedZone(direction: 1 | -1): boolean {
setTreeGroupHeaderHidden(group.id, false)
}
return true
return nextId
}
/** Remove a pane from the tree WITHOUT a dismissal record — for surfaces
@@ -1244,25 +1248,61 @@ export function applyTree(tree: LayoutNode, presetId: string) {
}
/**
* Shift-drag span: merge the highlighted zones into one holding `paneId`. Falls
* back to a single-zone move at `fallbackGroupId` when the set can't merge
* (non-rectangular selection).
* Move a multi-tab SELECTION in one commit (drag any selected tab): the lead
* pane takes the drop geometry, the rest stack in behind it in strip order,
* and `activeId` (the pressed tab) fronts in the landing group.
*/
export function mergeTreeZones(groupIds: string[], paneId: string, fallbackGroupId: string | null) {
export function moveTreePanes(
paneIds: readonly string[],
target: { groupId: string; pos: DropPosition; before?: null | string },
activeId?: string
) {
const tree = $layoutTree.get()
if (!tree) {
return
}
const next = movePanesOp(tree, paneIds, target, activeId)
if (next !== tree) {
commit(next)
markActivePreset('custom')
for (const paneId of paneIds) {
markPaneUserPlaced(paneId)
}
}
}
/**
* Shift-drag span: merge the highlighted zones into one holding `paneId`. Falls
* back to a single-zone move at `fallbackGroupId` when the set can't merge
* (non-rectangular selection).
*/
export function mergeTreeZones(
groupIds: string[],
paneId: string | readonly string[],
fallbackGroupId: null | string
) {
const tree = $layoutTree.get()
if (!tree) {
return
}
const paneIds = typeof paneId === 'string' ? [paneId] : paneId
const merged = mergeZonesWithPaneOp(tree, groupIds, paneId)
if (merged) {
commit(merged)
markActivePreset('custom')
markPaneUserPlaced(paneId)
for (const id of paneIds) {
markPaneUserPlaced(id)
}
} else if (fallbackGroupId) {
moveTreePane(paneId, { groupId: fallbackGroupId, pos: 'center' })
moveTreePanes(paneIds, { groupId: fallbackGroupId, pos: 'center' })
}
}
@@ -1274,11 +1314,13 @@ export function activateTreePane(groupId: string, paneId: string) {
}
}
export function reorderTreePane(groupId: string, paneId: string, toIndex: number) {
/** Reorder a tab block (multi-tab selection, or a single tab) within its
* group's strip — the block keeps its own order. */
export function reorderTreePanes(groupId: string, paneIds: readonly string[], toIndex: number) {
const tree = $layoutTree.get()
if (tree) {
commit(reorderPaneInGroupOp(tree, groupId, paneId, toIndex))
commit(reorderPanesInGroupOp(tree, groupId, paneIds, toIndex))
markActivePreset('custom')
}
}
@@ -0,0 +1,99 @@
/**
* Multi-tab selection on a zone's tab strip — Chrome's tab-selection grammar:
*
* - ⌥-click (Ctrl-click off-Mac) → toggle the tab in/out of the selection;
* - Shift-click → select the range from the anchor
* (the last explicitly clicked tab, else
* the active one) to the clicked tab;
* - plain click → collapse back to a single tab.
*
* ⌘-click stays CLOSE (middle-click.ts) and ⌃-click stays the macOS context
* menu, so the toggle chord is ⌥ on Mac / Ctrl elsewhere. One selection at a
* time, scoped to one zone — dragging any selected tab carries the whole set
* (drag-session resolves it), and ids are validated against the strip's
* current tabs at use time, so closed/moved panes fall out on their own.
*/
import { atom } from 'nanostores'
export interface TabSelection {
groupId: string
ids: ReadonlySet<string>
/** Range anchor: the last explicitly clicked tab (Chrome semantics). */
anchor: string
}
export const $tabSelection = atom<null | TabSelection>(null)
const isMac = typeof navigator !== 'undefined' && /Mac|iP(hone|ad|od)/.test(navigator.platform)
/** The toggle-select chord: ⌥-click on Mac (⌘ closes, ⌃ is the context menu),
* Ctrl-click elsewhere — ⌥ is accepted everywhere for one muscle memory. */
export const isToggleSelectClick = (event: { altKey: boolean; button: number; ctrlKey: boolean; metaKey: boolean }) =>
event.button === 0 && !event.metaKey && (event.altKey || (!isMac && event.ctrlKey))
export function clearTabSelection() {
if ($tabSelection.get()) {
$tabSelection.set(null)
}
}
/** ⌥/Ctrl-click: toggle `paneId`. A fresh selection seeds with the active tab
* (it is implicitly selected, as in Chrome); collapsing to ≤1 dissolves the
* selection entirely — a single "selected" tab is just a tab. */
export function toggleTabSelected(groupId: string, paneId: string, activeId: string) {
const current = $tabSelection.get()
const ids = new Set(current?.groupId === groupId ? current.ids : [activeId])
if (ids.has(paneId)) {
ids.delete(paneId)
} else {
ids.add(paneId)
}
if (ids.size <= 1) {
$tabSelection.set(null)
return
}
$tabSelection.set({ anchor: paneId, groupId, ids })
}
/** Shift-click: select the contiguous range anchor→`paneId` in strip order,
* replacing the previous range (the anchor holds, Chrome-style). */
export function selectTabRange(groupId: string, orderedPanes: readonly string[], paneId: string, activeId: string) {
const current = $tabSelection.get()
const anchor = current?.groupId === groupId && orderedPanes.includes(current.anchor) ? current.anchor : activeId
const a = orderedPanes.indexOf(anchor)
const b = orderedPanes.indexOf(paneId)
if (a === -1 || b === -1) {
return
}
const ids = new Set(orderedPanes.slice(Math.min(a, b), Math.max(a, b) + 1))
if (ids.size <= 1) {
$tabSelection.set(null)
return
}
$tabSelection.set({ anchor, groupId, ids })
}
/** The selection as an ordered slice of `orderedPanes` — but only when the
* pressed tab rides it (dragging an unselected tab is a single-tab drag).
* Stale ids (closed panes) drop out here. */
export function selectionFor(groupId: string, orderedPanes: readonly string[], paneId: string): null | string[] {
const current = $tabSelection.get()
if (current?.groupId !== groupId || !current.ids.has(paneId)) {
return null
}
const ids = orderedPanes.filter(id => current.ids.has(id))
return ids.length > 1 ? ids : null
}
@@ -48,13 +48,13 @@ describe('activateTreeTabSlot indexes shown panes only', () => {
it('⌘1 is workspace and ⌘2 is the first SESSION tab when files is hidden', async () => {
const { activeOf, tree } = await setup()
expect(tree.activateTreeTabSlot(1)).toBe(true)
expect(tree.activateTreeTabSlot(1)).toBe('workspace')
expect(activeOf()).toBe('workspace')
expect(tree.activateTreeTabSlot(2)).toBe(true)
expect(tree.activateTreeTabSlot(2)).toBe('session-tile:a')
expect(activeOf()).toBe('session-tile:a')
expect(tree.activateTreeTabSlot(3)).toBe(true)
expect(tree.activateTreeTabSlot(3)).toBe('session-tile:b')
expect(activeOf()).toBe('session-tile:b')
})
@@ -63,19 +63,19 @@ describe('activateTreeTabSlot indexes shown panes only', () => {
// Shown: workspace + A + B → 3. Slot 4 would have been `B` on the raw array
// (workspace, files, A, B) before this fix — now it correctly refuses.
expect(tree.activateTreeTabSlot(4)).toBe(false)
expect(tree.activateTreeTabSlot(4)).toBeNull()
})
it('⌃Tab cycles only visible chips', async () => {
const { activeOf, tree } = await setup()
expect(tree.cycleTreeTabInFocusedZone(1)).toBe(true)
expect(tree.cycleTreeTabInFocusedZone(1)).toBe('session-tile:a')
expect(activeOf()).toBe('session-tile:a')
expect(tree.cycleTreeTabInFocusedZone(1)).toBe(true)
expect(tree.cycleTreeTabInFocusedZone(1)).toBe('session-tile:b')
expect(activeOf()).toBe('session-tile:b')
expect(tree.cycleTreeTabInFocusedZone(1)).toBe(true)
expect(tree.cycleTreeTabInFocusedZone(1)).toBe('workspace')
expect(activeOf()).toBe('workspace')
})
})
@@ -31,12 +31,21 @@ const TAB_ACTIVE_UNDERLINE = 'shadow-[inset_0_-2px_0_var(--pane-tab-active-accen
const TAB_IDLE =
'text-(--ui-text-tertiary) [--tab-bg:var(--pane-tab-strip-bg,var(--ui-sidebar-surface-background))] hover:shadow-[inset_0_0_0_100vmax_color-mix(in_srgb,#000_var(--ui-tab-hover-darken),transparent)] hover:text-(--ui-text-secondary)'
// A tab riding a multi-tab selection: an accent wash over whatever surface the
// tab sits on. A background-image gradient (not a shadow) so it stacks cleanly
// over `--tab-bg` without fighting the active underline / hover shadows.
const TAB_SELECTED =
'[background-image:linear-gradient(color-mix(in_srgb,var(--ui-accent)_14%,transparent),color-mix(in_srgb,var(--ui-accent)_14%,transparent))] text-foreground'
interface PaneTabProps extends React.ComponentProps<'div'> {
active?: boolean
dirty?: boolean
/** Close gesture, no hover X (too easy to hit on small tabs): middle-click,
* or ⌘-click as the trackpad-friendly Mac equivalent. */
onClose?: () => void
/** Part of a multi-tab selection (⌥/Ctrl-click, Shift-click) — an accent
* wash marks every tab that a drag would carry, Chrome-style. */
selected?: boolean
/** Vertical rail form (collapsed sidebar zones). */
vertical?: boolean
/** Content-facing edge of a vertical rail — the strip line the active tab cuts. */
@@ -59,6 +68,7 @@ export const PaneTab = React.forwardRef<HTMLDivElement, PaneTabProps>(function P
onPointerDown,
onPointerUp,
onClickCapture,
selected = false,
vertical = false,
side = 'left',
children,
@@ -81,9 +91,11 @@ export const PaneTab = React.forwardRef<HTMLDivElement, PaneTabProps>(function P
active
? cn(TAB_ACTIVE, !vertical && TAB_ACTIVE_UNDERLINE)
: cn(TAB_IDLE, edge && `${edge}-(--ui-stroke-tertiary)`),
selected && TAB_SELECTED,
className
)}
data-active={active}
data-selected={selected || undefined}
data-vertical={vertical || undefined}
onClickCapture={event => {
// Sites whose tab activates on the label's own onClick (the preview
+42 -1
View File
@@ -1,7 +1,11 @@
import { describe, expect, it } from 'vitest'
import { describe, expect, it, vi } from 'vitest'
import { dispatchPluginNativeNotification } from '@/store/native-notifications'
import { createPluginContext } from './plugin'
vi.mock('@/store/native-notifications', () => ({ dispatchPluginNativeNotification: vi.fn() }))
describe('createPluginContext.onDispose', () => {
it('collects arbitrary cleanups so the host runs them on deactivate', () => {
const disposers: Array<() => void> = []
@@ -19,3 +23,40 @@ describe('createPluginContext.onDispose', () => {
expect(cleaned).toBe(true)
})
})
describe('createPluginContext.os', () => {
it('dispatches a native notification attributed to the plugin', () => {
const ctx = createPluginContext('demo')
ctx.os.notify({ body: 'b', title: 't' })
expect(dispatchPluginNativeNotification).toHaveBeenCalledWith('demo', { body: 'b', title: 't' })
})
it('resolves false (never throws) when the desktop bridge is missing', async () => {
const ctx = createPluginContext('demo')
// jsdom has no window.hermesDesktop — the exact older-shell/browser case.
await expect(ctx.os.openExternal('https://example.com')).resolves.toBe(false)
await expect(ctx.os.revealPath('/tmp')).resolves.toBe(false)
await expect(ctx.os.writeClipboard('hi')).resolves.toBe(false)
})
it('routes through the bridge and turns a bridge throw into false', async () => {
const bridge = {
openExternal: vi.fn().mockResolvedValue(undefined),
revealPath: vi.fn().mockResolvedValue(true),
writeClipboard: vi.fn().mockRejectedValue(new Error('nope'))
}
;(window as unknown as { hermesDesktop: unknown }).hermesDesktop = bridge
try {
const ctx = createPluginContext('demo')
await expect(ctx.os.openExternal('https://example.com')).resolves.toBe(true)
expect(bridge.openExternal).toHaveBeenCalledWith('https://example.com')
await expect(ctx.os.revealPath('/tmp')).resolves.toBe(true)
await expect(ctx.os.writeClipboard('hi')).resolves.toBe(false)
} finally {
delete (window as unknown as { hermesDesktop?: unknown }).hermesDesktop
}
})
})
+59
View File
@@ -15,11 +15,13 @@
import { pluginRest, type PluginRestOptions, pluginSocket } from '@/hermes'
import { createPluginI18n, type PluginI18n } from '@/i18n'
import { readKey, writeKey } from '@/lib/storage'
import { dispatchPluginNativeNotification, type PluginNativeNotificationInput } from '@/store/native-notifications'
import { registry } from './registry'
import type { Contribution } from './types'
export type { PluginRestOptions } from '@/hermes'
export type { PluginNativeNotificationInput } from '@/store/native-notifications'
/** A contribution as a plugin author writes it — provenance + id scoping are
* the host's job, so those fields are off-limits here. */
@@ -33,6 +35,27 @@ export interface PluginStorage {
remove(key: string): void
}
/** The curated OS door — every way a plugin reaches outside the app window,
* in one attributed namespace instead of the raw `window.hermesDesktop`
* bridge. Every member resolves a result instead of throwing when the
* capability can't apply (no Electron shell, older desktop build), so
* callers branch on the return value rather than sniffing the bridge. */
export interface PluginOs {
/** Native OS notification (Electron), attributed to this plugin. Gated by
* Settings ▸ Notifications ▸ "Plugin notifications" and fires only while
* the user is away from Hermes — use `host.notify` for the in-app toast.
* Throttled per plugin; reserve it for genuinely notable events. */
notify: (input: PluginNativeNotificationInput) => void
/** Open a URL with the OS default handler (browser, mail client, custom
* schemes like `spotify:`). Resolves false when the shell can't. */
openExternal: (url: string) => Promise<boolean>
/** Reveal a path in the OS file manager (Finder / Explorer). Resolves
* false when unavailable. */
revealPath: (path: string) => Promise<boolean>
/** Write text to the system clipboard. Resolves false when unavailable. */
writeClipboard: (text: string) => Promise<boolean>
}
export interface PluginContext {
/** The resolved plugin source tag, e.g. `'plugin:cost-meter'`. */
readonly source: string
@@ -54,6 +77,10 @@ export interface PluginContext {
* returned. Resolves to a no-op on OAuth remotes — treat it as an
* accelerator over your polling, never a replacement. */
socket: (path: string, onMessage: (data: unknown) => void) => () => void
/** The curated OS door: native notification, open-external, reveal-in-file-
* manager, clipboard — attributed to this plugin, result-shaped (never
* throws for a missing capability). */
os: PluginOs
/** Plugin-scoped persistence. */
storage: PluginStorage
/** Plugin-scoped i18n: ship + register locale bundles under this plugin,
@@ -96,6 +123,37 @@ function createPluginStorage(pluginId: string): PluginStorage {
}
}
// Never throws for a missing capability: the renderer can outlive an older
// Electron shell (or run in a plain browser), so every door degrades to a
// false result the plugin can branch on.
function createPluginOs(pluginId: string): PluginOs {
const attempt = async (run: (bridge: NonNullable<typeof window.hermesDesktop>) => Promise<boolean>) => {
const bridge = typeof window === 'undefined' ? undefined : window.hermesDesktop
if (!bridge) {
return false
}
try {
return await run(bridge)
} catch {
return false
}
}
return {
notify: input => dispatchPluginNativeNotification(pluginId, input),
openExternal: url =>
attempt(async bridge => {
await bridge.openExternal(url)
return true
}),
revealPath: path => attempt(async bridge => (bridge.revealPath ? bridge.revealPath(path) : false)),
writeClipboard: text => attempt(bridge => bridge.writeClipboard(text))
}
}
/** Build the scoped context handed to a plugin's `register`. `onDispose`
* receives every registration's disposer (the loader's unload/reload hook). */
export function createPluginContext(pluginId: string, onDispose?: (dispose: () => void) => void): PluginContext {
@@ -115,6 +173,7 @@ export function createPluginContext(pluginId: string, onDispose?: (dispose: () =
onDispose: fn => void track(fn),
rest: <T>(path: string, opts?: PluginRestOptions) => pluginRest<T>(pluginId, path, opts),
socket: (path, onMessage) => track(pluginSocket(pluginId, path, onMessage)),
os: createPluginOs(pluginId),
storage: createPluginStorage(pluginId),
i18n: createPluginI18n(pluginId, track)
}
+8
View File
@@ -129,6 +129,12 @@ declare global {
}
readFileText: (filePath: string) => Promise<HermesReadFileTextResult>
selectPaths: (options?: HermesSelectPathsOptions) => Promise<string[]>
/** Native save dialog; returns the chosen path or null on cancel. */
selectSavePath?: (options?: {
defaultPath?: string
filters?: Array<{ extensions: string[]; name: string }>
title?: string
}) => Promise<null | string>
writeClipboard: (text: string) => Promise<boolean>
readClipboard: () => Promise<string>
saveImageFromUrl: (url: string) => Promise<boolean>
@@ -769,6 +775,8 @@ export interface HermesNotification {
silent?: boolean
kind?: string
sessionId?: string
/** Dedupe discriminator for session-less notifications (e.g. plugin id). */
tag?: string
actions?: { id: string; text: string }[]
}
+32
View File
@@ -47,6 +47,7 @@ import type {
PairingResponse,
PairingUser,
ProfileCreatePayload,
ProfileDesktopOverlay,
ProfileSetupCommand,
ProfileSoul,
ProfilesResponse,
@@ -186,6 +187,7 @@ export type {
PairingResponse,
PairingUser,
ProfileCreatePayload,
ProfileDesktopOverlay,
ProfileInfo,
ProfileSetupCommand,
ProfileSoul,
@@ -1431,6 +1433,36 @@ export function getProfileSetupCommand(name: string): Promise<ProfileSetupComman
})
}
/** Export a profile to a shareable .tar.gz on the backend's filesystem.
* `extraFiles` stages extra root-level files (desktop.json — the appearance/
* interface overlay) into the archive alongside the profile's own artifacts. */
export function exportProfileArchive(
name: string,
opts: { extraFiles?: Record<string, string>; output?: string } = {}
): Promise<{ archive: string; ok: boolean }> {
return window.hermesDesktop.api<{ archive: string; ok: boolean }>({
path: `/api/profiles/${encodeURIComponent(name)}/export`,
method: 'POST',
body: { extra_files: opts.extraFiles ?? {}, output: opts.output ?? '' },
timeoutMs: STARTUP_REQUEST_TIMEOUT_MS
})
}
/** Import a profile .tar.gz as a new profile. Returns the bundled desktop
* appearance overlay too (when the archive carried one) so the caller can
* apply theme/layout without another round-trip. */
export function importProfileArchive(
archive: string,
name?: string
): Promise<{ desktop: null | ProfileDesktopOverlay; name: string; ok: boolean; path: string }> {
return window.hermesDesktop.api<{ desktop: null | ProfileDesktopOverlay; name: string; ok: boolean; path: string }>({
path: '/api/profiles/import',
method: 'POST',
body: { archive, name: name || null },
timeoutMs: STARTUP_REQUEST_TIMEOUT_MS
})
}
export function getUsageAnalytics(days = 30): Promise<AnalyticsResponse> {
return window.hermesDesktop.api<AnalyticsResponse>({
...profileScoped(),
+8 -1
View File
@@ -1293,6 +1293,12 @@ export const ar = defineLocale({
count: count => `${count} ملف شخصي`,
loading: 'جار التحميل...',
newProfile: 'ملف شخصي جديد',
importProfile: 'استيراد ملف شخصي…',
exportProfile: 'تصدير ملف شخصي…',
imported: 'تم استيراد الملف الشخصي',
exported: 'تم تصدير الملف الشخصي',
failedImport: 'فشل استيراد الملف الشخصي',
failedExport: 'فشل تصدير الملف الشخصي',
allProfiles: 'كل الملفات الشخصية',
showAllProfiles: 'إظهار كل الملفات الشخصية',
switchToProfile: name => `التبديل إلى ${name}`,
@@ -2265,7 +2271,8 @@ export const ar = defineLocale({
layoutNamePlaceholder: fallback => `اسم التخطيط (${fallback})`,
saveApply: 'حفظ وتطبيق',
notExpressible: 'هذا الترتيب متشابك — لا يمكن تمثيله كتقسيمات متداخلة بعد',
zoneCount: count => `${count} مناطق`
zoneCount: count => `${count} مناطق`,
tabCount: count => `${count} تبويبات`
},
assistant: {
thread: {
+17 -1
View File
@@ -395,6 +395,10 @@ export const en: Translations = {
credits: {
label: 'Credit alerts',
description: 'Credit access is paused or restored.'
},
plugin: {
label: 'Plugin notifications',
description: 'A desktop plugin sent a notification while Hermes was in the background.'
}
},
test: 'Send test notification',
@@ -1559,6 +1563,12 @@ export const en: Translations = {
search: 'Search profiles...',
loading: 'Loading profiles...',
newProfile: 'New profile',
importProfile: 'Import profile…',
exportProfile: 'Export profile…',
imported: 'Profile imported',
exported: 'Profile exported',
failedImport: 'Failed to import profile',
failedExport: 'Failed to export profile',
allProfiles: 'All profiles',
showAllProfiles: 'Show all profiles',
switchToProfile: name => `Switch to ${name}`,
@@ -1872,6 +1882,11 @@ export const en: Translations = {
menuAddFolder: 'Add folder',
menuSetActive: 'Set active',
menuDelete: 'Delete',
moveToProject: 'Move to project',
movedTo: name => `Moved to ${name}`,
moveFailed: 'Could not move session',
moveNoFolder: 'That project has no folder to move into',
moveNoProjects: 'No other projects',
reveal: 'Reveal in folder',
copyPath: 'Copy path',
removeFromSidebar: 'Hide from sidebar',
@@ -2693,7 +2708,8 @@ export const en: Translations = {
layoutNamePlaceholder: fallback => `Layout name (${fallback})`,
saveApply: 'Save & apply',
notExpressible: 'this arrangement interlocks (pinwheel) — not expressible as nested splits yet',
zoneCount: count => `${count} zones`
zoneCount: count => `${count} zones`,
tabCount: count => `${count} tabs`
},
assistant: {
+12 -1
View File
@@ -269,6 +269,10 @@ export const ja = defineLocale({
credits: {
label: 'クレジット通知',
description: 'クレジットの利用が停止または復旧しました。'
},
plugin: {
label: 'プラグイン通知',
description: 'Hermes がバックグラウンドの間に、デスクトッププラグインが通知を送信しました。'
}
},
test: 'テスト通知を送信',
@@ -1396,6 +1400,12 @@ export const ja = defineLocale({
search: 'プロファイルを検索...',
loading: 'プロファイルを読み込み中...',
newProfile: '新しいプロファイル',
importProfile: 'プロファイルをインポート…',
exportProfile: 'プロファイルをエクスポート…',
imported: 'プロファイルをインポートしました',
exported: 'プロファイルをエクスポートしました',
failedImport: 'プロファイルのインポートに失敗しました',
failedExport: 'プロファイルのエクスポートに失敗しました',
allProfiles: 'すべてのプロファイル',
showAllProfiles: 'すべてのプロファイルを表示',
switchToProfile: name => `${name} に切り替え`,
@@ -2519,7 +2529,8 @@ export const ja = defineLocale({
layoutNamePlaceholder: fallback => `レイアウト名(${fallback})`,
saveApply: '保存して適用',
notExpressible: 'この配置は互いに噛み合っています(風車型)— 入れ子の分割では表現できません',
zoneCount: count => `${count} ゾーン`
zoneCount: count => `${count} ゾーン`,
tabCount: count => `${count} 個のタブ`
},
assistant: {
+13 -1
View File
@@ -324,7 +324,7 @@ export interface Translations {
enableAllDesc: string
focusedHint: string
kinds: Record<
'approval' | 'backgroundDone' | 'credits' | 'input' | 'turnDone' | 'turnError',
'approval' | 'backgroundDone' | 'credits' | 'input' | 'plugin' | 'turnDone' | 'turnError',
{ label: string; description: string }
>
test: string
@@ -1305,6 +1305,12 @@ export interface Translations {
search: string
loading: string
newProfile: string
importProfile: string
exportProfile: string
imported: string
exported: string
failedImport: string
failedExport: string
allProfiles: string
showAllProfiles: string
switchToProfile: (name: string) => string
@@ -1573,6 +1579,11 @@ export interface Translations {
menuAddFolder: string
menuSetActive: string
menuDelete: string
moveToProject: string
movedTo: (name: string) => string
moveFailed: string
moveNoFolder: string
moveNoProjects: string
reveal: string
copyPath: string
removeFromSidebar: string
@@ -2293,6 +2304,7 @@ export interface Translations {
saveApply: string
notExpressible: string
zoneCount: (count: number) => string
tabCount: (count: number) => string
}
assistant: {
+12 -1
View File
@@ -263,6 +263,10 @@ export const zhHant = defineLocale({
credits: {
label: '額度提醒',
description: '額度存取被暫停或恢復。'
},
plugin: {
label: '外掛通知',
description: 'Hermes 在背景時,桌面外掛傳送了通知。'
}
},
test: '傳送測試通知',
@@ -1345,6 +1349,12 @@ export const zhHant = defineLocale({
search: '搜尋設定檔…',
loading: '正在載入設定檔…',
newProfile: '新增設定檔',
importProfile: '匯入設定檔…',
exportProfile: '匯出設定檔…',
imported: '設定檔已匯入',
exported: '設定檔已匯出',
failedImport: '匯入設定檔失敗',
failedExport: '匯出設定檔失敗',
allProfiles: '全部設定檔',
showAllProfiles: '顯示全部設定檔',
switchToProfile: name => `切換至 ${name}`,
@@ -2439,7 +2449,8 @@ export const zhHant = defineLocale({
layoutNamePlaceholder: fallback => `版面名稱(${fallback})`,
saveApply: '儲存並套用',
notExpressible: '此排列互相咬合(風車形)——暫時無法表示為巢狀分割',
zoneCount: count => `${count} 個區域`
zoneCount: count => `${count} 個區域`,
tabCount: count => `${count} 個分頁`
},
assistant: {
+17 -1
View File
@@ -387,6 +387,10 @@ export const zh: Translations = {
credits: {
label: '额度提醒',
description: '额度访问被暂停或恢复。'
},
plugin: {
label: '插件通知',
description: 'Hermes 在后台时,桌面插件发送了通知。'
}
},
test: '发送测试通知',
@@ -1753,6 +1757,12 @@ export const zh: Translations = {
search: '搜索配置档案…',
loading: '正在加载配置档案…',
newProfile: '新建配置档案',
importProfile: '导入配置档案…',
exportProfile: '导出配置档案…',
imported: '配置档案已导入',
exported: '配置档案已导出',
failedImport: '导入配置档案失败',
failedExport: '导出配置档案失败',
allProfiles: '全部配置档案',
showAllProfiles: '显示全部配置档案',
switchToProfile: name => `切换到 ${name}`,
@@ -2066,6 +2076,11 @@ export const zh: Translations = {
menuAddFolder: '添加文件夹',
menuSetActive: '设为活动',
menuDelete: '删除',
moveToProject: '移动到项目',
movedTo: name => `已移动到 ${name}`,
moveFailed: '无法移动会话',
moveNoFolder: '该项目没有可移入的文件夹',
moveNoProjects: '没有其他项目',
reveal: '在文件夹中显示',
copyPath: '复制路径',
removeFromSidebar: '从侧边栏移除',
@@ -2871,7 +2886,8 @@ export const zh: Translations = {
layoutNamePlaceholder: fallback => `布局名称(${fallback})`,
saveApply: '保存并应用',
notExpressible: '此排列互相咬合(风车形)——暂无法表示为嵌套拆分',
zoneCount: count => `${count} 个区域`
zoneCount: count => `${count} 个区域`,
tabCount: count => `${count} 个标签页`
},
assistant: {
+84
View File
@@ -0,0 +1,84 @@
/**
* Render cost of one message's content parts, in budget units.
*
* Two layers bound long transcripts and both spend the same currency: the
* store window (how many messages reach assistant-ui at all) and the DOM page
* budget (how many of those actually render). Neither can be a message COUNT —
* counting only parts underpriced a 51KB tool result as "1", so a handful of
* huge results let a 600KB transcript through the old 300-part cap and drove
* Chromium's renderer into a GC crash. Characters approximate markdown
* parsing, text-node allocation, and tool-result formatting; parts approximate
* component/node count.
*
* Shared so a heavy-but-short session is bounded by the same rule as a
* long-but-light one (#55191).
*/
export const RENDER_WEIGHT_CHARS = 512
// Stop traversing once a single message has enough text to consume a complete
// DOM render page. Going further cannot change which whole turn crosses any
// budget, and avoiding an unbounded walk matters for deeply nested tool
// payloads.
const MAX_MEASURED_MESSAGE_CHARS = 300 * RENDER_WEIGHT_CHARS
const contentWeightCache = new WeakMap<object, number>()
const NON_RENDERED_CONTENT_FIELDS = new Set(['id', 'role', 'toolCallId', 'toolName', 'type'])
/**
* Estimate the synchronous renderer cost of one message's content array.
*
* A WeakMap keeps settled history O(message count) on later store updates;
* both assistant-ui and the session store publish a new content array when a
* streaming message changes, so the live tail still receives a fresh weight.
*/
export function messageRenderWeight(content: unknown): number {
if (!Array.isArray(content)) {
return 1
}
const cached = contentWeightCache.get(content)
if (cached !== undefined) {
return cached
}
const seen = new WeakSet<object>()
const pending: unknown[] = [...content]
let characters = 0
while (pending.length > 0 && characters < MAX_MEASURED_MESSAGE_CHARS) {
const value = pending.pop()
if (typeof value === 'string') {
characters += Math.min(value.length, MAX_MEASURED_MESSAGE_CHARS - characters)
continue
}
if (!value || typeof value !== 'object' || seen.has(value)) {
continue
}
seen.add(value)
if (Array.isArray(value)) {
for (const nested of value) {
pending.push(nested)
}
continue
}
for (const [key, nested] of Object.entries(value)) {
if (!NON_RENDERED_CONTENT_FIELDS.has(key)) {
pending.push(nested)
}
}
}
const weight = Math.max(1, content.length) + Math.ceil(characters / RENDER_WEIGHT_CHARS)
contentWeightCache.set(content, weight)
return weight
}
+2
View File
@@ -200,6 +200,8 @@ export type {
HermesPlugin,
PluginContext,
PluginContribution,
PluginNativeNotificationInput,
PluginOs,
PluginRestOptions,
PluginStorage
} from '@/contrib/plugin'
@@ -3,6 +3,7 @@ import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { $gateway } from './gateway'
import {
dispatchNativeNotification,
dispatchPluginNativeNotification,
NATIVE_NOTIFICATION_KINDS,
respondToApprovalAction,
sendTestNativeNotification,
@@ -170,6 +171,34 @@ describe('dispatchNativeNotification post-connect baseline', () => {
})
})
describe('dispatchPluginNativeNotification', () => {
it('fires while the user is away and tags the plugin id for dedupe', () => {
dispatchPluginNativeNotification('index-network', { body: 'New match', title: 'Opportunity' })
expect(notify).toHaveBeenCalledWith(
expect.objectContaining({ body: 'New match', kind: 'plugin', tag: 'index-network', title: 'Opportunity' })
)
})
it('suppresses while the window is focused (the in-app toast covers foreground)', () => {
setWindowState({ focused: true, hidden: false })
dispatchPluginNativeNotification('focused-plugin', { title: 'Opportunity' })
expect(notify).not.toHaveBeenCalled()
})
it('is gated by the "plugin" kind preference', () => {
setNativeNotifyKind('plugin', false)
dispatchPluginNativeNotification('muted-plugin', { title: 'Opportunity' })
expect(notify).not.toHaveBeenCalled()
})
it('throttles per plugin, so two plugins cannot collapse each other', () => {
dispatchPluginNativeNotification('plugin-a', { title: 'a' })
dispatchPluginNativeNotification('plugin-a', { title: 'a again' })
dispatchPluginNativeNotification('plugin-b', { title: 'b' })
expect(notify).toHaveBeenCalledTimes(2)
})
})
describe('dispatchNativeNotification throttle', () => {
it('collapses duplicate kind+session within the throttle window', () => {
const sessionId = freshSession()
+44 -4
View File
@@ -9,7 +9,14 @@ import { $activeSessionId } from './session'
// Native OS notifications (Electron `Notification`), separate from the in-app
// toast feed in `notifications.ts`. Each kind toggles independently.
export type NativeNotificationKind = 'approval' | 'backgroundDone' | 'credits' | 'input' | 'turnDone' | 'turnError'
export type NativeNotificationKind =
| 'approval'
| 'backgroundDone'
| 'credits'
| 'input'
| 'plugin'
| 'turnDone'
| 'turnError'
export const NATIVE_NOTIFICATION_KINDS: readonly NativeNotificationKind[] = [
'approval',
@@ -17,7 +24,8 @@ export const NATIVE_NOTIFICATION_KINDS: readonly NativeNotificationKind[] = [
'turnDone',
'turnError',
'backgroundDone',
'credits'
'credits',
'plugin'
]
// Blocking prompts — surface even while focused if they're for another session.
@@ -32,7 +40,15 @@ const STORAGE_KEY = 'hermes:native-notifications'
const DEFAULT_PREFS: NativeNotificationPrefs = {
enabled: true,
kinds: { approval: true, backgroundDone: true, credits: true, input: true, turnDone: true, turnError: true }
kinds: {
approval: true,
backgroundDone: true,
credits: true,
input: true,
plugin: true,
turnDone: true,
turnError: true
}
}
function readPrefs(): NativeNotificationPrefs {
@@ -152,6 +168,12 @@ export interface NativeNotificationInput {
global?: boolean
silent?: boolean
actions?: NativeNotificationAction[]
/**
* Extra throttle/dedupe discriminator for session-less notifications (e.g.
* the plugin id), so unrelated emitters of the same kind don't collapse
* into one another. Never drives click-to-focus like `sessionId` does.
*/
tag?: string
}
export function dispatchNativeNotification(input: NativeNotificationInput): void {
@@ -169,7 +191,7 @@ export function dispatchNativeNotification(input: NativeNotificationInput): void
return
}
if (throttled(`${input.kind}:${input.sessionId ?? (input.global ? 'global' : '')}`, Date.now())) {
if (throttled(`${input.kind}:${input.sessionId ?? input.tag ?? (input.global ? 'global' : '')}`, Date.now())) {
return
}
@@ -179,10 +201,28 @@ export function dispatchNativeNotification(input: NativeNotificationInput): void
kind: input.kind,
sessionId: input.sessionId ?? undefined,
silent: input.silent,
tag: input.tag,
title: input.title
})
}
// -- the plugin door (`ctx.os.notify`) ----------------------------------------
export interface PluginNativeNotificationInput {
title: string
body?: string
silent?: boolean
}
/** Native OS notification on behalf of a plugin. One "Plugin notifications"
* preference gates all plugins; the plugin id keys throttling/dedupe so two
* plugins can't collapse each other's notifications. Fires only while the
* user is away from Hermes — the in-app toast (`host.notify`) covers the
* foreground case. */
export function dispatchPluginNativeNotification(pluginId: string, input: PluginNativeNotificationInput): void {
dispatchNativeNotification({ ...input, global: true, kind: 'plugin', tag: pluginId })
}
// Resolve a pending approval from a notification button, mirroring the in-app
// Run/Reject bar. Keyed by session id — a background approval has no local guard.
export async function respondToApprovalAction(sessionId: null | string, actionId: string): Promise<void> {
@@ -0,0 +1,126 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import type { DesktopTheme } from '@/themes/types'
import type { ProfileDesktopOverlay } from '@/types/hermes'
// Keep side-effecting transitive imports inert (gateway sockets, REST).
vi.mock('@/store/gateway', async () => {
const { atom } = await import('nanostores')
return {
$gateway: atom<unknown>(null),
ensureGatewayForProfile: vi.fn(async () => undefined),
openGatewayForProfile: vi.fn(async () => undefined)
}
})
vi.mock('@/hermes', () => ({
exportProfileArchive: vi.fn(async () => ({ archive: '/tmp/out.tar.gz', ok: true })),
getProfiles: vi.fn(async () => ({ profiles: [] })),
importProfileArchive: vi.fn(async () => ({ desktop: null, name: 'imported', ok: true, path: '/tmp/p' })),
setApiRequestProfile: vi.fn()
}))
vi.mock('@/lib/query-client', () => ({ invalidateProfileScopedQueries: vi.fn() }))
vi.mock('@/store/starmap', () => ({ resetStarmapGraph: vi.fn() }))
const { applyDesktopOverlay, buildDesktopOverlay, exportProfileBundle } = await import('./profile-share')
const { $profileColors, setProfileColor } = await import('./profile')
const { modePref, skinPref } = await import('@/themes/context')
const { $userThemes } = await import('@/themes/user-themes')
const { $layoutTree } = await import('@/components/pane-shell/tree/store')
const { exportProfileArchive } = await import('@/hermes')
// isValidTheme only requires background/foreground/primary at runtime; the
// static type wants the full palette, hence the cast.
const roseTheme = {
name: 'rose-quartz',
label: 'Rose Quartz',
description: 'test theme',
colors: { background: '#fff0f5', foreground: '#221122', primary: '#e91e63' }
} as unknown as DesktopTheme
beforeEach(() => {
window.localStorage.clear()
$userThemes.set({})
$profileColors.set({})
})
afterEach(() => {
vi.clearAllMocks()
})
describe('buildDesktopOverlay', () => {
it('snapshots skin, mode, rail color, and the layout tree for the profile', () => {
skinPref.assign('glam', 'mono')
modePref.assign('glam', 'dark')
setProfileColor('glam', '#e91e63')
const overlay = buildDesktopOverlay('glam')
expect(overlay.version).toBe(1)
expect(overlay.skin).toBe('mono')
expect(overlay.mode).toBe('dark')
expect(overlay.profileColor).toBe('#e91e63')
// Built-in skin → no bundled theme definitions.
expect(overlay.themes).toBeUndefined()
})
it('bundles the full definition of a non-built-in skin', () => {
$userThemes.set({ 'rose-quartz': roseTheme })
skinPref.assign('glam', 'rose-quartz')
const overlay = buildDesktopOverlay('glam')
expect(overlay.skin).toBe('rose-quartz')
expect(overlay.themes).toEqual({ 'rose-quartz': roseTheme })
})
})
describe('applyDesktopOverlay', () => {
it('installs bundled themes and assigns skin/mode/color to the new profile', () => {
applyDesktopOverlay('glam-copy', {
version: 1,
skin: 'rose-quartz',
mode: 'dark',
themes: { 'rose-quartz': roseTheme },
profileColor: '#e91e63'
})
expect($userThemes.get()['rose-quartz']).toEqual(roseTheme)
expect(skinPref.resolve('glam-copy')).toBe('rose-quartz')
expect(modePref.resolve('glam-copy')).toBe('dark')
expect($profileColors.get()['glam-copy']).toBe('#e91e63')
})
it('ignores a skin that resolves to nothing and junk layout trees', () => {
const before = $layoutTree.get()
applyDesktopOverlay('glam-copy', {
skin: 'no-such-skin',
layoutTree: { bogus: true }
} as ProfileDesktopOverlay)
// Unresolvable skin → pref falls back to the default resolution.
expect(skinPref.resolve('glam-copy')).toBe(skinPref.resolve('some-unassigned'))
expect($layoutTree.get()).toBe(before)
})
it('is a no-op for a plain CLI archive (no overlay)', () => {
expect(() => applyDesktopOverlay('glam-copy', null)).not.toThrow()
expect(() => applyDesktopOverlay('glam-copy', undefined)).not.toThrow()
})
})
describe('exportProfileBundle', () => {
it('stages desktop.json into the archive through extra_files', async () => {
skinPref.assign('glam', 'mono')
const archive = await exportProfileBundle('glam', '/tmp/glam.tar.gz')
expect(archive).toBe('/tmp/out.tar.gz')
const call = vi.mocked(exportProfileArchive).mock.calls[0]
expect(call[0]).toBe('glam')
const overlay = JSON.parse(call[1]?.extraFiles?.['desktop.json'] ?? '{}') as ProfileDesktopOverlay
expect(overlay.skin).toBe('mono')
expect(call[1]?.output).toBe('/tmp/glam.tar.gz')
})
})
+219
View File
@@ -0,0 +1,219 @@
/**
* Profile share: export/import a profile as a portable bundle.
*
* The archive is the CLI's own `hermes profile export` tar.gz (config, skills,
* SOUL.md, cron — credentials always excluded), plus one desktop-only file at
* the root: `desktop.json`, the appearance/interface overlay (skin + mode,
* any user-theme definitions the skin needs, the profile rail color, and the
* layout tree). A CLI import of the same archive simply carries the file
* along; the desktop import applies it so the receiving user gets the whole
* look — theme, layout, skills — as a ready-to-use profile.
*
* Paths, not bytes, cross the renderer↔backend boundary: the native save/open
* dialogs and the backend share the filesystem for local and pooled backends.
*/
import { isLayoutNode, normalize } from '@/components/pane-shell/tree/model'
import { $layoutTree, markActivePreset, persistTree } from '@/components/pane-shell/tree/store'
import { exportProfileArchive, importProfileArchive } from '@/hermes'
import { translateNow } from '@/i18n'
import { modePref, skinPref, type ThemeMode } from '@/themes/context'
import { BUILTIN_THEMES } from '@/themes/presets'
import type { DesktopTheme } from '@/themes/types'
import { $userThemes, installUserTheme, resolveTheme } from '@/themes/user-themes'
import type { ProfileDesktopOverlay } from '@/types/hermes'
import { notify, notifyError } from './notifications'
import {
$activeGatewayProfile,
$profileColors,
normalizeProfileKey,
refreshActiveProfile,
selectProfile,
setProfileColor
} from './profile'
/** Filename of the overlay inside the archive (profile root). */
export const DESKTOP_OVERLAY_FILENAME = 'desktop.json'
const OVERLAY_VERSION = 1
/**
* Snapshot the desktop appearance/interface for `profile` into the overlay.
* The layout tree is global (one window layout, not per-profile) — it rides
* along so the receiver can opt into the sender's whole interface.
*/
export function buildDesktopOverlay(profile: string): ProfileDesktopOverlay {
const key = normalizeProfileKey(profile)
const skin = skinPref.resolve(key)
const mode = modePref.resolve(key)
// Bundle the full definition of any non-built-in theme the skin points at,
// so the receiver's picker can resolve it. Built-ins resolve by name.
const themes: Record<string, unknown> = {}
const userTheme = BUILTIN_THEMES[skin] ? undefined : $userThemes.get()[skin]
if (userTheme) {
themes[userTheme.name] = userTheme
}
return {
version: OVERLAY_VERSION,
skin,
mode,
...(Object.keys(themes).length ? { themes } : {}),
profileColor: $profileColors.get()[key] ?? null,
layoutTree: $layoutTree.get()
}
}
/** Export `profile` (backend archive + desktop overlay) to `output` (or the
* backend's staging dir when omitted). Returns the archive path. */
export async function exportProfileBundle(profile: string, output?: string): Promise<string> {
const overlay = buildDesktopOverlay(profile)
const { archive } = await exportProfileArchive(profile, {
extraFiles: { [DESKTOP_OVERLAY_FILENAME]: JSON.stringify(overlay, null, 2) },
output
})
return archive
}
const isThemeMode = (value: unknown): value is ThemeMode =>
value === 'light' || value === 'dark' || value === 'system'
/**
* Apply an imported overlay: install bundled themes, assign the new profile's
* skin + mode + rail color, and (when present) adopt the sender's layout tree.
* Every step is independent and best-effort — a malformed half never blocks
* the rest, and a missing overlay is a plain CLI-exported archive (no-op).
*/
export function applyDesktopOverlay(profile: string, overlay: null | ProfileDesktopOverlay | undefined): void {
if (!overlay || typeof overlay !== 'object') {
return
}
const key = normalizeProfileKey(profile)
// 1. Bundled theme definitions. installUserTheme validates shape and refuses
// built-in collisions; a bad entry just doesn't install.
for (const theme of Object.values(overlay.themes ?? {})) {
try {
installUserTheme(theme as DesktopTheme)
} catch {
// Invalid/colliding theme — the skin assignment below falls back.
}
}
// 2. Appearance assignment for the new profile. Only assign a skin that
// actually resolves so the pref never points at nothing.
if (typeof overlay.skin === 'string' && resolveTheme(overlay.skin)) {
skinPref.assign(key, overlay.skin)
}
if (isThemeMode(overlay.mode)) {
modePref.assign(key, overlay.mode)
}
// 3. Rail color.
if (typeof overlay.profileColor === 'string' && overlay.profileColor) {
setProfileColor(key, overlay.profileColor)
}
// 4. Layout tree — global by design (one window layout). Normalize through
// the same canonicalizer the boot load uses; a null result means the
// tree was junk, so the current layout stays.
if (overlay.layoutTree != null && isLayoutNode(overlay.layoutTree)) {
const tree = normalize(overlay.layoutTree)
if (tree) {
$layoutTree.set(tree)
persistTree()
markActivePreset('custom')
}
}
}
/** Import an archive, apply its desktop overlay, return the new profile name. */
export async function importProfileBundle(archive: string, name?: string): Promise<string> {
const result = await importProfileArchive(archive, name)
applyDesktopOverlay(result.name, result.desktop)
return result.name
}
/** The profile the export pickers should default to — the active one. */
export function activeProfileKey(): string {
return normalizeProfileKey($activeGatewayProfile.get())
}
// ── Dialog-driven flows ──────────────────────────────────────────────────────
// One store function per user verb (⌘K row, rail button, and any future menu
// item all funnel here). Toasts via the shared notification store; strings via
// translateNow so the flows stay callable from non-React surfaces.
const ARCHIVE_FILTERS = [{ extensions: ['tar.gz', 'tgz'], name: 'Hermes profile' }]
/** Pick a save location and export `profile` (default: the active one).
* Returns the archive path, or null when the user cancelled. */
export async function runExportProfileFlow(profile?: string): Promise<null | string> {
const target = normalizeProfileKey(profile ?? activeProfileKey())
const pick = window.hermesDesktop?.selectSavePath
if (!pick) {
return null
}
const output = await pick({
title: translateNow('profiles.exportProfile'),
defaultPath: `${target}.tar.gz`,
filters: ARCHIVE_FILTERS
})
if (!output) {
return null
}
try {
const archive = await exportProfileBundle(target, output)
notify({ kind: 'success', title: translateNow('profiles.exported'), message: archive })
return archive
} catch (error) {
notifyError(error, translateNow('profiles.failedExport'))
return null
}
}
/** Pick an archive and import it as a new profile; lands the user in it on a
* fresh chat. Returns the new profile name, or null when cancelled/failed. */
export async function runImportProfileFlow(): Promise<null | string> {
const paths = await window.hermesDesktop?.selectPaths?.({
title: translateNow('profiles.importProfile'),
multiple: false,
filters: ARCHIVE_FILTERS
})
const archive = paths?.[0]
if (!archive) {
return null
}
try {
const name = await importProfileBundle(archive)
notify({ kind: 'success', title: translateNow('profiles.imported'), message: name })
// Same landing as CreateProfileDialog's onCreated: refresh the list, then
// switch into the new profile on a fresh chat.
await refreshActiveProfile()
selectProfile(name)
return name
} catch (error) {
notifyError(error, translateNow('profiles.failedImport'))
return null
}
}
+41 -1
View File
@@ -22,6 +22,7 @@ import {
$sessions,
idsShareLineage,
sessionMatchesStoredId,
setSessions,
workspaceCwdForNewSession
} from '@/store/session'
import { $focusedSessionState, $focusedStoredSessionId } from '@/store/session-states'
@@ -173,7 +174,7 @@ export function exitProjectScope(): void {
// one. Empty for the path-less Home bucket. (The sidebar's `projectTreeCwd` is
// the same rule over the same tree — this is the store-side copy so the store
// doesn't reach into the sidebar's React module.)
const projectRootCwd = (project: SidebarProjectTree | undefined): string =>
export const projectRootCwd = (project: SidebarProjectTree | undefined): string =>
(project?.path || project?.repos.find(repo => repo.path)?.path || '').trim()
// ⌘K "go to project": flip the sidebar into grouped mode and enter the project
@@ -520,6 +521,45 @@ export async function fetchProjectSessions(projectId: string): Promise<SidebarPr
}
}
interface WorkspaceMovePayload {
branch?: null | string
cwd?: string
git_repo_root?: null | string
}
// Re-home a stored session into another project's root folder — the fix for a
// chat created in the wrong directory. The backend replaces cwd + git identity
// (so the tree's grouping follows) and re-anchors any live agent bound to the
// row; here we mirror the move into the `$sessions` cache so both the flat list
// and the grouped tree reflect it before the next authoritative refresh.
export async function moveSessionToProject(
sessionId: string,
projectId: string,
profile?: null | string
): Promise<void> {
const cwd = projectRootCwd($projectTree.get().find(node => node.id === projectId))
if (!cwd) {
throw new Error(translateNow('sidebar.projects.moveNoFolder'))
}
const res = await gatewayRequest<WorkspaceMovePayload>('session.workspace.move', {
cwd,
session_key: sessionId,
...(profile ? { profile } : {})
})
const moved = res.cwd || cwd
setSessions(prev =>
prev.map(s =>
sessionMatchesStoredId(s, sessionId)
? { ...s, cwd: moved, git_branch: res.branch ?? null, git_repo_root: res.git_repo_root ?? null }
: s
)
)
void refreshProjectTree()
}
export interface RepoDiscoveryPolicy {
enabled: boolean
roots: string[]
+20
View File
@@ -871,6 +871,26 @@ export interface ProfileSetupCommand {
command: string
}
// The desktop appearance/interface overlay bundled into a profile export as
// `desktop.json`. Everything optional — an archive exported by an older (or
// non-desktop) Hermes simply carries none of it. See store/profile-share.ts.
export interface ProfileDesktopOverlay {
/** Overlay schema version (1). */
version?: number
/** Skin name (built-in or bundled user theme). */
skin?: string
/** Light/dark/system preference. */
mode?: string
/** Full user-theme definitions the skin may reference (DesktopTheme JSON). */
themes?: Record<string, unknown>
/** Rail color override for this profile. */
profileColor?: null | string
/** Layout tree (hermes.desktop.layoutTree.v2 shape). */
layoutTree?: unknown
/** Active layout preset id. */
layoutPreset?: string
}
// ── Projects ───────────────────────────────────────────────────────────────
// A first-class, per-profile, human-named workspace spanning one or more
// folders. Mirrors hermes_cli/projects_db.Project.to_dict().
+4
View File
@@ -10249,6 +10249,10 @@ class HermesCLI(CLIAgentSetupMixin, CLICommandsMixin, CLIBillingMixin):
self._handle_rollback_command(cmd_original)
elif canonical == "snapshot":
self._handle_snapshot_command(cmd_original)
elif canonical == "export":
self._handle_export_command(cmd_original)
elif canonical == "import":
self._handle_import_command(cmd_original)
elif canonical == "stop":
self._handle_stop_command()
elif canonical == "agents":
+2
View File
@@ -0,0 +1,2 @@
zyz619963502zyz
# PR #73449 salvage → #78815
+1
View File
@@ -0,0 +1 @@
rapsealk
+74
View File
@@ -363,6 +363,80 @@ class CLICommandsMixin:
print(f" Unknown subcommand: {subcmd}")
print(" Usage: /snapshot [list|create [label]|restore <id>|prune [N]]")
def _handle_export_command(self, command: str):
"""Handle /export — export a profile to a shareable .tar.gz archive.
Syntax:
/export — export the active profile
/export <profile> — export a named profile
/export [profile] -o <path> — choose the output path
"""
from hermes_cli.profiles import export_profile, get_active_profile_name
parts = command.split()[1:]
output = None
if "-o" in parts:
idx = parts.index("-o")
if idx + 1 >= len(parts):
print(" Usage: /export [profile] [-o output.tar.gz]")
return
output = parts[idx + 1]
parts = parts[:idx] + parts[idx + 2:]
name = parts[0] if parts else (get_active_profile_name() or "default")
if not output:
output = f"{name}.tar.gz"
try:
result = export_profile(name, output)
print(f" ✓ Exported '{name}' to {result}")
print(" Share it: the other user runs /import or `hermes profile import <archive>`.")
except (ValueError, FileNotFoundError) as e:
print(f" Error: {e}")
def _handle_import_command(self, command: str):
"""Handle /import — import a shared profile archive as a new profile.
Syntax:
/import <archive.tar.gz> [--name <name>]
"""
from hermes_cli.profiles import (
check_alias_collision, create_wrapper_script, import_profile,
)
parts = command.split()[1:]
name = None
if "--name" in parts:
idx = parts.index("--name")
if idx + 1 >= len(parts):
print(" Usage: /import <archive.tar.gz> [--name <name>]")
return
name = parts[idx + 1]
parts = parts[:idx] + parts[idx + 2:]
if not parts:
print(" Usage: /import <archive.tar.gz> [--name <name>]")
return
archive = " ".join(parts) # paths may contain spaces
try:
profile_dir = import_profile(archive, name=name)
except (ValueError, FileExistsError, FileNotFoundError) as e:
print(f" Error: {e}")
return
imported = profile_dir.name
print(f" ✓ Imported profile '{imported}' at {profile_dir}")
try:
if not check_alias_collision(imported):
wrapper_path = create_wrapper_script(imported)
if wrapper_path:
print(f" Wrapper created: {wrapper_path}")
except Exception:
pass
print(f" Use it: hermes -p {imported}")
def _handle_stop_command(self):
"""Handle /stop — kill all running background processes and
background (async) delegations.
+4
View File
@@ -133,6 +133,10 @@ COMMAND_REGISTRY: list[CommandDef] = [
args_hint="[number]"),
CommandDef("snapshot", "Create or restore state snapshots of Hermes config/state", "Session",
cli_only=True, aliases=("snap",), args_hint="[create|restore <id>|prune]"),
CommandDef("export", "Export a profile (config, skills, theme) to a shareable archive", "Configuration",
cli_only=True, args_hint="[profile] [-o output.tar.gz]"),
CommandDef("import", "Import a shared profile archive as a new profile", "Configuration",
cli_only=True, args_hint="<archive.tar.gz> [--name <name>]"),
CommandDef("stop", "Kill all running background processes", "Session",
busy_policy="interrupt_then_dispatch", busy_handler="stop"),
CommandDef("approve", "Approve a pending dangerous command", "Session",
+34 -1
View File
@@ -372,6 +372,35 @@ def _primary_log_path(log_name: str) -> Optional[Path]:
return (get_hermes_home() / "logs" / filename) if filename else None
# Logs written by a client process rather than by this backend. When the
# desktop app talks to a remote/docker/SSH backend, `hermes debug share` runs
# on the *backend* and can never see them — a bare "(file not found)" then
# reads as "the app logged nothing" and sends triage down a dead end, which is
# exactly the wrong answer when the client is the thing being debugged.
_CLIENT_SIDE_LOGS = {
"desktop": (
"written by Hermes Desktop on the machine running the app, not by this "
"backend. If the desktop connects to a remote/docker/SSH backend, collect "
"it on that client machine"
),
}
def _missing_log_note(log_name: str) -> str:
"""Explain a missing log instead of stating a bare absence.
For a client-side log the absence is expected on a remote backend, so the
note names the writer and the path to collect by hand.
"""
reason = _CLIENT_SIDE_LOGS.get(log_name)
if reason is None:
return "(file not found)"
primary = _primary_log_path(log_name)
where = f" — expected at {primary}" if primary else ""
return f"(not on this host: {reason}{where})"
def _resolve_log_path(log_name: str) -> Optional[Path]:
"""Find the log file for *log_name*, falling back to the .1 rotation.
@@ -432,7 +461,11 @@ def _capture_log_snapshot(
log_path = _resolve_log_path(log_name)
if log_path is None:
primary = _primary_log_path(log_name)
tail = "(file empty)" if primary and primary.exists() else "(file not found)"
tail = (
"(file empty)"
if primary and primary.exists()
else _missing_log_note(log_name)
)
return LogSnapshot(path=None, tail_text=tail, full_text=None)
try:
+3
View File
@@ -6876,6 +6876,9 @@ def _desktop_linux_needs_no_sandbox() -> bool:
unprivileged desktop user on an AppArmor-restricted host. The root case
should remain an explicit user choice.
"""
if os.environ.get("ELECTRON_DISABLE_SANDBOX", 0) == "1":
return True
if sys.platform != "linux":
return False
if hasattr(os, "geteuid") and os.geteuid() == 0:
+76
View File
@@ -109,6 +109,19 @@ _MATCHING_PREFIX_STRIP_PROVIDERS: frozenset[str] = frozenset({
"xai",
})
# Providers whose API serves ``vendor/model`` ids but whose endpoint can also
# front arbitrary self-hosted models, so a bare name cannot be prefixed
# blindly. A bare id is repaired only when the curated catalogue for that
# provider holds exactly one entry ending in ``/<name>`` — a lookup, not a
# guess. NVIDIA NIM is the case in hand: build.nvidia.com serves
# ``nvidia/nemotron-…`` (and third-party ``z-ai/glm-…``), while the same
# provider id also points at local NIM containers with their own naming.
# Without this repair a bare ``nemotron-3-ultra-550b-a55b`` reaches the API
# and returns a bare ``404 page not found`` that never names the model (#78796).
_CATALOGUE_PREFIX_REPAIR_PROVIDERS: frozenset[str] = frozenset({
"nvidia",
})
# Providers whose APIs require lowercase model IDs. Xiaomi's
# ``api.xiaomimimo.com`` rejects mixed-case names like ``MiMo-V2.5-Pro``
# that users might copy from marketing docs — it only accepts
@@ -350,6 +363,63 @@ def _prepend_vendor(model_name: str) -> str:
return model_name
def _repair_prefix_from_catalogue(model_name: str, provider: str) -> str:
"""Restore a dropped ``vendor/`` prefix using the provider's catalogue.
Unlike :func:`_prepend_vendor`, this never guesses from the model's name
shape — it only repairs a bare id that matches **exactly one** curated
entry for this provider modulo the prefix. That keeps self-hosted models
behind the same provider id (local NIM containers, proxies) untouched,
since they aren't in the catalogue.
Examples::
>>> _repair_prefix_from_catalogue("nemotron-3-ultra-550b-a55b", "nvidia")
'nvidia/nemotron-3-ultra-550b-a55b'
>>> _repair_prefix_from_catalogue("my-local-nim", "nvidia")
'my-local-nim'
"""
if "/" in model_name:
return model_name
try:
from hermes_cli.models import _PROVIDER_MODELS
except Exception:
return model_name
catalogue = _PROVIDER_MODELS.get(provider) or []
# Compare against the catalogue's own suffix, tag included: a bare
# ``…:free`` id must resolve to the ``:free`` entry, not its paid sibling.
needle = model_name.strip().lower()
matches = {
entry
for entry in catalogue
if "/" in entry and entry.split("/", 1)[1].strip().lower() == needle
}
if len(matches) == 1:
return matches.pop()
return model_name
def suggest_prefixed_model_id(provider: str, model_name: str) -> Optional[str]:
"""Return the prefixed catalogue id for a bare *model_name*, if unambiguous.
The diagnostic counterpart to :func:`_repair_prefix_from_catalogue`: used
to explain a provider's content-free 404 when the configured id lost its
``vendor/`` prefix. Returns ``None`` when the name already has a prefix,
the provider has no curated catalogue, or nothing matches — so callers can
stay silent rather than guess (#78796).
"""
name = (model_name or "").strip()
if not name or "/" in name:
return None
try:
canonical = _normalize_provider_alias(provider)
except Exception:
return None
repaired = _repair_prefix_from_catalogue(name, canonical)
return repaired if repaired != name else None
# ---------------------------------------------------------------------------
# Main normalisation entry point
# ---------------------------------------------------------------------------
@@ -492,6 +562,12 @@ def normalize_model_for_provider(model_input: str, target_provider: str) -> str:
result = result.lower()
return result
# --- Catalogue-backed prefix repair: restore a dropped ``vendor/`` on a
# bare id that matches exactly one curated entry. Unknown names (a
# local NIM container, a proxied model) pass through untouched. ---
if provider in _CATALOGUE_PREFIX_REPAIR_PROVIDERS:
return _repair_prefix_from_catalogue(name, provider)
# --- Authoritative native providers: preserve user-facing slugs as-is ---
if provider in _AUTHORITATIVE_NATIVE_PROVIDERS:
return name
+37 -5
View File
@@ -30,7 +30,7 @@ import sys
import time
from dataclasses import dataclass
from pathlib import Path, PurePosixPath, PureWindowsPath
from typing import List, Optional, Tuple
from typing import Dict, List, Optional, Tuple
from agent.skill_utils import is_excluded_skill_path
@@ -237,6 +237,9 @@ _DEFAULT_EXPORT_INCLUDE_ROOT = frozenset({
# Configuration / persona
"config.yaml", "SOUL.md", "MEMORY.md", "USER.md", "todo.json",
"system_prompt.md", "AGENTS.md", "CLAUDE.md", ".cursorrules",
# Desktop appearance/interface overlay (written by the desktop app's
# profile export; applied by its import — see desktop.json handling).
"desktop.json",
# User-facing skill, cron, and session artifacts
"skills", "cron", "scripts", "sessions",
# Plugin / memory surfaces (per-profile overrides live here)
@@ -1898,9 +1901,29 @@ def _default_export_ignore(root_dir: Path):
return _ignore
def export_profile(name: str, output_path: str) -> Path:
def _make_profile_archive(base: str, root_dir: str, base_dir: str) -> str:
"""Create ``<base>.tar.gz`` of ``root_dir/base_dir`` — GNU tar format.
Not :func:`shutil.make_archive`: that writes PAX (Python's tarfile default
since 3.8), whose fractional-mtime records macOS Archive Utility rejects —
double-clicking an exported profile threw "Error 94 - Bad message." GNU
format keeps long paths working (longlink extensions) and stays integer-
mtime, so Finder, bsdtar, and gnutar all extract it.
"""
import tarfile
archive_path = f"{base}.tar.gz"
with tarfile.open(archive_path, "w:gz", format=tarfile.GNU_FORMAT) as tf:
tf.add(str(Path(root_dir) / base_dir), arcname=base_dir)
return archive_path
def export_profile(name: str, output_path: str, extra_files: Optional[Dict[str, str]] = None) -> Path:
"""Export a profile to a tar.gz archive.
``extra_files`` maps root-relative filenames (e.g. ``desktop.json``) to
text content staged into the archive alongside the profile's own files —
the desktop app uses it to bundle its appearance/interface overlay.
Returns the output file path.
"""
import tempfile
@@ -1912,9 +1935,16 @@ def export_profile(name: str, output_path: str) -> Path:
raise FileNotFoundError(f"Profile '{canon}' does not exist.")
output = Path(output_path)
# shutil.make_archive wants the base name without extension
# Archive base name without extension (.tar.gz appended by the writer).
base = str(output).removesuffix(".tar.gz").removesuffix(".tgz")
def _stage_extras(staged: Path) -> None:
for rel, content in (extra_files or {}).items():
parts = _normalize_profile_archive_parts(rel)
target = staged.joinpath(*parts)
target.parent.mkdir(parents=True, exist_ok=True)
target.write_text(content, encoding="utf-8")
if canon == "default":
# The default profile IS ~/.hermes itself — its parent is ~/ and its
# directory name is ".hermes", not "default". We stage a clean copy
@@ -1927,7 +1957,8 @@ def export_profile(name: str, output_path: str) -> Path:
symlinks=True,
ignore=_default_export_ignore(profile_dir),
)
result = shutil.make_archive(base, "gztar", tmpdir, "default")
_stage_extras(staged)
result = _make_profile_archive(base, tmpdir, "default")
return Path(result)
# Named profiles — stage a filtered copy to exclude credentials
@@ -1940,7 +1971,8 @@ def export_profile(name: str, output_path: str) -> Path:
symlinks=True,
ignore=lambda d, contents: _CREDENTIAL_FILES & set(contents),
)
result = shutil.make_archive(base, "gztar", tmpdir, canon)
_stage_extras(staged)
result = _make_profile_archive(base, tmpdir, canon)
return Path(result)
+16
View File
@@ -578,6 +578,22 @@ class ProfileRename(BaseModel):
new_name: str
class ProfileExport(BaseModel):
# Optional extra root-level files to stage into the archive, filename →
# text content (e.g. desktop.json — the desktop appearance overlay).
extra_files: Dict[str, str] = {}
# Where to write the archive. Empty → a staging path under HERMES_HOME.
output: str = ""
class ProfileImport(BaseModel):
# Path to a profile .tar.gz on the backend's filesystem (the desktop's
# local/pooled backends share the machine with the picker dialog).
archive: str
# Override the profile name inferred from the archive root.
name: Optional[str] = None
class ProfileSoulUpdate(BaseModel):
content: str
+104
View File
@@ -26,6 +26,8 @@ from hermes_cli.web_deps import late
from hermes_cli.web_models import (
ProfileCreate,
ProfileActiveUpdate,
ProfileExport,
ProfileImport,
ProfileRename,
ProfileSoulUpdate,
ProfileDescriptionUpdate,
@@ -687,3 +689,105 @@ async def describe_profile_auto_endpoint(name: str, body: ProfileDescribeAuto):
# auto-generated.
"description_auto": bool(outcome.ok),
}
# ── Export / Import ──────────────────────────────────────────────────────────
# Profile sharing for the desktop: wraps hermes_cli.profiles.export_profile /
# import_profile (the same machinery behind `hermes profile export|import`).
# Paths are exchanged, not bytes — the desktop's local and pooled backends
# share the filesystem with the native save/open dialogs that produce them.
@router.post("/api/profiles/{name}/export")
async def export_profile_endpoint(name: str, body: ProfileExport):
from hermes_cli import profiles as profiles_mod
output = (body.output or "").strip()
if not output:
from hermes_constants import get_hermes_home
staging = get_hermes_home() / "profile-exports"
try:
staging.mkdir(parents=True, exist_ok=True)
except OSError as exc:
raise HTTPException(status_code=500, detail=f"Could not create export directory: {exc}")
stamp = time.strftime("%Y%m%d-%H%M%S")
output = str(staging / f"{profiles_mod.normalize_profile_name(name)}-{stamp}.tar.gz")
loop = asyncio.get_running_loop()
try:
result = await loop.run_in_executor(
None,
lambda: profiles_mod.export_profile(name, output, extra_files=body.extra_files or None),
)
except FileNotFoundError as e:
raise HTTPException(status_code=404, detail=str(e))
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
except Exception as e:
_log.exception("POST /api/profiles/%s/export failed", name)
raise HTTPException(status_code=500, detail=str(e))
return {"ok": True, "archive": str(result)}
@router.post("/api/profiles/import")
async def import_profile_endpoint(body: ProfileImport):
from hermes_cli import profiles as profiles_mod
archive = (body.archive or "").strip()
if not archive:
raise HTTPException(status_code=400, detail="archive path is required")
loop = asyncio.get_running_loop()
try:
profile_dir = await loop.run_in_executor(
None,
lambda: profiles_mod.import_profile(archive, name=(body.name or "").strip() or None),
)
except FileNotFoundError as e:
raise HTTPException(status_code=404, detail=str(e))
except (ValueError, FileExistsError) as e:
raise HTTPException(status_code=400, detail=str(e))
except Exception as e:
_log.exception("POST /api/profiles/import failed")
raise HTTPException(status_code=500, detail=str(e))
imported = profile_dir.name
# Match the CLI import flow: create the wrapper alias when it's safe.
try:
if not profiles_mod.check_alias_collision(imported):
profiles_mod.create_wrapper_script(imported)
except Exception:
_log.exception("Creating wrapper for imported profile %s failed", imported)
# Surface the bundled desktop appearance overlay (if the archive carried
# one) so the desktop can apply theme/interface prefs without re-reading
# the file over another round-trip.
desktop_overlay = None
overlay_path = profile_dir / "desktop.json"
if overlay_path.is_file():
try:
import json as _json
desktop_overlay = _json.loads(overlay_path.read_text(encoding="utf-8"))
except Exception:
_log.exception("Reading desktop.json from imported profile %s failed", imported)
return {
"ok": True,
"name": imported,
"path": str(profile_dir),
"desktop": desktop_overlay,
}
@router.get("/api/profiles/{name}/desktop-overlay")
async def get_profile_desktop_overlay(name: str):
"""The desktop appearance/interface overlay bundled with an imported
profile (``desktop.json`` at the profile root), or ``exists: false``."""
overlay_path = _resolve_profile_dir(name) / "desktop.json"
if not overlay_path.is_file():
return {"exists": False, "desktop": None}
try:
import json as _json
return {"exists": True, "desktop": _json.loads(overlay_path.read_text(encoding="utf-8"))}
except Exception as e:
raise HTTPException(status_code=500, detail=f"Could not read desktop.json: {e}")
+96 -5
View File
@@ -3649,7 +3649,12 @@ class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)
return False
def update_session_cwd(
self, session_id: str, cwd: str, git_branch: str = None, git_repo_root: str = None
self,
session_id: str,
cwd: str,
git_branch: str = None,
git_repo_root: str = None,
replace_git_meta: bool = False,
) -> None:
"""Persist the session working directory when a frontend knows it.
@@ -3664,6 +3669,11 @@ class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)
every surface reads the same membership instead of re-probing git in the
GUI over a partial page. Each field is only written when non-empty so a
probe failure never clobbers a previously-captured value.
``replace_git_meta`` inverts that non-empty rule: a deliberate workspace
MOVE (re-homing a session into another project) must overwrite the old
repo identity even when the new cwd resolves to none — keeping the stale
root would leave the session grouped under the project it just left.
"""
if not session_id or not cwd:
return
@@ -3673,12 +3683,12 @@ class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)
sets = ["cwd = ?"]
params: List[Any] = [cwd]
if branch:
if branch or replace_git_meta:
sets.append("git_branch = ?")
params.append(branch)
if repo_root:
params.append(branch or None)
if repo_root or replace_git_meta:
sets.append("git_repo_root = ?")
params.append(repo_root)
params.append(repo_root or None)
params.append(session_id)
def _do(conn):
@@ -5478,6 +5488,80 @@ class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)
rowcount = self._execute_write(_do)
return rowcount > 0
def set_session_read(self, session_id: str, read: bool = True) -> bool:
"""Mark a session read or unread (and its whole compression lineage).
Read state is a watermark, not a flag: ``last_read_at`` records when
the conversation was last read, and it counts as unread when activity
postdates that watermark (the derived ``unread`` key on
:meth:`list_sessions_rich` rows). New messages therefore flip a read
conversation back to unread without any write on the message path.
Three states:
* NULL — never tracked (every pre-feature row): treated as read, so
shipping the column doesn't badge a user's entire history at once.
* 0 — explicitly marked unread: any activity postdates it.
* timestamp — read up to that moment.
Like :meth:`set_session_archived` / :meth:`set_session_pinned`, the
whole compression chain is stamped as a unit, so reading the surfaced
tip clears the root (and vice-versa) no matter which id the caller
holds. Returns True when at least one row changed.
"""
def _do(conn):
cursor = conn.execute(
"""
WITH RECURSIVE
ancestors(id) AS (
SELECT ?
UNION
SELECT parent.id
FROM ancestors a
JOIN sessions child ON child.id = a.id
JOIN sessions parent ON parent.id = child.parent_session_id
WHERE parent.end_reason = 'compression'
),
descendants(id) AS (
SELECT ?
UNION
SELECT child.id
FROM descendants d
JOIN sessions parent ON parent.id = d.id
JOIN sessions child ON child.parent_session_id = parent.id
WHERE parent.end_reason = 'compression'
),
lineage(id) AS (
SELECT id FROM ancestors
UNION
SELECT id FROM descendants
)
UPDATE sessions
SET last_read_at = ?
WHERE id IN (SELECT id FROM lineage)
""",
(session_id, session_id, time.time() if read else 0.0),
)
rowcount = cursor.rowcount
if rowcount is None or rowcount < 0:
rowcount = conn.execute("SELECT changes()").fetchone()[0]
return rowcount
rowcount = self._execute_write(_do)
return rowcount > 0
@staticmethod
def session_unread(session_row: Dict[str, Any]) -> bool:
"""Derive unread from a session row's watermark and activity.
Shared by ``list_sessions_rich`` and any future surface that holds a
row (or projected row) with ``last_read_at`` and ``last_active``.
NULL watermark = never tracked = read.
"""
last_read = session_row.get("last_read_at")
if last_read is None:
return False
last_active = session_row.get("last_active") or session_row.get("started_at")
return float(last_active or 0) > float(last_read)
def get_session_by_title(self, title: str) -> Optional[Dict[str, Any]]:
"""Look up a session by exact title. Returns session dict or None."""
with self._read_ctx() as conn:
@@ -5994,6 +6078,13 @@ class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)
projected.append(merged)
sessions = projected
# Derive read state per surfaced conversation. ``last_read_at`` is
# lineage-stamped by set_session_read, so a projected row's root
# watermark and its tip's are the same value — comparing it against
# the tip's last_active is correct either way.
for s in sessions:
s["unread"] = self.session_unread(s)
return sessions
# =========================================================================
+1
View File
@@ -245,6 +245,7 @@ CREATE TABLE IF NOT EXISTS sessions (
rewind_count INTEGER NOT NULL DEFAULT 0,
archived INTEGER NOT NULL DEFAULT 0,
pinned INTEGER NOT NULL DEFAULT 0,
last_read_at REAL,
FOREIGN KEY (parent_session_id) REFERENCES sessions(id),
FOREIGN KEY (system_prompt_hash) REFERENCES system_prompts(hash)
);
+1 -4
View File
@@ -27,10 +27,7 @@
mkdir -p $out/bin
install -Dm755 ${../hermes} $out/bin/hermes
'')
(pkgs.runCommand "dev-sandbox" { } ''
mkdir -p $out/bin
install -Dm755 ${../scripts/dev-sandbox.sh} $out/bin/sandbox
'')
self'.packages.sandbox
uv
# Headless Wayland compositor for E2E tests (test:e2e:visual).
# cage renders a single client with no window management, so
+5
View File
@@ -9,6 +9,9 @@
...
}:
let
sandbox = pkgs.callPackage ./sandbox.nix { };
minimal = pkgs.callPackage ./hermes-agent.nix {
inherit (inputs) uv2nix pyproject-nix pyproject-build-systems;
npm-lockfile-fix = inputs'.npm-lockfile-fix.packages.default;
@@ -51,6 +54,8 @@
}).node-gyp;
default = full;
inherit sandbox;
inherit minimal;
# Ships discord.py + python-telegram-bot + slack-sdk so a plain
+124
View File
@@ -0,0 +1,124 @@
{
# electron deps
alsa-lib,
at-spi2-atk,
atk,
cairo,
cups,
dbus,
expat,
fontconfig,
freetype,
glib,
gtk3,
libdrm,
libgbm,
libxkbcommon,
mesa,
nspr,
nss,
pango,
systemd,
libX11,
libXcomposite,
libXdamage,
libXext,
libXfixes,
libXrandr,
libXrender,
libXtst,
libxcb,
# sandbox deps
bash,
bubblewrap,
cacert,
coreutils,
curl,
gawk,
git,
glibc,
gnumake,
gnugrep,
gnused,
gzip,
nodejs_22,
openssl,
python3,
slirp4netns,
stdenv,
gnutar,
util-linux,
# etc
writeShellApplication,
lib,
}:
let
electronRuntime = [
alsa-lib
at-spi2-atk
atk
cairo
cups
dbus
expat
fontconfig
freetype
glib
gtk3
libdrm
libgbm
libxkbcommon
mesa
nspr
nss
pango
systemd
libX11
libXcomposite
libXdamage
libXext
libXfixes
libXrandr
libXrender
libXtst
libxcb
];
in
writeShellApplication {
name = "sandbox";
runtimeInputs = [
bash
bubblewrap
cacert
coreutils
curl
gawk
git
glibc.bin
gnumake
gnugrep
gnused
gzip
nodejs_22
openssl
python3
slirp4netns
stdenv.cc
gnutar
util-linux
]
++ electronRuntime;
text = ''
export DEV_SANDBOX_REAL_CA_CERT=${cacert}/etc/ssl/certs/ca-bundle.crt
export DEV_SANDBOX_DYNAMIC_LINKER=${stdenv.cc.bintools.dynamicLinker}
export DEV_SANDBOX_NODE_DIR=${nodejs_22}
export DEV_SANDBOX_ELECTRON_LD_LIBRARY_PATH=${lib.makeLibraryPath electronRuntime}
# The script is imported into the store as a single file, so its own
# directory has no scripts/sandbox/ beside it. Point it at the assets
# (fake-internet proxy, ssh shim) explicitly.
export DEV_SANDBOX_ASSETS=${../scripts/sandbox}
exec ${../scripts/dev-sandbox.sh} "$@"
'';
}
+24
View File
@@ -23,6 +23,7 @@ import subprocess
import tempfile
import threading
import time
import traceback
from collections import defaultdict
from contextlib import suppress
from typing import Callable, Dict, List, Optional, Any, Tuple
@@ -3040,6 +3041,29 @@ class DiscordAdapter(BasePlatformAdapter):
"""
if not self._client:
return SendResult(success=False, error="Not connected")
if not (content or "").strip():
logger.warning(
"[%s] Dropped empty message to chat=%s (caller bug). Call site:\n%s",
self.name,
chat_id,
"".join(traceback.format_stack(limit=12)[:-1]),
)
result = SendResult(
success=False,
error="Refusing to send empty message",
)
# Mirror the exception path's recovery bookkeeping. Missed-message
# backfill decides what to replay from this table, so a dropped
# final reply must be recorded as failed — otherwise the reply is
# both never sent and never retried.
await asyncio.to_thread(
self._record_discord_response,
reply_to=reply_to,
result=result,
content=content,
final=bool(metadata and metadata.get("notify")),
)
return result
try:
# Determine target channel: thread_id in metadata takes precedence.
+175 -1
View File
@@ -5390,6 +5390,168 @@ class AIAgent:
return True
def _try_refresh_env_client_credentials(self) -> bool:
"""Adopt ~/.hermes/.env credential/base-url edits at the turn boundary.
A Settings save (desktop ``PUT /api/env``, ``hermes setup``) updates
``.env`` and the *saving* process's os.environ, but a live session
worker keeps the base_url/api_key captured at agent init until it
restarts — so an open chat silently keeps calling the old endpoint
(#67821). Called at the start of each conversation turn, this
re-resolves the provider's env-sourced credentials (load_env() is
mtime-memoized, so an unchanged file costs one stat()) and rebuilds
the client when the user edited them.
Reacts only to env *edits* (resolved values changed since the last
look), never to mere divergence from the agent's current values —
credential-pool rotation and failover legitimately move the session
off the env credential, and stomping those back every turn would
flap. A config.yaml ``model.base_url`` (or a pool entry with a
custom endpoint) also wins: edits are only adopted while the
session's current base_url is still the registry default or the
previously-seen env value.
Covers api-key registry providers and named custom providers with a
``key_env`` (#67935) — the latter resolve to ``provider="custom"``
with no registry entry, so they are matched through the runtime
provider's config lookup instead.
"""
if self.api_mode != "chat_completions":
return False
if getattr(self, "_fallback_activated", False):
return False
try:
from agent.credential_pool import get_env_prefer_dotenv
from hermes_cli.auth import PROVIDER_REGISTRY
except ImportError:
return False
pconfig = PROVIDER_REGISTRY.get(self.provider)
if (
pconfig
and getattr(pconfig, "auth_type", "") == "api_key"
and getattr(pconfig, "api_key_env_vars", ())
):
api_key = ""
for env_var in pconfig.api_key_env_vars:
api_key = get_env_prefer_dotenv(env_var).strip()
if api_key:
break
if not api_key:
return False
env_url = ""
if pconfig.base_url_env_var:
env_url = get_env_prefer_dotenv(pconfig.base_url_env_var).strip().rstrip("/")
default_base = (pconfig.inference_base_url or "").strip().rstrip("/")
base_url = env_url or default_base
if self.provider == "kimi-coding":
from hermes_cli.auth import _resolve_kimi_base_url
base_url = _resolve_kimi_base_url(
api_key, pconfig.inference_base_url, env_url
).rstrip("/")
elif self.provider == "zai":
from hermes_cli.auth import _resolve_zai_base_url
base_url = _resolve_zai_base_url(
api_key, pconfig.inference_base_url, env_url
).rstrip("/")
elif self.provider == "custom":
# Named custom provider (#67935): identity lives in config
# (``providers.<name>`` / ``custom_providers``), the credential in
# the env var it names via ``key_env``. Re-resolve through the
# same config lookup the runtime resolver uses; entries without
# ``key_env`` (inline ``api_key``, pool-backed) have no
# env-sourced credential to watch.
try:
from hermes_cli.runtime_provider import _get_named_custom_provider
except ImportError:
return False
custom_provider = _get_named_custom_provider(
getattr(self, "requested_provider", "") or ""
)
if not custom_provider:
return False
key_env = str(custom_provider.get("key_env") or "").strip()
if not key_env:
return False
api_key = get_env_prefer_dotenv(key_env).strip()
if not api_key:
return False
# Custom providers pin their endpoint in config, not env — the
# config base_url is both the resolved and the "default" base, so
# only key edits are ever adopted here.
default_base = str(custom_provider.get("base_url") or "").strip().rstrip("/")
base_url = default_base
else:
return False
if not base_url:
return False
resolved = (base_url, api_key)
prev = getattr(self, "_env_creds_seen", None)
current_base = (self.base_url or "").strip().rstrip("/")
if prev is None:
# First look — no baseline to diff against. Adopt only the
# boot-default case (worker spawned before the user saved an
# override); anything else is unattributable on turn one.
adopt = current_base == default_base and not (
base_url == current_base and api_key == self.api_key
)
else:
# Env unchanged → no-op; any drift from self.* is rotation/
# failover or config precedence — leave it alone. An edit is
# only adopted while the session still runs on the registry
# default or the previously-seen env value.
adopt = (
resolved != prev
and current_base in {default_base, prev[0]}
and not (base_url == current_base and api_key == self.api_key)
)
if not adopt:
self._env_creds_seen = resolved
return False
from hermes_cli.route_identity import normalize_route_base_url
route_changed = normalize_route_base_url(self.base_url) != normalize_route_base_url(
base_url
)
prior_api_key = self.api_key
prior_base_url = self.base_url
prior_client_kwargs = dict(self._client_kwargs)
self.api_key = api_key
self.base_url = base_url
self._client_kwargs["api_key"] = self.api_key
self._client_kwargs["base_url"] = self.base_url
# A base-url change moves the route: TLS material and default
# headers derived from the old endpoint must be recomputed, exactly
# as on credential-pool rotation.
self._reapply_route_client_config(route_changed=route_changed)
if not self._replace_primary_openai_client(reason="env_credential_refresh"):
# Leave the baseline un-advanced so the unchanged edit is
# retried next turn, and roll the agent back so its state keeps
# matching the still-live old client.
self.api_key = prior_api_key
self.base_url = prior_base_url
self._client_kwargs.clear()
self._client_kwargs.update(prior_client_kwargs)
return False
self._env_creds_seen = resolved
logger.info(
"Applied updated .env credentials for %s: endpoint %s",
self.provider,
self.base_url,
)
return True
def _try_refresh_vertex_client_credentials(self) -> bool:
"""Re-mint the Vertex OAuth2 access token and rebuild the OpenAI client.
@@ -5743,6 +5905,19 @@ class AIAgent:
self.base_url = runtime_base.rstrip("/") if isinstance(runtime_base, str) else runtime_base
self._client_kwargs["api_key"] = self.api_key
self._client_kwargs["base_url"] = self.base_url
self._reapply_route_client_config(route_changed=route_changed)
self._replace_primary_openai_client(reason="credential_rotation")
def _reapply_route_client_config(self, *, route_changed: bool) -> None:
"""Recompute route-derived client kwargs for the current ``self.base_url``.
TLS material (``ssl_verify``/``ssl_ca_cert``) and default headers are
derived from the endpoint, not the credential — any client rebuild
that may have moved ``base_url`` must recompute them or the new
endpoint inherits configuration computed for the old one. Shared by
credential-pool rotation and the per-turn env refresh so the two
paths cannot drift.
"""
self._client_kwargs.pop("ssl_verify", None)
self._client_kwargs.pop("ssl_ca_cert", None)
try:
@@ -5766,7 +5941,6 @@ class AIAgent:
self.base_url,
apply_user_headers=not route_changed,
)
self._replace_primary_openai_client(reason="credential_rotation")
def _recover_with_credential_pool(
self,
+535 -142
View File
@@ -1,198 +1,591 @@
#!/usr/bin/env bash
# Run a Hermes instance in an isolated sandbox — separate HERMES_HOME,
# separate Electron userData, and a distinct Desktop app name so it doesn't compete
# with your main desktop instance's single-instance lock.
# Run a command in a disposable, network-isolated fake Internet.
#
# By default the sandbox is throwaway: a temp dir is created and removed on
# exit. Use --persistent to keep the sandbox across restarts (stored under
# .hermes-sandbox/ in the worktree git root).
#
# Usage:
# scripts/dev-sandbox.sh python -m hermes_cli.main
# scripts/dev-sandbox.sh hermes desktop
# scripts/dev-sandbox.sh electron .
# scripts/dev-sandbox.sh -- npm run dev # from apps/desktop/
# scripts/dev-sandbox.sh --persistent hermes desktop
# scripts/dev-sandbox.sh --persistent -- npm run dev
#
# Seed the sandbox HERMES_HOME from an existing directory (e.g. your main
# ~/.hermes) so config, sessions, skills, etc. are pre-populated:
# scripts/dev-sandbox.sh --from ~/.hermes hermes desktop
#
# Override the app name (default: HermesSandbox):
# HERMES_DEV_SANDBOX_NAME=Staging scripts/dev-sandbox.sh hermes desktop
#
# Override the persistent sandbox dir name (default: .hermes-sandbox):
# HERMES_DEV_SANDBOX_DIR=.staging-sandbox scripts/dev-sandbox.sh --persistent hermes desktop
# The command runs in private user, mount, PID, and network namespaces. This
# script is stage 1: it builds the sandbox tree, mints the fake CA, and creates
# the user+network namespaces with `unshare` (see the namespace plan further
# down), then re-execs into scripts/sandbox/stage2-run.sh, which adds the
# mount/pid namespaces with bubblewrap and runs the payload. Its only writable
# filesystem is SANDBOX_ROOT. HTTP(S) goes to a local static MITM proxy;
# github.com SSH uses a sandbox-local git-upload-pack shim; neither transport
# can reach the host network.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Helper files the sandbox needs: the stage-2 script it re-execs into, plus the
# files it copies in (the fake-internet proxy, the ssh shim, the openssl config).
# They sit next to this script in the repo, but the Nix wrapper installs the
# script into the store on its own, so it exports DEV_SANDBOX_ASSETS to point
# here.
SANDBOX_ASSETS="${DEV_SANDBOX_ASSETS:-$SCRIPT_DIR/sandbox}"
for asset in proxy.py ssh-shim.sh openssl.cnf stage2-run.sh; do
[ -f "$SANDBOX_ASSETS/$asset" ] || {
echo "error: missing sandbox asset: $SANDBOX_ASSETS/$asset" >&2
exit 1
}
done
print_help() {
cat <<'EOF'
Usage: dev-sandbox.sh [--persistent] [--from DIR] [--] <command...>
Usage: dev-sandbox.sh [options] [--] <command...>
dev-sandbox.sh install [options] [--] [installer arguments...]
Run a Hermes instance in an isolated sandbox.
Run COMMAND in a throwaway chroot-like bubblewrap sandbox. The sandbox has no
writable host mounts: only its own root, mounted at /work, is writable.
Options:
--persistent Keep the sandbox dir across restarts (under the worktree
git root, in .hermes-sandbox/). Without this flag the
sandbox is a temp dir that is removed on exit.
--from DIR Copy DIR into the sandbox HERMES_HOME as the starting
point (config, sessions, skills, etc.).
Ignored if the sandbox HERMES_HOME already has content
(e.g. reusing a --persistent sandbox) to avoid clobbering.
--delete Delete the existing persistent sandbox in .hermes-sandbox.
-h, --help Show this help message.
--persistent Keep the whole sandbox under .hermes-sandbox/.
--delete Delete the persistent sandbox (asks first).
--root Install as uid 0 with the root FHS layout: code in
/usr/local/lib/hermes-agent, command in
/usr/local/bin. Default is the user-level layout.
--from DIR One-time copy of DIR into the sandbox's $HOME.
Existing persistent sandboxes are never overwritten.
--http-root DIR Copy DIR into the fake web server root for this run.
Requests map to DIR/<host>/<path>; no URL is forwarded.
--installer PATH With `install`, serve PATH at the canonical install.sh
URL. Default: scripts/install.sh in this worktree.
--from-main With `install`, fetch the real upstream main installer
and repository, then advance fake main to this folder
after a successful install for update testing.
Shorthand for --install-ref refs/heads/main.
--install-ref REF Like --from-main, but installs REF instead of main:
a branch, a tag (v2026.7.7), or a SHA reachable from main.
Use it to test updating from an older release, not just
from the tip.
-h, --help Show this help.
Option order matters: every option above is consumed by THIS script, and
parsing stops at the first argument it does not recognize. Everything from
that point on is passed through to the command (or, with `install`, to the
installer). Put sandbox options first and separate installer arguments with
`--`, otherwise they arrive here and fail:
# WRONG — --from-main reaches install.sh, which rejects it
scripts/dev-sandbox.sh install --skip-setup --from-main
# RIGHT
scripts/dev-sandbox.sh install --from-main -- --skip-setup
Install layout: `install.sh` picks its layout from `id -u` alone, so uid is what
separates the two real-world Linux installs. By default the sandbox runs as an
unprivileged `hermes` user, giving the layout most people have —
$HERMES_HOME/hermes-agent plus a ~/.local/bin launcher. Pass --root for the FHS
one. Both are worth testing; they differ in more than paths (root also relocates
uv's Python to /usr/local/share for world-readability).
The fake web server signs certificates with a CA trusted only inside this
sandbox. HTTP_PROXY/HTTPS_PROXY send fixture URLs there first; other HTTP(S)
requests pass through the sandbox's rootless outbound network. SSH to github.com
runs a sandbox-local upload-pack shim, never your SSH config, agent,
known-hosts file, or authorized keys.
Fake github main always comes from this folder. If it has staged, unstaged, or
non-ignored untracked changes, the sandbox warns and creates a temporary local
commit containing them; it never stages or commits the real worktree.
Environment:
HERMES_DEV_SANDBOX_NAME Override the app name (default: HermesSandbox)
HERMES_DEV_SANDBOX_DIR Override the persistent dir name (default: .hermes-sandbox)
HERMES_DEV_SANDBOX_DIR Sandbox directory name, relative to the repo root
(default: .hermes-sandbox).
Examples:
dev-sandbox.sh hermes desktop
dev-sandbox.sh --persistent hermes desktop
dev-sandbox.sh --from ~/.hermes hermes desktop
dev-sandbox.sh -- npm run dev
# create a sandbox, install this branch as `main`, and then drop to a shell,
# skipping `hermes setup` & the browser tools for speed.
scripts/dev-sandbox.sh install --persistent -- --skip-setup --skip-browser
# Install the official upstream main. You're dropped into a shell where
# you can run `hermes update`.
scripts/dev-sandbox.sh install --persistent --from-main
EOF
}
PERSISTENT=false
DELETE=false
RUN_AS_USER=true
SEED_DIR=""
HTTP_ROOT=""
INSTALL_SHORTCUT=false
INSTALLER_PATH=""
# Which upstream commit the sandbox installs before the update routes run.
# Empty means "install this worktree's own installer" (no upstream fetch); set,
# it is anything git can resolve -- a branch, a tag (v2026.7.7), or a SHA
# reachable from main -- so "can a user two releases back still update?" is
# expressible. --from-main is shorthand for refs/heads/main.
INSTALL_REF=""
UPSTREAM_URL="${HERMES_DEV_SANDBOX_UPSTREAM:-https://github.com/NousResearch/hermes-agent.git}"
if [ "${1:-}" = install ]; then
INSTALL_SHORTCUT=true
shift
fi
while [ "$#" -gt 0 ]; do
case "$1" in
--persistent)
PERSISTENT=true
shift
;;
--persistent) PERSISTENT=true; shift ;;
--delete) DELETE=true; shift ;;
--root) RUN_AS_USER=false; shift ;;
--user) RUN_AS_USER=true; shift ;; # the default; accepted for symmetry
--from)
if [ "$#" -lt 2 ] || [[ "$2" == -* ]]; then
echo "error: --from requires a directory argument" >&2
exit 1
fi
SEED_DIR="$2"
shift 2
;;
--from=*)
SEED_DIR="${1#--from=}"
if [ -z "$SEED_DIR" ]; then
echo "error: --from requires a directory argument" >&2
exit 1
fi
shift
;;
--delete)
DELETE=true
shift
;;
-h|--help)
print_help
exit 0
;;
--)
shift
break
;;
*)
break
;;
[ "$#" -ge 2 ] || { echo 'error: --from needs a directory' >&2; exit 1; }
SEED_DIR="$2"; shift 2 ;;
--http-root)
[ "$#" -ge 2 ] || { echo 'error: --http-root needs a directory' >&2; exit 1; }
HTTP_ROOT="$2"; shift 2 ;;
--installer)
[ "$#" -ge 2 ] || { echo 'error: --installer needs a file' >&2; exit 1; }
INSTALLER_PATH="$2"; shift 2 ;;
--from-main) INSTALL_REF="refs/heads/main"; shift ;;
--install-ref)
[ "$#" -ge 2 ] || { echo 'error: --install-ref needs a value' >&2; exit 1; }
INSTALL_REF="$2"
shift 2 ;;
--from=*|--http-root=*|--installer=*|--install-ref=*)
key="${1%%=*}"; value="${1#*=}"
[ -n "$value" ] || { echo "error: $key needs a value" >&2; exit 1; }
case "$key" in
--from) SEED_DIR="$value" ;;
--http-root) HTTP_ROOT="$value" ;;
--installer) INSTALLER_PATH="$value" ;;
--install-ref) INSTALL_REF="$value" ;;
esac
shift ;;
-h|--help) print_help; exit 0 ;;
--) shift; break ;;
*) break ;;
esac
done
if [ -n "$SEED_DIR" ]; then
if [ ! -d "$SEED_DIR" ]; then
echo "error: --from dir '$SEED_DIR' does not exist" >&2
exit 1
fi
# Resolve to absolute path so it's valid after we cd later.
SEED_DIR="$(cd "$SEED_DIR" && pwd)"
fi
if [ "$#" -eq 0 ]; then
if [ "$INSTALL_SHORTCUT" = false ] && [ "$#" -eq 0 ]; then
print_help >&2
exit 1
fi
if [ -n "$INSTALLER_PATH" ] && [ "$INSTALL_SHORTCUT" = false ]; then
echo 'error: --installer is only valid with the install shortcut' >&2
exit 1
fi
if [ -n "$INSTALL_REF" ] && [ "$INSTALL_SHORTCUT" = false ]; then
echo 'error: --from-main / --install-ref are only valid with the install shortcut' >&2
exit 1
fi
if [ -n "$INSTALL_REF" ] && [ -n "$INSTALLER_PATH" ]; then
echo 'error: --from-main / --install-ref cannot be combined with --installer' >&2
exit 1
fi
SANDBOX_DIR_NAME="${HERMES_DEV_SANDBOX_DIR:-.hermes-sandbox}"
GIT_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || echo "$SCRIPT_DIR/..")"
for dir in "$SEED_DIR" "$HTTP_ROOT"; do
[ -z "$dir" ] || [ -d "$dir" ] || { echo "error: directory '$dir' does not exist" >&2; exit 1; }
done
GIT_ROOT="${HERMES_SANDBOX_SOURCE_ROOT:-$(git rev-parse --show-toplevel)}"
GIT_ROOT="$(cd "$GIT_ROOT" && pwd)"
PERSISTENT_SANDBOX_ROOT="$GIT_ROOT/$SANDBOX_DIR_NAME"
if [ "$INSTALL_SHORTCUT" = true ] && [ -z "$INSTALL_REF" ] && [ -z "$INSTALLER_PATH" ]; then
INSTALLER_PATH="$GIT_ROOT/scripts/install.sh"
fi
if [ -n "$INSTALLER_PATH" ] && [ ! -f "$INSTALLER_PATH" ]; then
echo "error: installer '$INSTALLER_PATH' does not exist" >&2
exit 1
fi
COMMIT="$(git -C "$GIT_ROOT" rev-parse --verify 'HEAD^{commit}')" || {
echo "error: current folder has no HEAD commit" >&2
exit 1
}
SANDBOX_DIR_NAME="${HERMES_DEV_SANDBOX_DIR:-.hermes-sandbox}"
PERSISTENT_ROOT="$GIT_ROOT/$SANDBOX_DIR_NAME"
if [ "$DELETE" = true ]; then
if [ -d "$PERSISTENT_SANDBOX_ROOT" ]; then
read -r -p "[sandbox] delete $PERSISTENT_SANDBOX_ROOT? [y/N] " REPLY
case "$REPLY" in
[yY]|[yY][eE][sS])
echo "[sandbox] deleting $PERSISTENT_SANDBOX_ROOT" >&2
rm -rf -- "$PERSISTENT_SANDBOX_ROOT"
;;
*)
echo "[sandbox] aborted" >&2
exit 1
;;
esac
else
echo "[sandbox] nothing to delete at $PERSISTENT_SANDBOX_ROOT" >&2
if [ ! -d "$PERSISTENT_ROOT" ]; then
echo "[sandbox] nothing to delete at $PERSISTENT_ROOT" >&2
exit 0
fi
read -r -p "[sandbox] delete $PERSISTENT_ROOT? [y/N] " reply
case "$reply" in
y|Y|yes|YES) rm -rf -- "$PERSISTENT_ROOT" ;;
*) echo '[sandbox] aborted' >&2; exit 1 ;;
esac
exit 0
fi
# Derive a per-worktree app name so multiple checkouts don't collide.
# Each worktree has its own toplevel path even though they share one repo,
# so we hash that path into a short, stable suffix.
WORKTREE_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || echo "$SCRIPT_DIR/..")"
WORKTREE_ROOT="$(cd "$WORKTREE_ROOT" && pwd)"
WORKTREE_HASH="$(printf '%s' "$WORKTREE_ROOT" | cksum | cut -d' ' -f1)"
WORKTREE_NAME="$(basename "$WORKTREE_ROOT")"
DEFAULT_SANDBOX_NAME="HermesSandbox-${WORKTREE_NAME}-${WORKTREE_HASH}"
SANDBOX_NAME="${HERMES_DEV_SANDBOX_NAME:-$DEFAULT_SANDBOX_NAME}"
if [ "$PERSISTENT" = true ]; then
SANDBOX_ROOT="$PERSISTENT_SANDBOX_ROOT"
SANDBOX_ROOT="$PERSISTENT_ROOT"
else
SANDBOX_ROOT="$(mktemp -d -t hermes-sandbox.XXXXXX)"
cleanup() { chmod -R u+w "$SANDBOX_ROOT"; rm -rf -- "$SANDBOX_ROOT"; }
trap cleanup EXIT INT TERM
fi
export HERMES_HOME="$SANDBOX_ROOT/hermes-home"
export HERMES_DESKTOP_USER_DATA_DIR="$SANDBOX_ROOT/user-data"
export HERMES_DESKTOP_APP_NAME="$SANDBOX_NAME"
mkdir -p "$HERMES_HOME" "$HERMES_DESKTOP_USER_DATA_DIR"
if [ -n "$SEED_DIR" ]; then
# Only seed when the sandbox HERMES_HOME is empty — avoids clobbering an
# existing persistent sandbox on re-run.
if [ -z "$(ls -A "$HERMES_HOME" 2>/dev/null)" ]; then
echo "[sandbox] seeding HERMES_HOME from $SEED_DIR" >&2
cp -a "$SEED_DIR/." "$HERMES_HOME/"
mkdir -p "$SANDBOX_ROOT"/{root,home,etc}
UPSTREAM_REPO=""
UPSTREAM_COMMIT=""
if [ -n "$INSTALL_REF" ]; then
echo "[sandbox] fetching upstream $INSTALL_REF for installer/update test" >&2
UPSTREAM_REPO="$(mktemp -d -t hermes-sandbox-upstream.XXXXXX)"
git -C "$UPSTREAM_REPO" init -q
# Fetch the ref as given. A branch or tag name resolves on its own; a raw SHA
# needs the remote to allow fetching it directly, so fall back to fetching
# main and resolving the SHA locally (which works for any commit that is an
# ancestor of main -- the interesting case for "update from N versions ago").
#
# Peel to ^{commit} in both cases: an annotated tag fetches as a tag OBJECT,
# and using it directly fails later with "trying to write non-commit object
# ... to branch 'refs/heads/main'".
if git -C "$UPSTREAM_REPO" fetch -q "$UPSTREAM_URL" "$INSTALL_REF" 2>/dev/null; then
UPSTREAM_COMMIT="$(git -C "$UPSTREAM_REPO" rev-parse "FETCH_HEAD^{commit}")"
elif git -C "$UPSTREAM_REPO" fetch -q "$UPSTREAM_URL" refs/heads/main \
&& UPSTREAM_COMMIT="$(git -C "$UPSTREAM_REPO" rev-parse --verify -q "$INSTALL_REF^{commit}")"; then
:
else
echo "[sandbox] --from ignored: $HERMES_HOME already has content" >&2
rm -rf -- "$UPSTREAM_REPO"
echo "error: could not resolve upstream ref: $INSTALL_REF" >&2
echo ' Use a branch (main), a tag (v2026.7.7), or a SHA reachable from main.' >&2
exit 1
fi
fi
if [ ! -e "$SANDBOX_ROOT/root/repo/.sandbox-source" ]; then
mkdir -p "$SANDBOX_ROOT/root/repo"
# Persistent roots live under the worktree, so copying with cp would recurse
# into the sandbox itself. tar also lets us exclude a worktree's .git file,
# which can point at the host's shared worktree metadata.
tar -C "$GIT_ROOT" --exclude='./.git' --exclude="./$SANDBOX_DIR_NAME" -cf - . \
| tar -C "$SANDBOX_ROOT/root/repo" -xf -
: > "$SANDBOX_ROOT/root/repo/.sandbox-source"
fi
echo "[sandbox] HERMES_HOME=$HERMES_HOME" >&2
echo "[sandbox] userData=$HERMES_DESKTOP_USER_DATA_DIR" >&2
echo "[sandbox] appName=$HERMES_DESKTOP_APP_NAME" >&2
if [ "$PERSISTENT" = true ]; then
echo "[sandbox] persistent: $SANDBOX_ROOT" >&2
if [ -n "$SEED_DIR" ] && [ ! -e "$SANDBOX_ROOT/.seeded" ]; then
echo "[sandbox] seeding home from $SEED_DIR" >&2
cp -a "$SEED_DIR/." "$SANDBOX_ROOT/home/"
: > "$SANDBOX_ROOT/.seeded"
fi
rm -rf "$SANDBOX_ROOT/root/http"
mkdir -p "$SANDBOX_ROOT/root/http"
if [ -n "$HTTP_ROOT" ]; then
cp -a "$HTTP_ROOT/." "$SANDBOX_ROOT/root/http/"
fi
if [ "$INSTALL_SHORTCUT" = true ]; then
mkdir -p "$SANDBOX_ROOT/root/http/hermes-agent.nousresearch.com"
if [ -n "$INSTALL_REF" ]; then
git -C "$UPSTREAM_REPO" show "$UPSTREAM_COMMIT:scripts/install.sh" \
> "$SANDBOX_ROOT/root/http/hermes-agent.nousresearch.com/install.sh"
else
cp -a "$INSTALLER_PATH" "$SANDBOX_ROOT/root/http/hermes-agent.nousresearch.com/install.sh"
fi
set -- bash -c '
set +e
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash -s -- "$@"
install_status=$?
if [ "$install_status" -eq 0 ] && [ -f /work/promote-main ]; then
next_main=$(cat /work/promote-main)
if git --git-dir=/work/repos/hermes-agent.git update-ref refs/heads/main "$next_main"; then
rm -f /work/promote-main
printf "[sandbox] fake main advanced to this folder for update testing\n" >&2
else
printf "[sandbox] failed to advance fake main after install\n" >&2
install_status=1
fi
fi
if [ "$DEV_SANDBOX_INTERACTIVE" = true ]; then
printf "\n[sandbox] installer exited %s; entering sandbox shell\n" "$install_status" >&2
exec </dev/tty >/dev/tty 2>&1
exec bash -i
fi
exit "$install_status"
' sandbox-installer "$@"
fi
mkdir -p "$SANDBOX_ROOT/root"/{bin,certs,lib64,logs,repos,ssh,usr/bin,usr/local}
REAL_CA_CERT="${DEV_SANDBOX_REAL_CA_CERT:-}"
if [ -z "$REAL_CA_CERT" ]; then
for candidate in /etc/ssl/certs/ca-certificates.crt /etc/ssl/cert.pem; do
if [ -f "$candidate" ]; then
REAL_CA_CERT="$candidate"
break
fi
done
fi
if [ ! -f "$REAL_CA_CERT" ]; then
echo 'error: no system CA bundle found for outbound sandbox HTTPS' >&2
exit 1
fi
if [ ! -f "$SANDBOX_ROOT/root/certs/real-ca.pem" ]; then
cp "$REAL_CA_CERT" "$SANDBOX_ROOT/root/certs/real-ca.pem"
fi
printf 'nameserver 10.0.2.3\n' > "$SANDBOX_ROOT/etc/resolv.conf"
SANDBOX_SHELL="$(command -v bash)"
DYNAMIC_LINKER="${DEV_SANDBOX_DYNAMIC_LINKER:-}"
if [ -z "$DYNAMIC_LINKER" ]; then
# Nix store first: NixOS also ships a /lib64/ld-linux-x86-64.so.2 compat stub,
# so probing FHS paths first would quietly switch which loader a bare script
# invocation uses on this host. Globs that match nothing expand to themselves,
# so every candidate is -f tested. The FHS paths cover Debian/Ubuntu (where
# the loader is under /lib64 or a multiarch /lib dir), which is what CI runs.
for candidate in \
/nix/store/*-glibc-*/lib/ld-linux-*.so.* \
/lib64/ld-linux-x86-64.so.2 \
/lib/ld-linux-aarch64.so.1 \
/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 \
/lib/aarch64-linux-gnu/ld-linux-aarch64.so.1
do
if [ -f "$candidate" ]; then
DYNAMIC_LINKER="$candidate"
break
fi
done
fi
if [ ! -f "$DYNAMIC_LINKER" ]; then
echo 'error: no glibc dynamic linker found for sandboxed release binaries' >&2
echo ' Set DEV_SANDBOX_DYNAMIC_LINKER to its path.' >&2
exit 1
fi
ln -sf "$SANDBOX_SHELL" "$SANDBOX_ROOT/root/bin/sh"
ln -sf "$(command -v ls)" "$SANDBOX_ROOT/root/bin/ls"
ln -sf "$(command -v env)" "$SANDBOX_ROOT/root/usr/bin/env"
ln -sf "$DYNAMIC_LINKER" "$SANDBOX_ROOT/root/lib64/$(basename "$DYNAMIC_LINKER")"
# Identity inside the sandbox. install.sh chooses its layout from `id -u`
# alone (see resolve_install_layout), so the uid here is what decides between
# the root FHS install and a user-level one.
if [ "$RUN_AS_USER" = true ]; then
SANDBOX_UID=1000
SANDBOX_GID=1000
SANDBOX_USER=hermes
SANDBOX_HOME=/home/hermes
else
echo "[sandbox] ephemeral (will be cleaned up on exit)" >&2
SANDBOX_UID=0
SANDBOX_GID=0
SANDBOX_USER=root
SANDBOX_HOME=/root
fi
{
printf 'root:x:0:0:Sandbox Root:/root:%s\n' "$SANDBOX_SHELL"
if [ "$RUN_AS_USER" = true ]; then
printf '%s:x:%s:%s:Sandbox User:%s:%s\n' \
"$SANDBOX_USER" "$SANDBOX_UID" "$SANDBOX_GID" "$SANDBOX_HOME" "$SANDBOX_SHELL"
fi
} > "$SANDBOX_ROOT/etc/passwd"
{
printf 'root:x:0:\n'
if [ "$RUN_AS_USER" = true ]; then
printf '%s:x:%s:\n' "$SANDBOX_USER" "$SANDBOX_GID"
fi
} > "$SANDBOX_ROOT/etc/group"
# A user-level install writes the `hermes` launcher to ~/.local/bin and the
# checkout to $HERMES_HOME; both live under the sandbox HOME, which is bound
# from $SANDBOX_ROOT/home. bwrap maps our real uid to $SANDBOX_UID, so the
# host-side ownership of that directory is what the sandbox sees as its own.
printf 'hosts: files dns\n' > "$SANDBOX_ROOT/etc/nsswitch.conf"
printf '127.0.0.1 localhost\n' > "$SANDBOX_ROOT/etc/hosts"
SOURCE_REPO="$GIT_ROOT"
SOURCE_REF="$COMMIT"
SNAPSHOT_REPO=""
FAKE_REPO="$SANDBOX_ROOT/root/repos/hermes-agent.git"
git -C "$SANDBOX_ROOT/root/repos" init --bare -q hermes-agent.git
if [ -n "$INSTALL_REF" ]; then
git --git-dir="$FAKE_REPO" fetch -q --force "$UPSTREAM_REPO" \
"$UPSTREAM_COMMIT:refs/heads/main"
fi
if [ -n "$(git -C "$GIT_ROOT" status --porcelain)" ]; then
echo '[sandbox] warning: current folder is dirty; creating a temporary fake commit for main' >&2
SNAPSHOT_REPO="$(mktemp -d -t hermes-sandbox-snapshot.XXXXXX)"
git -C "$SNAPSHOT_REPO" init -q
git -C "$SNAPSHOT_REPO" fetch -q "$GIT_ROOT" "$COMMIT"
git -C "$SNAPSHOT_REPO" config user.name 'Hermes sandbox'
git -C "$SNAPSHOT_REPO" config user.email 'sandbox@invalid'
GIT_DIR="$SNAPSHOT_REPO/.git" GIT_WORK_TREE="$GIT_ROOT" git read-tree "$COMMIT"
GIT_DIR="$SNAPSHOT_REPO/.git" GIT_WORK_TREE="$GIT_ROOT" \
git add -A -- .
SNAPSHOT_TREE="$(GIT_DIR="$SNAPSHOT_REPO/.git" git write-tree)"
SNAPSHOT_PARENT="$COMMIT"
if EXISTING_MAIN="$(git --git-dir="$FAKE_REPO" rev-parse --verify refs/heads/main 2>/dev/null)"; then
git -C "$SNAPSHOT_REPO" fetch -q "$FAKE_REPO" "$EXISTING_MAIN"
SNAPSHOT_PARENT="$EXISTING_MAIN"
fi
SOURCE_REF="$(GIT_DIR="$SNAPSHOT_REPO/.git" git commit-tree "$SNAPSHOT_TREE" -p "$SNAPSHOT_PARENT" \
-m 'sandbox snapshot of dirty worktree')"
SOURCE_REPO="$SNAPSHOT_REPO"
fi
if [ "$PERSISTENT" = false ]; then
cleanup() {
chmod -R u+w "$SANDBOX_ROOT"
rm -rf -- "$SANDBOX_ROOT"
if [ -n "$INSTALL_REF" ]; then
git --git-dir="$FAKE_REPO" fetch -q --force "$SOURCE_REPO" \
"$SOURCE_REF:refs/hermes-sandbox/next"
printf '%s\n' "$SOURCE_REF" > "$SANDBOX_ROOT/root/promote-main"
else
git --git-dir="$FAKE_REPO" fetch -q --force "$SOURCE_REPO" \
"$SOURCE_REF:refs/heads/main"
fi
git --git-dir="$FAKE_REPO" symbolic-ref HEAD refs/heads/main
if [ -n "$SNAPSHOT_REPO" ]; then
# Best-effort: it is a mktemp directory the OS will reap, and failing the whole
# run over a leftover object file would be worse than leaking it. Concurrent
# git activity in the worktree can still be writing here as we delete.
rm -rf -- "$SNAPSHOT_REPO" 2>/dev/null || true
fi
if [ -n "$UPSTREAM_REPO" ]; then
rm -rf -- "$UPSTREAM_REPO"
fi
# openssl reads a config even for `req -addext`, and its compiled-in path is a
# symlink into /etc/ssl on Debian/Ubuntu -- which the sandbox replaces. Ship our
# own and point OPENSSL_CONF at it, both here and inside the sandbox.
cp "$SANDBOX_ASSETS/openssl.cnf" "$SANDBOX_ROOT/root/certs/openssl.cnf"
if [ ! -f "$SANDBOX_ROOT/root/certs/ca.pem" ]; then
if ! ca_error="$(OPENSSL_CONF="$SANDBOX_ROOT/root/certs/openssl.cnf" \
openssl req -x509 -newkey rsa:2048 -nodes -days 2 \
-subj '/CN=Hermes dev sandbox CA' \
-extensions sandbox_ca_ext \
-keyout "$SANDBOX_ROOT/root/certs/ca.key" \
-out "$SANDBOX_ROOT/root/certs/ca.pem" 2>&1 >/dev/null)"; then
echo 'error: could not create the sandbox CA:' >&2
printf '%s\n' "$ca_error" >&2
exit 1
fi
fi
GIT_UPLOAD_PACK="$(command -v git-upload-pack)"
sed "s|@GIT_UPLOAD_PACK@|$GIT_UPLOAD_PACK|" "$SANDBOX_ASSETS/ssh-shim.sh" \
> "$SANDBOX_ROOT/root/usr/bin/ssh"
chmod 700 "$SANDBOX_ROOT/root/usr/bin/ssh"
# The fake-internet proxy and the ssh shim are real files under
# scripts/sandbox/ rather than heredocs, so they can be linted, syntax-checked
# and diffed like any other source. Copy them into the sandbox tree.
cp "$SANDBOX_ASSETS/proxy.py" "$SANDBOX_ROOT/root/proxy.py"
if [ -n "$INSTALL_REF" ]; then
echo "[sandbox] fake main: upstream $INSTALL_REF ($UPSTREAM_COMMIT)" >&2
echo "[sandbox] prepared update: current folder ($SOURCE_REF)" >&2
else
echo "[sandbox] fake main: current folder ($SOURCE_REF)" >&2
fi
echo "[sandbox] root: $SANDBOX_ROOT" >&2
echo "[sandbox] http root: $SANDBOX_ROOT/root/http" >&2
if [ "$RUN_AS_USER" = true ]; then
echo "[sandbox] identity: $SANDBOX_USER (uid $SANDBOX_UID) — installs are user-level under $SANDBOX_HOME" >&2
else
echo '[sandbox] identity: root (uid 0) — installs use the /usr/local FHS layout' >&2
fi
[ "$PERSISTENT" = true ] && echo '[sandbox] persistent' >&2 || echo '[sandbox] ephemeral' >&2
for command in awk bash bwrap curl git openssl python3 slirp4netns tar unshare; do
command -v "$command" >/dev/null || {
echo "error: missing required command: $command" >&2
exit 1
}
trap cleanup EXIT
trap 'cleanup; exit 130' INT TERM
done
INTERACTIVE=false
if [ -t 0 ] && [ -t 1 ]; then
INTERACTIVE=true
fi
NODE_DIR="${DEV_SANDBOX_NODE_DIR:-}"
if [ -z "$NODE_DIR" ] && command -v node >/dev/null; then
NODE_DIR="$(dirname "$(dirname "$(command -v node)")")"
fi
WAYLAND_SOCKET=""
if [ -n "${XDG_RUNTIME_DIR:-}" ] && [ -n "${WAYLAND_DISPLAY:-}" ] \
&& [ -S "$XDG_RUNTIME_DIR/$WAYLAND_DISPLAY" ]; then
WAYLAND_SOCKET="$XDG_RUNTIME_DIR/$WAYLAND_DISPLAY"
fi
"$@"
rc=$?
exit $rc
# Namespace plan (stage 1 -> stage 2).
#
# slirp4netns joins the target's userns and setuids to root before configuring
# the netns, so the userns MUST map a uid 0. bwrap's own --unshare-user maps
# exactly one uid, so it cannot both run the payload as uid 1000 and offer slirp
# a root to become: that combination fails with
# setns(CLONE_NEWNET): Operation not permitted.
#
# So stage 1 builds the namespaces here with two ranges:
# inner 0 <- a subuid, unused by the payload, present only so slirp can
# become root inside the namespace
# inner $SANDBOX_UID <- our real host uid, so everything the sandbox writes
# stays owned by us and `rm -rf` on a persistent sandbox needs
# no privileges or chown dance
# The payload then runs in stage 2, where bwrap adds the mount/pid namespaces
# without creating a userns at all.
#
# The root layout needs no subuid at all: inner 0 IS the host uid there.
netns_args=(--user --net)
if [ "$RUN_AS_USER" = true ]; then
host_user="$(id -un)"
subuid_base="$(awk -F: -v u="$host_user" '$1 == u {print $2; exit}' /etc/subuid)"
subgid_base="$(awk -F: -v u="$host_user" '$1 == u {print $2; exit}' /etc/subgid)"
if [ -z "$subuid_base" ] || [ -z "$subgid_base" ]; then
echo "error: no /etc/subuid or /etc/subgid range for $host_user" >&2
echo ' A user-level sandbox needs one spare subordinate id to host' >&2
echo " its internal root. Add e.g. '$host_user:100000:65536' to both," >&2
echo ' or use --root.' >&2
exit 1
fi
netns_args+=(
--map-users="0:$subuid_base:1" --map-users="$SANDBOX_UID:$(id -u):1"
--map-groups="0:$subgid_base:1" --map-groups="$SANDBOX_GID:$(id -g):1"
)
else
netns_args+=(--map-root-user)
fi
sandbox_pid_file="$SANDBOX_ROOT/root/logs/sandbox.pid"
slirp_ready="$SANDBOX_ROOT/root/logs/slirp.ready"
slirp_log="$SANDBOX_ROOT/root/logs/slirp.log"
: > "$sandbox_pid_file"
: > "$slirp_ready"
env \
DEV_SANDBOX_ROOT="$SANDBOX_ROOT" \
DEV_SANDBOX_BASH="$(command -v bash)" \
DEV_SANDBOX_REAL_CA_CERT="$REAL_CA_CERT" \
DEV_SANDBOX_INTERACTIVE="$INTERACTIVE" \
DEV_SANDBOX_USER="$SANDBOX_USER" \
DEV_SANDBOX_HOME="$SANDBOX_HOME" \
DEV_SANDBOX_NODE_DIR="$NODE_DIR" \
DEV_SANDBOX_ELECTRON_LD_LIBRARY_PATH="${DEV_SANDBOX_ELECTRON_LD_LIBRARY_PATH:-}" \
DEV_SANDBOX_XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-}" \
DEV_SANDBOX_WAYLAND_DISPLAY="${WAYLAND_DISPLAY:-}" \
DEV_SANDBOX_WAYLAND_SOCKET="$WAYLAND_SOCKET" \
unshare "${netns_args[@]}" \
"$SANDBOX_ASSETS/stage2-run.sh" "$@" &
sandbox_launcher=$!
for _ in $(seq 1 200); do
[ -s "$sandbox_pid_file" ] && break
if ! kill -0 "$sandbox_launcher" 2>/dev/null; then
wait "$sandbox_launcher"
exit $?
fi
sleep 0.05
done
sandbox_pid="$(tr -dc '0-9' < "$sandbox_pid_file")"
if [ -z "$sandbox_pid" ]; then
echo 'error: sandbox did not report its PID' >&2
exit 1
fi
slirp4netns --configure --disable-host-loopback --ready-fd=3 \
--userns-path="/proc/$sandbox_pid/ns/user" "$sandbox_pid" tap0 \
3>"$slirp_ready" >"$slirp_log" 2>&1 &
slirp_pid=$!
cleanup_slirp() {
kill "$slirp_pid" 2>/dev/null || true
wait "$slirp_pid" 2>/dev/null || true
}
trap cleanup_slirp EXIT INT TERM
for _ in $(seq 1 200); do
[ -s "$slirp_ready" ] && break
if ! kill -0 "$slirp_pid" 2>/dev/null; then
cat "$slirp_log" >&2 || true
exit 1
fi
sleep 0.05
done
if [ ! -s "$slirp_ready" ]; then
echo 'error: timed out waiting for sandbox network setup' >&2
exit 1
fi
wait "$sandbox_launcher"
exit $?
+1
View File
@@ -1577,6 +1577,7 @@ LEGACY_AUTHOR_MAP = {
"17683456+wanazhar@users.noreply.github.com": "wanazhar",
"26782336+cixuuz@users.noreply.github.com": "cixuuz",
"aleksandr.pasevin@openzeppelin.com": "pasevin",
"pasevin@gmail.com": "pasevin",
"ubuntu@localhost.localdomain": "holynn-q",
"holynn@placeholder.local": "holynn-q",
"agent@hermes.local": "jacdevos",
+43
View File
@@ -0,0 +1,43 @@
# Minimal openssl config for the dev sandbox.
#
# The sandbox replaces /etc wholesale, and on Debian/Ubuntu
# /usr/lib/ssl/openssl.cnf (openssl's compiled-in OPENSSLDIR) is a symlink into
# /etc/ssl -- so the config openssl insists on reading disappears and every
# `openssl req` fails with:
#
# Can't open "/usr/lib/ssl/openssl.cnf" for reading
#
# which surfaces to the payload as a bare `curl: (35) Recv failure`. Rather than
# reconstruct each distro's /etc/ssl, point OPENSSL_CONF at this file: the proxy
# only needs enough config for `req -addext` and `x509 -copy_extensions`.
[ req ]
distinguished_name = req_distinguished_name
[ req_distinguished_name ]
# Used by `req -x509` for the sandbox's own CA. Without an explicit
# basicConstraints the generated certificate is not a CA, and every leaf it
# signs is rejected by the client with "invalid CA certificate (79)".
[ sandbox_ca_ext ]
basicConstraints = critical,CA:true
keyUsage = critical,keyCertSign,cRLSign
subjectKeyIdentifier = hash
[ ca ]
default_ca = sandbox_ca
[ sandbox_ca ]
default_md = sha256
policy = policy_anything
email_in_dn = no
preserve = no
[ policy_anything ]
commonName = optional
countryName = optional
stateOrProvinceName = optional
localityName = optional
organizationName = optional
organizationalUnitName = optional
emailAddress = optional
+110
View File
@@ -0,0 +1,110 @@
#!/usr/bin/env bash
# Pick the release tags the install/update E2E should update FROM.
#
# Emits a JSON array of tag names on stdout, suitable for a GitHub Actions
# matrix (`fromJSON`). Choosing at runtime rather than hardcoding keeps the
# matrix honest as releases land: a pinned list silently stops covering the
# newest release the day after it ships, and pins the "oldest" forever even
# after it stops being a version anyone still runs.
#
# Selection: the newest tag, the oldest tag, and evenly spaced tags in between.
# Newest catches "did the last release break updating?", oldest is the longest
# upgrade jump anyone can still make, and the spread samples the migrations in
# between (config-schema bumps, venv layout changes, dependency floors).
#
# Usage:
# scripts/sandbox/pick-release-tags.sh [--count N] [--repo DIR]
#
# --count how many tags to emit (default 5, minimum 1). Fewer tags than
# requested emits all of them.
# --repo repository to read tags from (default: this checkout).
#
# Reads tags from the local checkout, so it needs one fetched with tags
# (actions/checkout with fetch-depth: 0, or `fetch-tags: true`). A shallow
# checkout has no tags and this exits non-zero rather than silently emitting an
# empty matrix.
#
# Only vYYYY.M.D[.N] release tags are considered; the repo also carries
# backup/* and one-off tags that are not releases.
set -euo pipefail
COUNT=5
# Default to the repository containing this script, resolved through its real
# path so a symlinked or copied script still reads the checkout it lives in
# rather than whatever repo the caller happens to be standing in.
REPO=""
while [ "$#" -gt 0 ]; do
case "$1" in
--count)
[ "$#" -ge 2 ] || { echo 'error: --count needs a value' >&2; exit 1; }
COUNT="$2"; shift 2 ;;
--repo)
[ "$#" -ge 2 ] || { echo 'error: --repo needs a value' >&2; exit 1; }
REPO="$2"; shift 2 ;;
-h|--help) sed -n '2,30p' "$0"; exit 0 ;;
*) echo "error: unknown argument: $1" >&2; exit 1 ;;
esac
done
case "$COUNT" in
''|*[!0-9]*) echo "error: --count must be a positive integer: $COUNT" >&2; exit 1 ;;
esac
[ "$COUNT" -ge 1 ] || { echo 'error: --count must be at least 1' >&2; exit 1; }
# Resolve the script's own location through symlinks, then ask git which
# worktree that path belongs to. Deriving the repo from the script rather than
# from $PWD means a copied script cannot silently report a different checkout's
# tags, and --show-toplevel keeps it correct when invoked from a subdirectory.
if [ -z "$REPO" ]; then
script_path="${BASH_SOURCE[0]}"
if command -v readlink >/dev/null 2>&1; then
script_path="$(readlink -f "$script_path" 2>/dev/null || printf '%s' "$script_path")"
fi
script_dir="$(cd "$(dirname "$script_path")" && pwd)"
REPO="$(git -C "$script_dir" rev-parse --show-toplevel 2>/dev/null || printf '%s' "$script_dir")"
fi
# sort -V orders v2026.4.8 before v2026.4.13 (numeric), which a plain
# lexicographic sort gets wrong.
mapfile -t tags < <(
git -C "$REPO" tag --list 'v*' \
| grep -E '^v[0-9]{4}\.[0-9]+\.[0-9]+(\.[0-9]+)?$' \
| sort -V
)
total="${#tags[@]}"
if [ "$total" -eq 0 ]; then
echo "error: no release tags found in $REPO" >&2
echo ' A shallow clone has no tags: fetch with tags (actions/checkout' >&2
echo ' with fetch-depth: 0, or fetch-tags: true).' >&2
exit 1
fi
if [ "$total" -le "$COUNT" ]; then
picked=("${tags[@]}")
elif [ "$COUNT" -eq 1 ]; then
# One slot means the newest release; there is no span to spread across.
picked=("${tags[$((total - 1))]}")
else
# Evenly spaced indices across [0, total-1], endpoints included, so the
# oldest and newest are always present and the rest are spread between them.
picked=()
for slot in $(seq 0 $((COUNT - 1))); do
# Round to nearest rather than truncate, so the spacing does not bunch
# toward the oldest end.
index=$(( (slot * (total - 1) * 2 + (COUNT - 1)) / ((COUNT - 1) * 2) ))
candidate="${tags[$index]}"
# Guard against a duplicate if rounding lands twice on the same tag.
case " ${picked[*]-} " in
*" $candidate "*) continue ;;
esac
picked+=("$candidate")
done
fi
printf '['
for i in "${!picked[@]}"; do
[ "$i" -eq 0 ] || printf ','
printf '"%s"' "${picked[$i]}"
done
printf ']\n'
+237
View File
@@ -0,0 +1,237 @@
"""MITM proxy backing the dev sandbox's fake Internet.
Listens on 127.0.0.1:8080 and is pointed at by http_proxy/https_proxy inside
the sandbox. For each request it either serves a fixture from the filesystem or
forwards to the real host:
* ``<root>/<host>/<path>`` exists -> serve it. This is how the sandbox answers
the canonical install URL with the installer under test, so the payload can
run the true ``curl -fsSL https://…/install.sh | bash`` one-liner.
* otherwise -> forward upstream, verifying against the real CA bundle. The
sandbox is isolated from the *host*, not from the internet: a real install
still has to reach PyPI and npm.
HTTPS is intercepted by minting a per-host certificate from the sandbox's own
throwaway CA, which the payload trusts via CURL_CA_BUNDLE / SSL_CERT_FILE.
Usage: proxy.py <fixture-root> <certs-dir> <real-ca-bundle>
"""
import os
import pathlib
import socket
import ssl
import subprocess
import sys
import threading
from urllib.parse import unquote, urlsplit
ROOT, CERTS, REAL_CA = map(pathlib.Path, sys.argv[1:])
LISTEN_ADDRESS = ('127.0.0.1', 8080)
MAX_REQUEST_BYTES = 65536
UPSTREAM_TIMEOUT_SECONDS = 30
CERT_VALIDITY_DAYS = 2
def read_request(conn):
data = b""
while b"\r\n\r\n" not in data and len(data) < MAX_REQUEST_BYTES:
part = conn.recv(4096)
if not part:
return b""
data += part
return data
def run_openssl(args):
"""Run openssl, raising with its stderr when it fails.
Discarding stderr here costs real debugging time: the caller sees only a
dropped connection (``curl: (35) Recv failure``) and the log holds nothing
but the argv, so an unwritable directory, a missing CA key, and an option
the host's openssl rejects all look identical.
"""
done = subprocess.run(
['openssl', *args], stdout=subprocess.DEVNULL, stderr=subprocess.PIPE
)
if done.returncode != 0:
detail = done.stderr.decode('utf-8', 'replace').strip()
raise RuntimeError(
f'openssl {args[0]} failed (exit {done.returncode}): {detail}'
)
_CERT_LOCK = threading.Lock()
def cert_for(host):
"""Return a (cert, key) pair for host, minting it from the sandbox CA.
Minting is serialized and published atomically. The proxy is threaded, so
two concurrent requests for the same host would otherwise both run openssl
into the same paths, and a reader could pick up a finished certificate
beside a key from the other writer -- which TLS rejects as
``[X509: KEY_VALUES_MISMATCH] key values mismatch``.
"""
safe = ''.join(char if char.isalnum() or char in '.-' else '_' for char in host)
cert, key = CERTS / f'{safe}.pem', CERTS / f'{safe}.key'
if cert.exists() and key.exists():
return cert, key
with _CERT_LOCK:
# Re-check: another thread may have finished while we waited.
if cert.exists() and key.exists():
return cert, key
# Build under unique temp names, then rename into place. os.replace is
# atomic, so a reader sees either the old pair or the new one, never a
# half-written mix. The key lands first: the certificate's existence is
# what everything else keys off.
stamp = f'{os.getpid()}.{threading.get_ident()}'
tmp_key = CERTS / f'{safe}.key.{stamp}'
tmp_cert = CERTS / f'{safe}.pem.{stamp}'
csr = CERTS / f'{safe}.csr.{stamp}'
run_openssl([
'req', '-newkey', 'rsa:2048', '-nodes',
'-subj', f'/CN={host}',
'-addext', f'subjectAltName=DNS:{host}',
'-keyout', str(tmp_key), '-out', str(csr),
])
run_openssl([
'x509', '-req', '-days', str(CERT_VALIDITY_DAYS), '-in', str(csr),
'-CA', str(CERTS / 'ca.pem'), '-CAkey', str(CERTS / 'ca.key'),
'-CAcreateserial', '-copy_extensions', 'copy', '-out', str(tmp_cert),
])
csr.unlink(missing_ok=True)
os.replace(tmp_key, key)
os.replace(tmp_cert, cert)
return cert, key
def file_for(host, target):
"""Resolve a request to a fixture file, or None to forward upstream."""
path = urlsplit(target).path or '/'
parts = pathlib.PurePosixPath(unquote(path)).parts
if '..' in parts:
return None
candidate = ROOT / host / pathlib.PurePosixPath(*[p for p in parts if p != '/'])
if candidate.is_dir():
candidate /= 'index.html'
return candidate if candidate.is_file() else None
def respond_fixture(conn, found):
body = found.read_bytes()
headers = (
f'Content-Length: {len(body)}\r\nConnection: close\r\n\r\n'.encode()
)
conn.sendall(b'HTTP/1.1 200 OK\r\n' + headers + body)
def close_request(request, target=None):
"""Rewrite a proxied request for a direct upstream connection."""
headers, separator, body = request.partition(b'\r\n\r\n')
lines = headers.split(b'\r\n')
if target is not None:
method, _, version = lines[0].split(b' ', 2)
lines[0] = b' '.join((method, target.encode(), version))
lines = [
line for line in lines
if not line.lower().startswith(b'proxy-connection:')
]
lines.append(b'Connection: close')
return b'\r\n'.join(lines) + separator + body
def relay(source, destination):
while True:
chunk = source.recv(MAX_REQUEST_BYTES)
if not chunk:
return
destination.sendall(chunk)
def forward_https(conn, host, port, request):
context = ssl.create_default_context(cafile=str(REAL_CA))
with socket.create_connection((host, port), timeout=UPSTREAM_TIMEOUT_SECONDS) as raw:
with context.wrap_socket(raw, server_hostname=host) as upstream:
upstream.sendall(close_request(request))
relay(upstream, conn)
def forward_http(conn, host, port, request, target):
parsed = urlsplit(target)
path = parsed.path or '/'
if parsed.query:
path += f'?{parsed.query}'
with socket.create_connection((host, port), timeout=UPSTREAM_TIMEOUT_SECONDS) as upstream:
upstream.sendall(close_request(request, path))
relay(upstream, conn)
def handle_connect(conn, target):
"""Intercept a CONNECT tunnel, terminating TLS with a minted cert."""
host, _, port_text = target.rpartition(':')
port = int(port_text or '443')
conn.sendall(b'HTTP/1.1 200 Connection Established\r\n\r\n')
cert, key = cert_for(host)
context = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
context.load_cert_chain(cert, key)
with context.wrap_socket(conn, server_side=True) as tls:
nested = read_request(tls)
if not nested:
return
line = nested.split(b'\r\n', 1)[0].decode('iso-8859-1')
nested_target = line.split(' ', 2)[1]
found = file_for(host, nested_target)
if found is not None:
respond_fixture(tls, found)
else:
forward_https(tls, host, port, nested)
def host_from_headers(request):
for header in request.split(b'\r\n')[1:]:
if header.lower().startswith(b'host:'):
value = header.split(b':', 1)[1].strip().decode()
return value.split(':', 1)[0]
return None
def handle_request(conn):
with conn:
request = read_request(conn)
if not request:
return
line = request.split(b'\r\n', 1)[0].decode('iso-8859-1')
method, target, _ = line.split(' ', 2)
if method.upper() == 'CONNECT':
handle_connect(conn, target)
return
parsed = urlsplit(target)
host = parsed.hostname or host_from_headers(request) or 'unknown'
found = file_for(host, target)
if found is not None:
respond_fixture(conn, found)
else:
forward_http(conn, host, parsed.port or 80, request, target)
def handle(conn):
try:
handle_request(conn)
except Exception as error:
print(f'proxy request failed: {error!r}', file=sys.stderr, flush=True)
def main():
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as server:
server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
server.bind(LISTEN_ADDRESS)
server.listen()
while True:
conn, _ = server.accept()
threading.Thread(target=handle, args=(conn,), daemon=True).start()
if __name__ == '__main__':
main()
+13
View File
@@ -0,0 +1,13 @@
#!/usr/bin/env bash
# Stand-in for ssh inside the dev sandbox.
#
# install.sh and `hermes update` clone over ssh first (git@github.com:...), so
# the sandbox needs an `ssh` that answers. Rather than run a real sshd, this
# ignores the host, user, and command git asked for and speaks the
# upload-pack protocol directly against the sandbox's bare repo -- which is
# what makes the ssh-first code path exercisable with no keys, no known_hosts,
# and no network.
#
# GIT_UPLOAD_PACK is substituted by dev-sandbox.sh when it installs this shim,
# because the host's git-upload-pack is not necessarily on the sandbox PATH.
exec @GIT_UPLOAD_PACK@ /work/repos/hermes-agent.git
+251
View File
@@ -0,0 +1,251 @@
#!/usr/bin/env bash
# Stage 2 of the dev sandbox: build the mounts and run the payload.
#
# Not called directly. scripts/dev-sandbox.sh (stage 1) creates the user and
# network namespaces with `unshare` and re-execs into this script inside them,
# so by the time this runs we are already at the target uid with a private
# netns. bwrap therefore does NOT create a userns here -- it only adds the
# mount and pid namespaces. (`unshare --user` grants its creator full
# capabilities in the new userns regardless of which uid it maps, which is what
# lets bwrap mount as a non-root uid.)
#
# The whole interface with stage 1 is the DEV_SANDBOX_* environment, asserted
# below: there are no shared functions or variables between the two stages.
# Stage 1 locates this script alongside the other sandbox assets (see
# DEV_SANDBOX_ASSETS in dev-sandbox.sh), so the Nix wrapper's store copy and a
# plain repo checkout both work.
set -euo pipefail
: "${DEV_SANDBOX_ROOT:?missing DEV_SANDBOX_ROOT}"
: "${DEV_SANDBOX_BASH:?missing DEV_SANDBOX_BASH}"
: "${DEV_SANDBOX_INTERACTIVE:?missing DEV_SANDBOX_INTERACTIVE}"
: "${DEV_SANDBOX_USER:?missing DEV_SANDBOX_USER}"
: "${DEV_SANDBOX_HOME:?missing DEV_SANDBOX_HOME}"
# Announce our pid so stage 1 can point slirp4netns at these namespaces,
# then hold until it reports the network is up.
slirp_ready="$DEV_SANDBOX_ROOT/root/logs/slirp.ready"
printf '%s\n' "$$" > "$DEV_SANDBOX_ROOT/root/logs/sandbox.pid"
for _ in $(seq 1 200); do
[ -s "$slirp_ready" ] && break
sleep 0.05
done
if [ ! -s "$slirp_ready" ]; then
echo 'error: timed out waiting for sandbox network setup' >&2
cat "$DEV_SANDBOX_ROOT/root/logs/slirp.log" >&2 || true
exit 1
fi
# The sandbox HOME is /root for a root install and /home/<user> for a
# user-level one. Only the latter needs its parent created first; --dir /
# is not a thing bwrap accepts.
home_mounts=()
home_parent="$(dirname "$DEV_SANDBOX_HOME")"
if [ "$home_parent" != / ]; then
home_mounts+=(--dir "$home_parent")
fi
home_mounts+=(--bind "$DEV_SANDBOX_ROOT/home" "$DEV_SANDBOX_HOME")
node_env=()
if [ -n "${DEV_SANDBOX_NODE_DIR:-}" ]; then
node_env+=(--setenv npm_config_nodedir "$DEV_SANDBOX_NODE_DIR")
fi
electron_env=()
if [ -n "${DEV_SANDBOX_ELECTRON_LD_LIBRARY_PATH:-}" ]; then
electron_env+=(
--setenv LD_LIBRARY_PATH "$DEV_SANDBOX_ELECTRON_LD_LIBRARY_PATH"
--setenv HERMES_DESKTOP_DISABLE_GPU 1
)
fi
gui_mounts=()
if [ -n "${DEV_SANDBOX_WAYLAND_SOCKET:-}" ]; then
runtime_dir="${DEV_SANDBOX_XDG_RUNTIME_DIR:?missing DEV_SANDBOX_XDG_RUNTIME_DIR}"
runtime_parent="$(dirname "$runtime_dir")"
runtime_grandparent="$(dirname "$runtime_parent")"
gui_mounts+=(
--dir "$runtime_grandparent"
--dir "$runtime_parent"
--dir "$runtime_dir"
--bind "$DEV_SANDBOX_WAYLAND_SOCKET" "$DEV_SANDBOX_WAYLAND_SOCKET"
--setenv XDG_RUNTIME_DIR "$runtime_dir"
--setenv WAYLAND_DISPLAY "${DEV_SANDBOX_WAYLAND_DISPLAY:?missing DEV_SANDBOX_WAYLAND_DISPLAY}"
)
fi
# How the sandbox gets a usable runtime, and where its own shims go.
#
# On Nix, every binary lives under /nix/store, so the sandbox can own /bin,
# /lib64 and /usr/bin outright and fill them with symlinks into the store.
#
# Elsewhere the runtime IS /usr, /bin, /lib, /lib64 -- so binding the
# sandbox's near-empty versions over them hides the real thing, and bwrap
# dies with `execvp /usr/bin/bash: No such file or directory`. Keep the host
# directories read-only and override only the individual files we shim.
#
# The same answer decides how /etc is handled further down.
if [ -d /nix ] && [[ "$(readlink -f "$DEV_SANDBOX_BASH")" == /nix/* ]]; then
USE_HOST_RUNTIME=false
else
USE_HOST_RUNTIME=true
fi
runtime_mounts=()
shim_mounts=()
if [ "$USE_HOST_RUNTIME" = false ]; then
runtime_mounts+=(--ro-bind /nix /nix)
shim_mounts+=(
--dir /usr
--dir /bin
--dir /lib64
--bind "$DEV_SANDBOX_ROOT/root/bin" /bin
--bind "$DEV_SANDBOX_ROOT/root/lib64" /lib64
--bind "$DEV_SANDBOX_ROOT/root/usr/bin" /usr/bin
)
else
for path in /usr /bin /sbin /lib /lib64; do
[ -e "$path" ] && runtime_mounts+=(--ro-bind "$path" "$path")
done
# The git-upload-pack shim standing in for github.com is the only file that
# must beat the host's copy; sh/ls/env are already there for real.
shim_mounts+=(--bind "$DEV_SANDBOX_ROOT/root/usr/bin/ssh" /usr/bin/ssh)
fi
# /etc: start from a copy of the host's and overwrite only the files we fake.
#
# Replacing the whole directory with a five-file one is the tempting shortcut
# and it is wrong: a distro puts things under /etc that binaries outside /etc
# depend on, so hiding all of it breaks tools that look fine on PATH. Two real
# examples, both Debian/Ubuntu: openssl's compiled-in openssl.cnf is a symlink
# into /etc/ssl, and /usr/bin/awk is a symlink to /etc/alternatives/awk -- with
# /etc replaced, openssl cannot mint a certificate and awk reports "not found".
# Those are two symptoms of one cause, and nothing says there are only two.
#
# Copying rather than mount-overlaying the individual files, because several of
# these are symlinks in the wild (resolv.conf -> ../run/systemd/... on Ubuntu,
# hosts and nsswitch.conf -> /etc/static/... on NixOS) and bwrap cannot bind a
# file onto a symlink whose target does not exist inside the sandbox.
#
# Symlinks are copied as symlinks, never dereferenced: on NixOS /etc/static
# points into the store and following it would copy gigabytes per sandbox. The
# store is already mounted at /nix on that path, and the host runtime dirs are
# mounted at their own paths, so absolute symlinks still resolve.
#
# The five we override, and why each must differ from the host's:
# passwd, group the sandbox identity, which does not exist on the host
# resolv.conf slirp4netns's DNS, not the host resolver
# nsswitch.conf files+dns only, so nothing consults host NSS modules
# hosts minimal, so no host entry leaks in
#
# os-release is removed rather than replaced. Installers branch on it to reach
# for a package manager -- `install.sh` reads ID from it and, on debian/ubuntu,
# offers to apt-get build tools, prompting on /dev/tty when sudo exists but is
# not passwordless. That prompt cannot be satisfied here (no terminal) and it is
# fatal under `set -e`. Inheriting the host's file would make the sandbox claim
# to be a distro whose package manager it cannot actually use; absent means
# DISTRO="unknown" and the apt path is skipped, which is the truth.
etc_mounts=()
if [ "$USE_HOST_RUNTIME" = true ] && [ -d /etc ]; then
sandbox_etc="$DEV_SANDBOX_ROOT/etc-merged"
rm -rf -- "$sandbox_etc"
mkdir -p "$sandbox_etc"
# -a keeps symlinks as symlinks; unreadable entries (shadow, sudoers) are
# skipped rather than failing the run.
cp -a /etc/. "$sandbox_etc/" 2>/dev/null || true
for etc_file in passwd group resolv.conf nsswitch.conf hosts; do
[ -f "$DEV_SANDBOX_ROOT/etc/$etc_file" ] || continue
rm -f "$sandbox_etc/$etc_file"
cp "$DEV_SANDBOX_ROOT/etc/$etc_file" "$sandbox_etc/$etc_file"
done
rm -f "$sandbox_etc/os-release" "$sandbox_etc/lsb-release"
etc_mounts+=(--ro-bind "$sandbox_etc" /etc)
else
etc_mounts+=(--bind "$DEV_SANDBOX_ROOT/etc" /etc)
fi
# /dev without a tty, so a script guarding on `[ -e /dev/tty ]` takes its
# no-terminal path.
#
# bwrap's --dev creates a /dev/tty NODE, but nothing in here has a controlling
# terminal, so opening it fails with "No such device or address". That is the
# worst of both: the guard passes and the read then fails. Under `set -e` --
# which install.sh uses -- a failed read inside a function aborts the whole
# installer, which is exactly how older releases died here while prompting for
# sudo to install ripgrep/ffmpeg.
#
# Making the tty real is not the fix: with an openable terminal that prompt
# blocks forever waiting for input nobody will type. Absent is what a headless
# machine looks like, and what every prompt in here should assume.
#
# --dev cannot be used with the node removed afterwards (bwrap refuses to mount
# a directory over a device node), so /dev is assembled explicitly.
dev_mounts=(
--tmpfs /dev
--dev-bind /dev/null /dev/null
--dev-bind /dev/zero /dev/zero
--dev-bind /dev/full /dev/full
--dev-bind /dev/random /dev/random
--dev-bind /dev/urandom /dev/urandom
--symlink /proc/self/fd /dev/fd
--symlink /proc/self/fd/0 /dev/stdin
--symlink /proc/self/fd/1 /dev/stdout
--symlink /proc/self/fd/2 /dev/stderr
)
if [ "$DEV_SANDBOX_INTERACTIVE" = true ]; then
# An interactive shell is deliberately given a terminal; keep bwrap's /dev.
dev_mounts=(--dev /dev)
fi
exec bwrap \
--unshare-pid \
--die-with-parent --proc /proc --tmpfs /tmp \
"${dev_mounts[@]}" \
"${gui_mounts[@]}" \
"${runtime_mounts[@]}" \
--bind "$DEV_SANDBOX_ROOT/root" /work \
"${shim_mounts[@]}" \
--bind "$DEV_SANDBOX_ROOT/root/usr/local" /usr/local \
"${home_mounts[@]}" \
"${etc_mounts[@]}" \
--chdir /work/repo \
--clearenv \
--setenv PATH "$DEV_SANDBOX_HOME/.local/bin:/usr/local/bin:/usr/bin:$PATH" \
--setenv HOME "$DEV_SANDBOX_HOME" \
--setenv USER "$DEV_SANDBOX_USER" \
--setenv LOGNAME "$DEV_SANDBOX_USER" \
--setenv CURL_CA_BUNDLE /work/certs/ca.pem \
--setenv SSL_CERT_FILE /work/certs/ca.pem \
--setenv GIT_SSL_CAINFO /work/certs/ca.pem \
--setenv NODE_EXTRA_CA_CERTS /work/certs/real-ca.pem \
--setenv OPENSSL_CONF /work/certs/openssl.cnf \
--setenv HTTP_PROXY http://127.0.0.1:8080 \
--setenv HTTPS_PROXY http://127.0.0.1:8080 \
--setenv ALL_PROXY http://127.0.0.1:8080 \
--setenv NO_PROXY '' \
--setenv DEV_SANDBOX_INTERACTIVE "$DEV_SANDBOX_INTERACTIVE" \
--setenv ELECTRON_DISABLE_SANDBOX 1 \
"${node_env[@]}" \
"${electron_env[@]}" \
-- "$DEV_SANDBOX_BASH" -ceu '
python3 /work/proxy.py /work/http /work/certs /work/certs/real-ca.pem >/work/logs/proxy.log 2>&1 &
proxy_pid=$!
cleanup() {
kill "$proxy_pid" 2>/dev/null || true
wait "$proxy_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
# Bash opens /dev/tcp itself, so the readiness probe needs no netcat --
# one less binary the sandbox has to find on the host (GitHub runners
# ship no `nc`).
proxy_up() { (exec 3<>/dev/tcp/127.0.0.1/8080) 2>/dev/null; }
for _ in $(seq 1 100); do
proxy_up && break
sleep 0.05
done
if ! proxy_up; then
echo "error: the sandbox fake-internet proxy never came up" >&2
cat /work/logs/proxy.log >&2 || true
exit 1
fi
"$@"
' sandbox-command "$@"
@@ -71,6 +71,14 @@ The ONLY import surface is `@hermes/plugin-sdk` (plus `react` /
(renders below Artifacts, lights up at the route) — and/or a
`PALETTE_AREA` command calling `host.navigate('/my-page')`.
- `ctx.storage.get/set/remove` — persistence namespaced to your plugin.
- `ctx.os` — the curated OS door, attributed to your plugin:
`ctx.os.notify({ title, body?, silent? })` posts a native OS notification.
Fires only while the user is away from Hermes (use `host.notify` for the
in-app toast); gated by Settings ▸ Notifications ▸ "Plugin notifications"
and throttled per plugin — reserve it for genuinely notable events.
`ctx.os.openExternal(url)`, `ctx.os.revealPath(path)`, and
`ctx.os.writeClipboard(text)` resolve `false` (never throw) when the
capability isn't available.
- `ctx.i18n.register({ en, ja, ... })` — ship your OWN locale bundles, scoped
to your plugin (never edit core `en.ts`). Values are literal strings or
interpolator functions; nested trees are addressed by dot-path. Read them
@@ -0,0 +1,106 @@
"""Tests for credential pool upsert — key rotation clears exhaustion state."""
from __future__ import annotations
import json
def _write_auth_store(tmp_path, payload: dict) -> None:
hermes_home = tmp_path / "hermes"
hermes_home.mkdir(parents=True, exist_ok=True)
(hermes_home / "auth.json").write_text(json.dumps(payload, indent=2))
def test_key_rotation_clears_exhausted_status(tmp_path, monkeypatch):
"""Replacing an exhausted API key via _upsert_entry resets last_status.
Regression: `hermes setup` saves a new OPENROUTER_API_KEY to .env, which
triggers _seed_from_env → _upsert_entry. If the existing pool entry was
marked exhausted (e.g. from a rate-limit on the old key), the stale status
was preserved on the new key — making the pool appear unusable even though
a fresh valid key was present.
"""
monkeypatch.setenv("HERMES_HOME", str(tmp_path / "hermes"))
_write_auth_store(
tmp_path,
{
"version": 1,
"credential_pool": {
"openrouter": [
{
"id": "cred-1",
"label": "OPENROUTER_API_KEY",
"auth_type": "api_key",
"priority": 0,
"source": "env:OPENROUTER_API_KEY",
"access_token": "old-key",
"last_status": "exhausted",
"last_status_at": 1000.0,
"last_error_code": 429,
"last_error_reason": "rate_limit",
"last_error_message": "Too many requests",
"last_error_reset_at": 2000.0,
}
]
},
},
)
# Simulate the user rotating their key (new value in env)
monkeypatch.setenv("OPENROUTER_API_KEY", "new-rotated-key")
from agent.credential_pool import load_pool
pool = load_pool("openrouter")
entry = pool.select()
assert entry is not None, "Pool should have a usable entry after key rotation"
assert entry.access_token == "new-rotated-key"
assert entry.last_status is None, "last_status should be cleared after key rotation"
assert entry.last_status_at is None
assert entry.last_error_code is None
assert entry.last_error_reason is None
assert entry.last_error_message is None
assert entry.last_error_reset_at is None
def test_same_key_preserves_exhausted_status(tmp_path, monkeypatch):
"""If the key has NOT changed, _upsert_entry does not clear exhaustion state."""
monkeypatch.setenv("HERMES_HOME", str(tmp_path / "hermes"))
from agent.credential_pool import PooledCredential, _upsert_entry
existing = PooledCredential.from_dict(
"openrouter",
{
"id": "cred-1",
"label": "OPENROUTER_API_KEY",
"auth_type": "api_key",
"priority": 0,
"source": "env:OPENROUTER_API_KEY",
"access_token": "same-key",
"last_status": "exhausted",
"last_status_at": 1000.0,
"last_error_code": 429,
"last_error_reason": "rate_limit",
"last_error_message": "Too many requests",
"last_error_reset_at": 2000.0,
},
)
entries = [existing]
# Upsert with the same token — should NOT clear exhaustion
_upsert_entry(
entries,
"openrouter",
"env:OPENROUTER_API_KEY",
{
"source": "env:OPENROUTER_API_KEY",
"auth_type": "api_key",
"access_token": "same-key",
},
)
assert entries[0].last_status == "exhausted", (
"last_status should not be cleared when the key is unchanged"
)
+54 -3
View File
@@ -210,7 +210,7 @@ class TestPoolRotationCycle:
)
assert recovered is True
assert has_retried is False # reset after rotation
pool.mark_exhausted_and_rotate.assert_called_once_with(status_code=429, error_context=None, api_key_hint="test-api-key")
pool.mark_exhausted_and_rotate.assert_called_once_with(status_code=429, error_context=None, api_key_hint="test-api-key", failure_reason="rate_limit")
agent._swap_credential.assert_called_once_with(entries[1])
def test_pool_exhaustion_returns_false(self):
@@ -236,7 +236,7 @@ class TestPoolRotationCycle:
)
assert recovered is True
assert has_retried is False
pool.mark_exhausted_and_rotate.assert_called_once_with(status_code=402, error_context=None, api_key_hint="test-api-key")
pool.mark_exhausted_and_rotate.assert_called_once_with(status_code=402, error_context=None, api_key_hint="test-api-key", failure_reason="billing")
def test_api_key_hint_from_pool_current_when_agent_key_missing(self):
@@ -273,7 +273,8 @@ class TestPoolRotationCycle:
)
assert recovered is True
pool.mark_exhausted_and_rotate.assert_called_once_with(
status_code=402, error_context=None, api_key_hint="pool-current-key"
status_code=402, error_context=None, api_key_hint="pool-current-key",
failure_reason="billing",
)
@@ -492,3 +493,53 @@ class TestFailureAttribution:
assert self._statuses(pool)["cred-0"] != "exhausted"
agent._swap_credential.assert_not_called()
def test_classified_billing_403_recorded_on_entry(self, tmp_path, monkeypatch):
"""A billing-classified 403 must reach the pool as `billing`, not a bare 403.
`error_classifier` maps OpenRouter's `key limit exceeded` 403 (and xAI
spending-limit blocks) to FailoverReason.billing, but the pool only
ever saw the raw status — so a sole-credential pool gave a spent
account the 60s transient cooldown and re-failed every minute. The
recovery path now forwards the classified reason so the pool can size
the bench correctly.
"""
from agent.error_classifier import FailoverReason
pool = self._make_pool(
tmp_path, monkeypatch,
[self._entry(0, "key-a"), self._entry(1, "key-b")],
)
agent = self._agent(pool, failing_key="key-b")
agent._is_entitlement_failure = MagicMock(return_value=False)
from agent.agent_runtime_helpers import recover_with_credential_pool
recover_with_credential_pool(
agent,
status_code=403,
has_retried_429=False,
classified_reason=FailoverReason.billing,
)
failed = {e.id: e for e in pool.entries()}["cred-1"]
assert failed.last_status == "exhausted"
assert failed.failure_reason == "billing"
def test_unclassified_403_records_no_billing_reason(self, tmp_path, monkeypatch):
"""An unclassified 403 stays transient — no billing verdict is invented."""
pool = self._make_pool(
tmp_path, monkeypatch,
[self._entry(0, "key-a"), self._entry(1, "key-b")],
)
agent = self._agent(pool, failing_key="key-b")
agent._is_entitlement_failure = MagicMock(return_value=False)
from agent.agent_runtime_helpers import recover_with_credential_pool
recover_with_credential_pool(
agent, status_code=403, has_retried_429=False
)
failed = {e.id: e for e in pool.entries()}["cred-1"]
assert failed.failure_reason != "billing"
@@ -0,0 +1,163 @@
"""Sole-credential cooldown: a pool with nothing to rotate to should not bench
its only key for an hour on a transient throttle (429/403/5xx).
Regression for the case where removing fallbacks / running a single API key
turned a transient rate-limit into an hour of hard failures.
"""
from __future__ import annotations
import json
import time
import pytest
def _write_auth_store(tmp_path, payload: dict) -> None:
hermes_home = tmp_path / "hermes"
hermes_home.mkdir(parents=True, exist_ok=True)
(hermes_home / "auth.json").write_text(json.dumps(payload, indent=2))
def _entry(
error_code: int,
*,
age_seconds: float,
cred_id: str = "cred-1",
priority: int = 0,
failure_reason: str | None = None,
) -> dict:
entry = {
"id": cred_id,
"label": cred_id,
"auth_type": "api_key",
"priority": priority,
"source": "manual",
"access_token": "***",
"base_url": "https://openrouter.ai/api/v1",
"last_status": "exhausted",
"last_status_at": time.time() - age_seconds,
"last_error_code": error_code,
}
if failure_reason is not None:
entry["failure_reason"] = failure_reason
return entry
def _load(tmp_path, monkeypatch, entries: list[dict]):
monkeypatch.setenv("HERMES_HOME", str(tmp_path / "hermes"))
monkeypatch.delenv("OPENROUTER_API_KEY", raising=False)
_write_auth_store(
tmp_path,
{"version": 1, "credential_pool": {"openrouter": entries}},
)
from agent.credential_pool import load_pool
return load_pool("openrouter")
def test_sole_credential_429_recovers_after_short_cooldown(tmp_path, monkeypatch):
"""A single 429-throttled key recovers within ~1 min, not 1 hour.
Exhausted 90s ago: under the old 1-hour TTL this stays benched (and a
single-key pool would have nothing to return); the sole-credential short
cooldown lets it recover.
"""
pool = _load(tmp_path, monkeypatch, [_entry(429, age_seconds=90)])
entry = pool.select()
assert entry is not None
assert entry.id == "cred-1"
assert entry.last_status == "ok"
def test_sole_credential_403_recovers_after_short_cooldown(tmp_path, monkeypatch):
"""403 (edge-throttle variant, hits the catch-all default TTL) also recovers."""
pool = _load(tmp_path, monkeypatch, [_entry(403, age_seconds=90)])
entry = pool.select()
assert entry is not None
assert entry.last_status == "ok"
def test_sole_credential_billing_403_keeps_full_bench(tmp_path, monkeypatch):
"""A 403 classified as BILLING must keep the full bench, not the 60s cooldown.
Providers overload 403: OpenRouter returns it for `key limit exceeded` and
xAI for a spending-limit block, both of which `error_classifier` maps to
FailoverReason.billing. Status alone can't tell those from an edge
throttle, so retrying a spent account every 60s just re-fails forever.
The classified reason rides along on the entry and wins over the status.
"""
pool = _load(
tmp_path,
monkeypatch,
[_entry(403, age_seconds=90, failure_reason="billing")],
)
assert pool.has_available() is False
assert pool.select() is None
def test_sole_credential_billing_403_survives_reload(tmp_path, monkeypatch):
"""The classified reason persists, so a restart can't downgrade the bench.
`failure_reason` is written to auth.json with the entry; without that, a
process restart would re-read a bare 403 and hand the spent key back after
60 seconds.
"""
from agent.credential_pool import _exhausted_ttl
pool = _load(
tmp_path,
monkeypatch,
[_entry(403, age_seconds=90, failure_reason="billing")],
)
entry = pool.entries()[0]
assert entry.failure_reason == "billing"
assert _exhausted_ttl(403, sole_credential=True, failure_reason="billing") == 60 * 60
assert _exhausted_ttl(403, sole_credential=True) == 60
def test_sole_credential_402_keeps_full_bench(tmp_path, monkeypatch):
"""402 (billing/quota) is genuine exhaustion — a quick retry can't help, so
the sole-credential short cooldown must NOT apply."""
pool = _load(tmp_path, monkeypatch, [_entry(402, age_seconds=90)])
assert pool.has_available() is False
assert pool.select() is None
def test_sole_credential_next_available_at_uses_short_cooldown(tmp_path, monkeypatch):
"""next_available_at must also honour the sole-credential short cooldown.
Without this, the fallback restore gate in agent_runtime_helpers waits an
hour for a 60s cooldown, keeping the agent on a fallback provider far
longer than necessary.
"""
from agent.credential_pool import EXHAUSTED_TTL_SOLE_CREDENTIAL_SECONDS
pool = _load(tmp_path, monkeypatch, [_entry(429, age_seconds=10)])
next_at = pool.next_available_at()
assert next_at is not None
# Should be ~60s from exhaustion, not ~3600s. The entry was exhausted 10s
# ago, so the remaining wait is ~50s.
remaining = next_at - time.time()
assert remaining < EXHAUSTED_TTL_SOLE_CREDENTIAL_SECONDS, (
f"next_available_at returned {remaining:.0f}s remaining — expected < "
f"{EXHAUSTED_TTL_SOLE_CREDENTIAL_SECONDS}s (sole-credential cooldown)"
)
assert remaining < 300, (
f"next_available_at returned {remaining:.0f}s — should be seconds, not hours"
)
def test_multi_key_429_keeps_full_bench(tmp_path, monkeypatch):
"""With more than one non-DEAD entry there IS something to rotate to, so the
short cooldown must not kick in — both recently-throttled keys stay benched."""
pool = _load(
tmp_path,
monkeypatch,
[
_entry(429, age_seconds=90, cred_id="cred-1", priority=0),
_entry(429, age_seconds=90, cred_id="cred-2", priority=1),
],
)
assert pool.has_available() is False
assert pool.select() is None
+33
View File
@@ -371,6 +371,39 @@ class TestClassifyApiError:
assert result.retryable is True
assert result.should_fallback is False
def test_404_bare_model_id_missing_prefix_is_model_not_found(self):
"""A bare id the provider only serves as ``vendor/id`` is malformed.
Regression for #78796: NVIDIA NIM answers a prefix-less
``nemotron-3-ultra-550b-a55b`` with a naked ``404 page not found``.
Without the catalogue check this fell into the generic branch and
burned three retries on a deterministic failure, reporting what
looked like an outage.
"""
e = MockAPIError("404 page not found", status_code=404)
result = classify_api_error(
e, provider="nvidia", model="nemotron-3-ultra-550b-a55b"
)
assert result.reason == FailoverReason.model_not_found
assert result.retryable is False
def test_404_correctly_prefixed_model_stays_generic(self):
"""A properly prefixed id hitting a 404 is a real endpoint problem —
it must keep the retryable generic classification."""
e = MockAPIError("404 page not found", status_code=404)
result = classify_api_error(
e, provider="nvidia", model="nvidia/nemotron-3-ultra-550b-a55b"
)
assert result.reason == FailoverReason.unknown
assert result.retryable is True
def test_404_unknown_bare_model_stays_generic(self):
"""A local NIM container isn't in the catalogue — no verdict invented."""
e = MockAPIError("404 page not found", status_code=404)
result = classify_api_error(e, provider="nvidia", model="my-local-nim")
assert result.reason == FailoverReason.unknown
assert result.retryable is True
# ── Provider policy-block (OpenRouter privacy/guardrail) ──
+35
View File
@@ -47,6 +47,41 @@ _ensure_discord_mock()
from plugins.platforms.discord.adapter import DiscordAdapter # noqa: E402
@pytest.mark.asyncio
async def test_send_rejects_whitespace_and_records_failed_final_reply(
caplog, monkeypatch, tmp_path
):
monkeypatch.setenv("HERMES_HOME", str(tmp_path))
monkeypatch.setenv("DISCORD_MISSED_MESSAGE_BACKFILL", "true")
adapter = DiscordAdapter(PlatformConfig(enabled=True, token="***"))
channel = SimpleNamespace(send=AsyncMock())
get_channel = MagicMock(return_value=channel)
adapter._client = SimpleNamespace(
get_channel=get_channel,
fetch_channel=AsyncMock(),
)
with caplog.at_level("WARNING"):
result = await adapter.send(
"555",
" \n\t ",
reply_to="123",
metadata={"notify": True},
)
assert result.success is False
assert result.error == "Refusing to send empty message"
get_channel.assert_not_called()
channel.send.assert_not_awaited()
row = adapter._with_discord_recovery_db(
lambda conn: conn.execute(
"SELECT status, replied, outage_response, response_message_id "
"FROM discord_messages WHERE message_id='123'"
).fetchone()
)
assert tuple(row) == ("failed", 0, 0, None)
assert "Dropped empty message to chat=555" in caplog.text
def _voice_adapter(reference_obj, *, native_result=None, native_error=None):
adapter = DiscordAdapter(PlatformConfig(enabled=True, token="***"))
ref_msg = SimpleNamespace(id=99, to_reference=MagicMock(return_value=reference_obj))
+57
View File
@@ -146,6 +146,63 @@ class TestCaptureLogSnapshot:
assert len(kept) == 10
class TestMissingLogNote:
"""A missing log explains itself when the writer isn't this backend.
`hermes debug share` runs on the backend, so a desktop connected to a
remote/docker/SSH backend can never contribute desktop.log. Reporting a
bare absence sends triage after a client-side bug it cannot see.
"""
def test_backend_written_log_reports_plain_absence(self, hermes_home):
from hermes_cli.debug import _capture_log_snapshot
(hermes_home / "logs" / "agent.log").unlink()
snap = _capture_log_snapshot("agent", tail_lines=10)
assert snap.full_text is None
assert snap.tail_text == "(file not found)"
def test_client_written_log_names_its_writer_and_path(self, hermes_home):
from hermes_cli.debug import _capture_log_snapshot
(hermes_home / "logs" / "desktop.log").unlink()
snap = _capture_log_snapshot("desktop", tail_lines=10)
assert snap.full_text is None
assert "not on this host" in snap.tail_text
assert "Hermes Desktop" in snap.tail_text
# The reader needs the path to collect by hand on the client machine.
assert str(hermes_home / "logs" / "desktop.log") in snap.tail_text
def test_present_client_log_is_captured_normally(self, hermes_home):
"""A local backend still reads desktop.log — the note is only for a miss."""
from hermes_cli.debug import _capture_log_snapshot
snap = _capture_log_snapshot("desktop", tail_lines=10)
assert "backend spawned" in snap.tail_text
assert "not on this host" not in snap.tail_text
def test_empty_client_log_is_empty_not_absent(self, hermes_home):
"""An empty file means the app ran and logged nothing — a different fact."""
from hermes_cli.debug import _capture_log_snapshot
(hermes_home / "logs" / "desktop.log").write_text("")
snap = _capture_log_snapshot("desktop", tail_lines=10)
assert snap.tail_text == "(file empty)"
def test_report_carries_the_note_for_a_remote_backend(self, hermes_home):
"""The uploaded report — what people paste into support — must explain it."""
from hermes_cli.debug import collect_debug_report
(hermes_home / "logs" / "desktop.log").unlink()
report = collect_debug_report(log_lines=10, dump_text="dump\n")
assert "--- desktop.log" in report
assert "not on this host" in report
# ---------------------------------------------------------------------------
+50
View File
@@ -137,3 +137,53 @@ class TestDeepseekCanonicalAndReasonerMapping:
def test_reasoner_keywords_map_to_v4_flash(self, model):
assert _normalize_for_deepseek(model) == "deepseek-v4-flash"
# ── Regression: issue #78796 ───────────────────────────────────────────
class TestIssue78796NvidiaPrefixRepair:
"""A bare NVIDIA model id must regain its ``vendor/`` prefix.
build.nvidia.com serves ``nvidia/nemotron-…``; a bare
``nemotron-3-ultra-550b-a55b`` returns a naked ``404 page not found``
that never names the model, so the failure reads like an outage.
"""
@pytest.mark.parametrize("model,expected", [
("nemotron-3-ultra-550b-a55b", "nvidia/nemotron-3-ultra-550b-a55b"),
("nemotron-3-super-120b-a12b", "nvidia/nemotron-3-super-120b-a12b"),
(
"nemotron-3-nano-omni-30b-a3b-reasoning",
"nvidia/nemotron-3-nano-omni-30b-a3b-reasoning",
),
])
def test_bare_nemotron_regains_prefix(self, model, expected):
assert normalize_model_for_provider(model, "nvidia") == expected
def test_third_party_model_gets_its_own_vendor(self):
"""NIM also hosts third-party models — the prefix is the catalogue's,
not a hardcoded ``nvidia/``."""
assert normalize_model_for_provider("glm-5.2", "nvidia") == "z-ai/glm-5.2"
@pytest.mark.parametrize("model", [
"nvidia/nemotron-3-ultra-550b-a55b",
"z-ai/glm-5.2",
])
def test_already_prefixed_is_untouched(self, model):
assert normalize_model_for_provider(model, "nvidia") == model
@pytest.mark.parametrize("model", [
"my-local-nim-container",
"some-finetune-v2",
])
def test_unknown_names_pass_through(self, model):
"""The same provider id fronts local NIM containers. An id absent from
the catalogue is a lookup miss, not a guess — leave it alone."""
assert normalize_model_for_provider(model, "nvidia") == model
def test_other_providers_unaffected(self):
assert normalize_model_for_provider("my-model", "custom") == "my-model"
assert (
normalize_model_for_provider("claude-sonnet-4.6", "openrouter")
== "anthropic/claude-sonnet-4.6"
)
@@ -0,0 +1,105 @@
import time
import pytest
from hermes_state import SessionDB
@pytest.fixture
def db(tmp_path):
database = SessionDB(tmp_path / "state.db")
try:
yield database
finally:
database.close()
def _last_read(db, sid):
row = db._conn.execute(
"SELECT last_read_at FROM sessions WHERE id = ?", (sid,)
).fetchone()
return row["last_read_at"] if row is not None else None
def _row(db, sid):
rows = db.list_sessions_rich(include_archived=True)
return next(s for s in rows if s["id"] == sid)
def test_untracked_sessions_are_read(db):
"""NULL watermark = never tracked = read, so shipping the column doesn't
badge a user's entire pre-feature history at once."""
db.create_session(session_id="s1", source="cli")
db.append_message(session_id="s1", role="user", content="hi")
assert _last_read(db, "s1") is None
assert _row(db, "s1")["unread"] is False
def test_mark_read_then_new_activity_flips_back_to_unread(db):
db.create_session(session_id="s1", source="cli")
db.append_message(session_id="s1", role="user", content="hi")
assert db.set_session_read("s1") is True
assert _row(db, "s1")["unread"] is False
# New activity postdating the watermark makes it unread again without
# any write on the message path.
time.sleep(0.01)
db.append_message(session_id="s1", role="assistant", content="reply")
assert _row(db, "s1")["unread"] is True
def test_mark_unread_explicitly(db):
db.create_session(session_id="s1", source="cli")
db.append_message(session_id="s1", role="user", content="hi")
db.set_session_read("s1")
assert db.set_session_read("s1", read=False) is True
assert _last_read(db, "s1") == 0.0
assert _row(db, "s1")["unread"] is True
def test_missing_session_returns_false(db):
assert db.set_session_read("nope") is False
def _compression_pair(db: SessionDB):
base = time.time() - 100
db.create_session("root", source="cli")
db.create_session("tip", source="cli", parent_session_id="root")
db._conn.execute(
"UPDATE sessions SET started_at = ?, ended_at = ?, end_reason = 'compression', message_count = 1 WHERE id = 'root'",
(base, base + 10),
)
db._conn.execute(
"UPDATE sessions SET started_at = ?, message_count = 1 WHERE id = 'tip'",
(base + 20,),
)
db._conn.commit()
def test_reading_compression_tip_stamps_whole_lineage(db):
_compression_pair(db)
assert db.set_session_read("tip") is True
root_read = _last_read(db, "root")
assert root_read is not None and root_read > 0
assert root_read == _last_read(db, "tip")
# The projected conversation row (root surfaced as tip) derives read.
rows = db.list_sessions_rich(order_by_last_active=True)
assert [s["id"] for s in rows] == ["tip"]
assert rows[0]["unread"] is False
def test_marking_root_unread_marks_projected_conversation(db):
_compression_pair(db)
db.set_session_read("tip")
assert db.set_session_read("root", read=False) is True
rows = db.list_sessions_rich(order_by_last_active=True)
assert [s["id"] for s in rows] == ["tip"]
assert rows[0]["unread"] is True
+293
View File
@@ -0,0 +1,293 @@
#!/usr/bin/env bash
# Prove a user on some earlier commit can reach this one.
#
# Installs a real, earlier Hermes the way a user does, applies ONE update route,
# and requires the checkout to land on this commit with a working `hermes`.
#
# Nothing here is mocked. scripts/dev-sandbox.sh provides the fake Internet --
# a bubblewrap sandbox with no writable host mounts, a MITM proxy serving the
# canonical install.sh URL, and a git-upload-pack shim standing in for
# github.com -- so `install.sh` really installs uv, a managed Python, Node and
# the venv, cloning "github.com" over the ssh-first path a user hits.
#
# One route per run, on a sandbox built from scratch, because the routes are only
# meaningful from a pristine install. Sharing one install across routes -- or
# rewinding the checkout with `git reset --hard` between them -- leaves the
# second route running against a tree the first already updated (same venv, same
# installed console script, same __pycache__), which is not the state any real
# user is in: a route can then pass only because its predecessor did the work,
# and a failure in the first leaves the second exercising something undefined.
# If you add a route, give it its own run.
#
# Usage:
# tests/install/install-update-e2e.sh --route update|installer
# [--install-ref REF] [--keep]
#
# --route which update path to exercise (required):
# update `hermes update`
# installer re-running the curl one-liner over the checkout
# --install-ref what to install first; anything git resolves (a branch, a
# tag like v2026.7.7, or a SHA reachable from main).
# Default: refs/heads/main.
#
# Requires a CLEAN worktree: every dev-sandbox invocation re-derives fake main
# from the working copy, so uncommitted changes move the update target between
# the call that installs and the call that verifies.
set -euo pipefail
ROUTE=""
INSTALL_REF="refs/heads/main"
KEEP=false
while [ "$#" -gt 0 ]; do
case "$1" in
--route)
[ "$#" -ge 2 ] || { echo 'error: --route needs a value' >&2; exit 1; }
ROUTE="$2"; shift 2 ;;
--install-ref)
[ "$#" -ge 2 ] || { echo 'error: --install-ref needs a value' >&2; exit 1; }
INSTALL_REF="$2"; shift 2 ;;
--keep) KEEP=true; shift ;;
-h|--help) sed -n '2,35p' "$0"; exit 0 ;;
*) echo "error: unknown argument: $1" >&2; exit 1 ;;
esac
done
case "$ROUTE" in
update|installer) ;;
'') echo 'error: --route is required (update or installer)' >&2; exit 1 ;;
*) echo "error: unknown route: $ROUTE (want update or installer)" >&2; exit 1 ;;
esac
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
cd "$REPO_ROOT"
# Keep sandbox state out of the default .hermes-sandbox so a run never clobbers
# a developer's own sandbox, and scope it per route so two routes can run
# concurrently (CI runs them as parallel matrix legs). dev-sandbox.sh joins this
# onto the worktree root and feeds it to `tar --exclude`, so it MUST be a
# relative directory name.
SANDBOX_DIR_NAME=".hermes-sandbox-e2e-$ROUTE"
export HERMES_DEV_SANDBOX_DIR="$SANDBOX_DIR_NAME"
SANDBOX_ROOT="$REPO_ROOT/$SANDBOX_DIR_NAME"
INSTALL_DIR="/home/hermes/.hermes/hermes-agent" # user-level layout (sandbox default)
FAKE_REMOTE="/work/repos/hermes-agent.git"
# Only used to fetch an old install.sh for the flag probe below; the sandbox does
# its own fetching. Same override dev-sandbox.sh honours, so a fork can retarget
# both together.
UPSTREAM_URL="${HERMES_DEV_SANDBOX_UPSTREAM:-https://github.com/NousResearch/hermes-agent.git}"
# Installer transcripts live outside the sandbox root: the sandbox is recreated
# and (unless --keep) deleted, and these logs are the most useful artifact when
# a real install breaks. Created after the dirty check below, so that a log dir
# pointed inside the repo cannot be the thing that makes the tree dirty.
LOG_DIR="${HERMES_E2E_LOG_DIR:-$(mktemp -d -t hermes-install-e2e-logs.XXXXXX)}"
step() { printf '\n\033[1;36m▶ %s\033[0m\n' "$*"; }
ok() { printf '\033[1;32m ✓ %s\033[0m\n' "$*"; }
fail() { printf '\n\033[1;31m✗ %s\033[0m\n' "$*" >&2; exit 1; }
# The sandbox's internal logs (fake-internet proxy, slirp) explain failures that
# happen BEFORE install.sh gets to say anything -- a TLS handshake the proxy
# rejected looks like a bare `curl: (35)` from outside. Copy them out where a CI
# artifact upload can find them, and echo the proxy log since it is the usual
# culprit.
collect_sandbox_logs() {
# Separate `local` statements on purpose: a single `local a=$1 b="$a"` does
# NOT see the earlier assignment, so under `set -u` the second expansion dies
# with "a: unbound variable".
local tag="$1"
local src="$SANDBOX_ROOT/root/logs"
local dest="$LOG_DIR/sandbox-$tag"
[ -d "$src" ] || return 0
mkdir -p "$dest"
cp -a "$src/." "$dest/" 2>/dev/null || true
# Print it, not just archive it: a rejected TLS handshake here is the whole
# explanation for a failure that otherwise reads as a bare `curl: (35)`, and
# whoever is reading the job log should not have to download an artifact to
# see it. In full, not tailed -- the file is short, and the useful line is not
# reliably at the end.
if [ -s "$dest/proxy.log" ]; then
echo "--- sandbox proxy.log ---" >&2
cat "$dest/proxy.log" >&2
echo "--- end proxy.log ---" >&2
fi
}
# ── preflight ──────────────────────────────────────────────────────────────
# Prefer the `sandbox` wrapper from the Nix devShell: it supplies both the PATH
# (bwrap, slirp4netns, openssl, ...) and the DEV_SANDBOX_* variables the script
# needs -- notably DEV_SANDBOX_DYNAMIC_LINKER, without which it cannot find a
# glibc loader on NixOS. Off Nix, the script is the entry point and finds its
# dependencies on the system PATH.
if command -v sandbox >/dev/null 2>&1; then
SANDBOX=(sandbox)
elif command -v bwrap >/dev/null 2>&1; then
SANDBOX=("$REPO_ROOT/scripts/dev-sandbox.sh")
else
fail 'no usable sandbox: enter the Nix devShell (for `sandbox`) or install bubblewrap'
fi
if [ -n "$(git status --porcelain)" ]; then
printf '\033[1;31m✗ working tree is dirty:\033[0m\n' >&2
git status --porcelain | sed 's/^/ /' >&2
fail 'Every sandbox invocation re-snapshots the working copy into a new
fake-main commit, so the update target would move mid-run. Commit or stash
first. (If a path above is build or log output, it needs gitignoring or to
live outside the repo.)'
fi
mkdir -p "$LOG_DIR"
if [ "$KEEP" = false ]; then
trap 'rm -rf -- "$SANDBOX_ROOT"' EXIT INT TERM
fi
rm -rf -- "$SANDBOX_ROOT"
# ── helpers ────────────────────────────────────────────────────────────────
# Does the INSTALLED hermes accept FLAG on `hermes update`?
#
# Asked of the installed binary rather than parsed out of a release's source:
# the update subcommand has lived in main.py, subcommands/update.py, and
# update_cmd.py across the releases we sample, so any static parse is a guess
# that silently rots. `hermes update --help` is the same surface a user meets,
# and argparse prints every option it accepts.
update_supports() {
local flag="$1"
in_sandbox "hermes update --help 2>&1" | grep -qF -- "$flag"
}
# Does the installer at REF accept FLAG? Read it out of that ref's own
# install.sh rather than assuming this checkout's flag set: the point of the
# matrix is to install releases from months back, whose installers predate
# options we take for granted. (Unlike the updater, the installer runs before
# anything is installed, so there is no --help to ask yet.)
#
# The ref may not be local -- the sandbox does its own fetching -- so fall back
# to fetching just that blob. Unresolvable means "flag absent", which costs a
# more conservative invocation, never a wrong one.
installer_supports() {
local ref="$1"
local flag="$2"
local script=""
script="$(git show "$ref:scripts/install.sh" 2>/dev/null)" || {
git fetch -q --depth 1 "$UPSTREAM_URL" "$ref" 2>/dev/null || return 1
script="$(git show FETCH_HEAD:scripts/install.sh 2>/dev/null)" || return 1
}
printf '%s' "$script" | grep -qF -- "$flag"
}
# Run the real install one-liner inside the sandbox. `ref` non-empty installs
# that upstream commit and promotes THIS checkout to fake main afterwards,
# leaving the state a user is in when an update is waiting; empty serves this
# worktree's own installer and points fake main here.
install_in_sandbox() {
local what="$1"
local ref="$2"
local tag="$3"
local log="$LOG_DIR/$tag.log"
local args=(install --persistent)
[ -n "$ref" ] && args+=(--install-ref "$ref")
# Installer flags have to match the installer being run, not this checkout's.
# Older releases reject options added later ("Unknown option: --skip-browser"),
# and this test deliberately installs releases from months back. --skip-setup
# goes back further than any tag we sample; anything newer is probed for.
local installer_flags=(--skip-setup)
if [ -z "$ref" ] || installer_supports "$ref" --skip-browser; then
installer_flags+=(--skip-browser)
fi
# Sandbox flags must precede `--`; the rest goes to install.sh.
args+=(-- "${installer_flags[@]}")
# Stream the installer's output to stdout AND keep a copy on disk. It is the
# substance of this test -- a real install of uv, a managed Python, Node and
# the venv -- so it belongs in the job log where anyone reading the run can
# see it, not only in an artifact they have to download. The file copy is what
# the artifact upload keeps and what the failure paths grep.
#
# `set -o pipefail` is load-bearing here: without it the pipeline reports
# tee's status and a failed install looks like a pass.
local status=0
"${SANDBOX[@]}" "${args[@]}" 2>&1 | tee "$log" || status=$?
if [ "$status" -ne 0 ]; then
collect_sandbox_logs "$tag"
fail "$what failed (exit $status)"
fi
grep -q 'Installation Complete' "$log" \
|| { collect_sandbox_logs "$tag"; \
fail "$what did not report a completed install"; }
ok "$what completed (log: $log)"
}
in_sandbox() { "${SANDBOX[@]}" --persistent bash -lc "$1"; }
# fake main's SHA is read fresh whenever it is needed, never cached across a
# sandbox invocation: each invocation re-derives it from the worktree.
sandbox_target() { in_sandbox "git --git-dir=$FAKE_REMOTE rev-parse main" | tr -d '[:space:]'; }
sandbox_head() { in_sandbox "cd $INSTALL_DIR && git rev-parse HEAD" | tr -d '[:space:]'; }
require_landed_on_target() {
local what="$1" head target
head="$(sandbox_head)"
target="$(sandbox_target)"
[ "$head" = "$target" ] || fail "$what left HEAD at $head, wanted $target"
ok "$what landed on ${head:0:12}"
}
# The real smoke test: goes through the venv launcher and imports the app, so it
# fails if the venv, dependencies, or entry point are broken.
require_hermes_works() {
local when="$1" out
out="$(in_sandbox "hermes --version" 2>&1)" \
|| { printf '%s\n' "$out" >&2; fail "hermes --version failed $when"; }
printf '%s\n' "$out" | sed 's/^/ /'
ok "hermes runs $when"
}
# ── install the earlier Hermes ─────────────────────────────────────────────
step "installing upstream $INSTALL_REF (real curl | install.sh: uv, Python, Node, venv)"
install_in_sandbox "install of upstream $INSTALL_REF" "$INSTALL_REF" install
BASE="$(sandbox_head)"
TARGET="$(sandbox_target)"
[ -n "$BASE" ] || fail "could not read the installed commit"
[ "$BASE" != "$TARGET" ] \
|| fail "install landed on the update target ($BASE); base and target must differ"
ok "installed ${BASE:0:12}; update target is ${TARGET:0:12}"
require_hermes_works 'after install'
# ── apply exactly one update route ─────────────────────────────────────────
case "$ROUTE" in
update)
step 'ROUTE: hermes update'
# `--yes` reaches the update subcommand only in later releases, and argparse
# rejects the whole invocation when it does not exist. Ask the installed
# hermes which it accepts; older ones read the prompt from stdin, so close it.
if update_supports --yes; then
update_cmd="hermes update --yes"
else
update_cmd="hermes update </dev/null"
fi
if ! in_sandbox "cd $INSTALL_DIR && $update_cmd"; then
collect_sandbox_logs update
fail "hermes update failed ($update_cmd)"
fi
require_landed_on_target 'hermes update'
require_hermes_works 'after hermes update'
;;
installer)
step 'ROUTE: installer re-run over the existing checkout'
# No ref: serves this worktree's installer and points fake main at this
# checkout, which is what the re-run must land on.
install_in_sandbox 'installer re-run' '' reinstall
require_landed_on_target 'installer re-run'
require_hermes_works 'after installer re-run'
;;
esac
printf '\n\033[1;32m✓ install/update E2E passed (route: %s, from: %s)\033[0m\n' \
"$ROUTE" "$INSTALL_REF"
[ "$KEEP" = true ] && echo " sandbox kept at $SANDBOX_ROOT"
exit 0
@@ -456,9 +456,14 @@ def test_recover_with_credential_pool_rotates_on_xai_spending_limit_403():
status_code,
error_context=None,
api_key_hint=None,
failure_reason=None,
):
assert status_code == 403
assert api_key_hint == "test-key"
# An xAI spending-limit 403 classifies as billing, and the pool
# must be told so — otherwise a sole-credential pool gives a spent
# account the transient 60s cooldown instead of the full bench.
assert failure_reason == "billing"
assert error_context == {
"reason": "personal-team-blocked:spending-limit",
"message": (
@@ -22,6 +22,10 @@ def test_credential_rotation_replaces_route_scoped_tls_settings():
_apply_client_headers_for_base_url=MagicMock(),
_replace_primary_openai_client=MagicMock(),
)
agent._reapply_route_client_config = MethodType(
AIAgent._reapply_route_client_config,
agent,
)
entry = SimpleNamespace(
runtime_api_key="new",
access_token="",
@@ -70,6 +74,10 @@ def test_credential_rotation_does_not_carry_global_headers_across_routes():
AIAgent._apply_user_default_headers,
agent,
)
agent._reapply_route_client_config = MethodType(
AIAgent._reapply_route_client_config,
agent,
)
entry = SimpleNamespace(
runtime_api_key="new",
access_token="",
@@ -101,5 +109,3 @@ def test_credential_rotation_does_not_carry_global_headers_across_routes():
headers = agent._client_kwargs["default_headers"]
assert "Authorization" not in headers
assert headers["X-Route"] == "b"
@@ -0,0 +1,262 @@
"""Per-turn adoption of ~/.hermes/.env credential edits (#67821).
A Settings save (desktop ``PUT /api/env``, ``hermes setup``) updates .env and
the saving process's os.environ, but a live session worker keeps the
base_url/api_key captured at agent init until restart — an open chat silently
kept calling the old endpoint (e.g. a local-server key sent to
api.openai.com → opaque 401).
``AIAgent._try_refresh_env_client_credentials`` re-resolves env-sourced
credentials at the start of each conversation turn and rebuilds the client
when the user edited them. It must react only to env *edits*, never to mere
divergence from the agent's current values: credential-pool rotation and
failover legitimately move the session off the env credential, and config
``model.base_url`` has higher precedence than the env override.
"""
import os
import sys
from unittest.mock import MagicMock
import pytest
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", ".."))
from run_agent import AIAgent
DEFAULT_BASE = "https://api.openai.com/v1"
LOCAL_BASE = "http://127.0.0.1:39080"
def _make_agent(*, provider="openai-api", base_url=DEFAULT_BASE, api_key="sk-old"):
agent = object.__new__(AIAgent)
agent.provider = provider
agent.requested_provider = provider
agent.api_mode = "chat_completions"
agent.base_url = base_url
agent.api_key = api_key
agent._client_kwargs = {"base_url": base_url, "api_key": api_key}
agent._fallback_activated = False
agent._replace_primary_openai_client = MagicMock(return_value=True)
agent._reapply_route_client_config = MagicMock()
return agent
@pytest.fixture
def env(monkeypatch):
"""Dict-driven stand-in for the .env/os.environ resolution chain."""
values = {}
import agent.credential_pool as cp
monkeypatch.setattr(
cp, "get_env_prefer_dotenv", lambda key: values.get(key, "")
)
return values
class TestAdoptsEnvEdits:
def test_boot_default_adopts_override_on_first_look(self, env):
"""The reported scenario: worker spawned before the user saved the
override — first turn after the save must switch to the local URL."""
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is True
assert agent.base_url == LOCAL_BASE
assert agent._client_kwargs["base_url"] == LOCAL_BASE
agent._replace_primary_openai_client.assert_called_once_with(
reason="env_credential_refresh"
)
def test_edit_between_turns_is_adopted(self, env):
"""No-op first turn, then the user saves an override → next turn
rebuilds the client onto the new endpoint."""
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is True
assert agent.base_url == LOCAL_BASE
def test_key_rotation_in_env_is_adopted(self, env):
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
env["OPENAI_API_KEY"] = "sk-new"
assert agent._try_refresh_env_client_credentials() is True
assert agent.api_key == "sk-new"
assert agent._client_kwargs["api_key"] == "sk-new"
class TestLeavesNonEnvStateAlone:
def test_unchanged_env_is_a_noop(self, env):
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
agent._replace_primary_openai_client.assert_not_called()
def test_pool_rotation_is_not_stomped(self, env):
"""After the pool rotates the session onto a different key, an
unchanged env must not flap the session back every turn."""
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
agent.api_key = "sk-rotated-pool-entry"
assert agent._try_refresh_env_client_credentials() is False
assert agent.api_key == "sk-rotated-pool-entry"
def test_custom_endpoint_wins_over_env_edit(self, env):
"""A session running on a config/pool custom endpoint (not the
registry default, not a previously-seen env value) keeps it."""
agent = _make_agent(base_url="https://my-proxy.corp.example/v1")
env["OPENAI_API_KEY"] = "sk-old"
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is False
assert agent.base_url == "https://my-proxy.corp.example/v1"
def test_skipped_while_failed_over(self, env):
agent = _make_agent()
agent._fallback_activated = True
env["OPENAI_API_KEY"] = "sk-old"
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is False
def test_skipped_for_non_api_key_provider(self, env):
agent = _make_agent(provider="openai-codex")
assert agent._try_refresh_env_client_credentials() is False
def test_skipped_for_non_chat_completions_api_mode(self, env):
agent = _make_agent()
agent.api_mode = "anthropic_messages"
assert agent._try_refresh_env_client_credentials() is False
def test_skipped_when_no_key_resolves(self, env):
agent = _make_agent()
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is False
class TestFailedRebuildRetries:
def test_failed_rebuild_rolls_back_and_retries_next_turn(self, env):
"""A failed client rebuild must not advance the edit baseline: the
agent rolls back to the still-live old client's state and the same
unchanged edit is retried on the next turn."""
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
env["OPENAI_BASE_URL"] = LOCAL_BASE
agent._replace_primary_openai_client.return_value = False
assert agent._try_refresh_env_client_credentials() is False
# Rolled back: agent state still matches the old client.
assert agent.base_url == DEFAULT_BASE
assert agent.api_key == "sk-old"
assert agent._client_kwargs == {"base_url": DEFAULT_BASE, "api_key": "sk-old"}
agent._replace_primary_openai_client.return_value = True
assert agent._try_refresh_env_client_credentials() is True
assert agent.base_url == LOCAL_BASE
assert agent._client_kwargs["base_url"] == LOCAL_BASE
class TestRouteConfigRefresh:
def test_base_url_change_recomputes_route_tls_and_headers(self, env):
"""Moving to a new endpoint must recompute route-derived TLS material
and default headers, exactly as credential-pool rotation does."""
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
env["OPENAI_BASE_URL"] = LOCAL_BASE
assert agent._try_refresh_env_client_credentials() is True
agent._reapply_route_client_config.assert_called_once_with(route_changed=True)
def test_key_only_change_keeps_route_config(self, env):
agent = _make_agent()
env["OPENAI_API_KEY"] = "sk-old"
assert agent._try_refresh_env_client_credentials() is False
env["OPENAI_API_KEY"] = "sk-new"
assert agent._try_refresh_env_client_credentials() is True
agent._reapply_route_client_config.assert_called_once_with(route_changed=False)
CUSTOM_BASE = "https://api.longcat.example/openai/v1"
@pytest.fixture
def named_custom_provider(monkeypatch):
"""Register a named custom provider (config `providers.longcat` block)."""
block = {"name": "longcat", "base_url": CUSTOM_BASE, "key_env": "LONGCAT_API_KEY"}
import hermes_cli.runtime_provider as rp
monkeypatch.setattr(
rp,
"_get_named_custom_provider",
lambda requested: block if requested == "longcat" else None,
)
return block
class TestNamedCustomProviders:
"""#67935: named custom providers resolve to provider="custom" with no
PROVIDER_REGISTRY entry — their `key_env` credential must refresh too."""
def _make_custom_agent(self, *, api_key="no-key-required"):
agent = _make_agent(provider="custom", base_url=CUSTOM_BASE, api_key=api_key)
agent.requested_provider = "longcat"
return agent
def test_key_added_mid_session_is_adopted(self, env, named_custom_provider):
"""The #67935 repro: long-lived worker spawned before the key_env var
was written to .env — the first turn after the save must pick it up."""
agent = self._make_custom_agent()
env["LONGCAT_API_KEY"] = "lc-fresh"
assert agent._try_refresh_env_client_credentials() is True
assert agent.api_key == "lc-fresh"
assert agent._client_kwargs["api_key"] == "lc-fresh"
agent._replace_primary_openai_client.assert_called_once_with(
reason="env_credential_refresh"
)
def test_key_rotation_between_turns_is_adopted(self, env, named_custom_provider):
agent = self._make_custom_agent(api_key="lc-old")
env["LONGCAT_API_KEY"] = "lc-old"
assert agent._try_refresh_env_client_credentials() is False
env["LONGCAT_API_KEY"] = "lc-new"
assert agent._try_refresh_env_client_credentials() is True
assert agent.api_key == "lc-new"
def test_unchanged_env_is_a_noop(self, env, named_custom_provider):
agent = self._make_custom_agent(api_key="lc-old")
env["LONGCAT_API_KEY"] = "lc-old"
assert agent._try_refresh_env_client_credentials() is False
agent._replace_primary_openai_client.assert_not_called()
def test_skipped_without_key_env(self, env, named_custom_provider):
"""Inline `api_key` / pool-backed entries have no env-sourced
credential to watch."""
named_custom_provider.pop("key_env")
agent = self._make_custom_agent()
env["LONGCAT_API_KEY"] = "lc-fresh"
assert agent._try_refresh_env_client_credentials() is False
def test_skipped_for_unknown_custom_provider(self, env, named_custom_provider):
agent = self._make_custom_agent()
agent.requested_provider = "someone-else"
env["LONGCAT_API_KEY"] = "lc-fresh"
assert agent._try_refresh_env_client_credentials() is False
+5 -1
View File
@@ -4771,10 +4771,12 @@ class TestCredentialPoolRecovery:
status_code,
error_context=None,
api_key_hint=None,
failure_reason=None,
):
assert status_code == 402
assert error_context is None
assert api_key_hint == agent.api_key
assert failure_reason == "billing"
return next_entry
agent._credential_pool = _Pool()
@@ -4801,11 +4803,13 @@ class TestCredentialPoolRecovery:
return []
def mark_exhausted_and_rotate(
self, *, status_code, error_context=None, api_key_hint=None
self, *, status_code, error_context=None, api_key_hint=None,
failure_reason=None,
):
assert status_code == 429
assert error_context is None
assert api_key_hint == agent.api_key
assert failure_reason == "rate_limit"
return next_entry
agent._credential_pool = _Pool()
+74
View File
@@ -747,6 +747,80 @@ def _(rid, params: dict) -> dict:
return _ok(rid, info)
@method("session.workspace.move")
def _(rid, params: dict) -> dict:
"""Re-home a STORED session's workspace into another folder/project.
Unlike ``session.cwd.set`` (which acts on a live runtime session by its UI
id), this targets a persisted row by ``session_key`` so the desktop can fix
a session that was created in the wrong directory — no live agent required.
The git branch/root columns are REPLACED (not merely enriched), because the
whole point of the move is to change which project claims the session; a
stale ``git_repo_root`` would keep it grouped under the project it left.
A live agent bound to the row follows through the runtime path too, so its
terminal/file tools re-anchor immediately; a mid-turn session refuses the
move rather than yanking the workspace out from under its tools.
"""
target = str(params.get("session_key") or "").strip()
if not target:
return _err(rid, 4007, "session_key required")
raw = str(params.get("cwd", "") or "").strip()
if not raw:
return _err(rid, 4016, "cwd required")
from hermes_constants import translate_cwd_for_wsl_backend
resolved = os.path.abspath(os.path.expanduser(translate_cwd_for_wsl_backend(raw)))
if not os.path.isdir(resolved):
return _err(rid, 4017, f"working directory does not exist: {raw}")
# Snapshot under the lock — concurrent RPCs mutate _sessions (same pattern
# as _cwd_for_session_key).
live = None
live_sid = ""
with _sessions_lock:
for sid, sess in list(_sessions.items()):
if sess.get("session_key") == target:
live, live_sid = sess, sid
break
if live is not None and live.get("running"):
return _err(rid, 4009, "session busy")
branch = _git_branch_for_cwd(resolved)
root = _git_common_repo_root_for_cwd(resolved)
with _profile_db(params) as db:
if db is None:
return _db_unavailable_error(rid, code=5007)
# A brand-new draft has no persisted row yet; the live re-home below
# still applies and the row inherits the cwd when it is first written.
row_exists = bool(db.get_session(target))
if not row_exists and live is None:
return _err(rid, 4007, "session not found")
if row_exists:
try:
db.update_session_cwd(
target, resolved, branch, root, replace_git_meta=True
)
except Exception as e:
return _err(rid, 5007, f"move failed: {e}")
if live is not None:
try:
_set_session_cwd(live, resolved)
except ValueError as e:
return _err(rid, 4017, str(e))
agent = live.get("agent")
info = _session_info(agent, live) if agent is not None else {
"cwd": resolved,
"branch": branch,
"project": _project_info_for_cwd(resolved),
"lazy": True,
}
_emit("session.info", live_sid, info)
return _ok(rid, {"cwd": resolved, "branch": branch, "git_repo_root": root})
@method("session.active_list")
def _(rid, params: dict) -> dict:
"""Return live TUI sessions in this gateway process.
+3
View File
@@ -274,6 +274,9 @@ _LONG_HANDLERS = frozenset(
"session.compress",
"session.list",
"session.resume",
# Workspace re-home runs git branch/root subprocess probes against an
# arbitrary folder — inline they'd stall the reader on a slow mount.
"session.workspace.move",
"shell.exec",
"skills.manage",
"slash.exec",
@@ -168,6 +168,8 @@ interface PluginContext {
rest: <T>(path: string, opts?: PluginRestOptions) => Promise<T>
/** Live WebSocket to this plugin's own namespace. Returns a disposer. */
socket: (path: string, onMessage: (data: unknown) => void) => () => void
/** The curated OS door: native notification, open-external, reveal-in-file-manager, clipboard. */
os: PluginOs
/** Plugin-scoped JSON persistence (keys live under `hermes.plugin.<id>.`). */
storage: PluginStorage
}
@@ -370,6 +372,10 @@ host.state.viewport // ReadableAtom<{ width, height, narrow }>
host.notify({ kind, message, title?, detail?, action? }) // toast; returns id
host.notifyError(error, fallbackMessage) // toast an error
ctx.os.notify({ title, body?, silent? }) // native OS notification (attributed to your plugin)
ctx.os.openExternal(url) // OS default handler (browser, mail, spotify:) → Promise<boolean>
ctx.os.revealPath(path) // reveal in Finder / Explorer → Promise<boolean>
ctx.os.writeClipboard(text) // system clipboard → Promise<boolean>
host.navigate('/route') // hash-route navigation
host.onEvent(type, fn) // gateway event stream ('*' = all); returns disposer
host.logs(...) // tail an app log file
@@ -385,6 +391,18 @@ listener can't affect app dispatch. Every `host` door is async-safe: a sync thro
from an internal helper (e.g. no desktop bridge in a plain browser) becomes a
rejection your `.catch()` sees, never an error-boundary crash.
`ctx.os` is the curated OS door — every way a plugin reaches outside the app
window, in one namespace attributed to your plugin. `ctx.os.notify` posts a
**native OS notification** — the same Electron pipeline the app's own
approval/turn alerts use. It fires only while the user is away from Hermes
(backgrounded / unfocused); use `host.notify` for the in-app toast when
they're looking at the app. Users can silence it per device under Settings ▸
Notifications ▸ "Plugin notifications", and repeats from the same plugin are
throttled, so treat it as a signal for genuinely notable events — not a log.
The other doors (`openExternal`, `revealPath`, `writeClipboard`) resolve
`false` instead of throwing when the capability isn't available (older desktop
shell, plain browser) — branch on the result rather than sniffing the bridge.
## Data layer — React Query + nanostores
Plugins share the app's single `QueryClient`, so plugin queries cache, dedupe,
@@ -597,7 +615,7 @@ not treat this pipeline as a trust boundary.
| Category | Exports |
|----------|---------|
| Host | `host` (`.state.*`, `.notify`, `.notifyError`, `.navigate`, `.onEvent`, `.logs`, `.status`, `.restartGateway`, `.request`) |
| Plugin contract | `HermesPlugin`, `PluginContext`, `PluginContribution`, `PluginStorage`, `PluginRestOptions`, `Contribution` |
| Plugin contract | `HermesPlugin`, `PluginContext`, `PluginContribution`, `PluginStorage`, `PluginOs`, `PluginRestOptions`, `PluginNativeNotificationInput`, `Contribution` |
| Area constants | `PANES_AREA`, `ROUTES_AREA`, `SIDEBAR_NAV_AREA`, `STATUSBAR_AREAS`, `TITLEBAR_AREAS`, `PALETTE_AREA`, `KEYBINDS_AREA`, `THEMES_AREA`, `COMPOSER_AREAS` |
| Area payloads | `RouteContribution`, `SidebarNavContribution`, `StatusbarItem`, `TitlebarTool`, `PaletteContribution`, `KeybindContribution`, `ComposerMiddleware`, `ComposerAttachmentProvider` |
| React / state | `useValue`, `atom`, `computed`, `useQuery`, `useMutation`, `useQueryClient`, `queryClient`, `Contribute` |