2770f93064aa41c19f001a0efa8cfccb3deafbf8
83 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e7657792df |
refactor(themes): web dashboard presets derive from the desktop palette table
The desktop and the web dashboard each carried a private copy of the cyberpunk / ember / midnight / mono palettes and they had drifted: the dashboard's cyberpunk canvas was #040608 with a mint #9bffcf accent while the desktop's was #000a00 with #00ff41, ember and midnight disagreed on both canvas and accent, mono agreed only by luck. Move the raw palette table for every built-in preset into @hermes/shared (`THEME_PRESET_PALETTES`, apps/shared/src/theme-presets.ts) and make it the single source of truth: - apps/desktop/src/themes/presets.ts spreads its `colors` / `darkColors` from the shared table; the OKLCH synthesis, terminal palettes and typography stay in the desktop. Serialised BUILTIN_THEMES are byte-identical to before, so the existing `--dt-primary-solid` parity pins stay green untouched. - web/src/themes/presets.ts projects each shared preset onto its 3-slot model through one pure function, `webPresetFromShared` (background <- background, midground <- primary, warmGlow <- the midground/ring accent), so cyberpunk / ember / midnight / mono now render the desktop's palette. Web-only presets (default, default-large, nous-blue, rose) are untouched. - Invariant test (web): for every preset shared by both surfaces the dashboard canvas equals the shared background and the projected text colour keeps >= 3:1 contrast against it. Red on the previous hexes, green now. Why: one edit in one place should recolour a preset on every surface; two hand-maintained tables guarantee the drift the audit found. |
||
|
|
988d471479 | style(ts): sort imports/exports the way perfectionist wants after the rebase | ||
|
|
c3edad29ba |
docs(shared): declare M as compactNumber's top rung
The 1e9 → '1000M' test row read as a snapshot of a missing rung; the formatter comment now states the cap (token/cost figures stay well under a billion; a B suffix would collide with the bytes reading) and the row cites it. |
||
|
|
c764d6d354 |
fix(shared): ensureContrast keeps the desktop's 0.2-step ladder; TUI chain opts into 0.05
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb → #3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR body said no preset VALUE changed. The ladder is now the desktop's original algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to 1.0001, re-mixed from the source colour — with `step` as a parameter. The only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so the terminal palette is byte-identical too. Test: apps/desktop context.test.tsx iterates every builtin preset × mode, paints it through ThemeProvider and asserts --dt-primary-solid equals the value a reference copy of the old desktop algorithm computes. Sabotage (default step 0.05): 11/30 rows fail. Docs: the SDK table now lists contrastRatio as `number | null` under sRGB measures, not OKLCH. |
||
|
|
3f02259518 |
fix(shared): fuzzyRank folds [-_.] to space on both sides like the desktop picker did
The desktop model picker moved from foldIncludes (searchFold: lower-case + `[-_.]` → space on text AND query) to the shared fuzzyRank, which only lower-cased. `gpt.4o`, `claude_3` and `qwen3-8` returned zero rows where main listed gpt-4o / claude-3-opus / qwen3.8-flash, while HighlightMatches still folded — filter and highlight disagreed. fuzzyScore now folds both sides with a length-preserving fold, so positions still index the original target and all three surfaces rank a separator variant identically. Tests: three separator rows in fuzzy.test.ts and in the desktop picker test. Sabotage (lower-case only): all six fail. |
||
|
|
057c2c85fc |
refactor(slash): delete the dead web slash re-implementation; one slash parser + command.dispatch narrowing in @hermes/shared
web/src/lib/slashExec.ts and web/src/components/SlashPopover.tsx had zero
importers since the React composer was replaced by the PTY-embedded TUI
(
|
||
|
|
35022e02ed |
refactor(themes): one sRGB color-math module in @hermes/shared; measured readableOn + fine ensureContrast ladder on both surfaces
ui-tui/src/lib/color.ts called itself "the twin of the desktop app's src/themes/color.ts" and the two had already drifted: the desktop measured readableOn but used a coarse 0.2x5 ensureContrast ladder and returned 0 for unparseable luminance; the TUI had the fine 0.05x20 ladder and null-for-garbage but a luminance>0.5 threshold readableOn. Both now import the primitives from apps/shared/src/color.ts (`@hermes/shared/color`, also exported from the root index); each surface keeps only what is specific to it. No palette / preset / skin VALUE changes anywhere — only math. Sites (path::symbol → canonical): apps/desktop/src/themes/color.ts::hexToRgb → @hermes/shared/color::parseColor (deleted) apps/desktop/src/themes/color.ts::rgbToHex → @hermes/shared/color::toHex (deleted) apps/desktop/src/themes/color.ts::mix → @hermes/shared/color::mix apps/desktop/src/themes/color.ts::relativeLuminance → @hermes/shared/color::relativeLuminance apps/desktop/src/themes/color.ts::contrastRatio → @hermes/shared/color::contrastRatio apps/desktop/src/themes/color.ts::readableOn → @hermes/shared/color::readableOn (desktop wrapper readableInk pins ['#161616','#ffffff']) apps/desktop/src/themes/color.ts::ensureContrast → @hermes/shared/color::ensureContrast ui-tui/src/lib/color.ts::{Rgb,parseColor,toHex,mix,relativeLuminance,contrastRatio,readableOn,ensureContrast,lighten,darken} → @hermes/shared/color (same names) Stays desktop-only (apps/desktop/src/themes/color.ts): luminance, normalizeHex, readableInk, OKLCH set (hexToOklch, oklchToHex, oklchToSrgb255, maxChroma, hueDelta, harmonize, mixOklab, withHue, ensureContrastOklch). Stays TUI-only (ui-tui/src/lib/color.ts): liftForContrast, grayOf, desaturate, toHsl, fromHsl, retone, boostSaturation, color()/ColorChain. Importers repointed (17): apps/desktop/src/{sdk/index.ts, themes/context.tsx, themes/retint.ts, themes/retint.test.ts, themes/skin.ts, themes/vscode.ts, themes/vscode.test.ts}; ui-tui/src/{theme.ts, sdk/index.ts, sdk/apps/weather.tsx, app/createGatewayEventHandler.ts, components/agentsPanel.tsx, components/branding.tsx, components/loaders.tsx, components/overlayPrimitives.tsx, lib/color.ts, lib/color.test.ts}. Wiring: apps/shared/package.json exports './color'; apps/shared/src/index.ts re-exports; apps/desktop/tsconfig.json paths + vite.config.ts alias for '@hermes/shared/color' (ui-tui resolves the subpath via the workspace package exports, like './billing'). Behavior change (1): relativeLuminance / contrastRatio return null for unparseable input on the desktop too (previously 0, which made garbage measure like pure black). Desktop SDK export `contrastRatio` therefore widens to `number | null`. Only ensureContrastOklch relied on the number: it now treats null as "already passing / can't measure" and returns the input unchanged. Every other desktop caller passes 6-digit hex. Behavior change (2): readableOn MEASURES both candidate inks and returns the one with the higher contrast ratio (desktop semantics; the threshold version got mid-lightness accents wrong: white on #4f9e5e is 3.29:1 vs near-black 5.50:1). Signature is readableOn(bg, inks = ['#000000', '#ffffff']); the desktop passes its own pair via `readableInk` so desktop output is byte-identical. The TUI switches from the luminance>0.5 threshold to measurement: over the 185 distinct hexes in ui-tui/src/theme.ts (DARK/LIGHT seeds + built palettes) and hermes_cli/skin_engine.py, 56 flip from '#ffffff' to '#000000' — all mid-lightness accents (L 0.18–0.49, e.g. #cd7f32, #4caf50, #ef5350, #ffa726, #4dabf7) where black measures 4.6–10.8:1 against white's 1.9–4.5:1. Note the TUI never called readableOn directly; it only reaches ensureContrast's pole choice (below), and ensureContrast is only reachable via the color() chain and the theme.ts re-export (no production caller today). Behavior change (3): ensureContrast steps 0.05 x 20 from the ORIGINAL color toward the measured readableOn pole (TUI semantics). The desktop previously stepped 0.2 x 5 toward a threshold-chosen pole, so desktop-derived accents that needed a lift (skin/VS Code imports whose accent fails 4.5:1 on the sidebar, and --dt-primary-solid) may now land up to 0.15 closer to their original hue — they stop at the first passing rung. Palette VALUES are unchanged; only synthesized colors move. Also: parseColor accepts #rgb shorthand where desktop hexToRgb rejected it — strictly more permissive; the only desktop path fed raw user hex is normalizeHex, which already expands shorthand itself. Tests: apps/shared/src/color.test.ts (moved TUI parse/mix/contrast cases + two invariants): - "readableOn(%s) returns the ink with the higher measured contrast" — computes contrastRatio for each candidate in the test and asserts the returned ink is the max (a contract, not a hardcoded hex) over #4f9e5e (both ink pairs), #cba6f7, #ffffff, #101014. Sabotage: reverted readableOn to the luminance threshold → 3 red (#4f9e5e x2, #cba6f7); restored → green. - "ensureContrast(%s on %s) clears %s" — 5 failing pairs end ≥ min; plus "leaves passing and unparseable colors byte-identical". Sabotage: truncated the ladder to 3 rungs → 5 red; restored → green. ui-tui/src/lib/color.test.ts keeps only the color() chain case. Validation: apps/shared: npx tsc -p . --noEmit (0) && npx vitest run → 3 files, 30 tests passed; npm run lint clean apps/desktop: npx tsc -p . --noEmit (0); npx vitest run --project ui → 798/800 files, 7563/7572 tests; the 9 failures (src/app/messaging/index.test.tsx x8 12s-timeouts, src/lib/markdown-blocks.test.ts property fuzz 36s) are load-induced flakes under the full parallel run: both files pass in isolation on this branch (16/16) and on origin/main; neither imports color math. npm run lint 0 errors ui-tui: npm run build:ink; npx tsc -p . --noEmit (0) && npx vitest run → 168 files, 1764 tests passed; npm run lint 0 errors git diff --check clean; no new gitignored .d.ts. Handoff: desktop vs web preset palettes diverge for the four shared ids (web presets carry a 3-slot palette {background, midground, foreground(alpha 0)} + warmGlow, not the desktop's 24-slot set, so only the comparable slots are listed; web `foreground` is #ffffff alpha 0 on all four — a glow/overlay slot, not text ink). Design call for Teknium; nothing changed here. preset slot desktop web cyberpunk background #000a00 #040608 cyberpunk accent #00ff41 (primary/ring/mid) #9bffcf (midground) cyberpunk foreground #00ff41 #ffffff (alpha 0) ember background #160800 #1a0a06 ember accent #d97316 (ring/midground) #ffd8b0 (midground = desktop fg/primary) ember foreground #ffd8b0 #ffffff (alpha 0) midnight background #08081c #0a0a1f midnight accent #8b80e8 (ring/midground) #d4c8ff (midground) midnight foreground #ddd6ff #ffffff (alpha 0) mono background #0e0e0e #0e0e0e (match) mono accent #9a9a9a (ring/midground) #eaeaea (midground = desktop fg/primary) mono foreground #eaeaea #ffffff (alpha 0) |
||
|
|
65ca7eac5f |
refactor(i18n): shared define-locale/RTL/endonym scaffolding in @hermes/shared; desktop+web forward to it
Desktop and web each re-implemented the same locale plumbing: the
TranslationOverride<T> partial-catalog type, isRecord (four copies across
the two apps), mergeTranslations, the RTL_LOCALES={'ar'} set with the
documentElement.lang/dir effect, and the endonym table for the language
picker (6 entries on desktop, 17 on web, overlapping and hand-synced).
The generic parts now live once in apps/shared/src/i18n.ts (exported from
the root index and the `@hermes/shared/i18n` subpath). It is generic over
the catalog type — no Translations, no `en` — so translation catalogs stay
per-app (content decision, deliberately not merged here).
Sites (path::symbol → canonical):
apps/desktop/src/i18n/define-locale.ts::TranslationOverride, isRecord,
mergeTranslations → @hermes/shared/i18n; defineLocale is a one-liner
web/src/i18n/define-locale.ts::TranslationOverride, isRecord,
mergeTranslations → @hermes/shared/i18n; defineLocale is a one-liner
apps/desktop/src/i18n/runtime.ts::isRecord → shared isRecord
apps/desktop/src/i18n/context.tsx::isRecord, RTL_LOCALES,
applyDocumentLocale → shared isRecord / applyDocumentLocale
web/src/i18n/context.tsx::RTL_LOCALES + inline lang/dir effect
→ shared applyDocumentLocale
web/src/i18n/context.tsx::LOCALE_META literal (17 names)
→ derived from shared LOCALE_ENDONYMS (same exported shape)
apps/desktop/src/i18n/languages.ts::LOCALE_OPTIONS.name (6 names)
→ LOCALE_ENDONYMS.<id>; englishName/configValue columns stay
The six desktop endonyms were byte-identical to web's before the move.
Tests: apps/shared/src/i18n.test.ts — mergeTranslations keeps untouched
sibling keys under a nested partial override and replaces functions/arrays
wholesale without mutating the base; RTL_LOCALES ⊆ keys(LOCALE_ENDONYMS);
applyDocumentLocale is a no-op without a document. The existing desktop
context.test.tsx RTL/lang assertions keep covering the effect.
Behavior change: none.
|
||
|
|
a3d259019b |
refactor(ts): one stripAnsi in @hermes/shared (TUI's OSC/DCS/partial-CSI coverage); desktop adopts it
Three TS surfaces each carried their own ANSI stripper with different
coverage. The TUI's (OSC, DCS/SOS/PM/APC strings, complete and truncated
CSI, multi-byte non-CSI ESC sequences, stray ESC, C0 controls) is now the
single implementation at apps/shared/src/ansi.ts, exported from the root
index and the new `@hermes/shared/ansi` subpath (ui-tui has no DOM lib, so
it imports the subpath like it does for billing/skin).
Sites (path::symbol → canonical):
ui-tui/src/lib/text.ts::stripAnsi, sanitizeAnsiForRender, hasAnsi
→ moved to apps/shared/src/ansi.ts (text.ts now imports stripAnsi
from '@hermes/shared/ansi' for its own trail helpers)
ui-tui: 13 importers repointed from '../lib/text.js' to
'@hermes/shared/ansi' (createGatewayEventHandler.ts,
components/messageLine.tsx, 11 __tests__ files)
apps/desktop/src/lib/ansi.ts::stripAnsi (2 regexes) → deleted;
parseAnsi/ansiColorClass/hasAnsiCodes stay (styled-segment parser)
apps/desktop/src/app/session/hooks/use-prompt-actions/index.ts
→ imports stripAnsi from '@hermes/shared/ansi'
apps/desktop/src/components/assistant-ui/tool/fallback-model/index.ts
private SGR-only stripAnsi → deleted; imports the shared one
Tests: the TUI 'ANSI sanitizers' cases move from
ui-tui/src/__tests__/text.test.ts to apps/shared/src/ansi.test.ts, plus
one invariant: an OSC-8 hyperlink + DCS string + SGR + partial CSI tail
strips to exactly the visible text with no ESC/BEL left.
Behavior change: desktop chat system messages (use-prompt-actions) and
inline-diff chrome (stripInlineDiffChrome) now also lose OSC hyperlink
payloads, DCS strings, truncated CSI tails and C0 control bytes that the
weaker regexes let through. TUI behavior is unchanged.
|
||
|
|
172b2a722b |
refactor(ts): one compactNumber and one reasoning-effort value set in @hermes/shared
Three hand-rolled compact-number formatters and two mirrored copies of the
reasoning-effort value set collapse into apps/shared/src/format.ts and
apps/shared/src/reasoning-effort.ts, exported from the package root and as
the subpaths `@hermes/shared/format` / `@hermes/shared/reasoning-effort`
(the TUI compiles with lib ES2023 and imports subpaths only). Surfaces keep
their own label maps and UI helpers. No re-export shims remain.
Convention for compactNumber (desktop's implementation, moved verbatim):
lowercase 'k', uppercase 'M', promotion-guarded thresholds (>= 999.5 -> k,
>= 999_950 -> M) so rounding can never print "1000k", trailing ".0"
stripped, non-finite / <= 0 -> "0".
Sites (path::symbol -> canonical):
apps/desktop/src/lib/format.ts::compactNumber -> apps/shared/src/format.ts::compactNumber (moved; file deleted)
web/src/lib/format.ts::formatTokenCount -> deleted
ui-tui/src/lib/text.ts::fmtK -> deleted (text.ts's own callers use compactNumber)
apps/desktop/src/app/agents/index.tsx -> @hermes/shared
apps/desktop/src/app/chat/sidebar/chrome.tsx -> @hermes/shared
apps/desktop/src/app/chat/sidebar/session-row.tsx -> @hermes/shared
apps/desktop/src/app/command-center/index.tsx -> @hermes/shared
apps/desktop/src/app/shell/context-usage-panel.tsx -> @hermes/shared
apps/desktop/src/app/shell/titlebar-controls.tsx -> @hermes/shared
apps/desktop/src/app/skills/index.tsx -> @hermes/shared
apps/desktop/src/app/skills/mcp-tab.tsx -> @hermes/shared
apps/desktop/src/components/ui/tab-dropdown.tsx -> @hermes/shared
apps/desktop/src/lib/statusbar.tsx -> @hermes/shared
apps/desktop/src/sdk/index.ts::compactNumber -> re-exported from @hermes/shared (plugin SDK surface unchanged)
apps/desktop/src/plugins/kanban/{board,drawer}.tsx -> unchanged (import via @hermes/plugin-sdk)
web/src/components/ModelInfoCard.tsx::formatTokenCount -> @hermes/shared::compactNumber
web/src/pages/ModelsPage.tsx::formatTokenCount -> @hermes/shared::compactNumber
ui-tui/src/components/appChrome.tsx::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/components/thinking.tsx::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/app/slash/commands/session.ts::fmtK -> @hermes/shared/format::compactNumber
ui-tui/src/__tests__/text.test.ts::fmtK suite -> apps/shared/src/format.test.ts (table incl. promotion guard)
apps/desktop/src/lib/reasoning-effort.ts::REASONING_EFFORTS/REASONING_EFFORT_VALUES/
DEFAULT_REASONING_EFFORT/ReasoningEffort/isReasoningEffort -> apps/shared/src/reasoning-effort.ts
(SHORT_LABELS, reasoningEffortLabel, isThinkingEnabled, resolveReasoningEffort stay local)
apps/desktop/src/app/settings/constants.ts -> @hermes/shared
apps/desktop/src/app/settings/model-settings.tsx -> @hermes/shared
apps/desktop/src/app/shell/model-catalog-menu.tsx -> @hermes/shared (+ local reasoningEffortLabel)
apps/desktop/src/app/shell/model-edit-submenu.tsx -> @hermes/shared (+ local UI helpers)
apps/desktop/src/app/shell/model-menu-panel.tsx -> @hermes/shared
apps/desktop/src/lib/model-status-label.ts -> @hermes/shared (+ local reasoningEffortLabel)
apps/desktop/src/sdk/index.ts -> value set re-exported from @hermes/shared; label helper stays from '@/lib/reasoning-effort'
apps/desktop/src/lib/reasoning-effort.test.ts -> value-set + isReasoningEffort cases moved to apps/shared/src/reasoning-effort.test.ts
web/src/lib/reasoning-effort.ts::EFFORT_OPTIONS -> labels mapped over shared REASONING_EFFORT_VALUES (same order: none, then 7 levels)
web/src/lib/reasoning-effort.ts::VALID_EFFORTS -> Set(REASONING_EFFORT_VALUES); normalizeEffort falls back to DEFAULT_REASONING_EFFORT
Semantics kept: web `none` is selectable; desktop `none` resolves to ''
(thinking off); desktop isReasoningEffort still trims + lowercases.
Behavior change:
- web: token counts on the Models page and ModelInfoCard now print a
lowercase 'k' and are promotion-guarded: 128_000 "128K" -> "128k",
999_999 "1000.0K" -> "1M", 1_500 "1.5K" -> "1.5k". 'M' is unchanged.
- TUI: fmtK used Intl compact notation; compactNumber differs only in
suffix case and the guard: 1_000_000 "1m" -> "1M", and billions no
longer get a 'b' suffix (1_000_000_000 "1b" -> "1000M"). Sub-million
values are identical ("999", "1k", "1.5k"). Non-positive values now
print "0" instead of "-1k".
- desktop: none (its formatter moved verbatim).
Tests: apps/shared/src/format.test.ts::"compactNumber" (table incl.
999_999 -> "1M", 999_949 -> "999.9k"; fails when the promotion guard is
removed) and apps/shared/src/reasoning-effort.test.ts::"reasoning-effort"
(no duplicate values, `none` is the only non-level, default is a member;
fails on a duplicated level or a `none`-accepting isReasoningEffort).
|
||
|
|
a2ae8f229d |
refactor(ts): one fuzzy + model-search-text helper in @hermes/shared; desktop picker ranks with fuzzyRank
Three byte-identical (modulo prettier and a "keep in sync" header comment)
copies of model-search-text.ts and two of fuzzy.ts collapse into one copy
each under apps/shared/src, exported from the package root and as the
subpaths `@hermes/shared/fuzzy` / `@hermes/shared/model-search-text` (the
TUI compiles with lib ES2023 and imports subpaths, never the DOM-typed
root). The vitest suites move with the code; no re-export shims remain.
Sites (path::symbol -> canonical):
ui-tui/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank -> apps/shared/src/fuzzy.ts (moved)
web/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank -> deleted
ui-tui/src/lib/model-search-text.ts::modelSearchText -> apps/shared/src/model-search-text.ts (moved)
web/src/lib/model-search-text.ts::modelSearchText -> deleted
apps/desktop/src/lib/model-search-text.ts::modelSearchText -> deleted
ui-tui/src/lib/fuzzy.test.ts -> apps/shared/src/fuzzy.test.ts (moved)
ui-tui/src/lib/model-search-text.test.ts -> apps/shared/src/model-search-text.test.ts (moved)
ui-tui/src/components/modelPicker.tsx::fuzzyRank, modelSearchText -> @hermes/shared/fuzzy, @hermes/shared/model-search-text
web/src/components/ModelPickerDialog.tsx::fuzzyRank, modelSearchText -> @hermes/shared
web/src/lib/model-picker-filter.ts::fuzzyScoreMulti -> @hermes/shared
apps/desktop/src/components/model-picker.tsx::modelSearchText -> @hermes/shared (+ fuzzyRank, see below)
The header comment now names only the cross-language twin
(hermes_cli/model_search.py) as the thing to keep in sync.
Behavior change (desktop only): the desktop model picker used to filter
model rows with `foldIncludes` substring matching and keep the curated
order; it now ranks them with the same `fuzzyRank(models, query,
modelSearchText)` the web and TUI pickers use. What a user sees
differently while typing a query:
- subsequence queries match: "g4o" now finds "gpt-4o" (previously only
a literal substring such as "gpt-4" or "4o" matched);
- the best match floats to the top instead of rows staying in curated
order (exact > prefix > word-boundary > contiguous > scattered);
- a query that matches the provider name/slug still shows that
provider's full curated list in order, exactly as before;
- an empty query still shows the curated list verbatim.
The in-row highlight is unchanged (substring emphasis via HighlightMatches),
so a fuzzy-only hit renders without emphasis rather than mis-highlighting.
Tests: apps/desktop/src/components/model-picker.test.tsx::"orders model
rows exactly as the shared fuzzyRank does" asserts the rendered row order
equals the shared fuzzyRank order for the same inputs (fails on both the
old substring filter and a reversed ranking).
|
||
|
|
c1e0fd83f9 |
fix(shared): GatewayEventMap drops phantom keys and types child_session_id
Re-verified against the tui_gateway emitters:
- SubagentEventPayload.cost_usd / .iteration: not in
tool_progress.py::_SUBAGENT_FIELDS, never emitted → removed; the TUI's
turnController no longer copies them (its SubagentProgress keeps the
fields for spawn-history persistence).
- SubagentEventPayload.child_session_id: emitted (in _SUBAGENT_FIELDS, read
by agent_callbacks.py::_mirror_subagent_to_child) but untyped → added.
- ToolCompletePayload.error: _on_tool_complete never sets it → removed;
the TUI's completeTool drops its dead `error` parameter and renders the
trail line as non-error (which is what it always did on the wire).
- ToolStartPayload.todos: not on the wire either, but the TUI handler and
its fixtures exercise recordTodos from tool.start; kept with a comment
saying so rather than churning the handler.
- MessageCompletePayload.failure_reason: prompt_turn.py passes
result.get("failure_reason") through → `string | null`.
|
||
|
|
2435131573 |
fix(shared): JSON-RPC channel ignores non-object frames and keeps the TUI's pong-based liveness
handleFrame guarded JSON.parse but then read `frame.id` on whatever came back, so a stdout line of `null`/`42`/`"str"` threw a TypeError out of the readline handler — an uncaughtException in the Ink process, where main's TUI had caught and logged it. Non-object frames now return null (the owner logs a protocol error, as for non-JSON). The shared heartbeat counted ANY inbound frame as liveness, silently dropping the TUI's original contract (fail on an unanswered gateway.ping): a backend whose request loop is wedged but still streams deltas never tripped the deadline. `heartbeatLiveness` now selects the contract: 'response' (default, TUI) — only a pong or a response to our own request resets the deadline; 'any-inbound' — the desktop/web WebSocket client's original behaviour, which JsonRpcGatewayClient passes explicitly so that surface is unchanged. The dead 'error' branch comment in connect()'s onClose is corrected to describe the onSocketClose-intercept case it actually serves. Tests: it.each over 'null'/'42'/'"str"'/'true' asserts no throw and null; 'response' mode: pongs and request responses keep it alive, streaming deltas with unanswered pings fire onHeartbeatFailure. Sabotage (remove the object check + count any inbound): 5 tests fail with the original TypeError. |
||
|
|
bab5cece78 |
refactor(ts): one reconnect backoff in apps/shared; web events feed rides the shared client and survives reconnects
Four backoff formulas (ui-tui 1000/30s, desktop 300/15s jittered, web events
1000/30s, web PTY inline 250/3s cap 5 — untested) collapse into
apps/shared/src/reconnect-backoff.ts::reconnectBackoffDelayMs(attempt,
{baseDelayMs, capMs, jitter}). Every caller keeps its own parameters
(table in the PR body); the PTY ladder gains a test.
web/src/components/ChatSidebar.tsx hand-rolled a third WebSocket frame
dispatcher (`new WebSocket` + JSON.parse + `frame.method === "event"` switch
+ a private RpcEnvelope re-declaring shared JsonRpcFrame) for /api/events.
That socket now goes through EventsFeedClient, a notification-only subclass
of the shared JsonRpcGatewayClient (replay off, heartbeat off, connect
timeout covering ticket minting); the effect keeps only the retry ladder and
the banner. Both sidebar clients are now created once per component instead
of per `version` bump, so the shared client's seq watermarks survive a drop
and its `session.events.since` gap replay can actually fire for web
(previously the client was rebuilt on every reconnect and replay never ran).
Behavior change: web sidecar reconnects reuse the same JsonRpcGatewayClient
(gap replay now runs); the events feed's handshake `error`+`close` pair is one
`closed` transition (one retry timer, as before); no parameter of any
backoff ladder changed.
|
||
|
|
6b406f1c89 |
refactor(ts): ui-tui rides apps/shared's JSON-RPC request channel; one pending map, one heartbeat, typed RPC errors
Two independent JSON-RPC client cores existed for one backend: apps/shared's
JsonRpcGatewayClient (desktop, web) and ui-tui/src/gatewayClient.ts, which
re-implemented request ids, the pending map with timeouts, response->error
mapping, event decoding and the gateway.ping heartbeat (~200 LOC, drifted).
Split the transport-agnostic half out of the shared client into
JsonRpcRequestChannel (apps/shared/src/json-rpc-channel.ts): the owner binds a
JsonRpcTransport { send(text) } per connection generation and feeds inbound
text through handleFrame(). JsonRpcGatewayClient keeps only the WebSocket
lifecycle, seq replay and the typed event hub on top of it; the Ink TUI keeps
only its two transports (spawned child stdio, attached socket) and its
mount-order event buffering, and delegates everything else.
Behavior change:
- TUI RPC errors now carry the JSON-RPC `code` / `data` (JsonRpcGatewayError)
instead of a bare Error(message); the TUI's timeout text is now the shared
"request timed out after Ns: <method>" (was "timeout: <method>", matched by
no caller) and callers may pass a per-call timeout.
- TUI heartbeat liveness counts any inbound frame (shared semantics) rather
than tracking one in-flight ping id; the interval/deadline are unchanged
and pings no longer carry the unread `last_activity_ms` param.
- Desktop isMissingRpcMethod reads the -32601 code first and only regexes the
message for code-less (IPC-flattened) errors, so a tool result that merely
mentions "unknown method" no longer reads as a capability verdict.
- Shared connect() now settles on a `close` during the handshake (auth-gate
4401/4403) instead of waiting out the 15s connect timeout, and
invalidate()/close() drop the socket generation before calling close() so a
synchronous close event cannot run the closed-path twice.
|
||
|
|
36773e0d78 |
refactor(ts): one GatewayEventMap in apps/shared typed from tui_gateway emitters; drop never-emitted tool.progress
Three TypeScript clients each declared their own copy of the tui_gateway wire
types and had drifted apart: apps/shared had a partial GatewayEventName union
with a `(string & {})` escape hatch, ui-tui/gatewayTypes.ts a 150-line
discriminated union, and apps/desktop an `RpcEvent<T>` that was field-for-field
the shared GatewayEvent with `type: string`. None matched the emitter:
message.complete lacked warning/status/error/recoverable/error_surface,
tool.start/tool.complete lacked args/result, SessionResumeResponse lacked
session_key/messages_omitted/hydrating/auto_continue/todo_state, three
different ModelOptionProvider shapes disagreed on fields, and all three unions
handled a `tool.progress` event that no Python emitter has ever produced.
Now:
* `apps/shared/src/gateway-events.ts` is the single home: payload interfaces
typed from the Python emitters (file::symbol cited per interface),
`BackendGatewayEventMap` (89 backend names) + `ClientLocalGatewayEventMap`
(5 TUI-synthetic transport events, clearly marked, excluded from the
contract) merged into `GatewayEventMap`; `GatewayEvent<K>` is discriminated
on `type` with `seq` typed. RPC shapes shared by 2+ surfaces live beside it
(ModelOptionProvider = union of every field hermes_cli/inventory.py sets,
incl. pricing_pending/free_tier_pending; SessionResumeResponse<Info>;
SessionListItem with resolved_id; Usage).
* `JsonRpcGatewayClient.on<K>` is keyed by event name; the gateway.ready
heartbeat/replay_epoch and per-frame `seq` reads are typed instead of cast.
* ui-tui and apps/desktop import the shared names; their local duplicates are
deleted (no re-export shims — importers are repointed; the desktop plugin
SDK barrel keeps its public `RpcEvent` name as an alias of GatewayEvent).
web/src repoints ModelOptionProvider/ModelOptionsResponse.
* `tool.progress` handling is removed from the TUI handler/turnController,
desktop event sets/tools handler, shared union, tests, and two docs
(`grep '"tool.progress"' tui_gateway/` = 0 hits; the `display.tool_progress`
config mode is unrelated and untouched).
* `message.complete.warning` (history-commit note from
prompt_turn.py::_complete_turn_payload) is typed and surfaced on both
surfaces through their existing notice paths (TUI pushActivity 'warn',
desktop notify kind 'warning').
Contract: `apps/shared/src/gateway-events.json` is the sorted list of
backend-emitted names. `tests/tui_gateway/test_gateway_event_contract.py`
collects names from the Python emitter side (emit-helper literals, the
`.request → .expire` table, change-watcher table, child delta mirror,
subagent relay, desktop_ui tool emitters, gateway.ready/setup.ready/
browser-controller frames) and asserts emitted == JSON in both directions.
`apps/shared/src/gateway-events.test.ts` asserts BACKEND_EVENT_NAMES (which
the map type is `satisfies`-checked against) == JSON. Sabotage-verified: a
fake JSON name fails both tests; a fake TS name fails tsc + vitest; a fake
Python `_emit("...")` fails pytest.
|
||
|
|
0dcadf6f41 |
revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces. Later non-Wisdom work on shared files (guest onboarding i18n, dashboard startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call sites and config were stripped from those files. |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
809654d488 |
refactor(desktop): keep main's boot classification and share the ws-URL guard
Follow-up to the renderer-dial retry: - Drop the `!connectionDescriptorResolved` gate so `bootFailureIsRetryable()` still decides every failure main can see (#82679 contract preserved for post-descriptor mint failures); the renderer-owned dial is OR'd in as the one failure main cannot classify. The budget check now short-circuits before the IPC, as on main. - Four booleans → one `stage` variable; the `isGatewayReauthRequired` term is dropped (reauth is raised before the dial stage, so it was unreachable). - `isGatewayWebSocketUrl` moves to `apps/shared/src/json-rpc-gateway.ts` and `JsonRpcGatewayClient.connect()` uses it, so the hook's "valid dial" predicate cannot drift from what `connect()` accepts. - Tests: keep the three that bind behaviour (first dial retries under a stale ready snapshot; invalid URL terminal; post-connect failure terminal), drop the four that pinned the removed gate or duplicated the existing #82679 bound test. 46/46 green; reverting the hook to main makes the retry test fail. |
||
|
|
1175ac4bb9 | feat(desktop): default glass to 29% tint on the sidebar | ||
|
|
393af4a310 |
fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core: - TodoStore gains a monotonic in-memory revision; the todo tool result returns it so clients can reject stale updates - tui_gateway emits a dedicated todo.updated full-snapshot event that bypasses optional tool-progress display settings - session resume/activate responses attach the authoritative todo snapshot; renderer restores it with revision arbitration - desktop store tracks per-session revisions and rejects regressions The session_todo_state DB table from the original PR is intentionally dropped: canonical todo tool results already persist in conversation history, so resume paths derive the snapshot from the stored transcript instead of a parallel store. |
||
|
|
9f05b06589 |
Revert "Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay"
This reverts commit |
||
|
|
df7d7f6e8d |
Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay
feat(gateway): slim WS-only server — remove FastAPI/uvicorn from desktop boot path |
||
|
|
693dd5d042 | fix(desktop): guard stale refresh ownership | ||
|
|
874fab0ce0 |
feat(chat-plane): trace_id + turn telemetry, transient-delta split, seq-namespace epoch
Chat/event-plane quality work for the amended Phase 1 scope of #94484 (maintainer restructure: lean chat/event plane, no control-plane changes). Three fixes came out of a source-level comparison against OpenHands, Chainlit, VS Code, Zed, LangGraph, and Goose. 1. Per-turn trace_id + active-turn telemetry: _start_inflight_turn mints a 12-hex trace_id; _event_frame stamps it on every event frame in the turn, so a client can correlate the full lifecycle (dispatch -> first token -> tool calls -> complete) from one identifier — none of the six surveyed projects has frame-level turn correlation. session.events.stats now reports active_turns (session_id, trace_id, elapsed_s, streaming). 2. Transient vs durable events (OpenHands StreamingDeltaEvent pattern): message.delta / thinking.delta are stamped with seqs (live ordering holds) but never buffered — one streaming turn emitted hundreds of delta frames and evicted every durable control event from the 512-slot ring, defeating replay for the exact reconnect window it exists to cover. The ring now evicts manually and records the highest DURABLE seq dropped, so truncated means real data loss and delta-only gaps no longer false-positive. 3. Seq-namespace epoch (Goose stale-cursor recovery): event_replay.EPOCH (8-hex per boot) is announced in gateway.ready and echoed by session.events.since; the client drops its seq watermarks when the epoch changes, so a stale HIGH watermark from a previous gateway process can no longer suppress replay/gap-detection forever. Legacy backends without an epoch are unaffected (client keeps watermarks). Validation: Python 85/85 across replay/entry_ws/keepalive/protocol (12 replay tests, 3 new); vitest 9/9 shared (2 new epoch tests), 66/66 desktop; full tui_gateway sweep 614/615 (1 known ordering flake, passes in isolation). Live e2e on this tree: 12-event turn -> seqs contiguous 1..12, 8 durable frames buffered + 4 deltas live-only, single trace_id on all frames, active_turns elapsed_s matches the real turn duration. Research provenance: NousResearch/hermes-agent#94484 (comparative-scan comment); techniques credited to OpenHands (transient split), Goose (epoch/stale-cursor), per maintainer-restructured plan. |
||
|
|
7c97343950 |
fmt(js): npm run fix on merge (#94410)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
beb7941236 |
fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219)
The #94219 replay was a production no-op: the server returned full JSON-RPC envelopes from session.events.since while the client's replay loop dispatches only elements with a top-level 'type' — every replayed event was silently skipped. Each side's tests validated its own assumption, so both suites stayed green. - server: events_since() now returns bare event objects (the frame's params), the exact shape the live dispatch path consumes; ring stores params directly; cross-language contract test added on both sides. - client: live frames racing an in-flight replay are parked and flushed seq-gated afterward — no double dispatch of deltas, no gap-skip from a watermark advanced past the replay window. - restart poisoning: seq counters are in-process, so a backend restart reset them while clients kept high watermarks (replay forever empty, truncated=false). New replay_epoch advertised in gateway.ready and echoed by session.events.since; the client clears watermarks on epoch change. - methods_session no longer reaches into event_replay privates (is_truncated() accessor). Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff clean. |
||
|
|
03b87d666d |
fmt(js): npm run fix on merge (#94230)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
c7577403f8 | test: drop unused afterEach import (CI eslint) | ||
|
|
87631bd8ae |
feat(tui-gateway): seq-stamped event replay for lossless desktop reconnect
Server: per-session monotonic seq on every routed event frame, bounded 512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs — session.events.since (replay newer-than-watermark, reports latest_seq + truncated so clients detect gaps) and session.events.stats (telemetry). Client: per-session seq watermarks recorded from live frames; after any successful reconnect a fire-and-forget fetchReplay() drains missed events through the normal dispatch path (recordSeq ignores non-increasing seqs, so stale replay can never regress a watermark); focus-triggered reconnect nudge in use-gateway-boot for the Electron unfocused case where macOS wake skips visibilitychange. Replay failures are swallowed by design: lossless resume is an upgrade over the previous lossy reconnect, never a new failure mode. |
||
|
|
9153be2a51 |
feat(shared): heartbeat and socket-generation invalidation in JsonRpcGatewayClient
The shared-client half of the gateway.ping heartbeat contract (#89958); tracks lastInboundAt, sends pings, invalidates a silently-dead socket. Part of #83166. |
||
|
|
b8c6547ca4 |
fix(desktop): off means off when glass is turned down
The light default carries a single point of fade so the window edge reads as glass rather than as paint. That point followed anyone who dragged the tint to zero, leaving a window that asked to be opaque sitting at 0.9999. Fade now applies only while glass is actually active, not merely selected. |
||
|
|
be3166607e |
feat(desktop): glass ships on, tuned per appearance and platform
Translucency was one number serving both appearances and both platforms, resting at zero. A lever that starts at zero is a feature nobody finds, and one number cannot serve four situations: a tint that reads as a whisper over a dark palette is a milky sheet over a light one, and the same numbers that read as frost on macOS vibrancy read as a washed sheet over Windows acrylic, which composites its own tint in DWM before the page is drawn. So the state splits. `mode` stays global — clear versus glass is a choice about the window, not the palette — while the values resolve through a ladder, per key: the appearance you are looking at, then a shared base, then the platform default. Tuning light mode stays in light mode; an untouched dark keeps inheriting. A v1 state lands in base, so a window someone already tuned crosses the upgrade with exactly what was on screen. Main reads the same defaults at window creation, because a window born opaque cannot reliably be swapped to glass afterwards. The chat backdrop goes off by default in the same pass: it was competing with the glass field for the same surface. |
||
|
|
7f9e79b2e1 |
feat(desktop): back the HUD band with the same window material the app uses
The HUD asked for vibrancy directly and always with the 'hud' material — one of the two rungs the macOS census rejected, because it collapses into under-window on blur and so changed the frost the moment another app took focus. It also ignored the translucency setting entirely: Glass off still frosted, and Windows got nothing at all. hudFrostFor is the mapping for a transparent window, beside vibrancyFor in the shared module both processes read. Two gates give it its answer: the renderer's report that the band actually covers the window, and the user's Glass setting. Off resolves to no material rather than a resting one, since a transparent window has no opaque page to hide an unwanted frost behind. Windows 11 rides setBackgroundMaterial through the same call, so the HUD follows the frost ladder on both platforms. Main self-diffs and keys the latch to the window, so a Settings change re-frosts a live HUD, a tint drag touches nothing native, and a HUD respawned on another profile is not mistaken for the window that already carried the material. |
||
|
|
00e5a361b6 |
Merge pull request #89611 from NousResearch/bb/desktop-godfiles
refactor(desktop): decompose god files into atomic modules |
||
|
|
9ab5d92fac |
fix(desktop): sort translucency named exports for eslint
Perfectionist wants values before types in the glass/Windows barrel. |
||
|
|
5a027a0081 |
refactor(desktop): give duplicated helpers one owner
Splitting the god files made a pile of copy-paste helpers visible and, for the first time, fixable — sharing them previously meant importing a god file. Hashing function bodies through the TypeScript AST found twelve groups desktop-wide; production code is now at zero duplicates. Each helper went to the module that already owns its concern: firstStringField to lib/text, the two REST 404 predicates to lib/gateway-rpc beside isMissingRpcMethod, useDebounced and prefersReducedMotion to their hooks, the superseded-bootstrap guard to electron/ssh-connection, the composer keyup handler to the trigger hook that owns the rest of that state machine, and clampDataUrlReadMaxMb to apps/shared, replacing a "keep these in sync" comment between two copies. Only helpers with no existing owner got a new file: lib/mcp-servers, lib/audio-context, lib/keyed-timeouts, lib/pointer-drag, and the command palette's status row. Error-shape predicates are the worst thing to copy — when the backend changes how it reports a missing route, every copy has to be found. |
||
|
|
eb52328857 |
feat(desktop): back window glass with Windows 11 system materials
Glass was macOS-only because it rode setVibrancy. Windows 11 22H2 has a first-party equivalent in setBackgroundMaterial, so the mode now resolves its backing per platform instead of per-OS-check: macOS keeps vibrancy, Windows 11 gets DWM acrylic / tabbed / mica, and everything older stays on Clear. No third-party native addon. Two Windows-specific details the mapping has to respect. DWM only paints the client area of a transparent window (electron#49443), so glass-capable Windows chat windows are born transparent with the opaque themed backgroundColor covering them while glass is off — a live Clear/Glass toggle then needs no window recreate. And Windows exposes three backdrops for four frost rungs, so the two heaviest both resolve to mica; the mapping stays total so a frost saved on a Mac still renders. Glass support is computed once from os.release() and shared: main uses it for the persisted default and every window, preload publishes it to the renderer so the UI can't offer a mode the window can't back. |
||
|
|
e9cefa4d75 |
feat(desktop): pre-select glass on macOS
Window Translucency shipped defaulting to Clear, which means the mode worth finding is the one nobody sees — Glass is the better-looking half and the reason the feature exists. A fresh macOS profile now starts with Glass selected. Nothing turns on. The intensity still defaults to 0, so the window is byte-for-byte what it is today until the user moves the lever; the default only decides which mode that lever will drive. windowOpacityFor stays 1 and the window is still born with its opaque backing. The one profile that must NOT flip is one already carrying a non-zero intensity with no mode recorded: it predates the setting, has been rendering as clear the whole time, and defaulting it to glass would change a window someone deliberately tuned. normalizeMode takes the saved intensity and keeps those on clear. The renderer store was hand-rolling its own copy of this rule, so it now routes through the shared normalizer and the two can't disagree. The store's default test only passed because a beforeEach reset the atom before it looked — it asserted the post-reset value, not the default, so it would have stayed green through this change. It now snapshots the atom at import time. All three mutations (default back to clear, escape hatch removed, glass leaking onto non-mac) fail the suite. |
||
|
|
efa15b5eb4 |
feat(desktop): frost picker, sidebar glass, full-range tint and peek
Restores the four features SHL0MS built on #84329 that an earlier pass on this branch had carved out, reconciled onto the shared translucency state rather than the four-atom store they were written against. - Frost picker. macOS exposes no blur-radius knob, so the vibrancy material IS the frost control. The four in the ladder come from a pixel census on macOS 26: the 14 Electron materials collapse to 9 distinct looks, and these four are the widest separations that stay distinct in BOTH appearances. sidebar/hud collapse into under-window when unfocused, which is why they're deliberately absent -- normalizeMaterial rejects them, with a test saying why. - Sidebar-only glass, the Finder shape. <body> stays the single painter and splits at the rail's live-measured edge with a hard gradient stop, so there's no smear across the seam and no per-layer tint stacking. RTL mirrors. - Full-range tint. glassSurfaceKeep runs linear to zero, so the top of the lever is bare untinted blur instead of stopping at a 30% wash. Text, cards and the composer keep their own opaque tokens, which is what makes 100% usable rather than unreadable. - Peek. The settings overlay covers the very effect its slider controls, so holding the slider ghosts the whole overlay layer and the live window becomes the preview. A counter, not a boolean: a held drag and a timed pulse from a picker click overlap, and the drag must not be cancelled by a pulse expiring underneath it. One change from the original: the peek's transition is scoped with :has() to the overlay that arms it. The version on #84329 shipped a bare `[data-overlay-surface] { transition: opacity 420ms }`, which gave every overlay in the app -- command center, cron, agents, model picker -- a 420ms opacity transition for the life of the process to serve one slider. Verified by running the candidate rules through lightningcss: the scoped selectors survive minification and no un-gated overlay transition remains. visualEffectState is pinned to 'active' at each chat window, because several materials collapse to a shared inactive look on blur -- without it the frost choice silently erases itself whenever the user clicks another app. Co-authored-by: SHL0MS <SHL0MS@users.noreply.github.com> |
||
|
|
d6adef6991 |
refactor(desktop): one owner for the translucency mapping
The mapping had grown three copies: the clear-mode ramp in electron/window-opacity.ts, a second clamp + mode normalizer in electron/translucency.ts, and a third clamp in the renderer store, with a "keep in sync" comment standing in for a shared type. Anything the two processes must agree on -- what a mode is, where the lever clamps, what intensity means as an opacity -- now lives in apps/shared/src/translucency.ts and both ends import it. electron/translucency.ts keeps only the piece that needs a BrowserWindow to mean anything (the constructor backing), and re-exports the rest so main.ts has a single import. Its relative specifier is deliberate: the electron bundle is built by esbuild with no tsconfig path resolution, so a bare @hermes/shared/translucency would typecheck and then fail to bundle -- the same constraint connection-registry.test.ts documents for backendScopeKey. The renderer's tsconfig drops its reference to the electron project. With both projects claiming the shared file, that edge made the renderer resolve it through the electron project's build output and demand a prior `tsc --build`. Nothing in src/ consumes electron's emitted types, so the reference bought nothing; `npm run typecheck` still checks both projects. |
||
|
|
decd6a73fa | fix(desktop): preserve registry route identity | ||
|
|
9caff74408 |
fix(desktop): key fan-out event consumption by (connectionId, profile)
Secondary-gateway events were tagged with connectionId (store/gateway fan-out) but no consumer read it: working/attention tracking, the pruneSecondaryGateways keep-set, and the profile-scoped event gates (skin.changed / change-watcher broadcasts / approval-mode reconcile) all keyed by session id + bare profile name. Every registered source exposes a 'default' profile (the roster force-unshifts it), so two connected gateways collided — gateway B's 'default' activity was attributed to gateway A's 'default', keeping the wrong socket alive and applying the wrong source's config/skin/cron changes. Thread connectionId through consumption using the existing composite backendScopeKey helper: - session-states records each registry-tagged event's (connectionId, profile) scope per runtime session; liveSessionScopes() projects the busy/needs-input ones as composite keys for the gateway keep-set. - recomputeKeptGateways (use-gateway-boot) seeds the keep-set with those scopes; pruneSecondaryGateways matches registry-scoped entries ONLY on their composite key, while local entries keep matching bare profile names (single-source path unchanged). - gateway-event's 'from the active profile' gates now compare the event's composite scope against the active gateway's connection via the new activeGatewayConnectionId(); untagged local/primary events behave byte-identically. Display-only surfaces that already use roster handles are untouched. |
||
|
|
e14b095f20 |
fix(tui): map lineage edit ordinals past compression prefix
Desktop/TUI count full displayed lineage after compression, but prompt.submit validated truncate ordinals against tip-only history. Translate via display_history_prefix and recover stale 4018s on Desktop. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
aff19d0251 |
feat(desktop): multi-source agents end-to-end — sockets, roster, SDK, fan-out updates
Phases 3-5 of the multi-connection campaign in one PR (per Teknium), on top of the registry (#86679) and composite-key backend routing (#86839). Agents from every registered connection are now usable side by side. Renderer socket registry (phase 3): - backendScopeKey moves to apps/shared (@hermes/shared) so main-process pool keys and renderer socket keys derive from ONE rule; the electron module keeps a byte-identical twin (tsconfig project boundaries) pinned by a cross-copy contract test. - store/gateway secondaries are scope-keyed: entries carry (connectionId, profile); registry-scoped entries dial through getConnectionFor + getGatewayWsUrlFor (fresh per-connect OAuth tickets against the right host); events keep the bare profile plus a connectionId tag; touch/idle keepalive uses the scope key; pruning keeps entries whose PROFILE has live work. New ensureGatewayForAgent/openGatewayForAgent fall through to the profile path for local/null sources — single-source behavior byte-identical. Union roster + plugin SDK (phases 3+4, the Bot Mode door): - hermes:agents:roster enumerates every connection's /api/profiles concurrently (eager REST, lazy sockets; unreachable sources report per-row; undialed ssh boxes stay connect-on-demand) and flattens through buildAgentRoster — the @name-device duplicate-handle rule applied once across all sources, pure + tested. - SDK: host.connections(), host.agents(), host.warmAgent(), host.ensureAgent() — feature-detected so plugins degrade cleanly on older Desktop builds. Fan-out updates (phase 5): - hermes:connections:update-all dispatches hermes update to every eligible source in parallel: local via the app's own applyUpdates pipeline, remote/ssh via the backend's own POST /api/hermes/update; cloud skipped as platform-managed (updateEligibility, pure + tested); per-connection result rows so one dead box can't wedge the batch. Settings → Connections gains the "Update all instances" button (shown with 2+ connections). Also: getJsonForBackend/postJsonForBackend helpers with the token/OAuth-cookie auth split; docs section updated from "staged rollout" to live behavior. Tests: +4 pure cases (cross-copy contract, roster handles, unreachable sources, update eligibility); FULL desktop suite 5115 passed; tsc renderer + electron + shared clean; eslint clean. |
||
|
|
f9d64b9a9d |
fix(cron): add reliable trigger feedback in Web and Desktop clients
- Shared per-job trigger controller (apps/shared) coalesces duplicate clicks for the same profile+job inside a mounted client while letting unrelated jobs run independently; the backend durable claim remains authoritative across windows/processes. - Two-phase feedback everywhere: the action stays disabled/spinning while the request is in flight and the terminal success/error is reported once, after the HTTP response — no premature success toast (Web), matching the Desktop info notification. - Desktop keeps the 24h trigger timeout for the synchronous long operation and fences stale profile/list responses and unmounted surfaces; the sidebar trigger button shows a spinner while busy. |
||
|
|
2cabeba563 | fix(tui,desktop): refresh context usage live during active turns | ||
|
|
cab8673ea6 |
fix(dashboard): reload loopback tabs after stale session-token closes
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and #54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart. |
||
|
|
fabc2d7d33 | fix(js): hoist eslint shared devDeps to workspace root | ||
|
|
515e88a80f | fix(sec): pin exact npm package versions everywhere |