Commit Graph

26157 Commits

Author SHA1 Message Date
Teknium f89f0a2eaa feat(prompt): default identity rewritten as a behavior spec — sizing rule, named prohibitions, anti-sycophancy, earned depth; exploration-thrift line deliberately removed (models under-explore) (#97926) 2026-08-29 07:00:53 -07:00
Teknium 5241df3d4a fix(prompt): skills-section cleanup — drop '(mandatory)' header, delete hermes-agent paragraph duplicating the help guidance, gate the skill pointer on the skill actually being installed, cut the 'when the two differ' dead clause (#97918) 2026-08-29 06:53:02 -07:00
Teknium e387cbc0aa refactor(prompt): platform-hint diet — 7 heavies compressed, −657 tok across the map (facts probe-pinned) (#97899)
* refactor(prompt): platform-hint diet — shared _MEDIA_NATIVE spine; seven heavies compressed with every verified fact intact (3,175 -> ~2,520 map total, -657)

* refactor(prompt): steer-channel note diet 225 -> 155 — marker is self-describing since its own provenance+replay clauses; prompt keeps only anti-lookalike + authority + latest-results scope (#40240/#76805 archaeology in comment)
2026-08-29 06:37:06 -07:00
Teknium 2c232761e4 fix(skills): restore grill-me name, keep frontier-rounds upgrade
Reverts the plan-interrogation rename from #97831 per maintainer decision —
the skill keeps its original grill-me name. The content upgrade (design-tree
frontier-rounds interview mechanic from mattpocock/skills' grilling) stays.
Docs page, catalog row, and sidebar entry renamed back.
2026-08-29 06:30:12 -07:00
hermes-seaeye[bot] 5831d8365a fmt(js): npm run fix on merge (#97896)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-29 13:11:00 +00:00
Hukla 447217b4cc fix(desktop): reserve space for pane tab close button (#96880) 2026-08-29 06:05:34 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00
Teknium b1ff8722a5 feat(optional-skills): add decision-questionnaire — turn a blocked decision into an async questionnaire
Ports the MIT-licensed 'to-questionnaire' skill from mattpocock/skills as
decision-questionnaire (optional-skills/productivity). Interviews the user
about the send only (recipient + needed outcomes), then drafts a
most-important-first discovery questionnaire with answer stubs to
decision-questionnaire-<slug>.md. Includes staging tests for frontmatter
standards, template sections, and de-upstreaming.
2026-08-29 05:04:13 -07:00
Teknium 7b17d02a09 feat(optional-skills): add setup-wizard-generator — bash wizard for human-only setup steps
Ports the MIT-licensed 'wizard' skill from mattpocock/skills as
setup-wizard-generator (optional-skills/devops). Generates an interactive
bash wizard that walks a human through manual procedures: opens dashboard
URLs, captures values (hidden entry for secrets), writes .env / GitHub
secrets idempotently, and confirms each stage. Vendors upstream's
template.sh library verbatim (bash -n verified) plus a staging test suite
covering frontmatter, template integrity, and de-upstreaming.
2026-08-29 05:03:52 -07:00
Teknium bfeeb50257 feat(skills): rename grill-me to plan-interrogation, fold in frontier-rounds interview mechanic
Renames the optional grill-me skill to the more descriptive plan-interrogation
and upgrades it with the design-tree frontier-rounds mechanic from
mattpocock/skills' MIT-licensed 'grilling' skill: batch all currently-askable
questions per round in dependency order, agent finds facts itself, decisions
stay with the user. Docs page, catalog row, and sidebar entry renamed.
2026-08-29 05:03:20 -07:00
Teknium e2037a6c71 fix: follow-up for salvaged PR #95943
- Widen opencode_zen_free_runtime healing to the union of the static floor,
  the in-process live memo, and the SWR disk cache — a newly-live free model
  now heals opencode-go/zen selections without a release (sibling site the
  original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
  included) so direct provider_model_ids() validation callers don't each
  block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
  list (offline fallback must not offer a model that 401s); add the newly
  live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
  memoization, negative caching, and union healing; docs note in providers.md.
2026-08-29 05:02:54 -07:00
Jackal991 d9d6112aa7 fix(opencode): revalidate keyless opencode-free catalog live against the Zen relay
opencode-free (keyless) models were served exclusively from a hardcoded
in-repo snapshot (_PROVIDER_MODELS["opencode-free"]). The SWR disk cache
only revalidated AUTHED providers — its entries were keyed by a credential
fingerprint, which keyless providers have none of — so the catalog never
refreshed against GET /zen/v1/models. When the relay delisted a free model
(e.g. x-preview-f-free, 2026-08-26) the picker kept offering it and
selecting it 401'd: "Model x-preview-f-free is not supported".

Now provider_model_ids("opencode-free") fetches the live /zen/v1/models
catalog anonymously, filters it to the anonymous-servable free tier
(excluding KEYED suffix-fakes like Go's ox-alpha-free), and falls back to
the curated static floor only when the live fetch fails or is empty. The
keyless provider gets a stable disk-cache fingerprint so the picker's SWR
path serves stale immediately while refreshing off-thread — the same
behavior authed providers already get.

Regression tests prove the fix: the delisted/newly-live model assertions
fail when the live-fetch wiring is reverted.

Closes #95914
2026-08-29 05:02:54 -07:00
Teknium 9d9f44d638 refactor(prompt): desktop hint diet — recipe-first widget teaching verified against the renderer source; setup_mcp sentence delegated to its schema (442 -> 307 tok) (#97850) 2026-08-29 04:35:08 -07:00
witcheer 299c652a66 docs(user-stories): drop the Tokyo itinerary entry
The quote in x-azed-tokyo-trip names no agent, and the linked post
credits the itinerary to MyClaw, mentioning Hermes only as another agent
that account manages. The collage promises stories about how people use
Hermes, so an entry that cannot be attributed to Hermes does not belong
in it.

Removed rather than reworded: the source does not support a
Hermes-attributed version of this story. 64 new entries remain.
2026-08-29 03:45:57 -07:00
witcheer 067e58238e docs(user-stories): add 65 new community stories
Adds 65 user stories gathered from Reddit and X, prepended to the front of
the list so the newest entries surface first.

Every quote was re-fetched from its live source and programmatically verified
to be a verbatim substring of the source text. No existing entries are
modified, reordered, or removed - this is a pure prepend.

45 Reddit, 20 X. All categories and sources already exist in the file; no
schema change.
2026-08-29 03:45:57 -07:00
Teknium 57746cbb84 fix(desktop): ::preview inline frame works on remote/URL connections — read via the mode-aware fs bridge instead of bailing to a card (vestigial gate predated /api/fs) (#97829) 2026-08-29 03:34:08 -07:00
Teknium 3b362acf48 feat(desktop): Download button on preview file cards — save any delivered file via the authenticated backend bridge (works local and remote) (#97816) 2026-08-29 03:34:04 -07:00
teknium1 71c823bbf7 refactor(skills): move grill-me to optional-skills
Per the 'when in doubt, optional' rule — plan-interview is an
on-request capability, not a weekly daily-driver for most users.

Install via: hermes skills install official/software-development/grill-me
2026-08-29 03:12:20 -07:00
teknium1 260adeb075 chore(skills/grill-me): contributor mapping + docs catalog/sidebar regen 2026-08-29 03:12:20 -07:00
rafaumeu a0590001b1 fix: shorten description to 47 chars, reformat to modern outline, add author 2026-08-29 03:12:20 -07:00
rafaumeu 21f86655bc feat: add grill-me skill — adversarial plan interview before coding
New bundled skill that stress-tests plans through structured adversarial
questioning. One question at a time, each with a recommendation, resolving
the full decision tree before any code is written.

Four-phase structure: Understanding -> Technical Decisions -> Edge Cases -> Synthesis.
Integrates with plan, subagent-driven-development, and requesting-code-review.
2026-08-29 03:12:20 -07:00
Teknium fae063fc74 fix(desktop): MEDIA: non-media files get the preview file card, not a degraded 'Open' anchor (#97812)
* fix(desktop): MEDIA:-delivered non-media files route to the preview pipeline — PDFs/data files get the file card instead of a dead 'Open' anchor (extends #84951 to every extension)

* docs(prompt): desktop guidance aligned with any-file MEDIA: delivery — preview card truth, local-markdown-image block warning
2026-08-29 02:53:42 -07:00
Teknium a2e19d484c refactor(prompt): memory/skills guidance — one builder, positive posture, session-accurate wording (537 → 255 served, −282/call) (#97760)
* refactor(prompt): diet the memory/skills guidance block — schema-taught curricula removed, form rule + pruning contract kept (537 -> 223 tok in the combined block)

* polish: literal check-glyphs in source; memory capacity posture — save proactively, replace/consolidate when full

* refactor: single spine for memory/profile guidance — form rule + capacity posture written once, variants differ only in opening frame

* refactor: ONE memory-guidance builder — frame adapts to enabled stores, body written once, positive posture leads (maintainer direction)

* fix wording: memory is loaded per SESSION, not injected per turn (maintainer correction)
2026-08-29 02:30:44 -07:00
Osham Wahab aff5125f8e fix(gateway): document turn-hold commit overshoot + route deferred-notice through i18n (#92318 review)
Reviewer feedback on #92318:
1. Bounded overshoot in the commit-in-flight branch is now documented at
   the await site — the turn can exceed _hyg_max_turn_hold_seconds by up
   to commit duration, and aborting mid-commit would corrupt the
   message-store transaction. Prevents a future 'fix' into mid-commit
   cancellation.
2. The user-facing 'Context compression deferred' bubble now routes
   through t() as gateway.compress.turnhold_deferred (en.yaml entry
   added), coordinating with the i18n surface expanding in #92338 so
   non-English users don't get hardcoded English copy.
2026-08-29 13:02:13 +05:30
Kshitij Kapoor 2abf72ad97 fix(gateway): register hygiene_max_turn_hold_seconds + flat retry-after on turn-hold abandonment
Review follow-ups on the salvaged #90845:

- hygiene_max_turn_hold_seconds registered in config_defaults next to
  its sibling hygiene knobs (run.py already read it; the key was
  undiscoverable).
- Turn-hold abandonment now records a flat 60s retry-after via the
  existing cooldown column. Without it, sustained traffic re-spawned,
  held, and cancelled a fresh compressor on EVERY turn — a per-turn
  summary-model token burn that never commits. Deliberately outside the
  x1/x3/x9 failure ladder: the compressor is healthy, so the failure
  streak must not advance (witness updated to assert exactly that
  boundary: no streak increment, flat <=120s spacing, turn-hold reason).
2026-08-29 13:02:13 +05:30
Machan-Army 31b974dfb7 fix(gateway): split turn-hold expiry from idle-timeout failure path
Introduce HygieneTurnHoldExceeded exception so turn-hold budget expiry
no longer collapses into the generic asyncio.TimeoutError handler.

- Add HygieneTurnHoldExceeded exception (availability boundary, not a failure)
- Add dedicated handler: stamps AGENT_COMPRESSION_TURNHOLD provenance,
  sends deferral notice, does NOT increment failure cooldown
- Preserve #87011 contract: idle timeout still sends 'no output' message
  and takes failure path
- Add behavior witnesses: turn-hold ≠ idle timeout, cooldown untouched

Fixes semantic boundary collapse flagged in PR #90845 review.
2026-08-29 13:02:13 +05:30
Machan-Army 543c85acbf fix(gateway): bound the hygiene-compression turn-hold so a streaming summary cannot freeze the turn
Session hygiene auto-compression runs inline on the incoming-message path
and awaits the summary worker with a progress-aware inactivity budget
(hygiene_timeout_seconds) that extends up to hygiene_total_ceiling_seconds
(default 600s). A summary model that keeps streaming tokens keeps resetting
the inactivity slice, so the wait can stretch toward the ceiling while zero
bytes reach the user — chat transports (Telegram ~30s idle-timeout) drop the
connection and the turn appears frozen, even though the gateway is healthy.

Add hygiene_max_turn_hold_seconds (default 10), a turn-hold budget that caps
the wall-clock the incoming message waits on hygiene compression. The wait
slice is additionally capped at the remaining budget so the budget is
re-evaluated even when the worker keeps the inactivity slice large. On
exceeding the budget the gateway abandons the inline wait and proceeds on
the uncompressed transcript via the existing timeout path, which revokes the
worker's commit admission (CompressionCommitFence) and defers cleanup — so a
stale compression finishing later can never overwrite the turns appended
after the wait was abandoned.

Well under the typical transport idle-timeout, this guarantees the message
is answered promptly while the detached compression completes in the
background. Configurable via compression.hygiene_max_turn_hold_seconds.

Adds a regression test: a worker that streams progress continuously (so the
inactivity slice never fires) must be abandoned once it exceeds the
turn-hold budget, the turn proceeds uncompressed, and the stale commit is
fenced (no session mutation, role alternation intact).
2026-08-29 13:02:13 +05:30
Jakub Wolniewicz 23bae43cfa fix(agent): normalize list-shaped streaming content deltas 2026-08-29 12:45:43 +05:30
kshitijk4poor c03c72a19e refactor(vertex): fold _creds_cache_key into _sa_snapshot (simplify pass)
Final /simplify-code pass on the full diff: after the content-digest
rework, _creds_cache_key had become a wrapper with zero production
callers — get_vertex_credentials inlined the same read/except dance —
kept alive only by its own tests. One helper now owns the
(bytes, cache_key) resolution for all three cases (ADC sentinel,
readable file, unreadable fallback); prod calls it, tests target it.
No behavior change: 10/10 green, same mutation-check results.
2026-08-29 12:33:44 +05:30
kshitijk4poor 91608eb20e docs: note the MEASURED_1H_PROVIDERS exception in effective_cache_ttl's docstring
Review finding: the docstring still claimed all Qwen/Alibaba routes clamp
to 5m, contradicting the new allow-list one paragraph of code below.
2026-08-29 12:23:41 +05:30
kshitijk4poor 264ae5b2c8 chore: map jack@powries.com -> powriej (attribution for #97582) 2026-08-29 12:23:41 +05:30
kshitijk4poor a377b33ffb test: replace the source-scanning call-site test with a behavior-level planner test
test_every_marker_emitting_call_site_goes_through_the_central_clamp read
production .py files with a regex and pinned an exact call-site count --
the change-detector / reads-source-in-tests antipattern AGENTS.md bans
outright. Replaced with a test that drives the real
plan_cache_sections_for_destination fan-in and asserts the emitted wire
markers: ttl=1h on the measured opencode-go route, ttl stripped on the
unmeasured opencode route. Mutation-checked: emptying
MEASURED_1H_PROVIDERS turns it red.
2026-08-29 12:23:41 +05:30
Jack Powrie d51d66e869 fix(cache): preserve the configured 1h TTL on the OpenCode Go route
`effective_cache_ttl` evaluates the generic `is_qwen_model` clamp before any
route-level allowance, so a configured `prompt_caching.cache_ttl: 1h` is
silently rewritten to `5m` for every Qwen model on opencode-go. Reproduced on
this commit's parent, no provider traffic:

    effective_cache_ttl('1h', 'opencode-go', 'qwen3.7-plus')  -> '5m'
    _build_marker('5m')                                       -> {'type': 'ephemeral'}

A repair for this exists historically -- payload d6b33faae1, merged as
a43fe4918d -- but neither commit is reachable from current main
(`git merge-base --is-ancestor` returns non-zero for both), so the regression
is live on this lineage.

Restores the precedence fix: MEASURED_1H_PROVIDERS (an allow-list holding only
opencode-go, the one route measured with a delayed read past five minutes) is
consulted ahead of the generic Qwen clamp, with NO_1H_TIER_MODELS nested inside
it for models measured to ignore the tier even on a capable route.

opencode-go stays in ALIBABA_FAMILY_PROVIDERS. That set is also the
cache-marker-layout opt-in read by
agent_runtime_helpers.anthropic_prompt_cache_policy, so narrowing it would turn
working five-minute caching into *no* caching rather than extending the window.
The two sets are kept separate on purpose and a test pins the separation.

One deliberate divergence from the historical payload, found by independent
review: that patch checked NO_1H_TIER_MODELS globally, ahead of the provider
gate. MiniMax on its own Anthropic-compatible endpoint is a separate and
genuinely cache-eligible route, and the global check regressed its configured
1h to 5m off the back of an opencode-go observation -- an unrelated-provider
change this repair must not make. Measured:

    base      effective_cache_ttl('1h', 'minimax', 'MiniMax-M2.5') -> '1h'
    payload                                                        -> '5m'
    here                                                           -> '1h'

Scope note on the evidence: the delayed-read run covered qwen3.8-max and
glm-5.2. The rule is keyed on the route, not the model, so the deployed
qwen3.7-plus is covered by it but has never itself been measured. This change
restores *sending* the requested 1h marker; it does not establish that the
provider honours 1h retention. Those stay separate claims, and the provider
labels every write `ephemeral_5m_input_tokens` whatever ttl was requested, so
only a delayed read past five minutes with no intervening call can settle it.

Tests: 11 new cases covering the deployed model, marker shape, precedence
against the generic clamp, cache eligibility surviving the repair, negative
controls for unrelated routes, provider case normalization, and a closed-set
audit of every marker-emitting call site. Ten mutations were applied to scratch
copies and each turned the suite red, including hoisting the generic clamp back
above the allowance, dropping opencode-go from ALIBABA_FAMILY_PROVIDERS,
un-nesting the model denial, and adding a new unclamped sender.

Known risk, not discharged here: this marker shape has never been sent on the
Go relay. #77217 records the sibling Zen relay returning HTTP 400 on an
unexpected marker shape, so field validation must check HTTP status before it
looks at any cache counter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QV9Ld4f5ZTqdZWncpfZcQh
2026-08-29 12:23:41 +05:30
kshitijk4poor 913b910434 fix(vertex): content-digest cache key — stat signature is not credential identity
Security review on #97701 (unsupportedpastels) found and reproduced a
metadata-signature collision: atomic replacement that preserves size
and mtime (deployment tools that restore metadata; equal-length JSON)
rotates the private key under an IDENTICAL (path, mtime_ns, size)
signature, so the cache kept serving the old identity — exactly the
failure the PR set out to close.

The stat idiom is right for config caches (guards a parse, mtime
collisions are harmless). It is wrong for a credential cache (guards
an identity). The key is now (path, sha256(content)):

- _read_sa_file() reads the file ONCE per call, returning both the
  bytes and the digest key; on a miss the credentials are built from
  that same snapshot via from_service_account_info — closing the
  stat->read TOCTOU the reviewer also flagged (key and credentials can
  never describe different bytes).
- Read failure degrades to the bare-path key and the SDK's own file
  read: byte-for-byte the pre-signature behavior.
- Cost: one read + sha256 of a ~2KB JSON per cache probe, noise next
  to the OAuth token mint the cache exists to avoid.

New regression test per the review: equal-length content swap with
mtime restored and atomic replace — asserts a NEW cache key and a NEW
Credentials object (fails on the stat-keyed version, reproducing the
reviewer's cache_keys_equal=True). Fake google-auth harness gains
from_service_account_info.

Suite 10/10; ruff green.
2026-08-29 12:21:49 +05:30
kshitijk4poor ef4cd77ffc fix(vertex): restore ADC->SA retry killed by tuple cache keys (review finding)
The /simplify-code reviewer caught a real regression in the signature-
keyed cache commit: the ADC-failure fallback still compared
`cache_key == "__adc__"` (string), but keys are tuples now — the
comparison is always False, silently disabling the retry that picks up
a service-account file added after startup. The lead's pre-verification
had checked sentinel collision, failure-path pop, and TOCTOU, but
missed this consumer of the OLD key shape.

Guard now tests the actual condition (`not resolved_path` — this
attempt was ADC) instead of a key literal, so it can't rot again if
the key shape changes. New regression test drives the full path: ADC
raises, SA file appears on re-resolution, retry succeeds — fails on
the tuple-comparison version AND on any future key-shape change that
breaks the guard.
2026-08-29 12:21:49 +05:30
kshitijk4poor 56a2623316 fix(vertex): pick up rotated service-account files — signature-keyed creds cache
Pattern-D fix (stale cache after out-of-band change): the Vertex
credentials cache was keyed on the service-account file PATH alone, so
rotating the file on disk (key revoked and re-issued, new identity)
kept serving tokens minted from the OLD Credentials object for the
life of the process. Operators rotate compromised keys precisely when
they most need the new identity to take effect.

The cache key is now the file's (path, mtime_ns, size) signature — the
established idiom (agent/skill_utils.py:414, hermes_cli/config.py:3343,
and the shape #89792 applies to model overrides). Rotation bumps the
signature, forcing one re-read; the superseded entry for the same path
is evicted on insert so the cache stays bounded at one Credentials per
file. ADC keeps a stable sentinel key ("__adc__",) and its existing
expiry/refresh handling; a stat failure degrades to the bare-path key,
i.e. exactly the pre-signature behavior.

Tests (tests/agent/test_vertex_adapter.py):
- rotation invalidates: rewrite + mtime bump -> new key, new
  Credentials object, old entry evicted (fails on main: main serves
  the first identity's object after rotation)
- stat failure falls back to bare-path key, never raises
- ADC sentinel stable across None/empty resolved paths

Suite 8/8 green; ruff green.
2026-08-29 12:21:49 +05:30
Teknium 1d8946b40b fix(prompt-caching): tool-using sessions no longer 400 behind LiteLLM Anthropic proxies (#89886)
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).

New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
2026-08-29 12:06:48 +05:30
Clifford Garwood 10e93c6ab9 fix(skills): drop redundant identical-strings guard and its vacuous tests
The guard and the tests around it pinned behavior that already existed.
fuzzy_find_and_replace rejects old_string == new_string at
tools/fuzzy_match.py:69-70, returning "old_string and new_string are
identical" — so main already answered success=False, and the three tests
asserting "identical" in the error passed with the guard deleted.

The earlier claim that this case "silently applies a no-op the model reports
as success" was wrong. Verified against main:

  {"success": false, "error": "old_string and new_string are identical",
   "file_preview": "..."}

The duplicate guard was also strictly worse: it fired before the skill
lookup and returned no file_preview, shadowing the richer message.

Removes the guard, the two identical-strings tests, and the third
parametrize case. What remains is the genuinely new behavior: an actionable
missing-old_string error, reachable through the public tool.
2026-08-28 23:17:35 -07:00
Clifford Garwood 1ff8ec9a49 test(skills): cover the patch recovery loop end to end
Asserting an error string contains "read" proves the wording, not that a
model obeying it reaches a working patch. TestPatchRecoveryLoop walks the
loop through the public tool — half-formed patch, read the error, follow it,
patch succeeds — and pins the invariants that centralizing validation could
have broken: ordering ahead of the skill lookup, new_string='' still
deleting, missing new_string still rejected, and a rejected patch leaving
the file byte-identical.

The two recovery tests fail against the pre-fix dispatcher; the rest pass
either way as regression guards.
2026-08-28 23:17:35 -07:00
Clifford Garwood 4f4e778db8 fix(skills): make skill_manage patch failures recoverable instead of a dead end
`skill_manage(action='patch')` rejected a missing `old_string` with:

    old_string is required for 'patch'. Provide the text to find.

That is a dead end. The model cannot tell whether it omitted the argument or
supplied text that did not match, so it retries blindly — and then escapes to
the neighbour that always works: `action='write_file'`, which rewrites the
entire skill file and destroys unrelated content. `skill_manage`'s own action
enum puts that destructive path one token away from the failing one.

The error now names the recovery route: `old_string` must be the EXACT text
currently in the file, read the target first (the skill's SKILL.md, or the
file named by `file_path`), copy the snippet verbatim, and do not fall back
to `action='write_file'`.

Validation lives in `_patch_skill` rather than the dispatcher. `skill_manage()`
previously returned its own bare missing-argument error before ever calling
the helper, which would leave the new guidance unreachable through the public
tool. Removing that duplicate makes the helper the single source of truth; its
{"success": False, "error": ...} flows through the same json.dumps path, so
the serialized shape is unchanged, and validation still precedes the skill
lookup — a missing old_string on an unknown skill still reports the argument
error rather than "skill not found".

Fixes #33064
2026-08-28 23:17:35 -07:00
Teknium 217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
hermes-seaeye[bot] 70c6d78cae fmt(js): npm run fix on merge (#97713)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-29 06:08:13 +00:00
StanleyStetson 24e54b55f5 fix(desktop): preserve streamed assistant text and unify atomic persistence (#95514)
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.

- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.

- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.

- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.

- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
2026-08-29 11:33:05 +05:30
Amandeep Khurana a7e7de6407 fix(desktop): make gateway file saves failure-atomic so a failed download never destroys an existing file
`pumpStreamToFile` opened the user-chosen destination with
`fs.createWriteStream`, which truncates the target the instant it opens,
and its error path then unlinked that same path. When a user picked an
existing file in the Save dialog (and confirmed the overwrite) and the
gateway dropped mid-stream, the original was gone: truncated first,
deleted second, with nothing written in its place. The data-URL
compatibility fallback (`saveGatewayFileViaDataUrl`) had the same class
of bug via `fs.promises.writeFile`, which truncates before the write
completes.

Both paths now go through one failure-atomic primitive. Bytes land in a
short, randomly named sibling temp file (`.hermes-download-<hex>.part`,
same directory so the final step is a same-volume rename), created with
`flags: 'wx'`, and are renamed onto the destination only after the whole
body has been written and the descriptor released. The destination is
never opened before that point, so a failed download leaves whatever was
there untouched.

- Ownership-gated cleanup: the temp file is unlinked only after the
  stream's 'open' event proved THIS operation created it. An exclusive
  create that fails before open (EEXIST collision, EACCES, missing
  parent) never removes a file that belongs to someone else.
- `WriteStream.close(cb)` rather than `end(cb)` before renaming: `end`'s
  callback fires on 'finish' while the fd may still be open, and Windows
  refuses to rename a file with an open handle. Falls back to `end` for
  stream shapes without `close`.
- The failure path waits for 'close' (bounded by a 2s grace period)
  before unlinking, for the same reason: `destroy()` releases the fd
  asynchronously and an unlink racing the open handle would leak the
  `.part` file on Windows.
- A rename failure (destination locked, permissions) removes the owned
  temp file and rejects; nothing is left behind.
- Fixed-length temp name so a long user-chosen filename cannot push it
  past the filesystem's name limit.
- `fsPumpDeps()` is the single production deps factory (`'wx'` create,
  `fs.promises.rename`, `fs.promises.unlink`); `writeBufferToFile()`
  routes the data-URL fallback through the same pump. `PumpDeps` gains
  `rename` and a `tempPathFor` test seam.

Tests. Fakes: temp-then-rename on success, close-before-rename ordering,
the regression itself (destination neither opened nor unlinked when the
response fails mid-stream), write-error cleanup, close-before-unlink
ordering, rename-failure cleanup, pre-open EEXIST leaves the colliding
file alone, `writeBufferToFile` success and post-open write failure, the
temp-name length bound, and the `main.ts` wiring. Real filesystem
(`gateway-file-download.fs.test.ts`, exact production deps in a scratch
dir): completed download replaces the destination with no temp left;
mid-stream failure leaves the pre-existing destination byte-for-byte
with no `.part`; failure into a fresh name leaves nothing; seeded temp
path survives a pre-open EEXIST with no rename; rename failure (directory
at the destination) cleans the owned temp; data-URL fallback success and
missing-directory failure.

Adds the contributor email mapping required by the attribution check.

Fixes #96597

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015u8q2pHVPZmxpSrgkt94jC
2026-08-29 11:33:05 +05:30
joaomarcos 6f689d0b05 test(cache): pin the prompt-cache scope isolation invariant for per-response session ids
Squash of the three commits on PR #96768 (net diff is tests-only: the
mid-series production hunk in agent/transports/codex.py was reverted
within the PR after review). Pins the cache-scope isolation invariant
for hosts that mint one physical session per response, plus the
system-prompt write-path lifecycle under a per-response session (#96570).
2026-08-29 11:33:05 +05:30
joaomarcos 88a78ecc96 fix(bedrock): recover from server-side cachePoint rejections per placement
Bedrock's cachePoint rules are per-model-family AND per-field. Amazon Nova
accepts a cachePoint block in `system` and `messages` but rejects it inside
`toolConfig.tools`, failing the whole request with

    ValidationException: Malformed input request: #/toolConfig/tools/18:
    extraneous key [cachePoint] is not permitted

so every tool-enabled Nova turn fails, with no retry path and no way for the
user to turn cache markers off (#97281).

The adapter decided placement from one static allowlist that answers only
"does this model cache at all", never "in which section". Any family whose
placement rules differ breaks 100% of turns until someone edits the table and
ships a release — the same maintenance trap the `_NON_TOOL_CALLING_PATTERNS`
comment already admits to ("if a model fails with a tool-related
ValidationException, add it here").

Make Bedrock's own verdict authoritative alongside the table: classify the
rejection by the JSON pointer AWS returns, drop the marker for that one
section, retry the request once, and remember the verdict for the rest of the
process so later turns are built clean. The other sections keep their cache
markers, so Nova still gets system/messages caching instead of losing prompt
caching wholesale. This mirrors the module's existing self-heal idiom
(`is_streaming_access_denied_error` → non-streaming `converse()`).

Applied at all four boto3 call sites: `call_converse`, `call_converse_stream`,
and both Bedrock dispatch sites in `chat_completion_helpers` (the streaming
one is the path in the report). A rejection with no marker to strip returns
None so the caller re-raises instead of looping.

Fixes #97281

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 11:33:05 +05:30
Teknium 3951ead838 fix(cache): over-length caller prompt_cache_key no longer 400s Chat Completions requests
Port from anomalyco/opencode#44571: OpenAI caps prompt_cache_key at 64
chars (DeepSeek and Zai inherit the same limit via their OpenAI-compatible
APIs) and rejects longer values with HTTP 400. The Responses transport
already bounds keys via _bounded_prompt_cache_key, but the Chat
Completions transport passed caller-supplied keys (request_overrides,
top-level or extra_body) through unmodified on both the profile and
legacy kwargs paths.

Bound caller keys with the same pck_<sha256[:24]> hash shape codex.py
uses so both transports behave identically; blank keys are dropped
instead of sent empty. Hermes-generated keys were already safe
(content-addressed pck_ hashes).
2026-08-29 11:33:05 +05:30
ericmaddox f0d5f1298b fix(caching): prevent whitespace-only text blocks in prompt cache prefix splits 2026-08-29 11:33:05 +05:30
kshitijk4poor 9c137e6163 chore: map eric.maddox@outlook.com -> ericmaddox (attribution for #97618) 2026-08-29 11:33:05 +05:30
hermes-seaeye[bot] d7c0fb9d66 fmt(js): npm run fix on merge (#97706)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-29 05:56:24 +00:00