_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
Follow-up to the tombstone salvage (review point 1): named_profile_home()
treated ANY path whose ancestor's parent dir was literally named
'profiles' as a named profile until a '.hermes' ancestor appeared. An
unrelated custom HERMES_HOME like /srv/profiles/buildcache would resolve
as profile 'buildcache', and setup_logging would raise FileNotFoundError
instead of creating logs/ — a behavior regression for non-profile users.
Recognition now requires the 'profiles' directory's parent to BE a
Hermes home: the classic ~/.hermes layout, a root carrying home marker
files (config.yaml / .env / state.db — covers Docker/custom roots), a
profiles/.deleted tombstone dir (only ever created by profile delete),
or the process's resolved default Hermes root.
Adds regression tests for the /srv/profiles/<x> false-positive shape
(named_profile_home is None, mkdir_under_hermes_home and setup_logging
still create dirs) plus the tombstone-dir and ~/.hermes anchors. Also
adds encoding= to a bare write_text in the salvaged test file
(Windows-footgun gate).
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.
Adds a regression test covering the assistant + tool_calls +
image-only-content case.
Closes#40463
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:
Replaying the pre-rewrite sidecar would resend exactly what the rewrite
removed, so it must be dropped — the cost is one cache boundary miss,
never wrong content.
`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:
agent._vision_supported = False
_imgs_removed = _strip_images_from_messages(messages) # history
if isinstance(api_messages, list):
_strip_images_from_messages(api_messages)
and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.
Reproduced with the real functions:
history content after strip : [{'type': 'text', 'text': 'look'}]
sidecar still present : True
NEXT TURN sends : 'look<IMAGE BYTES SENT LAST TURN>'
This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.
Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.
tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).
Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged
CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
The salvaged proactive validation (#53307/#76896) rejects the
b'\x89PNG...' + zero-padding fixtures test_vision_tools.py used.
Replace all 11 sites with a real 1x1 PNG (+ ignored trailing padding
for size-sensitive tests), and pin _resize_image_for_vision passthrough
in the size-limit test so it exercises rejection, not resize recovery.
With both salvaged validators in place, the resolver-boundary verify()
(#53307) catches a truncated PNG before the full-decode gate (#76896)
runs. Either error message means the bytes stayed out of history.
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.
That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.
Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.
Fixes#76039
Second review round: honoring base64Encoded recursively let any nested
dict spoof {"base64Encoded": true, "data": "<secret>"} past the redactor
(Runtime.evaluate returns arbitrary by-value JSON), and two carriers were
missed entirely — Network.streamResourceContent returns unflagged binary
bufferedData, and Network.getRequestPostData's postData was not covered.
Replace the ambient field sets with per-method exact result-path specs:
_CDP_ALWAYS_BINARY_PATHS for declared-binary paths (screenshots, PDFs,
streamResourceContent, beginFrame screenshotData, the nested
CacheStorage.requestCachedResponse.response.body) and
_CDP_FLAGGED_BINARY_PATHS for paths whose carrier object's base64Encoded
sibling gates the exemption (Network/Fetch.getResponseBody body, IO.read
data, getRequestPostData postData). Path suffixes propagate only into the
matching subtree, so base64Encoded is type information solely on trusted
carrier objects — never ambient trust in nested JSON.
Architecture-review follow-up: the method-scoped binary_payload flag skipped
redaction for every string anywhere in the result of the two listed methods,
and the same corruption stayed reachable through Network.getResponseBody /
Fetch.getResponseBody / IO.read / Network.streamResourceContent. Make the
exemption field-scoped instead: an explicit schema exempts exactly
Page.captureScreenshot.result.data and Page.printToPDF.result.data (carriers
with no flag of their own), and any dict whose base64Encoded sibling is
exactly True exempts its body/data/bufferedData string (the protocol's own
discriminator — text bodies with base64Encoded: false stay redacted). Every
other string in every result keeps full secret redaction.
_redact_cdp_output applied redact_sensitive_text(force=True) to every string
in CDP results, including the base64 screenshot/PDF payload of
Page.captureScreenshot and Page.printToPDF. The Fernet pattern (gAAAA + base64
alphabet) matches arbitrary spans inside such payloads wherever gAAAA follows a +
or /, collapsing them to first6...last4: decoded PNGs came out corrupt (valid header,
CRC failures mid-IDAT, no IEND), the persisted full copies in tool_result_storage
were redacted too, and vision_analyze then embedded corrupt images that the provider
rejected with 400 invalid_image - killing resume sessions with a misleading provider
error. Skip redaction for the two binary-payload methods: the payload is binary,
not free text, so there is no secret to protect there. Every other method keeps
full redaction.
_runtime_model_config merges the agent's current identity onto the row's
existing model_config JSON. For model and provider it only SET the key
when the agent attribute was truthy, while base_url/api_mode/
reasoning_config/service_tier already deleted stale values when falsy.
When an agent rebuilt with an empty provider (inheriting the profile
default) was persisted, the previous provider/endpoint survived in
model_config while _persist_live_session_runtime updated the model
column separately. Resume then read the fresh model from the column but
the STALE provider from model_config, silently routing the resumed chat
to the wrong endpoint (e.g. a VeniceAI/empero route under a model that
should run on the profile default).
Apply the same delete-on-falsy rule to model and provider, mirroring the
or-None deletion the CLI path (_persist_model_switch_to_session) already
uses, so a stale session state can never survive into a resume override.
Existing desynced rows self-heal on the next live metadata persist.
Adds regression tests: merge drops stale provider/model when the agent
attribute is falsy, a truthy provider overwrites the stale value, resume
overrides fall back to the billing provider instead of the stale
endpoint, a real-DB round trip heals an already-desynced row, and a
first write (existing=None) reflects only the agent's current identity.
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.
Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes#76884.
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.
Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.
Reported by ryuhaneul in #69078.
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The prior commit removed the too-blunt generic fallback (any
non-retryable 400 with image parts present -> strip and retry), but
nothing in the test suite would fail if someone brought it back --
every existing test either targets a recognized image_corrupt wording
or has no image content in the request at all.
Add test_unrelated_400_with_image_parts_does_not_strip_or_retry:
same image-bearing setup as the positive recovery test, but the 400 is
an unrelated unsupported-parameter rejection with no image wording.
Asserts exactly one provider call (no retry), the error surfaces to
the caller, and -- the regression guard -- the single call's messages
still carry the image_url part. A reintroduced generic fallback would
strip it before surfacing this unrelated error, which is exactly the
P1 that got this design reverted in the first place.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.
Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.
Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.
Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.
Two recovery layers, deliberately separate:
- Semantic split: new FailoverReason.image_corrupt with
_IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
_classify_by_message. Corrupt bytes route to strip-and-retry, never
to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
still contain image parts gets one strip-and-retry via the existing
_strip_images_from_messages helper, guarded by a new
stripped_images_this_turn one-shot flag on TurnRetryState. This
un-bricks the session for any current or future provider wording
without adding another pattern list to maintain.
Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The original ^(?!\s*#) prefix only skipped full-line comments starting
with '#'. An inline comment like:
cfg = environ.get('HOME') # os.environ available
still triggered python_os_environ because the regex matched the code part
before the '#'.
Two complementary fixes:
1. Replace ^(?!\s*#) with ^[^#\n]* in the regex — this rejects any line
where a '#' comment marker appears anywhere before os.environ.
2. Add _compute_docstring_lines() — a state machine that pre-computes
lines inside triple-quoted strings (docstrings) and skips them during
pattern matching. Also handles single-line self-contained docstrings.
6 new regression tests covering: inline comments, multi-line docstrings,
single-line docstrings, full-line comments, and a verification that real
bare dict(os.environ) code still triggers. All 85 tests pass.
Five targeted fixes for #60709 (reported by @mvanhorn):
1. ruby_env_secret: scope ENV[] to case-sensitive Ruby constant
((?-i:ENV)) — no longer matches Python env[key] dict access.
2. python_environ_get_secret: downgrade critical→medium — reading
an API key via os.environ.get() is normal auth, not exfiltration.
3. python_os_environ: skip comment lines with ^(?!\s*#) — no longer
flags os.environ references in docstrings or code comments.
4. deception_hide: downgrade critical→high + negative lookahead for
UX guidance context (unless/except/until/confirm/diagnose/verify).
5. oversized_skill: downgrade high→low + raise cap 1MB→5MB — large
skills are legitimate; structural size is informational only.
All 80 existing tests pass. 6 new verification tests added for each fix.
Review-fold from the 3-angle simplify pass:
- sed -Ei / -iE / --in-place now match the shell-critical tier (the
bare '\s-i\b' token missed combined short flags and the GNU long
form); read-only sed stays unflagged. Regression tests added.
- agent_config_contract joins plugin_guard's CODE_EXEMPT_PATTERN_IDS:
content-contract prose in plugin code files (docstrings/comments)
is the same false-positive class the existing agent_config_mod
exemption suppresses. Doc/config files keep the full pattern set.
Efficiency reviewer: 1.24x full-scan cost (+3.4ms/file, install-time
only), worst-case adversarial line 55us — no ReDoS exposure.
Follow-up hardening on top of #92249's tiered scoring:
- Shell-critical tier now also catches tee, and cp/mv with the config
file in destination position (cp/mv reads and .bak backups excluded).
A single '>' redirect must be preceded by a word/quote character so
markdown blockquotes and '->' arrows no longer match.
- Prose tier catches mid-line imperatives behind directive markers
('you must modify...', 'please update...', 'make sure to append...'),
which previously bypassed the line-start anchor.
- Prose instructions aimed at AGENT config files score critical again:
project-skill quarantine acts only on 'dangerous', so high/caution
silently converted 'quarantined' into 'allowed' for exactly the
sentence shape persistence attacks use (concern raised in #88952).
Hermes/other-agent config prose stays high/caution (setup docs
legitimately instruct config.yaml edits).
- New content-contract tier ('AGENTS.md should contain ...') at
high/caution — the shape is shared by authoring guides and attacks.
- .claude/settings and .codex/config gain the same shell-critical tier.
Verified against a 595-skill corpus: 0 skills blocked by these tiers
(main blocked 44 legitimate ones), all mattpocock repro skills from
#92021 install, and 20/20 attack corpus lines keep their verdicts.
Under #92021 the scanner no longer emits critical agent_config_mod /
hermes_config_mod findings for bare mentions - the migration script's
legitimate references now score as informational _ref findings and the
verdict is "safe" with zero modification-intent findings. Tighten the
assertion accordingly (verdict must be safe, not merely non-dangerous)
and swap the known-false-positive set to the new _ref ids.
The skills-guard-v1 scanner flagged ANY mention of AGENTS.md / CLAUDE.md /
.cursorrules / .clinerules as critical/persistence. Any critical finding
forces a dangerous verdict, and community installs cannot be overridden
with --force — so legitimate meta-skills that merely DISCUSS agent config
files (authoring guides, setup docs, cross-references) were permanently
blocked. Three popular community skills were hit in the wild.
skills-guard-v2 scores the persistence category in three tiers:
- Mechanical persistence (shell redirection or sed -i targeting an agent
config file) stays critical -> dangerous. An unambiguous write path.
- Modification language in imperative position (verb at line/bullet start
within 80 chars of the filename) is high -> caution. Regexes cannot
separate "Edit AGENTS.md to inject instructions" from descriptive prose,
but imperative verbs are the shape real instructions take. Caution keeps
the install confirmable instead of irreversibly blocked.
- Bare references drop to low/informational for auditability without
driving the verdict.
The verb-proximity shape matches the existing convention in
tools/threat_patterns.py, and the tiering mirrors how allowed_tools_field
was already handled. The pattern id agent_config_mod is preserved so
plugin_guard.CODE_EXEMPT_PATTERN_IDS stays valid; hermes_config_mod /
other_agent_config get parallel _shell / _ref splits fixing the whole bug
class. SCANNER_VERSION bumps to v2 so cached v1 dangerous verdicts are
invalidated and re-scanned on next install attempt.
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)
* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
Drives the REAL session.resume -> _make_agent -> AIAgent path (eager_build)
through tui_gateway.server.handle_request against a real seeded state.db in
an isolated HERMES_HOME — no mocks. Three scenarios from the live report:
deleted provider falls back to default, renamed provider heals, and a legacy
canonical Bot Chat (no follow_profile_config marker) follows the profile's
current config.
A/B verified: all 3 FAIL on origin/main with the reported
"resume failed: Unknown provider '<name>'"; 3/3 pass with the salvaged
fixes + backfill.
Bot Chats created before the follow_profile_config marker existed carry no
contract in model_config, so they would stay pinned to a stale stored
provider until deleted — the exact shape of the live reports (#89497,
#94818). Mirror the plugin's own identity rule (the profile's session
titled exactly 'Bot Chat') as a legacy fallback in
_stored_session_runtime_overrides, matching the room-plumbing legacy
'Group:' title fallback.
Follow-up to the salvaged #90343 (@curator8888) and #96111 (@lorzl).
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.
That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.
- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
coverage for both session kinds
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.
Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
created by older desktop builds that never sent the marker; hidden
non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
per-member room sessions
Fixes#89497