Follow-ups on top of the salvaged #97712 foundation:
- read_events() pages now include the room's authority stamp
(authority.gateway_id + authority.epoch), so a replicating participant
can persist lineage with every page and a future takeover layer can
fence stale authorities from replayed state alone.
- New regression test proves the replay page bound counts UTF-8 BYTES,
not characters: sabotaging LENGTH(CAST(.. AS BLOB)) back to
LENGTH(TEXT) and dropping the .encode('utf-8') guard fails the test;
the pre-existing multibyte test passed under that sabotage.
- _raise_room_not_found typed NoReturn so narrowing survives closures.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Simplify-code pass findings:
- The docstring claimed 'Matching bare-name suffix' but the fast path
matches the exact directory name (parent.name) — reworded to say
what the code actually does.
- _local_root() swallowed every Exception silently; narrowed to
OSError (what resolve() raises) with a logger.debug breadcrumb so a
recurring resolve failure is diagnosable instead of degrading every
categorized lookup to a silent not-found.
Review feedback (kokhlo): the categorized-name match ran
resolve().relative_to() for every SKILL.md in the walk even when the
bare-name branch already matched — 50+ resolve calls per invocation on
a bare-name lookup in a large profile.
Restructure so the bare directory-name check stays first and the
resolve/relative_to machinery only runs when the lookup name actually
contains a path separator. The skills root is resolved once, lazily,
only when a categorized lookup happens at all. Also compare the
relative path via as_posix() so 'category/skill' lookups work on
Windows, where str(Path) renders backslashes.
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):
1. skill_view(name, file_path='references') returned a raw
'[Errno 21] Is a directory' OS error. The local-skill branch gated
on target_file.exists(); a directory passes exists(), fell through
to read_text(), and raised. The plugin-skill sibling branch already
gated on is_file() — this aligns the local branch so a directory
request gets the same helpful not-found payload with
available_files listing instead of an OS error.
2. skill_manage rejected categorized names ('category/skill-name')
with 'not found in active profile'. _find_skill matched only the
bare directory name, while skill_view's own ambiguity hint tells
the caller to use exactly the categorized form — every call that
followed the hint failed. _find_skill now also matches the full
relative path of the skill dir, giving skill_manage resolution
parity with skill_view across edit/patch/delete/write_file/
remove_file.
Both fixes are covered by regression tests that fail on main.
test_explicit_blank_masks_leaked_cron_env_for_gateway_classification
used platform=api_server as an arbitrary gateway platform; api_server
is now intentionally excluded from gateway approval contexts
(unattended class). Switch to telegram — the test's subject is the
blank-cron-ContextVar masking, not platform policy.
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies. This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.
Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).
Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
Follow-up to imsuperseller's #96740 (cherry-picked as the previous commit).
Widens _CACHE_BUSTING_CONFIG_KEYS with the other construction-baked
compaction-routing settings that had the same stale-cache shape:
compression.in_place, checkpoint_required, micro_compact,
micro_compact_every_n_turns, micro_compact_defrag_threshold_tokens.
Without these, a messaging-gateway session cached before a config edit
keeps the old compaction routing forever.
Not added (reported instead): abort_on_summary_failure, max_attempts,
protect_first_n, codex_gpt55_autoraise_notice, idle_compact_after_seconds
— behavior-tuning rather than routing, left for a deliberate pass.
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
Folded from PR #98345 (@ericmaddox): the one scenario its suite covered
that #91557's did not — an assistant message between the image-only user
message and the checkpoint, asserting post-prune ordering.
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.
Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.
Refs #90976 and #91477.
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
Bundled by #83063 into the always-shipped set, but it serves only
multi-agent kanban campaigns — too niche for every install's skills
index (~20 tok/call for all users). Moved to the optional catalog
(installable via /skills search + hub, official source) and renamed so
the trigger is legible at a glance: 'merge-reconciler' read like a
generic git helper; 'agent-merge-conflict-arbiter' says who it is for.
- skills/autonomous-ai-agents/merge-reconciler -> optional-skills/autonomous-ai-agents/agent-merge-conflict-arbiter
- frontmatter name + description updated (description within the 60-char hardline its own contract test enforces)
- contract test moved/renamed, 9/9 green
- kanban docs (en + zh-Hans) repointed; zero merge-reconciler refs remain
The supporting-files block emitted every file twice per line
(relative -> absolute), duplicating an identical directory prefix
hundreds of times on reference-heavy skills. The absolute base is
already stated once in the [Skill directory: ...] header and the
footer example, so each line now carries only the relative path.
On hermes-agent-dev (462 supporting files) the activation message
drops from 42,145 to 29,042 o200k tokens (-13,103, -31%) — paid on
every session that preloads or invokes the skill.
* feat(compaction): always rebuild the system prompt at the commit boundary — keep-prompt now gated on byte equality of the LIVE builder output; plugin sections re-render with fail-open to last good bytes
* feat(clock): 'Conversation started' resolves through the session-lineage ROOT — a compacted/rotated session keeps its original birth date (Bot Mode forever-chats know when they were first born)
* test: retire old-contract pins — plugin sections re-render at invalidate (freeze stays restore-only), commit boundary always runs the live builder, byte-equal keep preserves object identity
Four gateway test files build module-level fake telegram.ext trees; a fake
missing InlineQueryHandler makes the adapter's lazy-install import fail
mid-suite, leaving ParseMode/ChatType bound to None for every subsequent
telegram test in the same worker (the CI-only MARKDOWN_V2/SUPERGROUP
NoneType cascade). Swept all fakes via grep, not just the one CI named.
The lazy-install placeholder test failing mid-run also left the adapter
module's ParseMode as None, cascading into test_telegram_thread_fallback
failures in the same CI process — all downstream of the same root.
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
Resume was attaching an unused store as {todos: [], revision: 0} and the desktop rejected tool.start updates that have no revision. A merge:true start after reconnect never patched the list until complete.
Skip unused empty snapshots. Apply unversioned updates without moving the watermark so a later todo.updated can still win.
Real-profile browsing (browser.use_real_profile) is meant to drive a COPY of
the user's profile headlessly in the background so they can keep working while
the agent acts on their behalf. Instead, on any host with a display it launched
the user's real browser binary HEADED, popping a window that grabbed focus on
every turn.
Root cause: the native launch (which bypasses agent-browser's own launcher to
avoid --use-mock-keychain dropping keychain-encrypted cookies) only added
--headless=new when Linux had no DISPLAY/WAYLAND_DISPLAY. On a normal desktop
the guard was false, so Chrome opened a visible window.
Fix: launch headless by default on every platform. Chrome's NEW headless mode
shares the profile's normal cookie store (unlike legacy --headless), and the
cookie drop we guard against comes from --use-mock-keychain, not from headless
— so real-profile auth still loads. Users who want to watch can opt in via the
existing browser.headed / AGENT_BROWSER_HEADED toggle (honored for real-profile
now, same as the rest of the browser stack); display-less hosts stay headless
regardless so the launch doesn't die at startup.
Live A/B on a real X seat: old argv mapped a Chrome window (focus steal),
--headless=new mapped zero windows while still exposing a working CDP port.
Updated the stale test that pinned 'never passes --headless' (a legacy-headless
premise) to positively assert the chrome launch is --headless=new by default.
Same bug class as the salvaged terminal-scanner fix: skills_guard's
env_exfil_curl/wget/fetch used unanchored \w*(KEY|TOKEN|...|API)
alternations, so any var with API/KEY/TOKEN mid-name
($TRILLIUM_ETAPI_URL) scored a critical exfiltration finding. Applied
the same \b anchor + plural tolerance, dropped mid-name API (every real
secret it caught already ends in KEY/TOKEN), kept CREDENTIAL, and kept
the loopback exemption from #98246. httpx/requests patterns unchanged —
their (KEY|TOKEN|...) alternation is unanchored-by-design against
argument text, not var-name suffixes.
Anchor env var name matches with \b to avoid matching legitimate
env vars that contain KEY/TOKEN/API as substrings (e.g.,
$TRILLIUM_ETAPI_URL). The patterns now require KEY/TOKEN/SECRET/PASSWORD
to appear at the END of the env var name, reducing false positives on
common API-usage documentation in SOUL.md while still catching actual
exfiltration attempts.
Fixes#63977
Follow-up reconciling the four cherry-picked contributor fixes with the
full-dir GitHubSource.fetch() that landed in #98246:
- Missing SKILL.md-linked support paths now warn and install without the
file at all three sources (GitHub full-dir, GitHub fallback, UrlSource) —
dangling links are prose over-matches or repo-only dev tools, not install
blockers (#66760/#90081). A referenced path present in the tree as a
SYMLINK stays a hard rejection.
- Extension requirement dropped from the glob/placeholder filter:
references/LICENSE is a legitimate support file (82236's tests pin this).
Truncated prose placeholders (references/type-<name>.md -> 'type-') are
still rejected via the trailing-separator shape.
- percent-quoted Contents-API path (82236) merged with revision pinning
(96336) in _fetch_file_bytes.
- Fixture typo fix: four cherry-picked test strings used '\---' where
'\n---' was meant (DeprecationWarning + frontmatter never parsed).
Validation: 167/167 across tests/tools/{skills_hub,skills_guard,
skill_bundle_provenance} + tests/hermes_cli/test_skills_hub.py; live GitHub
fetches (impeccable 163 files rev-pinned; anthropics frontend-design).