Commit Graph

35107 Commits

Author SHA1 Message Date
teknium1 1a2b5a37c1 test(session-search): collapse OR-relaxed tests to three invariants; comment says sort applies
The retry passes ``route`` through unchanged, so the caller's ``sort`` (newest/oldest) still
drives ORDER BY; the comment claimed bm25 ranking unconditionally. Say what the code does
rather than force rank order — a user who asked for newest-first should get newest-first
from the relaxed hits too.

Tests: the seven ``_or_relaxed_query`` helper cases become one parametrized test; one DB
recovery test (exact untouched, paraphrase recovered, all partial rows, role_filter honoured)
and one negative (explicit NOT not relaxed, true miss stays empty, CJK route never reaches
the rewrite). Drops the upstream product name from module prose (credit stays in the PR body).
2026-09-15 04:01:17 -07:00
teknium1 d0bdf39dad test: rewind-exclusion assertion inspects the real results key
The discovery payload exposes hits under `results`, not `matches`, so the
loop over `result_rewind.get("matches", [])` never iterated and the
invariant (the rewound row never surfaces, even via the OR-relaxed retry)
was not being checked. Assert on the anchor message id and snippet of each
returned result instead.
2026-09-15 04:01:17 -07:00
Teknium 2a2c3aa588 feat(session-search): OR-relaxed retry recovers paraphrased recall
Port from nearai/ironclaw#7553 (Filter::FtsRanked): FTS5's implicit AND
between terms means a paraphrased multi-word query misses a stored
sentence that lacks even one of the words. When the exact-match search
and the substring fallbacks all return zero rows, retry the same
unicode61 FTS index with the terms OR-joined, ranked by bm25 so rows
covering more terms surface first.

Strictly additive: gated on a zero-result miss, so successful searches
keep exact-match semantics and ordering. Queries with explicit OR/NOT,
single-term queries, and CJK-routed queries are left untouched. Quoted
phrases relax as whole units.

Adapted for hermes-agent: implemented inside SessionSearchMixin's
zero-result fallback chain (after the CJK-bigram/trigram substring
retries) rather than as a separate filter variant, reusing the already-
built SQL/params so all source/role/sort filters apply to the retry.
2026-09-15 04:01:17 -07:00
teknium1 56e563d3a4 fix: keep reading scripts named inside masked heredoc bodies
Masking inert heredoc bodies before the referenced-script walk (previous
commit) also hid every path the body names, so a Python body that does
os.system('/x/restart.sh'), a bare '/x/restart.sh' line, or a nested cat
heredoc naming the script all went undetected — main caught them because
the walk read the script and found the lifecycle command inside.

The interpreter does execute those paths at runtime, so the walk must still
read them. Only the fail-closed verdicts (cloud placeholder, oversized or
binary file) stay restricted to the masked view: a >1 MiB data file merely
mentioned inside an inert body is not a script, which is the false positive
the previous commit fixed and its tests keep pinning.

Review finding: masking the heredoc body from the walk widened fail-open —
scripts executed by path from within the body were no longer scanned.
2026-09-15 04:00:34 -07:00
teknium1 1fd0ef7520 test(cron): trim the heredoc-walk regression to its two invariants
Keep the inert-body path (false) and the unquoted-body path (still true);
the `sh -c`-in-`cat` heredoc and plain script-reference cases are already
pinned by the existing lifecycle-guard suites. Comment shortened to the WHY.
2026-09-15 04:00:34 -07:00
Kevin Rajan 59a1403fa2 fix(cron): mask inert heredoc bodies before the referenced-script walk
The direct lifecycle scan masks provably-inert heredoc bodies
(strip_inert_heredoc_bodies), but the referenced-script and -c payload
walks ran on the unmasked command. A path inside such a body is never
shell-executed, so walking it was a pure false positive: a >1 MiB path
mentioned in a quoted heredoc body (e.g. python3 - <<'PY') failed closed
and hard-blocked an innocent command.

The walks now use the same masked view as the direct scan. Masking only
fires under the stripper's conservative contract (quoted delimiters,
exact terminator, single simple command, allowlisted consumer, no command
substitution); unquoted/shell-consumed/ambiguous bodies stay visible and
fail-closed.

Fixes #110422.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
2026-09-15 04:00:34 -07:00
teknium1 928e5993ee test(tui_gateway): drop removed GoalGate.last_failed_fingerprint from control-card fixture
The fix removed the fingerprint field from GoalGate (failed gates re-run every
boundary), so the structured-read fixture can no longer construct one with it.
No other reader/writer of last_failed_fingerprint/workspace_fingerprint remains
in tests/, tui_gateway/, apps/desktop or web/.
2026-09-15 03:59:50 -07:00
teknium1 bb745a0e9b fix(goals): failed quality gates re-run every boundary instead of replaying a status fingerprint
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
2026-09-15 03:59:50 -07:00
teknium1 b91ee8c72d docs: remove a stray conflict marker from the optional-skills catalog
47c029927f landed with a leftover ">>>>>>>" line and put the dream-loop
and mono-color rows under autonomous-ai-agents. Move both rows into the
creative table (alphabetical) and drop the marker.
2026-09-15 03:59:20 -07:00
teknium1 b55767be82 fix: judge wait_on_pid race no longer raises out of evaluate_after_turn
`_apply_wait_directive` probed `_pid_alive` and then called `wait_on`,
which re-checks liveness and raises ValueError. A pid exiting between
the two checks surfaced as an exception from `evaluate_after_turn`; all
three callers swallow it, so the goal silently lost that boundary's
continuation. Do a single check by catching wait_on's own ValueError
and falling through to the continue decision.

Review finding: TOCTOU between _pid_alive pre-check and wait_on re-check drops the continuation.
2026-09-15 03:59:10 -07:00
teknium1 ee07fcd4d7 fix(goals): a judge wait_on_pid naming an unobservable pid continues instead of parking
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
2026-09-15 03:59:10 -07:00
KoNit-K c70db196eb fix(goals): reject dead wait-on PIDs 2026-09-15 03:59:10 -07:00
teknium1 8d66d2dd30 test: pass encoding to read_text in the import-agent tests (windows-footguns gate) 2026-09-15 03:58:44 -07:00
teknium1 9f30a0a269 fix(import-sync): never clobber a locally edited imported skill; keep digest on errors
The docs promised "skills you created or modified under the import category
yourself are never clobbered", but sync replaced any destination whose name was
in `imported_skills`, regardless of what was there now — an imported skill the
user had since edited was silently overwritten on the next source change.

The manifest now stores `imported_skills` as {name: digest-of-the-copy-we-wrote}
and `--sync` refreshes a destination only while it still matches that digest;
a locally modified copy records a `conflict` ("modified locally — not
refreshed") and is skipped. Pre-digest manifests (a plain list) keep the old
trusted behaviour for one more cycle and are upgraded on the next import.

`sync_imported_agents` also refreshed the source digest after a run with
errors, so the failed items were never retried; the previous digest is kept
whenever the report has errors.
2026-09-15 03:58:44 -07:00
teknium1 57acaa50a2 test: fold 13 sync tests into 3 invariant tests
Salvage bar is <=2 invariant tests per behaviour, not one test per branch.
The three kept each pin a contract: import registers the source and an
unchanged (credential-only) change is a byte-identical no-op; a changed
source re-imports memory + Hermes-owned skills while a user-created skill
is never clobbered; --sync --dry-run writes nothing and leaves the digest.
Corrupt-manifest recovery, vanished-source skip and parser wiring are
exercised implicitly by the E2E path and don't merit their own tests.
2026-09-15 03:58:44 -07:00
Teknium 6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
teknium1 c164e12bd8 docs(cron): skill-backed jobs receive the skill config block 2026-09-15 03:58:27 -07:00
Hukla 43891e9f8e fix(cron): inject skill config into scheduled runs 2026-09-15 03:58:27 -07:00
teknium1 681a28b06a test(session_search): fold the temporal-narrowing tests to four invariants
Parse-unit cases (relative units, ISO, empty, invalid) become one parametrized test; one
end-to-end per parameter: ISO window enforced in SQL (survives a truncated FTS scan),
relative after/before + exclude_session_ids driven through INLINE_TOOL_EXECUTORS so the
production dispatch is what is tested (red on the executor before this fix), and lineage
exclusion. Drops the schema-membership and unbounded-equals-None change-detectors.
2026-09-15 03:58:05 -07:00
teknium1 18da11720d fix(session_search): forward after/before/exclude_session_ids through the inline executor
INLINE_TOOL_EXECUTORS["session_search"] is the production dispatch on every surface
(tool_executor + agent_runtime_helpers) and maps schema args to kwargs explicitly, so the
schema advertised after/before/exclude_session_ids while the executor silently dropped them
and every time-bounded call ran unbounded. Add the three mappings.

Also drop the stranded `as_exclusive_end` parameter on _parse_iso_bound (accepted, never
read; exclusivity is the SQL `<` predicate) and the Python re-check of the time window over
FTS rows in _discover — _search_filter_clauses already bounds sessions.started_at on every
route, so the only place the window still needs a Python check is the title-match branch,
which bypasses that query.
2026-09-15 03:58:05 -07:00
Teknium 324da01573 chore: map contributor email for Moodtuner997 (PR #86067 salvage) 2026-09-15 03:58:05 -07:00
Teknium e819846b10 Inspired by Amp: relative time bounds (7d/24h/2w) + wrapper forwarding for session_search after/before
Amp's thread feed supports relative time filters (`after:7d`,
`updated_before:7d`) alongside ISO dates. Extend the salvaged
after/before bounds (PR #86067 by @Moodtuner997) the same way:

- `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw`
  (case-insensitive) meaning "now minus N", alongside ISO
  dates/datetimes. Clearer error message names both accepted forms.
- Forward after/before/exclude_session_ids through the public
  `session_search()` wrapper (the PR predates the wrapper/impl split;
  without this the SQL bounds were unreachable from the registry
  handler — same class as the earlier `detail` forwarding fix).
  Appended after `detail` to preserve positional compatibility.
- Tool schema descriptions teach both forms.
- Tests: relative after/before against the discovery shape, unit
  checks for h/d/w math, case-insensitivity, and bad-unit rejection.
- Docs: tools-reference row mentions time bounds + exclude_session_ids.
2026-09-15 03:58:05 -07:00
Théo 6f793ddbdc fix(session_search): apply after/before in SQL WHERE
Push the session-start bounds into search_messages so FTS LIMIT
cannot be filled by out-of-range hits. Covers FTS5, CJK, trigram,
LIKE fallback, and the unindexed-gap supplement.

Refs #86021.
2026-09-15 03:58:05 -07:00
Théo 5655920f9a feat(session_search): add after/before bounds and exclude_session_ids
Discovery-only filters for issue #86021. sort remains a ranking
bias. Date-only before is an exclusive midnight UTC bound.
exclude_session_ids drops the named session and its lineage (cap 20).
2026-09-15 03:58:05 -07:00
teknium1 dddefaefae fix: paged, extracted and post-compaction reads count as a write_file baseline
The stale-overwrite refusal made write_file permanently unusable for any
existing file it could not show in one read_file page: every >2000-line
(or >100K-char) page was recorded as partial, no full baseline ever
existed, and the refusal told the model to "re-read the whole file", which
the tool cannot do. Track the line ranges each task pages through per path
at one mtime; contiguous pages from line 1 to total_lines are a full read
(a new mtime between pages restarts the coverage). The same gap hit two
siblings: the extracted-document branch (.ipynb, text-authorable) returned
before any read bookkeeping, so an existing notebook could never be
overwritten; and reset_file_dedup dropped every baseline on compaction
while keeping read_timestamps, so every write after compaction was refused
even for files unchanged on disk. Baselines now survive compaction exactly
like the dedup mtime map does — only while the recorded mtime still matches.

Refusal texts no longer embed the pre-PR "Warning: … Consider re-reading"
copy inside "Refusing to overwrite", and every refusal names a recovery the
model can perform: read the remaining pages, or use patch.
2026-09-15 03:57:24 -07:00
teknium1 4613f895ab test: give the rewrite-hint fixtures a read baseline; align docs row with schema
The stale-write guard now refuses write_file on an existing file the task
never read in full, so test_write_file_rewrite_hint's overwrite-without-read
fixtures were refused before the hint could be computed. Reading first is
the exact read->whole-file-rewrite pattern the hint exists for.

tools-reference.md's write_file row now mirrors the WRITE_FILE_SCHEMA
description (one-sentence contract + the recovery step) instead of a
longer paraphrase.
2026-09-15 03:57:24 -07:00
Teknium 6569651b87 fix: harden salvage of #65605 — redaction-gated test, sibling test baseline, docs
- test_file_staleness redacted-read case now force-enables redaction
  (matches tests/agent/test_redact.py convention) so it exercises the
  sentinel path in hermetic CI where security.redact_secrets is unset.
- test_write_verification CRLF case establishes a read baseline first
  (the new guard refuses unread existing-file overwrites by design).
- tools-reference.md documents the read-before-overwrite contract.
- contributors/emails mapping for DanSpicyTaco.
2026-09-15 03:57:24 -07:00
DanSpicyTaco a8af57fcd2 fix: block stale write_file overwrites
Require an explicit full-file baseline before replacing existing host-visible files with write_file, and fail closed when that baseline is stale. This prevents stale conversation context from clobbering manual or external edits.\n\nRefs #65604
2026-09-15 03:57:24 -07:00
Teknium 166dc1290f Merge pull request #111601 from NousResearch/pr/ux-messages-desktop-tui
Desktop and TUI errors say what happened and offer one-click fixes (message audit, GUI)
2026-09-15 03:57:13 -07:00
teknium1 04fcf9159c fix: keep api_server approval bridge and cron self-scheduling after presence strip
Widening the _presence() clearing from single-query to every unattended
context also cleared is_ask for platform=api_server. That surface answers
approvals through the /v1/runs bridge (approval.request ->
POST /v1/runs/{id}/approval), so a dangerous command that used to park in
waiting_for_approval became an instant BLOCK with no approval.request.
Restrict the clearing to single-query + cron, where nobody can answer.

Stripping HERMES_INTERACTIVE/HERMES_GATEWAY_SESSION/HERMES_EXEC_ASK from
the external worker env also made check_cronjob_requirements() False, so
the cronjob toolset vanished for every job on a managed-systemd gateway
even with cron.allow_agent_scheduling: true. Accept the existing
HERMES_CRON_SESSION marker (set by run_one_job's context) as well.

Review finding: _presence() over-widening broke the /v1/runs approval bridge; env strip hid the cronjob toolset in external workers.
2026-09-15 03:56:17 -07:00
teknium1 b6b7802447 fix(approval): every unattended context clears leaked presence vars
Widen the cron-only clearing to `_unattended_contexts()`: a webhook /
api_server session running inside a gateway inherits HERMES_EXEC_ASK=1
exactly like an external cron worker does, and `_presence()` returning
is_ask=True sent it to the gateway-decision branch with no notifier — a
pending card nobody can answer — instead of `approvals.unattended_mode`.
Same class as #110932, one predicate.

Test trimmed to two invariants (cron / webhook leak → cleared; interactive
keeps presence); the launch-path comment in cron/scheduler.py names the
env-fallback consumers instead of an internal incident log.
2026-09-15 03:56:17 -07:00
fabiantax 2db47cc9fb fix(cron): strip interactive presence vars from external worker env
Gateway sets HERMES_EXEC_ASK=1 (interactive launches set HERMES_INTERACTIVE / HERMES_GATEWAY_SESSION) at runtime; systemd-run cron workers inherited them, _is_interactive_cli() then bypassed approvals.cron_mode for terminal and every run hung on a pending card nobody could answer (fab-swarm #105: ms197 lane left 6 claims stranded, 10-30s hangs). Local mitigation; upstream report to follow.
2026-09-15 03:56:17 -07:00
fabiantax 2a630671d7 fix(approval): cron context is never interactive, even with leaked presence env
_presence() cleared is_cli/is_gateway/is_ask for single-query sessions but
not for cron, so a cron worker that inherited HERMES_INTERACTIVE /
HERMES_EXEC_ASK from its launching gateway resolved as an interactive CLI
and blocked on an approval card nobody could answer (measured: 31
pending_approval hangs/hour, 6 stranded claims — #110932).

Mirror the single-query clearing for _is_cron_approval_context(), matching
the cron exclusion already inside _is_gateway_approval_context(). Layer 1
(#110942) strips the vars at the launch path; this makes the gate robust
to any other leak route.
2026-09-15 03:56:17 -07:00
Teknium 225d53f953 Port from cline/cline#12876: classify Anthropic output-cap errors 2026-09-15 03:54:58 -07:00
teknium1 123db98635 docs(gemini): scope the base-URL normalization claim to the Google host and TTS
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
2026-09-15 03:54:01 -07:00
Teknium f030c03970 Port from cline/cline#13329: normalize host-root Gemini base URLs to /v1beta
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.

normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).

Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
2026-09-15 03:54:01 -07:00
teknium1 aa75d3724f fix(stream-json): verbatim text deltas, closed protocol on init failure, per-call tool keys
- on_text_delta dropped whitespace-only deltas, so concatenating the `text`
  events no longer reproduced the answer (a newline between paragraphs was
  lost). Only None/"" (the turn-end sentinel) is skipped now.
- The emitter was attached only after credentials + agent init succeeded, so a
  missing key or unknown provider exited 1 with an EMPTY stdout and the
  provider error rendered through ChatConsole (stdout). The emitter is now
  built before _ensure_runtime_credentials/_init_agent; that path closes the
  protocol with init + a failed `result` (exit_code 1, error) and the
  credential error goes to stderr whenever stdout is machine-readable
  (tool_progress_mode == "off", i.e. -Q and stream-json).
- _tool_started was keyed by tool name, so concurrent same-name calls
  clobbered each other's start time; key on tool_call_id when the caller
  passes one and surface it on tool_use/tool_result.

Live: `hermes chat -q … --format stream-json` with no provider and with a dead
custom base_url both yield pure JSONL (`system` + `result`, exit 1).
2026-09-15 03:53:13 -07:00
Teknium ee0666bf0e chore: map contributor email for attribution audit 2026-09-15 03:53:13 -07:00
Alan 1657a1ce2d feat(cli): add --format stream-json for structured JSONL output
Adds a --format flag to hermes chat single-query mode. stream-json
emits newline-delimited JSON events (init, text, tool_use, tool_result,
result envelope with token stats + exit code) to stdout for CI
pipelines and external tooling. Session ID stays on stderr.

Salvaged from PR #12278 by @ProDrifterDK onto current main, including
the follow-up commit enforcing the single-query contract (implies
quiet, rejects --tui, emits a final result record with exit code 130
on interrupt).
2026-09-15 03:53:13 -07:00
teknium1 39fbfbe4e6 test: drop prose change-detector for terminal description
The test asserted three verbatim substrings of TERMINAL_TOOL_DESCRIPTION,
so any future rewording of the guidance would fail it without a behavior
change. The PR's change is prose-only guidance; the description text is
not a stable interface worth pinning.
2026-09-15 03:52:24 -07:00
teknium1 8d8f68a0c6 docs: drop committed PR infographic
The SVG was a stand-in for the PR infographic attachment; repo docs/ is not
the place for per-PR artwork, and main has no docs/pr-infographics/ directory.
2026-09-15 03:52:24 -07:00
Teknium 325d73ce07 docs(tools): clarify terminal background waits
Port from Kilo-Org/kilocode#13224: fixed waits belong in foreground terminal calls, while background mode is reserved for independently running processes.
2026-09-15 03:52:24 -07:00
teknium1 cb84e7d94e fix(codex): import kill_process_tree so the SIGTERM-timeout path actually kills
close() escalated to kill_process_tree() but never imported it; the NameError
was swallowed by contextlib.suppress, so on the TimeoutExpired path neither the
kill nor the post-kill wait ran and a root codex ignoring SIGTERM leaked (a
regression vs the previous self._proc.kill()). Import it from agent.deadline
and add a test forcing the timeout path that asserts the tree kill and the
follow-up wait both run. Also drop the upstream product reference from the
test docstring (credit stays in the PR body) and pass encoding= to the PID
file reads flagged by the Windows footgun scanner.
2026-09-15 03:51:44 -07:00
Teknium a180274a55 fix(codex): reap app-server descendant processes
Port from openclaw/openclaw#126285: snapshot Codex app-server descendants before root retirement and sweep the proven process identities after close so independently grouped stdio MCP children cannot survive client shutdown.
2026-09-15 03:51:44 -07:00
teknium1 616a3ce036 fix(kanban): actor sentinel for synthesized runs; two invariant tests
A plain ``profile=None`` default could not tell "caller named the actor"
from "read the card" — an unassigned card is a legitimate None actor, and
the row re-read would silently kick back in for it. Use a module sentinel
so only callers that did not pass an actor fall back to the card row.

Trims the salvaged test file to two invariants in the existing review
lifecycle suite: the never-claimed handoff names the implementer (red on
main) and an unassigned card's synthesized run keeps ``profile=NULL``.
2026-09-15 03:51:28 -07:00
Kevin Rajan ff8f725c5a fix(kanban): attribute synthesized review-handoff run to the implementer
request_review captures the implementer before rewriting tasks.assignee to
the reviewer, but _synthesize_ended_run re-reads assignee off the mutated
row, so the zero-duration review_requested run names the reviewer instead
of the handoff's actor. Pass the captured implementer through a new
optional profile keyword on _end_or_synthesize_run/_synthesize_ended_run.

Fixes #111064
2026-09-15 03:51:28 -07:00
teknium1 3cad439de7 docs(sessions): describe the timings block in JSONL exports
User-visible export shape changed with no docs hunk. One paragraph in the JSONL section: what
the block holds (ids/roles/counts/durations, text-free), why complete is always false,
available=false when no message carries a timestamp, and that import ignores it.
2026-09-15 03:51:07 -07:00
teknium1 1685ffcb15 test(session-export): one lineage+import invariant replaces the duplicate export_all case
The export_all assertion already lives in tests/hermes_cli/test_session_export_batch.py. Replace
it with the lineage test (timings over merged messages, red on the previous head) which also
proves the derived block is excluded from the import size budget.
2026-09-15 03:51:07 -07:00
teknium1 d566442b56 fix(session-export): lineage timings span the merged messages; import ignores derived size
export_session_lineage spread segments[-1] over the merged dict, so the top-level `timings`
described only the last compression segment while `messages` spanned the whole lineage — a
reader would see a 500 ms wall clock over a lineage that ran for hours. Compute the block over
the merged message list (segments keep their own).

_validate_import_session measured the raw session JSON, so the derived `timings.intervals`
(one entry per message pair) counted toward the 5 MiB per-session limit and could reject a
long lineage whose actual content fits. Strip `timings` before measuring; it is rebuilt from
the messages on the next export anyway.
2026-09-15 03:51:07 -07:00
teknium1 a4b3ba5e40 test(export): batching contract compares rows without the derived timings block
export_all rows now carry a computed `timings` key; the batching test built
its expectation from search_sessions + get_messages, so equality failed on
the extra key (CI red). Strip it for the row comparison and assert it is
present on every row.
2026-09-15 03:51:07 -07:00