The retry passes ``route`` through unchanged, so the caller's ``sort`` (newest/oldest) still
drives ORDER BY; the comment claimed bm25 ranking unconditionally. Say what the code does
rather than force rank order — a user who asked for newest-first should get newest-first
from the relaxed hits too.
Tests: the seven ``_or_relaxed_query`` helper cases become one parametrized test; one DB
recovery test (exact untouched, paraphrase recovered, all partial rows, role_filter honoured)
and one negative (explicit NOT not relaxed, true miss stays empty, CJK route never reaches
the rewrite). Drops the upstream product name from module prose (credit stays in the PR body).
The discovery payload exposes hits under `results`, not `matches`, so the
loop over `result_rewind.get("matches", [])` never iterated and the
invariant (the rewound row never surfaces, even via the OR-relaxed retry)
was not being checked. Assert on the anchor message id and snippet of each
returned result instead.
Port from nearai/ironclaw#7553 (Filter::FtsRanked): FTS5's implicit AND
between terms means a paraphrased multi-word query misses a stored
sentence that lacks even one of the words. When the exact-match search
and the substring fallbacks all return zero rows, retry the same
unicode61 FTS index with the terms OR-joined, ranked by bm25 so rows
covering more terms surface first.
Strictly additive: gated on a zero-result miss, so successful searches
keep exact-match semantics and ordering. Queries with explicit OR/NOT,
single-term queries, and CJK-routed queries are left untouched. Quoted
phrases relax as whole units.
Adapted for hermes-agent: implemented inside SessionSearchMixin's
zero-result fallback chain (after the CJK-bigram/trigram substring
retries) rather than as a separate filter variant, reusing the already-
built SQL/params so all source/role/sort filters apply to the retry.
Masking inert heredoc bodies before the referenced-script walk (previous
commit) also hid every path the body names, so a Python body that does
os.system('/x/restart.sh'), a bare '/x/restart.sh' line, or a nested cat
heredoc naming the script all went undetected — main caught them because
the walk read the script and found the lifecycle command inside.
The interpreter does execute those paths at runtime, so the walk must still
read them. Only the fail-closed verdicts (cloud placeholder, oversized or
binary file) stay restricted to the masked view: a >1 MiB data file merely
mentioned inside an inert body is not a script, which is the false positive
the previous commit fixed and its tests keep pinning.
Review finding: masking the heredoc body from the walk widened fail-open —
scripts executed by path from within the body were no longer scanned.
Keep the inert-body path (false) and the unquoted-body path (still true);
the `sh -c`-in-`cat` heredoc and plain script-reference cases are already
pinned by the existing lifecycle-guard suites. Comment shortened to the WHY.
The direct lifecycle scan masks provably-inert heredoc bodies
(strip_inert_heredoc_bodies), but the referenced-script and -c payload
walks ran on the unmasked command. A path inside such a body is never
shell-executed, so walking it was a pure false positive: a >1 MiB path
mentioned in a quoted heredoc body (e.g. python3 - <<'PY') failed closed
and hard-blocked an innocent command.
The walks now use the same masked view as the direct scan. Masking only
fires under the stripper's conservative contract (quoted delimiters,
exact terminator, single simple command, allowlisted consumer, no command
substitution); unquoted/shell-consumed/ambiguous bodies stay visible and
fail-closed.
Fixes#110422.
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
The fix removed the fingerprint field from GoalGate (failed gates re-run every
boundary), so the structured-read fixture can no longer construct one with it.
No other reader/writer of last_failed_fingerprint/workspace_fingerprint remains
in tests/, tui_gateway/, apps/desktop or web/.
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
47c029927f landed with a leftover ">>>>>>>" line and put the dream-loop
and mono-color rows under autonomous-ai-agents. Move both rows into the
creative table (alphabetical) and drop the marker.
`_apply_wait_directive` probed `_pid_alive` and then called `wait_on`,
which re-checks liveness and raises ValueError. A pid exiting between
the two checks surfaced as an exception from `evaluate_after_turn`; all
three callers swallow it, so the goal silently lost that boundary's
continuation. Do a single check by catching wait_on's own ValueError
and falling through to the continue decision.
Review finding: TOCTOU between _pid_alive pre-check and wait_on re-check drops the continuation.
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
The docs promised "skills you created or modified under the import category
yourself are never clobbered", but sync replaced any destination whose name was
in `imported_skills`, regardless of what was there now — an imported skill the
user had since edited was silently overwritten on the next source change.
The manifest now stores `imported_skills` as {name: digest-of-the-copy-we-wrote}
and `--sync` refreshes a destination only while it still matches that digest;
a locally modified copy records a `conflict` ("modified locally — not
refreshed") and is skipped. Pre-digest manifests (a plain list) keep the old
trusted behaviour for one more cycle and are upgraded on the next import.
`sync_imported_agents` also refreshed the source digest after a run with
errors, so the failed items were never retried; the previous digest is kept
whenever the report has errors.
Salvage bar is <=2 invariant tests per behaviour, not one test per branch.
The three kept each pin a contract: import registers the source and an
unchanged (credential-only) change is a byte-identical no-op; a changed
source re-imports memory + Hermes-owned skills while a user-created skill
is never clobbered; --sync --dry-run writes nothing and leaves the digest.
Corrupt-manifest recovery, vanished-source skip and parser wiring are
exercised implicitly by the E2E path and don't merit their own tests.
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:
- Every successful import registers its source + a content digest of
everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
files changed since the last run (digest compare; unchanged = no-op).
Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
sync; user-created skills under the import category keep conflict
semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
trigger (or leak into) a sync.
Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
Parse-unit cases (relative units, ISO, empty, invalid) become one parametrized test; one
end-to-end per parameter: ISO window enforced in SQL (survives a truncated FTS scan),
relative after/before + exclude_session_ids driven through INLINE_TOOL_EXECUTORS so the
production dispatch is what is tested (red on the executor before this fix), and lineage
exclusion. Drops the schema-membership and unbounded-equals-None change-detectors.
INLINE_TOOL_EXECUTORS["session_search"] is the production dispatch on every surface
(tool_executor + agent_runtime_helpers) and maps schema args to kwargs explicitly, so the
schema advertised after/before/exclude_session_ids while the executor silently dropped them
and every time-bounded call ran unbounded. Add the three mappings.
Also drop the stranded `as_exclusive_end` parameter on _parse_iso_bound (accepted, never
read; exclusivity is the SQL `<` predicate) and the Python re-check of the time window over
FTS rows in _discover — _search_filter_clauses already bounds sessions.started_at on every
route, so the only place the window still needs a Python check is the title-match branch,
which bypasses that query.
Amp's thread feed supports relative time filters (`after:7d`,
`updated_before:7d`) alongside ISO dates. Extend the salvaged
after/before bounds (PR #86067 by @Moodtuner997) the same way:
- `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw`
(case-insensitive) meaning "now minus N", alongside ISO
dates/datetimes. Clearer error message names both accepted forms.
- Forward after/before/exclude_session_ids through the public
`session_search()` wrapper (the PR predates the wrapper/impl split;
without this the SQL bounds were unreachable from the registry
handler — same class as the earlier `detail` forwarding fix).
Appended after `detail` to preserve positional compatibility.
- Tool schema descriptions teach both forms.
- Tests: relative after/before against the discovery shape, unit
checks for h/d/w math, case-insensitivity, and bad-unit rejection.
- Docs: tools-reference row mentions time bounds + exclude_session_ids.
Push the session-start bounds into search_messages so FTS LIMIT
cannot be filled by out-of-range hits. Covers FTS5, CJK, trigram,
LIKE fallback, and the unindexed-gap supplement.
Refs #86021.
Discovery-only filters for issue #86021. sort remains a ranking
bias. Date-only before is an exclusive midnight UTC bound.
exclude_session_ids drops the named session and its lineage (cap 20).
The stale-overwrite refusal made write_file permanently unusable for any
existing file it could not show in one read_file page: every >2000-line
(or >100K-char) page was recorded as partial, no full baseline ever
existed, and the refusal told the model to "re-read the whole file", which
the tool cannot do. Track the line ranges each task pages through per path
at one mtime; contiguous pages from line 1 to total_lines are a full read
(a new mtime between pages restarts the coverage). The same gap hit two
siblings: the extracted-document branch (.ipynb, text-authorable) returned
before any read bookkeeping, so an existing notebook could never be
overwritten; and reset_file_dedup dropped every baseline on compaction
while keeping read_timestamps, so every write after compaction was refused
even for files unchanged on disk. Baselines now survive compaction exactly
like the dedup mtime map does — only while the recorded mtime still matches.
Refusal texts no longer embed the pre-PR "Warning: … Consider re-reading"
copy inside "Refusing to overwrite", and every refusal names a recovery the
model can perform: read the remaining pages, or use patch.
The stale-write guard now refuses write_file on an existing file the task
never read in full, so test_write_file_rewrite_hint's overwrite-without-read
fixtures were refused before the hint could be computed. Reading first is
the exact read->whole-file-rewrite pattern the hint exists for.
tools-reference.md's write_file row now mirrors the WRITE_FILE_SCHEMA
description (one-sentence contract + the recovery step) instead of a
longer paraphrase.
- test_file_staleness redacted-read case now force-enables redaction
(matches tests/agent/test_redact.py convention) so it exercises the
sentinel path in hermetic CI where security.redact_secrets is unset.
- test_write_verification CRLF case establishes a read baseline first
(the new guard refuses unread existing-file overwrites by design).
- tools-reference.md documents the read-before-overwrite contract.
- contributors/emails mapping for DanSpicyTaco.
Require an explicit full-file baseline before replacing existing host-visible files with write_file, and fail closed when that baseline is stale. This prevents stale conversation context from clobbering manual or external edits.\n\nRefs #65604
Widening the _presence() clearing from single-query to every unattended
context also cleared is_ask for platform=api_server. That surface answers
approvals through the /v1/runs bridge (approval.request ->
POST /v1/runs/{id}/approval), so a dangerous command that used to park in
waiting_for_approval became an instant BLOCK with no approval.request.
Restrict the clearing to single-query + cron, where nobody can answer.
Stripping HERMES_INTERACTIVE/HERMES_GATEWAY_SESSION/HERMES_EXEC_ASK from
the external worker env also made check_cronjob_requirements() False, so
the cronjob toolset vanished for every job on a managed-systemd gateway
even with cron.allow_agent_scheduling: true. Accept the existing
HERMES_CRON_SESSION marker (set by run_one_job's context) as well.
Review finding: _presence() over-widening broke the /v1/runs approval bridge; env strip hid the cronjob toolset in external workers.
Widen the cron-only clearing to `_unattended_contexts()`: a webhook /
api_server session running inside a gateway inherits HERMES_EXEC_ASK=1
exactly like an external cron worker does, and `_presence()` returning
is_ask=True sent it to the gateway-decision branch with no notifier — a
pending card nobody can answer — instead of `approvals.unattended_mode`.
Same class as #110932, one predicate.
Test trimmed to two invariants (cron / webhook leak → cleared; interactive
keeps presence); the launch-path comment in cron/scheduler.py names the
env-fallback consumers instead of an internal incident log.
Gateway sets HERMES_EXEC_ASK=1 (interactive launches set HERMES_INTERACTIVE / HERMES_GATEWAY_SESSION) at runtime; systemd-run cron workers inherited them, _is_interactive_cli() then bypassed approvals.cron_mode for terminal and every run hung on a pending card nobody could answer (fab-swarm #105: ms197 lane left 6 claims stranded, 10-30s hangs). Local mitigation; upstream report to follow.
_presence() cleared is_cli/is_gateway/is_ask for single-query sessions but
not for cron, so a cron worker that inherited HERMES_INTERACTIVE /
HERMES_EXEC_ASK from its launching gateway resolved as an interactive CLI
and blocked on an approval card nobody could answer (measured: 31
pending_approval hangs/hour, 6 stranded claims — #110932).
Mirror the single-query clearing for _is_cron_approval_context(), matching
the cron exclusion already inside _is_gateway_approval_context(). Layer 1
(#110942) strips the vars at the launch path; this makes the gate robust
to any other leak route.
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.
normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).
Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
- on_text_delta dropped whitespace-only deltas, so concatenating the `text`
events no longer reproduced the answer (a newline between paragraphs was
lost). Only None/"" (the turn-end sentinel) is skipped now.
- The emitter was attached only after credentials + agent init succeeded, so a
missing key or unknown provider exited 1 with an EMPTY stdout and the
provider error rendered through ChatConsole (stdout). The emitter is now
built before _ensure_runtime_credentials/_init_agent; that path closes the
protocol with init + a failed `result` (exit_code 1, error) and the
credential error goes to stderr whenever stdout is machine-readable
(tool_progress_mode == "off", i.e. -Q and stream-json).
- _tool_started was keyed by tool name, so concurrent same-name calls
clobbered each other's start time; key on tool_call_id when the caller
passes one and surface it on tool_use/tool_result.
Live: `hermes chat -q … --format stream-json` with no provider and with a dead
custom base_url both yield pure JSONL (`system` + `result`, exit 1).
Adds a --format flag to hermes chat single-query mode. stream-json
emits newline-delimited JSON events (init, text, tool_use, tool_result,
result envelope with token stats + exit code) to stdout for CI
pipelines and external tooling. Session ID stays on stderr.
Salvaged from PR #12278 by @ProDrifterDK onto current main, including
the follow-up commit enforcing the single-query contract (implies
quiet, rejects --tui, emits a final result record with exit code 130
on interrupt).
The test asserted three verbatim substrings of TERMINAL_TOOL_DESCRIPTION,
so any future rewording of the guidance would fail it without a behavior
change. The PR's change is prose-only guidance; the description text is
not a stable interface worth pinning.
The SVG was a stand-in for the PR infographic attachment; repo docs/ is not
the place for per-PR artwork, and main has no docs/pr-infographics/ directory.
Port from Kilo-Org/kilocode#13224: fixed waits belong in foreground terminal calls, while background mode is reserved for independently running processes.
close() escalated to kill_process_tree() but never imported it; the NameError
was swallowed by contextlib.suppress, so on the TimeoutExpired path neither the
kill nor the post-kill wait ran and a root codex ignoring SIGTERM leaked (a
regression vs the previous self._proc.kill()). Import it from agent.deadline
and add a test forcing the timeout path that asserts the tree kill and the
follow-up wait both run. Also drop the upstream product reference from the
test docstring (credit stays in the PR body) and pass encoding= to the PID
file reads flagged by the Windows footgun scanner.
Port from openclaw/openclaw#126285: snapshot Codex app-server descendants before root retirement and sweep the proven process identities after close so independently grouped stdio MCP children cannot survive client shutdown.
A plain ``profile=None`` default could not tell "caller named the actor"
from "read the card" — an unassigned card is a legitimate None actor, and
the row re-read would silently kick back in for it. Use a module sentinel
so only callers that did not pass an actor fall back to the card row.
Trims the salvaged test file to two invariants in the existing review
lifecycle suite: the never-claimed handoff names the implementer (red on
main) and an unassigned card's synthesized run keeps ``profile=NULL``.
request_review captures the implementer before rewriting tasks.assignee to
the reviewer, but _synthesize_ended_run re-reads assignee off the mutated
row, so the zero-duration review_requested run names the reviewer instead
of the handoff's actor. Pass the captured implementer through a new
optional profile keyword on _end_or_synthesize_run/_synthesize_ended_run.
Fixes#111064
User-visible export shape changed with no docs hunk. One paragraph in the JSONL section: what
the block holds (ids/roles/counts/durations, text-free), why complete is always false,
available=false when no message carries a timestamp, and that import ignores it.
The export_all assertion already lives in tests/hermes_cli/test_session_export_batch.py. Replace
it with the lineage test (timings over merged messages, red on the previous head) which also
proves the derived block is excluded from the import size budget.
export_session_lineage spread segments[-1] over the merged dict, so the top-level `timings`
described only the last compression segment while `messages` spanned the whole lineage — a
reader would see a 500 ms wall clock over a lineage that ran for hours. Compute the block over
the merged message list (segments keep their own).
_validate_import_session measured the raw session JSON, so the derived `timings.intervals`
(one entry per message pair) counted toward the 5 MiB per-session limit and could reject a
long lineage whose actual content fits. Strip `timings` before measuring; it is rebuilt from
the messages on the next export anyway.
export_all rows now carry a computed `timings` key; the batching test built
its expectation from search_sessions + get_messages, so equality failed on
the extra key (CI red). Strip it for the row comparison and assert it is
present on every row.