Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
test_prompt_stash_cli.py stubs prompt_toolkit with a bare module, so the
module-level 'from prompt_toolkit.enums import EditingMode' crashed every
test importing cli. Import it with the same ImportError fallback as
CursorShape and pass editing_mode via extra_kw only when available.
The owner-only pre-create helper ran before sqlite3.connect() and turned
a directory-as-state.db misconfiguration into IsADirectoryError instead
of the sqlite OperationalError the open path (and its lock-patience
classifier) expects. A directory leaks no row data, so skip it and let
sqlite fail canonically. Also map the salvage carry-commit author email
for the attribution gate.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).
Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.
The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.
Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.
Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
noninteractive_git_env() now spawns `git config --get-all safe.directory` before
building the env. Eight tests fake subprocess.run/Popen with a fixed sequence of
expected git calls (update check, plugin pull, MCP install, bounded probe) and the
extra spawn tripped them in CI. An autouse fixture stubs the read to "no entries";
the two carve-out invariant tests opt back in with @pytest.mark.real_safe_directory
(and were confirmed to still exercise the real read: the ordering test would fail
against the stub).
contributors/emails: pry@privacydied.net -> privacydied (check-attribution).
`hermes sessions repair --check-only` printed the corruption reason and
exited 0, so scripts and the console wrapper gating on the status read a
broken state.db as healthy. Return 1 from the CLI handler (main.py already
sys.exits a truthy return) and from the console handler, where
`_capture_output` turns the status into a ConsoleCommandError carrying the
printed reason.
Salvaged from PR #103321 (the check-only reporting part only; the probe
rewrite and connection-tracking changes were not taken). Refs #63386.
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.
One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".
Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).
Fixes#97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
A SIGKILL mid-write (OOM killer, #106667) can leave state.db with a garbage
page-1 header. SQLite refuses the file outright ("file is not a database",
SQLITE_NOTADB) and the sqlite3 shell's .recover opens the file like any other
client, so the lost_and_found lane failed with the same rc=26 twice although
every data page after the header survived. The #106587 quarantine now
preserves such a file as state.db.notadb-<ts>-<pid>.bak; this makes that
preserved file recoverable with `hermes sessions recover --source <bak>
--allow-partial`.
When both .recover attempts fail with "not a database", zero the 100-byte
header of the lane's private snapshot copy and rerun them. .recover trips
only on the magic check and infers page size and layout from the pages
themselves, so a zeroed header is enough; a spliced donor header (the PR's
original mechanism) instead advertises a database size / freelist that
contradicts the file and yields "database disk image is malformed" on a
direct open — verified live on a 139-page fixture, which also showed the
zeroed header recovers 60/60 sessions and 300/300 messages whether the
damage covers 100 bytes or the whole first page. The user's file is never
written; the report carries `sqlite3_cli.header_zeroed` and a warning about
the WAL boundary.
Live repro (sqlite3 shell 3.53.1 on PATH, header overwritten with random
bytes): BEFORE "page-level .recover salvage failed: ... file is not a
database (26)"; AFTER "Recovered 60 sessions and 300 messages", source md5
unchanged.
Salvaged from PR #102808 (intent; trimmed from 302 to ~40 source LOC by
dropping the donor-header/page-size sweep and the redundant preopen probe —
the shell's own refusal is the detector). Independent review on the PR by
@strzhao.
Refs #106667
Refs #106587
Reported-by: TaoMasterCoder
Cross-referenced-by: kshitijk4poor
`hermes sessions recover --allow-partial` walks damaged tables by rowid range
and falls back to an exact-rowid lookup for the last cell of a broken range.
A phantom row produced by page damage (a `sessions` row with a NULL
`started_at`, an `async_delegations` row with a NULL `state`) reads fine from
the source but violates the destination's NOT NULL constraint, and the
resulting sqlite3.IntegrityError escaped the exact-lookup boundary and aborted
the whole recovery — losing every healthy row behind it.
Catch IntegrityError at that boundary, count it under
`destination_rejected_rows` and record the rowid as a skipped singleton
("destination constraint rejected row: ...") so the run completes, verifies,
and the report shows exactly which rows were dropped.
Live repro (real fixture: source schema's NOT NULL relaxed via
writable_schema, phantom `sessions` row inserted): BEFORE
sqlite3.IntegrityError "NOT NULL constraint failed: sessions.started_at" at
session_recovery.py recover_exact_rowid; AFTER status=partial copied=3
destination_rejected_rows=1, verified=True. Reporter @i8ei confirmed the same
patch recovers the field database (22 sessions / 2,248 messages) in #102240.
Salvaged from PR #91413 (rebased onto the _RowidRangeSalvage refactor; test
reduced to one invariant).
Refs #102240
Reported-by: i8ei
* fix(desktop): filter expired Windows CAs and deduplicate trust roots
* test(desktop): verify CA expiry and deduplication with real certificates
* test(desktop): cover Linux alongside macOS CA no-op policy
Co-authored-by: Gabriel Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
* chore: map CA trust contributors for release attribution
* style(desktop): match CA filtering to lint rules
---------
Co-authored-by: wliu-dev <249166551+wliu-dev@users.noreply.github.com>
Co-authored-by: Gabriel Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
finish_reason='length' has two causes: the answer was long (max_tokens reached),
or the prompt itself left no room to generate. _continue_text treated both the
same: append the fragment + a continuation nudge and retry, up to 4 times. In the
second case every retry sends a strictly longer prompt, so each attempt is worse
(Ollama n_ctx=32768: 32,638 -> 32,685 -> 32,732 prompt tokens, all truncated), the
user is told "model hit max output tokens", and max_tokens is not the lever.
The response's usage already carries prompt_tokens and the compressor already
resolves the model's context window; compare them once per truncation. Under
_MIN_CONTINUATION_HEADROOM (512) free tokens the turn ends on the first
truncation, keeps the partial text, names the context window as the cause and
points at /compress or a larger window. Unknown usage or window keeps today's
behaviour. max_tokens semantics untouched.
Co-authored-by: gaoanze888 <214786078+gaoanze888@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Collapse the four class-based tests into two parametrized invariants over both
refunding restart flags and move them to tests/agent/ (the phase modules live in
agent/): a single restart still refunds-and-continues; a re-armed restart breaks
after max_retries refunds. The stub grows the redirect seam the follow-up commit
uses so the queued-correction contract is covered by the same test.
Drop the sentinel-only batch test: a batch sentinel is already rejected by the
shared _is_clarify_non_response_sentinel list check that the existing sentinel
tests pin, so the case adds no new contract. Also add the contributor email
mapping for the cherry-picked commit so release CI can attribute it.