Commit Graph

23426 Commits

Author SHA1 Message Date
negroni334 cf2d28a2ff fix(desktop): render Markdown in Bot Mode group chat messages (fixes #88391) 2026-08-17 11:29:47 -07:00
Teknium e3eddbb934 test(desktop): cover DiffusionCanvas 15fps budget and instance cap
Extends the scheduling test suite from #88407 with behavioral coverage for
the salvaged #79327 work: frames inside the 1000/15 budget reschedule
without repainting, instances over MAX_ANIMATED_INSTANCES draw one static
frame and skip the loop, and the counter releases on unmount. Both tests
verified to fail against the pause-controller-only version.
2026-08-17 11:29:28 -07:00
Chen Jin ee616b4b64 fix(desktop): throttle DiffusionCanvas rAF loop and cap concurrent instances
- Throttle frame rate to ~15fps via timestamp check instead of
  redrawing on every requestAnimationFrame callback
- Pause animation when document.hidden, resume on visibilitychange
- Cap concurrent animated instances at 2; extras render a single
  static frame instead of starting another animation loop
- Fixes #79077
2026-08-17 11:29:28 -07:00
Jakub Wolniewicz c776bdd244 fix(desktop): pause inactive diffusion rendering 2026-08-17 11:29:28 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
Jakub Wolniewicz 36d5dd3aee fix(desktop): sleep idle pixel egg animation 2026-08-17 11:26:23 -07:00
Ayush Nangia eb63c254bc test(desktop): remote-sub-profile routing cases
- remote primary + no own entry -> shared primary, scoped per request
- remote primary + own local entry -> local pooled backend
2026-08-17 10:58:48 -07:00
Ayush Nangia 15dd3bf586 fix(desktop): route remote sub-profiles through the primary's remote gateway
A URL-remote desktop whose PRIMARY profile has a per-profile remote
override lists the gateway's sub-profiles in the Bots pane, but
clicking one fell through the routing table's last case and spawned a
fresh local backend that shared nothing but the name (#88296).

resolveProfileBackendRoute now consults primaryRemoteActive: when the
primary's own backend is remote and the sub-profile has no stored
entry of its own, it routes through the primary gateway with profile
scoping (the same shared-primary flow global remote uses). Profiles
with their own local entries still pool locally.
2026-08-17 10:58:48 -07:00
David Dudok de Wit ed20a6f01a fix(desktop): preserve Bot Mode source routing 2026-08-17 10:58:48 -07:00
Teknium ce1f5dd30d fix(desktop): park the bot face clock when idle and tear it down on plugin disable
Follow-up to the salvaged #88219 visibility/15fps work:

- Dormancy: the rAF loop stops scheduling frames when no faces are mounted
  or none are visible, instead of running the 1Hz whole-document shadow-root
  scan forever. A mounting BotFace or a face scrolling into view wakes it.
- Teardown: register() now hooks ctx.onDispose so disabling the Bots plugin
  (or a hot reload) cancels the animation frame, disconnects the
  IntersectionObserver, drops cached nodes, and clears window.__hbFaceClock.
  Previously the loop ran until app restart even with the plugin disabled.
- Behavioral tests for park/wake/stop via a vm-extracted clock harness,
  verified to fail against the pre-fix source.
2026-08-17 10:58:45 -07:00
digitalbase ba2fb191c6 fix(desktop): bound bot face animation work 2026-08-17 10:58:45 -07:00
Teknium c86197e607 fix: make the self-repo git guard Windows-only
The live-checkout git mutation guard blocked history-rewriting git ops
(checkout, reset --hard, rebase, cherry-pick, ...) in the running source
checkout and its worktrees on every platform. The hazard it protects
against is only real on Windows, where NTFS locks loaded module files and
an in-place rewrite can corrupt the running process. On POSIX, open file
handles pin the old inodes, so a checkout swap under a running process is
safe, and the guard mostly taxed normal dev/salvage workflows with clone
workarounds.

- tools/self_repo_guard.py: add guard_active() -> os.name == "nt"
- tools/terminal_tool.py: consult guard_active() before running the
  detector; detector logic and block message unchanged for Windows
- tests: wiring tests force the guard on; new tests cover the POSIX
  pass-through and the platform predicate
2026-08-17 10:03:36 -07:00
Teknium 0ea430ab84 chore: map contributor email for @29206394 2026-08-17 10:02:32 -07:00
Teknium ece487ceee fix(bot-mode): live active-connection id for roster classification + source-qualified row keys
Layers on the salvaged #88489 (@29206394) and #88341 (@frizikk):

- sdk: host.activeConnectionId() — registry id of the LIVE active gateway.
  The salvaged fix classifies against the registry primary; after the user
  activates a non-primary source's agent, profiles.list answers from THAT
  source and primary-based matching would duplicate the active source's
  agents again. Live id wins, primaryConnectionId is the fallback, the
  legacy kind==='local' rule covers older desktops.
- plugin: roster/chip/picker list keys are botRowKey(bot) — source-qualified
  (connectionId, name) — so same-named agents on two genuine sources can
  never collide as duplicate React keys (the render half of the dupe-bots
  smear: name-keyed rows + duplicate names = repeated blocks every poll).
  Annotated active-source rows keep the plain-name key, so nothing remounts
  when a desktop gains the union roster.
- tests: live-id-beats-primary regression, botRowKey stability, source-shape
  anchor refresh.
2026-08-17 10:02:32 -07:00
Jakub Wolniewicz 7dd945b920 fix(desktop): deduplicate Bots roster and reuse canonical chats 2026-08-17 10:02:32 -07:00
netease 185209242e fix(desktop): stop duplicating active-gateway bots in the multi-source roster
The union agent roster (host.agents) enumerates EVERY registered connection,
including the active gateway that already answered profiles.list. The plugin
merger treated the active gateway's own agents as rows from other sources
because a remote-primary desktop reports them with connectionKind 'remote',
so every bot appeared twice (baseline) and kept growing with each refetch.

Match union agents to the active gateway via the new primaryConnectionId
field on the roster RPC response and annotate the local rows in place
instead of appending phantom copies. Same-named profiles on genuinely
separate sources (This device, other remotes) still get their own tagged
rows, preserving the @name-device disambiguation rule.

Fall back to the legacy connectionKind==='local' rule when
primaryConnectionId is absent (older Electron builds), so single-source
behavior is byte-identical.

Fixes #88344
2026-08-17 10:02:32 -07:00
kshitij 4323c67dcc fix(delegate): disclaim only the fields a failed probe actually left unmeasured
/simplify-code residual. The note hard-coded "'commits' and 'dirty' are
UNKNOWN", but the two probes fail independently: a bad base_commit fails
rev-list while `git status` still succeeds, so `dirty` is a REAL measurement
being reported as unknown. Safety was never affected (the worktree is preserved
either way), but telling the parent a measured value is untrustworthy is its own
kind of misreport — and it would push a human toward re-inspecting something
already proven.

`mark_worktree_payload_unproven()` now takes an `unmeasured` argument, and
finalize tracks which probe actually failed. The raising path still disclaims
both, because which probe raised is unknowable there.

Validation: 22/22 tests/tools/test_subagent_worktree.py; ruff + ty clean. New
guard mutation-checked (hard-coding "commits/dirty" back fails it).
2026-08-17 19:41:32 +05:30
kshitij ce93a398e8 refactor(delegate): extract the unproven-payload factory; drop the source-reading test
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.

- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
  `unproven_worktree_payload()`. Both producers of this schema now call them,
  so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
  hand-building the dict (-16 lines). The re-import is guarded: the outer
  `except` can be entered because the `from tools import subagent_worktree`
  itself failed, in which case the name is unbound -- an inline fallback keeps
  the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
  compares its key set against live `finalize_subagent_worktree()` output. Same
  contract, no source reading, refactor-proof, and it actually executes the
  code.

Also folded from the same review:

- Fail-closed on an unmeasurable commit count. With no `base_commit` the
  rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
  tree still reached `git worktree remove --force` + `git branch -D` -- the
  exact bug class #88113 is about, on a public function that takes a
  caller-supplied dict. Now returns un-inspected instead, with a test driving a
  real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
  names WHICH probe failed, its exit code, and a bounded git stderr tail, so
  the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
  removed the duplicated index-corruption block in favor of the existing
  `_break_git_index()` helper.

Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
2026-08-17 19:41:32 +05:30
kshitij 97c4f9eeec test(delegate): assert the unproven-state contract, not its prose
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.

- The distinguishability test asserted the failure payload was byte-identical
  to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
  That freezes the AMBIGUITY as a required property: emitting `commits: None`
  for "unknown" would improve exactly what #88113 is about and fail the test.
  Now asserts what the parent actually depends on -- both keep the worktree,
  and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
  happy path, forbidding an always-present-but-False flag (a legitimately
  better JSON contract: stable key set for serializers). Now
  `assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
  of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
  was not even a cross-producer contract: delegate_tool's note said "state
  unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
  assert the note names the worktree AND branch -- the actionable part for a
  human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
  before any git call would keep it green while proving nothing). Now checks
  `call_count` and mirrors the branch-survival + note-names-path legs its
  sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
  parent reads, but delegate_tool's fallback -- the second producer -- was
  verified only by reading. It now AST-parses the real fallback dict literal
  and compares against live `finalize_subagent_worktree()` output, so the two
  producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
  commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
  raising, handled in delegate_tool), and the module docstring listed
  `inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
  `_break_git_index()` beside the file's other module-level helpers.

Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
2026-08-17 19:41:32 +05:30
kshitij 38ea711fd0 fix(delegate): tell the parent when a worktree was preserved un-inspected
The preserved worktree is invisible to the only consumer that can act on it.

Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":

  inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
  inspected OK, child produced nothing     -> {commits: 0, dirty: False, pruned: False}

The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.

Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
  naming the worktree/branch, warns, and returns the payload. Both unproven
  exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
  non-numeric rev-list stdout) produced the same unproven payload but logged at
  DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
  exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
  creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
  schema missing commits/dirty/pruned. It now emits the same flagged shape, and
  logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
  affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
  by restoring the unconditional prune and reintroducing this P1.

Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.

Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
  suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
  the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
  work intact on disk; proven-clean still prunes (pruned=true).
2026-08-17 19:41:32 +05:30
liuhao1024 2b490a0513 fix(delegate): keep the worktree when git inspection fails
finalize_subagent_worktree() treated a non-zero exit from its rev-list
or status probes as proof of the payload defaults (commits=0, clean),
then pruned on them: git worktree remove --force plus branch -D
permanently deleted a child's uncommitted work whenever git could not
inspect the tree (e.g. a corrupted index) (#88113).

A destructive cleanup now requires affirmative proof of zero commits
plus a clean tree. Any non-zero inspection result keeps the worktree
and branch for manual review, with a warning naming both.
2026-08-17 19:41:32 +05:30
Jack Lau cf64ca20c5 fix(compression): stop an aborted rotation from growing the parent it could not publish
The rotation path flushes its un-persisted transcript to the parent (#47202)
and only then calls publish_compression_child. The abort handler rolls back
the in-memory transcript and keeps agent.session_id on the parent - its own
comment says "keep the parent live and discard the stale compacted snapshot" -
but the rows the flush just wrote are not part of what it discards. Every
failed rotation therefore leaves the parent transcript longer than it found
it, whatever the failure was.

That is survivable for a one-off failure and pathological for a sticky one.
A parent row carrying ended_at fails the publish on every attempt and nothing
in this path clears it, so each auto-compaction appends another copy of the
current turn to the transcript it was supposed to shrink. Worse, the growth
then satisfies conversation_compression's own len(durable_parent) >
len(messages) check, so the next attempt adopts the inflated snapshot as if it
were genuine concurrent activity and the in-memory transcript doubles too.

Check that one precondition before writing. It is a plain read of the row the
publish is about to read anyway, and it raises the publish's own message, so
split_status=aborted, failure_class=session_split_failed and the rollback path
are all unchanged; a live parent reaches the flush exactly as before.
Deliberately not extended to the compression lease, which is re-acquirable - a
transient miss there would abort a rotation that would otherwise have
committed. old_session_id moves above the flush so a failure raised from here
takes the same in-memory rollback as any other pre-publish failure.

Scope: this fixes the amplification for every abort cause. It does not fix
what marks a live session as ended in the first place (#88197 Bug 1), which
needs a maintainer decision on end-reason taxonomy and is tracked on the
issue; an affected session still aborts every attempt, it just stops making
itself larger while it does.

Refs #88197
2026-08-17 18:15:46 +05:30
kshitij 979ca57a50 Merge pull request #88244 from kshitijk4poor/fix/handoff-cleanup-race
fix: prevent handoff leg data loss + surface state.db corruption to users
2026-08-17 17:41:44 +05:30
kshitij 3baf6c14a5 fix: guard against string schedule in _clear_run_claim_best_effort
The PR's guard used `(job.get('schedule') or {}).get('kind')` which
crashes with AttributeError when schedule is a raw string (e.g.
'every 5m'), as happens in test_parallel_pool.py fixtures and any
job created via create_job(schedule='every 1h'). Use the
isinstance guard pattern already used at lines 5181 and 5325.
2026-08-17 17:31:26 +05:30
kshitij baa1cfb89d perf(cron): skip run_claim clear for recurring jobs on dispatch failure
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.

Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.

9/9 tests green; ruff clean.
2026-08-17 17:31:26 +05:30
kshitij 70fc5a5ee2 fix: guard claim cleanup best-effort + add #86522 regression tests
Follow-ups on the #87591 salvage:

- cron/scheduler.py: wrap the three clear_run_claim call sites in a
  best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
  and on the interpreter-shutdown path (or with a corrupt store) it could
  itself raise, defeating the skip-cleanly purpose of these early exits.
  A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
  clear_run_claim unit contract (one-shot cleared / already-clear noop /
  recurring never touched / unknown id), all three dispatch-failure paths
  through a real tick() clear the claim, and a raising clear_run_claim
  does not crash the tick. Mutation-verified: reverting the fix makes the
  suite fail.
2026-08-17 17:31:26 +05:30
RelaxJonh 6d85f214ac fix(cron): clear run_claim for one-shot jobs on dispatch failure (#86522)
get_due_jobs() stamps a run_claim on one-shot jobs before returning
them as due, and mark_job_run() clears it on successful completion.
When dispatch itself fails (interpreter shutdown, executor submit
error, execution-creation error) the job never reaches mark_job_run
and the stale claim blocks re-dispatch until the TTL expires
(default 30 min).

Add clear_run_claim() to jobs.py and call it on every early-exit
path in _submit_with_guard so the job stays due and fires on the
next healthy tick — matching the existing scheduler comment's
promise.

Fixes #86522
2026-08-17 17:31:26 +05:30
kshitij 0c3452bf75 Merge pull request #88327 from kshitijk4poor/chore/author-map-kstawiski
chore: add kstawiski to AUTHOR_MAP
2026-08-17 17:18:38 +05:30
kshitij eb4bc1513f fix(telegram): log first-choice IPv4 stick as info, not warning
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
2026-08-17 17:13:27 +05:30
kshitij 55e2dffe16 refactor(cron): two-phase ledger query + reuse shared terminal-state/timestamp helpers
/simplify-code findings on the salvage stack:
- efficiency HIGH: latest_executions() ran a SQLite connect + DDL + query
  every tick for the whole duration of ANY running job, even when every
  claim had a live future and the result was never consulted. Two-phase
  now: snapshot (job_id, future) under _running_lock, query the ledger
  only for claims whose future is missing/pending/done — the healthy
  steady state pays zero DB work per tick.
- quality: inline ("completed", "failed", "unknown") tuple duplicated
  cron/executions._TERMINAL_STATES (drift risk) — import the constant.
- reuse: hand-rolled naive-timestamp normalization in _row_belongs_to_claim
  duplicated cron.jobs._ensure_aware's legacy-naive policy — reuse it.
- quality: dropped the tautological 'if fut is None or pending or done'
  re-check (control only reaches it after the live-future continue) and
  collapsed the two copy-pasted release blocks into one with a computed
  reason.

24/24 tests green; mutation check re-verified on the final stack
(defeating the ownership guard fails exactly the 2 race-guard tests).
2026-08-17 16:56:54 +05:30
kshitij 6370134bfe fix: ledger-terminal release only for rows belonging to THIS claim
Follow-ups on the #87259 salvage:

- cron/scheduler.py: the ledger-terminal reconciliation now requires the
  terminal execution row's claimed_at to be >= the in-memory claim's
  registration time (_running_since). Without this, the latest terminal
  row for a recurring job is usually the PREVIOUS run's outcome — a fresh
  claim in the try_register_running_job -> create_execution window (or a
  finished run whose worker finally block hasn't released yet) would be
  force-released and the job double-dispatched. Unparseable/missing
  claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
  query under _running_lock — list() over a set concurrently mutated by
  try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
  two new race-guard tests (previous-run terminal row never releases a
  fresh claim; missing claimed_at fails closed). Mutation-verified:
  removing the ownership guard fails both.
2026-08-17 16:56:54 +05:30
devops 9b9cfbc1fd fix(cron): reconcile stale in-flight claim against executions ledger (t_8b5480b3)
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.

sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).

Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
2026-08-17 16:56:54 +05:30
kshitij 0a8e703701 refactor(cron): dedup EMFILE tick-failure handling; share the fd-exhaustion text matcher
/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
  loops (_start and _start_multiplex) along with duplicated function-local
  imports — extracted _note_tick_failure() next to _backoff_wait_seconds
  so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
  _is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
  _is_fd_exhaustion_text() out and use it from both.

11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
2026-08-17 16:55:25 +05:30
kshitij 80dc1836c4 fix: single-owner fd reclamation, shared backoff helper, clean CLI tick failure
Follow-ups on the #87796 salvage:

- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
  lock-failure raise site — the ticker loop's except handler already runs
  reclamation once per failed tick, so the raise-site call doubled the
  gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
  duplicated verbatim in start() and _start_multiplex() into a module-level
  _backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
  cleanly (exit 1) instead of dumping a traceback — tick() raising on real
  lock-acquisition failures is new behavior from this fix.
2026-08-17 16:55:25 +05:30
webtecnica 815934ae52 fix(cron): scheduler self-heals after EMFILE instead of stalling silently (#87644)
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.

- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
2026-08-17 16:55:25 +05:30
kshitij d153bfe6cb refactor(cron): single cadence-measurement implementation + bounded cadence cache
/simplify-code findings on the salvage stack:
- reuse HIGH: _compute_grace_seconds duplicated the exact croniter
  two-fire period measurement _schedule_cadence_seconds implements
  (interval minutes*60 branch included) — grace is now derived from the
  shared helper, so cadence is measured in exactly one place (and grace
  computations now benefit from the per-expr cache too).
- efficiency: _cron_cadence_cache was unbounded in principle (deleted/
  edited exprs never evicted) — hard 256-entry bound with full clear;
  rebuild cost is two croniter evals per live expr.

80 recovery/rearm/jobs tests + 94 scheduler tests green; ruff clean.
2026-08-17 16:55:00 +05:30
kshitij 8a4b28f439 fix: re-arm wedged cron jobs to next legal occurrence, cache cadence
Follow-ups on the #87261 salvage:

- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
  Re-arming to `now` fired CRON jobs at times their expression excludes —
  a weekday-only 9am job whose Friday run errored would fire on SATURDAY
  (croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
  the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
  — the next LEGAL occurrence — and only when that actually moves
  next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
  the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
  expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
  _jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
  errored Friday re-arms to Monday (not Saturday), correctly-parked cron
  value untouched, interval job still due immediately.
2026-08-17 16:55:00 +05:30
devops 122bfad535 fix(cron): persisted-state recovery re-arms recurring job stuck in stale last_status=error (t_8b5480b3)
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.

cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.

Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
2026-08-17 16:55:00 +05:30
kshitij b52b725f62 refactor(telegram): collapse sticky state onto one sentinel
_has_sticky + _sticky_ip=None overloaded "unset" and "sticky hostname".
One _UNSET sentinel is enough. Drop the unused _SEED_FALLBACK_IPS alias.
2026-08-17 16:07:05 +05:30
kshitij bd5565650d fix(telegram): try IPv4 API IPs before the dual-stack hostname
A blackholed IPv6 path to api.telegram.org never errors, so
_await_with_thread_deadline never fires and connect hangs at
"attempt 1/8". Known A-record IPs connect over IPv4 immediately.

DoH timeout now fail-opens to the seed IPv4 list instead of the
hostname. Hostname stays last for IPv6-only hosts.

Closes #87015
2026-08-17 16:07:05 +05:30
Teknium df68cc1c1a test: adapt stale dedup test to outstanding-call semantics, add cross-turn coverage
Follow-up to the salvaged #70734 fix:

- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
  global-uniqueness assumption (its second assistant call reused the id AFTER
  the first call was answered, which is now a legitimate new call). The
  replayed call now precedes the result, making it a true duplicate of a
  still-outstanding call, preserving the intended empty-tool_calls key-drop
  coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
  turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
  llama.cpp) — stock main kept 1/50 tool results, now 50/50.
2026-08-17 03:25:07 -07:00
flaviovargasbrandao 0b8fd04bea fix(agent): keep tool results when a server reuses one tool_call_id
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).

Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).

Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.

repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.

Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
2026-08-17 03:25:07 -07:00
kshitij 458f3bfac5 chore: add kstawiski to AUTHOR_MAP
Salvage of PR #73877 (cron: warn agent when scheduler is unavailable)
maps contributor email konrad.stawiski@umed.lodz.pl to GitHub login kstawiski.
2026-08-17 15:41:29 +05:30
hermes-seaeye[bot] 046a868b7f fmt(js): npm run fix on merge (#88321)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-17 10:09:48 +00:00
Teknium 2c0035d3b4 fix(desktop): bound the pre-start settle hold so a restore/edit arm can't latch the composer shut
The no-payload settle gate in gateway-event.ts held session.info
running=false off unconditionally while an optimistically armed turn
(busy/awaitingResponse from restore/edit/submit) had not gone live
backend-side. When the turn never went live at all — a rewind refused
after the optimistic arm, a submit response lost to a gateway bounce, a
terminal error event that never arrived — busy latched forever:
isTargetSessionBusy refused every send, the composer queued each message
('moves to the send area'), and the queue drain (gated on busy→false)
never fired. Only an app restart cleared it (#86795).

Bound the hold to PRE_TURN_LIVE_SETTLE_GRACE_MS (15s) measured from
turnStartedAt; past the window (or with no clock) the gateway's
running=false is authoritative and settles the session. Seed the clock +
reset turnLive in applyRewindOptimistic/applyReloadOptimistic (the
restore/edit/regenerate arm sites), and clear both on every rewind
rollback path in use-prompt-actions and session-tile-actions so a failed
rewind can't leave a stale seed.

Fixes #86795
2026-08-17 03:04:00 -07:00
Teknium 5975aff8a1 test: assert sprawl as a behavior contract, not a pack-count snapshot
CI git consolidates during incremental pack creation differently per
build (4 packs from 6 attempts on ubuntu-latest, 6 locally, 3 on the
previous run) — even pack-objects counts drift with auto-maintenance.
The fixture now only guarantees strictly-more-packs-than-threshold and
the test asserts consolidation strictly decreases the count.
2026-08-17 03:03:47 -07:00
Teknium 31fe024e93 test: make pack fixture deterministic via git pack-objects
Incremental 'git repack' consolidates small packs on newer git builds
(CI produced 3 packs from 6 commits), making the sprawl fixture count
nondeterministic. pack-objects with an explicit sha per commit creates
exactly one pack each on every git version.
2026-08-17 03:03:47 -07:00
Teknium 25c051d480 fix(cli): self-heal git health around worktree creation
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:

1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
   partially-materialized directory plus a LOCKED admin entry under
   .git/worktrees/ (lock pid = the live hermes process that timed out),
   which the startup pruner's dead-pid unlock never reaps — retries of
   the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
   admin entry, and orphaned branch on every failure path (timeout,
   nonzero exit, remote-base retry).

2. Pack maintenance: nothing consolidated the object store; on a
   multi-agent box packs sprawl (39 packs / 638MB at the incident) and
   every object lookup scans all pack indexes until worktree creation
   blows its timeout. _maintain_pack_health repacks (niced, background,
   fail-soft) when *.pack count reaches 15, wired into the existing
   startup maintenance thread on both the CLI (-w) and TUI paths.
   gc --auto doesn't cover this: its threshold is 50 packs.

Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
2026-08-17 03:03:47 -07:00
Teknium 382060f022 feat(mcp): speak the 2026-07-28 stateless protocol
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):

- Protocol-era negotiation (_negotiate_session): per-server `protocol`
  config key — auto (default, handshake-first with server/discover
  fallback on -32022/-32601), stateless (discover-first), legacy
  (handshake only). Auto is handshake-first deliberately: zero extra
  round-trips and zero behavior change for the entire existing server
  fleet, while 2026-07-28-only servers now connect via the fallback.
  All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
  route through the one choke point, so the CLI/desktop probe path
  inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
  during discovery and bound to the lazy-startup schema cache — TTL'd
  entries expire and force a live re-probe; hint-less (pre-2026)
  servers keep the never-expires behavior. Pagination continuation now
  speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
  (config-overridable), with a fallback for 1.x-era metadata models.
  (RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
  native to SDK 2.0's OAuthClientProvider — verified, no client-side
  gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
  Sampling feature as upstream-deprecated (12-month window) — kept
  fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
2026-08-17 03:03:35 -07:00
Teknium 9adc607190 fix(providers): give each CommandCode profile its own base-URL var for desktop parity
The provider-parity contract requires every CANONICAL provider to render a
card on the desktop Keys tab. /api/env rows are keyed by env var, and both
CommandCode profiles shared the single COMMANDCODE_API_KEY — so the
commandcode-anthropic profile had no row of its own and
test_provider_parity failed on CI (slice 6/12).

Fix: both profiles keep the shared API key, but each declares its own
base-URL override var (COMMANDCODE_BASE_URL / COMMANDCODE_ANTHROPIC_BASE_URL),
matching the sibling-provider pattern, so each renders its own card.

Verified: test_provider_parity.py + commandcode + providers suites green
locally (86 passed); PROVIDER_REGISTRY splits key vs base-URL vars correctly
for both profiles.
2026-08-17 02:56:17 -07:00