Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
Retain partial-line output, UTF-8 decoding, failure output and cancellation cleanup. Based on streaming investigations by Artemonim (#101850) and lEWFkRAD (#104843); gateway tee adapted from fangliquanflq (#97402). Live Linux child/tee probe: withheld or dropped on base, visible in 0.02 seconds after. Campaign-locked tests and native Windows proof are pending.
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
Reuse the focused pane derivation for visible sidebar activity rather than trusting a stale workspace route. Live Electron reproduction: Kanban stayed highlighted while a resumed session tab was visible; the highlight now clears without changing the retained route. Adapted the rendered sidebar invariant and focus derivation from MarcoFernstaedt's PR #88691 for #88517. Directory suite queued under the campaign test lock.
Native Windows run 34096838164 reports false success for absent and corrupt executables, missing bundle files, missing chunks and missing or stale stamps. Reuse the existing build identity and PE validators, and check interpreter presence before waiting for Desktop. Preserve dependency recovery and exit-2 refusal behavior.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Port the runtime verification portion of #104692 after native run 34095483533 reproduced ok=true for a zero-exit controlled child that removed its runtime module. Artifact/build-stamp validation remains unaddressed.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Salvage the missing-target guard from #104692. Native Windows run 34094671567 returned exit zero for the absent maintained script. Full runtime/artifact completion remains separate.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Preserve profile ownership proven at the secondary socket boundary rather
than trusting arbitrary wire profile fields. Retire transient local owners
with the profile pool, while keeping durable and exact remote ownership first.
Salvage #103774 with two invariant tests and routing documentation. The
profile-only fallback was also identified in the earlier #103770; this
version retains the producer provenance and retirement boundary.
Live Desktop renderer with two isolated serve backends: Reject previously
failed after clearing durable bindings, leaving approval pending. With this
change, the same action sends deny to the owning backend and pending clears.
Existing-binding controls pass on both sides. No vendor inference used.
Fixes#103755
Salvaged-from: 8a3c545e255b66b6ca4bc4b99cd725c5f7a08632
Keep compact completion labels while rendering only result and job output bodies, including legacy and batch deliveries. Credit Gyarados4157's #101083 investigation; avoid rendering its full model instruction envelope.
Remove whole-document whitespace cleanup that erased hard breaks and fenced-code spacing across all five ingestion paths. Limit generated-image gap joining to the removed image itself. Keep soft newlines as soft breaks.
Investigated #97117 and the approaches in #97431 (@wooyongbin3-cpu) and #97175 (@Jackal991); both leave fenced-code whitespace and sibling ingestion paths exposed, so remove the destructive normalization instead.
Salvage #99095 (e7ea53074ab2b64a1530641659399d9f1bb4b435), completing one-time migration for both boolean values and using the existing storage helpers. Hydration and local toggles never edit backend configuration. Fixes#99076.
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Carve the history-only implementation and regression from #104754
(b4bfa76facb43111a863b1256c89875e3f01fde0) by PLASMA-FR; omit
the unrelated locale-picker changes. Complement merged #104523.
Real serve + WebSocket reconnect: SQLite retains two rows; before,
resume/activate/history return only the user; after, both survive.
Plain-content control retains both rows in both arms. Native macOS
sleep and the full desktop symptom are not established by this probe.
Refs #68321
Wire the existing consent-aware, profile-keyed hook registrar at agent
construction, where the correct session home is already bound. This covers
serve and TUI agent construction without a startup-only registration or a
new helper that swallows registration failures.
Live isolated serve/WebSocket probes reproduce the missing registration on
base for both write_file and terminal, then verify each consented profile
blocks its configured tool while the unapproved profile still runs normally.
Repeated alpha construction does not duplicate hook callbacks.
Slim implementation of the agent-build placement proposed in #57020;
thanks also to the profile-scoped analysis in #102691.
Co-authored-by: grimmjoww578 <willies578@gmail.com>
The real closed-WebSocket resident variant still wedges after the missing-timer repair. Re-enter existing transport cleanup before rearming orphan timers, preserving viewer transfer and the reconnect/delegation fence rather than deleting registry rows from a stale snapshot. Extend the same invariant to both dead-transport shapes. Live class investigation informed by #104710; no direct-vouch reclaim machinery imported.
Salvage #104704 (e9423d2d0bbe3e795c5eaccb86a913f1d95444ba, 5fa2b98fc833fc4e1ee7f1aaa7eb45cdd3cd7cde). Reuse guarded orphan teardown instead of deleting ownership fences. Real two-backend WebSocket probe reproduces the missing-timer wedge on base and proves reconnect/delegation protection and recovery. Add reusable probe and user documentation.
Salvage #101453 (03a3f466d38134ba416764185884b3d655197a1d). Preserve its opt-out and first-run behavior; replace predicate-mocked tests with one native config/filesystem invariant and clarify XDG docs. Real venv/XDG probe: base clobbers custom entry, fix preserves it; targeted suite 94 passed.
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.
Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
Slim rework of despotak's modified-keypad fix in #97290. Mirror existing
non-keypad mappings for modified keypad keys, including lock-state variants,
so Alt+keypad Enter reaches the existing newline handler rather than leaking
[57414;3u into the draft. Preserve installed twin mappings before consulting
pending aliases, matching first-writer-wins registration.
Replace the source PR's keyed branch ladder with a format table and verify
parser parity plus real buffer insertion with two invariant tests. Document
keypad multiline support in English and Chinese.
Live PTY: the exact doubled leak after a real collapsed paste reproduces on
main; all 21 editor cases pass with the fix, including ordinary Enter and
legacy Alt+Enter controls. Whitespace also reproduces on main: adjacent
characters are not the root cause.
Co-authored-by: Christos Despotakis <christos@despotak.is>
Keep parsing, contracts, gates and persisted goal mutations in one dispatcher. Adapters retain authorization, rendering and scheduling; TUI drafting resolves the target session profile off the RPC reader. Document ACP as unsupported rather than implying a goal loop exists.
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
evals/token_accounting/replay_gates.py runs the REAL AIAgent turn loop against a local fake
chat-completions server with scripted usage.prompt_tokens, in three shapes (CLI same-object
history, gateway JSON-reloaded history, SessionDB close/reopen + fresh agent) x two arms
(transcript 3x threshold by bytes/4 while real usage is under; transcript tiny while real usage is
over). Acceptance: no gate fires when real usage is under threshold regardless of estimate
inflation; every gate fires once real usage is over.
origin/main 5f406d88ea: cli/gateway/restore inflated all FAIL (1 local compaction each on the
estimate). This branch: 6/6 PASS, anchor restored in the fresh process.
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.