Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Two test-only corrections for the failures on run 34105317301:
- tests/agent/test_anthropic_adapter.py::TestResolveAnthropicToken: the new
skip_borrowed branch reads ``entry.source``. ``PooledCredential.source`` is a
required dataclass field that ``from_dict`` always materializes (defaults to
SOURCE_MANUAL) and ``_available_entries`` returns only PooledCredential, so a
production entry can never lack it. The three SimpleNamespace doubles were the
incomplete side; build them via ``PooledCredential.from_dict`` instead of
duck-typing production with getattr.
- tests/agent/test_auxiliary_client.py::test_stale_anthropic_fallback_refreshes_and_retries:
the PR itself now passes ``failed_api_key=<client.api_key>`` into
``_refresh_provider_credentials`` so an unrelated borrowed login never owns the
refresh; the assertion still expected the bare ``("anthropic")`` call. Give the
stale client an explicit api_key and assert the request-bound call. Main's
auxiliary changes since the PR base (ebe4e7bb44, b40998bc3c..cb1a42d33b) did not
move this call.
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
Reuse the effective gateway config and shared session platform policy before hashing model-facing metadata. Preserve original routing state and cover enabled/disabled redaction across all busy injection routes.
Address independent review: include existing alternate and parent routing fields, and distinguish event message ID from source message ID without inferring a reply target. Expanded live probe first failed on parent_chat_id, then passed with the additional fields.
Trim salvage of #104047: preserve the shared gateway injection boundary, but keep lossless JSON entirely per message. Do not change the system prompt, steer ABI, or synthesize a delivery target. Live handler-to-agent A/B verified; serialized directory tests and independent review pending.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Salvaged from #104272. Preserve restart and fatal-exit policy while classifying the planned restart code as success. Earlier analysis in #13604 by Justin Kausel.
The descendant fence now keeps HERMES_KANBAN_DB/BOARD/WORKSPACE so a
fenced child can still read the board it belongs to; only worker identity
(TASK, RUN_ID, CLAIM_LOCK) is scrubbed. Three pre-existing tests still
asserted the DB var was dropped and went red on CI.
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904
Adapt the earliest routing repair in scroasdale PR #44268 to the current media helpers, retaining metadata in URL fallbacks and refusing successful text-only receipts for failed local uploads. Also informed by jasondschoeman-pixel issue #104357 and ericmaddox PR #104760.
Co-authored-by: scroasdale <67333169+scroasdale@users.noreply.github.com>
ClawHub's detail endpoint now answers a slug claimed by multiple owners
with 409 AMBIGUOUS_SKILL_SLUG; the bare GET in _skill_detail returned
None for every such slug, so 'skills install clawhub/@owner/slug' (and
the owner/skills/slug URL form) failed at fetch time even though the
requester already knew the owner (#104117).
- _skill_detail forwards expected_owner as the ?owner= query param on
the detail GET (params already flows through _get_json's **kwargs).
- _parse_identifier also accepts the clawhub/@owner/slug combination:
the @ surfaces only after the clawhub/ prefix is stripped, so the
had_at check now re-runs on the stripped form. GitHub-style
owner/repo/skill paths stay rejected.
_paginate_full_list wrapped the paginated list call in try/except TypeError
to detect the mcp 1.x calling convention. The same except also caught
TypeErrors raised INSIDE the modern list call — e.g. a server response
decode failure — and retried with the legacy cursor= keyword, replacing the
real error with a misleading 'unexpected keyword argument cursor' and
making genuine MCP pagination failures undiagnosable.
Probe list_method's signature instead (_list_method_accepts_params): the
legacy cursor= fallback fires only when the method genuinely doesn't accept
the mcp 2.0 params= keyword (or takes **kwargs), so a TypeError from inside
the list call propagates to the caller. Regression tests: the decode
TypeError surfaces and the legacy retry doesn't run; a genuinely 1.x-shaped
method keeps using the cursor fallback.
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:
1. delegation.independent_completions (new, default false). #104299 made every
ungrouped task its own completion message, so a 15-task call woke the
orchestrator up to 15 times; one chain received 132 notices and answered
130 of them with "already incorporated". A multi-task call now returns as
ONE consolidated message unless the flag is on; `group` is inert until then.
2. Queued units were killed before they started. Units of one call share a
pool slot but the executor was still sized by slots, so with 15 units live
a new unit queued behind a full pool; the stale monitor's clock ran from
dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
when its thread finally came up (13 such lanes in one session). The
executor now grows to the number of live units and the stall clock arms
when the runner actually starts.
3. The tool text said "do not wait or poll — just continue" without saying
that completions are delivered only BETWEEN turns. A model that never ends
its turn (one 203-minute turn, 717 API calls) never received 40 finished
results. Tool description, dispatch note and completion header now say to
finish independent work, give a one-line status, and end the turn.
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Slim adaptation of #83772 to the current schema and failure-hint table.
Generated helpers are module exports on every execution path, not globals.
Correct schema, recovery hints and CLI tip rather than injecting names or
changing the execution boundary. Two registry-driven invariants reproduce
both misleading instructions on main and execute the corrected guidance.
Additional tool fix discovered during campaign #104904.
Original diagnosis and correction: @yuzilongleif-collab (#83772).
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>