Commit Graph

2431 Commits

Author SHA1 Message Date
Teknium 0d6106eab8 test: consolidate Gemini array regressions into two invariants 2026-09-07 21:10:50 -07:00
Teknium bbcf1ee180 fix: preserve native Gemini union constraints
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.

Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
2026-09-07 21:10:50 -07:00
Max Freedom Pollard 6a04ea67c0 fix(gemini): collapse array-typed tool schemas instead of crashing translation
The enum-compatibility check evaluated `[...] in {...}` on an array `type`,
raising TypeError: unhashable type: 'list' and aborting translation of the whole
tool catalog rather than the one offending tool.

Addresses both review points:

- Reuses tools.schema_sanitizer._normalize_type_array instead of picking the
  first non-null member, so a real union becomes an anyOf of single-type
  branches and no branch is dropped.
- Derives the type outside the key loop and sets nullable after it, so the
  flag implied by "null" in the array beats an input nullable: false whichever
  key the producer emitted first. Both orders are pinned by a parametrized test.
2026-09-07 21:10:50 -07:00
Teknium e9313f6458 fix: let provider evidence adjudicate past-window preflight estimates 2026-09-07 14:11:41 -07:00
kshitijk4poor ed063034ad test(compressor): pin custom_providers threading at the get_model_context_length seam
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
2026-09-08 02:27:01 +05:30
Teknium b2c394bcad test: derive process_manage verbs from the handler table; widen aux-cancellation start-up bound
The schema-diet test froze the verb list (a change-detector); it now asserts the enum
matches _SESSION_ACTIONS plus the by-name verbs. The cancellation harness gave a worker
thread 1 s to reach its transport, which a loaded CI runner missed twice this week; the
cancellation-latency assertion is unchanged, only the start-up wait is wider.
2026-09-07 12:50:29 -07:00
kshitijk4poor c5ff900761 fix: resolve the 33 F821 undefined names outside tui_gateway / feishu / godmode
Sweep of `ruff check . --select F821 --target-version py311`: 2,234 hits. 2,201 are left
alone on purpose: tui_gateway (2,169; bind_module rebinds bodies onto server.py globals,
all names verified to resolve there), the Feishu adapter (27; globals().update() SDK
binding) and the godmode script (5; dead standalone script). The other 33 were all
genuine defects. No lint config change; no TYPE_CHECKING escape hatches — every
annotation names a real, imported type; ty on the touched files: 0 new diagnostics.

- gateway/slash_commands.py: HISTORY_UNREADABLE never imported after #102117
  → NameError on the /btw error branch (same one-liner as #102952).
- gateway/platforms/whatsapp_common.py: `-> Path` return annotation with no Path
  import (the body uses `_Path`). Never raised at runtime thanks to
  `from __future__ import annotations`, but `typing.get_type_hints()` and ty
  both fail on it.
- gateway/run.py: ActivityProvenance imported at module level
  (agent.session_activity has no gateway deps); stringly annotation and the
  lazy in-function import are gone.
- tools/patch_parser.py: PatchResult imported at module level; real return
  annotation. The "avoid circular import" lazy import guarded a cycle that
  does not exist (file_operations_common never imports patch_parser).
- gateway/platforms/helpers.py: base.py imports helpers at module level, so
  MessageEvent cannot be named here; TextBatchAggregator only reads .text and
  .source, so it is typed by a BatchableEvent Protocol that MessageEvent
  satisfies structurally.
- tools/mcp_tool_sampling.py: mcp_tool imports this module, so MCPServerTask
  cannot be named here; ElicitationHandler only reads
  owner._pending_call_context, typed by an ElicitationOwner Protocol.
- plugins/platforms/sms/adapter.py: aiohttp is an optional dep ([messaging] extra) →
  module-level try/except ImportError binding `aiohttp = web = None`, the pattern the
  homeassistant / webhook / whatsapp_cloud adapters already use. Retires three lazy
  in-function imports and the `_aiohttp_available()` wrapper; `_handle_webhook` typed
  `web.Request -> web.Response`.
- plugins/platforms/teams/summary_writer.py: plain module-level `import httpx` — httpx is a
  hard core dependency (pyproject `httpx[socks]==0.28.1`), so the lazy import and the
  "imported on every CLI start" docstring premise were both wrong (plugin discovery never
  imports this module; it is reached only via the Teams adapter / meeting pipeline).

Tests:
- tests/hermes_cli/test_config.py: a test body orphaned by the wave-1 prune
  (6b81590c55) sat inside the class as dead code with self/tmp_path unbound
  — header restored, so the v11→12 custom_providers migration is covered.
- tests/tools/test_mcp_tool.py: @staticmethod recursing on `self` in the
  win32 branch; call portalocker directly.
- tests/test_background_review_list_shapes.py: main() still ran 3 pruned tests.
- tests/agent/test_cursor_optimizations_parity.py: bench() used names only
  imported inside a sibling test.
- GatewayRunner / FeishuAdapter / Dict / Optional: missing imports.
2026-09-07 22:47:33 +05:30
kshitijk4poor 233757037d test(openai): Astra pricing/capability tests assert contracts, not snapshots
``amount_usd == Decimal("5.450025")`` and ``pricing_version == "openai-gpt-6-astra-2026-09"``
fail on the next price or version bump without catching a bug. The whole-request tier is the
contract: above 272K prompt tokens every component — including cache writes, the field this PR
adds — is billed at its ``*_above`` rate, so derive the expected total from the entry's own rates
and assert below < above. Likewise the builtin-metadata test keeps only what
``_UNKNOWN_MODEL_BASE`` could not have supplied (vision) plus the openai-api == openai parity.
2026-09-07 21:43:54 +05:30
kshitijk4poor e583d68c14 perf(models): stop revalidating a fresh provider cache just because it lists Astra
``gated_cache`` bypassed both the fresh-hit and the stale-while-revalidate branches of
``cached_provider_model_ids`` whenever the disk entry contained Astra. For an entitled
account that is every entry: live discovery rewrites the entry with Astra → next call is gated
again → a blocking /models round-trip on every picker open, for exactly the providers the user
is most likely on. The parallel prefetch's staleness check didn't know about the gate either, so
it skipped the slug and the blocking call landed in the serial picker loop the prefetch exists to
avoid.

Gate on provenance, not contents: only live discovery ever writes Astra into a same-credential
entry (the static/offline paths filter it out), so a fresh entry IS the entitlement record. The
one filter that matters stays — a stale entry served because the refresh failed drops Astra, and
the entry itself is left intact so the next successful fetch restores it.

The regression test now pins both halves: fresh entry served with Astra and zero live calls;
failed refresh past the fresh window serves the entry minus Astra.
2026-09-07 21:43:54 +05:30
kshitijk4poor 170a73637d fix(openai): drop prompt_cache_options from Astra requests — not an SDK kwarg, 30m is the server default
Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.

OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.

Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.

Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
2026-09-07 21:43:54 +05:30
Eva fdb76ce643 fix(openai): include canonical API picker identity in Astra discovery gate
(cherry picked from commit de57e601193a955a69201e1a274d5ccab83a7634)
2026-09-07 21:43:54 +05:30
Eva 328542d807 fix(openai): revalidate Astra at cached and saved-model picker boundaries
(cherry picked from commit f46e2f4c0609c1543d5636c780c6da0d5ce09a56)
2026-09-07 21:43:54 +05:30
Eva 299d86851c fix(openai): keep Astra 900K alias gated and wire-compatible
(cherry picked from commit c7cd27d7f050598b9dc052c73fa1dc78c045d3c6)
2026-09-07 21:43:54 +05:30
Eva 3825d25191 fix(openai): cap Astra Codex OAuth fallback
(cherry picked from commit bb4156c3881cbaf20736f0fcfd6c3cc5c861a6fb)
2026-09-07 21:43:54 +05:30
Eva 850680fcdf fix(openai): require canonical host for Astra cache
(cherry picked from commit 6e73cc5cc66e4003276c4472e6ce2d77e602f130)
2026-09-07 21:43:54 +05:30
Eva 5a82258626 fix(openai): keep Astra rules on eligible routes
(cherry picked from commit f92fb9d9964e6e8ea118035c548aae812054f504)
2026-09-07 21:43:54 +05:30
Eva c990017482 test(openai): close Astra baseline review gaps
(cherry picked from commit 8c27b7c9316ba675e4d9ebcc7a659754f37f18ef)
2026-09-07 21:43:54 +05:30
Eva 2c315ff59b feat(openai): add GPT-6 Astra baseline support
(cherry picked from commit a8c53d20c6b16cc35745e364e16bb7259166a1d3)
2026-09-07 21:43:54 +05:30
Teknium a7198a8855 fix: keep budget checkpoints out of cancelled tool results
Skip checkpoint evaluation when a turn is interrupted so cancellation rows
remain durable without urging continued execution. The existing minimal-agent
interrupt regression also avoids dereferencing an absent iteration budget.

Consolidate the warning coverage into two invariants, including real SQLite
readback and dispatcher/child scope controls. Cold-start tool availability
between construction cases to model independent worker processes. Place ratio
normalization beside the existing iteration budget instead of growing init.

Real cancelled-tool A/B against current main, draft, and fix: three cancelled
rows and zero writes on all arms; persisted checkpoint notices 0 / 1 / 0.
Repeated scripted HTTP/SQLite loop A/B preserves completion opportunity,
ordinary default-off behavior, and blocked/two-failure exhaustion behavior.

Local targeted run initially passed 15 cases with one fixture cache-isolation
failure; corrected target and inherited affected suites remain queued behind
the campaign lock. This commit is not a CI-green or merge-ready claim.
2026-09-07 08:28:43 -07:00
Teknium 35104c20f5 test: pin checkpoint transcript persistence before the next request 2026-09-07 08:28:43 -07:00
Teknium 93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium 37fb7adfd6 fix: structured reasoning no longer breaks chat consumers
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.

Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.

Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.

Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:24:54 -07:00
Teknium bbbcde8173 fix(process): stop late forks escaping deadline tree cleanup 2026-09-07 08:19:32 -07:00
Teknium 62f64ae5b0 fix(display): retain provider provenance for persisted fallback readings 2026-09-07 08:13:01 -07:00
Teknium 04767e7aaa fix(display): distinguish estimated context from provider usage 2026-09-07 08:13:01 -07:00
Teknium 1eb1b795b5 test(gemini): verify alias routing and compatible endpoint controls 2026-09-07 08:10:36 -07:00
Teknium a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium 7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium bbbccd3935 fix: failed probe credentials cannot fall through to configured auth 2026-09-07 08:08:04 -07:00
Teknium 7d44fe9c74 fix: capability probes send minted credentials instead of callable representations
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:08:04 -07:00
Teknium 07ddfaa2d6 test: fix CI red — pool fixtures carry source, refresh assertion binds the failed key
Two test-only corrections for the failures on run 34105317301:

- tests/agent/test_anthropic_adapter.py::TestResolveAnthropicToken: the new
  skip_borrowed branch reads ``entry.source``. ``PooledCredential.source`` is a
  required dataclass field that ``from_dict`` always materializes (defaults to
  SOURCE_MANUAL) and ``_available_entries`` returns only PooledCredential, so a
  production entry can never lack it. The three SimpleNamespace doubles were the
  incomplete side; build them via ``PooledCredential.from_dict`` instead of
  duck-typing production with getattr.

- tests/agent/test_auxiliary_client.py::test_stale_anthropic_fallback_refreshes_and_retries:
  the PR itself now passes ``failed_api_key=<client.api_key>`` into
  ``_refresh_provider_credentials`` so an unrelated borrowed login never owns the
  refresh; the assertion still expected the bare ``("anthropic")`` call. Give the
  stale client an explicit api_key and assert the request-bound call. Main's
  auxiliary changes since the PR base (ebe4e7bb44, b40998bc3c..cb1a42d33b) did not
  move this call.
2026-09-07 08:07:26 -07:00
Teknium b14315ab6b test: exercise owned OAuth recovery through the unchanged ladder surface 2026-09-07 08:07:26 -07:00
Teknium bdf45abd7c fix: prefer owned Anthropic grants and bind auxiliary refresh to request credentials
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>

Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
2026-09-07 08:07:26 -07:00
Teknium de25786dc7 fix: place reauthenticated credentials by their saved identity 2026-09-07 08:06:48 -07:00
Teknium cb1a42d33b refactor: isolate custom health identity and trim redundant alias tests 2026-09-07 07:06:58 -07:00
fangliquanflq 84bbf176a3 fix(agent): preserve route URL path identity 2026-09-07 07:06:58 -07:00
fangliquanflq e3ff3e78d1 fix(agent): quarantine failed fallback destination 2026-09-07 07:06:58 -07:00
fangliquanflq e033bbaa3c fix(agent): quarantine bare custom aliases by endpoint 2026-09-07 07:06:58 -07:00
fangliquanflq 3fcbf13ddb fix(agent): scope named custom health by endpoint 2026-09-07 07:06:58 -07:00
fangliquanflq b40998bc3c fix(agent): isolate custom endpoint billing health 2026-09-07 07:06:58 -07:00
Teknium fe04d5b36d test: preserve metadata and passthrough assertions after cap removal 2026-09-07 06:15:43 -07:00
Teknium 27f32bd50b test: exercise output-cap removal across native and child surfaces 2026-09-07 06:15:43 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 4270e9ad10 fix: distinguish callable Anthropic bearer sources from OAuth text 2026-09-07 06:07:16 -07:00
Teknium b4464f6fac fix: retain rotating credentials during fallback handoff
Adapt the callable-source diagnosis from snipecoder (#102244) and fallback slice from BGwill-OUTLOOK (#102721), without unrelated reasoning or override changes.

Co-authored-by: CloudWishOS <99405975+snipecoder@users.noreply.github.com>\nCo-authored-by: BGwill-OUTLOOK <bgwillwork@outlook.com>
2026-09-07 06:07:16 -07:00
Teknium 4fbb253904 fix: mirror successful memory alias writes to providers 2026-09-07 06:05:13 -07:00
Teknium 7ef4187e11 test: pin memory alias persistence without dispatch mocks 2026-09-07 06:05:13 -07:00
liuhao1024 fb0fc7e140 fix(agent): forward the new_text alias in the memory inline executor
The memory_tool schema advertises new_text as an alias for content, and
memory_tool resolves it when content is None. But the table-driven inline
executor's arg_specs (agent/inline_tool_executors.py) did not list new_text,
so _call_tool's allowlist silently dropped it: a replace call using the
documented alias reached memory_tool with both fields None and failed with
"content is required for 'replace' action." — even though the caller
supplied the value. Forward new_text alongside content/old_text so the
documented alias fires and content still wins when both are set, matching
what the batch path (op.get("content") or op.get("new_text")) already
accepts.
2026-09-07 06:05:13 -07:00
Teknium 11082f603e test: consolidate video rejection variants into one invariant 2026-09-07 06:04:44 -07:00
fangliquanflq 00b4c49939 test(agent): cover all rejected video part types
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-07 06:04:44 -07:00