4bdd64b334
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
597 lines
26 KiB
Python
597 lines
26 KiB
Python
"""Integration tests for the AWS Bedrock provider wiring.
|
|
|
|
Verifies that the Bedrock provider is correctly registered in the
|
|
provider registry, model catalog, and runtime resolution pipeline.
|
|
These tests do NOT require AWS credentials or boto3 — all AWS calls
|
|
are mocked.
|
|
|
|
Note: Tests that import ``hermes_cli.auth`` or ``hermes_cli.runtime_provider``
|
|
require Python 3.10+ due to ``str | None`` type syntax in the import chain.
|
|
"""
|
|
|
|
from unittest.mock import MagicMock, patch
|
|
|
|
import pytest
|
|
|
|
|
|
_BOTO_PREFIXES = ("botocore", "boto3")
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _boto_sys_modules_hygiene():
|
|
"""Snapshot/restore boto* sys.modules around every test.
|
|
|
|
Tests here plant fake botocore/boto3 modules; a fake that leaks (or a
|
|
real submodule first-imported inside a stub window) poisons later
|
|
imports of the real ``botocore.exceptions`` with
|
|
``No module named 'botocore.vendored'`` (PR #92617 CI flake). This
|
|
fixture makes stub windows airtight regardless of test ordering.
|
|
"""
|
|
import sys as _sys
|
|
|
|
saved = {
|
|
name: mod
|
|
for name, mod in _sys.modules.items()
|
|
if name.split(".", 1)[0] in _BOTO_PREFIXES
|
|
}
|
|
yield
|
|
for name in [n for n in _sys.modules if n.split(".", 1)[0] in _BOTO_PREFIXES]:
|
|
_sys.modules.pop(name, None)
|
|
_sys.modules.update(saved)
|
|
|
|
|
|
|
|
class TestProviderRegistry:
|
|
"""Verify Bedrock is registered in PROVIDER_REGISTRY."""
|
|
|
|
def test_bedrock_in_registry(self):
|
|
from hermes_cli.auth import PROVIDER_REGISTRY
|
|
assert "bedrock" in PROVIDER_REGISTRY
|
|
|
|
|
|
def test_bedrock_has_no_api_key_env_vars(self):
|
|
"""Bedrock uses the AWS SDK credential chain, not API keys."""
|
|
from hermes_cli.auth import PROVIDER_REGISTRY
|
|
pconfig = PROVIDER_REGISTRY["bedrock"]
|
|
assert pconfig.api_key_env_vars == ()
|
|
|
|
|
|
|
|
class TestProviderAliases:
|
|
"""Verify Bedrock aliases resolve correctly."""
|
|
|
|
def test_aws_alias(self):
|
|
from hermes_cli.models import _PROVIDER_ALIASES
|
|
assert _PROVIDER_ALIASES.get("aws") == "bedrock"
|
|
|
|
|
|
|
|
|
|
|
|
class TestProviderLabels:
|
|
"""Verify Bedrock appears in provider labels."""
|
|
|
|
def test_bedrock_label(self):
|
|
from hermes_cli.models import _PROVIDER_LABELS
|
|
assert _PROVIDER_LABELS.get("bedrock") == "AWS Bedrock"
|
|
|
|
|
|
class TestModelCatalog:
|
|
"""Verify Bedrock has a static model fallback list."""
|
|
|
|
def test_bedrock_has_curated_models(self):
|
|
from hermes_cli.models import _PROVIDER_MODELS
|
|
models = _PROVIDER_MODELS.get("bedrock", [])
|
|
assert len(models) > 0
|
|
|
|
def test_bedrock_models_include_claude(self):
|
|
from hermes_cli.models import _PROVIDER_MODELS
|
|
models = _PROVIDER_MODELS.get("bedrock", [])
|
|
claude_models = [m for m in models if "anthropic.claude" in m]
|
|
assert len(claude_models) > 0
|
|
|
|
def test_bedrock_models_include_nova(self):
|
|
from hermes_cli.models import _PROVIDER_MODELS
|
|
models = _PROVIDER_MODELS.get("bedrock", [])
|
|
nova_models = [m for m in models if "amazon.nova" in m]
|
|
assert len(nova_models) > 0
|
|
|
|
|
|
class TestResolveProvider:
|
|
"""Verify resolve_provider() handles bedrock correctly."""
|
|
|
|
def test_explicit_bedrock_resolves(self, monkeypatch):
|
|
"""When user explicitly requests 'bedrock', it should resolve."""
|
|
# bedrock is in the registry, so resolve_provider should return it
|
|
from hermes_cli.auth import resolve_provider
|
|
result = resolve_provider("bedrock")
|
|
assert result == "bedrock"
|
|
|
|
def test_aws_alias_resolves_to_bedrock(self):
|
|
from hermes_cli.auth import resolve_provider
|
|
result = resolve_provider("aws")
|
|
assert result == "bedrock"
|
|
|
|
|
|
def test_auto_detect_with_aws_credentials(self, monkeypatch):
|
|
"""When AWS credentials are present and no other provider is configured,
|
|
auto-detect should find bedrock."""
|
|
from hermes_cli.auth import resolve_provider
|
|
|
|
# Clear all other provider env vars
|
|
for var in ["OPENAI_API_KEY", "OPENROUTER_API_KEY", "ANTHROPIC_API_KEY",
|
|
"ANTHROPIC_TOKEN", "GOOGLE_API_KEY", "DEEPSEEK_API_KEY"]:
|
|
monkeypatch.delenv(var, raising=False)
|
|
|
|
# Set AWS credentials
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
|
|
# The Nous free tier counts as a configured provider and sits above the Bedrock chain
|
|
# (NS-829); this test's contract is the chain itself, so switch the free tier off.
|
|
monkeypatch.setattr("hermes_cli.anon_auth.guest_enabled", lambda: False)
|
|
# Mock the auth store to have no active provider
|
|
with patch("hermes_cli.auth._load_auth_store", return_value={}):
|
|
result = resolve_provider("auto")
|
|
assert result == "bedrock"
|
|
|
|
|
|
class TestRuntimeProvider:
|
|
"""Verify resolve_runtime_provider() handles bedrock correctly."""
|
|
|
|
|
|
|
|
def test_bedrock_runtime_no_credentials_raises_on_auto_detect(self, monkeypatch):
|
|
"""When bedrock is auto-detected (not explicitly requested) and no
|
|
credentials are found, runtime resolution should raise AuthError."""
|
|
from hermes_cli.runtime_provider import resolve_runtime_provider
|
|
from hermes_cli.auth import AuthError
|
|
|
|
# Clear all AWS env vars
|
|
for var in ["AWS_ACCESS_KEY_ID", "AWS_SECRET_ACCESS_KEY", "AWS_PROFILE",
|
|
"AWS_BEARER_TOKEN_BEDROCK", "AWS_CONTAINER_CREDENTIALS_RELATIVE_URI",
|
|
"AWS_WEB_IDENTITY_TOKEN_FILE"]:
|
|
monkeypatch.delenv(var, raising=False)
|
|
|
|
# Mock both the provider resolution and boto3's credential chain
|
|
mock_session = MagicMock()
|
|
mock_session.get_credentials.return_value = None
|
|
with patch("hermes_cli.runtime_provider.resolve_provider", return_value="bedrock"), \
|
|
patch("hermes_cli.runtime_provider._get_model_config", return_value={"provider": "bedrock"}), \
|
|
patch("hermes_cli.runtime_provider.resolve_requested_provider", return_value="auto"), \
|
|
patch.dict("sys.modules", {"botocore": MagicMock(), "botocore.session": MagicMock()}):
|
|
import botocore.session as _bs
|
|
_bs.get_session = MagicMock(return_value=mock_session)
|
|
with pytest.raises(AuthError, match="No AWS credentials"):
|
|
resolve_runtime_provider(requested="auto")
|
|
|
|
def test_bedrock_runtime_explicit_skips_credential_check(self, monkeypatch):
|
|
"""When user explicitly requests bedrock, trust boto3's credential chain
|
|
even if env-var detection finds nothing (covers IMDS, SSO, etc.)."""
|
|
from hermes_cli.runtime_provider import resolve_runtime_provider
|
|
|
|
# No AWS env vars set — but explicit bedrock request should not raise
|
|
for var in ["AWS_ACCESS_KEY_ID", "AWS_SECRET_ACCESS_KEY", "AWS_PROFILE",
|
|
"AWS_BEARER_TOKEN_BEDROCK"]:
|
|
monkeypatch.delenv(var, raising=False)
|
|
|
|
with patch("hermes_cli.runtime_provider.resolve_provider", return_value="bedrock"), \
|
|
patch("hermes_cli.runtime_provider._get_model_config", return_value={"provider": "bedrock"}):
|
|
result = resolve_runtime_provider(requested="bedrock")
|
|
assert result["provider"] == "bedrock"
|
|
assert result["api_mode"] == "bedrock_converse"
|
|
|
|
def test_bedrock_openai_models_route_to_mantle_responses(self, monkeypatch):
|
|
"""Bedrock's OpenAI models (GPT-5.5 / GPT-5.6 family) are not Converse
|
|
models — they only answer on the Mantle /openai/v1 Responses surface.
|
|
Every allowlisted ID must route there, with the aws-sdk IAM sentinel."""
|
|
from agent.bedrock_adapter import BEDROCK_OPENAI_RESPONSES_MODEL_IDS
|
|
from hermes_cli.runtime_provider import resolve_runtime_provider
|
|
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
monkeypatch.setenv("AWS_REGION", "us-east-2")
|
|
|
|
assert "openai.gpt-5.5" in BEDROCK_OPENAI_RESPONSES_MODEL_IDS
|
|
for suffix in ("sol", "terra", "luna"):
|
|
assert f"openai.gpt-5.6-{suffix}" in BEDROCK_OPENAI_RESPONSES_MODEL_IDS
|
|
|
|
for model_id in BEDROCK_OPENAI_RESPONSES_MODEL_IDS:
|
|
with patch("hermes_cli.runtime_provider.resolve_provider", return_value="bedrock"), \
|
|
patch("hermes_cli.runtime_provider._get_model_config", return_value={
|
|
"provider": "bedrock",
|
|
"default": model_id,
|
|
}):
|
|
result = resolve_runtime_provider(requested="bedrock")
|
|
|
|
assert result["api_mode"] == "codex_responses", model_id
|
|
assert result["model"] == model_id
|
|
assert result["base_url"] == "https://bedrock-mantle.us-east-2.api.aws/openai/v1"
|
|
assert result["api_key"] == "aws-sdk"
|
|
assert result["bedrock_openai"] is True, model_id
|
|
|
|
def test_bedrock_openai_context_length_is_272k(self):
|
|
"""AWS model cards list a 272K context window for the Mantle OpenAI
|
|
models; make sure we do not fall back to the 128K default."""
|
|
from agent.bedrock_adapter import (
|
|
BEDROCK_OPENAI_RESPONSES_MODEL_IDS,
|
|
get_bedrock_context_length,
|
|
)
|
|
for model_id in BEDROCK_OPENAI_RESPONSES_MODEL_IDS:
|
|
assert get_bedrock_context_length(model_id) == 272_000
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# providers.py integration
|
|
# ---------------------------------------------------------------------------
|
|
|
|
class TestProvidersModule:
|
|
"""Verify bedrock is wired into hermes_cli/providers.py."""
|
|
|
|
def test_bedrock_alias_in_providers(self):
|
|
from hermes_cli.providers import ALIASES
|
|
assert ALIASES.get("bedrock") is None # "bedrock" IS the canonical name, not an alias
|
|
assert ALIASES.get("aws") == "bedrock"
|
|
assert ALIASES.get("aws-bedrock") == "bedrock"
|
|
|
|
|
|
def test_determine_api_mode_from_bedrock_url(self):
|
|
from hermes_cli.providers import determine_api_mode
|
|
assert determine_api_mode(
|
|
"unknown", "https://bedrock-runtime.us-east-1.amazonaws.com"
|
|
) == "bedrock_converse"
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Error classifier integration
|
|
# ---------------------------------------------------------------------------
|
|
|
|
class TestErrorClassifierBedrock:
|
|
"""Verify Bedrock error patterns are in the global error classifier."""
|
|
|
|
def test_throttling_in_rate_limit_patterns(self):
|
|
from agent.error_classifier import _RATE_LIMIT_PATTERNS
|
|
assert "throttlingexception" in _RATE_LIMIT_PATTERNS
|
|
|
|
def test_context_overflow_patterns(self):
|
|
from agent.error_classifier import _CONTEXT_OVERFLOW_PATTERNS
|
|
assert "input is too long" in _CONTEXT_OVERFLOW_PATTERNS
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# pyproject.toml bedrock extra
|
|
# ---------------------------------------------------------------------------
|
|
|
|
class TestPackaging:
|
|
"""Verify Bedrock remains a declared lazy optional dependency."""
|
|
|
|
@staticmethod
|
|
def _optional_dependencies():
|
|
import tomllib
|
|
from pathlib import Path
|
|
|
|
content = (Path(__file__).parent.parent.parent / "pyproject.toml").read_text()
|
|
return tomllib.loads(content)["project"]["optional-dependencies"]
|
|
|
|
def test_bedrock_extra_exists(self):
|
|
extras = self._optional_dependencies()
|
|
assert "bedrock" in extras
|
|
assert any(dep.startswith("boto3==") for dep in extras["bedrock"])
|
|
|
|
def test_bedrock_is_not_eager_installed_by_all_extra(self):
|
|
extras = self._optional_dependencies()
|
|
assert "hermes-agent[bedrock]" not in extras["all"]
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Model ID dot preservation — regression for #11976
|
|
# ---------------------------------------------------------------------------
|
|
# AWS Bedrock inference-profile model IDs embed structural dots:
|
|
#
|
|
# global.anthropic.claude-opus-4-7
|
|
# us.anthropic.claude-sonnet-4-5-20250929-v1:0
|
|
# apac.anthropic.claude-haiku-4-5
|
|
#
|
|
# ``agent.anthropic_message_convert.normalize_model_name`` converts dots to hyphens
|
|
# unless the caller opts in via ``preserve_dots=True``. Before this fix,
|
|
# ``AIAgent._anthropic_preserve_dots`` returned False for the ``bedrock``
|
|
# provider, so Claude-on-Bedrock requests went out with
|
|
# ``global-anthropic-claude-opus-4-7`` (all dots mangled to hyphens) and
|
|
# Bedrock rejected them with:
|
|
#
|
|
# HTTP 400: The provided model identifier is invalid.
|
|
#
|
|
# The fix adds ``bedrock`` to the preserve-dots provider allowlist and
|
|
# ``bedrock-runtime.`` to the base-URL heuristic, mirroring the shape of
|
|
# the opencode-go fix for #5211 (commit f77be22c), which extended this
|
|
# same allowlist.
|
|
|
|
|
|
class TestBedrockPreserveDotsFlag:
|
|
"""``AIAgent._anthropic_preserve_dots`` must return True on Bedrock so
|
|
inference-profile IDs survive the normalize step intact."""
|
|
|
|
def test_bedrock_provider_preserves_dots(self):
|
|
from types import SimpleNamespace
|
|
agent = SimpleNamespace(provider="bedrock", base_url="")
|
|
from run_agent import AIAgent
|
|
assert AIAgent._anthropic_preserve_dots(agent) is True
|
|
|
|
|
|
|
|
def test_non_bedrock_aws_url_does_not_preserve_dots(self):
|
|
"""Unrelated AWS endpoints (e.g. ``s3.us-east-1.amazonaws.com``)
|
|
must not accidentally activate the dot-preservation heuristic —
|
|
the heuristic is scoped to the ``bedrock-runtime.`` substring
|
|
specifically."""
|
|
from types import SimpleNamespace
|
|
agent = SimpleNamespace(
|
|
provider="custom",
|
|
base_url="https://s3.us-east-1.amazonaws.com",
|
|
)
|
|
from run_agent import AIAgent
|
|
assert AIAgent._anthropic_preserve_dots(agent) is False
|
|
|
|
|
|
|
|
class TestBedrockModelNameNormalization:
|
|
"""End-to-end: ``normalize_model_name`` + the preserve-dots flag
|
|
reproduce the exact production request shape for each Bedrock model
|
|
family, confirming the fix resolves the reporter's HTTP 400."""
|
|
|
|
def test_global_anthropic_inference_profile_preserved(self):
|
|
"""The reporter's exact model ID."""
|
|
from agent.anthropic_message_convert import normalize_model_name
|
|
assert normalize_model_name(
|
|
"global.anthropic.claude-opus-4-7", preserve_dots=True
|
|
) == "global.anthropic.claude-opus-4-7"
|
|
|
|
|
|
|
|
def test_bedrock_prefix_preserved_without_preserve_dots(self):
|
|
"""Bedrock inference profile IDs are auto-detected by prefix and
|
|
always returned unmangled -- ``preserve_dots`` is irrelevant for
|
|
these IDs because the dots are namespace separators, not version
|
|
separators. Regression for #12295."""
|
|
from agent.anthropic_message_convert import normalize_model_name
|
|
assert normalize_model_name(
|
|
"global.anthropic.claude-opus-4-7", preserve_dots=False
|
|
) == "global.anthropic.claude-opus-4-7"
|
|
|
|
|
|
|
|
class TestBedrockBuildAnthropicKwargsEndToEnd:
|
|
"""Integration: calling ``build_anthropic_kwargs`` with a Bedrock-
|
|
shaped model ID and ``preserve_dots=True`` produces the unmangled
|
|
model string in the outgoing kwargs — the exact body sent to the
|
|
``bedrock-runtime.`` endpoint. This is the integration-level
|
|
regression for the reporter's HTTP 400."""
|
|
|
|
def test_bedrock_inference_profile_survives_build_kwargs(self):
|
|
from agent.anthropic_adapter import build_anthropic_kwargs
|
|
kwargs = build_anthropic_kwargs(
|
|
model="global.anthropic.claude-opus-4-7",
|
|
messages=[{"role": "user", "content": "hi"}],
|
|
tools=None,
|
|
max_tokens=1024,
|
|
reasoning_config=None,
|
|
preserve_dots=True,
|
|
)
|
|
assert kwargs["model"] == "global.anthropic.claude-opus-4-7", (
|
|
"Bedrock inference-profile ID was mangled in build_anthropic_kwargs: "
|
|
f"{kwargs['model']!r}"
|
|
)
|
|
|
|
def test_bedrock_model_preserved_without_preserve_dots(self):
|
|
"""Bedrock inference profile IDs survive ``build_anthropic_kwargs``
|
|
even without ``preserve_dots=True`` -- the prefix auto-detection
|
|
in ``normalize_model_name`` is the load-bearing piece.
|
|
Regression for #12295."""
|
|
from agent.anthropic_adapter import build_anthropic_kwargs
|
|
kwargs = build_anthropic_kwargs(
|
|
model="global.anthropic.claude-opus-4-7",
|
|
messages=[{"role": "user", "content": "hi"}],
|
|
tools=None,
|
|
max_tokens=1024,
|
|
reasoning_config=None,
|
|
preserve_dots=False,
|
|
)
|
|
assert kwargs["model"] == "global.anthropic.claude-opus-4-7"
|
|
|
|
|
|
class TestBedrockModelIdDetection:
|
|
"""Tests for ``_is_bedrock_model_id`` and the auto-detection that
|
|
makes ``normalize_model_name`` preserve dots for Bedrock IDs
|
|
regardless of ``preserve_dots``. Regression for #12295."""
|
|
|
|
def test_bare_bedrock_id_detected(self):
|
|
from agent.anthropic_message_convert import _is_bedrock_model_id
|
|
assert _is_bedrock_model_id("anthropic.claude-opus-4-7") is True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_bare_bedrock_id_preserved_without_flag(self):
|
|
"""The primary bug from #12295: ``anthropic.claude-opus-4-7``
|
|
sent to bedrock-mantle via auxiliary clients that don't pass
|
|
``preserve_dots=True``."""
|
|
from agent.anthropic_message_convert import normalize_model_name
|
|
assert normalize_model_name(
|
|
"anthropic.claude-opus-4-7", preserve_dots=False
|
|
) == "anthropic.claude-opus-4-7"
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# auxiliary_client Bedrock resolution — fix for #13919
|
|
# ---------------------------------------------------------------------------
|
|
# Before the fix, resolve_provider_client("bedrock", ...) fell through to the
|
|
# "unhandled auth_type" warning and returned (None, None), breaking all
|
|
# auxiliary tasks (compression, memory, summarization) for Bedrock users.
|
|
|
|
|
|
class TestAuxiliaryClientBedrockResolution:
|
|
"""Verify resolve_provider_client handles Bedrock's aws_sdk auth type."""
|
|
|
|
def test_bedrock_returns_client_with_credentials(self, monkeypatch):
|
|
"""With valid AWS credentials, Bedrock should return a usable client."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
monkeypatch.setenv("AWS_REGION", "us-west-2")
|
|
|
|
mock_anthropic_bedrock = MagicMock()
|
|
with patch("agent.anthropic_adapter.build_anthropic_bedrock_client",
|
|
return_value=mock_anthropic_bedrock):
|
|
from agent.auxiliary_client import resolve_provider_client, AnthropicAuxiliaryClient
|
|
client, model = resolve_provider_client("bedrock", None)
|
|
|
|
assert client is not None, (
|
|
"resolve_provider_client('bedrock') returned None — "
|
|
"aws_sdk auth type is not handled"
|
|
)
|
|
assert isinstance(client, AnthropicAuxiliaryClient)
|
|
assert model is not None
|
|
assert client.api_key == "aws-sdk"
|
|
assert "us-west-2" in client.base_url
|
|
|
|
def test_bedrock_returns_none_without_credentials(self, monkeypatch):
|
|
"""Without AWS credentials, Bedrock should return (None, None) gracefully."""
|
|
with patch("agent.bedrock_adapter.has_aws_credentials", return_value=False):
|
|
from agent.auxiliary_client import resolve_provider_client
|
|
client, model = resolve_provider_client("bedrock", None)
|
|
|
|
assert client is None
|
|
assert model is None
|
|
|
|
|
|
|
|
|
|
def test_bedrock_default_model_is_haiku(self, monkeypatch):
|
|
"""Default auxiliary model for Bedrock should be Haiku (fast, cheap)."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
|
|
with patch("agent.anthropic_adapter.build_anthropic_bedrock_client",
|
|
return_value=MagicMock()):
|
|
from agent.auxiliary_client import resolve_provider_client
|
|
_, model = resolve_provider_client("bedrock", None)
|
|
|
|
assert "haiku" in model.lower()
|
|
|
|
|
|
|
|
|
|
|
|
def test_bedrock_converse_shim_stream_returns_complete_response(self, monkeypatch):
|
|
"""stream=True is not supported by the shim — a complete response comes
|
|
back and call_llm's streaming consumer downgrades gracefully."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIO...MPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
|
|
from agent.auxiliary_client import BedrockAuxiliaryClient
|
|
|
|
client = BedrockAuxiliaryClient("us-east-1", "openai.gpt-oss-20b-1:0")
|
|
sentinel = object()
|
|
with patch("agent.bedrock_adapter.call_converse", return_value=sentinel) as mock_converse:
|
|
resp = client.chat.completions.create(
|
|
model="openai.gpt-oss-20b-1:0",
|
|
messages=[{"role": "user", "content": "hi"}],
|
|
stream=True,
|
|
)
|
|
# Non-streaming call_converse is still used; the caller's
|
|
# got-final-object downgrade path handles the rest.
|
|
assert resp is sentinel
|
|
assert mock_converse.call_count == 1
|
|
|
|
def test_bedrock_shim_uncapped_when_caller_omits_max_tokens(self, monkeypatch):
|
|
"""No caller max_tokens → the shim passes None through and the wire
|
|
request carries no inferenceConfig.maxTokens, so Bedrock uses the
|
|
model's maximum allowed output (#10809 on the Bedrock wire).
|
|
|
|
Guards against the shim's old hardcoded ``else 4096`` fallback, which
|
|
kept aux vision descriptions capped after the vision call sites
|
|
dropped their own caps."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIO...MPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
|
|
from agent.auxiliary_client import BedrockAuxiliaryClient
|
|
|
|
client = BedrockAuxiliaryClient("us-east-1", "openai.gpt-oss-20b-1:0")
|
|
boto3_client = MagicMock()
|
|
with patch("agent.bedrock_adapter._get_bedrock_runtime_client",
|
|
return_value=boto3_client), \
|
|
patch("agent.bedrock_adapter.normalize_converse_response"):
|
|
# Aux vision-style call: no max_tokens key at all.
|
|
client.chat.completions.create(
|
|
model="openai.gpt-oss-20b-1:0",
|
|
messages=[{"role": "user", "content": "describe"}],
|
|
temperature=0.1,
|
|
)
|
|
wire_kwargs = boto3_client.converse.call_args.kwargs
|
|
assert "maxTokens" not in wire_kwargs.get("inferenceConfig", {})
|
|
|
|
# An explicit caller cap still lands on the wire unchanged.
|
|
client.chat.completions.create(
|
|
model="openai.gpt-oss-20b-1:0",
|
|
messages=[{"role": "user", "content": "describe"}],
|
|
max_tokens=1234,
|
|
)
|
|
wire_kwargs = boto3_client.converse.call_args.kwargs
|
|
assert wire_kwargs["inferenceConfig"]["maxTokens"] == 1234
|
|
|
|
def test_bedrock_mantle_config_region_beats_env_region(self, monkeypatch):
|
|
"""bedrock.region in config.yaml must win over AWS_REGION for auxiliary
|
|
Mantle calls — the same priority the main runtime resolver uses (#65076
|
|
review: aux resolution previously derived its region env-first and
|
|
could leave the primary runtime's configured region)."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
monkeypatch.setenv("AWS_REGION", "eu-central-1")
|
|
monkeypatch.setenv("AWS_BEARER_TOKEN_BEDROCK", "test-bearer")
|
|
|
|
captured = {}
|
|
|
|
class _FakeOpenAI:
|
|
def __init__(self, **kwargs):
|
|
captured.update(kwargs)
|
|
self.api_key = kwargs.get("api_key")
|
|
self.base_url = kwargs.get("base_url")
|
|
|
|
def close(self):
|
|
pass
|
|
|
|
with patch("hermes_cli.config.load_config_readonly",
|
|
return_value={"bedrock": {"region": "us-west-2"}}), \
|
|
patch("agent.auxiliary_client.OpenAI", _FakeOpenAI):
|
|
from agent.auxiliary_client import resolve_provider_client
|
|
client, model = resolve_provider_client("bedrock", "openai.gpt-5.6-sol")
|
|
|
|
assert client is not None
|
|
assert model == "openai.gpt-5.6-sol"
|
|
assert "us-west-2" in captured.get("base_url", ""), (
|
|
"Mantle auxiliary base_url ignored config.yaml bedrock.region"
|
|
)
|
|
|
|
def test_bedrock_openai_aux_uses_responses_client(self, monkeypatch):
|
|
"""Auxiliary tasks on Bedrock GPT models use the Mantle Responses
|
|
path (SigV4 http client + aws-sdk sentinel), not the Anthropic shim."""
|
|
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "AKIAIOSFODNN7EXAMPLE")
|
|
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
|
|
monkeypatch.setenv("AWS_REGION", "us-east-2")
|
|
|
|
with patch("agent.auxiliary_client.OpenAI", return_value=MagicMock()) as mock_openai, \
|
|
patch("agent.bedrock_adapter.build_bedrock_openai_http_client", return_value=MagicMock()):
|
|
from agent.auxiliary_client import resolve_provider_client, CodexAuxiliaryClient
|
|
client, model = resolve_provider_client("bedrock", "openai.gpt-5.5")
|
|
|
|
assert model == "openai.gpt-5.5"
|
|
assert isinstance(client, CodexAuxiliaryClient)
|
|
kwargs = mock_openai.call_args.kwargs
|
|
assert kwargs["api_key"] == "aws-sdk"
|
|
assert kwargs["base_url"] == "https://bedrock-mantle.us-east-2.api.aws/openai/v1"
|
|
assert "http_client" in kwargs
|