fix: Bedrock Guardrails enforced on the Claude route and blocks surface as refusals

bedrock.guardrail was only attached on the Converse route (guardrailConfig in the
body). Claude on Bedrock goes through the AnthropicBedrock SDK, i.e. InvokeModel,
whose body has no guardrailConfig, so the default Claude route ran with no guardrail
at all (#52179; live-verified by JiaDe-Wu: the blocked word came back through Hermes).

Bedrock reads the guardrail for InvokeModel from X-Amzn-Bedrock-GuardrailIdentifier /
-GuardrailVersion / -Trace headers. Attach them as default_headers in
build_anthropic_bedrock_client so every AnthropicBedrock client Hermes builds
(primary init, /model switch, fallback, per-request rebuild, auxiliary) enforces the
same guardrail, with prompt caching / thinking / 1M context kept (the reason Claude is
not routed through Converse).

InvokeModel blocks do NOT change stop_reason (stays end_turn) and return the guardrail's
canned text as an ordinary assistant reply, flagged only by
amazon-bedrock-guardrailAction=INTERVENED in the body (SDK: response.model_extra).
AnthropicTransport.response_finish_reason maps that to content_filter so the loop runs
its refusal handling instead of reasoning over the canned text; _derive_finish_reason
uses it for the anthropic_messages branch.

Mantle (openai.gpt-5.x) is documented by AWS as not supporting Guardrails on the
Responses endpoint; the docs now say so instead of promising "all model invocations".

Header mechanism proposed in #52312 by @JoaoMarcos44 (stale base, 7-file conflict,
detection keyed on a Converse-only stopReason); reimplemented on current main.

Live probe (local sink, SigV4 fake creds): before, no X-Amzn-Bedrock-* header on the
InvokeModel request; after, headers present, SigV4 intact, INTERVENED → content_filter.
This commit is contained in:
Teknium
2026-09-10 18:20:14 -07:00
parent d15ed44452
commit f8456248d9
6 changed files with 79 additions and 4 deletions
+5 -2
View File
@@ -416,14 +416,17 @@ def build_anthropic_bedrock_client(region: str):
"""AnthropicBedrock client for Bedrock Claude models (boto3 default credential chain). The
SDK's native Bedrock adapter gives full Claude feature parity (prompt caching, thinking
budgets, adaptive thinking, fast mode) that Converse lacks. The common betas plus
``context-1m-2025-08-07`` are attached: without the latter Bedrock caps Opus 4.6/4.7 at 200K."""
``context-1m-2025-08-07`` are attached: without the latter Bedrock caps Opus 4.6/4.7 at 200K.
A configured ``bedrock.guardrail`` rides as InvokeModel headers so every client built here
(primary, auxiliary, per-request rebuild) enforces it."""
from agent.bedrock_adapter import bedrock_guardrail_headers
sdk = _require_sdk("the Bedrock provider")
if not hasattr(sdk, "AnthropicBedrock"):
raise ImportError("anthropic.AnthropicBedrock not available. Upgrade with: pip install 'anthropic>=0.39.0'")
return sdk.AnthropicBedrock(
aws_region=region, timeout=_client_timeout(None),
max_retries=0, # retry belongs to hermes's outer loop (honors Retry-After)
default_headers=_beta_header([*_COMMON_BETAS, _CONTEXT_1M_BETA]),
default_headers={**_beta_header([*_COMMON_BETAS, _CONTEXT_1M_BETA]), **bedrock_guardrail_headers()},
)
+27
View File
@@ -321,6 +321,33 @@ def bedrock_guardrail_config(config: Optional[Dict[str, Any]] = None) -> Optiona
return out
def bedrock_guardrail_headers(config: Optional[Dict[str, Any]] = None) -> Dict[str, str]:
"""InvokeModel/Messages-wire form of the configured guardrail. The AnthropicBedrock SDK speaks
InvokeModel, which has no ``guardrailConfig`` body field; Bedrock reads the guardrail from these
headers instead (same enforcement, keeps prompt caching / thinking / 1M context)."""
gr = bedrock_guardrail_config(config)
if not gr:
return {}
headers = {
"X-Amzn-Bedrock-GuardrailIdentifier": str(gr["guardrailIdentifier"]),
"X-Amzn-Bedrock-GuardrailVersion": str(gr["guardrailVersion"]),
}
if str(gr.get("trace", "")).lower() in {"enabled", "enabled_full", "true"}:
headers["X-Amzn-Bedrock-Trace"] = "ENABLED"
return headers
GUARDRAIL_ACTION_FIELD = "amazon-bedrock-guardrailAction"
def anthropic_response_guardrail_intervened(response: Any) -> bool:
"""True when Bedrock substituted the InvokeModel reply with guardrail messaging. Unlike Converse
(``stopReason=guardrail_intervened``), InvokeModel keeps ``stop_reason=end_turn`` and signals the
block only via an unmodelled body field the Anthropic SDK keeps in ``model_extra``."""
extra = getattr(response, "model_extra", None) or {}
return str(extra.get(GUARDRAIL_ACTION_FIELD, "")).upper() == "INTERVENED"
def bind_bedrock_runtime(agent, base_url: str, api_mode: str) -> None:
"""Point *agent* at a non-Mantle Bedrock wire: ``bedrock_converse`` (boto3 direct, no SDK client) or
``anthropic_messages`` (AnthropicBedrock SDK, SigV4 via the boto3 chain). ``aws-sdk`` is a sentinel,
+10 -1
View File
@@ -97,11 +97,20 @@ class AnthropicTransport(ProviderTransport):
provider_data["anthropic_content_blocks"] = ordered_blocks
return NormalizedResponse(
content="\n".join(text_parts) if text_parts else None, tool_calls=tool_calls or None,
finish_reason=self.map_finish_reason(response.stop_reason),
finish_reason=self.response_finish_reason(response),
reasoning="\n\n".join(reasoning_parts) if reasoning_parts else None, usage=None,
provider_data=provider_data or None,
)
def response_finish_reason(self, response: Any) -> str:
"""``stop_reason`` mapped to the OpenAI vocabulary. Bedrock InvokeModel guardrail blocks keep
``stop_reason=end_turn`` and hand back the guardrail's canned text as an ordinary reply; they
must surface as ``content_filter`` so the loop treats them as a refusal, not model output."""
from agent.bedrock_adapter import anthropic_response_guardrail_intervened
if anthropic_response_guardrail_intervened(response):
return "content_filter"
return self.map_finish_reason(response.stop_reason)
def validate_response(self, response: Any) -> bool:
"""Structural check; empty content is legitimate for ``end_turn``/``refusal`` (retrying
either would loop forever)."""
+1 -1
View File
@@ -67,7 +67,7 @@ def _derive_finish_reason(agent: Any, response: Any, messages: Any) -> str:
return _codex_finish_reason(response)
transport = agent._get_transport()
if agent.api_mode == "anthropic_messages":
return transport.map_finish_reason(response.stop_reason)
return transport.response_finish_reason(response)
normalized = transport.normalize_response(response) # Bedrock already normalized at dispatch
finish_reason = normalized.finish_reason
if agent.api_mode != "bedrock_converse" and agent._should_treat_stop_as_truncated(
@@ -0,0 +1,32 @@
"""``bedrock.guardrail`` must be enforced on the Claude route, not only Converse. The AnthropicBedrock
SDK speaks InvokeModel, whose body has no ``guardrailConfig``; Bedrock reads the guardrail from
``X-Amzn-Bedrock-Guardrail*`` headers. Blocks come back with ``stop_reason=end_turn`` and the
guardrail's canned text, flagged only by ``amazon-bedrock-guardrailAction`` in the body."""
from types import SimpleNamespace
from unittest.mock import patch
CFG = {"bedrock": {"guardrail": {"guardrail_identifier": "gr-1", "guardrail_version": "DRAFT", "trace": "enabled"}}}
def test_anthropic_bedrock_client_carries_configured_guardrail_headers():
with patch("hermes_cli.config.load_config_readonly", return_value=CFG):
from agent.anthropic_adapter import build_anthropic_bedrock_client
client = build_anthropic_bedrock_client("us-east-2")
headers = {k.lower(): v for k, v in client.default_headers.items()}
assert headers["x-amzn-bedrock-guardrailidentifier"] == "gr-1"
assert headers["x-amzn-bedrock-guardrailversion"] == "DRAFT"
assert headers["x-amzn-bedrock-trace"] == "ENABLED"
assert "anthropic-beta" in headers, "guardrail headers must not displace the beta header"
def test_guardrail_intervened_invokemodel_reply_is_content_filter_not_model_text():
from agent.transports.anthropic import AnthropicTransport
transport = AnthropicTransport()
blocked = SimpleNamespace(
stop_reason="end_turn", content=[SimpleNamespace(type="text", text="BLOCKED_BY_GUARDRAIL_INPUT")],
model_extra={"amazon-bedrock-guardrailAction": "INTERVENED"},
)
clean = SimpleNamespace(stop_reason="end_turn", content=[SimpleNamespace(type="text", text="hi")], model_extra={})
assert transport.response_finish_reason(blocked) == "content_filter"
assert transport.response_finish_reason(clean) == "stop"
+4
View File
@@ -86,6 +86,10 @@ bedrock:
trace: "disabled" # "enabled", "disabled", or "enabled_full"
```
The guardrail is attached on the Converse route (`guardrailConfig`) and on the Claude route (InvokeModel headers via the Anthropic Bedrock SDK, so prompt caching and thinking are kept). A blocked request surfaces as a content-filter refusal rather than as model text. `stream_processing_mode` only applies to Converse.
AWS does not apply Guardrails to the Mantle Responses endpoint used by `openai.gpt-5.x` models ([AWS docs](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html)); use a Converse-served model when a guardrail is required.
### Model Discovery
Hermes auto-discovers available models via the Bedrock control plane. You can customize discovery: