561e161123
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
- Gateway internal identity: when a service token is configured, reject wrong/missing tokens even from loopback (closes SSRF/local bypass). - Terminal metering: classified AgentControlError propagates without retry; exhausted retries raise BILLING_UNAVAILABLE instead of a generic RuntimeError, keeping error attribution accurate.
Model Runtime Layout
The model runtime has three configuration and execution boundaries.
| Layer | Source | Owns | Must not own |
|---|---|---|---|
| Provider | configuration/provider.py |
Adapter identity, credentials, endpoints, headers, connection defaults | Model capabilities, model token limits, derived tool transport |
| Model | configuration/model.py |
Provider model ID, capabilities, limits, canonical parameters, access, billing | Credentials, base URL, SDK client options, derived tool transport |
| Invocation | invocation/contract.py |
Immutable API mode, output parameter, tool transport, streaming flag, final SDK parameters | Admin persistence, credentials, routing decisions |
Supporting modules have narrower responsibilities:
model_config_v4.pynormalizes and persists the Provider + ModelProfile admin contract, then projects it to the stable runtime schema.model_config.pyparses and validates the runtime schema. It re-exports the provider and model contracts for compatibility with existing integrations.adapter_registry.pydeclares provider/model-family support and converts canonical model parameters into provider SDK parameters.runtime.pyselects a frozen route, asks its adapter to compile parameters, compiles anInvocationPlan, and constructs the provider client from that plan only.
The call chain is fixed:
V4 Provider + ModelProfile
-> normalize and validate
-> V3 runtime projection
-> select provider endpoint and model profile
-> merge canonical model parameters
-> provider adapter compilation
-> immutable InvocationPlan validation
-> provider SDK call
Important invariants:
- Environment variables may provide secrets, proxy settings, and timeouts; they cannot select an API protocol or rewrite a compiled invocation.
tool_call_transportis not administrator configuration. It is derived asnativewhencapabilities.tools=true, otherwisedisabled.- Exactly one provider output-limit parameter is allowed in a compiled plan:
max_output_tokens,max_completion_tokens, ormax_tokens. - Provider-specific parameter names are selected by the adapter. Gateway, frontend, and generic runtime code must not guess them from model names.
- Runtime logs report the final non-secret plan and parameter names. They must never include credentials, authorization headers, or raw secret values.
- Provider input projection removes assistant history that has neither final
text nor a tool call. A newly completed empty response receives one bounded
same-route repair attempt, then fails as
MODEL_PROVIDER_RESPONSE_INVALID.