vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)

* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
This commit is contained in:
Teknium
2026-08-28 12:15:24 -07:00
committed by GitHub
parent 72874b0675
commit 93de1d3430
3 changed files with 62 additions and 42 deletions
+27 -21
View File
@@ -17,17 +17,18 @@ It reads ``agent.image_input_mode`` from config.yaml (``auto`` | ``native``
| ``text``, default ``auto``) and the active model's capability metadata.
In ``auto`` mode:
- If the active model reports ``supports_vision=True`` (via config
override or models.dev metadata), we attach natively — vision-capable
main models should always see the original pixels, even when an
auxiliary vision backend is configured. That auxiliary backend then
acts as a *fallback* for sessions whose main model can't take images.
- Otherwise, if the user has explicitly configured ``auxiliary.vision``
(provider/model/base_url not ``auto``/empty), we route through the
text pipeline so the auxiliary vision backend can describe the image
for the text-only main model.
- Otherwise (non-vision model, no explicit override), we fall back to
text via the default vision_analyze flow.
- If the user has explicitly configured ``auxiliary.vision``
(provider/model/base_url not ``auto``/empty), images route through
that backend — the DE-FACTO choice: a user who named a dedicated
vision model wants it used, even when the main model has native
vision (maintainer decision 2026-08-28, reversing #29135's
fallback-only posture).
- Otherwise, if the active model reports ``supports_vision=True`` (via
config override or models.dev metadata), we attach natively.
- Otherwise (non-vision model, no aux backend), text via the default
vision_analyze flow.
``agent.image_input_mode: native`` remains the absolute override for
users who want native attach despite a configured aux backend.
This keeps ``vision_analyze`` surfaced as a tool in every session — skills
and agent flows that chain it (browser screenshots, deeper inspection of
@@ -448,10 +449,11 @@ def _coerce_mode(raw: Any) -> str:
def _explicit_aux_vision_override(cfg: Optional[Dict[str, Any]]) -> bool:
"""True when the user configured a specific auxiliary vision backend.
An explicit override means the user has a dedicated vision backend
available; it's used as a *fallback* when the main model can't take
images natively. In ``auto`` mode, native vision on a vision-capable
main model still wins over this fallback — see issue #29135.
An explicit backend is the DE-FACTO image route in ``auto`` mode —
the user named a dedicated vision model, so images go through it even
when the main model could take them natively (maintainer decision,
reversing #29135). ``agent.image_input_mode: native`` still forces
native; unset/auto aux config leaves native as the default.
"""
if not isinstance(cfg, dict):
return False
@@ -586,10 +588,16 @@ def decide_image_input_mode(
if mode_cfg == "text":
return "text"
# auto: prefer native vision when the main model supports it. An
# explicit auxiliary.vision config acts as a *fallback* for text-only
# main models — it should not preempt native vision on a model that
# can natively inspect the pixels (issue #29135).
# auto: an explicitly configured auxiliary.vision backend is the
# DE-FACTO choice — the user named a dedicated vision model, so that's
# what they want images to go through, even when the main model has
# native vision (maintainer decision, 2026-08-28, reversing #29135's
# fallback-only posture: config that only takes effect when the main
# model gets worse is a trap, not a setting). Native vision remains
# the default for unconfigured installs, and the fallback when the
# aux backend is unset.
if _explicit_aux_vision_override(cfg):
return "text"
if requested_provider:
supports = _lookup_supports_vision(
provider,
@@ -603,8 +611,6 @@ def decide_image_input_mode(
supports = _lookup_supports_vision(provider, model, cfg)
if supports is True:
return "native"
if _explicit_aux_vision_override(cfg):
return "text"
return "text"
+20 -6
View File
@@ -64,13 +64,27 @@ class TestDecideImageInputMode:
with patch("agent.image_routing._lookup_supports_vision", return_value=None):
assert decide_image_input_mode("openrouter", "brand-new-slug", {}) == "text"
def test_auto_prefers_native_for_vision_capable_main_model_even_with_aux_configured(self):
"""Regression #29135: vision-capable main model wins over aux fallback.
Auxiliary.vision is a fallback for text-only main models; it must
not preempt native vision on a vision-capable main model.
"""
def test_auto_explicit_aux_backend_is_the_defacto_route(self):
"""Maintainer decision (2026-08-28, reverses #29135): a user who
NAMED a dedicated vision backend wants it used — even when the
main model has native vision. Config that only takes effect when
the main model gets worse is a trap, not a setting."""
cfg = {"auxiliary": {"vision": {"provider": "openrouter", "model": "google/gemini-2.5-flash"}}}
with patch("agent.image_routing._lookup_supports_vision", return_value=True):
assert decide_image_input_mode("anthropic", "claude-sonnet-4", cfg) == "text"
def test_auto_unset_aux_backend_native_remains_default(self):
"""No configured aux backend -> native for vision-capable mains
(the unconfigured-install default is unchanged)."""
for cfg in ({}, {"auxiliary": {}}, {"auxiliary": {"vision": {"provider": "auto"}}}):
with patch("agent.image_routing._lookup_supports_vision", return_value=True):
assert decide_image_input_mode("anthropic", "claude-sonnet-4", cfg) == "native"
def test_image_input_mode_native_overrides_aux_backend(self):
"""agent.image_input_mode: native stays the absolute escape hatch —
forces native attach even with an explicit aux backend."""
cfg = {"agent": {"image_input_mode": "native"},
"auxiliary": {"vision": {"provider": "openrouter", "model": "google/gemini-2.5-flash"}}}
with patch("agent.image_routing._lookup_supports_vision", return_value=True):
assert decide_image_input_mode("anthropic", "claude-sonnet-4", cfg) == "native"
+15 -15
View File
@@ -1807,15 +1807,16 @@ from tools.registry import registry, tool_error
VISION_ANALYZE_SCHEMA = {
"name": "vision_analyze",
# Dieted (#95681): routing mechanics (native attach vs aux-model text
# fallback) removed — the route is automatic and the native path's own
# tool result says "you can see it natively now"; the schema doesn't
# need to predict plumbing. Region keeps its flow teaching: it's
# pre-effect guidance (a model that doesn't know crops keep full
# resolution never zooms).
"description": (
"Load an image into the conversation so you can see it. Accepts a "
"URL, local file path, or data URL. When your active model has "
"native vision, the image is attached to your context directly "
"and you read the pixels yourself on the next turn — call this "
"any time the user references an image (filepath in their message, "
"URL in tool output, screenshot from the browser, etc.). For "
"non-vision models, falls back to an auxiliary vision model that "
"returns a text description."
"Load an image into the conversation so you can see it. Call it "
"any time the user references an image — then answer from what "
"you see."
),
"parameters": {
"type": "object",
@@ -1826,7 +1827,7 @@ VISION_ANALYZE_SCHEMA = {
},
"question": {
"type": "string",
"description": "Your specific question or request about the image. Optional context the model uses on the next turn after seeing the image."
"description": "Your question or request about the image."
},
"region": {
"type": "array",
@@ -1834,12 +1835,11 @@ VISION_ANALYZE_SCHEMA = {
"minItems": 4,
"maxItems": 4,
"description": (
"Optional [x1, y1, x2, y2] crop region in pixel coordinates "
"of the ORIGINAL image, applied before any downscaling so "
"the region keeps full resolution. Intended flow: load the "
"full image first, then call again with a region to zoom "
"into a detail (small text, UI element, fine print). "
"Coordinates are clamped to the image bounds."
"Optional [x1, y1, x2, y2] crop in ORIGINAL-image pixel "
"coordinates, applied before any downscaling — the crop "
"keeps full resolution. Load the full image first, then "
"re-call with a region to zoom into small text or fine "
"detail."
)
}
},