fix(ollama-cloud): capability-gate reasoning_effort + correct disable semantics
Three follow-up fixes to the salvaged reasoning_effort support, all verified
live against ollama.com /v1/chat/completions + /api/show on deepseek-v4-pro,
gemma3, and qwen3-coder:
1. Capability-gate on /api/show 'thinking'. The original ignored the
supports_reasoning flag and emitted reasoning_effort for every model. Now
gated: only models whose native /api/show capabilities list contains
'thinking' (deepseek-v4 yes; gemma3 / qwen3-coder no) get reasoning_effort.
Mirrors the LM Studio pattern — capability resolved once per (model,
base_url) in run_agent._supports_reasoning_extra_body via a cached probe
(hermes_cli.models.ollama_model_supports_thinking), threaded into the
profile hook as supports_reasoning. No live HTTP in the per-request path.
2. Disable actually disables. Ollama Cloud defaults to thinking ON and IGNORES
the extra_body.thinking:{type:disabled} shape (verified: still returned
reasoning). The only working off switch is top-level reasoning_effort:'none'.
The salvaged code returned ({}, {}) for enabled:false / effort:none, leaving
thinking ON. Now emits {'reasoning_effort': 'none'}.
3. Omit unrecognized effort. The original forwarded any unknown string verbatim
including 'minimal' (a real Hermes effort level). Ollama Cloud rejects
unrecognized values with a hard HTTP 400 (accepted set: low/medium/high/
max/none), so forwarding 'minimal' would break the request. Now omitted.
Core touches (run_agent.py, hermes_cli/models.py) add the capability probe;
the plugin profile only consumes the resolved flag. 24/24 profile tests green;
194 provider/transport tests unaffected.
This commit is contained in:
@@ -3306,6 +3306,54 @@ def lmstudio_model_reasoning_options(
|
||||
return []
|
||||
|
||||
|
||||
def ollama_model_supports_thinking(
|
||||
model: str,
|
||||
base_url: Optional[str],
|
||||
api_key: Optional[str] = None,
|
||||
timeout: float = 5.0,
|
||||
) -> Optional[bool]:
|
||||
"""Return True if an Ollama (Cloud or local) model advertises ``thinking``.
|
||||
|
||||
Probes the native ``/api/show`` endpoint and checks the ``capabilities``
|
||||
list, which Ollama populates from the model's metadata (e.g.
|
||||
``deepseek-v4-pro`` → ``["completion", "tools", "thinking"]`` while
|
||||
``gemma3:27b`` → ``["completion", "vision"]``). This is the authoritative
|
||||
capability source — the OpenAI-compat ``/v1/models`` endpoint omits it.
|
||||
|
||||
Returns:
|
||||
True — the model declares the ``thinking`` capability.
|
||||
False — ``/api/show`` succeeded but the model has no ``thinking`` cap.
|
||||
None — the probe failed (unreachable / non-Ollama / error); the caller
|
||||
decides the fallback (we treat None as "don't emit").
|
||||
"""
|
||||
import httpx
|
||||
|
||||
server_url = (base_url or "").strip().rstrip("/")
|
||||
if server_url.endswith("/v1"):
|
||||
server_url = server_url[:-3]
|
||||
if not server_url:
|
||||
return None
|
||||
|
||||
bare_model = _strip_ollama_cloud_suffix((model or "").strip())
|
||||
if not bare_model:
|
||||
return None
|
||||
|
||||
token = str(api_key or "").strip()
|
||||
headers = {"Authorization": f"Bearer {token}"} if token else {}
|
||||
|
||||
try:
|
||||
with httpx.Client(timeout=timeout, headers=headers) as client:
|
||||
resp = client.post(f"{server_url}/api/show", json={"name": bare_model})
|
||||
if resp.status_code != 200:
|
||||
return None
|
||||
caps = resp.json().get("capabilities")
|
||||
if isinstance(caps, list):
|
||||
return "thinking" in caps
|
||||
except Exception:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _fetch_github_models(api_key: Optional[str] = None, timeout: float = 5.0) -> Optional[list[str]]:
|
||||
catalog = fetch_github_model_catalog(api_key=api_key, timeout=timeout)
|
||||
if not catalog:
|
||||
|
||||
Reference in New Issue
Block a user