fix(ollama-cloud): capability-gate reasoning_effort + correct disable semantics

Three follow-up fixes to the salvaged reasoning_effort support, all verified
live against ollama.com /v1/chat/completions + /api/show on deepseek-v4-pro,
gemma3, and qwen3-coder:

1. Capability-gate on /api/show 'thinking'. The original ignored the
   supports_reasoning flag and emitted reasoning_effort for every model. Now
   gated: only models whose native /api/show capabilities list contains
   'thinking' (deepseek-v4 yes; gemma3 / qwen3-coder no) get reasoning_effort.
   Mirrors the LM Studio pattern — capability resolved once per (model,
   base_url) in run_agent._supports_reasoning_extra_body via a cached probe
   (hermes_cli.models.ollama_model_supports_thinking), threaded into the
   profile hook as supports_reasoning. No live HTTP in the per-request path.

2. Disable actually disables. Ollama Cloud defaults to thinking ON and IGNORES
   the extra_body.thinking:{type:disabled} shape (verified: still returned
   reasoning). The only working off switch is top-level reasoning_effort:'none'.
   The salvaged code returned ({}, {}) for enabled:false / effort:none, leaving
   thinking ON. Now emits {'reasoning_effort': 'none'}.

3. Omit unrecognized effort. The original forwarded any unknown string verbatim
   including 'minimal' (a real Hermes effort level). Ollama Cloud rejects
   unrecognized values with a hard HTTP 400 (accepted set: low/medium/high/
   max/none), so forwarding 'minimal' would break the request. Now omitted.

Core touches (run_agent.py, hermes_cli/models.py) add the capability probe;
the plugin profile only consumes the resolved flag. 24/24 profile tests green;
194 provider/transport tests unaffected.
This commit is contained in:
kshitijk4poor
2026-06-23 23:40:41 +05:30
committed by Teknium
parent 4759362188
commit 5d9a72b7c2
4 changed files with 241 additions and 21 deletions
+48
View File
@@ -3306,6 +3306,54 @@ def lmstudio_model_reasoning_options(
return []
def ollama_model_supports_thinking(
model: str,
base_url: Optional[str],
api_key: Optional[str] = None,
timeout: float = 5.0,
) -> Optional[bool]:
"""Return True if an Ollama (Cloud or local) model advertises ``thinking``.
Probes the native ``/api/show`` endpoint and checks the ``capabilities``
list, which Ollama populates from the model's metadata (e.g.
``deepseek-v4-pro`` → ``["completion", "tools", "thinking"]`` while
``gemma3:27b`` → ``["completion", "vision"]``). This is the authoritative
capability source — the OpenAI-compat ``/v1/models`` endpoint omits it.
Returns:
True — the model declares the ``thinking`` capability.
False — ``/api/show`` succeeded but the model has no ``thinking`` cap.
None — the probe failed (unreachable / non-Ollama / error); the caller
decides the fallback (we treat None as "don't emit").
"""
import httpx
server_url = (base_url or "").strip().rstrip("/")
if server_url.endswith("/v1"):
server_url = server_url[:-3]
if not server_url:
return None
bare_model = _strip_ollama_cloud_suffix((model or "").strip())
if not bare_model:
return None
token = str(api_key or "").strip()
headers = {"Authorization": f"Bearer {token}"} if token else {}
try:
with httpx.Client(timeout=timeout, headers=headers) as client:
resp = client.post(f"{server_url}/api/show", json={"name": bare_model})
if resp.status_code != 200:
return None
caps = resp.json().get("capabilities")
if isinstance(caps, list):
return "thinking" in caps
except Exception:
return None
return None
def _fetch_github_models(api_key: Optional[str] = None, timeout: float = 5.0) -> Optional[list[str]]:
catalog = fetch_github_model_catalog(api_key=api_key, timeout=timeout)
if not catalog: