Files
hermes-agent/tests/gateway/test_model_command_reasoning_flag.py
T
teknium1 2c0bec33f9 feat(model-pickers): reasoning effort selection on every model picker
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.

One request now carries a model pick AND its effort on every surface:

- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
  <level>` (validated against `parse_reasoning_effort`; unknown level ->
  `MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
  `ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
  high` applies the effort AFTER the agent swap (`switch_model` re-resolves
  `reasoning_config` from config.yaml, so an earlier write is clobbered) with the
  pick's scope (session; config on `--global`; `--once` snapshots and restores it).
  The `/model` picker gains a third stage, "Reasoning effort for <model>", built
  from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
  inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
  `config.set model "X --reasoning high"` applies after the swap; session pin
  (`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
  one-turn restore carries `reasoning_config`; re-emits `session_info` so the
  status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
  `<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
  label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
  `_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
  the Copilot-only inline prompt; Copilot keeps its per-model level set via
  `github_model_reasoning_efforts`, other routes get the ladder, catalog
  `supports_reasoning=False` skips it) plus a "Reasoning effort for the current
  model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
  with the same step (+ "Provider default"), stored as
  `auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
  the task list ("openrouter · model · high"), cleared by "Reset all to auto";
  tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
  it.

Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
  "Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
  and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
  writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
  errors "Model names cannot contain spaces"; after switches and `config.get
  reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
  effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
  "reasoning: high", status bar "fable 5.1 high".
2026-09-13 16:43:50 -07:00

54 lines
2.2 KiB
Python

"""Gateway ``/model <m> --reasoning <level>``: the effort rides with the pick through the same
applier ``/reasoning`` uses (session override by default, ``agent.reasoning_effort`` on --global)."""
from types import SimpleNamespace
from unittest.mock import AsyncMock
import pytest
from gateway.config import Platform
from gateway.slash_commands_model import _ModelSwitchContext
def _runner():
from gateway.run import GatewayRunner
runner = object.__new__(GatewayRunner)
calls = {}
runner._switch_cached_agent_model = lambda *_a, **_k: None
runner._record_model_switch = AsyncMock()
runner._model_switch_confirmation = AsyncMock(return_value="switched")
runner._apply_reasoning_selection = (
lambda session_key, platform_key, value, persist_global=False:
calls.setdefault("applied", (session_key, platform_key, value, persist_global)) and "effort set")
return runner, calls
@pytest.mark.asyncio
async def test_reasoning_flag_applies_after_the_switch_with_the_pick_scope():
runner, calls = _runner()
ctx = _ModelSwitchContext(session_key="telegram:c1", source=None, config_path=None,
persist_global=True, reasoning_effort="high")
result = SimpleNamespace(new_model="m", target_provider="nous")
source = SimpleNamespace(platform=Platform.TELEGRAM)
reply = await runner._commit_model_switch(result, ctx, source=source)
assert calls["applied"] == ("telegram:c1", "telegram", "high", True)
assert reply == "switched\neffort set"
@pytest.mark.asyncio
async def test_no_flag_and_once_leave_reasoning_untouched():
runner, calls = _runner()
source = SimpleNamespace(platform=Platform.TELEGRAM)
result = SimpleNamespace(new_model="m", target_provider="nous")
await runner._commit_model_switch(
result, _ModelSwitchContext(session_key="k", source=None, config_path=None, persist_global=False),
source=source)
await runner._commit_model_switch(
result, _ModelSwitchContext(session_key="k", source=None, config_path=None, persist_global=False,
one_turn=True, reasoning_effort="high"),
source=source)
assert "applied" not in calls