refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model

One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.

Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.

Sites -> canonical:
  hermes_cli/cli_model_switch_mixin.py::_persist_global_switch          -> deleted; _commit_model_switch calls persist_model_selection
  hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
  gateway/slash_commands_model.py::_persist_model_switch_to_config       -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
  tui_gateway/model_switch.py::_persist_model_switch                     -> deleted; _apply_model_switch calls persist_model_selection
  hermes_cli/web_server_config.py::_apply_main_model_assignment          -> apply_model_selection(result) (+ explicit custom api_key)
  hermes_cli/web_server_config.py::_validated_main_model_selection       -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
  hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
  acp_adapter/server.py::_resolve_model_selection                        -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError

Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.

Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.

Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
This commit is contained in:
teknium1
2026-09-12 21:50:37 -07:00
committed by Teknium
parent 3e066dfedd
commit 11576390fe
21 changed files with 394 additions and 277 deletions
+42 -36
View File
@@ -4,7 +4,7 @@
import logging
import os
from fastapi import HTTPException
from typing import Any, Dict, List, Optional, Tuple
from typing import Any, Dict, List, Optional, Tuple, TYPE_CHECKING
from agent.model_metadata import is_local_endpoint
from hermes_cli.config import (
DEFAULT_CONFIG,
@@ -17,6 +17,9 @@ from hermes_cli.config import (
)
from hermes_cli.web_server_memory import _normalize_memory_provider_name
if TYPE_CHECKING:
from hermes_cli.model_switch import ModelSwitchResult
# Same logger the code used before extraction (record parity).
_log = logging.getLogger("hermes_cli.web_server")
@@ -442,43 +445,43 @@ def _normalize_main_model_assignment(provider: str, model: str) -> tuple[str, st
return prov_in, model_in
def _apply_main_model_assignment(
model_cfg: "Any", provider: str, model: str, base_url: str = "", api_key: str = ""
) -> dict:
"""Apply a main-slot model assignment to a ``model`` config dict in place.
def _validated_main_model_selection(
cfg: dict, provider: str, model: str, base_url: str = "", api_key: str = ""
) -> "ModelSwitchResult":
"""Route a dashboard main-slot pick through ``switch_model`` (catalog/alias/credential
validation) seeded with the configured route, exactly like a ``/model <model> --provider
<provider> --global``. A bare ``custom`` target carries the submitted endpoint as the current
one, which is how ``switch_model`` binds a custom base_url/key. Rejections become 400s."""
from hermes_cli.config import get_compatible_custom_providers
from hermes_cli.model_switch import switch_model
Sets ``provider``/``default``, then reconciles endpoint fields. ``base_url`` and the
endpoint key share one lifecycle: an explicit value is always persisted; an existing
value is cleared ONLY when switching to a *different* provider (it belonged to the old
endpoint); a same-provider re-pick preserves it — re-picking a model used to wipe a
user's custom host (e.g. a Xiaomi MiMo Token Plan URL) and break their keys. The
runtime resolver reads ``model.base_url`` from config and only honors it when the
configured provider matches, so preserving it here is what lets the override route.
A stale secret may live under the legacy ``api`` alias with no ``api_key``, so the
switch-clears-the-key path triggers on either field. ``context_length`` is always
dropped (the new model may have a different window).
model_cfg = cfg.get("model") if isinstance(cfg.get("model"), dict) else {}
is_bare_custom = provider.strip().lower() in {"custom", "local"}
result = switch_model(
raw_input=model, explicit_provider=provider, is_global=True,
current_provider=str(model_cfg.get("provider") or ""), current_model=str(model_cfg.get("default") or ""),
current_base_url=base_url if is_bare_custom else str(model_cfg.get("base_url") or ""),
current_api_key=api_key if is_bare_custom else "",
user_providers=cfg.get("providers") if isinstance(cfg.get("providers"), dict) else {},
custom_providers=get_compatible_custom_providers(cfg))
if not result.success:
raise HTTPException(status_code=400, detail=result.error_message or "model switch rejected")
return result
Returns the same dict (a fresh dict if the input wasn't one).
"""
if not isinstance(model_cfg, dict):
model_cfg = {}
prev_provider = str(model_cfg.get("provider") or "").strip().lower()
new_provider = provider.strip().lower()
switched = new_provider != prev_provider
model_cfg["provider"] = provider
model_cfg["default"] = model
if base_url.strip():
model_cfg["base_url"] = base_url.strip()
elif model_cfg.get("base_url") and switched:
model_cfg["base_url"] = ""
def _apply_main_model_assignment(model_cfg: "Any", result: "ModelSwitchResult", api_key: str = "") -> dict:
"""Apply a main-slot selection to a ``model`` config dict via the canonical /model shape
(``hermes_cli.model_switch.apply_model_selection``). An explicit key for a custom endpoint is
the one inline credential the runtime reads (``model.api_key``); the legacy ``api`` alias is
dropped so a stale secret cannot shadow it.
Returns a new dict."""
from hermes_cli.model_switch import apply_model_selection
model_cfg = apply_model_selection(model_cfg, result)
if api_key.strip():
model_cfg["api_key"] = api_key.strip()
model_cfg.pop("api", None)
elif (model_cfg.get("api_key") or model_cfg.get("api")) and switched:
clear_model_endpoint_credentials(model_cfg, clear_api_mode=False)
if switched:
clear_model_endpoint_credentials(model_cfg, clear_api_key=False)
model_cfg.pop("context_length", None)
return model_cfg
@@ -656,7 +659,9 @@ def _apply_main_assignment_sync(cfg: dict, provider: str, model: str, base_url:
provider_entry = providers_cfg.get(provider) if isinstance(providers_cfg, dict) else None
if not base_url and isinstance(provider_entry, dict) and provider_entry.get("base_url"):
base_url = str(provider_entry.get("base_url") or "").strip()
model_cfg = _apply_main_model_assignment(cfg.get("model", {}), provider, model, base_url, api_key)
result = _validated_main_model_selection(cfg, provider, model, base_url, api_key)
provider, model = result.target_provider, result.new_model
model_cfg = _apply_main_model_assignment(cfg.get("model", {}), result, api_key)
_resolve_assignment_credentials(model_cfg, provider, provider_entry)
cfg["model"] = model_cfg
@@ -826,8 +831,9 @@ def _denormalize_config_from_web(config: Dict[str, Any]) -> Dict[str, Any]:
new_provider, resolved_model = _infer_provider_on_model_change(model_val, prev_provider)
if new_provider and new_provider.strip().lower() != prev_provider.lower():
norm_provider, norm_model = _normalize_main_model_assignment(new_provider, resolved_model)
disk_model = _apply_main_model_assignment(disk_model, norm_provider, norm_model)
model_val = norm_model
result = _validated_main_model_selection(load_config(), norm_provider, norm_model)
disk_model = _apply_main_model_assignment(disk_model, result)
model_val = result.new_model
disk_model["default"] = model_val
if ctx_sent:
if ctx_override > 0: