refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
This commit is contained in:
@@ -4,7 +4,7 @@
|
||||
import logging
|
||||
import os
|
||||
from fastapi import HTTPException
|
||||
from typing import Any, Dict, List, Optional, Tuple
|
||||
from typing import Any, Dict, List, Optional, Tuple, TYPE_CHECKING
|
||||
from agent.model_metadata import is_local_endpoint
|
||||
from hermes_cli.config import (
|
||||
DEFAULT_CONFIG,
|
||||
@@ -17,6 +17,9 @@ from hermes_cli.config import (
|
||||
)
|
||||
from hermes_cli.web_server_memory import _normalize_memory_provider_name
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from hermes_cli.model_switch import ModelSwitchResult
|
||||
|
||||
# Same logger the code used before extraction (record parity).
|
||||
_log = logging.getLogger("hermes_cli.web_server")
|
||||
|
||||
@@ -442,43 +445,43 @@ def _normalize_main_model_assignment(provider: str, model: str) -> tuple[str, st
|
||||
return prov_in, model_in
|
||||
|
||||
|
||||
def _apply_main_model_assignment(
|
||||
model_cfg: "Any", provider: str, model: str, base_url: str = "", api_key: str = ""
|
||||
) -> dict:
|
||||
"""Apply a main-slot model assignment to a ``model`` config dict in place.
|
||||
def _validated_main_model_selection(
|
||||
cfg: dict, provider: str, model: str, base_url: str = "", api_key: str = ""
|
||||
) -> "ModelSwitchResult":
|
||||
"""Route a dashboard main-slot pick through ``switch_model`` (catalog/alias/credential
|
||||
validation) seeded with the configured route, exactly like a ``/model <model> --provider
|
||||
<provider> --global``. A bare ``custom`` target carries the submitted endpoint as the current
|
||||
one, which is how ``switch_model`` binds a custom base_url/key. Rejections become 400s."""
|
||||
from hermes_cli.config import get_compatible_custom_providers
|
||||
from hermes_cli.model_switch import switch_model
|
||||
|
||||
Sets ``provider``/``default``, then reconciles endpoint fields. ``base_url`` and the
|
||||
endpoint key share one lifecycle: an explicit value is always persisted; an existing
|
||||
value is cleared ONLY when switching to a *different* provider (it belonged to the old
|
||||
endpoint); a same-provider re-pick preserves it — re-picking a model used to wipe a
|
||||
user's custom host (e.g. a Xiaomi MiMo Token Plan URL) and break their keys. The
|
||||
runtime resolver reads ``model.base_url`` from config and only honors it when the
|
||||
configured provider matches, so preserving it here is what lets the override route.
|
||||
A stale secret may live under the legacy ``api`` alias with no ``api_key``, so the
|
||||
switch-clears-the-key path triggers on either field. ``context_length`` is always
|
||||
dropped (the new model may have a different window).
|
||||
model_cfg = cfg.get("model") if isinstance(cfg.get("model"), dict) else {}
|
||||
is_bare_custom = provider.strip().lower() in {"custom", "local"}
|
||||
result = switch_model(
|
||||
raw_input=model, explicit_provider=provider, is_global=True,
|
||||
current_provider=str(model_cfg.get("provider") or ""), current_model=str(model_cfg.get("default") or ""),
|
||||
current_base_url=base_url if is_bare_custom else str(model_cfg.get("base_url") or ""),
|
||||
current_api_key=api_key if is_bare_custom else "",
|
||||
user_providers=cfg.get("providers") if isinstance(cfg.get("providers"), dict) else {},
|
||||
custom_providers=get_compatible_custom_providers(cfg))
|
||||
if not result.success:
|
||||
raise HTTPException(status_code=400, detail=result.error_message or "model switch rejected")
|
||||
return result
|
||||
|
||||
Returns the same dict (a fresh dict if the input wasn't one).
|
||||
"""
|
||||
if not isinstance(model_cfg, dict):
|
||||
model_cfg = {}
|
||||
prev_provider = str(model_cfg.get("provider") or "").strip().lower()
|
||||
new_provider = provider.strip().lower()
|
||||
switched = new_provider != prev_provider
|
||||
model_cfg["provider"] = provider
|
||||
model_cfg["default"] = model
|
||||
if base_url.strip():
|
||||
model_cfg["base_url"] = base_url.strip()
|
||||
elif model_cfg.get("base_url") and switched:
|
||||
model_cfg["base_url"] = ""
|
||||
|
||||
def _apply_main_model_assignment(model_cfg: "Any", result: "ModelSwitchResult", api_key: str = "") -> dict:
|
||||
"""Apply a main-slot selection to a ``model`` config dict via the canonical /model shape
|
||||
(``hermes_cli.model_switch.apply_model_selection``). An explicit key for a custom endpoint is
|
||||
the one inline credential the runtime reads (``model.api_key``); the legacy ``api`` alias is
|
||||
dropped so a stale secret cannot shadow it.
|
||||
|
||||
Returns a new dict."""
|
||||
from hermes_cli.model_switch import apply_model_selection
|
||||
|
||||
model_cfg = apply_model_selection(model_cfg, result)
|
||||
if api_key.strip():
|
||||
model_cfg["api_key"] = api_key.strip()
|
||||
model_cfg.pop("api", None)
|
||||
elif (model_cfg.get("api_key") or model_cfg.get("api")) and switched:
|
||||
clear_model_endpoint_credentials(model_cfg, clear_api_mode=False)
|
||||
if switched:
|
||||
clear_model_endpoint_credentials(model_cfg, clear_api_key=False)
|
||||
model_cfg.pop("context_length", None)
|
||||
return model_cfg
|
||||
|
||||
|
||||
@@ -656,7 +659,9 @@ def _apply_main_assignment_sync(cfg: dict, provider: str, model: str, base_url:
|
||||
provider_entry = providers_cfg.get(provider) if isinstance(providers_cfg, dict) else None
|
||||
if not base_url and isinstance(provider_entry, dict) and provider_entry.get("base_url"):
|
||||
base_url = str(provider_entry.get("base_url") or "").strip()
|
||||
model_cfg = _apply_main_model_assignment(cfg.get("model", {}), provider, model, base_url, api_key)
|
||||
result = _validated_main_model_selection(cfg, provider, model, base_url, api_key)
|
||||
provider, model = result.target_provider, result.new_model
|
||||
model_cfg = _apply_main_model_assignment(cfg.get("model", {}), result, api_key)
|
||||
_resolve_assignment_credentials(model_cfg, provider, provider_entry)
|
||||
cfg["model"] = model_cfg
|
||||
|
||||
@@ -826,8 +831,9 @@ def _denormalize_config_from_web(config: Dict[str, Any]) -> Dict[str, Any]:
|
||||
new_provider, resolved_model = _infer_provider_on_model_change(model_val, prev_provider)
|
||||
if new_provider and new_provider.strip().lower() != prev_provider.lower():
|
||||
norm_provider, norm_model = _normalize_main_model_assignment(new_provider, resolved_model)
|
||||
disk_model = _apply_main_model_assignment(disk_model, norm_provider, norm_model)
|
||||
model_val = norm_model
|
||||
result = _validated_main_model_selection(load_config(), norm_provider, norm_model)
|
||||
disk_model = _apply_main_model_assignment(disk_model, result)
|
||||
model_val = result.new_model
|
||||
disk_model["default"] = model_val
|
||||
if ctx_sent:
|
||||
if ctx_override > 0:
|
||||
|
||||
Reference in New Issue
Block a user