8b6931393e
Providers key prompt caches per model, so a mid-session /model switch makes the next reply re-read the entire conversation at full input price. deepagents gates user-initiated switches behind a confirmation once the active thread exceeds a configurable token threshold; this ports the same protection into Hermes' unified selection-guard registry so it renders on every surface at once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard). - hermes_cli/model_selection_guards.py: new context_cache guard + SelectionContext carrier + selection_context_for_agent() helper; registry threads live-session facts to guards (6-arg signature with a TypeError fallback for externally patched 5-arg guards). - config: model.switch_context_confirm_tokens (default 100000, 0 disables). - cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the live agent's measured context into the guard call. - docs: configuring-models.md mid-session switch section. - tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).