Port from langchain-ai/deepagents#5829: confirm mid-session model switches that abandon a large cached context

Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).

- hermes_cli/model_selection_guards.py: new context_cache guard +
  SelectionContext carrier + selection_context_for_agent() helper;
  registry threads live-session facts to guards (6-arg signature with a
  TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
  live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
This commit is contained in:
Teknium
2026-08-25 22:27:43 -07:00
parent 73a9c34529
commit 8b6931393e
6 changed files with 259 additions and 16 deletions
+4 -2
View File
@@ -463,11 +463,13 @@ class CLIModelSwitchMixin:
if not getattr(result, "success", False):
return True
try:
from hermes_cli.model_selection_guards import combined_selection_warning
from hermes_cli.model_selection_guards import (
combined_selection_warning, selection_context_for_agent)
warning = combined_selection_warning(
result.new_model, provider=result.target_provider,
base_url=result.base_url or self.base_url or "",
api_key=result.api_key or self.api_key or "", model_info=result.model_info)
api_key=result.api_key or self.api_key or "", model_info=result.model_info,
selection_context=selection_context_for_agent(getattr(self, "agent", None)))
except Exception:
warning = None
if warning is None: