fix(deepseek): recognise the version-less deepseek-flash model id

DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
This commit is contained in:
YipTszkwan
2026-09-10 13:41:28 +08:00
committed by Teknium
parent 5172f22df2
commit 8435a3ae00
8 changed files with 52 additions and 5 deletions
@@ -108,6 +108,10 @@ class TestDeepSeekModelGating:
"deepseek-v4-flash",
"deepseek-v4-future-variant",
"DEEPSEEK-V4-PRO", # case-insensitive
# Version-less canonical ids (2026-09 Flash refresh) carry the
# same thinking-mode contract but no v<N> marker.
"deepseek-flash",
"DEEPSEEK-FLASH", # case-insensitive
],
)
def test_thinking_capable_models_emit_thinking(self, deepseek_profile, model):