139396995a
OpenCode pins requests sharing an x-opencode-session value to one upstream backend, which is what keeps its prompt cache warm across a conversation. Hermes never sent it, so cache ratios on OpenCode traffic were poor. - agent/opencode_affinity.py: single owner of the header — target detection (built-in zen/go/free, custom opencode-* providers, any opencode.ai URL) and the key (affinity scope → conversation root → session id, cron timestamp stripped), same resolution as OpenRouter/xAI affinity hints. - build_api_kwargs: merged once after the per-mode builder, so chat_completions, codex_responses and anthropic_messages all carry it. - auxiliary _build_call_kwargs: same key from the runtime-main session so compression/title/vision calls stay on the conversation's backend; the aux Codex and Anthropic adapters now forward extra_headers. Closes #81584, #81832 (deepseek-v4-flash 400 without the header).