d762ed9b3c
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring check and gives it its own injection gate, independent of tool_use_enforcement, controlled by config.yaml `agent.execution_guidance` (auto/true/false/list — same semantics as tool_use_enforcement). The "auto" list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm, minimax, mimo, and mistral. Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where competitors passed: financial math done in prose, no read-back after external writes, malformed identifiers "repaired", completeness claimed despite count mismatches. The discipline block existed but those models never received it. The block is extended with compact clauses distilled from that analysis: - external-write read-back (tool-call success is not task success; internal file edits already confirmed by the tool are not re-verified) - count reconciliation (declared totals/has_more are hard assertions) - literal preservation (never normalize identifiers that fail a stated format; lookup success does not validate a malformed token) - retry-differently (empty/partial/suspiciously narrow results get a broader retry before concluding) - completion gated on verification (done = every named acceptance criterion verified, never a plausible subset) The todo tool description now encourages enumeration-as-checklist for "all N items" tasks and gates completed status on verified work, never intent. Guidance is chosen once at session start keyed on model name, so the system prompt stays byte-stable for the life of a conversation. Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874 (MiMo), #53847 (GLM tool-calls-as-text stall). Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com> Co-authored-by: intelac <8803887+intelac@users.noreply.github.com> Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com> Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>