fix(gateway): dedupe user turns on transient failure (#47237)

When the gateway persists a user message after a transient provider
failure (429/timeout/auth error), subsequent retries of the same
Telegram message could stack duplicate user turns in the transcript,
causing the agent to fall behind by 1-2 messages.

Add has_platform_message_id() to SessionDB (using the existing
idx_messages_platform_msg_id partial index) and a SessionStore wrapper.
The gateway's transient-failure path checks this before
append_to_transcript -- if the platform_message_id is already
persisted, the duplicate write is skipped.

Salvaged from #47869 by @davidgut1982. Adapted to current main which
has additional append sites and an existing content-based dedupe in
the exception handler path.

Closes #47237
This commit is contained in:
davidgut1982
2026-06-25 12:09:17 +05:30
committed by kshitijk4poor
parent d6cf383d74
commit 6208d6b3be
5 changed files with 167 additions and 4 deletions
@@ -65,6 +65,9 @@ def _bootstrap(monkeypatch, tmp_path):
)
runner.session_store.load_transcript.return_value = []
runner.session_store.append_to_transcript = MagicMock()
# Mock has_platform_message_id to return False so the dedupe guard
# (#47237) in gateway/run.py does not skip the append_to_transcript call.
runner.session_store.has_platform_message_id.return_value = False
runner.session_store.update_session = MagicMock()
monkeypatch.setattr(gateway_run, "_hermes_home", tmp_path)