fix: print transform_llm_output appended content after CLI streaming

When the CLI streams a response token-by-token, it marks the response as
already displayed and skips re-printing after the tool loop. This means any
content appended by a transform_llm_output plugin fires after streaming — the
appended text is in the final response and stored in history, but never shown
to the user.

Fix by tracking the pre-transform response in finalize_turn() and including it
in the result dict as pre_transform_response. The CLI then checks whether the
response was transformed and, if so, prints only the appended suffix.

Previously the already_streamed branch was a no-op pass. Now it detects
post-stream plugin additions and outputs them without re-printing the streamed
body.
This commit is contained in:
Ken Weiner
2026-07-02 17:35:03 +00:00
committed by kshitij
parent a4af262638
commit 367dda813c
2 changed files with 10 additions and 1 deletions
+7 -1
View File
@@ -14445,7 +14445,13 @@ class HermesCLI(CLIAgentSetupMixin, CLICommandsMixin, CLIBillingMixin):
elif already_streamed:
# Response was already streamed token-by-token with box framing;
# _flush_stream() already closed the box. Skip Rich Panel.
pass
# If a plugin appended content (shout, trace) after streaming,
# print only the appended portion now.
if result and result.get("response_transformed"):
_pre = result.get("pre_transform_response") or ""
_appended = response[len(_pre):]
if _appended.strip():
_cprint(_appended)
else:
_chat_console = ChatConsole()
_chat_console.print(Panel(