Files
hermes-agent/tests/tools/test_mcp_result_size_limit.py
T
Teknium 09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00

70 lines
2.9 KiB
Python

"""Regression tests for the MCP hard result cap (#56059).
MCP tool results had no allocation bound — a buggy or malicious MCP server
could return multi-megabyte text that floods memory and context before the
budget/spillover layer sees it. The hard cap truncates only pathological
payloads (over 2M chars by default) with a 40% head / 60% tail split;
ordinary large results pass through untouched so the 50K MCP spillover
threshold (tools/budget_config.py) can preserve them in full on disk.
Test shape adapted from PR #56511 (Tranquil-Flow); cap semantics differ —
see _MCP_HARD_RESULT_CAP_CHARS in tools/mcp_tool.py.
"""
from __future__ import annotations
from tools.mcp_tool import _MCP_HARD_RESULT_CAP_CHARS, _truncate_mcp_text_result
class TestTruncateMcpTextResult:
def test_short_result_unchanged(self):
text = "x" * 100
assert _truncate_mcp_text_result(text) == text
def test_exact_limit_unchanged(self):
text = "y" * 100
assert _truncate_mcp_text_result(text, max_chars=100) == text
def test_spillover_sized_result_passes_untouched(self):
"""A 60K result (over the 50K spillover threshold) is NOT truncated
here — the budget layer must receive it intact so spillover can
preserve the full payload on disk."""
text = "z" * 60_000
assert _truncate_mcp_text_result(text) == text
def test_pathological_result_is_truncated(self):
text = "z" * (_MCP_HARD_RESULT_CAP_CHARS + 500_000)
result = _truncate_mcp_text_result(text)
assert len(result) < len(text)
assert "TRUNCATED" in result
def test_truncation_preserves_head_and_tail(self):
head_marker = "HEAD_MARKER_START"
tail_marker = "TAIL_MARKER_END"
text = head_marker + "x" * 5000 + tail_marker
result = _truncate_mcp_text_result(text, max_chars=200)
assert result.startswith(head_marker)
assert result.endswith(tail_marker)
def test_truncation_includes_omitted_count(self):
text = "a" * 5000
result = _truncate_mcp_text_result(text, max_chars=100)
assert "4,900" in result # 5000 - 100 omitted
assert "5,000" in result # total original length
def test_truncation_uses_40_60_head_tail_split(self):
text = "H" * 40 + "M" * 5000 + "T" * 60
result = _truncate_mcp_text_result(text, max_chars=100)
assert result[:40] == "H" * 40
assert result[-60:] == "T" * 60
def test_empty_result_unchanged(self):
assert _truncate_mcp_text_result("") == ""
def test_hard_cap_sits_above_spillover_threshold(self):
"""The hard cap must stay far above the MCP spillover threshold so
spillover, not lossy truncation, handles ordinary large results."""
from tools.budget_config import DEFAULT_MCP_RESULT_SIZE_CHARS
assert _MCP_HARD_RESULT_CAP_CHARS > DEFAULT_MCP_RESULT_SIZE_CHARS * 10