Files
hermes-agent/tests/run_agent/test_gemini_native_reservation.py
T
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00

97 lines
3.3 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Native Gemini output-token reservation at agent init (#57275 claim 4).
When ``model.max_tokens`` is unset, the native generateContent adapter still
sends ``maxOutputTokens=65,535`` (GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) — Gemini
treats an omitted cap as a low internal default, not "unlimited", so the
adapter always sends an explicit cap. The compressor's trigger is
``pct × (window − max_tokens)``; constructing it with ``max_tokens=None``
reserved 0 while the wire reserved 65,535: on a 128K window the trigger
landed at ~96K against a real safe input budget of ~65K, and the provider
400'd before compaction ever fired.
These tests assert the compressor's reservation mirrors the adapter default
on the native Gemini path, and ONLY there.
"""
from unittest.mock import patch
import agent.context_compressor as cc_mod
from agent.gemini_native_adapter import GEMINI_DEFAULT_MAX_OUTPUT_TOKENS
CFG = {"agent": {}}
def _build_agent(model, base_url, provider="", max_tokens=None, window=131072):
with (
patch("model_tools.get_tool_definitions", return_value=[]),
patch("model_tools.check_toolset_requirements", return_value={}),
patch("agent.process_bootstrap.OpenAI"),
patch("hermes_cli.config.load_config", return_value=CFG),
patch("hermes_cli.config.load_config_readonly", return_value=CFG),
patch(
"agent.model_metadata.get_model_context_length", return_value=window,
),
patch.object(cc_mod, "get_model_context_length", return_value=window),
):
from run_agent import AIAgent
return AIAgent(
model=model,
api_key="test-key-1234567890",
base_url=base_url,
provider=provider,
max_tokens=max_tokens,
quiet_mode=True,
skip_context_files=True,
skip_memory=True,
)
def test_native_gemini_unset_max_tokens_reserves_adapter_default():
agent = _build_agent(
"gemma-3-27b-it",
"https://generativelanguage.googleapis.com/v1beta",
)
cc = agent.context_compressor
assert cc.max_tokens == GEMINI_DEFAULT_MAX_OUTPUT_TOKENS
# Trigger must sit at/below the real safe input budget the wire leaves.
assert cc.threshold_tokens <= cc.context_length - GEMINI_DEFAULT_MAX_OUTPUT_TOKENS
def test_gemini_provider_name_also_reserves_default():
agent = _build_agent(
"gemini-3.7-flash", "https://example-proxy.invalid/v1", provider="google",
)
assert agent.context_compressor.max_tokens == GEMINI_DEFAULT_MAX_OUTPUT_TOKENS
def test_explicit_max_tokens_wins_over_adapter_default():
agent = _build_agent(
"gemma-3-27b-it",
"https://generativelanguage.googleapis.com/v1beta",
max_tokens=8192,
)
assert agent.context_compressor.max_tokens == 8192
def test_non_gemini_paths_keep_no_reservation():
agent = _build_agent(
"openai/gpt-4.1", "https://openrouter.ai/api/v1",
)
assert agent.context_compressor.max_tokens is None
def test_gemini_openai_compat_endpoint_not_treated_as_native():
# The /openai compatibility endpoint does not use the native adapter.
agent = _build_agent(
"gemma-3-27b-it",
"https://generativelanguage.googleapis.com/v1beta/openai",
)
assert agent.context_compressor.max_tokens is None