4 Commits

Author SHA1 Message Date
fangliquan 3755dca7d6 fix(agent): handle usage-less empty responses 2026-09-03 03:41:52 -07:00
Teknium c7727540f3 test: strengthen empty-response guard tests (salvage follow-up for #75115) 2026-08-16 22:06:08 -07:00
Shannon Sands d10f87245e refactor(agent): move empty-response guard settings from env vars to config.yaml
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:

  agent:
    empty_response_guard:
      enabled: true            # false = legacy fixed 3-retry behaviour
      cost_threshold_usd: 0.25 # per-attempt cost that halves the budget

- hermes_cli/config_defaults.py: new documented subsection under agent
  (additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
  section to (enabled, threshold) with fail-open tolerance for
  malformed values; guard_enabled()/_cost_threshold_usd() now read the
  init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
  agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
  following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
  covering malformed sections, YAML string booleans, bad thresholds,
  and a DEFAULT_CONFIG sync check; new integration test proving
  enabled:false restores the legacy 1+3-call behaviour.

Requested by isak-ialogics on PR #75115.
2026-08-16 22:06:08 -07:00
Shannon Sands ac06c2ff8b fix(agent): stop re-billing deterministic empty responses (NS-503)
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).

Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.

New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:

- Deterministic-empty detection: two consecutive empty attempts with
  usage present, output_tokens == 0 (reasoning tokens count as output),
  and identical (model, provider, finish_reason) skip the remaining
  retries and go straight to the fallback chain — a different model may
  well answer. Missing usage, nonzero output, or any signature change
  keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
  exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
  empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
  included/subscription routes are untouched.

At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.

Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.

Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.

Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.

Refs NS-503.
2026-08-16 22:06:08 -07:00