Commit Graph

163 Commits

Author SHA1 Message Date
Teknium 99f7d2b4df fix(honcho): isolate recall generations and reject unscoped fallback 2026-09-07 08:12:23 -07:00
funky-xamarin 707267095a feat(honcho): add bounded opt-in current-query recall 2026-09-07 08:12:23 -07:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium a0be177aac fix(compat): pointers resolve to the object that MOVED, not a same-named stranger; stdin checker binds stdin= to the splatted definition
Review findings on #102117 (independent reviewer + itsflownium):

* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
  board= parameter). The compat generator ranked candidate homes by path proximity when a name is
  defined in several modules. Now it requires shape compatibility with the BASE definition (same
  literal for constants, superset of parameter names for defs) and prefers the facade's own
  <stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
  and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
  Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
  plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
  the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
  test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
  and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
  and requires stdin inside that expression/body.

Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
2026-09-03 22:00:01 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium ecf760db96 simplify(compat): plugins/memory — drop 15 re-exports + 1 alias, repoint 13 test callers
hindsight: drop _PORT_HEALTH_GRACE_ENV/_sanitize_bank_segment facade re-exports (tests -> .embedded/.settings).
honcho: drop 5 tool-schema re-exports, _credential_fingerprint (-> client_cache), _is_auth_error (-> session_auth),
and the _redact_tokens alias (callers renamed to redact_tokens).
openviking: drop 5 _setup re-exports; tests call openviking_module._setup.* directly.
2026-09-03 13:04:16 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 23b9ffc4fa fix(integration): restore subprocess stdin=DEVNULL / utf-8 encoding guards and windows-footgun gates dropped by round-3 compaction
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
2026-09-03 02:46:19 -07:00
Teknium 4b4502a429 refactor(honcho): merge pre-build and cached OAuth refresh hooks into one _refresh_oauth 2026-09-03 00:09:52 -07:00
Teknium a905beac5e refactor(honcho): unify cli peer/tokens show-or-set skeleton, fold status env-fallback and wizard int-parsing 2026-09-03 00:06:52 -07:00
Teknium 0096a458a6 refactor(honcho): tighten cli clone/inherit/prompt helpers 2026-09-02 22:47:02 -07:00
Teknium a72073fff3 refactor(honcho): collapse defensive layers in client build path 2026-09-02 22:36:29 -07:00
Teknium 9f4535b757 refactor(honcho): fold identity-mapping shape branches and status/choice printers in cli 2026-09-02 22:28:39 -07:00
Teknium 41bd98309d refactor(honcho): compact tool/config schema literals (dumps byte-identical) 2026-09-02 22:26:23 -07:00
Teknium def8c585a8 refactor(honcho): tighten provider tool handlers and session peer-config sync 2026-09-02 22:19:33 -07:00
Teknium bd45e5f7b4 refactor(honcho): unify guarded-session prologue in session_context; compact dialectic 2026-09-02 22:07:48 -07:00
Teknium cc67af75c4 refactor(honcho): compact oauth + oauth_flow; drop _atomic_write_config for utils.atomic_json_write 2026-09-02 22:07:48 -07:00
Teknium 58f5d23f78 refactor(honcho): compact cli setup/status/identity bodies (help bytes identical) 2026-09-02 21:48:08 -07:00
Teknium 33f883219a refactor(honcho): compact client/session modules (no behavior change) 2026-09-02 21:36:44 -07:00
Teknium 93ee1c0e38 refactor(honcho): compact cli helpers and provider bodies (no behavior change) 2026-09-02 21:20:55 -07:00
Teknium 64f430a051 refactor(plugins/memory): honcho — extract session context/auth/peers/migration, dialectic, client cache, tool schemas; unify _parse_* config helpers; table-driven CLI 2026-09-02 13:30:09 -07:00
kshitij 1226970001 fix: track and join honcho-memwrite thread in shutdown
on_memory_write spawns a fire-and-forget daemon thread that was never
stored on self, so shutdown() couldn't join it — the exact problem the
PR fixes for the async writer thread. Store as self._memwrite_thread
and include it in the shutdown join loop.

Review follow-up for salvaged PR #83500.
2026-08-13 23:43:15 +05:30
Erosika 9e77d83354 fix(honcho): gate memory-file migration on the declared owner
The previous gate compared session.user_peer_id against a fresh
_resolve_user_peer_id() call on the same manager. Both values come from
the same resolver with the same inputs, so a non-owner triggering a new
session in a shared channel passed the check and received the owner's
MEMORY.md/USER.md under their peer.

The owner is now a config fact: _declared_owner_peer_id() returns the
sanitized peerName, and migration runs only when the session's user peer
is that peer. Without a declared peerName, migration runs only when no
runtime gateway identity is present (the single-operator CLI path).
Aliases still work: a platform ID mapped onto peerName resolves to the
owner peer before the comparison.

Tests now derive each session's user peer from the real resolver instead
of hand-picking mismatched ids, so the non-owner test fails against the
old gate.
2026-08-13 23:43:15 +05:30
Erosika 27021f5f84 fix(honcho): resolve migration owner gate through _resolve_user_peer_id
The owner gate from #82038 compared against config.peer_name directly,
which is None for most single-user setups — sanitizing None would raise
and the gate never accounted for pinned/runtime/aliased identities.
Resolve the owner the same way sessions do, and add the non-owner skip
regression test the original PR shipped without.

Co-authored-by: menhguin <menhguin@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
carnie[bot] 7bad91c51d fix(honcho): skip memory-file migration on non-owner sessions (task #00000801)
migrate_memory_files() uploads USER.md/MEMORY.md with peer=user_peer — the
session's runtime user. In shared channels, a non-owner's new thread uploads
the owner's full profile under the NON-OWNER's peer; Honcho's deriver then
attributes the owner's psychometrics/medical/biography to that person. This
was the root contamination vector (55/70 contaminated sessions carried the
payload). Skip migration unless the session user is the configured owner.
SOUL.md unaffected (uploads under assistant peer).

Co-authored-by: Minh Nguyen <menhguin@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
Erosika 756aa54b67 fix(honcho): honor writeFrequency in sync_turn by routing through manager.save()
sync_turn called manager._flush_session() directly, which flushes
synchronously every turn no matter what writeFrequency says — the
"async", "session", and every-N-turns modes were dead configuration
on the main turn path. Route through save(), the dispatcher that
actually implements those modes.

Same bug class reported in #19650 (starship-s) and #72708 (Diaspar4u);
this takes the minimal one-line routing fix without their broader
lifecycle refactors.

Co-authored-by: starship-s <45587122+starship-s@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
Erosika 9cfff1546d fix(honcho): join the session manager's async-writer thread on provider shutdown
Provider shutdown() only called manager.flush_all(), which drains the
queue but never joins the async-writer thread — manager.shutdown()
exists and nothing called it. The writer thread could still be blocked
in httpx I/O at interpreter exit (the #37632 crash class). Now
shutdown() calls manager.shutdown() (flush + join) when persistence is
enabled, and a new manager.stop_async_writer() (join only, no flush)
when saveMessages is false, so containment and clean teardown compose.
2026-08-13 23:43:15 +05:30
Erosika 08b3312031 fix(honcho): persist one-sided turns under the empty-content guard
The containment commit skipped the whole turn when either side was
empty, which would drop a real user message on interrupted or
tool-only turns. Keep the guard for fully-empty turns only and skip
empty sides individually inside the sync loop.
2026-08-13 23:43:15 +05:30
赵桂雄 d610b238c6 fix(honcho): extend saveMessages=false guard to shutdown() flush
Salvages #67559 — original gated sync_turn/on_memory_write/on_session_end but missed shutdown(), whose flush_all() still persisted on exit. hermes-sweeper review (salvageability=high) flagged this as the one gap.

Guard sits after the worker-thread joins, not at the top: cleanup is independent of persistence, and a top-of-method return would leak _prefetch_thread/_sync_thread. Adds TestShutdown and clarifies the saveMessages=false README row.

Credit @Matroskin86 (original PR author).
2026-08-13 23:43:15 +05:30
eapwrk 2042b3122b honcho: honor saveMessages=false across all automatic write paths
The saveMessages knob has been parsed by HonchoClientConfig since its
introduction but was never consumed: sync_turn, on_memory_write and
on_session_end persisted to Honcho regardless. With saveMessages=false the
provider now never writes automatically (raw turns, memory-write conclusion
mirroring, session-end flush) while read/tools paths stay fully functional.
Guard uses getattr with a True default so legacy/injected configs keep the
old behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018giroL5zeMPnPxERxAxXHY
2026-08-13 23:43:15 +05:30
Dillon Townsel ffdc0b0be6 fix(honcho): enforce saveMessages write containment + reject gateway-internal turns 2026-08-13 23:43:15 +05:30
Erosika c28c434706 fix(honcho): surface honcho_reasoning backend failures instead of 'No result'
dialectic_query collapsed every backend failure to an empty string, so
the explicit honcho_reasoning tool rendered timeouts, server errors,
and genuinely-empty answers identically as 'No result from Honcho.'
(#36098 issue 4). Operators debugging 'search works but reasoning does
not' were sent down representation/observation rabbit holes when the
real cause was a 30s timeout on a medium-reasoning dialectic call.

Add raise_errors to dialectic_query (default false — automatic
injection keeps its fail-quiet behavior and cadence backoff) and pass
it from the explicit tool call, returning a tool error that names the
failure and points at the timeout knob. Auth errors keep their
dedicated handler.
2026-08-13 23:43:15 +05:30
Erosika 606481586a fix(honcho): honor explicit top-level apiKey on local base_urls; warn on keyless profile host blocks
Two silent-auth-failure paths from #36098 (also #66125):

- the local-URL guard only escaped the 'local' placeholder when the
  HOST BLOCK had apiKey. A top-level apiKey in honcho.json — explicit
  user intent, and what 'hermes honcho setup' writes for single-host
  configs — was dropped on the floor, so AUTH_USE_AUTH self-hosts
  401'd on every request. Now any explicit key in honcho.json (host
  block or top level) is honored; only env-sourced keys are still
  treated as likely-cloud and skipped for local URLs.

- named-profile host blocks do not inherit the default host's apiKey
  (credential isolation is by design), but the failure was silent:
  the profile ran unauthenticated and every tool said 'no context'.
  Affirm isolation and warn loudly at config-resolution time instead,
  the outcome #66125 proposed if inheritance is rejected.
2026-08-13 23:43:15 +05:30
spfcraze 32238f9942 fix(honcho): resolve peers host keys via profile_host_key (underscore form) (#76414)
_all_profile_host_configs() built per-profile host keys inline as
f"{HOST}.{profile}" ("hermes.work") while profile_host_key() — used by
honcho status/enable/sync and the runtime memory plugin — produces the
underscore form ("hermes_work"). The lookup always missed, so
'hermes honcho peers' showed "(not set)" / leaked the raw malformed key
into the AI-peer column for every non-default profile. Profile names
needing sanitization (dots/spaces) were doubly broken.

Verified live: with hosts["hermes_work"] populated, cmd_peers showed
'work ... hermes.work' before the fix and 'work ... hermes' after.

Tests: host keys match the writer form, sanitized profile names resolve,
peers output shows populated identities with no key leak, and clean
fallback for profiles without a block.
2026-08-13 23:43:15 +05:30
Bartok9 41d77caf11 fix(honcho): drop non-printable base_url values before client init
Salvage of #2757 by @teyrebaz33 — rebased onto current Honcho plugin layout.

Stray control characters (e.g. terminal escapes pasted into HONCHO_BASE_URL
or config baseUrl) are dropped with a warning so SDK construction cannot
crash startup on Invalid non-printable ASCII character errors.
2026-08-13 23:43:15 +05:30
mohamedorigami-jpg c5f6f58d66 fix(honcho): use _host_block helper for dot-form legacy host key fallback (fixes #37436)
_resolve_or_create_client() used a plain dict.get(config.host) that
fails for dot-form profile host keys (e.g. "hermes.profile_a") even
though the _host_block() helper defined nearby handles the legacy
dot-form → underscore-form fallback correctly. The result:
_host_has_key evaluates to False for every authenticating user,
so effective_api_key is set to "local" and every Honcho API call
returns 401 Invalid JWT — cascade failure into silent data loss
for cross-peer queries and message sync.

Fixes by calling the existing _host_block() helper instead of
reimplementing the direct lookup. Local variable renamed from
_host_block → _host_block_local to avoid shadowing the function.

Closes #37436
2026-08-13 23:43:15 +05:30
LeonSGP43 a97d6747f3 fix(honcho): honor host-specific baseUrl 2026-08-13 23:43:15 +05:30
Rob Sherman ad588542ea fix(memory): read endpoint.baseUrl from Honcho config; accept HONCHO_URL
HonchoClientConfig.from_global_config() only consulted top-level
baseUrl / base_url / HONCHO_BASE_URL in ~/.honcho/config.json. The
Honcho SDK's native config format — and what Claude Desktop writes —
nests the URL at endpoint.baseUrl. Users with that config format had
their self-hosted Honcho container silently ignored: every honcho_*
call routed to https://api.honcho.dev with a workspace_id that does not
exist there, so tools returned empty data with no error anywhere.

Resolution order in from_global_config(), highest first:
  1. endpoint.baseUrl    (SDK-native, what Claude Desktop writes)
  2. baseUrl / base_url  (root-level, existing behavior)
  3. HONCHO_BASE_URL     (existing env var)
  4. HONCHO_URL          (the SDK's own env var, honcho/client.py:234)

HONCHO_URL is also read in from_env(). from_global_config() delegates to
from_env() whenever the config file is missing or unreadable, so an env
fallback wired into only one of the two would silently do nothing for
users with no config file.

A non-dict endpoint value falls through cleanly rather than raising.
Existing users are unaffected — the new sources are consulted only when
the existing ones resolve to None.

The INFO log for the base_url-unset case now says so explicitly instead
of printing only the host. The SDK resolves that case from its own
ENVIRONMENTS map (honcho/client.py:36-39), which for environment=
production means the public cloud; a self-hosted user whose config was
not picked up otherwise sees a healthy-looking startup line.

Closes #43800.
2026-08-13 23:43:15 +05:30
kshitij 5118692c25 fix: replace double-lambda with functools.partial, close from_env config_path gap
- _submit_background and _prefetch_provider: replace unreadable
  (lambda inner: (lambda: ctx.run(inner)))(fn) with functools.partial(ctx.run, fn)
- from_env(): set config_path=resolve_config_path() so bound_config_path()
  doesn't re-resolve from ContextVar on daemon threads (the exact bug
  the PR fixes for from_global_config)

Review follow-ups for salvaged PR #83525.
2026-08-13 23:43:15 +05:30
Erosika 3a7d29a8ad fix(honcho): drop unread _client_slot_timeouts bookkeeping
The dict was written on every build and popped/cleared on eviction and
reset, but no read site remained — timeout staleness detection moved
into the cache key itself (a timeout change produces a new identity and
_slot_for evicts the old slot), which the isolation tests already pin.
Flagged in review by @spfcraze.
2026-08-13 23:43:15 +05:30
Erosika 671f9cbafa fix(honcho): propagate contextvars to all plugin background threads
Profile isolation is a ContextVar; plain threading.Thread targets start
with an empty context, so the plugin's nine daemon threads (session
init, prewarm, first-turn base/prefetch, prefetch, sync, memwrite,
async writer, context prefetch) resolved ambient state — config path,
active host, hermes home, oauth token paths — against the DEFAULT
profile whenever they ran under a routed profile's turn.

Adds spawn_context_thread(), which copies the caller's context at spawn
time so the thread sees the profile scope it was created under, and
routes every plugin thread spawn through it. Defense-in-depth under the
bound-config work: even ambient resolution on these threads now lands
on the right profile.

The copy_context approach follows the gateway's own
_run_in_executor_with_context pattern; #81401 applied it to the init
thread, this extends it to all nine spawns.

Co-authored-by: angel12 <angel12@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
Erosika 696470ce8a fix(honcho): cache clients per identity with a rotation-stable credential fingerprint
Replaces the process-wide first-config-wins client singleton with a
per-identity slot map. The singleton baked the first profile's
workspace_id and bearer into one shared client, so in multi-profile
processes (gateway multiplexer, dashboard, cron) every profile's
memory landed in whichever workspace initialized first — cross-tenant
bleed with no error (#69123, #74065).

cache key: (host, workspace, base_url, environment, provenance paths,
effective timeout, credential fingerprint). the fingerprint hashes the
OAuth REFRESH token (stable across in-place access-token rotation,
changes on re-auth/account switch) or the static api key — so
re-running 'hermes honcho setup' to switch accounts produces a new
identity instead of silently reusing the old account's client and
writing tenant B's data with tenant A's bearer, a hole per-path keys
alone cannot close.

same-identity slots with a different fingerprint or timeout are
EVICTED on replacement, so credential churn can't accumulate pinned
clients — the replaced client's pools close when its last holder
drops. timeout changes rebuild via the key (the old explicit staleness
check is subsumed). failed in-place OAuth rotation resets only the
client's own slot. reset_honcho_client() clears everything, preserving
test and oauth-flow re-login semantics.

per-config-identity caching was first proposed in #69142; the
provenance-key shape follows #81401. this implementation adds the
credential fingerprint and eviction they lacked.

Co-authored-by: NaMinhyeok <NaMinhyeok@users.noreply.github.com>
Co-authored-by: angel12 <angel12@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
Erosika 8b7d0ed4c7 fix(honcho): bind config provenance so background threads stop resolving the wrong profile
Profile isolation in every multi-profile process (gateway multiplexer,
dashboard, cron) is a ContextVar (set_hermes_home_override) that
threading.Thread targets cannot see. The plugin's daemon threads —
async writer, prefetch, sync, first-turn, init — all funnel through
HonchoSessionManager.honcho, which called get_honcho_client() with NO
config, re-resolving resolve_config_path()/resolve_active_host() from
the ContextVar-blind thread context: every background memory access
landed on the DEFAULT profile. Worse, the OAuth paths did the same, so
a token refresh on a daemon thread could persist the rotated token
into the wrong profile's honcho.json, and a 401 recovery could burn
the wrong profile's single-use refresh token.

- HonchoClientConfig gains provenance (config_path, hermes_home)
  captured at resolution time inside the caller's profile scope, with
  bound_config_path() for consumers
- manager.honcho passes the bound config instead of re-resolving
- OAuth paths (_apply_fresh_oauth_token, _refresh_cached_oauth,
  _reauth_required, _force_reauth) use the bound path
- the honcho.json timeout memo becomes path-keyed instead of
  single-slot, so multi-profile processes stop thrashing it and
  returning profile A's timeout for profile B

Groundwork for per-identity client caching (#69123, #74065); the
provenance-field shape follows #81401.

Co-authored-by: angel12 <angel12@users.noreply.github.com>
2026-08-13 23:43:15 +05:30
kshitij 520a1e7812 fix(honcho): keep _pop_auth_notice tolerant of minimal fake managers; make fast-path test binding
Gap-fill from the follow-up commit's own review:

- __init__.py: restore getattr tolerance in _pop_auth_notice — test
  fixtures outside tests/honcho_plugin/ install minimal fake managers
  without pop_auth_notice (tests/test_honcho_startup_fail_open.py's
  SlowManager failed with AttributeError). Exceptions still propagate;
  only the blanket except was dropped.
- test_auth_recovery.py: the fast-path test used a raising stub, but
  _reauth_required swallows all exceptions — the test passed even with
  the fast path removed. Rewritten as a recording spy with a call-count
  assertion; mutation-verified (removing the fast path now fails it).
- test_auth_recovery.py: autouse fixture resetting oauth module dicts
  (_dead_grants, _refresh_failure_at, _reauth_check_cache,
  _expiry_cache) so state can't leak between tests.

honcho_plugin 293 + test_honcho_startup_fail_open 7 + plugins/memory
285 = 585 passed.
2026-08-08 14:40:46 +05:30
kshitij edfe4f5136 refactor(honcho): dedupe refresh-failure handling; harden exchange budget, dogpile cooldown, and rebuild race
Follow-ups from review of #80590:

- oauth.py: extract _rotate_and_persist() — the twin ~18-line
  OAuthRefreshError permanent/transient handling blocks in
  ensure_fresh_token and force_refresh_token were byte-identical
  except the log verb.
- oauth.py: cap the exchange cycle at _REFRESH_TOTAL_BUDGET_SECONDS
  (20s). The retry runs while holding the global refresh locks on the
  path to a memory call; a timed-out first attempt no longer earns a
  second full 15s exchange (~32s lock hold -> <=20s).
- oauth.py: transient-failure cooldown (_refresh_failure_at, 30s).
  Waiting threads and later turns fail open to the stale token instead
  of serializing their own full exchange cycles against an endpoint
  that just failed. Cleared on successful rotation and re-login.
- oauth.py: mtime-gate reauth_required()'s config read — the dead-grant
  state persists until re-login, and the verdict can only change when
  the config file is rewritten; drop the per-call read+parse.
- oauth.py: derive _TOKEN_VALUE_RE from ACCESS_TOKEN_PREFIX /
  REFRESH_TOKEN_PREFIX so a prefix change can't silently break
  redaction; promote redact_tokens to public (session.py imported the
  private name).
- session.py: fast path in _reauth_required — skip config-path
  resolution entirely while no grant is dead (runs before every SDK
  call).
- session.py: client-generation counter closes the fetch/store race in
  _sdk_session/_get_or_create_peer — an object resolved from the old
  client mid-rebuild is no longer cached (it would 401 forever and burn
  a token rotation per retry).
- __init__.py: drop the getattr/callable/except triple-guard in
  _pop_auth_notice; the manager is always None or HonchoSessionManager.

7 new tests (budget, cooldown x3, generation guard, fast path); all
mutation-checked (disabling each guard fails its test). honcho_plugin
293 passed; plugins/memory 285 passed; live E2E against a real HTTP
token endpoint re-verified.
2026-08-08 14:40:46 +05:30
Erosika 086dc8b880 fix(honcho): surface the auth notice when session init itself fails
An init-time HonchoAuthError discarded the manager that recorded it, so
context/hybrid prefetch returned nothing and tools mode returned the
generic init error. The provider now keeps the failure detail across the
manager discard, prefetch emits the one-time notice at the readiness
guard, tools mode returns an explicit authentication error, and a
successful re-login retry clears the stored failure. Non-auth init
failures keep failing open with no notice.
2026-08-08 14:40:46 +05:30
Erosika 864035b241 fix(honcho): route every authenticated sdk call through one 401-recovery helper
_authed_call checks the dead-grant marker before calling, retries a
confirmed auth failure once after a forced refresh, and records the
failure for the one-time notice. Operations re-resolve their peer and
session objects inside the call, so a retry after a client rebuild no
longer reuses objects bound to the old transport. Tool handlers now
return an explicit auth error instead of an empty result, and non-auth
failures keep their fail-open behavior.
2026-08-08 14:40:46 +05:30
Erosika da1f8779ef style(honcho): trim auth recovery comments to one line each 2026-08-08 14:40:46 +05:30
Erosika b1414baa09 fix(honcho): stop classifying bare '401' digits as auth errors; redact session-side auth logs
_is_auth_error matched the substring '401' anywhere in an error string,
so a latency figure ('retry after 4010 ms'), a request id, or a
workspace name containing those digits classified as an auth failure.
A false positive calls _force_reauth, which runs a real token exchange;
the server rotates the refresh token on every exchange, and a lost
rotation response leaves Hermes holding a superseded token whose later
replay revokes the whole grant — the exact wedge this branch fixes.

The status attribute check (SDK AuthenticationError carries status=401)
does the real work and stays first. A concrete non-401 status now wins
over ambiguous text. The text fallback keeps only specific markers:
'invalid or expired access token', 'authentication failed' (not bare
'authentication', which also matches auth-infrastructure outage
messages), 'unauthorized', and '401' only with HTTP context ('HTTP
401', 'status 401'), never as a bare number. The classifier is biased
toward false negatives: a missed auth error costs one un-recovered
call, a false positive spends a rotation.

Also redacts token values in _record_auth_failure, _auth_error_message,
and the two retry warnings, matching oauth.py. The SDK's auth errors
carry no token values today, but this is the one credential path where
an upstream regression would leak silently.

Tests: the four false-positive strings stay non-auth, HTTP-context 401s
still match, a concrete 429 status beats 'authentication failed' text,
and the recorded failure plus notice redact token values.
2026-08-08 14:40:46 +05:30
Erosika ecfc427b28 fix(honcho): skip memory calls while the oauth grant is dead
reauth_required() existed but nothing called it, so after a grant died
every dialectic fire and sync flush still sent a Honcho API call that
401ed. dialectic_query and _flush_session now check the dead-grant flag
first and skip the call: dialectic raises HonchoAuthError (exempt from
cadence backoff), sync returns False with the failure recorded so the
one-time notice still fires.

The check compares the on-disk refresh-token digest, so a re-login flips
it back with no network call and the next cadence resumes immediately.
Transient auth errors keep the existing force-refresh-and-retry path.

Four new tests: a dead grant issues no dialectic or sync call, and a
re-login resumes both without waiting.
2026-08-08 14:40:46 +05:30