Commit Graph

34251 Commits

Author SHA1 Message Date
kshitijk4poor b2ee58d24a fix(honcho): prune orphaned observation flags; share the local-platform set; drop test-shape defenses
A flush that rebuilds an evicted session's SDK session stores its observation
flags again, and the cap pass pruned every per-session dict but that one, so the
dict grew one entry per evicted-then-flushed session. The cap pass now prunes it
with the rest.

The peer-failure notice classified platforms with its own {"cli","tui","desktop"}
set, so an ACP session read the gateway wording ("do not suggest peerName"); it
now uses agent.coding_context.INTERACTIVE_CODING_PLATFORMS, which includes acp.

Three getattr/try-except guards existed only for test doubles (a bare __new__
provider, a SimpleNamespace config); the real types always carry the attribute.
Removed, and the two tests build real objects. An unset timeout resolves to the
client's 30s default, so the join-budget test expects 30, not the 5s floor.

Timing tests keep the >= 2s wall-clock bound the testing rules ask for.
2026-09-13 19:05:39 +05:30
kshitijk4poor eed52456f6 fix(honcho): an author peer joins with its session's synced observation flags
main's _join_observation_flags returned the manager-wide booleans with a comment
saying #103889 would plug the per-session flags in here; this is that plug. The
two main-side tests that stubbed the old two-tuple _get_or_create_honcho_session
return move to the three-tuple.
2026-09-13 19:05:39 +05:30
Erosika 58b117ed5b test(honcho): stub the three-tuple _get_or_create_honcho_session return
The session manager now returns observation flags as a third value.
The user-id scoping test still stubbed the old two-tuple and failed to unpack.
2026-09-13 19:05:39 +05:30
Erosika c04722102f fix(honcho): declare the injection block in config_schema so the desktop panel can pin sessionStart
The generic panel writes every field flat, and the plugin reads `injection` as one object, so a dotted `injection.sessionStart` field would never be read. Declare `injection` as a JSON field instead. Blank clears the pin.
2026-09-13 19:05:39 +05:30
Erosika 29113b24d8 fix(honcho): give the writer join and its drain the shutdown deadline
stop_async_writer bounded only the join and then drained the queue with no deadline, so a shutdown whose budget was already spent could still start uploads, and honcho-ai's add_messages has no per-call timeout. The join and the drain now share the shutdown deadline, no upload starts once it has passed, the writer skips its 2s retry once shutdown began, and shutdown logs one warning with the count left unsynced. The docstrings now say what the budget can do: stop new uploads and bound lock waits, while an upload already in flight runs to the client's HTTP timeout.
2026-09-13 19:05:39 +05:30
Erosika 3861167d63 fix(honcho): keep a failed late save reachable for flush_all after an eviction
A clean session can be evicted while its caller still holds it, and the caller's next save in turn mode flushed inline and never put the object back, so a failed upload left the batch nowhere flush_all() looks. Main reinserted after every flush. A failed save-time flush now reinserts the session when its key is free, or holds it in a retry list when a newer object owns the key, and flush_all() covers both; the collision path in _keep_until_flushed honors the flush result the same way.
2026-09-13 19:05:39 +05:30
Erosika 0b4df46255 fix(tui-gateway): stamp the login on the session record so rebuilds keep it
_session_auth_user_id read auth_identity off the transport slot at build time, but a second window replaces that slot with a FanoutTransport that has no identity, and branch and eager-resume builds run before the record exists, so those agents got user_id None. The record now carries auth_user_id from the creating transport, builds that run before the record pass it to _make_agent explicitly, and the compute host turn frame carries it too. A different login attaching to an existing session is logged once with both ids and the creator's id is kept.
2026-09-13 19:05:39 +05:30
Erosika 8523402db0 fix(honcho): bound the shutdown flush by the shutdown deadline
Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
2026-09-13 19:05:39 +05:30
Erosika 4b916022ab fix(honcho): surface the peer notice and audit the injection on the recall sync path
With recallSync on, prefetch popped only the auth notice and returned without writing the injection log. A session whose init failed for a missing user peer never told the model that memory was off, and the audit file stayed empty for every turn. The recall sync branch now pops the peer notice the way the async branch does and records each turn as injected or recall-sync-empty.
2026-09-13 19:05:39 +05:30
Erosika d532d83eca refactor(honcho): trim duplicated tests and long docstrings
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer
and deferred-save tests that differed only in their inputs. Share the blocking
remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the
added docstrings to the what and the one non-obvious why.
2026-09-13 19:05:39 +05:30
Erosika e9396ed8f4 fix(honcho): register the recall sync thread with its owner so shutdown joins it
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
2026-09-13 19:05:39 +05:30
Erosika beab8b6f27 fix(honcho): keep observation flags across a flush rebuild, never orphan an evicted session, namespace dashboard logins
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the
cached path returned none, so recall fell back to the config snapshot. Both paths now return and
store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or
flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the
writer lock, and the writer drains its queue after the join, so a put that raced shutdown is
written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining
budget instead of a fixed ten seconds.

The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop
passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a
gateway platform supplies no user id, the peer notice and tool error no longer recommend
peerName, which would merge every user of that gateway onto one peer. README documents
`injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution.
2026-09-13 19:05:39 +05:30
Erosika c284067d75 test(honcho): declare a peerName in the cache-bounds test
the manager under test had no peerName and no runtime identity. that
cohort now fails closed instead of minting a fallback peer, so the test
names an owner the way the auth-recovery tests do.
2026-09-13 19:05:39 +05:30
Erosika 73e9ecb7e7 test(honcho): drop the duplicated TestContextTokensForwarded class
the class arrived once with the adopted #92964 commit and once more from
the port; the second definition shadowed the first.
2026-09-13 19:05:39 +05:30
Erosika c5973cd540 feat(tui-gateway): build the agent with the authenticated dashboard user as user_id
the dashboard login already stamps {user_id, provider} on the websocket
at upgrade time and tui_gateway keeps it as WSTransport.auth_identity,
but _make_agent never read it. every dashboard and desktop session was
built with user_id=None, so memory providers saw no runtime user and fell
back to the configured peer, mixing all logins together (#89794).

_make_agent now reads the session transport's auth_identity and passes
its user_id to AIAgent, the kwarg gateway platforms already use. the
legacy ?token= path, stdio, and the server-internal credential the PTY
child connects with carry no human and pass None, so they keep resolving
to the configured peer. the change is provider-neutral: honcho and any
other memory provider receive the id through the existing initialize()
kwargs.
2026-09-13 19:05:39 +05:30
Erosika 3da6a80d50 fix(honcho): refuse to mint a user peer when no identity or peerName exists
a desktop or cli session with no peerName in honcho.json and no gateway
user id landed on a peer derived from the session key: user-default-<dir>
for per-directory sessions, user-<channel>-<chat> for keyed ones. every
directory got its own phantom peer, so the operator's turns and memory
never reached their real peer and the injected representation went stale
(#93326).

_resolve_user_peer_id now raises HonchoPeerUnresolvedError instead of
deriving a name. a peer is either the declared peerName or an identity
the transport supplied. the provider records the failure, tells the model
once that memory is off and which key to set, returns the same detail
from tool calls, and stops retrying init because a missing config key
does not heal mid-session. the memory-file migration gate loses its
"no owner and no runtime identity" branch: that cohort no longer has a
session to migrate into. whitespace-only peerName is treated as unset
rather than sanitized to "--".
2026-09-13 19:05:39 +05:30
Erosika cc3bfc2120 fix(honcho): read the sessionStart pin defensively in the first-turn formatter
tests and callers can build the provider without initialize(); the
formatter now treats a missing pin as unpinned instead of raising.
2026-09-13 19:05:39 +05:30
Erosika c3a5649aac fix(honcho): join every plugin thread on shutdown within one budget
provider shutdown joined the dialectic, sync and memwrite threads for 5s
each and then the async writer, but the session-init thread,
honcho-base-first and honcho-context-prefetch were never joined. any of
them still blocked in httpx when the interpreter finalized aborted the
process with SIGABRT 134 (#37632, #60616, #33485). the 5s join was also
shorter than the 30s http timeout a blocked call can hold (#33485).

spawn_context_thread now registers each thread under its owner (the
provider or the manager) in a weak registry, and shutdown joins every live
thread of both owners inside one deadline: at least 5s, or the configured
http timeout when longer. the manager refuses new prefetch threads once
shutdown began and flushes a late save() inline instead of respawning the
writer. threads that outlive the budget are named in a warning.

the sdk client exposes close() on its http pool. one client is shared by
every manager with the same identity in a gateway process, so a per-agent
shutdown cannot close it; close_honcho_clients() closes all pools and is
registered with atexit when the first client is built, the pattern the
hindsight and mem0 plugins use. follows #69070, #33701, #7627.
2026-09-13 19:05:39 +05:30
Erosika 0cb2977e84 fix(honcho): cap the manager caches and keep unsynced sessions out of eviction
the idle sweep from #71463 left three growth paths open. _peers_cache had no
bound at all, _session_observation (from #98941) grew one entry per session
id and kept orphans when an init failed after add_peers, and a burst of
distinct sessions inside one ttl window was not bounded. the sweep also
evicted sessions whose messages had not reached honcho yet, which in
"session" write mode drops the only copy, and _flush_session re-inserted an
evicted session into the cache with no observation flags, so recall for it
routed from the config snapshot instead of its server config.

_cache and _sessions_cache now cap at 128 entries and _peers_cache at 512,
evicting least recently used first (dict order, refreshed on every hit).
a session with unsynced messages is never evicted by the sweep or the cap.
_configure_session_peers returns the synced flags and get_or_create stores
them under _cache_lock next to the cache entry, so the observation dict can
hold no id the cache does not; eviction drops both. recall reads go through
_cached_session, which stamps updated_at, so a read-only session survives
the idle sweep. follows #71461, #71463, #98936.
2026-09-13 19:05:39 +05:30
Erosika ff29c27003 fix(honcho): hold a per-session lock across select, send and the _synced flip
_flush_session read the unsynced messages, posted them, and only then set
_synced. in a short-lived run the async writer draining save()'s queue and
the exit-time flush_all() both saw the same batch unsynced, so every turn
of a one-shot run landed in honcho twice (#92458).

each HonchoSession now carries an RLock that _flush_session holds around
the whole select, send, mark sequence. the second flusher enters after the
first marked the batch and sends nothing; different sessions still flush
in parallel. the lock lives on the session object instead of a manager
dict keyed by session id (#86094, #92787): it needs no eviction, and two
flushers of one message list can never hold different locks. tests
adapted from #86094 and #92787.
2026-09-13 19:05:39 +05:30
Aleksei Ivanov 280ac22663 fix(honcho): bound local session cache growth
HonchoSession.messages grew forever -- add_message() appended every
turn and _flush_session() only marked entries synced, never trimmed
them. HonchoSessionManager's four caches (_cache, _peers_cache,
_sessions_cache, _context_cache) had no eviction path besides an
explicit /new reset. A long-lived channel that is never manually
reset accumulates both for the gateway's entire uptime.

Honcho is the durable source of truth (get_or_create() already
re-fetches history from Honcho on a cache miss), so both bounds are
safe: evicting an idle local entry only costs one extra Honcho
round-trip next time that key is used.

- Trim already-synced messages beyond a retention window right after
  a successful flush; unsynced messages are never touched.
- Add a rate-limited idle-TTL sweep triggered opportunistically from
  get_or_create(), so no new background task/watcher wiring is
  needed.

Fixes #71461
2026-09-13 19:05:39 +05:30
686f6c61 2614b900b8 fix(honcho): omit reasoning-contaminated session summaries
Strip think blocks and drop planning-only Honcho summaries at fetch,
format, and cache so they cannot be reinjected as trusted memory.
2026-09-13 19:05:39 +05:30
liuhao1024 a4939af48b fix(memory): scope honcho observation flags per session (#98936)
The four observation booleans on HonchoSessionManager were manager-wide
mutable state initialized from per-session server configs: every session
setup overwrote them, so the last session to initialize retuned recall
routing for all other sessions the manager serves. Store the server-synced
flags under each session's own id instead; the manager-level fields stay
as the config snapshot and sessions that never synced fall back to them.
2026-09-13 19:05:39 +05:30
Erosika 5af0bf4111 test(honcho): pin sessionStart filtering and the injection audit; expose logging in setup
injection.sessionStart: unset renders every component in table order, an
empty list renders nothing, a pinned list renders only those names in table
order (not config order), a host block pin beats root, a non-list value is
treated as unset, and initialize() reads the pin.

logging: off by default, the logging key or HONCHO_LOGGING turns it on, a
host block can turn it back off, HONCHO_INJECTION_LOG overrides the path,
each record carries reason/turn/session_key/bytes/payload, an unwritable
path never raises, and a tools-mode prefetch logs its reason. the record
holds the user's representation verbatim, which is why the default stays off.

the two session-context tests that asserted exact context() kwargs now expect
tokens= as well (adopted from #92964). config_schema declares the logging
switch so hermes memory setup shows it.
2026-09-13 19:05:39 +05:30
Hector Suzanne 3b5d9116a7 fix(honcho): honour contextTokens cap on summary/peer context calls
Salvage of #70951 (Willkons / Alice-Willk-bot). Two of three
session.context() sites never forwarded the configured cap, so Honcho
always returned honcho_chat_summary_long.

Adds the missing get_session_context tokens= regression the sweeper
asked for on that PR.
2026-09-13 19:05:39 +05:30
Eugene Eisenstein 8744d7f7c8 repair injection.sessionStart 2026-09-13 19:05:39 +05:30
Eugene Eisenstein 61438268c6 fix(logging): repair the logging config key and HONCHO_LOGGING
The logging is also improved a little by saving reasons
2026-09-13 19:05:39 +05:30
kshitijk4poor 8e1a3b1d95 chore(contributors): map the three adopted authors' emails for the #103889 salvage 2026-09-13 19:05:39 +05:30
kshitijk4poor 135221e216 test(honcho): drop the assertion on the tolerant reader main folded into utils.read_json_or_empty 2026-09-13 19:05:28 +05:30
kshitijk4poor 8900fb2cf8 fix(honcho): adopt a sibling's on-disk rotation even inside our exchange cooldown
force_refresh_token gated the adopt-from-disk paths behind the failure cooldown.
After one of our exchanges failed transiently, a 401 within the next 30s returned
None even when a sibling process had already rotated and written a valid grant,
so the operation raised HonchoAuthError with a good token sitting on disk.
Adopting is a disk read, not an exchange; the cooldown exists to stop replaying a
single-use refresh token, so the two adopt checks now run before the gates.

_write_config parsed honcho.json twice under the lock and only wrapped the first
read into ConfigWriteRefused; _refuse_unparseable now returns the parsed dict and
the branches use it. The getattr/isinstance duck-typing collapses to one guard.
cli._read_config reuses oauth's tolerant reader (BOM-tolerant, like the strict
reader the write side uses) instead of its own utf-8 copy.
2026-09-13 19:05:28 +05:30
kshitijk4poor b0d3ec62da test(honcho): trim the strict-reader tests to one case per behavior
The unreadable/corrupt variants asserted the same two invariants (writers raise
and leave bytes untouched; rotation fails before the exchange) in five tests;
they are two parametrized tests now. The two lock-spy tests only asserted that
the implementation called the lock helper; the threaded save_config test is the
behavioral guard for the same property.
2026-09-13 19:05:28 +05:30
kshitijk4poor 24b33d42e0 fix(honcho): every honcho.json writer holds _refresh_lock and reads strictly
The CLI's _write_config took only the best-effort file lock, so an in-process
refresh thread could still interleave with a command's read-modify-write. The
dashboard's Honcho save (_write_provider_honcho) still seeded its whole-file
rewrite from a tolerant reader, so a honcho.json that exists but does not parse
was replaced by the active host's block alone from the UI - the same bug class
this PR closes on the CLI and refresh paths. `hermes profile create --clone`
swallowed the new ConfigWriteRefused as "plugin not installed".

One parametrized test covers the web writer for the corrupt and parseable cases.
2026-09-13 19:05:28 +05:30
Erosika 56231e51ad fix(honcho): advance the read baseline after each write so a revert reaches disk
_write_config applied a command's edits relative to the snapshot the read took, but never moved that snapshot after a successful write. A second write on the same object therefore compared A -> B -> A against A, saw no change, and left disk at B. After a write the snapshot and path now follow the caller's dict, so the next write applies only the edits made since.
2026-09-13 19:05:28 +05:30
Erosika 2f81f6a831 fix(honcho): keep a rotation on honcho.json when the read was seeded from another file
_read_config() falls back to ~/.honcho/config.json or the default profile when honcho.json does not exist, so cfg.path never matched the write path and _write_config() wrote the whole dict. That overwrote a refresh rotation a serve child had written onto the honcho.json the setup login created moments earlier. When the local file exists at write time, the seed snapshot is now overlaid with the local file and only the command's edits are applied onto it.
2026-09-13 19:05:28 +05:30
Erosika 36c98cb5eb fix(honcho): hold the refresh locks while save_config rewrites honcho.json
save_config read the file and wrote it back without the locks the token refresh holds around its own read, exchange and write. A refresh that landed between the two steps was overwritten and its consumed refresh token was gone. The read and the write now run under both locks.
2026-09-13 19:05:28 +05:30
Erosika be8ce602d0 fix(honcho): keep a rotation that lands while the setup wizard is still asking questions
install_grant writes the login grant to disk before the wizard's later prompts. _apply_grant_to_host wrote it only into the live cfg, so _write_config saw the grant as an edit and copied it over whatever a serve child rotated onto disk meanwhile. The snapshot now takes the grant too, so the final save leaves the newer on-disk grant alone.
2026-09-13 19:05:28 +05:30
Erosika 9aa2f77142 refactor(honcho): trim the oauth persistence tests to one case per behavior
Near-duplicate tests become one parametrized test each: the 401 adopt
cases, the invalid_grant race, the setup apikey answers, the clone and
enable credential checks, and the plain-dict writes. _point_cli_at
replaces the per-class monkeypatch helpers in test_cli. The removed-key
check folds into the merge test, and the missing-file bootstrap check
into the file-lock test; the command-level refusal test already drives
_write_config through ConfigWriteRefused, so the unit test for it goes.
The spy helpers in the install_grant lock test and the reauth bearer
test lose their duplicated closures. Every behavior the removed tests
asserted still has an assertion.
2026-09-13 19:05:28 +05:30
Erosika 498545bd23 refactor(honcho): trim the oauth persistence docstrings and helpers
_ReadConfig builds its snapshot in __init__; _read_config no longer
assigns the attributes after the fact. cmd_enable prints its hint
inline, since _no_credential_hint had one caller. The split
force_refresh_token signature and call fit on one line. Docstrings and
comments keep the what and the one non-obvious why. No behavior
changes; _apply_edits keeps its original body.
2026-09-13 19:05:28 +05:30
Erosika dd6b183c02 fix(honcho): only on-disk credentials count when a write enables a host block
clone and enable accepted HONCHO_API_KEY from the environment as proof the
block could authenticate and wrote enabled: true. the variable can be absent
from the next process, leaving an enabled block with nothing behind it, which
is the cohort the previous commit set out to remove.

_resolve_api_key takes env=False at both write sites, so a block is enabled
only when honcho.json itself holds a key, an oauth grant, or a self-hosted
baseUrl. status and setup keep the environment fallback for display.
2026-09-13 19:05:28 +05:30
Erosika 6966705751 fix(honcho): cli writes hold the refresh lock and merge only their edits onto disk
every hermes honcho command read honcho.json, changed a field, and wrote the
whole dict back with no lock. a token refresh in another process that landed
between the read and the write was overwritten with the old access and
refresh tokens, and the next refresh replayed a rotated single-use token.

_write_config now holds the same cross-process file lock the refresh path
holds, re-reads disk under it, and applies only the keys the command changed
since its _read_config(). untouched keys keep their on-disk value, so a
rotation survives; a credential the command set on purpose still wins. a
write with no prior read keeps today's whole-file behavior.
2026-09-13 19:05:28 +05:30
Erosika d3b599f374 docs(honcho): prune comments in the oauth persist commits
issue numbers move to the commit bodies; docstrings keep what and the one
non-obvious why.
2026-09-13 19:05:28 +05:30
Erosika 06e73dde64 fix(honcho): save_config refuses to rewrite a honcho.json that does not parse
hermes memory setup writes through HonchoMemoryProvider.save_config, which
merged the new values over {} when the existing file could not be parsed
and then replaced the file. same wipe as the oauth and cli write paths,
through a different door. it now reads through _read_config_strict and
raises, so the caller sees the error and the file stays as it was.
2026-09-13 19:05:28 +05:30
Erosika f87c915bf8 fix(honcho): write enabled: true only for a host block that can authenticate
a named profile cloned from a default profile that signed in with oauth
got a host block with enabled: true and nothing to authenticate with.
hosts.hermes holds the grant, its apiKey is not inherited (#66125), and
copying the oauth block would make two blocks replay one single-use
refresh token. 'hermes honcho enable' on a fresh profile wrote the same
shape. status then showed the profile as enabled while every honcho
call ran without memory and the plugin quietly stayed inactive.

clone_honcho_for_profile and cmd_enable now resolve a credential for the
target block (its own apiKey, the root apiKey, HONCHO_API_KEY, or a base
url) before writing enabled: true. _resolve_api_key takes the block to
check so both share one definition of "can authenticate".

cohorts:
  clone from an oauth default block, no root key: block written without
    enabled; the client's auto-enable rule turns it on once a credential
    appears (setup apikey writes the root key, or a per-profile login)
  clone from a default block with a host-level static key only: same
  clone with a root apiKey, an env key, or a base url: enabled as before
  enable on an empty or fresh block with no credential: refused, one
    message names the profile's setup command and the hosts.<name> key,
    nothing is written
  legacy blocks already on disk as enabled with no credential: nothing
    rewrites them; the plugin already treats them as unusable and stays
    inactive; enable now prints the same message instead of "already
    enabled"
  env-only key: counted as a credential at write time, as the client
    does at run time; if the variable later disappears the client still
    refuses to initialize the block
2026-09-13 19:05:28 +05:30
Erosika 3eb3dad994 fix(honcho): let the setup wizard's apikey answer replace a dead oauth grant
after the token endpoint revoked a grant (invalid_grant), running
'hermes honcho setup', choosing apikey and pasting a valid key changed
nothing. the wizard wrote the key to the root apiKey only. hosts.<name>
still held the dead access token under apiKey and the grant under oauth,
and the host block wins the lookup, so status kept reporting the revoked
grant and every call kept failing (#97990).

the apikey branch now drops the host's oauth block and writes the chosen
key onto the host block as well as the root. a dead access token is no
longer shown as the current key, so a blank answer with no other key
aborts instead of keeping the grant. a static host key without a grant
is kept as before.
2026-09-13 19:05:28 +05:30
Erosika 7113a6c18d fix(honcho): take the refresh locks around install_grant
a login finishing while a refresh was rotating the same honcho.json ran
its read-modify-write unlocked. the two writers could interleave: the
login read the file, the refresh persisted a rotated token, the login
wrote its own copy back and the rotated token was gone.

install_grant now holds _refresh_lock and _config_refresh_lock(path)
around the strict read, the config merge and the persist, the same way
force_refresh_token does. the token response is parsed before the locks
so a malformed grant never holds them.
2026-09-13 19:05:28 +05:30
Erosika bb06e1621e fix(honcho): refuse to rewrite a honcho.json that exists but does not parse
a truncated or hand-edited honcho.json reads as {} on the tolerant read
path. the next write then replaced the file with only the current host
block: an oauth refresh, a login, or any 'hermes honcho' command that
saves a setting wiped every other host and the root keys.

_read_config_strict now raises on a parse error the same way it raises
on a read error, and logs one sentence naming the file. _rotate_and_persist
treats that as a refresh failure and enters the cooldown without spending
the refresh token. install_grant raises into the setup flow. in cli.py
every write goes through _write_config, which runs the same check first
and raises ConfigWriteRefused; the honcho router, the setup wizard and the
profile sync print the sentence instead of a traceback. the wizard checks
before asking its questions. read paths keep the {} fallback.

follows #95860, which added the strict reader for unreadable files and
kept a .corrupt copy on a parse error. leaving the original file in place
keeps the same bytes without a second copy of the tokens on disk.
2026-09-13 19:05:28 +05:30
FestoneX 0bf9a9ff1b fix(honcho): a transient read failure is not an empty credential store
_persist_credential promises "leaving all else intact", but it seeded
its write from _read_config, which returns {} on ANY read failure, and
_atomic_write_config replaces the whole file via os.replace — which
needs only a writable parent, so a present-but-unreadable honcho.json
did not stop the overwrite. One OSError (EACCES after a root-owned
write, EIO, a stalled mount) during an automatic token refresh or a
fresh login therefore replaced the store with a single-host file,
destroying every other host's credentials and honcho.json's root
config. No user action is required to trigger the refresh path.

Same defect class as #75206 (P1, fixed for the core auth store in

Add _read_config_strict for the write paths: a missing file still
bootstraps as {}; an unreadable file raises with the store untouched;
genuine corruption still degrades but preserves a .corrupt copy first,
since a truncated store usually holds the other hosts' tokens verbatim.
_rotate_and_persist now takes its strict read BEFORE the exchange —
rotation is single-use, so an exchange whose result cannot be persisted
loses the grant — and threads the dict through to _persist_credential,
which also closes the re-read race between the locked read and the
persist. install_grant seeds its root-merge from the strict reader for
the same reason. Read paths keep their fail-open contract untouched;
both readers now use utf-8-sig so a BOM'd store is not misclassified
as corruption (the wipe vector needing no filesystem fault at all).

Adds TestPersistReadFailure: six tests, four of which fail against the
previous source; the rotate-ordering test additionally pins that no
exchange is attempted against an unreadable store.
2026-09-13 19:05:28 +05:30
Erosika 885e27fbf4 fix(honcho): snapshot the failing bearer before the operation runs
reading the live client's api_key inside _force_reauth races the in-place
rotation: a sibling waiter's apply_token_to_client (or the proactive
refresh entered via the honcho property) can swap the bearer before the
read, so force_refresh_token receives the already-rotated token, disk
matches it, and the adopt branch never fires — reproducing the very
force-refresh burst this PR removes.

_authed_call now snapshots the bearer before invoking the operation and
passes that exact token through to force_refresh_token.
2026-09-13 19:05:28 +05:30
Erosika 2250f1d1f3 fix(honcho): adopt on-disk oauth grant after a 401 instead of re-exchanging
desktop spawns multiple serve processes that share one rotating honcho
refresh token. force_refresh_token treated every 401 as "rotate now" even
when a sibling had already persisted a new grant, which can replay a
single-use refresh token and revoke the whole grant.

re-read honcho.json under the existing file lock and adopt if the token
that 401'd is no longer on disk. on invalid_grant, re-read once more
before marking the grant dead.
2026-09-13 19:05:28 +05:30
kshitijk4poor 3aa0db1713 chore(contributors): map FestoneX's email for the #97364 salvage 2026-09-13 19:05:28 +05:30