_render_peers_map_view called _sibling_resolutions per account, and each call
ran _all_profile_host_configs(), which parses every profile's config.yaml via
list_profiles() and re-reads its honcho.json; show() re-runs after every
workspace switch. cmd_peers_map now scans once and threads the rows through.
_seen_gateway_accounts turned a locked or corrupt state.db into the same [] that
means "no gateway traffic yet"; it now notes the error on stderr before returning
the empty list. The missing-table fallback for old schemas is unchanged.
Iterating a honcho SyncPage walks every following page, so _api_workspace_peers
pulled the whole workspace on its first call, never hit the `< 50` break, re-walked
pages 2..N and returned duplicates; the 200-peer cap that exists so a public bot's
workspace cannot stall the CLI was defeated and the `p<N>` picks pointed at the
wrong peers. It now reads `.items` per page and stops at the cap; `_api_workspaces`
reads `.items` too.
The wizard's `new_host` probe looked only at the host block, so an install that
keeps peerName/enabled/workspace at the root with no hosts.hermes block read as
fresh and Enter defaulted to pinning every account onto one peer. The probe
includes the root.
`_seen_gateway_accounts` dropped rows whose origin said `is_bot`, but
SessionSource.to_dict never serializes that field, so the filter was dead; removed
with its test row. `_sanitize_peer_id` was a copy of session_peers.sanitize_peer_id.
A confirmed workspace repoint is persisted, so the exit line no longer says
"Nothing changed" after one.
The identity step treated any host block without a mapping key as a new install and defaulted the choice to single peer. An install with enabled, workspace and peerName then had Enter write pinUserPeer: true and merge every gateway account onto the operator. cmd_setup now decides new-ness before its prompts populate the block, and only a block with none of the mapping, peerName, workspace or enabled keys defaults to single.
The docstring said grouping session rows by (source, user_id) enumerates every account the gateway handled. record_gateway_session_peer overwrites a row's user_id, so a shared thread keeps only its last author. The docstring now says that, and a test pins it.
clone_honcho_for_profile copied sessionPeerPrefix but not sessionAiPeerPrefix. A new profile cloned from a default block with the AI prefix on fell back to unprefixed session names and collided with the default profile's gateway sessions.
peers map showed sanitize(prefix + user_id) for prefixed accounts. The runtime appends a sha256 suffix when sanitizing changed the id or it collides with an explicit peer, so the preview named a peer the gateway never writes to. The preview now builds a HonchoSessionManager with no client and asks it.
peers map can repoint a host block's workspace, but the gateway agent cache signature did not include it. A live gateway kept writing to the old workspace until an unrelated eviction or restart. The signature now carries honcho.workspace.
Clearing the last alias in peers map popped userPeerAliases from the host block, and the host then inherited the root aliases again. The host block now keeps an empty map, which both config readers treat as an explicit override. The summary says when that empty map hides root aliases.
The seven _preview_peer_resolution tests become one parametrized test.
The seven _seen_gateway_accounts tests become two: one database whose
rows exercise grouping, ordering, bot and null-user filtering, legacy
rows without session_key, and both label sources; one for a missing
database or table.
The cmd_peers_map runner is a module function. It derives the profile
list from the config's host blocks, so the multi-profile tests no longer
build one by hand. Tests that share an assertion are parametrized: alias
written to the host block, dash clears an alias, nothing changed writes
nothing, and messages printed in the view. test_raw_runtime_id_entry
was a subset of the offline test and is folded into it. The two
_classify_workspace_peers tests merge into one assertion over the full
label map. The three setup-wizard tests lose their answer comments and
long docstrings; the two default-choice tests are parametrized.
_api_workspace_peers returns peer IDs instead of dicts. The created date
it carried was never printed. Callers index the list directly.
_preview_peer_resolution uses _sanitize_peer_id instead of a local copy.
_sibling_resolutions strips the display suffix itself; its one caller did.
_seen_gateway_accounts opens the database under contextlib.closing and
parses origin_json once into a dict. _save_alias_map picks the target
block first and writes it once. cmd_peers_map renders through one local
show() and computes the previous resolution on one path for both account
numbers and typed runtime IDs.
a browse used to be able to list whatever workspace the process built its
first client for. get_honcho_client now keys its cache on workspace_id, so
the override already gets its own client; this test keeps it that way.
peers map read seen accounts from state.db but required a non-empty
session_key. every telegram row on a long-lived install written before the
gateway stamped that column was dropped, so the picker showed no accounts
while the same rows grouped fine by (source, user_id). the key was never
used by the grouping.
Extends the read-only 'hermes honcho peers' view with a 'map' action
that joins two sources: workspace peers fetched from the Honcho API,
labeled from local config (your peer, each profile's AI peer, alias
targets, runtime peers of seen accounts, user-* fallback peers,
'unrecognized' otherwise), and the gateway accounts recorded in
state.db with what each currently resolves to. Targets are picked
from the workspace list so a typo cannot silently create a peer;
every assignment states its consequence (aliases move future
messages only; a runtime peer left behind keeps its history).
'w' lists every workspace the key can reach — the wrong-workspace
fallback — and can repoint the profile's workspace on explicit
confirmation. With multiple profiles, saving a root-cascading map
asks whether to write root or fork the host block, root writes warn
when sibling profiles sit on other workspaces, and the accounts
table marks siblings that resolve an account differently. Offline
the command degrades to typed targets over the local account list.
The setup wizard's gateway step closes by pointing at the command.
Address review feedback on #39130: the new sessionAiPeerPrefix setting
affects the resolved Honcho session key, which HonchoMemoryProvider freezes
at construction (self._session_key). Because it wasn't part of the gateway's
cached-agent signature, a live config flip left an existing gateway session
bound to its old, AI-peer-agnostic Honcho session until an unrelated eviction
or restart.
Add honcho.session_ai_peer_prefix to _HONCHO_CACHE_BUSTING_KEYS and the
_extract_honcho_cache_busting_config values so a flip rebuilds the cached
agent on the next turn, mirroring the existing aiPeer / pin_peer_name /
runtime_peer_prefix contracts.
Also close the symmetric gap for the pre-existing user-side sessionPeerPrefix:
it feeds the same resolve_session_name output (per-session/title/per-repo/
per-directory strategies) into the same frozen _session_key, so it had the
identical live-flip staleness bug and was likewise absent from the cache
signature. Fixing both keeps the two prefixes consistent.
Add one config-flip regression test covering both keys, alongside the
existing Honcho cache-signature test.
The gateway_session_key branch of resolve_session_name() returns an
AI-peer-agnostic name, so multiple AI peers sharing one workspace +
peerName + gateway chat key collide on a single Honcho session.
Add sessionAiPeerPrefix (symmetric counterpart to sessionPeerPrefix):
when set, the resolved session name is prefixed with {ai_peer}- on every
resolution path. The prefixed name is re-run through the session-id
length cap so the prefix can never exceed Honcho's limit.
- config field + host/root parsing in client.py
- public resolve_session_name() wraps a new _resolve_session_name_base()
- tests covering parsing, the gateway-key case, cross-peer disjointness,
the length cap, and a disabled-by-default regression guard
- README: config table + resolution notes
The step declares its scope up front: human mapping only, with each
Hermes profile bringing its own AI peer. A note explains aliases as
the join between platform accounts and named peers. Each shape now
says when it fits. Fresh configs default the choice to [1] single
peer — the common personal setup — instead of [3], which silently
fragmented a solo operator's gateway account away from their
peerName history. Configured setups keep their detected shape as
the default.
_SESSION_CACHE_MAX_SIZE was assigned twice with two comments describing one
constant. _retain_for_retry and _keep_until_flushed shared the has-unsynced /
current-owner / reclaim body and differed only in what to do when a newer object
owns the key; _reclaim_key_locked returns that owner and each caller keeps its
tail. _ensure_async_writer had no production caller (save() uses the _locked
form under _async_thread_lock); removed, tests retargeted, constructor comment
fixed. shutdown() sets _shutting_down under _async_thread_lock like the other
site so save()'s flag check and the enqueue cannot interleave with it.
tests/test_honcho_session_cache_bounds.py hoists its mid-file imports.
A flush that rebuilds an evicted session's SDK session stores its observation
flags again, and the cap pass pruned every per-session dict but that one, so the
dict grew one entry per evicted-then-flushed session. The cap pass now prunes it
with the rest.
The peer-failure notice classified platforms with its own {"cli","tui","desktop"}
set, so an ACP session read the gateway wording ("do not suggest peerName"); it
now uses agent.coding_context.INTERACTIVE_CODING_PLATFORMS, which includes acp.
Three getattr/try-except guards existed only for test doubles (a bare __new__
provider, a SimpleNamespace config); the real types always carry the attribute.
Removed, and the two tests build real objects. An unset timeout resolves to the
client's 30s default, so the join-budget test expects 30, not the 5s floor.
Timing tests keep the >= 2s wall-clock bound the testing rules ask for.
main's _join_observation_flags returned the manager-wide booleans with a comment
saying #103889 would plug the per-session flags in here; this is that plug. The
two main-side tests that stubbed the old two-tuple _get_or_create_honcho_session
return move to the three-tuple.
The generic panel writes every field flat, and the plugin reads `injection` as one object, so a dotted `injection.sessionStart` field would never be read. Declare `injection` as a JSON field instead. Blank clears the pin.
stop_async_writer bounded only the join and then drained the queue with no deadline, so a shutdown whose budget was already spent could still start uploads, and honcho-ai's add_messages has no per-call timeout. The join and the drain now share the shutdown deadline, no upload starts once it has passed, the writer skips its 2s retry once shutdown began, and shutdown logs one warning with the count left unsynced. The docstrings now say what the budget can do: stop new uploads and bound lock waits, while an upload already in flight runs to the client's HTTP timeout.
Provider shutdown handed the manager a remaining budget, but flush_all ran first with no deadline and blocked on each session's flush lock. An async upload still in flight held that lock, so shutdown waited the full HTTP timeout past its declared budget. flush_all and the queue drain now take the deadline, skip a session whose lock or budget is gone, and log one warning with the count of messages that stayed unsynced.
With recallSync on, prefetch popped only the auth notice and returned without writing the injection log. A session whose init failed for a missing user peer never told the model that memory was off, and the audit file stayed empty for every turn. The recall sync branch now pops the peer notice the way the async branch does and records each turn as injected or recall-sync-empty.
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer
and deferred-save tests that differed only in their inputs. Share the blocking
remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the
added docstrings to the what and the one non-obvious why.
recall_sync.py spawned honcho-recall-sync without owner=, so shutdown's
join_plugin_threads((self, manager), ...) never saw it. The worker now
registers under the provider like the other provider threads.
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the
cached path returned none, so recall fell back to the config snapshot. Both paths now return and
store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or
flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the
writer lock, and the writer drains its queue after the join, so a put that raced shutdown is
written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining
budget instead of a fixed ten seconds.
The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop
passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a
gateway platform supplies no user id, the peer notice and tool error no longer recommend
peerName, which would merge every user of that gateway onto one peer. README documents
`injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution.
a desktop or cli session with no peerName in honcho.json and no gateway
user id landed on a peer derived from the session key: user-default-<dir>
for per-directory sessions, user-<channel>-<chat> for keyed ones. every
directory got its own phantom peer, so the operator's turns and memory
never reached their real peer and the injected representation went stale
(#93326).
_resolve_user_peer_id now raises HonchoPeerUnresolvedError instead of
deriving a name. a peer is either the declared peerName or an identity
the transport supplied. the provider records the failure, tells the model
once that memory is off and which key to set, returns the same detail
from tool calls, and stops retrying init because a missing config key
does not heal mid-session. the memory-file migration gate loses its
"no owner and no runtime identity" branch: that cohort no longer has a
session to migrate into. whitespace-only peerName is treated as unset
rather than sanitized to "--".
provider shutdown joined the dialectic, sync and memwrite threads for 5s
each and then the async writer, but the session-init thread,
honcho-base-first and honcho-context-prefetch were never joined. any of
them still blocked in httpx when the interpreter finalized aborted the
process with SIGABRT 134 (#37632, #60616, #33485). the 5s join was also
shorter than the 30s http timeout a blocked call can hold (#33485).
spawn_context_thread now registers each thread under its owner (the
provider or the manager) in a weak registry, and shutdown joins every live
thread of both owners inside one deadline: at least 5s, or the configured
http timeout when longer. the manager refuses new prefetch threads once
shutdown began and flushes a late save() inline instead of respawning the
writer. threads that outlive the budget are named in a warning.
the sdk client exposes close() on its http pool. one client is shared by
every manager with the same identity in a gateway process, so a per-agent
shutdown cannot close it; close_honcho_clients() closes all pools and is
registered with atexit when the first client is built, the pattern the
hindsight and mem0 plugins use. follows #69070, #33701, #7627.
the idle sweep from #71463 left three growth paths open. _peers_cache had no
bound at all, _session_observation (from #98941) grew one entry per session
id and kept orphans when an init failed after add_peers, and a burst of
distinct sessions inside one ttl window was not bounded. the sweep also
evicted sessions whose messages had not reached honcho yet, which in
"session" write mode drops the only copy, and _flush_session re-inserted an
evicted session into the cache with no observation flags, so recall for it
routed from the config snapshot instead of its server config.
_cache and _sessions_cache now cap at 128 entries and _peers_cache at 512,
evicting least recently used first (dict order, refreshed on every hit).
a session with unsynced messages is never evicted by the sweep or the cap.
_configure_session_peers returns the synced flags and get_or_create stores
them under _cache_lock next to the cache entry, so the observation dict can
hold no id the cache does not; eviction drops both. recall reads go through
_cached_session, which stamps updated_at, so a read-only session survives
the idle sweep. follows #71461, #71463, #98936.
_flush_session read the unsynced messages, posted them, and only then set
_synced. in a short-lived run the async writer draining save()'s queue and
the exit-time flush_all() both saw the same batch unsynced, so every turn
of a one-shot run landed in honcho twice (#92458).
each HonchoSession now carries an RLock that _flush_session holds around
the whole select, send, mark sequence. the second flusher enters after the
first marked the batch and sends nothing; different sessions still flush
in parallel. the lock lives on the session object instead of a manager
dict keyed by session id (#86094, #92787): it needs no eviction, and two
flushers of one message list can never hold different locks. tests
adapted from #86094 and #92787.
The four observation booleans on HonchoSessionManager were manager-wide
mutable state initialized from per-session server configs: every session
setup overwrote them, so the last session to initialize retuned recall
routing for all other sessions the manager serves. Store the server-synced
flags under each session's own id instead; the manager-level fields stay
as the config snapshot and sessions that never synced fall back to them.
injection.sessionStart: unset renders every component in table order, an
empty list renders nothing, a pinned list renders only those names in table
order (not config order), a host block pin beats root, a non-list value is
treated as unset, and initialize() reads the pin.
logging: off by default, the logging key or HONCHO_LOGGING turns it on, a
host block can turn it back off, HONCHO_INJECTION_LOG overrides the path,
each record carries reason/turn/session_key/bytes/payload, an unwritable
path never raises, and a tools-mode prefetch logs its reason. the record
holds the user's representation verbatim, which is why the default stays off.
the two session-context tests that asserted exact context() kwargs now expect
tokens= as well (adopted from #92964). config_schema declares the logging
switch so hermes memory setup shows it.
Salvage of #70951 (Willkons / Alice-Willk-bot). Two of three
session.context() sites never forwarded the configured cap, so Honcho
always returned honcho_chat_summary_long.
Adds the missing get_session_context tokens= regression the sweeper
asked for on that PR.
force_refresh_token gated the adopt-from-disk paths behind the failure cooldown.
After one of our exchanges failed transiently, a 401 within the next 30s returned
None even when a sibling process had already rotated and written a valid grant,
so the operation raised HonchoAuthError with a good token sitting on disk.
Adopting is a disk read, not an exchange; the cooldown exists to stop replaying a
single-use refresh token, so the two adopt checks now run before the gates.
_write_config parsed honcho.json twice under the lock and only wrapped the first
read into ConfigWriteRefused; _refuse_unparseable now returns the parsed dict and
the branches use it. The getattr/isinstance duck-typing collapses to one guard.
cli._read_config reuses oauth's tolerant reader (BOM-tolerant, like the strict
reader the write side uses) instead of its own utf-8 copy.
The unreadable/corrupt variants asserted the same two invariants (writers raise
and leave bytes untouched; rotation fails before the exchange) in five tests;
they are two parametrized tests now. The two lock-spy tests only asserted that
the implementation called the lock helper; the threaded save_config test is the
behavioral guard for the same property.
_write_config applied a command's edits relative to the snapshot the read took, but never moved that snapshot after a successful write. A second write on the same object therefore compared A -> B -> A against A, saw no change, and left disk at B. After a write the snapshot and path now follow the caller's dict, so the next write applies only the edits made since.
_read_config() falls back to ~/.honcho/config.json or the default profile when honcho.json does not exist, so cfg.path never matched the write path and _write_config() wrote the whole dict. That overwrote a refresh rotation a serve child had written onto the honcho.json the setup login created moments earlier. When the local file exists at write time, the seed snapshot is now overlaid with the local file and only the command's edits are applied onto it.
install_grant writes the login grant to disk before the wizard's later prompts. _apply_grant_to_host wrote it only into the live cfg, so _write_config saw the grant as an edit and copied it over whatever a serve child rotated onto disk meanwhile. The snapshot now takes the grant too, so the final save leaves the newer on-disk grant alone.
Near-duplicate tests become one parametrized test each: the 401 adopt
cases, the invalid_grant race, the setup apikey answers, the clone and
enable credential checks, and the plain-dict writes. _point_cli_at
replaces the per-class monkeypatch helpers in test_cli. The removed-key
check folds into the merge test, and the missing-file bootstrap check
into the file-lock test; the command-level refusal test already drives
_write_config through ConfigWriteRefused, so the unit test for it goes.
The spy helpers in the install_grant lock test and the reauth bearer
test lose their duplicated closures. Every behavior the removed tests
asserted still has an assertion.
clone and enable accepted HONCHO_API_KEY from the environment as proof the
block could authenticate and wrote enabled: true. the variable can be absent
from the next process, leaving an enabled block with nothing behind it, which
is the cohort the previous commit set out to remove.
_resolve_api_key takes env=False at both write sites, so a block is enabled
only when honcho.json itself holds a key, an oauth grant, or a self-hosted
baseUrl. status and setup keep the environment fallback for display.
every hermes honcho command read honcho.json, changed a field, and wrote the
whole dict back with no lock. a token refresh in another process that landed
between the read and the write was overwritten with the old access and
refresh tokens, and the next refresh replayed a rotated single-use token.
_write_config now holds the same cross-process file lock the refresh path
holds, re-reads disk under it, and applies only the keys the command changed
since its _read_config(). untouched keys keep their on-disk value, so a
rotation survives; a credential the command set on purpose still wins. a
write with no prior read keeps today's whole-file behavior.
a named profile cloned from a default profile that signed in with oauth
got a host block with enabled: true and nothing to authenticate with.
hosts.hermes holds the grant, its apiKey is not inherited (#66125), and
copying the oauth block would make two blocks replay one single-use
refresh token. 'hermes honcho enable' on a fresh profile wrote the same
shape. status then showed the profile as enabled while every honcho
call ran without memory and the plugin quietly stayed inactive.
clone_honcho_for_profile and cmd_enable now resolve a credential for the
target block (its own apiKey, the root apiKey, HONCHO_API_KEY, or a base
url) before writing enabled: true. _resolve_api_key takes the block to
check so both share one definition of "can authenticate".
cohorts:
clone from an oauth default block, no root key: block written without
enabled; the client's auto-enable rule turns it on once a credential
appears (setup apikey writes the root key, or a per-profile login)
clone from a default block with a host-level static key only: same
clone with a root apiKey, an env key, or a base url: enabled as before
enable on an empty or fresh block with no credential: refused, one
message names the profile's setup command and the hosts.<name> key,
nothing is written
legacy blocks already on disk as enabled with no credential: nothing
rewrites them; the plugin already treats them as unusable and stays
inactive; enable now prints the same message instead of "already
enabled"
env-only key: counted as a credential at write time, as the client
does at run time; if the variable later disappears the client still
refuses to initialize the block
after the token endpoint revoked a grant (invalid_grant), running
'hermes honcho setup', choosing apikey and pasting a valid key changed
nothing. the wizard wrote the key to the root apiKey only. hosts.<name>
still held the dead access token under apiKey and the grant under oauth,
and the host block wins the lookup, so status kept reporting the revoked
grant and every call kept failing (#97990).
the apikey branch now drops the host's oauth block and writes the chosen
key onto the host block as well as the root. a dead access token is no
longer shown as the current key, so a blank answer with no other key
aborts instead of keeping the grant. a static host key without a grant
is kept as before.
a login finishing while a refresh was rotating the same honcho.json ran
its read-modify-write unlocked. the two writers could interleave: the
login read the file, the refresh persisted a rotated token, the login
wrote its own copy back and the rotated token was gone.
install_grant now holds _refresh_lock and _config_refresh_lock(path)
around the strict read, the config merge and the persist, the same way
force_refresh_token does. the token response is parsed before the locks
so a malformed grant never holds them.
a truncated or hand-edited honcho.json reads as {} on the tolerant read
path. the next write then replaced the file with only the current host
block: an oauth refresh, a login, or any 'hermes honcho' command that
saves a setting wiped every other host and the root keys.
_read_config_strict now raises on a parse error the same way it raises
on a read error, and logs one sentence naming the file. _rotate_and_persist
treats that as a refresh failure and enters the cooldown without spending
the refresh token. install_grant raises into the setup flow. in cli.py
every write goes through _write_config, which runs the same check first
and raises ConfigWriteRefused; the honcho router, the setup wizard and the
profile sync print the sentence instead of a traceback. the wizard checks
before asking its questions. read paths keep the {} fallback.
follows #95860, which added the strict reader for unreadable files and
kept a .corrupt copy on a parse error. leaving the original file in place
keeps the same bytes without a second copy of the tokens on disk.
_persist_credential promises "leaving all else intact", but it seeded
its write from _read_config, which returns {} on ANY read failure, and
_atomic_write_config replaces the whole file via os.replace — which
needs only a writable parent, so a present-but-unreadable honcho.json
did not stop the overwrite. One OSError (EACCES after a root-owned
write, EIO, a stalled mount) during an automatic token refresh or a
fresh login therefore replaced the store with a single-host file,
destroying every other host's credentials and honcho.json's root
config. No user action is required to trigger the refresh path.
Same defect class as #75206 (P1, fixed for the core auth store in
Add _read_config_strict for the write paths: a missing file still
bootstraps as {}; an unreadable file raises with the store untouched;
genuine corruption still degrades but preserves a .corrupt copy first,
since a truncated store usually holds the other hosts' tokens verbatim.
_rotate_and_persist now takes its strict read BEFORE the exchange —
rotation is single-use, so an exchange whose result cannot be persisted
loses the grant — and threads the dict through to _persist_credential,
which also closes the re-read race between the locked read and the
persist. install_grant seeds its root-merge from the strict reader for
the same reason. Read paths keep their fail-open contract untouched;
both readers now use utf-8-sig so a BOM'd store is not misclassified
as corruption (the wipe vector needing no filesystem fault at all).
Adds TestPersistReadFailure: six tests, four of which fail against the
previous source; the rotate-ordering test additionally pins that no
exchange is attempted against an unreadable store.
reading the live client's api_key inside _force_reauth races the in-place
rotation: a sibling waiter's apply_token_to_client (or the proactive
refresh entered via the honcho property) can swap the bearer before the
read, so force_refresh_token receives the already-rotated token, disk
matches it, and the adopt branch never fires — reproducing the very
force-refresh burst this PR removes.
_authed_call now snapshots the bearer before invoking the operation and
passes that exact token through to force_refresh_token.
desktop spawns multiple serve processes that share one rotating honcho
refresh token. force_refresh_token treated every 401 as "rotate now" even
when a sibling had already persisted a new grant, which can replay a
single-use refresh token and revoke the whole grant.
re-read honcho.json under the existing file lock and adopt if the token
that 401'd is no longer on disk. on invalid_grant, re-read once more
before marking the grant dead.