Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 01d16451f7 refactor(agent/compression): extract commit-phase helpers out of compress_context
- _fold_todo_snapshot(): stale-snapshot strip + live todo fold into the tail
- _rebuild_system_prompt_at_boundary(): tool refresh + byte-equal keep-prompt
- _salvage_or_refuse_grown_transcript(): commit-site anti-growth guard
- _publish_rotated_compaction(): parent flush, child publish, id re-point,
  goal/heartbeat/loop/title carry-over
- old_session_id is now an explicit Optional local instead of locals().get()
compress_context 1181 -> 794 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium cbd709124f refactor(agent/compression): extract pre-summary phases of compress_context
- _adopt_grown_durable_parent(): rotation-only durable-snapshot adoption
- _pre_compress_memory_context(): provider on_pre_compress / checkpoint gate
- _resolve_compress_call(): compress() + signature-filtered kwargs
- _run_summary_dispatch(): progress hook, stream deadline, interrupt guard
compress_context 1359 -> 1181 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium 42c41dfcff refactor(agent/compression): extract lease acquisition and lifecycle out of compress_context
- _CompactionLifecycle: the one-shot terminal status edge (was a nonlocal closure)
- _CompressionLease: holder/watermark/ttl + refresher, holder-only release,
  fence lock-setup bracket, release() (was 4 nested closures + 9 locals)
- _acquire_compression_lease(): the 190-line lock acquisition ladder
- _adopt_if_parent_rotated(): late-contender adoption check
compress_context 1767 -> 1359 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium 852f2a02d8 refactor(agent/compression): unify repeated abort-path snippets in compress_context
- _existing_system_prompt: 21 copies of cached-or-rebuild prompt
- _emit_aborted_attempt_telemetry: 11 copies of aborted/aborted telemetry
- _restore_messages_snapshot: 4 copies of deepcopy rollback
- _restore_prune_rearm_tokens: 3 copies of prune-runway restore
2026-09-02 13:29:37 -07:00
Teknium 682d48c48a refactor(agent/compression): compact comments and docstrings in conversation_compression (AST-identical) 2026-09-02 13:29:37 -07:00
Teknium 3e8977086f refactor(agent): compact Anthropic stream + retry-loop rationale comments/docstrings (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium 9aee90335b refactor(agent): compact Bedrock and chat_completions stream-loop rationale comments (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium d224e22507 refactor(agent): compact non-streaming call-path rationale comments/docstrings to their invariants (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium 89a5ce12dd refactor(agent): build_api_kwargs — shared chat kwargs for profile/legacy paths; xAI tool_search alias rewrite helper 2026-09-02 13:29:36 -07:00
Teknium c5c6d7e566 refactor(agent): lift non-streaming stale/codex watchdog resolution into _resolve_nonstream_watchdogs 2026-09-02 13:29:36 -07:00
Teknium 3c80d4fe77 refactor(agent): compact build_assistant_message rationale comments to their invariants 2026-09-02 13:29:36 -07:00
Teknium 24b6b0e439 refactor(agent): compact streaming-path rationale comments to their invariants; drop no-op stale-kill if/else 2026-09-02 13:29:36 -07:00
Teknium 0e1875e859 refactor(agent): try_activate_fallback — lift api_mode hint/resolution and credential-pool rebinding into helpers 2026-09-02 13:29:36 -07:00
Teknium 3f497cf3f8 refactor(agent): interruptible_api_call — shared watchdog abort + post-kill worker wait helpers 2026-09-02 13:29:36 -07:00
Teknium 3e2ab5b382 refactor(agent): shared relay stream identity/metadata builders for the three streaming wires 2026-09-02 13:29:36 -07:00
Teknium ce9efccbaa refactor(agent): extract _ToolCallAccumulator from the chat_completions stream loop 2026-09-02 13:29:36 -07:00
Teknium 2d93a8b127 refactor(agent): share SSE connection-drop phrase check and codex silent-hang hint lookup 2026-09-02 13:29:36 -07:00
Teknium 3dc2403c45 refactor(agent): handle_max_iterations — single per-wire summary attempt with one retry pass; drop dead summary_request literal 2026-09-02 13:29:36 -07:00
Teknium e5e491ef5a refactor(agent): unify per-request client registry into _RequestClientRegistry (streaming + non-streaming) 2026-09-02 13:29:36 -07:00
Teknium ce276cd864 refactor(agent): decompose interruptible_streaming_api_call into _StreamingCall + per-wire branch functions
Closures -> _StreamingCall methods (shared worker/monitor state on the
instance), codex passthrough and Bedrock Converse branches -> free
functions, _stream_final_text/_emit_stream_* -> module helpers. Bodies
are AST-identical modulo the self.<name> renames (verified by script);
stream event handling and wire shapes unchanged.
2026-09-02 13:29:36 -07:00
Teknium 695f7da3cf refactor(agent): error_classifier — rule-table verdicts, shared match helpers, compacted pattern rationale (2244 -> 1357 LOC) 2026-09-02 13:29:36 -07:00
Teknium 0edb835abb refactor(prompt_builder): structural simplification with byte-identical prompt output
- build_environment_hints: split into _local_host_hints / _remote_backend_hint /
  _embedder_environment_hint; backend probe split into _run_backend_probe +
  _format_backend_probe with image-key / container-config dispatch tables
  replacing the if/elif chain.
- Skills index: _SkillFilter (frozen dataclass) unifies the disabled+conditions
  check that was copied 4x (snapshot, scan, project, external);
  _collect_extra_skills dedupes the project/external scan loops;
  _read_category_descriptions dedupes DESCRIPTION.md reading;
  _label_visible_entries and _render_skills_index lift the org-labeling and
  rendering regions out of _build_skills_system_prompt_inner; snapshot and scan
  sources now feed one visibility pass.
- Context files: _read_context_file + _context_section unify the
  read/strip/scan/section/truncate sequence across .hermes.md, AGENTS.md,
  CLAUDE.md and .cursorrules loaders.
- Dead: _clear_backend_probe_cache (test-only helper; tests clear the dict
  directly), unused org_id_of_path re-export.
- Comments/docstrings hand-compacted; every rule, invariant, ordering and
  failure-mode rationale kept.

System prompt text verified byte-identical against origin/main over a 273-case
fixture corpus (env hints x backends/probe states, skills index x toolsets /
platforms / project / org / compact, context files x all loaders, full
AIAgent._build_system_prompt_parts x 9 configs). Tool schema byte-identical.
2026-09-02 13:29:35 -07:00
Teknium 4e65d089c0 refactor(agent): simplify moa_loop — dataclass accounting, extracted create() helpers, relay dispatch table, compacted docstrings
- _RefAccounting is a slotted dataclass (trace redaction uses dataclasses.replace)
- create(): preset resolution, fan-out cadence key, reference joining and
  usage/cost summing extracted into named helpers; guidance header shared
- reference pricing and pending-accounting folds unified (one lock path)
- build_moa_facade display-event relay routed via _RELAY_EVENTS table
- dict/object tool-call field access collapsed into _field; dead branches removed
- comments/docstrings compacted to their invariants and WHY content
2026-09-02 13:29:35 -07:00
Teknium 76062b1d6b refactor(agent): decompose init_agent into ordered phase helpers
init_agent (2711 LOC) becomes a ~280-line ordered orchestrator over
_resolve_api_mode / _finalize_routing / _init_* / _build_client /
_load_tools / _parse_compression_config -> CompressionSettings /
_resolve_context_length / _build_context_engine / ... phase helpers.
_build_client is further split per wire mode (_init_anthropic_client,
_init_moa_client, _init_bedrock_client, _init_openai_client with
_explicit_client_kwargs / _routed_client_kwargs). Statement order and
every side effect on the agent are preserved (AST body-parity checked
against origin/main).

Dedupe/dead code: drop _relay_moa_reference_event/_moa_reference_output_allowed
(zero callers; only their own test) and their test file; alias
_normalize_route_base_url; _parse_config_int replaces three copies of the
strict int parser; _cfg_flag replaces four inline truthy-set checks;
_client_kwargs_from_routed + _fallback_entries replace duplicated
routed-client/fallback-entry blocks; _warn_invalid_config_int unifies the
three log+stderr invalid-int warnings (byte-identical text);
_bedrock_region_from_url; _memory_provider_init_kwargs; the
host->default_headers if/elif chain becomes the _HOST_DEFAULT_HEADERS
dispatch table; callback params assigned from _CALLBACK_PARAMS.

Comments/docstrings hand-compacted to their rationale (invariants,
ordering, failure modes kept; issue numbers and narrative dropped).
test_pre_compress_checkpoint_contract source-check repointed at the
CompressionSettings field names.

Verified: tests/run_agent (2066 passed) + all agent_init-referencing tests
(1134 passed), get_tool_definitions() byte-identical vs origin/main, import
smokes for cli/run_agent/gateway.run/hermes_cli.main/agent.conversation_loop/
tui_gateway.server.
2026-09-02 13:29:34 -07:00
Teknium eb67765c58 refactor(agent): agent_runtime_helpers — drop dead predicates, dedupe runtime restore/switch/recovery, compact narratives
- Dead: agent_runtime_owns_post_tool_hook, intent_ack_continuation_enabled (only their
  own tests referenced them; tests removed).
- invoke_tool routes inline tools via INLINE_TOOL_EXECUTORS.
- switch_model normalizes provider names once (was 5x); restore_primary_runtime shares
  primary-pool load/match helpers; _apply_primary_runtime_fields and
  _build_anthropic_client_from_runtime shared by transport recovery and turn-start
  restore; recover_with_credential_pool rotate-and-swap helper (4 sites).
- Incident-narrative comments/docstrings compacted; rules, orderings, invariants kept.
5266 -> 3837 LOC.
2026-09-02 13:29:31 -07:00
Teknium 4dcbbc84ab refactor(agent): table-driven inline tool executors shared by both tool paths; dedupe tool_executor
- agent/inline_tool_executors.py: INLINE_TOOL_EXECUTORS dispatch table (13 agent-level
  tools) replaces two drifted if/elif chains (invoke_tool + execute_tool_calls_sequential);
  resolve_invoke_tool_executor preserves the concurrent path's historical precedence.
- tool_hook_ids / emit_terminal_post_tool_call: single owners for hook identity kwargs
  and the terminal post_tool_call emit (was 8 + 3 hand-copied sites).
- tool_executor: shared quiet-spinner start/stop, tool-search unwrap + alias mapping,
  tool-result finalization tail and completion-callback fan-out between the concurrent
  and sequential executors; quiet/default registry branches merged.
2939 -> 2327 LOC.
2026-09-02 13:29:31 -07:00
Teknium 68be67e0d0 refactor(agent): error_classifier — unify billing/rate-limit/overflow verdict helpers, drop dead _THINKING_SIG_PATTERNS, compact incident narratives
2244 -> 1612 LOC. Classification order, pattern lists, FailoverReason values and
ClassifiedError fields unchanged; _THINKING_SIG_PATTERNS had zero references.
2026-09-02 13:29:31 -07:00
Teknium 064bcaf3c1 refactor(cron): compact scheduler_provider and small cron modules
- scheduler_provider.py: _profile_entry/_profile_cron_scope replace four hand-rolled
  home-override+store blocks in the multiplex ticker; comments compacted to the WHY.
- incidents/executions/notepad/monitor/suggestions/blueprint_catalog/__init__: docstrings
  and comments compacted; no code change (AST-identical).
2026-09-02 13:29:30 -07:00
Teknium 3d3d2c89f5 refactor(cli): table-drive setup.py wizard sections; drop dead helpers
setup.py (3933 -> 3651):
- Dead code (zero references repo-wide, refs.py-verified): _model_config_dict,
  _get/_set_credential_pool_strategy, _supports_same_provider_pool_setup,
  _current/_set_reasoning_effort, _prompt_container_resources, _setup_telegram_auto.
- setup_terminal_backend: (slug, label) table + _TERMINAL_BACKEND_SETUP dispatch
  to one _setup_backend_* per backend; _ensure_sdk / _prompt_secret_env dedupe
  the modal/daytona install+token prompts.
- _setup_tts_provider: label/choice/api-key/local-engine tables, three step
  functions replace the 9-branch elif; _pip_install_tts_package dedupes the
  neutts/kittentts installer tail.
- _print_setup_summary: TTS/STT rows from _TTS/_STT_SUMMARY_ROWS via
  _voice_provider_status; banner and hint lines from small tables.
- setup_agent_settings: session-reset mode table + _prompt_int_setting.
- setup_gateway: _HOME_CHANNEL_CHECKS table.
Printed output and .env/config writes verified byte-identical against base
over exhaustive input matrices (summary 630, tts 3840, agent 3375 combos).
2026-09-02 13:29:29 -07:00
Teknium 5212a3077d refactor(cli): dedupe model_setup_flows boilerplate
model_setup_flows.py (3313 -> 2848):
- _load_config_model_section, _begin/_commit_model_config, _ensure_flow_api_key,
  _pick_model_or_prompt, _run_login, _models_dev_merged, _copilot_model_list,
  _show_curated replace ~15 copies of config-save / api-key / picker boilerplate.
- _gemini_tier_ok and _api_key_provider_model_list lift the two inline blocks
  out of _model_flow_api_key_provider; five-way provider branch -> early returns.
- Comments compacted, keeping every rationale (Bedrock geo routing, key_env
  hygiene, discover_models semantics, Nous free/paid partition, etc.).

Also drops two tests that only asserted the existence of setup.py helpers
removed in the next commit.
2026-09-02 13:29:29 -07:00
kshitijk4poor a3d33fe22f fix(copilot): single-flight the token exchange and close abandoned responses
Follow-up to the off-loop move: once the credential-pool handlers run on
worker threads, the dashboard's periodic /api/credentials/pool polls can
overlap, and during a DNS outage each poll would have started its own
exchange and abandoned its own hung resolver thread.

- Per-fingerprint threading.Lock around the exchange: concurrent callers
  wait on the one in-flight attempt, then hit the positive or negative
  cache (bounded worker count, no duplicate network calls).
- _urlopen_bounded: when the hard cap fires and the abandoned worker later
  succeeds, close the HTTPResponse instead of leaking the socket.
- Tests (none shipped with the original PR): hard cap + late-close,
  single-flight success and failure paths, and the pool endpoint running
  off-loop / keeping the loop responsive under a 200 ms blocking read.
2026-09-03 01:57:09 +05:30
peetteerr ff0afff0e4 fix(server): move blocking credential-pool calls off the event loop
Network off (unplugged) froze the backend 17 minutes: the async
/api/credentials/pool endpoints (GET/POST/DELETE) called load_pool()
synchronously on the event-loop thread -> Copilot token exchange ->
blocking urlopen -> getaddrinfo stuck in C for 1016s, immune to
urlopen(timeout=10). WS dropped (1006), sessions detached, even log
writes stalled.

Fix 1 (web_server.py): move the three endpoint bodies into
asyncio.to_thread, matching the file's existing _run pattern.

Fix 2 (copilot_auth.py): _urlopen_bounded() runs the request in a
daemon thread with a wall-clock hard cap (timeout+5s) so DNS hangs
can no longer block any caller indefinitely.

Measured: simulated DNS hang now raises TimeoutError after 6.0s
instead of freezing the loop.
2026-09-03 01:57:09 +05:30
kshitijk4poor 2b4e70ec07 refactor(cli): use the router's run_in_threadpool alias; offload profile model write
The module already binds run_in_threadpool (used by list_profiles_endpoint)
and every sibling router uses the same starlette helper; the nine new
loop.run_in_executor(None, _run) sites now go through that alias so the
file has one offload idiom. Behaviour-identical (both hand the callable to
a worker thread).

Also sweeps the one endpoint the PR left synchronous:
update_profile_model_endpoint's _write_profile_model reads and rewrites
the profile's config.yaml on the event loop.
2026-09-03 01:56:47 +05:30
briandevans beb2e91d04 test(cli): cover the profiles router off-loop sweep
Two assertions per offloaded site:

- a loop probe, where the stubbed callee records whether an event loop is
  running in its own thread — the idiom already used by
  tests/hermes_cli/test_cron_dashboard_off_loop.py; and
- a concurrency proof, where the stubbed callee blocks on a threading.Event
  while an unrelated request is timed. On the unfixed handlers that request
  waits out the whole block; served off the loop it returns in
  milliseconds.

The concurrency proof needs a single event loop across requests, so the
client fixture enters the TestClient context manager: that pins one
blocking portal for the whole fixture, where a bare TestClient(app) would
spin up a fresh loop per request and pass even unfixed.

Also covers the status-code mapping through the executor hop (404 on a
missing profile, 400 on a rename collision, 404 from the resolve that stays
on the loop ahead of describe-auto) and the _MISSING sentinel cases: a
desktop.json holding `null` still reports exists=true, an absent one
reports exists=false, and an empty SOUL.md is still distinguishable from a
missing one.

The client fixtures read web_server._SESSION_TOKEN from the module rather
than pinning a literal. web_server resolves that token once at import, so
whichever test file imports it first fixes the value for the session and a
later monkeypatch.setenv is silently ignored — two files hardcoding
different tokens would 401 depending on collection order.
2026-09-03 01:56:47 +05:30
briandevans 34a8e1dc50 fix(cli): run profile document I/O off the dashboard event loop
The remaining in-scope handlers in this router read and write profile
documents inline on the ASGI event loop:

- GET  /api/profiles/{name}/soul            reads SOUL.md
- PUT  /api/profiles/{name}/soul            atomic_write_text(SOUL.md)
- PUT  /api/profiles/{name}/description     write_profile_meta(profile.yaml)
- GET  /api/profiles/{name}/desktop-overlay reads desktop.json

The persona save is the sharpest of the four: atomic_write_text() writes a
temp file, fsyncs it and replaces the original, so the loop is parked for
however long the filesystem takes to durably commit — unbounded on a slow
or contended disk, and paid on every Save in the editor.

Each handler keeps its existing status-code mapping. The reads probe and
load in a single executor hop rather than two, which also avoids widening
the gap between the existence check and the read.

Both readers return a _MISSING sentinel rather than None for an absent
file. desktop.json may legitimately contain the document `null`; collapsing
that onto None would newly report an existing-but-empty overlay as absent.
The same distinction is what the SOUL.md durability tests rely on, where
"file missing" and "file empty" must not both read as never-set.

_resolve_profile_dir() stays on the loop in all four, as it does in the
rest of this sweep: it is a name check plus one stat, and it owns the
400/404 responses.
2026-09-03 01:56:47 +05:30
briandevans 63d42cd0e2 fix(cli): run profile rename and active-profile state off the event loop
Three more handlers in this router did filesystem work inline on the ASGI
event loop:

- PATCH /api/profiles/{name} calls rename_profile(), which stops a running
  gateway through the same 10-second _stop_gateway_process() poll that
  delete uses, then renames the profile directory, rewrites the Honcho
  host blocks and regenerates the wrapper script.
- GET /api/profiles/active reads the active_profile state file and
  resolves HERMES_HOME against the profiles root. The sidebar polls it.
- POST /api/profiles/active stats the target profile, creates the state
  directory and writes active_profile via a temp file plus replace.

Rename carries the same worst case as delete and belongs off the loop for
the same reason. The two active-profile handlers are individually cheap,
but they are the routes the dashboard polls, so they are the ones most
likely to be queued behind something slower — and leaving them inline is
what made the router inconsistent with list_profiles_endpoint, which
already offloads a plain directory listing eight lines above.

The two reads in GET share one executor hop rather than taking one each.
2026-09-03 01:56:47 +05:30
briandevans 4da5689b80 fix(cli): run auto-describe LLM round-trip off the dashboard event loop
POST /api/profiles/{name}/describe-auto called
profile_describer.describe_profile() inline. That function is a plain def;
it reaches agent.auxiliary_client.call_llm(), also a plain def, which makes
a synchronous provider request with a 60-second ceiling.

Held on the ASGI event loop that is six times the 10-second WebSocket
ready-probe threshold web_server.py records as the point where the desktop
app gives up (GH-73083). A single describe-auto on a slow or unreachable
auxiliary provider therefore takes the whole dashboard offline for up to a
minute, including the /api/ws and /api/pty sockets the desktop app and the
Chat tab run on.

Move the import and call into the default executor. _resolve_profile_dir()
deliberately stays on the loop ahead of the hop: it is a name validation
plus a single stat, and it owns the 400/404 responses that the handler's
`except Exception` would otherwise turn into a 500.
2026-09-03 01:56:47 +05:30
briandevans a2504a0a59 fix(cli): run profile deletion off the dashboard event loop
DELETE /api/profiles/{name} called profiles.delete_profile() inline on the
ASGI event loop. When the target profile has a gateway running, that call
stops it via _stop_gateway_process(), which polls the PID every 500 ms for
up to 10 s before escalating to a force kill, and then removes the profile
tree.

For the whole of that window the dashboard process serves nothing else.
web_server.py's own notes record what that costs: a stall of this length
"caus[ed] the Desktop's 10-second WebSocket ready-probe to time out
(GH-73083)", and both the desktop app and the dashboard's Chat tab drive the
agent over those WebSockets. Deleting a profile whose gateway is up is a
routine action that reliably reaches the full ten seconds — the handler's
own output announces "Gateway is running - it will be stopped".

Move the call into the default executor via run_in_executor, matching
list_profiles_endpoint, export_profile_endpoint and import_profile_endpoint
in this same module. The exception-to-status mapping is unchanged:
FileNotFoundError/ValueError are raised inside the worker and re-raised by
the await, so they still map to 404/400.
2026-09-03 01:56:47 +05:30
briandevans f786c8699b test(proxy): cover the retry credential blocking the proxy event loop
Extends the off-loop suite to the third and last blocking method on the
`UpstreamAdapter` contract, the 401/429 rotation.

As with the two existing pairs, the primary assertion is **thread identity**,
not latency: a latency assertion measured by an HTTP client on the blocked
loop is vacuous, because the client's own timer cannot advance until the
block ends and it therefore reports a fast response on provably frozen code.

  * `test_get_retry_credential_runs_off_the_event_loop` records
    `threading.get_ident()` inside the fake adapter and compares it to the
    loop thread, and checks the rotation still works end to end (rejected
    bearer forwarded first, rotated bearer second).
  * `test_event_loop_keeps_running_while_the_retry_credential_resolves`
    samples a loop-side heartbeat counter from inside the stalled adapter. On
    the unfixed handler it records exactly 0 loop iterations across a 0.5s
    rotation.
  * `test_retry_credential_failure_still_returns_the_upstream_rejection`
    guards the error contract the change must leave alone: a raising rotation
    is still swallowed and the upstream's own 401 is streamed back, with no
    second forward.

A new `_build_rejecting_upstream` harness drives the `status in {401, 429}`
branch by rejecting every bearer except the rotated one.
2026-09-03 01:56:32 +05:30
briandevans 8f6df4a82a fix(proxy): resolve the 401/429 retry credential off the event loop
`handle_proxy` already offloads the two credential reads on the happy path,
but the rotation inside the `upstream_resp.status in {401, 429}` branch still
called `adapter.get_retry_credential` inline on the event loop.

That is the most expensive of the three blocking methods on the
`UpstreamAdapter` contract, not the cheapest:

  * `NousPortalAdapter.get_retry_credential` routes into
    `_get_credential(force_refresh=True)`, so the token-refresh POST that
    `get_credential` performs only near expiry is unconditional here — and it
    runs under the same `_auth_store_lock()`, a cross-process advisory lock
    with a 15s timeout.
  * `XAIGrokAdapter.get_retry_credential` loads the key pool off disk and
    calls `try_refresh_current` / `mark_exhausted_and_rotate` under its lock.

So every upstream 401 or 429 froze the proxy's single event loop — and with
it every other in-flight streaming completion — for the whole rotation. A 429
is exactly when the proxy is busiest, which is the worst moment to stall.

Wrap it in `asyncio.to_thread`, matching the two sites above. The error
contract is unchanged: `to_thread` re-raises the worker's exception in the
awaiting frame, so the existing `except Exception -> retry_cred = None` still
swallows a failed rotation and streams the upstream's own rejection back.
2026-09-03 01:56:32 +05:30
briandevans 4b2f16d828 test(proxy): cover health responsiveness while the auth store is locked
Extends `test_proxy_off_loop.py` with the `/health` half, using the same
two-assertion shape as the credential tests:

- `test_is_authenticated_runs_off_the_event_loop` compares the thread the
  adapter's `is_authenticated` ran on against the loop thread. Before the
  fix they are the same ident.
- `test_event_loop_keeps_running_while_health_resolves_auth_state` reads a
  loop-side heartbeat counter sampled by the adapter across its own stall.
  Before the fix exactly 0 iterations run across 0.5s.

Both also assert the response is unchanged (`200`, `authenticated: true`),
so the offload cannot quietly alter what `/health` reports.
2026-09-03 01:56:32 +05:30
briandevans 985b034a51 fix(proxy): keep the health endpoint off the auth-store lock
`handle_health` called `adapter.is_authenticated()` inline from an
`async def`. `UpstreamAdapter.is_authenticated` is documented as
"Should be cheap — no network calls. Used by `proxy start` for a clear
up-front error before binding a port." (`adapters/base.py`), and that is
true of the `proxy start` preflight, which runs in a plain synchronous
CLI function. It is not true on the event loop:
`NousPortalAdapter.is_authenticated` goes through `_read_state()`, which
takes `_auth_store_lock()` — the same cross-process lock with a 15s
timeout as credential resolution — and `XAIGrokAdapter` reads its key
pool off disk.

`/health` is precisely what a supervisor, systemd unit, container
healthcheck or load balancer polls, on a fixed interval, so it is the
endpoint least able to afford a lock wait; and a wait here freezes every
concurrent proxied stream, not just the healthcheck.

Offload it with `asyncio.to_thread`. The response body is byte-identical;
only the scheduling changes.
2026-09-03 01:56:32 +05:30
briandevans d4c2c74ece test(proxy): cover credential resolution blocking the proxy event loop
Adds `tests/hermes_cli/test_proxy_off_loop.py`, mirroring the harness in
`test_proxy.py`: the proxy and a fake upstream run as real aiohttp
servers on ephemeral ports under a single `asyncio.run`, guarded by
`pytest.importorskip("aiohttp")` — no pytest-aiohttp dependency.

The primary assertion is thread identity, not latency. A latency
assertion measured with an HTTP client on the blocked loop is vacuous:
the client's own timer cannot advance until the block ends, so it reports
a fast response on code that was provably frozen.

- `test_get_credential_runs_off_the_event_loop` records
  `threading.get_ident()` inside the adapter and compares it to the loop
  thread. Before the fix both are the same ident.
- `test_event_loop_keeps_running_while_credentials_resolve` runs a
  heartbeat task on the loop and has the adapter sample its counter on
  entry and exit, so the reading is taken from the loop rather than
  through a client that shares it. Before the fix exactly 0 iterations
  run across a 0.5s stall; after it, ~50.
- `test_credential_failure_still_maps_to_401` pins the error contract
  across the change of call form. It is deliberately not in the
  red-before set — it guards behaviour the fix must leave alone.
2026-09-03 01:56:32 +05:30
briandevans 5b63c6174b fix(proxy): resolve upstream credentials off the event loop
`create_app` registers `handle_proxy` as an `async def`, and it called
`adapter.get_credential()` directly on the aiohttp event loop.

`UpstreamAdapter` is a synchronous contract (`adapters/base.py` — every
method is a plain `def`), and the shipped adapters implement it with
blocking I/O. `NousPortalAdapter.get_credential` takes
`_auth_store_lock()` — a cross-process advisory lock with
`AUTH_LOCK_TIMEOUT_SECONDS = 15.0` (`hermes_cli/auth.py:110`) — reads
`auth.json` off disk, and may issue a token-refresh POST; on a terminal
`AuthError` it takes that lock a second time to persist the quarantined
state. `XAIGrokAdapter.get_credential` reads its key pool off disk.

The proxy is one process with one loop, and `handle_proxy` streams with
`sock_read=300`, so long-lived completions are the normal case. Blocking
inside credential resolution therefore freezes *every* concurrent
in-flight stream mid-token for the duration — a concurrent `hermes auth`
command holding the auth-store lock is enough to do it. Every proxied
request goes through this path.

Dispatch through `asyncio.to_thread` instead. This is a pure scheduling
change: `to_thread` re-raises the worker's exception in the awaiting
frame, so the existing `except Exception` -> 401 `upstream_auth_failed`
mapping is unchanged, and the adapters' own `self._lock` still serialises
concurrent resolutions exactly as before. Fixing it at the handler also
leaves the synchronous `UpstreamAdapter` ABC untouched, so it covers
every adapter without conflicting with in-flight work that subclasses it.
2026-09-03 01:56:32 +05:30
kshitijk4poor 5f24f291c2 fix(curator): carry .git file pointers too and document the deleted-skill drop
Submodule and worktree checkouts store .git as a file (gitdir: pointer);
the carry-over only looked at directory names, so that form was still lost
on rollback. Handle files with the same guard. Docstring now states the
deliberate limit: an excluded entry whose skill dir the target snapshot
lacks is dropped with staging (no orphan .git) and is not undoable via the
safety snapshot, which excludes these paths as well.
2026-09-03 01:43:22 +05:30
kshitijk4poor 30be83ab72 fix(curator): carry nested excluded subtrees across rollback
Excluding nested .git from snapshots has a side effect on rollback: the
staging move takes the whole live skill dir (including its .git) into
.rollback-staging-*, the extract restores the snapshot without it, and the
staging dir is then deleted — so a skill that is itself a git checkout lost
its .git on any rollback. Reproduced: main preserves it, the exclusion-only
branch did not.

After a successful extract, move excluded subtrees from the staged copy back
under their restored skill dir (mirroring how a top-level .git survives by
never being staged). Regression test included.
2026-09-03 01:43:22 +05:30
kshitijk4poor 3bfa07fc5d docs(curator): carry over the #91449 growth rationale for the .git exclusion
Folds the incident explanation from #91458 (@liuhao1024) into the comment on
_EXCLUDE_TOP_LEVEL so the reason .git is excluded — compounding snapshot
growth, not just rollback safety — survives next to the set.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-03 01:43:22 +05:30
Jakub Wolniewicz 5b92b9f913 fix(agent): exclude .git and .curator_backups in curator snapshot_skills 2026-09-03 01:43:22 +05:30
kshitijk4poor ff7233b815 fix(credential_files): apply the same exclusions to the symlink-safe mount copy
_safe_skills_path() is the sibling of iter_skills_files(): when a symlink in
skills/ forces a sanitized copy for mount-based backends (Docker/Singularity),
it rglob-copied the whole tree — .hub, .curator_backups, node_modules and all.
Prune EXCLUDED_SKILL_DIRS before descending, same rule as the sync generator,
so the mounted copy never carries (or walks) the bookkeeping trees either.
2026-09-03 01:33:18 +05:30
kshitijk4poor 1d06ef3a5d refactor(credential_files): prune excluded dirs before descending in the sync walk
Replaces the three hand-copied rglob loops + post-hoc parts check with one
os.walk generator that drops EXCLUDED_SKILL_DIRS from dirnames before
recursing. Same file set as the cherry-picked fix (the test binds it), but
the walk no longer stats every file under .hub/.curator_backups/node_modules
on each 5s FileSyncManager tick.

Bench (synthetic skills tree: 20 skills + 400 .hub files + 5x8MB curator
tarballs + 50 archived files): iter_skills_files() 35ms -> 2.4ms.
2026-09-03 01:33:18 +05:30