acfda3759871e50a6899151069a564458604a721
3915 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f1ee8e0f39 |
fix(cli): release SessionDB handles when deleting a profile
delete_profile already force-closes holographic memory_store.db in this process, but the shared SessionDB registry kept state.db open. Recreate then failed with a replaced/locked database. Close every shared handle under the doomed directory, same contract as MemoryStore.release_all_under. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
c9b71ebaf7 |
fix(dashboard): startup schema reconcile opens state.db read-only first
The dashboard's `_eager_reconcile_own_session_db` did an unconditional writable `acquire()` at every startup. When the gateway shares that state.db the dashboard became a second long-lived writable SessionDB owner: a close-time WAL checkpoint plus a possible FTS rebuild in `_init_fts`, the two-writer vector behind the corruption reports in #107688 and #100896 ("5 live SessionDB handles" precursor, gateway + dashboard both holding the WAL). Route the startup reconcile through `_open_session_db_at_path(..., read_only=True)`, which already bootstraps a missing store and heals a stale/malformed schema through exactly ONE writable open before reopening read-only. A healthy store now gets zero writable opens from the dashboard while the #79531/#80037 "bring schema current before the first poll" contract is kept (existing heal test unchanged). Live repro (healthy store, count writable SessionDB.__init__ calls from the startup worker): before=1 after=0. Reported-by: #107688, #100896 (@kokhlo diagnosis) Refs #107688 #100896 |
||
|
|
b02f3aa00b |
fix(doctor): refuse the WAL checkpoint while a live writer holds state.db (#103339)
doctor --fix ran a raw sqlite3.connect + wal_checkpoint(PASSIVE) with no holder guard. Under a running gateway that second connection joins the live WAL and its close-time handling is the second-writer corruption class #103339 tracks. Route the checkpoint through live_writer_holds_db and skip with an actionable finding when the database is held. [Salvage note: the original PR also flipped the repair probe's DatabaseError branch to fail closed; that hunk is dropped here because it refuses every header-destroyed DB (the exact class repair exists for) — the probe's own connect raises before the EXCLUSIVE/BEGIN statements ever run.] |
||
|
|
5d2d5e906d |
fix(tests): banner ssh-fastpath tests patch the seam production reads (_github_branch_tip)
Since
|
||
|
|
9ce7547faf |
fix(update): bound network git in hermes update and prune shallow grafts on apply
- _git_run(network=True) now carries a 300s timeout; a dead-stalled fetch (HTTP/2 to GitHub on some networks, black-holed proxy) becomes a failed run whose stderr names the stall instead of an update pinned on 'Fetching updates...' forever (#93759, #95777). Local git stays unbounded. - The apply path prunes stale .git/shallow grafts alongside its existing lock/tmp_pack cleanup, so installs that already accumulated grafts from past depth-1 checks (#105951: 57 entries) heal on their next update, not only on --check. |
||
|
|
5bfa389e8a |
fix(cli): prune stale shallow grafts left by depth-1 update checks (#105951)
Every 'git fetch --depth 1' in 'hermes update --check' (and the past banner passive checks, before #107648 moved them to the GitHub API) appends the fetched tip to .git/shallow as a new graft and git never removes the previous one, so a long-lived shallow installer checkout accumulates one graft per check (57 observed). The stale grafts break merge-base and push 'hermes update' into the orphan-divergence reset path with a rescue ref on every run. prune_stale_shallow_grafts() now runs after each successful depth-1 fetch in 'hermes update --check' and clears the grafts already accumulated by past checks: it keeps only the boundaries still protecting referenced tips (HEAD, FETCH_HEAD, every ref tip) and atomically rewrites .git/shallow, restoring the original file if the trimmed set breaks history walking. The dropped commits are already unreachable; their objects are left for git gc. Rebased onto main after #107648: the banner.py hook is dropped (the passive check no longer git-fetches); the update --check prune and the cleanup of already-accumulated grafts are kept. (cherry picked from commit 6174837fc5b9f4cc3d4d46dc1b2d9a2f6b83c120) |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
c2d01f1012 |
fix(desktop): key the page under the stamp its rows already carry
A custom HERMES_HOME (outside the profiles tree) reported `profile: null`, so the tail entry landed under profile '' while the session's owner hint — built from list rows that _serving_profile() stamps "default" — looked up 'default'. Two identities for one backend; the exact-key lookup missed and "Show earlier" stayed hidden for the Desktop's own sandboxed and HERMES_HOME-overridden installs (electron/main.ts resolveHermesHome). The messages endpoint now returns the same _serving_profile() stamp the list rows carry, so the Desktop keys a page under the owner it already routes the session by; the inline resolver and its function-body import go away. On the renderer, the ambient profile used for backends that predate the field is captured at request time alongside the connection, so a profile switch mid-read cannot re-key the page. |
||
|
|
d45842d206 | fix(desktop): preserve older history pagination across scopes | ||
|
|
2423385c05 |
refactor(serve): one place maps "process gone" to orphaned in the watchdog
The inner ProcessLookupError handler around the marker probe made the outer one reachable only through an injected pid_exists that raises (gateway.status._pid_exists never does). Keep the single mapping; the degrade-to-liveness fixture no longer models a stderr the probe now classifies as a missing process. |
||
|
|
e7eed649a6 |
fix(serve): only "no such process" from ps means the Desktop parent is gone
The stderr sniff also matched "not found", which BSD ps prints for an
unsupported column ("ps: lstart: keyword not found", rc=1, verified on
macOS). That would have classified a healthy parent as dead and os._exit'd
the backend — the fail-unsafe direction the watchdog must never take. Keep
the explicit missing-process message only; the macos_only test now pins the
"keyword not found" case as a plain OSError. The standalone
ProcessLookupError test (already green on main) folds into the degrade test.
|
||
|
|
60562ed81a |
test(serve): watchdog probe tests stop faking the host platform
Fold the five #80204 tests into three contracts: OSError degrades to pid liveness, ProcessLookupError is conclusive, and the darwin `ps` probe classifies a missing-process message vs any other failure. The last one carries `macos_only` instead of monkeypatching sys.platform / os.name. |
||
|
|
3e72aef744 | fix(serve): reserve ProcessLookupError for known missing process signals | ||
|
|
2489f15040 | fix(serve): ensure parent-death watchdog detects exited Desktop parent (#80204) | ||
|
|
8d79c2ff57 |
Merge pull request #107959 from NousResearch/fix/local-fit-mtp-accounting
fix(local-runtime): unify memory accounting and effective-window MTP |
||
|
|
0316d3d404 |
fix(local-runtime): unify memory accounting and effective-window MTP
Price weights, context, runtime, projector and batch overhead consistently across catalog admission, initial launch, growth and restored windows. Keep MTP and the larger window when lean batches avoid unnecessary spill. Admit optional external drafts only when their complete footprint fits. Use preset-only model discovery so refused files cannot autoload, and preserve refusal/spill decisions atomically for desktop status read-back. Add regression coverage for complete-footprint boundaries, MTP restarts, growth admission, draft budgets and placement status transitions. Builds on the overhead-accounting contribution in #102993 and the restored-window MTP contribution in #106897. Does not adopt the 40% host-RAM reserve or resolve the remaining requests in #102865/#106895. Co-authored-by: infinitycrew39 <infinitycrew39@gmail.com> Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com> |
||
|
|
0c0983ab21 |
test(auth): anon-auth fixture resets the per-profile token memo as a dict
#107697 ( |
||
|
|
338bf9ea9a |
fix(cli): banner/TUI/dashboard update checks go through the GitHub API, cached 24h
The Python passive check (`banner.check_for_updates`, used by the CLI
banner, `hermes --tui`, every `tui_gateway` spawn and the dashboard's
/api/hermes/update/check) also ran `git fetch origin main` on every cache
miss, and never cached an inconclusive result so a flaky line retried on
every start. Same GitHub complaint, same fix:
- remote tip via GET /repos/{slug}/commits/main (vnd.github.sha), local tip
via rev-parse, exact count + changelog via the compare API when they
differ. HTTPS `ls-remote` remains only as the fallback when the API is
unreachable or the origin isn't on GitHub.
- cache TTL 6h -> 24h, failures cached 1h; the cache is keyed on HEAD so
`hermes update` invalidates it immediately.
- the dashboard's "what's changed" list comes from the memoized compare
payload (`upstream_commits_behind`) instead of `git log HEAD..origin/main`,
which was stale without a fetch.
Tests rewritten to the new contract: passive checks must not run
`git fetch`/`ls-remote` for a GitHub origin; the daily cache invalidates
when HEAD moves and re-asks after the failure window.
|
||
|
|
08830efd96 |
fix(secrets): secret-source re-pull no longer latches an empty snapshot or wipes sibling profiles
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).
Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (
|
||
|
|
89fb1d028e |
fix(mcp): retain profile secret scope during discovery
(cherry picked from commit f344095fcf16c5d154ef2d207cce9b13339bf2e4) |
||
|
|
2b4deeb32b |
fix(auth): key sibling per-process credential memos by profile home under multiplex
Same class as the resolve_nous_access_token memo: three more process-wide memos carried a credential resolved under one profile's HERMES_HOME override into another profile's turn for their TTL. - hermes_cli/nous_billing.py::_token_cache (30s (token, base) memo for the charge poll loop) was a single unkeyed slot -> dict keyed by hermes_home_key(); invalidate_cached_token() clears the dict. - agent/moa_loop.py::_runtime_cache carried api_key/base_url/api_mode keyed only (provider, model) for 5 min -> (hermes_home_key(), provider, model). - agent/auxiliary_client.py::_client_cache_key had no profile component, so callers that omit api_key (pool / Nous auth.json paths) could be handed a client built with another profile's bearer -> hermes_home_key() leads the key. WHY hermes_home_key(): it reads the per-turn HERMES_HOME override the multiplex gateway sets (falling back to the env var), and it is symlink-stable, so the memo key is exactly the credential home the resolution itself read from. Profiles stay independent islands; the default-profile process env never leaks into a secondary's turn. Tests: one invariant per site, proven red on origin/main. |
||
|
|
173105ce6f |
fix(auth): scope the resolve_nous_access_token memo to the active profile
resolve_nous_access_token()'s 5s startup-burst memo (#76930) cached the resolved Nous Portal access token in a single module-level slot keyed by nothing but wall-clock time. The underlying resolution is profile-scoped: _auth_file_path() reads get_hermes_home(), which checks the context-local _HERMES_HOME_OVERRIDE ContextVar before falling back to the HERMES_HOME env var — gateway/run.py and tui_gateway/server.py set that override per-profile for multiplex concurrency. In a multiplex gateway serving two profiles with different authenticated Nous accounts, if profile A's context resolves a token and profile B's context calls resolve_nous_access_token() within the next 5 seconds, profile B received profile A's cached token — used to authenticate against the managed tool gateway / relay self-provisioning under the wrong account. Key the memo by str(get_hermes_home()) instead of a single slot, so each profile's context reads only its own cached token. The lock around read/write is unchanged; only the cache's shape moved from a single (timestamp, token) tuple to a dict keyed by resolved home. (cherry picked from commit 9d8846b88cfd2ea74c6958d5f8f28a50880dda75) |
||
|
|
4bdd64b334 |
The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cbcf7b72f7 |
feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop * feat(gateway): /signin signs the free tier into a Nous account from a DM * feat(cli): chat surfaces name /signin as the sign-in verb * fix(auth): review follow-ups for the shared sign-in flow and /signin * fix(i18n): carry the /status free-tier line in every locale catalog * refactor(cli): the chat sign-in command is /login * fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules |
||
|
|
3b01b4ce0f |
feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status answers has_guest / enabled / carries_inference / notice_pending with zero network, and free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and billing.state carry free_tier (billing answers the free tier locally instead of a portal call that can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector transfer and returns its code and consent URL; the poller waits for the transfer before the token grant, persists the account, runs settle_after_upgrade, and the poll response gains reason, account_email and model. * feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding overlay opens on a ready screen when the free tier carries inference, else a one-time strip above the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model / Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens; Done settles billing, model options, providers and re-homes a session still on nous/welcome. The picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes. * fix(desktop): free_tier.status starts the free tier's background setup when no identity exists A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free tier was never set up on the desktop: no connectors, no notice strip. The first status read now starts the same one-attempt background setup; the call itself never waits. * fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in). The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free tier, instead of Nous Portal · Connected. * fix(desktop): Settings > Providers never files the free tier under Connected * fix(desktop): the intro's shape is keyed on the route, not on the identity free_tier.status reports available (an identity exists and the tier is on); whether inference runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready screen shows when that route is the free tier; the composer strip when the user's own provider carries inference. An own-key install used to get the ready screen. * docs(desktop): say what the free-tier chip is keyed on * fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds * fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the transfer wait, after the token grant, and once more under the session lock together with the save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an attempt generation that every continuation checks after each await, so a poll from a closed attempt cannot publish over the one on screen (and its backend session is cancelled). The ready screen comes down only after the backend recorded the acknowledgement. A composer still mounted takes over the notice claim when its owner unmounts. One thin test per hole. |
||
|
|
04a76c4109 |
fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs after the account is persisted: a config on the free tier's route moves to the account's inference host and the recommended default for the account's plan, through the same config write a plain Nous login uses; a config on the user's own model is left alone. The pick is the one GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default. * fix(auth): a sign-in completion with no eligible recommendation leaves no default model The static provider-wide default is not narrowed by the account's plan or org policy, so writing it as a fallback could persist a model the account may not use. When the recommendation cannot yield a model, the route still moves to the account's host but model.default is left unset; the CLI says so and points at `hermes model`. * docs(free-tier): say what happens when no recommendation is available after sign-in * fix(auth): sign-in completion moves the host and clears the default in one config write Two writes could fail between them and leave the account host paired with nous/welcome. _update_config_for_provider gains clear_default so the caller with no model to offer removes model.default in the same atomic write that sets the host. |
||
|
|
a2db110ccc |
feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous provider) instead of forcing the setup wizard. The identity is persisted through the same path a real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off. * test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin * docs(user-guide): free tier and signing in New page explaining what a fresh install gets before any key or sign-in (free inference on nous/welcome plus connectors), how the free tier coexists with a user's own API key, how to sign in with hermes auth upgrade and keep connectors, how to turn the free tier off with nous.guest, what hermes logout does in each state, a troubleshooting table, and a plain privacy note. Wired into the Using Hermes sidebar. * fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user is told they were never signed in. Logging out of a real Nous account now also clears the cross-profile store, so a profile logout is not silently re-adopted on the next boot. * fix(model): switching off the free tier points at signing in, never hops providers * Name the free tier in the gateway startup notice and tell explicit-provider installs about it once * Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token * fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free tier before declaring nothing configured. On a fresh install the first command lands in chat on nous/welcome; a failed setup still falls through to the existing guidance. * Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors The device-code flow runs as usual, with a promotion intent registered on the portal between the code request and the token poll so the account that approves the code inherits the free tier's connectors. The promotion status decides the outcome: only a completed one is followed by the token grant, which is persisted over the free-tier singleton and the shared store. Declined, superseded, retired and busy outcomes each print their own plain copy, and a retired identity is cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals. * Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off * fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the claim code, not the generic device page. A failed mint is attempted once per process so several bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check. * fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure A credential-pool entry can select a paid Nous key while the profile singleton is still the free tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and the pin in model normalization is removed since it had no route to look at. A failed background identity setup releases its latch so a later attempt in the same process can try again. * fix(auth): decide the Nous model together with the route on every credential-pool swap The credential pool can move a Nous agent between the welcome host and the portal host after init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model. * fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start * fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the shared account. Locks are taken in the documented order (profile, then shared). A minted credential is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and trigger a second mint. Retiring a dead credential removes only that credential from both stores. Guest exchange uses the resolver's canonical portal URL. * fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential The welcome host serves one model, so a rotation onto it is refused for any conversation on another model instead of silently switching that conversation to nous/welcome (the model pin applies only when a route is first chosen). The connector token path now treats the free tier as absent when nous.guest is false, including cached tokens, and shares the one dead-credential rule with inference: a retired identity is replaced once rather than returning its stale token. * fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only A free-tier identity in the shared store is not an OAuth credential to offer for import; a real sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted state (no token refresh at boot), so an expired free-tier token cannot stall the online message. |
||
|
|
8b2b83a906 | fix(providers): pin all Actual routes to chat completions (E-1047) | ||
|
|
d7b0a72c2a | fix(providers): route Actual through chat completions (E-1047) | ||
|
|
0ac631b6c0 |
feat(api): accept a turn author on session chat and runs; peer dm forwards it
POST /api/sessions/{id}/chat, /chat/stream and /v1/runs take an optional
"author" object. absent or null keeps today's behavior; a non-object is a 400
invalid_author. the value reaches run_conversation as turn_author and nothing
else: it is a claim by an API-key holder and grants no capability.
hermes peer dm and peer run add the author to the request body when the
message_agent runner launched them with HERMES_TURN_AUTHOR set, so a
cross-machine dm is attributed the same way a local one is.
|
||
|
|
ee35a4624f |
fix(routing): address review — session runtime decides, quarantine continues the chain, unclassified ids break ownership
Three defects found in review of the first head (@ehz0ah): * The mid-request hop read the PERSISTED provider from disk. A live `/model xai-oauth` session over `model.provider: auto` therefore still reached the discovery chain and billed Nous. _try_payment_fallback now takes the route's main_runtime snapshot; disk is the fallback only when no session runtime exists. * After a configured fallback was quarantined mid-request (401, refresh failed), the second pass went straight to the discovery chain, which the new gate refuses — so later CONFIGURED entries never ran and the original error was re-raised. The second pass now re-walks the task chain and main chain (the quarantined entry is unhealthy and skipped) before discovery. * current_provider_owns_vendor dropped ids detect_vendor could not classify, so Bedrock's 15-id catalog (14 unclassified `us.anthropic…`) looked exclusively DeepSeek and `/model deepseek-v4-pro` stuck on Bedrock. An unclassified id now counts as evidence of a multi-vendor catalog: ownership requires every id to classify to the one vendor. |
||
|
|
7e0b5cd235 |
fix(auxiliary): auto never bills a provider the user did not select
With a main provider selected, an unusable main route (expired xAI/Codex OAuth token, 401/402/429 mid-session) fell through the built-in discovery chain (OpenRouter -> Nous -> custom -> api-key) and quietly ran every compression, title and memory-flush call on whichever OTHER account was still logged in. Reported as "using Grok on my Premium+ sub, my Nous Portal balance kept draining" — the chat visibly stayed on Grok while the side tasks were billed elsewhere, and re-logging into X did not help because the aux side never consulted the selected provider. The discovery chain is now reserved for installs with no selected main provider (`model.provider: auto` / unset). Otherwise the ladder is main -> auxiliary.<task>.fallback_chain -> fallback_providers -> refuse with a warning naming the dead provider and the fix. Both entry points gate on the same predicate: the resolve-time route and the mid-request payment/auth hop (_try_payment_fallback). Existing chain tests that asserted the hop now pin `provider=auto`, the one case where discovery is still the contract. |
||
|
|
6964eebd35 |
fix(deepseek): pass custom model ids through instead of rewriting them
The deepseek provider allow-listed ids by shape (two canonicals plus a ^deepseek-v<N> regex) and rewrote everything else to deepseek-flash. That swallowed the vendor's own deepseek-flash the day it shipped (#107206) and would do the same to any future id without a v<N> marker, while still letting shape-matching guesses such as deepseek-v4.1-flash through to a 400. Only the two ids DeepSeek actually retired (deepseek-chat, deepseek-reasoner) are remapped now; every other id the user typed reaches the wire as typed and the API's own error names the valid models. |
||
|
|
70ed17e3d8 |
fix(cli): same-provider /model on a session-only custom endpoint stays on it (#74143)
A bare `custom`/`local` session whose base_url is not the trusted config `model.base_url` re-resolved credentials from config on a same-provider switch and fell through to the OpenRouter default: the next turn hit openrouter.ai with an empty (or the custom) key. Keep the session endpoint and key when resolution comes back empty or on OpenRouter; a config-backed custom URL still wins so key/endpoint rotation is not pinned to a stale session. Diagnosed by fangliquanflq in #74143 / #71693; this is the minimal form of that fix at the same-provider credential step. |
||
|
|
2466684db5 |
fix(models): never auto-switch to a provider the user has no credentials for
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).
detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.
Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
|
||
|
|
7f487df266 |
fix(gateway): hermes gateway restart without a service unit no longer bails on linger
The linger hint explains a FAILED systemd unit restart. It ran unconditionally, so on any Linux login session with linger off and no unit installed the command printed the hint and exited 0 without stopping or starting anything. The dashboard/Desktop "Restart gateway" action spawns exactly this command and polls its exit code, so it reported success while the gateway never came back with the newly saved credentials. Gate the hint on an installed systemd unit; the no-unit path falls through to the detached restart as intended. |
||
|
|
e1f1c8a935 |
fix(dashboard): stopped gateway no longer reports a dead process's platform error
gateway_state.json preserves per-platform entries across restarts, so a gateway that once ran without TELEGRAM_BOT_TOKEN and then stopped kept reporting 'fatal / No bot token configured' on the Channels/Messaging pages even after the user saved a token — the Desktop showed 'Saved' on both fields next to the error. Only a live gateway's verdict describes current config; when no gateway is running the platform reads gateway_stopped with no error. |
||
|
|
aeecb110f8 |
fix(deepseek): deepseek-flash is the canonical Flash id; retired names fold onto it
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's model name is now `deepseek-flash` and /v1/models lists only it. Hermes still folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash` on the DeepSeek provider was rewritten, then the validator "auto-corrected" it back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash" on every switch. Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated deepseek-v4-* ids still pass through untouched. Builds on YipTszkwan's #107126 (earliest fix in the cluster). |
||
|
|
215fd0ecb9 |
fix(models): a model the current provider serves never re-routes to OpenRouter
detect_provider_for_model() walked static catalogs and then the OpenRouter catalog. Providers whose live catalog outruns the static list (Codex early-access ids, Nous Portal slugs, Ollama Cloud) had nothing to stop the ladder, so `/model gpt-6-astra` from a Codex session, or `/model glm-5.3-flash` from a Nous-sub-only session, silently rebuilt the agent on metered OpenRouter (or on a vendor the user has no key for). The current provider's disk-cached live catalog (cached_provider_model_ids, 1h TTL + SWR) is now consulted first: an exact or bare-name hit stays put and resolves to the provider's own spelling. Fixes the whole class at the shared choke point, covering /model, ACP, oneshot and the dashboard. Refs #97487 (sj-unit72 diagnosed the ollama-cloud instance of this). |
||
|
|
8435a3ae00 |
fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id: GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the API accepts it directly, and the older `deepseek-v4-flash` is server-side aliased onto it. Every DeepSeek model-id gate in Hermes keys off the `deepseek-v<N>` prefix, so the new id silently missed all four: * DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and omitted `extra_body.thinking`. The server then defaults to thinking-on, so the user's thinking toggle and `reasoning_effort` were quietly ignored. * `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the V-series regex), so the id a user picked never reached the wire and the config stored a different model than the picker advertised. * `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all instead of the real 1M window, capping the model at an eighth of its context before compaction kicked in. * `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream detector at its 180s default instead of the 600s reasoning-model floor. Verified live against api.deepseek.com: `deepseek-flash` answers 200 with `model: deepseek-flash`, accepts image input (the refresh folds vision into the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected with "The supported API model names are deepseek-flash, deepseek-v4-pro". Adds the id to all four gates plus regression coverage for each site. |
||
|
|
8068c09432 |
fix(desktop): verify the Windows update receipt against the checkout, not cwd
The Windows hand-off script runs from the PRE-update checkout, spawned from
HERMES_HOME by the Desktop.
|
||
|
|
d783c312a7 |
fix(plugins): run gh auth token under the noninteractive git env
The MCP-catalog noninteractive contract test asserts every subprocess spawned during a git install carries GIT_TERMINAL_PROMPT=0 and a closed stdin; the gh token probe was spawned with the inherited env. Use the same hardened env (plus GH_PROMPT_DISABLED) so gh cannot open a browser/device flow either, and let the contract accept a stdin fed by input= (credential fill writes its request and closes). |
||
|
|
9886f6e53b |
feat(plugins): install plugins from private git repos using the user's stored credentials
`hermes plugins install owner/private-repo` failed for every private repository even when the same user could `git clone` it from their shell: the hardened `noninteractive_git_env` disables credential helpers, askpass and global config (so a hostile repo can't make our plumbing prompt or hang), which also blocks the user's own stored credential. The result was "could not read Username" or a 60s hang on a GUI askpass, with no hint about how to authenticate. New `hermes_cli/git_credentials.py` resolves a credential up front from sources the user already owns — GITHUB_TOKEN/GH_TOKEN (profile-scoped), `gh auth token`, then `git credential fill` against their configured helpers with prompting disabled (any host) — and hands it to git as a one-shot `http.<origin>/.extraheader` via the GIT_CONFIG_* env block. Nothing lands in the URL, `.git/config` or install metadata. The same path covers `plugins update`, catalog MCP git installs and profile-distribution staging, which share the same hardened env and the same failure. A private-repo clone with no credential now fails fast with an actionable hint. |
||
|
|
f6ddd89692 |
fix(dashboard-auth): offer every configured provider in native sign-in (#107018)
Desktop opens /auth/native/authorize without naming a provider. The empty-provider auto-select filtered password providers out of the candidate set, so a deployment with one OAuth provider plus username/password went straight to OAuth and never offered the password option — even though native sign-in brokers password providers through /login since |
||
|
|
6f3e630b47 |
fix(auth): honor Desktop-saved model.key_env for registry providers
The Desktop settings UI persists a registry provider's API key as a credential pointer -- model.key_env in config.yaml pointing at an env var in $HERMES_HOME/.env (e.g. HERMES_CUSTOM_LMSTUDIO_API_KEY) -- while keeping model.provider on the registry id. _resolve_api_key_provider_secret resolved registry providers exclusively from PROVIDER_REGISTRY[...].api_key_env_vars, so the UI-saved key was silently ignored; for lmstudio the runtime then substituted the no-auth placeholder (dummy-lm-api-key) and auth-enabled LM Studio servers rejected every request with an opaque, retryable-flagged 401. Fix: consult model.key_env (only when config.yaml's main model targets the provider being resolved, so the pointer never leaks across providers) before the registry env vars, running the value through the same _usable_declared_secret prefix validation. Precedence for the canonical env vars and the credential pool is unchanged; a valid canonical key still wins when no pointer is set. Complements #75373, which covers providers.<id>.key_env at the provider level; the Desktop save path writes the pointer at the model level, which that PR does not consult. Fixes #106336 |
||
|
|
b7bef04861 |
fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver every surface funnels through, `kanban_db.create_task`: board-project inheritance now runs only when the caller left `workspace_kind` open (`None`), and `workspace_kind` defaults to scratch after that check. The tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and `board=` scoping but drops its handler-local sentinel logic, since the resolver now owns the rule; the `self_task` project inheritance for dispatcher-owned workers is unchanged. Sibling surfaces had the same bug through the same line and are fixed by the same change: - CLI `hermes kanban create --workspace scratch` on a project-scoped board produced a project worktree; `--workspace` no longer defaults in argparse so the resolver can tell "omitted" from "scratch". - Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same; `CreateTaskBody.workspace_kind` defaults to `None` for the same reason. - `kanban_swarm.create_swarm` threads `None` through for consistency. Tests: one resolver invariant in test_kanban_board_project.py (explicit scratch stays scratch, omitted still inherits) and the salvaged tool test folded into a single parametrized matrix over scoped/unscoped target boards. |
||
|
|
53b9615ea6 |
test(hermes_cli): keep one opencode-go delisting invariant; record the delisting beside the keyed-suffix set
Drop the floor-pin test from #106290 (a snapshot of the static table); the end-to-end merge test through provider_model_ids with the real curated floor already fails if the slug is re-added. Note beside _OPENCODE_FREE_KEYED_SUFFIX_MODELS why ox-alpha-free stays excluded from the keyless catalog even though it left the curated floor. |
||
|
|
b2cae54035 |
fix(hermes_cli): drop delisted ox-alpha-free from opencode-go curated floor
The Go relay (GET /zen/go/v1/models) delisted ox-alpha-free 2026-09-09, but _PROVIDER_MODELS["opencode-go"] still carried it. _profile_live_catalog merges the curated floor into the live list (live-first for opencode-go), so the model picker kept offering a model that now 401s — the same failure class as #95914 (opencode-free / x-preview-f-free). Remove it from the floor and add two regression tests: an end-to-end merge test through provider_model_ids with the real floor (fails if a stale floor resurrects it) and a floor-pin test asserting the known-delisted model stays out of the offline fallback. (cherry picked from commit 091fd85865caa09e828928f93ee985dae7e9580d) |
||
|
|
e306374e20 |
test(models): pin native Anthropic picker to aggregator flagships
The curated fallback is the only list subscription tokens and lagged /v1/models responses see. Assert Opus 5 and Fable 5.1 stay on that list, newest-first, including when live discovery omits them. |
||
|
|
ac087e6ada |
fix(console): cancel/timeout interrupts the command's agent and waits for the worker
Hermes Console ran each command on a ThreadPoolExecutor and cancelled only the asyncio waiter. A command that forks an AIAgent (`curator run --consolidate`) kept its worker thread and its in-flight provider request alive after the prompt said "cancelled" — a llama.cpp generation kept decoding for 30+ minutes and held the inference slot (#106179). Root cause: the host owning the thread never knew about the agent created deep inside the synchronous command, so it could not call the existing cooperative `interrupt()` path that closes the request sockets. Fix: `agent/interrupt_scope.py` gives the host an `InterruptScope`; the console binds it around the worker (ContextVar), and every `AIAgent.run_conversation()` registers itself with the bound scope for the turn. On cancel, timeout and disconnect the console calls `scope.cancel()` (hard-interrupts every registered agent; an agent registering after the cancel is interrupted on entry so the cancel cannot lose the race with a turn that has not started) and awaits the worker Future with a 10s bound before reporting cancelled/timeout. Queued-but-unstarted work is dropped via `Future.cancel()` alone. Live repro (fake OpenAI-compatible provider blocking like llama.cpp, real /api/console, real curator dispatch, real AIAgent + direct request path): origin/main at the "cancelled" frame -> request_exited=false, worker_exited=false; with this change -> both true, provider saw the peer close, prompt reported cancelled 0.27s after the frame. Salvage of #106197 by @kyssta-exe (executor-future handle, cancel/timeout/ disconnect propagation) and #106320 by @Xixiartemis (deterministic lifecycle regression: terminal(cancelled) => no owned request or worker remains live; interrupt-on-late-registration). Both rebuilt slimmer: #106197 keyed its fallback on Future.cancel() returning False, but the handle it held was run_in_executor's asyncio wrapper, whose cancel() returns True while the thread keeps running, so its thread-name abort registry was never consulted; #106320's command-scoped ownership model is folded into one small module hooked at the turn facade instead of a per-caller `bind_agent`. Co-authored-by: kyssta-exe <218078013+kyssta-exe@users.noreply.github.com> Co-authored-by: Xixiartemis <182932319+Xixiartemis@users.noreply.github.com> |