Commit Graph

62 Commits

Author SHA1 Message Date
Austin Pickett f6ddd89692 fix(dashboard-auth): offer every configured provider in native sign-in (#107018)
Desktop opens /auth/native/authorize without naming a provider. The empty-provider
auto-select filtered password providers out of the candidate set, so a deployment
with one OAuth provider plus username/password went straight to OAuth and never
offered the password option — even though native sign-in brokers password
providers through /login since 56f1afc834. The filter's rationale ("a password
provider can never be the target of the native flow", ed5e17f4b8) predates that
change.

Render a provider chooser when more than one interactive session provider is
registered. Each link re-enters the same validated authorize route with an
explicit provider, so the choice never leaves the native PKCE flow. A single
provider still auto-selects, and the chooser is emitted before any broker state
is allocated or any cookie set.

The desktop only shell-opens the authorize URL and never parses its response
(apps/desktop/electron/native-oauth-login.ts), so the 200 chooser page is safe on
existing desktop builds.

Supersedes #101713, #76941.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
Co-authored-by: Phuong Lambert <vmphuongit@gmail.com>
2026-09-09 21:35:08 -04:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 7443080d5a refactor(hermes_cli): dashboard_auth docstring/comment compaction (keep WHY); native_flow pending lookup helper; prefix warn-once helper; commands registry repack 2026-09-03 00:07:25 -07:00
Teknium ca8f68774a refactor(hermes_cli): dashboard_auth routes share _audit/_complete_login tails; middleware refresh phase helpers; commands cfg_get + compaction 2026-09-02 23:55:41 -07:00
Teknium b90311f0d8 refactor(hermes_cli): compact dashboard_auth docstrings (keep WHY/invariants); drop _lax_bare_attrs dup 2026-09-02 21:18:03 -07:00
Teknium c04eb54f9b refactor(hermes_cli): AST-neutral layout compaction of dashboard_auth + commands_platforms 2026-09-02 21:01:32 -07:00
Teknium a3d7a0c25b refactor(hermes_cli): compact dashboard_auth package (registry/native_flow/audit/prefix/docs) 2026-09-02 20:33:40 -07:00
Teknium 28f098540b refactor(hermes_cli): dashboard_auth routes/middleware share one provider-scan helper; compact cookies 2026-09-02 20:10:27 -07:00
Teknium f0f7f55f76 refactor(web): dashboard_auth/web_routers/kanban dedupe + _common helpers (resume — verified partial work) 2026-09-02 13:30:05 -07:00
Teknium 2f44998353 fix(dashboard-auth): a non-JWT bearer is "not my token", not "provider unreachable" (#94558)
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.

Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.

Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).

This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.

Refs #94558
2026-09-02 01:15:58 -07:00
Ben Barclay 56916841b5 refactor(dashboard-auth): replace PKCE cookie payload with base64url(JSON) codec (#99210)
The PKCE cookie's payload has now needed three serialization fixes at
the same spot: the original flat 'k=v;k=v' string tripped http.cookies'
\073 quoted form (dropped whole by strict cookie parsers like Go's
net/http — #83832 field case), and #99176 URL-encoded the whole flat
payload to stay inside the RFC 6265 cookie-octet set. The stacked
layers (single-encoded next=, ';' joins, whole-payload encoding, legacy
discriminator) were the recurring defect source.

Kill the bug class instead of patching it again: the payload is a dict
end-to-end and goes on the wire as base64url(JSON) — the urlsafe
alphabet is a strict subset of cookie-octets, and JSON framing means no
segment value can ever collide with a delimiter. parse_pkce_payload
keeps a three-rung compatibility ladder (base64url(JSON) -> oldest flat
form split-as-is -> #99176 unquote-then-split) for in-flight cookies
during a rolling upgrade (10-minute TTL); a new cookie hitting an old
server fails the OAuth state check and the user just retries.

The 'next' segment is stored as its plain validated path — no extra
encoding layer, so the post-login redirect Location is byte-for-byte
the original target.

Refs #99176, #84065.
2026-09-01 11:47:38 +10:00
Ben Barclay a65a517d04 fix(dashboard-auth): url-encode the PKCE cookie value so strict proxy hops stop dropping it (#99176)
The PKCE payload is a flat 'provider=...;state=...;verifier=...;next=...'
string. A raw ';' is a cookie-attribute terminator, so Python's
http.cookies emits the value in RFC 6265 quoted form with each ';'
escaped as the backslash-octal '\073'. Mainstream browsers echo that
form back verbatim and Python parsers decode it — the browser round
trip is fine. But '"' and '\' are outside the plain cookie-octet set,
and non-Python hops that re-serialize the Cookie header reject the
value and drop the cookie entirely: Go's net/http (Traefik middleware,
Authentik outposts, other gateways) refuses any cookie value
containing a backslash. The OIDC callback then 400s with "Missing
PKCE state cookie" even though the browser sent the cookie.

Field reproduction: support thread "Still unable to use Authentik for
signin with traefik" — devtools showed the browser sending the intact
quoted \073 cookie on /auth/callback while Hermes logged
missing_pkce_cookie behind a Traefik+Authentik chain.

Fix: URL-encode the whole payload in set_pkce_cookie (quote(payload,
safe='') — ';' becomes '%3B') so the wire value contains only
cookie-octets and no parser in the chain has anything to reject, and
decode through a single shared inverse, cookies.parse_pkce_payload(),
in BOTH readers: the OAuth /auth/callback and the native
password-login path (routes.login_submit), whose broker/provider
binding check would otherwise parse zero segments from the
newly-encoded value and silently disable itself.

Regression coverage: the wire-shape test pins the full cookie-octet
set (the '"'/'\' assertions are the ones a Go-parser hop fails
pre-fix), the round-trip tests drive the real /auth/login →
/auth/callback path, and the next= test pins the exact post-login
redirect byte shape. Native-flow broker assertions updated to decode
through parse_pkce_payload instead of substring-matching the raw wire
value.

Salvaged from #84065 (rebased onto current main, which gained the
SameSite=None PKCE attrs and the RFC 8252 native password flow since
the PR branched): kept main's _pkce_attrs cookie shape, extended the
fix to the login_submit reader the original PR predated, and reframed
the rationale — browsers do NOT truncate at the first ';' (there is
no literal ';' on the wire in the quoted form); the failing hop is a
strict middlebox cookie parser.

Closes #83832

Co-authored-by: Kailigithub <12250313+Kailigithub@users.noreply.github.com>
2026-08-31 15:57:21 +10:00
Ben Barclay 53ab03dc4d fix(auth): thread use_https into the native password-login PKCE clear
Merging main brought in the RFC 8252 native sign-in path for password
providers (#75808), added while this PR was open. Its loopback-code
branch calls clear_pkce_cookie() without use_https, which is now a
required keyword-only argument — so /auth/native/password-login raised
TypeError on the success path.

This is the same call-site class the PR already fixed at the other three
sites: the deletion must mirror the shape the setter emitted for the
active origin, or the browser keeps the stale PKCE cookie.

Caught by CI running the merge commit against main's newer
test_dashboard_auth_native_flow.py suite, which does not exist on the
branch. Three tests failed there and pass with this change.
2026-08-24 16:58:51 +10:00
Ben Barclay 9a91f7058c Merge remote-tracking branch 'origin/main' into fix/pkce-samesite-none-salvage 2026-08-24 16:48:06 +10:00
Ben Barclay 91191123ed docs(auth): correct the SameSite contract in the cookie source docs
The SameSite=None change updated the website docs but left two
source-level contracts asserting the opposite:

  - base.py: LoginStart.cookie_payload said cookies set there "MUST"
    be SameSite=Lax.
  - cookies.py: the module docstring said all three cookies are
    SameSite=Lax.

Both now describe the actual behaviour: session cookies stay Lax, the
short-lived PKCE cookie is SameSite=None; Secure over HTTPS and Lax
over plain HTTP. A provider author following the old base.py contract
would have had a documented reason to undo the fix.

Also records the forwarded_allow_ips caveat in cookies.py: uvicorn only
honours X-Forwarded-Proto from a peer inside forwarded_allow_ips
(default 127.0.0.1), so a TLS terminator reaching the dashboard from a
non-loopback address (a reverse proxy in its own container) leaves the
request looking like HTTP and the cookies written in their HTTP shape.

Docstrings only; no behaviour change.
2026-08-24 16:41:13 +10:00
fangliquan bd13d593a9 fix(dashboard-auth): correct authentication docs anchor 2026-08-23 16:09:14 -07:00
fangliquan 7cc92cb134 fix(dashboard-auth): replace stale insecure guidance 2026-08-23 16:09:14 -07:00
rainbowgits b2ade2388d fix(dashboard-auth): keep the provider registry alive across per-home plugin-manager unloads
The dashboard auth registry is process-global, but a bundled auth provider
was registered under the per-home plugin manager's scope and enrolled in that
manager's reverse-order teardown. A per-home manager is unloaded routinely
(profile-scoped dashboard activity, forced re-discovery), and that teardown
disposed the registration — emptying the auth registry for the whole process
and permanently disabling sign-in until restart.

Register dashboard-auth providers in the process-global slot as persistent
host-owned registrations kept out of per-home manager teardown, so a routine
unload can no longer disable authentication process-wide. Registration upserts,
so a forced re-discovery (e.g. a password change) still rotates the provider in
place. The test-only manager reset now clears the auth registry too, since
persistent registrations deliberately survive unload.

Fixes #91701
2026-08-23 15:49:44 -07:00
Buff Pesos 56f1afc834 feat(dashboard-auth): extend RFC 8252 native sign-in to password providers (system-browser autofill) (#75808)
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers

The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.

The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.

Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:

* /auth/native/authorize now accepts a supports_password provider:
  register the pending broker authorization as usual, then 302 the
  system browser to the interactive /login form with the opaque
  broker_state in the gateway's PKCE cookie (the same server-controlled
  channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
  broker handle, a successful credential check completes the pending
  authorization exactly like the /auth/callback native branch — mint
  the one-time loopback code, return the loopback redirect (validated
  loopback-only at authorize time) as `next`, clear the PKCE cookie,
  and set NO session cookies. A lapsed broker is a clean 400 telling
  the user to restart sign-in; a failed credential attempt leaves the
  pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
  session provider is registered (previously only for non-password
  providers), so the desktop selects the system-browser strategy for
  password-only gateways.

Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.

Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dashboard-auth): bind native password completion to the authorize-time provider

Review follow-ups for #75808:

* /auth/password-login now enforces that body.provider matches the
  provider recorded in the server-set PKCE cookie by
  /auth/native/authorize before completing a pending native
  authorization. /login renders a form for every session provider, so
  without this a native flow started for provider A could be completed
  with provider B's credentials, binding B's session into A's pending
  entry. The mismatch is rejected BEFORE credential verification (no
  session minted, no oracle) and preserves both the pending entry and
  the cookie, so the user can still submit the correct provider's form.
  Covered by a two-password-provider E2E regression test.

* Update the two docs spots that still said password-only providers do
  not advertise native_pkce (website desktop-native-signin guide and the
  auth_flows type comment in web/src/lib/api.ts).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: map contributor email for #75808 (buffpesos)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
2026-08-14 21:39:01 +00:00
doncazper 85020f2238 fix(plugins): isolate ownership by profile 2026-08-12 19:13:32 -07:00
doncazper 22af80bcfd feat(plugins): add ownership ledger unload lifecycle 2026-08-12 19:13:32 -07:00
x7peeps ed5e17f4b8 fix(auth): /auth/native/authorize 空 provider 自动选择不再统计会被拒绝的密码 provider
Fix #78906

当部署同时启用 basic 密码 provider 与一个 OAuth/OIDC session provider 时,
list_session_providers() 会把密码 provider 也计入 "exactly one candidate"
判断(密码 provider 虽是 session provider,但下一行就会因 supports_password
被原生 OAuth broker 流程拒绝),导致 len == 2、自动选择被跳过,桌面端
空 provider 登录返回 404 "Unknown provider: ''"。

修复:自动选择只在可 broker 的 provider(supports_session 且非
supports_password)中计数,与 /api/status 的 native_pkce 能力宣告使用同一
"brokerable" 定义;当没有任何可 broker provider 时保留原有选择逻辑,
让显式的 400 错误继续解释密码 provider 不支持原生 OAuth。

新增回归测试:basic+OIDC 并存时自动选中 OIDC、单 OAuth provider 自动
选中、多 OAuth provider 歧义 404、纯密码部署保留 400。
2026-08-11 16:00:55 -04:00
Ben Barclay 52629a5de6 fix(auth): make prefixed cookie deletions valid per cookie-prefix rules
Same bug class as the PKCE clear-shape mismatch, sibling call paths:
clear_session_cookies() and clear_sso_attempt_cookie() emitted
__Host-/__Secure- name deletions without the Secure attribute. RFC
6265bis prefix rules make such a Set-Cookie invalid — browsers reject
the header outright — so the __Host- session cookie deletions on
logout were silently ignored on HTTPS origins (masked in practice by
the 15-minute access-token TTL, but logout did not actually delete
the cookies it claimed to).

Extract one _clear_cookie_variants() helper used by all three clear
functions: prefixed variants always carry the attributes their name
demands (Secure, and Path=/ for __Host-), while the bare-name deletion
mirrors the corresponding setter's shape so it works on plain-HTTP
origins too. Adds a contract test pinning the deletion shapes.
2026-08-11 10:11:28 +10:00
Ben Barclay 5e57bf19a2 fix(auth): mirror the PKCE setter's cookie shape in clear_pkce_cookie
The SameSite=None change left clear_pkce_cookie() deleting every name
variant with `SameSite=None; Secure` unconditionally. Over loopback
HTTP the setter emits a bare-name cookie with `SameSite=Lax` and no
`Secure` — a Secure deletion on a plain-HTTP origin can be ignored by
the browser, leaving a stale PKCE cookie behind after callback/logout.

Thread use_https from the request into clear_pkce_cookie() and derive
both the set and clear attribute shapes from one helper (_pkce_attrs)
so they cannot drift apart again. The __Host-/__Secure- variants keep
`Secure; SameSite=None` regardless of origin — those names require
Secure to be valid at all and only ever exist on HTTPS origins.

Adds contract tests pinning the full Set-Cookie shape on both origins
for set and clear, and updates the dashboard cookie documentation
(which still claimed all cookies are SameSite=Lax).
2026-08-11 10:03:51 +10:00
Keith Ammon 7c7cd24d00 fix(auth): use SameSite=None for PKCE cookie over HTTPS to fix cross-site redirect dropping
Chromium intermittently drops SameSite=Lax cookies set on a 302 redirect
in a cross-site redirect chain (crbug 40508226). The OAuth flow sets the
PKCE cookie on /auth/login (302 → IDP), and the callback returns from the
cross-site IDP — exactly the scenario where Chromium drops the cookie,
causing 'Missing PKCE state cookie' on /auth/callback.

SameSite=None + Secure is the correct attribute for cookies that must
survive cross-site redirect delivery. Chromium processes these reliably.

Loopback HTTP (dev) stays SameSite=Lax since SameSite=None requires Secure.

Fixes #56750
2026-08-11 10:00:05 +10:00
teknium1 5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
teknium1 19055492aa fix: route stray HERMES_HOME hardcodes through get_hermes_home() (profile + native-Windows safety) 2026-07-29 09:33:48 -07:00
izumi0uu ccab46ca43 fix(dashboard): add lightweight /api/health liveness endpoint
/api/status is the only public liveness route, and its handler loads the
gateway config, probes gateway health, and counts sessions before it can
answer. That work is wrong for a readiness probe: a caller that only needs
to know the process is up pays for a cold plugin import tree.

Add /api/health, which returns process liveness, version, and the auth-gate
shape and touches nothing else.
2026-07-24 19:43:44 -05:00
Teknium 4d23b2238e fix(dashboard-auth): harden the public native-authorize surface
Two tightenings on /auth/native/authorize (a public pre-auth route):

- Per-IP pending cap (8): the broker store is capacity-bounded fail-closed
  at 256 entries with a 600s TTL, so one unauthenticated spammer could fill
  it and deny native sign-in gateway-wide for the pending window. Pending
  entries now record the requester IP and each address is capped well above
  any legitimate concurrent-login count; other addresses keep signing in.
- Loopback redirect_uri accepts IP literals only (127.0.0.1 / ::1):
  'localhost' can be re-pointed via the hosts file or a hostile resolver
  (RFC 8252 \u00a78.3 says to use loopback IP literals); the desktop always
  sends 127.0.0.1, so nothing legitimate used the name.

3 new tests: per-IP cap enforced, cap frees on TTL expiry, localhost
redirect rejected at the route.
2026-07-22 06:50:50 -07:00
Ben Barclay 7857d8737c feat(dashboard-auth): RFC 8252 native desktop sign-in (system browser + PKCE, no webview/cookies)
The Desktop app can now sign in to a gated gateway using the user's SYSTEM
browser and OAuth 2.0 for Native Apps (RFC 8252) instead of an embedded
Electron BrowserWindow, and authenticates with bearer tokens it holds itself
instead of relying on HttpOnly browser session cookies.

Why brokered: the upstream IDP (Nous Portal) binds client_id to the gateway
instance and only permits redirect_uris on the gateway's own origin, so a
desktop loopback redirect can't be a direct Portal client. The gateway
therefore acts as the authorization server TO the desktop and an OAuth client
TO the Portal, reusing the existing PKCE start_login/complete_login provider
path unchanged.

Server (Ben's dashboard-auth lane):
- native_flow.py: in-memory broker — binds the desktop's PKCE challenge to a
  completed Session, mints a single-use, short-TTL, PKCE-verified gateway
  authorization code. Constant-time compare, single-use (consumed before the
  PKCE check so a wrong verifier can't be retried), capacity-bounded.
- routes.py: GET /auth/native/authorize (starts the brokered PKCE login,
  loopback-only redirect_uri, S256-only), POST /auth/native/token (loopback
  code + verifier -> tokens in the JSON body, never Set-Cookie), POST
  /auth/native/refresh (desktop-held RT rotation). /auth/callback branches to
  mint a loopback code + 302 to 127.0.0.1 when a broker_state rides the PKCE
  cookie; the cookie/SPA path is untouched.
- middleware.py: the gate accepts Authorization: Bearer <access_token>,
  verified via the same verify_session provider stack (no cookie set/read),
  with the same "provider unreachable -> 503, not logout" semantics.
- web_server.py /api/status: advertise auth_flows (["cookie","native_pkce"])
  so clients can detect the capability; native_pkce only when a brokerable
  OAuth provider is registered.

Desktop (Ben's lane):
- native-oauth.ts: pure PKCE/capability/URL/callback/token helpers.
- native-oauth-login.ts: loopback-listener orchestration (system browser via
  openExternal, ephemeral 127.0.0.1 listener, state/PKCE verification), all
  I/O injected for testability.
- main.ts: capability-gated oauth-login IPC — native flow when advertised,
  automatic fallback to the existing embedded-webview cookie flow otherwise;
  tokens stored encrypted (safeStorage/OS keychain), REST + ws-ticket
  authenticated by bearer, transparent refresh, logout clears both shapes.

Tests: 18 server pytest (broker unit + full authorize->callback->token E2E +
cookieless bearer auth of a gated route + ws-ticket mint + capability
advertisement + refresh); desktop node --test/vitest for both pure modules
(PKCE, capability detection, callback CSRF, loopback round trip, timeout,
browser-open failure). Electron project typechecks clean.

Docs: website/docs/guides/desktop-native-signin.md.
2026-07-22 06:50:50 -07:00
Ben Barclay 05dea7be04 fix(mcp): complete OAuth through hosted dashboards 2026-07-17 04:50:47 -07:00
mark 96a0708448 fix(auth): preserve provider fallback during refresh 2026-07-14 07:02:05 -07:00
mark f9e35e6e94 fix(auth): route session refresh with provider hint cookie 2026-07-14 07:02:05 -07:00
Shannon Sands 3e24b16f56 fix(dashboard): support mobile OAuth login 2026-07-09 12:21:16 +05:30
izumi0uu ef79ad014d fix(dashboard): accept HA ingress prefix paths
Allow mainstream reverse-proxy path mounts to keep their X-Forwarded-Prefix when Home Assistant Supervisor ingress already consumes nearly the old 64-character budget. Keep validation bounded and keep rejected non-empty prefixes diagnosable with a deduplicated warning.

Constraint: HA Supervisor ingress prefixes are 63 chars before add-on subpaths, so the old 64-char cap dropped valid dashboard deployments.

Rejected: remove the length cap entirely | a bounded header budget is still a conservative validation guard.

Confidence: high

Scope-risk: narrow

Directive: Keep prefix validation centralized in hermes_cli.dashboard_auth.prefix so auth routes, cookies, and SPA asset rewriting agree.

Tested: python probe for the 73-char HA ingress prefix; scripts/run_tests.sh tests/hermes_cli/test_dashboard_auth_prefix.py -q; .venv/bin/python -m pytest tests/hermes_cli/test_web_server.py -k 'spa_assets_are_read_as_utf8' -q; python -m ruff check hermes_cli/dashboard_auth/prefix.py tests/hermes_cli/test_dashboard_auth_prefix.py; git diff --check

Not-tested: full test suite
2026-07-06 03:18:02 -07:00
Shannon Sands 4493bba901 Add dashboard Hermes console websocket 2026-07-03 20:18:00 +05:30
teknium1 61f56d27db refactor(dashboard-auth): drop redundant _interactive_providers helper
list_session_providers() already filters on supports_session=True, so the
new helper re-filtered an already-filtered list. Call it directly at the
single auto-SSO call site.
2026-06-29 04:25:18 -07:00
Ben f5ecbe1ec6 feat(dashboard): auto-initiate portal SSO redirect on unauthenticated load
When the dashboard gateway has no local session cookie, it rendered a
click-through /login interstitial — even though the Nous portal's
/oauth/authorize auto-approves any current member of the dashboard's org
and is a silent 302 when the user already holds a portal session. For the
common case (clicking a hosted-agent dashboard link while signed in to the
portal) that interstitial click is pure friction.

This makes the gate auto-initiate the OAuth redirect on an unauthenticated
HTML document load instead of rendering the interstitial, when exactly one
interactive provider is registered. A one-shot loop-guard cookie
(hermes_sso_attempt, 60s TTL) ensures that a genuinely absent portal
session (the portal bounces back still-unauthenticated) falls back to the
/login page after exactly one bounce rather than ping-ponging forever. The
marker is cleared on a successful callback and whenever the gate falls back
to /login.

Security: this removes a human CLICK, not a security check. The redirect
lands on the existing /auth/login route and runs the unchanged PKCE
auth-code flow; token verification, audience checks, redirect-URI match,
and org-membership checks are all untouched. /api/* fetches still get the
401 JSON envelope (never a 302 a fetch() would follow opaquely), and with
two or more providers the /login chooser still renders.

Phase 1 of the cloud-auto-discovery work.
2026-06-29 04:25:18 -07:00
Nacho Avecilla dbe734beff fix(dashboard-auth): exclude non-interactive providers from interactive login surfaces (#53239)
* Return None instead of erroring on drain login failure

* Fix login on drain

* Remove login for drained endpoints flow and clean the code

* chore: drop unrelated credits changes from this PR

* Remove extra comments that were not really necessary
2026-06-27 10:08:13 +10:00
Ben cb9cb6ba1c feat(dashboard-auth): generic non-interactive API-token capability
Task 2.0a of the safe-shutdown drain-coordination plan. Widens the dashboard
auth framework GENERICALLY to support non-interactive (service-to-service)
bearer-token auth, mirroring the existing supports_password precedent. This is
a reusable capability — any future machine-credential provider plugs in without
core changes (decisions.md Q-C). The drain bearer-secret plugin (Task 2.0b) is
the first consumer, not the definition.

- base.py: add TokenPrincipal dataclass (the token analog of Session) +
  supports_token capability flag + verify_token() on the ABC (default raises
  NotImplementedError so a misconfigured provider fails loud). Contract mirrors
  verify_session stacking: return None for unrecognised tokens (never raise),
  raise ProviderError only on a genuine backing-store outage.
- registry.py: list_token_providers() — the supports_token subset, in
  registration order. Empty when none registered (token routes fail closed).
- token_auth.py (new): route-agnostic seam. Routes opt in via
  register_token_route(exact path); token_auth_middleware owns the auth
  decision for those routes only — authenticate via stacked providers, attach
  request.state.token_principal + token_authenticated, pass through. 401 on
  missing/unrecognised token, 503 when a provider was unreachable, untouched
  passthrough for non-token routes. Fails closed (never open).
- web_server.py: install the seam OUTERMOST (registered last → runs first).
  Both downstream gates (legacy auth_middleware + gated_auth_middleware) honour
  request.state.token_authenticated and skip enforcement, so a token-authed
  service request is never bounced to /login.
- audit.py: TOKEN_AUTH_SUCCESS / TOKEN_AUTH_FAILURE events.

Tests: tests/hermes_cli/test_dashboard_token_auth.py — ABC flag default,
verify_token NotImplementedError, registry filter, bearer extraction
(case-insensitive scheme, malformed/non-bearer → ""), provider stacking
(first-match-wins, unreachable-remembered, unreachable-then-valid, buggy
provider doesn't crash the gate), and the seam's passthrough/401/503/
fail-closed behaviour. 29 new tests; full dashboard-auth suite 169 passed.

Intentionally deferred:
- The concrete shared-bearer-secret provider plugin — Task 2.0b.
- The begin/cancel-drain endpoint that registers itself as a token route —
  Task 2.1.

Build status: dashboard-auth + plugin-hook suites green.
2026-06-26 00:47:19 -07:00
Ben c34840e22e fix(cron): serve /api/cron/fire on the dashboard app (hosted-agent surface)
Live-test finding: the Chronos fire webhook was only on the APIServerAdapter
(aiohttp), but hosted agents expose `hermes dashboard` (the FastAPI web_server
app on :9119) as their public URL — NOT the api_server adapter. So NAS's relay
callback to {callback_url}/api/cron/fire could never reach the verifier on a
hosted agent (the exact target environment). Two layers were wrong:

1. Wrong server: /api/cron/fire didn't exist on the dashboard app. Added
   cron_fire_webhook there, alongside the existing /api/cron/* dashboard routes.
   It resolves the job's profile (_find_cron_job_profile) and runs fire_due via
   the resolved provider under the cron-profile retarget lock
   (_fire_cron_job_for_profile, mirroring _call_cron_for_profile) so the CAS
   claim + run_one_job operate on the right profile's jobs.json. Runs with no
   live adapters (delivery falls back to the per-platform send path, like the
   desktop cron path). 202 + background so a long turn never trips NAS's
   timeout; the store CAS de-dupes a NAS retry. job-not-found -> 200 "gone".

2. Auth gate: the dashboard auth middleware 401s any non-cookie request before
   the handler runs. Added /api/cron/fire to the shared PUBLIC_API_PATHS so the
   NAS bearer-JWT callback reaches the verifier — the JWT (purpose=cron_fire),
   not the cookie, is the real gate. One shared frozenset feeds both the
   loopback and OAuth middlewares, so no drift.

Kept the APIServerAdapter route too (valid self-host api_server surface).
Contract doc updated to name the dashboard app as the hosted-agent callback
surface.

Tests: test_cron_fire_dashboard (6) — route registered on the dashboard app,
in PUBLIC_API_PATHS, 401 on bad token WITH the cookie gate engaged (proves it's
reachable past the gate + JWT is the gate), 400 missing job_id, 200 gone for
unknown job, 202 + fire_due invoked for the resolved profile on a valid token.
Full hermes_cli + cron + chronos + webhook suites green (7637).

Why the original tests missed it: the api_server webhook test built an
APIServerAdapter client directly and never asserted which server the hosted
public URL exposes — green-but-wrong-integration. The new test pins the route
to the dashboard app.
2026-06-19 12:43:30 +10:00
Ben Barclay 7df3aa34b1 fix(dashboard-auth): warn when public_url override is silently rejected (#43214)
A non-empty HERMES_DASHBOARD_PUBLIC_URL / dashboard.public_url value that
fails URL validation (overwhelmingly: a missing http(s):// scheme, e.g.
"hermes.domain.com") was silently discarded by resolve_public_url(),
falling back to reconstructing the OAuth redirect_uri from request
headers. Behind a reverse proxy that doesn't forward X-Forwarded-Proto
reliably, that yields an http:// callback even though the operator
explicitly set the public URL — with no signal as to why (#42780).

Emit a deduplicated operator-facing WARNING (once per distinct value,
since resolve_public_url runs per request) naming the offending value
and the required scheme. Turns a silent footgun into a self-diagnosing
one; behaviour is otherwise unchanged.

Tests assert the warning fires for a scheme-less value, is deduplicated
across repeated calls, and stays silent for a valid value — all three
fail without the fix.
2026-06-10 12:14:57 +10:00
Ben 439f53cab8 fix(desktop): gate OAuth remote connect on AT-or-RT, not access token alone
The desktop OAuth remote-gateway path gated connectivity on
hasOauthSessionCookie(), which checks only the access-token cookie
(hermes_session_at, ~15 min TTL). The moment that cookie's Max-Age
lapsed, Electron's cookie jar dropped it and both resolveRemoteBackend()
and sanitizeDesktopConnectionConfig() reported "not signed in" — forcing
a full IDP re-login every ~15 min — even though a valid 24h refresh-token
cookie (hermes_session_rt) was sitting in the same jar.

The desktop OAuth code (2026-06-04) was written against the obsolete
"contract v1 issues no refresh token" model, two days after #37247
re-introduced server-side transparent refresh: Portal now issues a 24h
rotating, reuse-detected refresh token, and the gateway middleware
(_attempt_refresh) rotates a fresh AT from the RT on the next
authenticated request. So an expired-AT/live-RT session is fully
connectable — the desktop just never let the request through.

Fix:
- connection-config.cjs: add RT_COOKIE_VARIANTS + cookiesHaveLiveSession()
  (true when EITHER a live AT or RT cookie is present). Keep
  cookiesHaveSession() AT-only for callers that need that specific signal.
- main.cjs: add hasLiveOauthSession(); resolveRemoteBackend()'s oauth
  branch now early-outs only when NEITHER cookie is present, otherwise
  uses the ws-ticket mint as the authoritative liveness probe (that POST
  carries the RT cookie and triggers the server-side AT rotation). A real
  401 still surfaces as needsOauthLogin. Settings indicator + oauth-logout
  report against the same AT-or-RT notion.
- Remove the stale "contract v1 / NO refresh token" docstrings in
  cookies.py and the verify_session comments in the Nous provider that
  contradicted #37247.

Tests: +57 lines in connection-config.test.cjs covering the RT-only
"still connectable" case. node --test: 32/32. dashboard-auth +
nous-provider Python suites: 223/223.

Note: server-side files (hermes_cli/dashboard_auth/, plugins/dashboard_auth/)
are comment/docstring-only here, but this touches outside apps/desktop/ so
it needs Teknium review.
2026-06-04 22:18:46 -07:00
Ben 616c0a36b6 fix(dashboard-auth): don't abort verify chain on one provider's ProviderError
The gated dashboard verifies a session cookie by trying each registered
DashboardAuthProvider's verify_session in turn (the session cookie stores
only the access token, not which provider issued it). A provider that
doesn't recognise a token returns None; a provider whose IDP/JWKS is
unreachable raises ProviderError.

The loop used to return HTTP 503 on the FIRST ProviderError, before any
later provider got a turn. With multiple providers stacked, that means an
unreachable IDP for a session you didn't even use blocks login through a
different, reachable provider.

Concrete repro: a self-hosted-OIDC session hits the 'nous' provider first
(registered earlier); nous tries to reach Nous Portal's JWKS, which is
unreachable in a self-hosted deployment, so it raises — and the gate
503s before the 'self-hosted' provider can verify the token. Hit live
while testing the new self-hosted OIDC plugin against a local Keycloak.

Fix: a ProviderError from one provider is logged and the loop continues
to the next. A 503 is returned only if NO provider verified the token
AND at least one was unreachable — distinguishing a transient IDP outage
(don't force a needless re-login) from a token that's genuinely invalid
(fall through to refresh/relogin). Single-provider behaviour is
unchanged.

Tests: adds an _UnreachableProvider stub and three cases — unreachable
provider first must not block a working second; all-unreachable still
503s; reachable-but-unrecognised falls through to 401/relogin (not 503).
Mutation-tested: reverting the fix makes the first case fail with the
exact 503 bug.
2026-06-04 03:23:45 -07:00
Ben ed9e8ba097 feat(dashboard-auth): add pluggable password (non-redirect) login
The dashboard auth gate was OAuth-only: a DashboardAuthProvider could
authenticate only via a redirect to an IDP (start_login -> /auth/callback
-> complete_login). There was no first-class path for username/password
auth, so self-hosters who just want a password on their dashboard had no
clean option short of an external OAuth IDP.

Extend the provider framework with a parallel, non-redirect front door
that converges on the same Session + cookie + refresh machinery:

  - base.py: add the optional supports_password flag and
    complete_password_login(username, password) -> Session (default
    raises NotImplementedError so an OAuth-only provider that forgets the
    flag fails loudly). Add InvalidCredentialsError. OAuth providers are
    unaffected (flag defaults False; the method is never called).
  - routes.py: add POST /auth/password-login, mirroring the cookie-minting
    tail of /auth/callback but skipping PKCE/state/code. Returns JSON
    {ok, next} (the form POSTs via fetch). Generic 401 for both unknown
    user and wrong password (no enumeration oracle); 404 hides whether a
    provider exists or supports passwords; per-IP sliding-window rate
    limit (10/min -> 429). /api/auth/providers now reports
    supports_password so the login page can branch.
  - middleware.py: allowlist /auth/password-login (a bootstrap route).
    verify/refresh/revoke/ws-tickets/logout need zero changes — a password
    session is just a Session with provider-minted opaque tokens.
  - login_page.py: render a credential form (instead of a redirect button)
    for supports_password providers, wired by a small inline script that
    POSTs to /auth/password-login and navigates on success. OAuth-only
    pages stay script-free.
2026-06-04 01:02:25 -07:00
kshitijk4poor e114b31eda test(dashboard): direct unit coverage for internal WS credential + docstring fix
Follow-up to Ben's PR #37892. Adds a TestInternalCredential block to
test_dashboard_auth_ws_tickets.py exercising the mint-once stability,
multi-use, unminted-rejection, empty-value, wrong-value, reset-and-remint,
and ticket-store-independence branches directly (previously only covered
indirectly via _ws_auth_ok, which left the unminted and empty-value
branches unexercised).

Also corrects the consume_internal_credential docstring: the returned
identity dict is discarded by the current _ws_auth_ok caller (which only
needs the boolean outcome), so the prior 'carry it into its session log'
wording over-promised.
2026-06-02 23:43:27 -07:00
Ben fd1ec8033d fix(dashboard): authenticate server-spawned PTY child WS with a process-internal credential
The embedded-TUI PTY child attaches to two server-internal WebSockets:
/api/ws (its primary JSON-RPC gateway backend) and /api/pub (the event
sidecar). Both URLs are built server-side in web_server.py and handed to
the child via its environment.

In OAuth-gated mode (auth_required=true, every hosted Fly agent), _ws_auth_ok
unconditionally rejects the legacy ?token=<_SESSION_TOKEN> path — a leaked
session token must not grant WS access once the gate is engaged. But
_build_gateway_ws_url() still only emitted ?token=, with no gated-mode
branch (its sibling _build_sidecar_url had been given a ticket branch; the
gateway-url builder was missed). So the TUI child's /api/ws upgrade was
rejected 4401 -> 'gateway websocket connection failed' -> 'gateway startup
timeout', leaving the embedded chat unusable on every gated deployment.

A single-use 30s browser ticket is the wrong shape for this link: the child
reads its attach URL once at startup and reuses it on every reconnect, and
on a slow cold boot it may not dial within the TTL. (_build_sidecar_url's
own docstring already flagged this fragility.)

Fix: add a process-lifetime, multi-use internal credential to
dashboard_auth.ws_tickets (internal_ws_credential / consume_internal_credential),
minted once per process and NEVER injected into the SPA — it only leaves the
process via a spawned child's env, so browser-side XSS can't read it, and a
leak grants no more than a ticket already does. _ws_auth_ok accepts it via
?internal= in gated mode only. Both _build_gateway_ws_url and
_build_sidecar_url now use it, so the child can reconnect both sockets.

Loopback / --insecure behavior is unchanged (still ?token=).

Needs review: touches _ws_auth_ok + dashboard_auth (core auth surface).
2026-06-02 23:43:27 -07:00
Ben Barclay c10ccaaf51 feat(dashboard-auth): rotate dashboard sessions via refresh token (#37247)
* feat(dashboard-auth): rotate dashboard sessions via refresh token

The dashboard auth-code grant now issues a 24h rotating refresh token
(server side: NousResearch/nous-account-service#293). This wires up the
Hermes client half so an expired access token is transparently refreshed
instead of bouncing the user to /login every 15 minutes.

plugins/dashboard_auth/nous:
- refresh_session() now POSTs grant_type=refresh_token to Portal's token
  endpoint and returns a Session carrying the ROTATED refresh token (was
  an unconditional RefreshExpiredError under the old "no RT in V1"
  contract). The RT is sent in BOTH the request body (Portal's schema
  requires it there) and the X-Refresh-Token header (log redaction) —
  verified against the #293 preview deploy: header-only is rejected as
  invalid_request, body is accepted.
- A 400 from Portal (expired / revoked / reuse-detected) maps to
  RefreshExpiredError so the middleware forces a clean re-login; network
  errors map to ProviderError; empty RT fast-fails without a network call.
- complete_login now captures the initial refresh token Portal returns
  (forward-tolerant: empty string if a deploy omits it).
- Extracted the shared token-response handling into
  _token_response_to_session, parameterised on the 400 exception type so
  the auth-code path raises InvalidCodeError and the refresh path raises
  RefreshExpiredError.
- revoke_session stays a best-effort no-op: Portal exposes no public
  token-endpoint revocation grant (revocation is the authenticated
  /sessions UI, keyed by sessionId+userId), so logout is cookie-clearing
  and the 24h session expires on its own. Documented for a future
  revoke grant.

hermes_cli/dashboard_auth/middleware:
- On an expired/invalid access token the gate now attempts refresh via
  the session's RT BEFORE forcing re-login. On success it serves the
  request and re-sets the rotated cookies on the response (mandatory:
  Portal rotates the RT every refresh and reuse-detects, so a stale RT
  cookie would revoke the whole session on the next refresh). On
  RefreshExpiredError (or no RT) it falls through to clear-and-relogin.
- ProviderError during refresh (Portal unreachable) forces a clean
  re-login rather than 500-ing the request.
- Uses the existing REFRESH_SUCCESS / REFRESH_FAILURE audit events.

Validation:
- 176 dashboard-auth unit/integration tests pass.
- Live E2E against the #293 preview deploy: refresh_session(bad rt) ->
  RefreshExpiredError through the real token endpoint; live JWKS fetch +
  RS256 verification rejects a forged token; empty-RT fast-fail. The
  successful happy-path rotation is covered by unit tests (a live run
  needs an interactive browser OAuth round trip + registered agent:*
  client).

Depends on: NousResearch/nous-account-service#293 (server-side RT issuance).

* fix(dashboard-auth): use Portal's x-nous-refresh-token header name

The refresh-token header must match Portal's REFRESH_TOKEN_HEADER exactly
("x-nous-refresh-token"); the initial cut used "X-Refresh-Token", which
Portal silently ignores (harmless since the RT is also in the body, which
is what the schema requires — but the header redaction was a no-op).
Confirmed against the NAS token route + re-validated live against the
#293 preview deploy.

* fix(dashboard-auth): refresh session when access-token cookie has been evicted

The gated middleware bounced users to /login the instant the access-token
cookie was absent, without ever consulting the refresh token:

    at, _rt = read_session_cookies(request)
    if not at:
        return _unauth_response(...)   # bailed here

This made transparent refresh effectively dead for the common case. The
access-token cookie is set with Max-Age = access_token_expires_in (~15 min),
so a real browser EVICTS hermes_session_at the moment the token lapses while
hermes_session_rt persists (30-day Max-Age). From that point the browser
sends only the refresh-token cookie — and the old guard rejected it before
_attempt_refresh could run. The _attempt_refresh path only fired for a
present-but-invalid access token, which never happens in a browser.

Fix: only hard-bounce when NEITHER cookie is present. A request carrying
just the refresh token now skips verification (no AT to verify) and flows
into the existing refresh path, which rotates both cookies and serves the
request transparently. A dead/expired RT still raises RefreshExpiredError
and falls through to clear-and-relogin.

This failure mode escaped the original tests + manual refresh button because
both kept the access-token cookie present; only a real browser evicting the
cookie at Max-Age exposes it. Added 3 regression tests covering: AT-evicted +
RT-present (transparent refresh), no-cookies (still bounces), and RT-only with
a dead RT (clean 401, no 500).
2026-06-02 21:16:41 +10:00