dfd4aa4a94c726c8dab9efcb6dfb2c375e986dd2
1985 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
70d0f556d7 |
fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme, locale strings and the docs — used ⚕, the staff of Asclepius (medicine). Hermes carries the Caduceus ☤. The ASCII-art logo was already correct. Mechanical swap across 60 files (no logic change); both glyphs are East-Asian-width Neutral so no layout shifts. Skins that set their own `response_label` / `goodbye` are unaffected. Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574, redone against current main. Fixes #9565 |
||
|
|
44ce128a27 |
fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes #9636
|
||
|
|
850c48cd84 |
feat(skills-hub): tap K-Dense and OpenScience scientific skills under one "science" bucket
~480 scientific research skills become searchable/installable through the Skills Hub with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0). A new optional tap-level `bucket` key stamps extra["category"] on every skill from a tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub category; a sidecar grouping still wins when present. Both repos stay at community trust (not in TRUSTED_REPOS) so the guard scans every install. Re-grafted from #60559 onto the post-split tools/skills_hub_github.py. |
||
|
|
53c57871d6 |
feat(plugin-catalog): add snyk — Snyk MCP server + snyk-security-scan skill
Snyk lands as a standalone Agent Plugins v1 package (NousResearch/hermes-plugin-snyk, pinned 2a41a07f) instead of an optional-mcps entry: the package carries the pinned `npx -y snyk@1.1306.0 mcp` stdio launch with CLI analytics disabled AND the workflow skill that tells the agent when to use which scanner, so a single `hermes plugins install snyk` gives both the tools and the playbook. Third-party product integrations ship outside the core tree per the contribution rubric. Supersedes the optional-mcps manifest from #73860 (same pin, same telemetry posture, tool-pruning rationale moved into the skill). |
||
|
|
be2f7e9c36 |
feat: curator prunes unused skills at 30 days (was 90), stale at 14
A skill nobody has loaded in a month is prompt weight, not knowledge, and archival is recoverable (`hermes curator restore`). Defaults move stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so an explicitly customized window is preserved. `hermes curator prune` now defaults --days to curator.archive_after_days instead of a hardcoded 90 so the manual and automatic paths agree. |
||
|
|
f364c19775 |
feat(plugins): touchdesigner ships as a catalog plugin (MCP server + skill), optional skill retired
The twozero TouchDesigner integration now lives in one installable unit: plugin-catalog/touchdesigner.yaml points at NousResearch/hermes-plugin-touchdesigner (portable Agent Plugins v1: mcp.json registers the twozero Streamable HTTP hub, skills/ carries touchdesigner-mcp). `hermes plugins install touchdesigner` + `hermes plugins enable td` replaces the optional skill whose setup.sh hand-wrote an mcp_servers block. The manifest name is `td` because Hermes names portable MCP tools mcp__agent_plugin_<name>_<hash>__<server>__<tool>; twozero's longest tool under a `touchdesigner` namespace is 71 chars, past the 64-char provider function-name cap. optional-skills/creative/touchdesigner-mcp and its generated docs (bundled, optional, zh-Hans) are removed; catalog tables, sidebar and kanban-video-orchestrator references are updated to point at the plugin. Supersedes #68607 (MCP-catalog-only approach). |
||
|
|
b7b35a84b7 | docs: remove the Nous guest free-tier guide (#108991) | ||
|
|
5c4e08e8df |
docs(multiplex): inbound-port platforms under the multiplexer
Replace the 'a secondary must not enable a port-binding platform' rule with the shared-listener contract and a per-platform URL table (Twilio, LINE, Teams, BlueBubbles, Microsoft Graph, WhatsApp Cloud, WeCom callback, Feishu webhook), plus the status/dashboard surfaces that print the URL. |
||
|
|
6cd4fbd640 |
docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch advance is provisional, restore once, never twice, grace, opt-out, paused never catches up, same on standalone and multiplexed); the user guide describes the catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md lists it as a hardening invariant. |
||
|
|
df51797e2e | feat(cron): let planned downtime skip missed recurring runs | ||
|
|
be82d52cb7 |
docs: migrating from per-profile gateways to the multiplexer
What `hermes update` does, blockers and fixes, URL change for inbound-port profiles, the post-create restart reminder, rollback, and the `gateway migrate` reference row. |
||
|
|
5aa17c0590 | docs(multiplex): list the newly per-profile caches in the isolation table | ||
|
|
bf867d3c74 |
chore(skills): ship dynamic-workflow as an optional skill
Install with hermes skills install official/autonomous-ai-agents/dynamic-workflow. Orchestration-campaign guidance is niche enough not to sit in every default prompt index. |
||
|
|
fb4ed7b284 |
docs(skills): dynamic-workflow v2 — background-first delegation, campaign lessons from #102117
Main moved under take-1: top-level delegate_task is background-forced (results
re-enter as messages; synthesizing on the same turn reads files that do not
exist yet), per-task `toolsets` is gone (children inherit the parent's set),
DELEGATE_BLOCKED_TOOLS is {delegate_task, clarify, memory, send_message,
cronjob_manage} (execute_code is NOT stripped), and the 457-char description
was truncated by the 60-char routing budget. All four bbopen review items fixed.
Adds the campaign shape learned running the 1,863-session / 19-hour
whole-codebase simplification fan-out (#102117): shared brief + exclusive
file ownership, commit-per-step as the only handoff, fleet ceiling before the
OAuth refresh stampede, per-round integration with a frozen base and a full
suite on the combined tree, `rev-list --count` per branch before declaring a
round done (168 late-slice commits were once left behind), one serialized
test runner, forward-port last, live QA as its own wave, refuting the parent's
own heuristics with the same attempt/refuter mechanic, HANDOFF.md on restart.
Frontmatter now meets the hardline standard (platforms, ≤60-char description,
modern section order); `/tmp` replaced by the terminal temp dir so the skill is
correct on Termux and Windows. Catalog + sidebar + generated page added.
|
||
|
|
284d220ba4 |
fix(multiplex): cron, kanban, /loop and completion paths for a served profile match its standalone gateway
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:
- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
timeout were bare os.getenv → the default profile's values; a job without a
model silently ran on the default's HERMES_MODEL instead of refusing.
cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
kanban worker inherited the launch profile's non-credential .env settings
and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
HERMES_MODEL) — X's worker ran in the default's docker image on the
default's model. tools/environments/local.py::strip_launch_profile_env drops
them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
assignee (toolset probes call get_secret without a scope → swallowed
UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
writing the completed tick into the DEFAULT profile's state.db and leaving
X's row awaiting_response forever; the --until judge ran with the default's
aux credentials.
- background processes: a secondary's processes.json (scope-relative since
|
||
|
|
9848e22ed6 |
feat(multiplex)!: drop gateway.multiplex_profile_allowlist — serve every profile
The multiplexing default gateway now serves default + every live named profile under profiles/. profiles_to_serve(multiplex=True) is a pure directory read (tombstoned profiles skipped, never mkdir); every reader — gateway served set, /p/<profile>/ prefixes for api_server + webhook, the named-profile standalone guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a running gateway is unchanged) — drops the allowlist parameter. Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG, GatewayConfig and the top-level yaml bridge no longer carry it. BREAKING: anyone who set an allowlist now has their excluded profiles served. Archive or delete a profile you do not want served (Teknium approved). |
||
|
|
9c9e7ab6e5 |
fix(multiplex): a served profile's turn sees its own cwd, approvals, redaction and tool policy
Under gateway.multiplex_profiles a secondary profile's turn ran with the LAUNCH profile's working directory, command allowlist, redact_secrets switch, credential file mounts, browser engine/headed flags, LSP service, auxiliary-provider health marks and MCP stderr log, and several TERMINAL_ENV consumers read the process env instead of the routed profile's terminal scope. A standalone `hermes -p X gateway run` never behaved that way. - tools/terminal_scope.py: resolve the terminal.cwd placeholder inside the profile scope with the same rule gateway/run.py applies at import (local -> $HOME, sandbox default otherwise) so the system prompt, context files and the terminal of a routed turn start where the profile's standalone gateway would. - tools/image_source.py, credential_files.py, image_generation_tool.py, skills_tool.py, delegate_tool_progress.py, agent/tool_executor.py: read TERMINAL_ENV / TERMINAL_CWD through the terminal scope. - tools/approval.py (+ approval_floors.py): one permanent allowlist per routed profile home; the unscoped module set stays for single-profile processes. - agent/redact.py: `_redact_enabled()` resolves security.redact_secrets for the routed profile (scope .env, then config); launch snapshot kept when unscoped. - tools/credential_files.py, agent/auxiliary_health.py, agent/lsp/__init__.py, tools/browser_tool_cloud.py, tools/mcp_tool_config.py, tools/tool_result_storage.py: key process caches by profile home (or bypass the slot under an override). Tests: tests/tools/test_multiplex_turn_parity.py (4, red on base). Docs: multi-profile-gateways.md isolation table. |
||
|
|
446f6f79a8 |
fix(multiplex): dashboard and gateway stop see a served profile's gateway as the multiplexer
A profile served by the default multiplexer owns no gateway.pid / gateway_state.json, so every surface that reads per-profile identity files called it stopped while the CLI status surfaces (hermes -p X status / gateway status / cron status) said "running via the default-profile multiplexer": - `/api/status?profile=X` and `/api/messaging/platforms?profile=X` reported gateway_running=false / state=None / "gateway_stopped" in the same body that listed X under gateways[].served_profiles. The shared ladder `resolve_gateway_liveness` gains a fourth rung for a named profile_dir: the live default multiplexer that records X in served_profiles IS X's gateway (pid = multiplexer pid, runtime = its record, X's `<X>:<platform>` entries re-keyed to the standalone shape). - `POST /api/gateway/stop?profile=X` spawned `hermes -p X gateway stop`, which printed "No gateway running for this profile" (exit 0) into the action log while the UI flipped to stopped and the multiplexer kept serving X; `/api/gateway/restart?profile=X` spawned a `-p X gateway restart` that only exits 78. start/stop now answer 409 with the multiplexer explanation (one helper shared with the existing start refusal) and restart targets the multiplexer, the process that actually serves X. A `--force`-started separate gateway for X (own pid file) keeps normal per-profile management. - CLI `hermes -p X gateway stop` refuses with exit 78 like run/start/install/restart when X has no gateway of its own, instead of a contradictory exit-0 "not running". Docs: multi-profile-gateways.md §1 and §5 describe stop + the dashboard behaviour. |
||
|
|
76a0c9d4c4 | docs(multiplex): list adapter settings in the per-profile isolation table | ||
|
|
f923faa0b8 |
feat(voice): GPT-Live voice chat mode — a full-duplex voice frontend that delegates to Hermes (Desktop)
`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and has no tools of its own. Every real request it hears becomes a normal Hermes turn on the open chat — any model/provider the session selected, full toolset, memory, approvals — and the voice paraphrases the reply aloud. Backend - tools/voice_live.py: mode/credential/persona resolution and the one server-side step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer, pinned to client delegation; the OpenAI key never reaches the renderer. Voice persona follows the vendor prompting guide (role, style, labelled delegation policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn model-input note (transcript in, speakable prose out). - REST: GET /api/audio/voice-live/status (mode + readiness, non-secret), POST /api/audio/voice-live/session (SDP exchange). The offer is passed byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF". - prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the note rides the model input via the existing _prepend_note seam, the persisted user row stays the user's words, the system prompt stays byte-stable. - config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*. Desktop - lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript accumulation, session.commentary/thinking/instructions appends (500-token chunking), mute, graceful close waiting for session.closed. - hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation; delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking appends; reply streamed back per sentence; spoken stop phrase ends the chat; a newer delegation interrupts an in-flight turn. - use-composer-voice mounts both engines and latches one at conversation start from the backend-resolved status; gpt-live without a key falls back to chained with a notice. Settings → Voice gets the mode dropdown, voice picker, persona. Live-verified on the worktree desktop build (headless Electron, CDP, synthetic mic): "what is 17 times 23 and which model are you on" → delegation → Hermes (Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through OpenRouter"; follow-up "double that" resolved from the spoken context → 782; "run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as quiet context → spoken kernel version; "stop" closed the session (reason=close_requested). Chained mode creates no RTCPeerConnection. |
||
|
|
fafb27ee5c |
feat(plugins): on_room_member_activity hook projects Group Chat member runtime events to plugins
A hosted room member runs on a hidden room_plumbing session with no client transport, so the tool.start/complete, approval.request, message.delta and reasoning.delta frames its turn already emits bottom out at stdio and vanish. Between turn.started and turn.settled in the durable room log a client sees a black box, and community clients (Hermes Crew) cannot render tool cards, approvals or live member status without inferring them from text. One seam in write_json (plus the connector bypass in tool_progress) re-routes those frames, stamped with the session's _hosted_room_task coordinates (room_id, thread_id, member_id, turn_id, task_id, execution_generation), to a new observer hook through the bounded per-consumer queues on_stream_* already use, so plugin code never runs on the token path. Nothing is written to the room log: deltas would exhaust a room's byte budget in minutes and checkpoint replay must stay a pure function of the durable events. The task stamp gains member_id (the driver already knows it; the proof did not carry it). Group Chat keeps execution, scheduling and persistence; plugins own presentation. |
||
|
|
27481a6338 | docs(whatsapp): quoted replies attach the quoted media (both adapters) | ||
|
|
bf51fee548 |
docs(multiplex): make the multiplexed-gateway page match what the code does
The "one gateway for all profiles" section had drifted from the runtime. Each
claim was re-verified at its defining symbol on current main and rewritten to
the behaviour users will actually see:
- named-profile guard: only `gateway run` refuses (exit 78 / EX_CONFIG,
systemd RestartPreventExitStatus; launchd KeepAlive still retries);
`start`/`install` do not refuse themselves and the message lands in the
service log; `--force` is a `run`-only flag
(hermes_cli/gateway.py::_guard_named_profile_under_multiplexer, _cmd_start)
- same (platform, token) in two profiles: the duplicate adapter is parked as
fatal/duplicate_credential and the gateway keeps running — it was described
as a fail-fast startup error (gateway/run_adapters.py::_refuse_duplicate_claim)
- status surfaces: one gateway_state.json under the default home with
`<profile>:<platform>` entries + served_profiles; nothing is written under a
secondary home (the page claimed a per-profile runtime_status.json), and
`hermes status` does not list served profiles — `gateway list`,
`-p X gateway status` and /api/status do (gateway/status.py::write_runtime_status,
hermes_cli/status.py::_render_gateway, hermes_cli/gateway.py::_cmd_status)
- API_SERVER_KEY in a secondary .env auto-enables api_server and trips the
port-binding skip; document the `enabled: false` pin
(gateway/config_env.py::_api_server / _enable_from_env)
- allowlist: it is a start-time snapshot, and the Desktop backend's cron
ticker enumerates every local profile regardless of it
(hermes_cli/web_server.py::_start_desktop_cron_ticker)
- routed-profile cron via the shared bot: only when the profile has no live
adapter of its own, and a route carrying guild_id never matches a cron
target because delivery matches on chat_id/thread_id only
(cron/scheduler_provider.py::tick_adapters_for,
cron/scheduler_preflight.py::SharedRouteAdapters.get)
- add a "What is isolated per profile" table (credentials, authorization,
endpoints, media denylist, MCP child env, outbound egress, session
namespace, logs, terminal) describing behaviour, not PR numbers
- configuration.md: the ${VAR} scoping paragraph now says where it applies
and links to the table
|
||
|
|
819517fbac |
fix(gateway): routed profiles get their own max_turns, fallback chain, hooks, aux auth and media policy
One multiplexed gateway process serves every profile, but several per-turn reads still went through state frozen from the LAUNCH profile: - `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS, so every secondary ran with the default profile's turn budget. A routed turn (HERMES_HOME override) now resolves `agent.max_turns` from its own config. - `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide slot, so secondaries fell back through the default's provider/model with their own keys. It now reads the active gateway home and keeps a last-known-good chain per home. - `_load_prefill_messages` resolved relative paths against the launch home. - `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a secondary's compression/title/vision calls authenticated to Nous with the default profile's token when it had no pool entry. Resolved per call via `hermes_cli.auth._auth_file_path()` (patched constant still wins in tests). - `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was loaded outside any profile scope, so secondaries' `hooks/` never ran and the default profile's handlers received every profile's messages, responses and user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the runner holds one registry per served home, picked from the active scope at emit time and front-loaded under each secondary's startup scope. - Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME and the default profile's secrets). They now get the routed HERMES_HOME via `build_subprocess_env`, scrubbed under multiplexing, and the stdin payload carries `profile` so one script can tell which profile fired it. - Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`, `trust_recent_files*`) was bridged once into env at startup and read from env per delivery; under a HERMES_HOME override the validator now reads the routed profile's config. Single-profile runs keep the env-bridge contract. Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro (temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7, fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values. |
||
|
|
651878e1e4 | docs(multiplex): per-profile config.yaml behaviour, env_passthrough and write guards | ||
|
|
d807a34c87 | docs(multiplex): cron ticker allowlist, named multiplexer, guild-scoped route delivery | ||
|
|
27768a9fa5 |
fix(gateway): profile_routes no longer hijack DMs to a dedicated secondary bot (#104933)
Telegram DM chat_id == user_id for every bot, so a `chat_id` route meant to pin a user's DM with the SHARED bot to profile `ops` also captured that user's DM with team_b's dedicated bot: the turn, session key, state.db and secrets were ops's while the reply left via team_b's transport. `ProfileRoute` gains an optional `bot_profile` discriminator (None = the default profile's bot) and `matches()` requires it to equal the receiving adapter's owner profile. `build_source` passes `_owner_profile` and stamps it as the fallback profile for secondary adapters; `_profile_name_for_source` derives it from the transport ref for sources built elsewhere. Secondary→secondary routes remain possible by naming the bot explicitly. Salvages the discriminator idea from #105040 and #105049 in a smaller shape (no bot-username sniffing across adapter internals; the owning profile is already declared at `set_owner_profile`). Co-authored-by: 0xAlyDev <agentai891@gmail.com> Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com> |
||
|
|
37dcc0a6e8 |
fix(gateway): a secondary API_SERVER_KEY no longer skips the profile; start/install/status honour the live multiplexer
Under gateway.multiplex_profiles the default gateway serves every profile, yet four startup/status paths still reasoned from the wrong source: * A secondary profile's API_SERVER_KEY (which the docs REQUIRE for /p/<profile>/ auth) auto-enabled api_server in that profile's config, so _load_secondary_profile_config raised SecondaryPortBindingConfigError and the whole profile was skipped. gateway/config_env.py::_enable_from_env now leaves `enabled` alone for port-binding platforms while a multiplexer loads a NON-default profile (home override + multiplex flag, the same signal gateway.config uses for scoped reads); the credential still lands in extra so the shared listener can authenticate the prefix. Default profile unchanged. * "Is this profile served?" was re-derived from the default config.yaml plus GATEWAY_MULTIPLEX_PROFILES as seen by the CLI process. `hermes -p coder ...` loads coder's .env, so an env-only opt-in on the default profile was invisible (guard never fired, status said stopped) and an allowlist edit flipped the answer before the restart. named_profile_served_by_running_multiplexer now reads the pid-verified default gateway_state.json served_profiles (written by _record_served_profiles) first and falls back to config derivation only when the key is absent. The record helpers live in hermes_cli/gateway_multiplex_served.py. * The served-profile guard ran only inside `gateway run`. `hermes -p X gateway start|install|restart` reached the service manager, whose unit then exited 78 forever (systemd parks it while the CLI prints "started"; launchd KeepAlive respawns every 30 s). The service verbs now run the same guard up front (exit 78, same message) and accept --force; the Desktop /api/gateway/start route returns 409 for a served profile instead of spawning a doomed child. * Status surfaces disagreed: `hermes -p coder status` said stopped, `hermes -p coder cron status` said "cron jobs will NOT fire" while `cron list` said fine, and the default `hermes status` never listed served profiles. Both now route through the probe / the recorded served set. The -p/--profile matcher in _scan_gateway_pids and gateway.status._command_line_belongs_to_profile compares the flag token for equality (`-p ops` no longer claims -- or lets `gateway stop` SIGTERM -- an `-p ops-2` gateway). Docs: multi-profile-gateways.md now describes the start/install refusal, the --force flags, the API_SERVER_KEY behaviour and the single default-home gateway_state.json (the per-profile runtime_status.json claim was wrong). Fixes #100397 Addresses #89726 #97360 #71344 (cherry picked from commit d002c1864a7b6a22c53758b16b7b0cc79aea2edf) |
||
|
|
ceaf622c6d |
fix(mcp): same-named MCP servers with different credentials connect per profile; owner /reload-mcp keeps adopters' tools
Under gateway.multiplex_profiles every connection ledger in tools/mcp_tool.py
(_servers, _server_scope_keys/_server_tool_scopes, connecting/error/cooldown
maps, the circuit breaker, lazy schema-cache configs, trust metadata) was keyed
by the bare server NAME. The common per-tenant layout — each profile names its
server `github`/`notion` with its own token — gave only the first profile a
connection: the second profile's register_mcp_servers saw the name as "already
connected", refused to adopt it (different credentials,
|
||
|
|
a9838c2100 |
fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.
Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.
Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.
Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.
Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.
Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.
Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.
session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.
Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.
Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
|
||
|
|
956793aeac | docs(multi-profile): profile sessions keep their own keys and store on every gateway path | ||
|
|
759024bdff |
fix(deepseek): honor Flash 1M window leftovers and native vision
Users on native DeepSeek were told to pin model.context_length and model.supports_vision in config.yaml. That is the wrong layer: the 1M window is already in DEFAULT_CONTEXT_LENGTHS, and a global supports_vision pin would also mark text-only deepseek-v4-pro as multimodal. Two catalog gaps still produced the reported symptoms: - A leftover context_length_cache.yaml entry of 128K (the old ``deepseek`` catch-all) outlived the 1M catalog keys because deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS. - When models.dev is empty/cold, Flash has no capability record, so image routing falls through to lossy text. Vendor docs (2026-09-10) mark deepseek-flash as vision-capable and deepseek-v4-pro as not. Discard those 128K leftovers, fill Flash (and retired Flash aliases) via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only. |
||
|
|
0dcadf6f41 |
revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces. Later non-Wisdom work on shared files (guest onboarding i18n, dashboard startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call sites and config were stripped from those files. |
||
|
|
cbd03e6e4c |
fix(gateway): secondary-profile adapters no longer inherit the default's allow-all / allowlists
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env. Several
adapter-owned authorization gates still read GATEWAY_ALLOW_ALL_USERS, GATEWAY_ALLOWED_USERS
or their platform allowlist/allow-all raw from os.environ, so the default profile opting
into open access opened every secondary email/QQ/WhatsApp/Matrix/Teams/Slack/LINE/DingTalk
bot to any sender (email additionally skipped From: authentication), the default's Matrix
allowlist decided who may approve tool calls on a secondary bot, and a secondary that
opted in only in its own .env was silently deny-all.
Every such read now goes through the adapter's existing module-local scoped reader
(gateway.platforms._shared.get_scoped_secret / matrix _startup_env_secret): profile
scope first, scoped miss = default, never os.environ; the unscoped default-profile and
single-profile paths keep the environ read, where it IS the profile's own value.
Sites: email _allow_all_senders/_allowlist_in_effect; qqbot _open_dm_opted_in;
whatsapp_common _open_dm_opted_in/_live_dm_allow_from; teams _card_action_denied;
matrix _is_authorized_user, MATRIX_ALLOWED_USERS, MATRIX_IGNORE_USER_PATTERNS,
_extra_csv_set (allowed/free-response rooms); slack _slack_allow_bots/_slack_api_human_users;
line _truthy_env/allowlist (allow-all, user/group/room allowlists); dingtalk _extra_get
(allowed_users/chats, free-response chats, require_mention).
Live repro (temp HERMES_HOME, multiplex on, default env GATEWAY_ALLOW_ALL_USERS=true,
secondary scope without opt-in): EmailAdapter._allow_all_senders() True -> False,
QQAdapter._open_dm_opted_in() True -> False, Matrix _is_authorized_user('@stranger')
True -> False, Teams card action allowed -> denied.
Co-authored-by: Drexuxux <drexux0@gmail.com>
Co-authored-by: MoonsvnLyn <FirmamentalSpring@users.noreply.github.com>
Co-authored-by: svector-anu <anuoluwakolapo94@gmail.com>
Co-authored-by: babatorik <durgun.ismail@gmail.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
|
||
|
|
94fb74fa97 |
docs(bot-mode): explain Warm Bot Backends, idle reaping, and slot waits
Fleet users read "Hermes backend for profile X exited (1)" as a crash and raise Warm Bot Backends past their profile count to make it stop. Name the setting, the defaults, what the exit line means, and which operations take a slot so the knob is tuned for active bots instead of total profiles. |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
8d79c2ff57 |
Merge pull request #107959 from NousResearch/fix/local-fit-mtp-accounting
fix(local-runtime): unify memory accounting and effective-window MTP |
||
|
|
0316d3d404 |
fix(local-runtime): unify memory accounting and effective-window MTP
Price weights, context, runtime, projector and batch overhead consistently across catalog admission, initial launch, growth and restored windows. Keep MTP and the larger window when lean batches avoid unnecessary spill. Admit optional external drafts only when their complete footprint fits. Use preset-only model discovery so refused files cannot autoload, and preserve refusal/spill decisions atomically for desktop status read-back. Add regression coverage for complete-footprint boundaries, MTP restarts, growth admission, draft budgets and placement status transitions. Builds on the overhead-accounting contribution in #102993 and the restored-window MTP contribution in #106897. Does not adopt the 40% host-RAM reserve or resolve the remaining requests in #102865/#106895. Co-authored-by: infinitycrew39 <infinitycrew39@gmail.com> Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com> |
||
|
|
05d705dd69 |
Merge pull request #105839 from victor-kyriazakos/fix/relay-explicit-disable
fix(relay): honor explicit profile disable before relay startup |
||
|
|
45a6101f36 |
fix(gateway): secondary-profile send_message, notices and /loop wakeups go out via their own bot
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved its live adapter by bare platform from runner.adapters — the DEFAULT profile's map — so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy, every plugin platform), its "Gateway shutting down/restarted" and /update notices, its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left through the default bot (Telegram DMs landed in the user's chat with the other bot). Every such door now resolves through the profile-aware, fail-closed resolver already used by the inbound reply path (authz_mixin: _adapters_for_profile / _authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear error, never the default bot. - tools/send_message_senders.py::_live_adapter — resolve via runner._authorization_adapter(platform, get_active_profile_name()); shared by _send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and the WeCom standalone sender. - gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay- aware resolve_delivery_transport callers); _authorization_adapter reuses it. - gateway/run_shutdown.py — shutdown/restart notice for a running session uses the session's source transport / agent:<profile>: key lane, never self.adapters. - gateway/slash_commands.py + run_notifications.py — /restart and /update markers persist `profile`; the restart notice, update result and update prompt resolve the requester's own adapter (legacy markers fall back to the session_key lane). /goal, /heartbeat, /approve, /deny confirmations use the source's own transport. - gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its route; the wakeup watcher scans every served profile's store under its own scope (same shape as _handoff_watcher) and fires through that profile's adapter map. - plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in the owning profile (its adapters and its home channels). - plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter. - docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity). Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py, tests/gateway/test_multiplex_notice_egress_profile_adapter.py. Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com> |
||
|
|
c3e01c753d |
fix(gateway): /p/<profile>/ webhook delivery and api_server callbacks stay on the routed profile
Under gateway.multiplex_profiles a /p/<profile>/ webhook route (or one with
`profile: <name>`) executed under that profile but delivered its reply as the
FIRST profile owning the target platform: `_find_adapter` took
`runner.adapters` then iterated `_profile_adapters` in dict order, and the
home-channel fallback read `runner.config` (the default profile's). The gh leg
for `github_comment` inherited the process environ, i.e. the default profile's
GH_TOKEN. Both directions leaked: a secondary route posted through the default
bot, and a default route borrowed a platform parked only on a secondary.
The api_server `/p/<profile>/api/platforms/<platform>/events` callback had the
same shape — `_get_platform_callback_adapter` read `runner.adapters` regardless
of `_api_request_profile`, so a secondary's Google Chat / Teams events were
verified and dispatched by the default adapter.
Now:
- webhook: `_delivery_info` / the deliver_only dict carry the resolved
profile; `_find_adapter(platform, profile)` resolves through the runner's
shared fail-closed `_authorization_adapter`; the delivery leg runs inside
`_profile_scope(profile)` and takes the home channel from that profile's
`load_gateway_config()`; `gh` gets GH_TOKEN/GITHUB_TOKEN from the profile
secret scope (default's values dropped from the child env when the profile
has none).
- api_server: the callback adapter resolves via
`_authorization_adapter(platform, _api_request_profile.get())`; a named
profile without the adapter is a 503, never the primary's adapter.
A profile without the target platform fails closed ("not connected" → 502 /
503) instead of a silent cross-profile send.
Fixes #65939
Fixes #84266
Co-authored-by: 604maestro <604maestro@protonmail.com>
Co-authored-by: mjshorty <mjshorty@users.noreply.github.com>
Co-authored-by: StellarisW <stellarisw@users.noreply.github.com>
|
||
|
|
338bf9ea9a |
fix(cli): banner/TUI/dashboard update checks go through the GitHub API, cached 24h
The Python passive check (`banner.check_for_updates`, used by the CLI
banner, `hermes --tui`, every `tui_gateway` spawn and the dashboard's
/api/hermes/update/check) also ran `git fetch origin main` on every cache
miss, and never cached an inconclusive result so a flaky line retried on
every start. Same GitHub complaint, same fix:
- remote tip via GET /repos/{slug}/commits/main (vnd.github.sha), local tip
via rev-parse, exact count + changelog via the compare API when they
differ. HTTPS `ls-remote` remains only as the fallback when the API is
unreachable or the origin isn't on GitHub.
- cache TTL 6h -> 24h, failures cached 1h; the cache is keyed on HEAD so
`hermes update` invalidates it immediately.
- the dashboard's "what's changed" list comes from the memoized compare
payload (`upstream_commits_behind`) instead of `git log HEAD..origin/main`,
which was stale without a fetch.
Tests rewritten to the new contract: passive checks must not run
`git fetch`/`ls-remote` for a GitHub origin; the daily cache invalidates
when HEAD moves and re-asks after the failure window.
|
||
|
|
580322ef1e |
fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for every name tagged in the process-global `_SECRET_SOURCES` map. That map is filled by EVERY served profile's secret-source hydration, while `os.environ` only ever holds the LAUNCH (default) profile's values — so once any profile's 1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio MCP server was started with the default profile's token. Resolve those names through the active profile's secret scope (`get_secret`) instead: the routed profile's value, or omitted when that profile has none. Under multiplex `get_secret` never falls through to environ; single-profile runs keep the .env overlay + environ behaviour, so the existing "vault vars reach MCP subprocesses" contract still holds there. `secret_source_names()` exposes the tagged NAMES only — values are never read from the shared map. Docs: the multi-profile guide's "MCP subprocesses only see their own profile's secrets" claim is now true for source-injected names too; say so explicitly. |
||
|
|
3b044261b6 |
fix(gateway): media denylist covers every profile's credentials, not just the launch home
Under `gateway.multiplex_profiles` one process serves every `<root>/profiles/*`,
but `_media_delivery_denied_paths` expanded `_ROOT_CREDENTIAL_PATHS` only under
the import-time `_HERMES_HOME` / `_HERMES_ROOT`. A `MEDIA:<root>/profiles/<B>/.env`
(or auth.json, state.db, config.yaml, sessions/, mcp-tokens/) emitted in ANY
profile's turn — including B's own, whose HERMES_HOME override was never
consulted — passed validation and was natively uploaded to the chat. The ALLOW
side (`_profile_cache_roots`) already enumerated profiles at check time; the
DENY side did not.
`_credential_home_roots()` now yields the active `get_hermes_home()`, the shared
root and every `<root>/profiles/*` at check time (shared `_profile_dirs()` with
the allow side), and the denylist is built from that. Profile cache artifacts
and plain agent-written files under a profile stay deliverable.
Live repro (/tmp/mux_audit/fix-media-denylist/repro.py): before, all five of
profiles/B/{.env,auth.json,state.db,config.yaml,sessions/s1.json} validated as
deliverable while <root>/.env was blocked; after, all None, cache/images/gen.png
and report.pdf still deliverable.
No prior report. Write-side analogue: #107327 / #107335 (memo keying, different
mechanism — left as is).
|
||
|
|
6f5dda1aee |
fix(relay): read the opt-out through the loader's own YAML reader
relay_explicitly_disabled() went through hermes_cli.config.read_raw_config,
which turns a malformed config.yaml into {} and then applies the managed
overlay. load_gateway_config() raises on the same file and falls back to
env + gateway.json WITHOUT the managed layer. With a managed
relay.enabled: false and an injected URL the two readers disagreed: the
loader kept relay enabled and suppressed native Slack, while startup
refused to register relay — a profile with no route at all.
Extract the loader's file selection into config_loader.read_yaml_layers()
and have both consumers use it; the predicate mirrors the loader's
fallback (no YAML layer) on a parse error. Regression test fails on the
previous predicate.
|
||
|
|
7ea67e97f4 |
docs(relay): describe the opt-out at its real boundary
Drop the claims about reconnects and media handling that no longer have code behind them; state where the verdict is read and what a gateway.json disable does. |
||
|
|
4bdd64b334 |
The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cbcf7b72f7 |
feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop * feat(gateway): /signin signs the free tier into a Nous account from a DM * feat(cli): chat surfaces name /signin as the sign-in verb * fix(auth): review follow-ups for the shared sign-in flow and /signin * fix(i18n): carry the /status free-tier line in every locale catalog * refactor(cli): the chat sign-in command is /login * fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules |
||
|
|
3b01b4ce0f |
feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status answers has_guest / enabled / carries_inference / notice_pending with zero network, and free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and billing.state carry free_tier (billing answers the free tier locally instead of a portal call that can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector transfer and returns its code and consent URL; the poller waits for the transfer before the token grant, persists the account, runs settle_after_upgrade, and the poll response gains reason, account_email and model. * feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding overlay opens on a ready screen when the free tier carries inference, else a one-time strip above the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model / Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens; Done settles billing, model options, providers and re-homes a session still on nous/welcome. The picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes. * fix(desktop): free_tier.status starts the free tier's background setup when no identity exists A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free tier was never set up on the desktop: no connectors, no notice strip. The first status read now starts the same one-attempt background setup; the call itself never waits. * fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in). The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free tier, instead of Nous Portal · Connected. * fix(desktop): Settings > Providers never files the free tier under Connected * fix(desktop): the intro's shape is keyed on the route, not on the identity free_tier.status reports available (an identity exists and the tier is on); whether inference runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready screen shows when that route is the free tier; the composer strip when the user's own provider carries inference. An own-key install used to get the ready screen. * docs(desktop): say what the free-tier chip is keyed on * fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds * fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the transfer wait, after the token grant, and once more under the session lock together with the save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an attempt generation that every continuation checks after each await, so a poll from a closed attempt cannot publish over the one on screen (and its backend session is cancelled). The ready screen comes down only after the backend recorded the acknowledgement. A composer still mounted takes over the notice claim when its owner unmounts. One thin test per hole. |
||
|
|
04a76c4109 |
fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs after the account is persisted: a config on the free tier's route moves to the account's inference host and the recommended default for the account's plan, through the same config write a plain Nous login uses; a config on the user's own model is left alone. The pick is the one GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default. * fix(auth): a sign-in completion with no eligible recommendation leaves no default model The static provider-wide default is not narrowed by the account's plan or org policy, so writing it as a fallback could persist a model the account may not use. When the recommendation cannot yield a model, the route still moves to the account's host but model.default is left unset; the CLI says so and points at `hermes model`. * docs(free-tier): say what happens when no recommendation is available after sign-in * fix(auth): sign-in completion moves the host and clears the default in one config write Two writes could fail between them and leave the account host paired with nous/welcome. _update_config_for_provider gains clear_default so the caller with no model to offer removes model.default in the same atomic write that sets the host. |