From 5aa798fecca2dafde69dd7be1a8a17d286235a17 Mon Sep 17 00:00:00 2001 From: Teknium <127238744+teknium1@users.noreply.github.com> Date: Wed, 5 Aug 2026 13:54:10 -0700 Subject: [PATCH] feat(skills): actual-setup optional skill + provider docs - optional-skills/devops/actual-setup: field-tested setup skill contributed by shl0ms, updated for the first-class 'actual' provider (the original targeted a custom-provider config that now collides with the built-in name) - docs: providers.md section + tables, environment-variables.md, quickstart.md - tests/skills: frontmatter + first-class-provider conformance checks --- optional-skills/devops/actual-setup/SKILL.md | 146 ++++++++++++++++++ .../actual-setup/references/opencode.md | 91 +++++++++++ tests/skills/test_actual_setup_skill.py | 68 ++++++++ website/docs/getting-started/quickstart.md | 1 + website/docs/integrations/providers.md | 33 +++- .../docs/reference/environment-variables.md | 2 + 6 files changed, 340 insertions(+), 1 deletion(-) create mode 100644 optional-skills/devops/actual-setup/SKILL.md create mode 100644 optional-skills/devops/actual-setup/references/opencode.md create mode 100644 tests/skills/test_actual_setup_skill.py diff --git a/optional-skills/devops/actual-setup/SKILL.md b/optional-skills/devops/actual-setup/SKILL.md new file mode 100644 index 0000000000..ae0fb98b29 --- /dev/null +++ b/optional-skills/devops/actual-setup/SKILL.md @@ -0,0 +1,146 @@ +--- +name: actual-setup +description: Set up Actual Computer (actual.inc) inference in Hermes. +version: 2.0.0 +author: shl0ms + Hermes Agent +license: MIT +metadata: + hermes: + tags: [actual, actual-inc, provider, local-inference, relay, gguf, setup] + category: devops +--- + +# Actual Computer Setup Skill + +Sets up [actual.inc](https://actual.inc) (Actual Computer) as a Hermes inference +provider. Actual turns the user's own hardware into a private inference cluster +and exposes an OpenAI-compatible API two ways: a hosted end-to-end-encrypted +relay at `https://api.actual.inc` (authenticated with an `ac_` key), and a local +on-device daemon at `http://127.0.0.1:8080` (no auth on loopback). This skill +does not install the Actual daemon for the user — device authorization requires +a human in a browser. + +## When to Use + +- User wants to add actual.inc as an inference provider (cloud relay or local). +- User has an `ac_` key and wants Hermes routed through their Actual cluster. +- User wants fully-local, on-device inference via the Actual daemon. +- Troubleshooting: Actual requests failing with cryptic 400s or empty streams. + +## Prerequisites + +- Hermes has **first-class `actual` provider support** (provider id `actual`, + aliases `actual-computer`, `actualcomputer`, `aci`). Do NOT configure Actual + as a `custom_providers` / `providers.actual.*` entry on current Hermes — the + built-in provider owns the name and handles base-url normalization, the + Responses transport, and local no-auth automatically. +- Relay mode: an Actual account and an `ac_` inference key from + https://actual.inc/user/keys. +- Local mode: the user has installed the daemon + (`curl -fsSL "https://actual.inc/install" | bash`) and completed device + authorization by running `actual` once and opening the printed + `https://actual.inc/device?code=...` URL in a browser. Relay that URL to the + user and WAIT — never invent an email or authorize on their behalf. Codes + expire in 5 minutes; re-run `actual` for a fresh one. + +## How to Run + +### Relay / API mode + +1. Put the key in `.env` (secrets only — never config.yaml): + append `ACTUAL_API_KEY=ac_...` to `~/.hermes/.env`. +2. Verify the key and discover models with `terminal`: + ```bash + curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY" + ``` +3. Select provider + model: + ```bash + hermes config set model.provider actual + hermes config set model.default "MODEL_ID_FROM_DISCOVERY" + ``` +4. Verify end-to-end: + ```bash + hermes chat -Q -q "Reply with exactly: ACTUAL_OK" --provider actual -m MODEL_ID + ``` + +### Local mode + +1. Human has installed + authorized the daemon (see Prerequisites). +2. Download and load a model (scriptable once authorized): + ```bash + actual models search "qwen2.5 0.5b instruct gguf" --limit 8 --no-prompt + # Downloads REQUIRE an explicit quantization (409 ambiguous_model_download otherwise): + actual models download "Qwen/Qwen2.5-0.5B-Instruct-GGUF/Q4_K_M" + actual models list # note the INSTALLED name (differs from download id) + actual models load "qwen2.5-0.5b-instruct-q4_k_m" # load by installed name + ``` +3. Point Hermes at the daemon. `ACTUAL_BASE_URL` with a loopback host flips the + built-in provider into local no-auth mode automatically — no key needed: + append `ACTUAL_BASE_URL=http://127.0.0.1:8080` to `~/.hermes/.env`, then: + ```bash + hermes config set model.provider actual + hermes config set model.default "INSTALLED_MODEL_NAME" + ``` +4. Verify (reduced toolset — see context-window pitfall below): + ```bash + hermes chat -Q -q "Reply with exactly: LOCAL_OK" --provider actual -m INSTALLED_NAME -t file,web + ``` + +## Quick Reference + +| Thing | Value | +|---|---| +| Hosted relay | `https://api.actual.inc/v1` (normalized from bare host automatically) | +| Local daemon | `http://127.0.0.1:8080/v1` (no auth on loopback) | +| Key env var | `ACTUAL_API_KEY` (`ac_...`) | +| Base URL env var | `ACTUAL_BASE_URL` (loopback host ⇒ local no-auth mode) | +| Provider id / aliases | `actual` / `actual-computer`, `actualcomputer`, `aci` | +| Transport | Responses API (`codex_responses`) — built-in, do not override | +| Cluster pinning | `X-Cluster-ID` header via `providers.actual.extra_headers` in config.yaml | +| Model size guide | 0.5B Q4_K_M ~470MB (toy), 7-8B Q4_K_M ~4.5GB (daily driver), 32B ~20GB | + +## Pitfalls + +1. **reasoning_effort trap (handled by Hermes since the first-class provider).** + Actual's SGLang/vLLM backends accept only `none/low/medium/high/max`; + `xhigh`/`ultra` used to fail with a cryptic + `Expecting value: line 1 column 1 (char 0)` (a wrapped HTTP 400). The + built-in provider clamps `xhigh→high` and `ultra→max` on the wire. If a + request still 400s this way on an old Hermes, set a per-model cap: + `agent.reasoning_overrides.: high` in config.yaml. +2. **Context-window overflow on small local models.** Hermes' default toolset + is ~26k tokens of schemas plus a ~9k-token system prompt. A model loaded + with a 32k context overflows before the first turn, and llama.cpp-family + servers emit a bare `data: [DONE]` — Hermes reports + `Provider returned an empty stream with no finish_reason`. This is NOT an + SSE bug. Fixes: restrict tools (`-t file,web`), load the model with a + larger `n_ctx`, or pick a >=64k-context model for the full toolset. + Upstream tracking: #51448 (do not file new issues; add evidence there). + Related but distinct: #65631 (HTTP-200 SSE carrying a 400), #56516 + (reasoning-only streams). +3. **Download ids vs installed names.** `actual models download` takes + `repo/QUANT` and 409s without an explicit quantization; + `actual models load` takes the INSTALLED name from `actual models list`. +4. **Reasoning models returning empty content.** GLM/Qwen reasoning variants + emit thinking in a separate `reasoning` field and can burn a small + `max_tokens` entirely on reasoning. Give generous max_tokens before + assuming failure. +5. **Do not create a custom provider named `actual`.** Older setup guides + (pre first-class support) wrote `providers.actual.*` config blocks. On + current Hermes the built-in provider wins the name; stale custom blocks + are ignored or conflict. Remove them and use the env vars + model.provider + flow above. + +## Verification + +```bash +# Relay: +hermes chat -Q -q "Reply with exactly: ACTUAL_OK" --provider actual -m MODEL +# Local (small model — reduced toolset): +hermes chat -Q -q "Reply with exactly: LOCAL_OK" --provider actual -m MODEL -t file,web +# Provider status (local no-auth shows key_source=local-offline): +hermes status +``` + +For other OpenAI-compatible clients (e.g. OpenCode), see +`references/opencode.md`. diff --git a/optional-skills/devops/actual-setup/references/opencode.md b/optional-skills/devops/actual-setup/references/opencode.md new file mode 100644 index 0000000000..88d29d9bb6 --- /dev/null +++ b/optional-skills/devops/actual-setup/references/opencode.md @@ -0,0 +1,91 @@ +# actual.inc as an OpenCode provider + +Verified end-to-end 2026-07 (OpenCode 1.18.3, macOS). Adds Actual's relay/GLM +cluster to OpenCode as a custom OpenAI-compatible provider. + +## Design: secret in auth.json, config in opencode.json + +OpenCode auto-injects a credential when the provider **id** in `opencode.json` +matches a credential **id** in `~/.local/share/opencode/auth.json`. So put the +key in auth.json and NOTHING sensitive goes in opencode.json. This is more robust +than `options.apiKey: "{env:ACTUAL_API_KEY}"`, because `{env:...}` only resolves +if the var is exported in the shell OpenCode launches from — and the Actual key +is typically only in `~/.hermes/.env`, not a shell profile, so the env form +breaks outside an inheriting terminal. + +### 1. Add the credential to auth.json + +File: `~/.local/share/opencode/auth.json`. Shape (preserve existing entries): +```json +{ + "anthropic": { "type": "api", "key": "..." }, + "actual": { "type": "api", "key": "ac_..." } +} +``` +Do this with a read-modify-write (json load, add the `actual` key, dump) so the +other credentials stay intact — don't overwrite the file. + +### 2. Add the provider to opencode.json + +File: `~/.config/opencode/opencode.json` (or `~/.opencode.json`). Add under +`provider` alongside anything already there. NO `apiKey` field — it comes from +auth.json by id match. +```json +{ + "$schema": "https://opencode.ai/config.json", + "provider": { + "actual": { + "npm": "@ai-sdk/openai-compatible", + "name": "Actual (GLM cluster)", + "options": { + "baseURL": "https://api.actual.inc/v1", + "headers": { + "X-Cluster-ID": "" + } + }, + "models": { + "glm-5.2-nvfp4": { + "name": "GLM-5.2 (b300x8)", + "limit": { "context": 1048576, "output": 65536 } + } + } + } + } +} +``` +- `npm`: `@ai-sdk/openai-compatible` for `/v1/chat/completions`. Use + `@ai-sdk/openai` only if the model needs `/v1/responses`. +- `options.headers.X-Cluster-ID`: pin to a specific cluster (optional; omit to + let the relay route). Get the hash from the Actual console URL + (`console/computers?cluster=`). +- `models.`: the id MUST match what `GET /v1/models` returns. Discover it + first: `curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer ac_..." -H "X-Cluster-ID: "`. +- `limit`: lets OpenCode track remaining context (custom providers don't get + this from models.dev). GLM-5.2 context = 1_048_576. + +### 3. Verify live (headless) + +```bash +opencode run -m actual/glm-5.2-nvfp4 "Reply with exactly this text: OPENCODE_ACTUAL_OK" +``` +OpenCode DOES use the `provider/model` slash form on the CLI (unlike Hermes, +where the slash form 404s custom providers). Expect the exact reply. Run a second +reasoning check (e.g. "What is 17 * 23?") since GLM-5.2 is a reasoning model. + +## Why no reasoning_effort trap here + +The Actual relay rejects `reasoning_effort: xhigh` with an HTTP 400 (see the +`hermes-custom-providers` skill, pitfall 2). Hermes hits this because it forwards +its global `agent.reasoning_effort`. OpenCode's ai-sdk does NOT send that param, +so Actual + OpenCode works with zero reasoning config. No `reasoning_overrides` +equivalent needed. + +## Gotchas + +- `auth.json` is the same store `/connect` writes; editing it directly is fine + and equivalent. `opencode auth list` should then show `actual` under + Credentials. +- If discovery/models don't appear, confirm the provider id in opencode.json + EXACTLY matches the auth.json credential id (`actual` == `actual`). +- No git-tracking risk on the default config dir (`~/.config/opencode` is not a + repo), but still keep the key in auth.json, not opencode.json, as the habit. diff --git a/tests/skills/test_actual_setup_skill.py b/tests/skills/test_actual_setup_skill.py new file mode 100644 index 0000000000..f2a110e418 --- /dev/null +++ b/tests/skills/test_actual_setup_skill.py @@ -0,0 +1,68 @@ +""" +Smoke tests for the actual-setup optional skill. + +Validates: + - SKILL.md frontmatter conforms to the ≤60-char description standard + - Frontmatter has required fields + - The skill references the first-class ``actual`` provider (not the + legacy custom-provider config path that conflicts with it) + - The bundled OpenCode reference exists +""" +from __future__ import annotations + +import re +from pathlib import Path + +import pytest +import yaml + +SKILL_DIR = ( + Path(__file__).resolve().parents[2] + / "optional-skills" + / "devops" + / "actual-setup" +) + + +@pytest.fixture(scope="module") +def skill_source() -> str: + return (SKILL_DIR / "SKILL.md").read_text(encoding="utf-8") + + +@pytest.fixture(scope="module") +def frontmatter(skill_source) -> dict: + m = re.search(r"^---\n(.*?)\n---", skill_source, re.DOTALL) + assert m, "SKILL.md missing YAML frontmatter" + return yaml.safe_load(m.group(1)) + + +def test_skill_dir_exists() -> None: + assert SKILL_DIR.is_dir(), f"missing skill dir: {SKILL_DIR}" + + +def test_description_under_60_chars(frontmatter) -> None: + desc = frontmatter["description"] + assert len(desc) <= 60, f"description is {len(desc)} chars (limit ≤60): {desc!r}" + + +def test_has_required_frontmatter_fields(frontmatter) -> None: + for field in ("name", "description", "version", "license"): + assert field in frontmatter, f"missing required field: {field}" + + +def test_credits_contributor(frontmatter) -> None: + assert "shl0ms" in str(frontmatter.get("author", "")), ( + "author must credit the human contributor first" + ) + + +def test_uses_first_class_provider_not_custom_provider(skill_source) -> None: + # The legacy setup configured Actual as providers.actual.* custom entries, + # which now collides with the built-in provider of the same name. + assert "--provider actual" in skill_source + assert "hermes config set providers.actual.api " not in skill_source + assert "key_env" not in skill_source + + +def test_opencode_reference_exists() -> None: + assert (SKILL_DIR / "references" / "opencode.md").is_file() diff --git a/website/docs/getting-started/quickstart.md b/website/docs/getting-started/quickstart.md index ecac3d64a3..390cdba2ac 100644 --- a/website/docs/getting-started/quickstart.md +++ b/website/docs/getting-started/quickstart.md @@ -119,6 +119,7 @@ Good defaults: | **Kimi / Moonshot China** | China-region Moonshot endpoint | Set `KIMI_CN_API_KEY` | | **Arcee AI** | Trinity models | Set `ARCEEAI_API_KEY` | | **GMI Cloud** | Multi-model direct API | Set `GMI_API_KEY` | +| **Actual Computer** | Your own hardware as a private inference cluster — hosted relay or local daemon | Set `ACTUAL_API_KEY` (relay) or `ACTUAL_BASE_URL=http://127.0.0.1:8080` (local, no key) | | **MiniMax (OAuth)** | MiniMax frontier model via browser OAuth — no API key needed (model name in `hermes_cli/models.py` may change between releases) | `hermes model` → MiniMax (OAuth) | | **MiniMax** | International MiniMax endpoint | Set `MINIMAX_API_KEY` | | **MiniMax China** | China-region MiniMax endpoint | Set `MINIMAX_CN_API_KEY` | diff --git a/website/docs/integrations/providers.md b/website/docs/integrations/providers.md index 57aeaa685b..a5b81f5660 100644 --- a/website/docs/integrations/providers.md +++ b/website/docs/integrations/providers.md @@ -28,6 +28,7 @@ You need at least one way to connect to an LLM. Use `hermes model` to switch pro | **Kimi / Moonshot (China)** | `KIMI_CN_API_KEY` in `~/.hermes/.env` (provider: `kimi-coding-cn`; aliases: `kimi-cn`, `moonshot-cn`) | | **Arcee AI** | `ARCEEAI_API_KEY` in `~/.hermes/.env` (provider: `arcee`; aliases: `arcee-ai`, `arceeai`) | | **GMI Cloud** | `GMI_API_KEY` in `~/.hermes/.env` (provider: `gmi`; aliases: `gmi-cloud`, `gmicloud`) | +| **Actual Computer** | `ACTUAL_API_KEY` in `~/.hermes/.env` for the hosted relay, or `ACTUAL_BASE_URL=http://127.0.0.1:8080` for the local daemon — no key needed on loopback (provider: `actual`; aliases: `actual-computer`, `actualcomputer`, `aci`) | | **MiniMax** | `MINIMAX_API_KEY` in `~/.hermes/.env` (provider: `minimax`) | | **MiniMax China** | `MINIMAX_CN_API_KEY` in `~/.hermes/.env` (provider: `minimax-cn`) | | **xAI (Grok) — Responses API** | `XAI_API_KEY` in `~/.hermes/.env` (provider: `xai`) | @@ -547,6 +548,35 @@ model: The base URL can be overridden with `GMI_BASE_URL` (default: `https://api.gmi-serving.com/v1`). +### Actual Computer + +Your own hardware as a private inference cluster via [Actual Computer](https://actual.inc). Two serving modes, both OpenAI-compatible (Hermes uses the Responses API transport): + +- **Hosted relay** — `https://api.actual.inc`, end-to-end encrypted, routes to *your* cluster. Authenticate with an `ac_` inference key from [actual.inc/user/keys](https://actual.inc/user/keys). +- **Local daemon** — on-device at `http://127.0.0.1:8080`, fully offline. No API key needed: Hermes detects the loopback base URL and authenticates with an internal placeholder automatically. + +```bash +# Hosted relay (ACTUAL_API_KEY in ~/.hermes/.env) +hermes chat --provider actual --model + +# Local daemon (ACTUAL_BASE_URL=http://127.0.0.1:8080 in ~/.hermes/.env, no key) +hermes chat --provider actual --model +``` + +Or set it permanently in `config.yaml`: +```yaml +model: + provider: "actual" + default: "" +``` + +Notes: +- Model IDs come from your cluster's `GET /v1/models` — discover with `hermes model` or `curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY"`. +- Bare hosts are normalized: `ACTUAL_BASE_URL=http://127.0.0.1:8080` becomes `http://127.0.0.1:8080/v1` automatically. +- Reasoning effort is clamped to Actual's supported range (`none/low/medium/high/max`) — a global `xhigh`/`ultra` setting will not 400 requests. +- Small local models: Hermes' full default toolset plus the system prompt can exceed a 32k context window, producing an empty-stream error from llama.cpp-family servers. Restrict the toolset (`-t file,web`) or load the model with a larger context. The optional `actual-setup` skill (`hermes skills install official/devops/actual-setup`) covers setup and troubleshooting in detail. +- Aliases: `actual-computer`, `actualcomputer`, `aci`. + ### StepFun Step-series models via [StepFun](https://platform.stepfun.com) — OpenAI-compatible API, API key authentication. @@ -1145,6 +1175,7 @@ Any service with an OpenAI-compatible API works. Some popular options: | [DeepSeek](https://deepseek.com) | `https://api.deepseek.com/v1` | DeepSeek models | | [Fireworks AI](https://fireworks.ai) | `https://api.fireworks.ai/inference/v1` | Fast open model hosting | | [GMI Cloud](https://www.gmicloud.ai/) | `https://api.gmi-serving.com/v1` | Managed OpenAI-compatible inference | +| [Actual Computer](https://actual.inc) | `https://api.actual.inc/v1` | Private relay to your own cluster; local daemon at `http://127.0.0.1:8080/v1` | | [Cerebras](https://cerebras.ai) | `https://api.cerebras.ai/v1` | Wafer-scale chip inference | | [Mistral AI](https://mistral.ai) | `https://api.mistral.ai/v1` | Mistral models | | [OpenAI](https://openai.com) | `https://api.openai.com/v1` | Direct OpenAI access | @@ -1524,7 +1555,7 @@ fallback_model: When activated, the fallback swaps the model and provider mid-session without losing your conversation. The chain is tried entry-by-entry; activation is one-shot per session. -Supported providers: `openrouter`, `nous`, `novita`, `openai-codex`, `copilot`, `copilot-acp`, `anthropic`, `gemini`, `qwen-oauth`, `huggingface`, `zai`, `kimi-coding`, `kimi-coding-cn`, `minimax`, `minimax-cn`, `minimax-oauth`, `deepseek`, `nvidia`, `xai`, `xai-oauth`, `ollama-cloud`, `bedrock`, `ai-gateway`, `azure-foundry`, `opencode-zen`, `opencode-go`, `kilocode`, `xiaomi`, `arcee`, `gmi`, `stepfun`, `lmstudio`, `alibaba`, `alibaba-coding-plan`, `tencent-tokenhub`, `custom`. +Supported providers: `openrouter`, `nous`, `novita`, `openai-codex`, `copilot`, `copilot-acp`, `anthropic`, `gemini`, `qwen-oauth`, `huggingface`, `zai`, `kimi-coding`, `kimi-coding-cn`, `minimax`, `minimax-cn`, `minimax-oauth`, `deepseek`, `nvidia`, `xai`, `xai-oauth`, `ollama-cloud`, `bedrock`, `ai-gateway`, `azure-foundry`, `opencode-zen`, `opencode-go`, `kilocode`, `xiaomi`, `arcee`, `gmi`, `actual`, `stepfun`, `lmstudio`, `alibaba`, `alibaba-coding-plan`, `tencent-tokenhub`, `custom`. :::tip Fallback is configured exclusively through `config.yaml` — or interactively via `hermes fallback`. For full details on when it triggers, how the chain advances, and how it interacts with auxiliary tasks and delegation, see [Fallback Providers](/user-guide/features/fallback-providers). diff --git a/website/docs/reference/environment-variables.md b/website/docs/reference/environment-variables.md index 39f6a057eb..bae4a6166c 100644 --- a/website/docs/reference/environment-variables.md +++ b/website/docs/reference/environment-variables.md @@ -45,6 +45,8 @@ Hermes reads environment variables from the process environment and, for user-ma | `ARCEE_BASE_URL` | Override Arcee base URL (default: `https://api.arcee.ai/api/v1`) | | `GMI_API_KEY` | GMI Cloud API key ([gmicloud.ai](https://www.gmicloud.ai/)) | | `GMI_BASE_URL` | Override GMI Cloud base URL (default: `https://api.gmi-serving.com/v1`) | +| `ACTUAL_API_KEY` | Actual Computer inference key (`ac_...`, [actual.inc/user/keys](https://actual.inc/user/keys)). Not needed for the local daemon. | +| `ACTUAL_BASE_URL` | Override Actual Computer base URL (default: `https://api.actual.inc/v1`). Set to `http://127.0.0.1:8080` for the local offline daemon — loopback hosts need no API key. | | `MINIMAX_API_KEY` | MiniMax API key — global endpoint ([minimax.io](https://www.minimax.io)). **Not used by `minimax-oauth`** (OAuth path uses browser login instead). | | `MINIMAX_BASE_URL` | Override MiniMax base URL (default: `https://api.minimax.io/anthropic` — Hermes uses MiniMax's Anthropic Messages-compatible endpoint). **Not used by `minimax-oauth`**. | | `MINIMAX_CN_API_KEY` | MiniMax API key — China endpoint ([minimaxi.com](https://www.minimaxi.com)). **Not used by `minimax-oauth`** (OAuth path uses browser login instead). |