docs: surface existing answers users can't find (migration, prompt-size, tool-call parsing, Desktop label)
This commit is contained in:
@@ -392,3 +392,4 @@ That sequence gets you from "broken vibes" back to a known state fast.
|
||||
- **[AI Providers](../integrations/providers.md)** — Full provider list and setup details
|
||||
- **[Skills System](../user-guide/features/skills.md)** — Reusable workflows and knowledge
|
||||
- **[Tips & Best Practices](../guides/tips.md)** — Power user tips
|
||||
- **[Moving to another machine](/reference/faq#exporting-hermes-to-another-machine)** — `hermes backup` migrates your whole setup (or [a single profile](/reference/faq#moving-a-single-profile-to-another-machine)); no need to rebuild from scratch
|
||||
|
||||
@@ -82,6 +82,10 @@ updates:
|
||||
|
||||
`updates.pre_update_backup` is a single knob with three modes: `quick` (default — the lightweight state snapshot described above), `full` (the quick snapshot plus a complete `HERMES_HOME` zip; can add minutes on large homes), and `off` (no pre-update backup at all — `--no-backup` does the same for a single run). Legacy boolean values still work: `true` means `full`, `false` means `off`.
|
||||
|
||||
:::tip Moving to a new machine instead?
|
||||
Update backups protect an in-place update. If you're migrating your whole setup to different hardware, use `hermes backup` + `hermes import` instead — see [Exporting Hermes to another machine](/reference/faq#exporting-hermes-to-another-machine) and [`hermes backup` vs `hermes profile export`](/reference/faq#hermes-backup-vs-hermes-profile-export).
|
||||
:::
|
||||
|
||||
### Windows: another `hermes.exe` is running
|
||||
|
||||
On Windows, `hermes update` will refuse to run if it detects another `hermes.exe` process holding the venv's entry-point executable open — most commonly the Hermes Desktop app's spawned backend, an open `hermes` REPL in another terminal, or a running gateway:
|
||||
|
||||
@@ -284,6 +284,8 @@ Smaller models (3B, 7B) sometimes ignore tool-call instructions and produce plai
|
||||
- **Hermes has auto-repair** — it detects malformed tool calls and attempts to fix them automatically.
|
||||
- **Set up a fallback** — if the local model fails 3 times, Hermes falls back to a cloud provider.
|
||||
|
||||
If the model prints raw JSON like `{"name": "web_search", ...}` in its reply instead of actually running the tool, that's usually the *server*, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in [Tool calls appear as text instead of executing](/integrations/providers#tool-calls-appear-as-text-instead-of-executing) (llama.cpp needs `--jinja`, vLLM needs `--enable-auto-tool-choice --tool-call-parser hermes`, and so on).
|
||||
|
||||
### Context window errors
|
||||
|
||||
The default Ollama context (2048 tokens) is too small for agentic work. See [Step 6](#step-6-optimize-for-speed) to increase it.
|
||||
|
||||
@@ -152,7 +152,7 @@ Instead of running terminal commands one at a time, ask the agent to write a scr
|
||||
Use `/model` to switch models mid-session. Use a frontier model (Claude Sonnet/Opus, GPT-4o) for complex reasoning and architecture decisions. Switch to a faster model for simple tasks like formatting, renaming, or boilerplate generation. Keep in mind each switch resets the prompt cache (see above), so on long sessions it's often cheaper to start a fresh session on the other model than to bounce back and forth.
|
||||
|
||||
:::tip
|
||||
Run `/usage` periodically to see your token consumption. Run `/insights` for a broader view of usage patterns over the last 30 days.
|
||||
Run `/usage` periodically to see your token consumption. Run `/insights` for a broader view of usage patterns over the last 30 days. To see what your *fixed* per-message cost is before any conversation — system prompt, skills index, memory, tool schemas — run [`hermes prompt-size`](/reference/cli-commands#hermes-prompt-size) (works offline).
|
||||
:::
|
||||
|
||||
## Messaging Tips
|
||||
|
||||
@@ -502,6 +502,10 @@ You can verify the plist has the correct PATH:
|
||||
|
||||
**Solution:**
|
||||
```bash
|
||||
# See exactly what the fixed prompt costs — breakdown by block
|
||||
# (system prompt, skills index, memory, tool schemas). Runs offline.
|
||||
hermes prompt-size
|
||||
|
||||
# Compress the conversation to reduce tokens
|
||||
/compress
|
||||
|
||||
@@ -509,6 +513,8 @@ You can verify the plist has the correct PATH:
|
||||
/usage
|
||||
```
|
||||
|
||||
If the baseline looks high before you've typed anything, that's the fixed prompt budget — the system prompt plus tool schemas sent on every call. Run [`hermes prompt-size`](/reference/cli-commands#hermes-prompt-size) to measure it, then trim: disable toolsets you don't use (`hermes tools`) and uninstall or disable skills you don't need (`hermes skills`).
|
||||
|
||||
:::tip
|
||||
Use `/compress` regularly during long sessions. It summarizes the conversation history and reduces token usage significantly while preserving context.
|
||||
:::
|
||||
|
||||
Reference in New Issue
Block a user