Files
hermes-agent/website/docs/user-guide/local-models.md
T
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00

135 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
sidebar_position: 4
title: Local Models
description: Run models entirely on your own machine — no account, no API key, nothing leaves your computer.
---
# Local Models
Hermes can run open models entirely on your own machine. It downloads and
manages the inference engine (llama.cpp), picks the right build of each
model for your hardware, and handles memory so you never configure context
sizes, GPU layers, or quantization. You pick a model; Hermes does the rest.
Nothing leaves your computer: no account, no API key, and no network access
after a model is downloaded.
## Getting started
1. Open **Settings → Providers → Local Models** (or choose **Run models
locally** during onboarding).
2. Click **Install runtime**. Hermes downloads the official llama.cpp
build for your hardware (a few hundred MB), verifies it, and keeps it
updated.
3. Pick a model from the catalog and click **Download**.
4. Click **Use**. New chats now run on the local model.
That's the whole flow. The server starts and stops with Hermes, restarts
survive app restarts, and switching back to a cloud provider is one click
in the model picker.
## How Hermes picks what to download
Every model in the catalog is priced against **your machine** before you
download anything. Each row shows:
- **Memory fit** — green (*Fits your GPU*: runs entirely in GPU memory),
amber (*Uses system RAM*: works, but slower), or red (*Too big for this
machine*).
- **Context** — the window the model starts with and the maximum it can
grow to.
- The download size of the build selected for your hardware.
Models ship in several quality grades (quantizations). Hermes picks the
highest-quality build that runs fully on your GPU; machines with less
memory get a more compact build of the same model with the same
guarantees. Below 4-bit the quality loss is too severe, so Hermes never
offers builds smaller than that — a machine that can't run the 4-bit
build spilled to system RAM simply can't run that model.
Models that don't fit stay visible with the reason, so you always know
what a hardware upgrade would unlock.
## How memory management works
Local models live or die by memory placement, so Hermes manages it
end-to-end and exposes no knobs:
- **Models start at a context window that fully fits your GPU** and grow
toward their native maximum as your conversation needs more room. You
may see "Context window grown" in the status feed during long sessions
— that's the window expanding, not an error.
- **Every recommended model gets at least a 64K context window.** When a
model is larger than your GPU's memory, Hermes deliberately places the
overflow in system RAM in the order that hurts least (expert weights
first, never the attention cache), trading some speed to protect the
context guarantee.
- **Conversation compression only kicks in at the model's maximum
window** — growth always comes first.
- Idle models are unloaded after 15 minutes to free GPU memory; they
reload automatically on the next message.
## The status bar
Right-click the status bar and enable **System resources** to see live GPU
utilization, GPU memory, and RAM while local models run. The context meter
always reflects the window the model is actually running with.
## Finding more models
The catalog is a curated starting point, not a boundary. The **Find more
models** section on the same page searches all of Hugging Face:
- Results show download counts and a per-file fit check sized to your
machine, so you know before downloading whether a build runs fully on
your GPU.
- Anything you download behaves exactly like a catalog model — Hermes
reads the model file itself to pick its context window and memory
placement. The only difference: community models don't carry our
"validated" testing badge.
- Already have a `.gguf` file on disk? **Add model file** links it into
your library without copying it (the original stays where it is), and
it's usable immediately.
## Using your own llama-server
If a llama-server is already running on your machine, Hermes detects it
and uses it instead of starting its own. Point a custom endpoint at any
OpenAI-compatible server for full manual control — the managed runtime is
a default, not a requirement. For manual setups (Ollama, MLX, custom
builds, headless CLI machines), see
[Run Hermes Locally with Ollama](/guides/local-ollama-setup) and
[Run Local LLMs on Mac](/guides/local-llm-on-mac).
## Configuration
The managed runtime is controlled by the `local_runtime` section of
`config.yaml`. The desktop UI writes these values for you; they're
documented for CLI and headless use:
```yaml
local_runtime:
enabled: false # true = start the managed server with Hermes.
# The desktop "Use" button sets this automatically.
backend: auto # auto | cuda | metal | vulkan | hip | cpu
tag: b10362 # pinned llama.cpp release; Hermes updates it with
# each release after re-validation
```
Models and runtime builds live under the Hermes home directory
(`models/` and `runtimes/llamacpp/`). Selecting a local model as your
main model uses the standard `model.provider: llamacpp` +
`model.default` settings — the same shape as every other provider.
## Requirements and limits
- **Windows and Linux:** NVIDIA GPU (CUDA) or CPU. **macOS:** Apple
Silicon (Metal). Vulkan builds serve AMD GPUs.
- A GPU with 8 GB+ of memory runs the small catalog models comfortably;
16 GB+ runs the 27–35B models at high quality.
- Model downloads are byte-size checked against the catalog during the
transfer; an incomplete download is deleted and reported, never
half-used. (Only the runtime engine zips are SHA-256 verified.)
- Deleting a model removes every file it staged, including vision
adapters and speculative-decoding companions.