feat: local models — managed llama.cpp runtime with one-click desktop setup

Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
This commit is contained in:
emozilla
2026-09-01 16:01:53 -04:00
parent 375ce8eee5
commit 43e67d872f
117 changed files with 16471 additions and 126 deletions
+47
View File
@@ -1228,6 +1228,53 @@ def _resolve_named_custom_runtime(
# `provider: ollama` with a LAN/WireGuard `base_url` doesn't silently
# fall through to OpenRouter.
requested_norm = (requested_provider or "").strip().lower()
# Managed llama.cpp runtime: a llamacpp-flavored alias with no explicit
# base_url resolves to the supervised server (or a detected external
# one) before the generic custom fallthrough. Explicit base_url always
# wins — a user pointing at a specific server means that server.
if requested_norm in ("llamacpp", "llama.cpp", "llama-cpp") and not explicit_base_url:
try:
from hermes_cli.local_runtime.endpoint import resolve_llamacpp_endpoint
endpoint = resolve_llamacpp_endpoint()
except Exception: # noqa: BLE001 — resolution is best-effort
endpoint = None
if endpoint:
return {
"provider": "custom",
"api_mode": "chat_completions",
"base_url": endpoint["base_url"],
"api_key": (explicit_api_key or "").strip()
or endpoint["api_key"] or "no-key-required",
"source": "local-runtime",
"requested_provider": requested_provider,
}
# No server to serve this model. Say so and stop — falling through
# to the generic custom path sends the request to whatever provider
# picks it up (OpenRouter with a placeholder key), and the user's
# "local server is off" surfaces as that provider's baffling
# "401 Invalid API key". The switch's own state picks the message:
# the user who turned the server off gets pointed at the switch,
# anyone else at the setup pane.
try:
from hermes_cli.config import load_config as _load_cfg
_lr_enabled = bool((_load_cfg().get("local_runtime") or {}).get("enabled"))
except Exception: # noqa: BLE001
_lr_enabled = False
if _lr_enabled:
raise ValueError(
"The local model server isn't running. It may still be "
"starting — try again in a moment, or check Settings → "
"Providers → Local models."
)
raise ValueError(
"The local model server is turned off. Turn it back on in "
"Settings → Providers → Local models, or switch to another "
"model."
)
if requested_norm and requested_norm != "custom":
try:
from hermes_cli.auth import resolve_provider as _resolve_provider