43e67d872f
Run models locally as a first-class provider. The CLI grows a managed llama.cpp runtime (engine install, model download, server supervision); the desktop app grows the full setup and management story on top of it. GUI surfaces ship behind the desktop --local launch flag (hermes desktop --local, or the flag on the packaged app); backend routes and the CLI are always live. Runtime (hermes_cli/local_runtime/): - curated GGUF catalog with per-machine variant selection: hardware probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice by context window - derived recommendation: quality-ranked picks gated by a predicted decode-speed floor, bandwidth-aware on unified memory; the decision table is pinned as a test (pick AND reason per memory class), and the Recommended badge explains its pick in a tooltip fed by the resolver's actual branch - engine install + model download with resumable split parts, cumulative plan-level progress, and staged-model integrity (a split GGUF counts only when every part is present) - server supervision: spawn/adopt/stop, router mode with per-model load progress relayed over SSE, abandoned-request cleanup Desktop: - Settings -> Providers -> Local models: one-click quickstart (install engine, download the recommended model, boot) plus per-model download/ activate/eject, fit-ranked catalog with context pills - model pickers (composer dropdown + Cmd+K) show staged local models, in-flight downloads as live progress rows, and load-into-memory bars - local-setup campaign tip for eligible hardware; System resources statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends - friendly dead-server errors, and failed agent builds retry on the next send instead of wedging the session Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
56 lines
2.1 KiB
Python
56 lines
2.1 KiB
Python
"""Managed llama.cpp runtime.
|
|
|
|
Hermes downloads, verifies, supervises, and updates one llama-server, and
|
|
decides per machine which model build and context window to run. Key
|
|
modules:
|
|
|
|
- ``binaries`` — resolve/download/verify official llama.cpp release zips
|
|
into ``$HERMES_HOME/runtimes/llamacpp/<tag>/``.
|
|
- ``supervisor``— spawn and supervise one llama-server in router mode;
|
|
readiness is a touch generation, never health-200 alone.
|
|
- ``detect`` — find an already-running llama-server (external or ours).
|
|
- ``estimator`` / ``context_policy`` / ``growth`` — price context memory
|
|
per architecture and run the window ladder (zero-spill start, grow
|
|
toward native max, compress only at the top).
|
|
- ``catalog`` / ``presets`` — the curated model list and the per-model
|
|
launch flags that carry policy decisions to the router.
|
|
|
|
Everything is driven by the ``local_runtime`` section of config.yaml.
|
|
"""
|
|
|
|
from hermes_cli.local_runtime.binaries import ( # noqa: F401
|
|
BinaryResolutionError,
|
|
ensure_runtime_installed,
|
|
resolve_assets,
|
|
select_backend,
|
|
)
|
|
from hermes_cli.local_runtime.bootstrap import ( # noqa: F401
|
|
ensure_local_runtime,
|
|
shutdown_local_runtime,
|
|
)
|
|
from hermes_cli.local_runtime.context_policy import ( # noqa: F401
|
|
FLOOR,
|
|
growth_decision,
|
|
initial_window,
|
|
ladder,
|
|
launch_args,
|
|
)
|
|
from hermes_cli.local_runtime.growth import ( # noqa: F401
|
|
clear_window_override,
|
|
load_window_overrides,
|
|
maybe_grow_window,
|
|
save_window_override,
|
|
)
|
|
from hermes_cli.local_runtime.detect import detect_server # noqa: F401
|
|
from hermes_cli.local_runtime.endpoint import resolve_llamacpp_endpoint # noqa: F401
|
|
from hermes_cli.local_runtime.estimator import ( # noqa: F401
|
|
HardwareBudget,
|
|
ctx_bytes,
|
|
physics_check,
|
|
profile_from_gguf,
|
|
)
|
|
from hermes_cli.local_runtime.gguf import read_gguf_header # noqa: F401
|
|
from hermes_cli.local_runtime.hardware import probe_budget # noqa: F401
|
|
from hermes_cli.local_runtime.presets import generate_presets # noqa: F401
|
|
from hermes_cli.local_runtime.supervisor import LlamaServerSupervisor # noqa: F401
|