Commit Graph

5 Commits

Author SHA1 Message Date
emozilla 5764ae3168 fix(local-models): keep automatic recommendations GPU-resident
Stop recommending a system-RAM spill when no curated model fits resident.
Preserve explicit model selection and the existing resident quality/speed
ranking, including the separate unified-memory policy.

Require a recommendation for automatic quickstart, expose Browse when
none exists, and rename Configure to Let me choose. Keep policy copy and
reason keys consistent across the four translated local-model sections.

Cover automatic refusal and explicit spilled setup against one budget,
plus the Browse, Download and Use interactions in the desktop pane.
2026-09-11 12:43:21 -04:00
Teknium a46ddf0f32 test: keep quickstart preflight independent of host acceleration 2026-09-07 08:23:08 -07:00
kshitijk4poor 5a0b1ba766 test(local-models): assert _assign_default reaches the real assignment body
Follow-up to the salvaged fix from #102994. The quickstart leg tests
stubbed web_deps.late itself, so the resolution path that regressed
(late() -> web_server facade -> getattr) was never exercised; the PR's
proposed tests pinned the module string instead, which fails on any
future move of the symbol even when the call site is updated.

Stub only the leaf (web_server_config._apply_model_assignment_sync) so
the real late() runs, and add one invariant test: _assign_default must
call it with ("main", "llamacpp", model_id, "", "", ""). Red on base
(AttributeError), green with the fix; not tied to where the symbol lives.
Drop the docstring parenthetical that described test mechanics.
2026-09-05 00:29:42 +05:30
Leanolf aef409343d fix(local-models): point _assign_default's late-bound model assignment at web_server_config
The whole-codebase refactor moved _apply_model_assignment_sync from the
web_server facade into web_server_config, but _assign_default in
web_routers/local_models.py still resolved it through the default late()
module (web_server), which no longer carries the symbol. A local-model
quickstart therefore failed its final 'making it your default' step with
AttributeError: module 'hermes_cli.web_server' has no attribute
'_apply_model_assignment_sync'.

Pass the defining module explicitly to web_deps.late() so the call
resolves against hermes_cli.web_server_config, matching how the sibling
models.py router imports the same helper. Update the two quickstart test
stubs to accept the module argument the real late(name, module=...) takes.
2026-09-05 00:29:42 +05:30
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00