Commit Graph

3 Commits

Author SHA1 Message Date
Teknium 4d1880e0bf fix(integration): restore subprocess encoding/stdin guards dropped in simplification
Simplification workers collapsed subprocess call sites into shared kwargs
helpers and dropped the Windows/TUI safety kwargs on the way:

- encoding='utf-8', errors='replace' restored on text=True runs in
  copilot_acp_client, hermes_cli/setup (vercel install), managed_uv
  (codesign steps), local_runtime/hardware._stdout, a2a adapter.
- stdin=subprocess.DEVNULL restored on copilot probe, verify/runner
  _SUBPROCESS_KW, iron_proxy._run, google_meet playwright/system_profiler,
  simplex convert, whatsapp _RUN_TEXT, mem0 ollama serve Popen.
  google_meet sudo/brew install keeps inherited stdin (user-confirmed,
  may prompt) — marked noqa: subprocess-stdin.
- Windows-safe SIGKILL: getattr(signal, 'SIGKILL', SIGTERM) in
  verify/runner; photon _kill call re-marked windows-footgun: ok
  (unreachable on win32).
- scripts/check_subprocess_stdin.py now recognizes **kwargs splats
  (**_KW / **_kw(...)) ONLY when the same-file definition provably sets
  stdin= — covers tui_gateway _capture_run_kwargs/run_kw. Parity test added.
2026-09-02 16:36:10 -07:00
Teknium 3ffd44acd3 refactor(hclib): models/runtime — models, inventory, runtime_provider, provider catalog, local_runtime, banner 2026-09-02 14:44:48 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00