804707bea6
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended with "Fleet version check returned no rows" (exit 1); the restarted gateway took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The gateway constructor was running `maybe_auto_prune_checkpoints` synchronously, before the control socket, adapters and the code_sha stamp, and on a 1.2 GB store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because the size-cap shrink gc'd again even when it could drop nothing. The same defect sat on the tool-call path: `CheckpointManager._take` ran `_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without dropping anything and triggered a 20-28s gc on the first file-mutating tool call of every turn once the store was over the cap. That loop also re-measured a pack size that cannot move without a gc, so a single over-cap checkpoint dropped 20 rounds of history and flattened every project to one snapshot. - `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap), drop at most one snapshot round, and mark the store `.gc-pending`. - `prune_checkpoints` gcs only when a ref moved (project deleted, or the pending marker), and its cap loop is drop -> gc -> re-measure. - `maybe_auto_prune_checkpoints` claims the interval marker before the run so a failing prune costs one day, not a gc per housekeeping tick. - `auto_prune_from_config` is the one config-driven entry point; the gateway calls it from the housekeeping tick (last chore), the CLI from a daemon thread. Nothing on either startup path waits for git. Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s -> 1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.