BasePlatformAdapter._acquire_platform_lock emits `{scope}_lock` with
retryable=True on purpose (#54167): a MID-RUN reconnect must be able to
recover once the live holder exits or a stale record is cleared. The
startup router keyed solely off that flag, so a live foreign holder of the
bot token at zero-connected startup landed in `_failed_platforms` with
gateway_state=running — alive, deaf, and retry-storming the token every
backoff — instead of the exit-78 (EX_CONFIG / startup_failed) contract
that #51228 established for single-writer conflicts.
Minimal class fix, salvaged from #83183 (@alexgunsberg) against current
main:
- gateway/restart.py: `is_global_startup_conflict(error_code)` — matches
the `*_lock` / `lock_conflict` code families every adapter emits for
scoped-lock and identity conflicts. Code only, never message text.
- gateway/run.py primary startup routing: a lock-conflict failure is
routed as non-retryable (parked `fatal`, not queued). Nothing else
connected → exit 78; alongside a transient peer → NS-609 mixed mode,
gateway stays alive and only the peer retries.
- gateway/run.py `_schedule_secondary_profile_startup_reconnect`: the same
contract for multiplex secondaries — park `<profile>:<platform>` fatal
like `duplicate_credential` instead of scheduling a reconnect storm.
- Mid-run behavior is untouched: `_handle_adapter_fatal_error_impl` and
the reconnect watcher still treat `*_lock` as retryable (#54167).
Not carried over from #83183 (superseded on main or out of scope): the
`degraded` lifecycle write only fires on the all-retryable path and the
runner immediately overwrites it with `running` (so busy/drain already
see `running`); the secondary retry bridge landed separately in
96489f3c1b (#92064); Buzz/IRC/LINE lock-tuple unpack and the reconnect
ownership registry are separate class fixes.
Live repro (real GatewayRunner.start(), isolated HERMES_HOME + lock dir,
live holder subprocess owning the lock via production
acquire_scoped_lock): before — exit_code=None, gateway_state=running,
telegram `retrying`, queued in _failed_platforms; after — exit_code=78,
gateway_state=startup_failed, telegram `fatal`, _failed_platforms={}.
Co-authored-by: alexgunsberg <alex@gunsberg.fi>
The generated unit only counted restart_drain_timeout, so a default
cron drain (30s + 10s cleanup) could still be inside budget when
systemd SIGKILLed the cgroup. Size TimeoutStopSec from max(drain,
cron floor + reserve) plus headroom so an in-budget stop is not killed.
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.
A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.
Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.
The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.
Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.
Relates to #82161 (complements #82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
request_restart was calling stop() immediately, so the requesting turn stayed
in the drain wait set and got force-killed at restart_drain_timeout. Wait for
active work to reach zero first, then stop against an idle gateway.
The restart-routing, systemd-support, and subprocess-HOME tests asserted
branch behavior but left part of the real probe surface unmocked, so they
fail when the suite itself runs inside a container (self-hosted CI) or a
launchd-descended shell:
- /restart routing tests: the handler also consults the real /.dockerenv —
extract the inline probe to gateway.restart.is_container_restart_context()
(patchable seam, no behavior change) and pin it False; scrub ALL four
supervisor env markers (ambient XPC_SERVICE_NAME on macOS flipped one).
- supports_systemd_services tests: pin shutil.which('systemctl') and
is_container() so the test asserts the branch, not the host.
- copilot ACP real-HOME test: pin is_container() (auto mode prefers profile
home in containers) and scrub ambient HERMES_REAL_HOME/TERMINAL_HOME_MODE.
97 tests green on macOS dev box AND inside a docker CI runner container.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Profiles without their own messaging token inherit the default
profile's token via os.getenv, hit a token collision, and exit with
startup_failed. s6 restarts them immediately, creating ~30MB tirith
sandbox dirs in /tmp each cycle — filling the disk in hours (#51228).
Changes:
- gateway/restart.py: add GATEWAY_FATAL_CONFIG_EXIT_CODE = 78
- gateway/run.py: set exit_code=78 on non-retryable startup errors
(token collision, no platforms)
- hermes_cli/service_manager.py: add _render_finish_script() that
translates exit 78 → exit 125 (s6 permanent failure)
- hermes_cli/container_boot.py: write finish script alongside run
script during profile registration
The s6 finish script pattern follows docker/s6-rc.d/dashboard/finish.
Closes#51228