The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:
- a routed desktop-ticker fire runs under multiplex semantics for exactly
its scope: a scope miss returns None instead of the launch credential and
the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
.env defined it or a launch external source supplied it (applied or lost
to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
profile's own value.
Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.
Review findings on f5f88d5058. Three are defects the previous round introduced.
Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.
Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.
Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.
Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.
Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.
(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
Review findings on d8c467f223, each reproduced through its production path.
Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.
Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.
Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).
Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.
(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).
Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (1aa62ceb45 isolated the routed reload but not the reset).
Change:
- env_loader: snapshot every name a source supplied (provenance + skipped_existing) from
the home's effective environment, so a shadowed re-apply keeps the values it had.
- reset_secret_source_cache(hermes_home=None): optional per-home reset; global clear kept
for tests/config edits.
- cron per-fire re-pull and plugins._refresh_secret_sources_after_discovery reset + reload
only the home they resolve to.
- Two invariant tests (red on base) in tests/test_env_loader_secret_sources.py; docs note in
secret-source-plugin.md.
Reported-by: luochen1990
Addresses #102041
Address teknium1 review on #64189:
- Re-pull gate now delegates to each source's is_enabled(cfg) via the
registry contract, so a plugin source with custom activation logic is
honored (previously only secrets.<name>.enabled was checked).
- Add BUILTIN_SOURCE_NAMES to the registry so plugin-vs-bundled is a
single source of truth instead of a hard-coded set at the call site.
- Reconcile docs: rewrite the timing :::note to describe both the
post-discovery re-pull and the remaining import-time limitation, and
cross-link the first-process bootstrap section.
- Tests: real SecretSource subclasses, custom is_enabled activation
(positive + negative), is_enabled-raises skip, builtin-only no-op,
and a discovery-registration end-to-end re-pull check.
After plugins register SecretSource backends, reset the env-loader cache
and re-run load_hermes_dotenv when an enabled plugin secret source is
configured. Closes the first-process bootstrap gap where import-time env
load stale-outs plugin vaults (tommck / Community ask). Fail-open, no-op
without plugin sources.
Docs: first-process bootstrap timing on secret-source plugin guide.
Tests: unit coverage for noop / enabled re-pull / discover hook.
Part of #64182 plugin-interface expansion.