0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Every hand-rolled writer this PR folded into utils._atomic_write created a
NEW file with write_text()/open("w"), i.e. at 0o666 masked by the umask
(0644 under 022). The canonical helper publishes through mkstemp, whose
temp is 0600, and with no explicit mode and no existing target to copy bits
from it left that 0600 in place - so debug, model_catalog, profiles,
breadcrumbs, worktree_ops, web_result_cache, plugin_compat, write_approval,
rich_sent_store, active_sessions and the google_meet state files were
silently tightened to owner-only, the volume-mount hazard
_restore_file_metadata's own docstring warns about. Undeclared in the PR.
Fix at the canonical: when mode is None and the target does not exist,
apply default_new_file_mode() (0o666 masked by the umask, read via the
umask two-call trick with a transient 0o077 so a racing thread can only get
a tighter file). The helper is hermes_cli/backup._default_new_file_mode
moved into utils and reused. Secret writers (mode=0o600) are 0600 before,
during and after as before; an existing target keeps its bits; on non-POSIX
the helper returns None so nothing is chmod'd.
The canonical writer defaults to ensure_ascii=False, but ~10 of the sites
repointed onto it (terminal breadcrumbs, shell-hook allowlist, active
sessions, debug pending, model-catalog cache, the credential writers)
previously used json's ensure_ascii=True default. A surrogate-escaped str
(os.fsdecode of a non-UTF-8 cwd/argv) that json used to persist as \udcff
now made the utf-8 text handle raise UnicodeEncodeError - a ValueError that
the callers' `except OSError` never catches, so breadcrumbs silently stopped
writing and the other sites leaked a new error type.
Fix at the canonical: serialize to a str first (so nothing lands in the
temp file on failure), and on UnicodeEncodeError retry the dump with
ensure_ascii=True. That escape round-trips - json.loads returns the same
str with the lone surrogate - whereas encoding with surrogateescape emits a
raw 0xFF byte the reader's utf-8 decode rejects. The happy path is
unchanged: normal content keeps its raw UTF-8 bytes on disk.
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.