fix(observability): attribute ACP and batch execution surfaces

Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:

1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
   but "acp" was absent from EXECUTION_SURFACES, so the contract's
   closed-schema fallback folded every editor session into "other" --
   the bucket meant for genuinely unclassifiable traffic.

2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
   "platform" entirely, so every batch task run reported "unknown"
   despite "batch" already being a first-class surface.

Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".

Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
  entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
  self.platform = "batch" on the runner, and defaulted at the worker call
  site so callers that build a config without it stay attributable

Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.

Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
  M1 revert acp from EXECUTION_SURFACES  -> 3 failed
  M2 revert acp entrypoint mapping only  -> 1 failed
  M3 revert batch passthrough            -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
This commit is contained in:
Ben Barclay
2026-09-09 11:27:36 +10:00
parent 72a3277cd7
commit 5a1246f830
4 changed files with 94 additions and 4 deletions
+11 -1
View File
@@ -58,6 +58,9 @@ _AGENT_PASSTHROUGH = (
"base_url", "api_key", "ephemeral_system_prompt", "providers_allowed", "providers_ignored",
"providers_order", "provider_sort", "openrouter_min_coding_score",
"reasoning_config", "prefill_messages",
# Without this, every batch task run is attributed to the "unknown" execution
# surface in shared metrics even though "batch" is a first-class surface.
"platform",
)
@@ -251,7 +254,11 @@ def _process_single_prompt(
log_prefix=f"[B{batch_num}:P{prompt_index}]",
skip_context_files=True, # Don't pollute trajectories with SOUL.md/AGENTS.md
skip_memory=True, # Don't use persistent memory in batch runs
**{key: config.get(key) for key in _AGENT_PASSTHROUGH},
**{key: config.get(key) for key in _AGENT_PASSTHROUGH if key != "platform"},
# Batch is a first-class execution surface. Defaulting here (rather than
# relying on the caller's config dict) keeps task-run telemetry attributable
# even for callers that build a config without it.
platform=config.get("platform") or "batch",
)
# task_id ensures each task gets its own isolated VM
@@ -442,6 +449,9 @@ class BatchRunner:
self.dataset_file = Path(dataset_file)
for name in _RUNNER_FIELDS:
setattr(self, name, params[name])
# Batch runs are their own execution surface; declaring it here keeps every
# worker's task-run telemetry attributable instead of falling back to "unknown".
self.platform = "batch"
if not validate_distribution(distribution):
raise ValueError(f"Unknown distribution: {distribution}. Available: {list(list_distributions().keys())}")