# Execution-freeze / run-brief (DRAFT) — RQ3 companion battery > **STATUS: PARAMS LOCKED 2026-07-21 — execution plan, NOT a new prereg.** RQ3's hypotheses, > gates, and analysis are **already frozen** in `sor-consent-prereg.md` (SHA > `f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b`, §3/§4/§6). Writing a > "new RQ3 prereg" would be redundant and risk HARKing. This brief (a) **pins the two open > `[APPROVAL]` execution parameters** (agent model, churn schedule — both locked below, > blind to RQ3 outcomes) and (b) lists the build + gate steps. **GO gate before the > confirmatory battery: the rebuild-classifier calibration gate + the churn-bites gate must > pass green (§3-4).** Build + synthetic tests MAY proceed now. ## 1. The frozen gates (verbatim — do not restate as new) Family stays the frozen size-7 {RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, **RQ3-P1-perf, RQ3-P1-latency, RQ3-P2**}. The three RQ3 tests were always in the family. | Test | Frozen gate | |---|---| | RQ3-P1-perf | throughput-retention(agent) − max(static,random): CI lower bound ≥ **10 pp** | | RQ3-P1-latency | added-latency(agent): CI upper bound ≤ **100 ms** | | RQ3-P2 | rebuild-classifier AUC: CI upper bound ≤ **0.60** | | RQ3-P3 | logical AND: CONFIRM iff P1 ∧ P2 (not a new test) | R = 30 runs/cell, C = 50 circuits/run, fixed stopping rule (no optional stopping / interim looks; uninformative cell → inconclusive). All inherited unchanged. ## 2. Open `[APPROVAL]` params to PIN before GO (the actual gate) The frozen prereg left two items tagged `[APPROVAL] pinned at freeze` — **the `[APPROVAL]` tag is still literally present in the frozen text, i.e. neither was concretely substituted.** These must be pinned now, blind to RQ3 data: **(A) Agent selector model ID + weights digest** (§3 L71-73) — **LOCKED 2026-07-21.** - Pin: **`qwen2.5:3b`** via local Ollama (`agent_selector.OllamaAgentPolicy`, temp 0, per-run seed, state-hash cache; heuristic fallback on query failure). Record the exact **model tag + SHA-256 weights digest** in every manifest (digest captured at run time). - Rails satisfied: **local / open-weight, $0**, no paid-frontier call. The Claude/frontier arm (`ClaudeExploratoryPolicy`) stays **inert / EXPLORATORY / budget-gated** — not wired. - **Reproducibility caveat (accepted; must be stated in the paper).** Ollama at temp 0 is **not bit-identical across machines** (quantization / GPU logit drift). The agent arm is reproducible via the **committed decision-log + state-hash cache replay**, *not* via independent model re-execution on other hardware — same honesty class as the RQ1 timing caveat. Commit the decision log + cache as the reproducibility anchor. **(B) Churn schedule set** (§3 L74, §4) — **LOCKED 2026-07-21.** `churn.churn_schedule(seed, nodes, steps, kill_prob_pct)` is implemented and deterministic. - Pin: a **seeded set** of schedules — one schedule per run, seed drawn from the same `SHA256(S0‖cell‖run)` family so every run's churn is reproducible and provenance-anchored. **`kill_prob_pct = 30`, `steps = 20`.** - **Churn must actually bite (gate, §4).** With these values, confirm on a dry pass that circuits genuinely lose hops and rebuild (non-zero `drops`/`rebuilds` per run). If drop rate is trivially zero, the throughput-retention and rebuild-classifier tests are degenerate — STOP and surface rather than report a zero-variance null. ## 3. Build gaps (hand to the driver in tmux; not built inline) Verified against the current worktree — what exists vs. what's missing: **Ready (no work):** `selector.py` static/random/agent all implemented + tested; `agent_selector.OllamaAgentPolicy` reproducible; `churn.py` deterministic schedules wired into `run_selection()`; `analysis/metrics.py` has `throughput_retention()` + `rebuild_classifier_auc()`. **Missing (build):** 1. **`battery.enumerate_cells()`** (`battery.py:73`) enumerates RQ1+RQ2 only (6 cells). Add RQ3 cells: `selector ∈ {static, random, agent} × churn-schedule(s)` at the 1house/bridge-off control; interleave the control arm like RQ1/RQ2. 2. **`executor.run_battery()`** — wire RQ3 collection: drive `run_selection()` over the pinned schedule; collect throughput-retention, **added-latency**, and rebuild-event gaps → classifier AUC. **Baselines LOCKED:** added-latency = median e2e latency(agent) − median latency(**min-latency baseline arm**, i.e. the faster of static/random); perf margin uses **max(static, random)** per the frozen gate. 3. **Rebuild-classifier calibration gate** — calibrate the classifier on labeled churned-vs-low-churn control signals **before** the confirmatory RQ3 battery (same discipline as the RQ1 correlator gate). **Low-churn baseline LOCKED: `kill_prob_pct = 5`.** Gate passes iff churned-vs-baseline is separable (AUC≈1) and baseline-vs-baseline is not (AUC≈0.5). Boolean; blocks the run. ## 4. Instrument-validation gate (boolean; blocks RQ3 confirmatory) Reuse the frozen 6-item gate spirit, plus RQ3-specific: - Rebuild classifier calibrated on control signals (AUC≈1 churned-vs-baseline separable, AUC≈0.5 baseline-vs-baseline) — **not fit to confirmatory cells**. - Agent selector **reproducibility check**: same seed + state-hash → byte-identical circuit choice (cache replay), model+digest echoed to manifest. - Isolation (`assert engine != local`), provenance integrity (sealed `events.jsonl`). ## 5. Analysis - Reuse `stage06_run` on the sealed RQ3 data; compute the three RQ3 tests with effect + BCa CI. - **Complete the Holm step-down over the full family of 7** now that all seven p-values exist. (The lead paper's 7/6/5/4 was a deliberately conservative *partial* embedding; the companion computes the exact Holm-7 — both remain valid, the partial never under-corrects.) **The companion's Holm-7 is the authoritative final correction over the whole family**, including the lead RQ1/RQ2 tests; RQ1-P1 and RQ2-P1 survive regardless (raw p ≈ 0), but state explicitly that the companion supersedes the lead's partial embedding. - Confirmatory vs exploratory labeled; nulls reported honestly (a selector that does **not** beat baselines, or a rebuild pattern that **is** a fingerprint, is the finding). ## 6. QUIC / ssh3 transport arm — stays EXPLORATORY Frozen as **EXPLORATORY, orthogonal to all three RQs, and not built** (design decision D3; §3 L75-78). Recommendation: **defer.** Promoting it doubles every cell + requires a new transport substrate for a weak lit-gap. If ever built, carry only as a labeled EXPLORATORY latency-sensitivity note — never in the Holm family. ## 7. Rails (immutable, inherited) Frozen prereg never edited; execution pins recorded here + in manifests; no HARKing; nulls are results. Containment: isolated docker only, self-traffic, lab-only. **Budget $0** — local open-weight agent only, no paid arm; surface before any spend. Worktree only on `feat/sor-consent-relay`. Raw data immutable + SHA-256. ## 8. Pin block (to complete at GO) ``` GO DATE: AGENT MODEL: qwen2.5:3b WEIGHTS DIGEST: CHURN: kill_prob_pct=30 steps= schedules= APPROVED BY: FROZEN PREREG SHA (unchanged): f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b ```