RQ2-P3 mechanism study (NEW prereg, approved-path/params-locked, not yet frozen):
pool-based willing-bridge instrument, concentration swept as a manipulated IV
(B∈{2,4,8}×α∈{0,1,2}), two-sided/direction-agnostic, run-level cluster bootstrap.
RQ3 companion run-brief (execution-freeze for the already-frozen RQ3): agent=qwen2.5:3b
local Ollama, churn kill_prob=30/steps=20, baselines + Holm-7 completion pinned.
note-unique-bridge-artifact.md: on-record analysis that the lead RQ2-P1 shrink is plausibly
a fresh-bridge-per-circuit artifact (unique exit-signature -> singleton set -> H~0), and that
realistic bridge reuse (a mix) may raise anonymity -> the mechanism study may qualify the lead
RQ2 headline. No edit to the frozen prereg or committed lead paper.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
7.2 KiB
Execution-freeze / run-brief (DRAFT) — RQ3 companion battery
STATUS: PARAMS LOCKED 2026-07-21 — execution plan, NOT a new prereg. RQ3's hypotheses, gates, and analysis are already frozen in
sor-consent-prereg.md(SHAf22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b, §3/§4/§6). Writing a "new RQ3 prereg" would be redundant and risk HARKing. This brief (a) pins the two open[APPROVAL]execution parameters (agent model, churn schedule — both locked below, blind to RQ3 outcomes) and (b) lists the build + gate steps. GO gate before the confirmatory battery: the rebuild-classifier calibration gate + the churn-bites gate must pass green (§3-4). Build + synthetic tests MAY proceed now.
1. The frozen gates (verbatim — do not restate as new)
Family stays the frozen size-7 {RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}. The three RQ3 tests were always in the family.
| Test | Frozen gate |
|---|---|
| RQ3-P1-perf | throughput-retention(agent) − max(static,random): CI lower bound ≥ 10 pp |
| RQ3-P1-latency | added-latency(agent): CI upper bound ≤ 100 ms |
| RQ3-P2 | rebuild-classifier AUC: CI upper bound ≤ 0.60 |
| RQ3-P3 | logical AND: CONFIRM iff P1 ∧ P2 (not a new test) |
R = 30 runs/cell, C = 50 circuits/run, fixed stopping rule (no optional stopping / interim looks; uninformative cell → inconclusive). All inherited unchanged.
2. Open [APPROVAL] params to PIN before GO (the actual gate)
The frozen prereg left two items tagged [APPROVAL] pinned at freeze — the [APPROVAL]
tag is still literally present in the frozen text, i.e. neither was concretely substituted.
These must be pinned now, blind to RQ3 data:
(A) Agent selector model ID + weights digest (§3 L71-73) — LOCKED 2026-07-21.
- Pin:
qwen2.5:3bvia local Ollama (agent_selector.OllamaAgentPolicy, temp 0, per-run seed, state-hash cache; heuristic fallback on query failure). Record the exact model tag + SHA-256 weights digest in every manifest (digest captured at run time). - Rails satisfied: local / open-weight, $0, no paid-frontier call. The Claude/frontier
arm (
ClaudeExploratoryPolicy) stays inert / EXPLORATORY / budget-gated — not wired. - Reproducibility caveat (accepted; must be stated in the paper). Ollama at temp 0 is not bit-identical across machines (quantization / GPU logit drift). The agent arm is reproducible via the committed decision-log + state-hash cache replay, not via independent model re-execution on other hardware — same honesty class as the RQ1 timing caveat. Commit the decision log + cache as the reproducibility anchor.
(B) Churn schedule set (§3 L74, §4) — LOCKED 2026-07-21. churn.churn_schedule(seed, nodes, steps, kill_prob_pct) is implemented and deterministic.
- Pin: a seeded set of schedules — one schedule per run, seed drawn from the same
SHA256(S0‖cell‖run)family so every run's churn is reproducible and provenance-anchored.kill_prob_pct = 30,steps = 20. - Churn must actually bite (gate, §4). With these values, confirm on a dry pass that
circuits genuinely lose hops and rebuild (non-zero
drops/rebuildsper run). If drop rate is trivially zero, the throughput-retention and rebuild-classifier tests are degenerate — STOP and surface rather than report a zero-variance null.
3. Build gaps (hand to the driver in tmux; not built inline)
Verified against the current worktree — what exists vs. what's missing:
Ready (no work): selector.py static/random/agent all implemented + tested;
agent_selector.OllamaAgentPolicy reproducible; churn.py deterministic schedules wired
into run_selection(); analysis/metrics.py has throughput_retention() +
rebuild_classifier_auc().
Missing (build):
battery.enumerate_cells()(battery.py:73) enumerates RQ1+RQ2 only (6 cells). Add RQ3 cells:selector ∈ {static, random, agent} × churn-schedule(s)at the 1house/bridge-off control; interleave the control arm like RQ1/RQ2.executor.run_battery()— wire RQ3 collection: driverun_selection()over the pinned schedule; collect throughput-retention, added-latency, and rebuild-event gaps → classifier AUC. Baselines LOCKED: added-latency = median e2e latency(agent) − median latency(min-latency baseline arm, i.e. the faster of static/random); perf margin uses max(static, random) per the frozen gate.- Rebuild-classifier calibration gate — calibrate the classifier on labeled
churned-vs-low-churn control signals before the confirmatory RQ3 battery (same
discipline as the RQ1 correlator gate). Low-churn baseline LOCKED:
kill_prob_pct = 5. Gate passes iff churned-vs-baseline is separable (AUC≈1) and baseline-vs-baseline is not (AUC≈0.5). Boolean; blocks the run.
4. Instrument-validation gate (boolean; blocks RQ3 confirmatory)
Reuse the frozen 6-item gate spirit, plus RQ3-specific:
- Rebuild classifier calibrated on control signals (AUC≈1 churned-vs-baseline separable, AUC≈0.5 baseline-vs-baseline) — not fit to confirmatory cells.
- Agent selector reproducibility check: same seed + state-hash → byte-identical circuit choice (cache replay), model+digest echoed to manifest.
- Isolation (
assert engine != local), provenance integrity (sealedevents.jsonl).
5. Analysis
- Reuse
stage06_runon the sealed RQ3 data; compute the three RQ3 tests with effect + BCa CI. - Complete the Holm step-down over the full family of 7 now that all seven p-values exist. (The lead paper's 7/6/5/4 was a deliberately conservative partial embedding; the companion computes the exact Holm-7 — both remain valid, the partial never under-corrects.) The companion's Holm-7 is the authoritative final correction over the whole family, including the lead RQ1/RQ2 tests; RQ1-P1 and RQ2-P1 survive regardless (raw p ≈ 0), but state explicitly that the companion supersedes the lead's partial embedding.
- Confirmatory vs exploratory labeled; nulls reported honestly (a selector that does not beat baselines, or a rebuild pattern that is a fingerprint, is the finding).
6. QUIC / ssh3 transport arm — stays EXPLORATORY
Frozen as EXPLORATORY, orthogonal to all three RQs, and not built (design decision D3; §3 L75-78). Recommendation: defer. Promoting it doubles every cell + requires a new transport substrate for a weak lit-gap. If ever built, carry only as a labeled EXPLORATORY latency-sensitivity note — never in the Holm family.
7. Rails (immutable, inherited)
Frozen prereg never edited; execution pins recorded here + in manifests; no HARKing; nulls
are results. Containment: isolated docker only, self-traffic, lab-only. Budget $0 —
local open-weight agent only, no paid arm; surface before any spend. Worktree only on
feat/sor-consent-relay. Raw data immutable + SHA-256.
8. Pin block (to complete at GO)
GO DATE: <date>
AGENT MODEL: qwen2.5:3b WEIGHTS DIGEST: <sha256>
CHURN: kill_prob_pct=30 steps=<S> schedules=<seeded set spec>
APPROVED BY: <operator>
FROZEN PREREG SHA (unchanged): f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b