analysis: seal RQ3 confirmatory battery + frozen analysis + authoritative Holm-7
RQ3 live-docker battery (90/90 runs) complete; un-blind and apply the frozen prereg §6 plan. No re-specification. Seal: SHA256SUMS over confirmatory-data/ (battery-results.json 5b61e461... + 90 rq3-run.json sidecars) — raw data immutable. New analyzer analysis/rq3_confirm.py drives the frozen stats/metrics (untouched) via a run-level multi-arm bootstrap mirroring two_sample_diff_ci (10k BCa, alpha=0.05). Effect+CI always, never bare p. RQ3 result (honest null): - RQ3-P1-perf: retention margin agent-max(static,random) = -0.6pp, CI [-1.58,+0.39]pp -> FAILS +10pp gate (all selectors heal ~all churn, ~99%). - RQ3-P1-latency: added-latency(agent-min-baseline) = -13.5ms, CI [-52.1,+34.9], upper <= 100ms -> within budget (agent not slower). - RQ3-P2: rebuild-classifier AUC(agent vs pooled baseline) = 0.587, CI [0.458,0.703], upper 0.703 > 0.60 -> fingerprint NOT excluded (underpowered). - RQ3-P3 = H0 (P1-perf fails and P2 fails). Authoritative Holm-7 over the frozen size-7 family (supersedes the lead's conservative partial embedding; RQ2-P3 slot carries the mechanism-corrected H1-pooled Spearman p=0, not the lead degenerate p=1). Survivors: RQ1-P1, RQ2-P1 (shrink), RQ2-P3 (mix). Non-survivors: RQ1-P2, RQ3-P2, RQ3-P1-perf, RQ3-P1-latency. Sealed: analysis/rq3-confirmatory-analysis.json (e09c66ef...). Tests: tests/test_sor_rq3_confirm.py 6 passed; full SOR suite 207 passed. Both prereg SHAs intact; $0/offline analysis; worktree only. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in: