analysis: seal RQ3 confirmatory battery + frozen analysis + authoritative Holm-7
RQ3 live-docker battery (90/90 runs) complete; un-blind and apply the frozen prereg §6 plan. No re-specification. Seal: SHA256SUMS over confirmatory-data/ (battery-results.json 5b61e461... + 90 rq3-run.json sidecars) — raw data immutable. New analyzer analysis/rq3_confirm.py drives the frozen stats/metrics (untouched) via a run-level multi-arm bootstrap mirroring two_sample_diff_ci (10k BCa, alpha=0.05). Effect+CI always, never bare p. RQ3 result (honest null): - RQ3-P1-perf: retention margin agent-max(static,random) = -0.6pp, CI [-1.58,+0.39]pp -> FAILS +10pp gate (all selectors heal ~all churn, ~99%). - RQ3-P1-latency: added-latency(agent-min-baseline) = -13.5ms, CI [-52.1,+34.9], upper <= 100ms -> within budget (agent not slower). - RQ3-P2: rebuild-classifier AUC(agent vs pooled baseline) = 0.587, CI [0.458,0.703], upper 0.703 > 0.60 -> fingerprint NOT excluded (underpowered). - RQ3-P3 = H0 (P1-perf fails and P2 fails). Authoritative Holm-7 over the frozen size-7 family (supersedes the lead's conservative partial embedding; RQ2-P3 slot carries the mechanism-corrected H1-pooled Spearman p=0, not the lead degenerate p=1). Survivors: RQ1-P1, RQ2-P1 (shrink), RQ2-P3 (mix). Non-survivors: RQ1-P2, RQ3-P2, RQ3-P1-perf, RQ3-P1-latency. Sealed: analysis/rq3-confirmatory-analysis.json (e09c66ef...). Tests: tests/test_sor_rq3_confirm.py 6 passed; full SOR suite 207 passed. Both prereg SHAs intact; $0/offline analysis; worktree only. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -106,3 +106,6 @@ RQ2 RATIFIED (operator Andre, 2026-07-20, while blind) — observation-consisten
|
|||||||
- RUNNING (do NOT trim): `executor.run_rq3_battery(live=True)` collecting the full frozen schedule — selector∈{static,random,agent} × pinned churn kp30/steps20, R=30 × C=50 = 4,500 real isolated-docker circuits (~12h est. at ~9.8s/circuit). Launch dir `output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/` (gitignored; force-add anchors on completion). Verified genuinely progressing: python PID 51087 alive + live 4-container circuit (client+3 hops) up on docker. **Completion-detection: `rq3-battery-results.json` written once at end; progress = count of `rq3-*/rq3-run.json` sidecars → 90 (3 cells × 30 runs).** Containment intact (isolated-docker only, self-traffic, $0 — local heuristic/random arms; no frontier spend). Both prereg SHAs intact; worktree-only. Analysis (RQ3-P1-perf/latency, RQ3-P2, Holm-7) runs AFTER completion on the sealed record.
|
- RUNNING (do NOT trim): `executor.run_rq3_battery(live=True)` collecting the full frozen schedule — selector∈{static,random,agent} × pinned churn kp30/steps20, R=30 × C=50 = 4,500 real isolated-docker circuits (~12h est. at ~9.8s/circuit). Launch dir `output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/` (gitignored; force-add anchors on completion). Verified genuinely progressing: python PID 51087 alive + live 4-container circuit (client+3 hops) up on docker. **Completion-detection: `rq3-battery-results.json` written once at end; progress = count of `rq3-*/rq3-run.json` sidecars → 90 (3 cells × 30 runs).** Containment intact (isolated-docker only, self-traffic, $0 — local heuristic/random arms; no frontier spend). Both prereg SHAs intact; worktree-only. Analysis (RQ3-P1-perf/latency, RQ3-P2, Holm-7) runs AFTER completion on the sealed record.
|
||||||
- 2026-07-21 COMPANION PAPER — RQ2-P3 HALF FILLED POST-SEAL (D3 combined paper). `docs/stage-07-companion-methods.md`: §3 un-marked FREEZE-PENDING → FROZEN 2026-07-21 (SHA 8db4e8a7…); §5 Results + §6 Discussion for RQ2-P3 filled FROM THE SEALED RECORD ONLY (H1 ρ=+0.6244 CI[+0.5941,+0.6545]; H2 β=+0.7052 CI[+0.6195,+0.7903] n=270; H3 RESOLVED=MIX; Holm own {H1,H2} both reject — effect+CI, never bare p). Finding stated plainly: shared-pool concentration RAISES anonymity (mix), REFUTES naive funnel; framed as honest QUALIFICATION/CORRECTION of lead RQ2-P1 "shrink" as a unique-bridge (fresh-bridge-per-circuit) artifact — cites `docs/note-unique-bridge-artifact.md` + frozen mechanism prereg SHA 8db4e8a7…. MANDATORY disclosure carried: §7 dry-pass previewed direction (ρ 0→+0.838); confirmatory quantifies effect already visible at calibration; pre-committed hypotheses were two-sided.
|
- 2026-07-21 COMPANION PAPER — RQ2-P3 HALF FILLED POST-SEAL (D3 combined paper). `docs/stage-07-companion-methods.md`: §3 un-marked FREEZE-PENDING → FROZEN 2026-07-21 (SHA 8db4e8a7…); §5 Results + §6 Discussion for RQ2-P3 filled FROM THE SEALED RECORD ONLY (H1 ρ=+0.6244 CI[+0.5941,+0.6545]; H2 β=+0.7052 CI[+0.6195,+0.7903] n=270; H3 RESOLVED=MIX; Holm own {H1,H2} both reject — effect+CI, never bare p). Finding stated plainly: shared-pool concentration RAISES anonymity (mix), REFUTES naive funnel; framed as honest QUALIFICATION/CORRECTION of lead RQ2-P1 "shrink" as a unique-bridge (fresh-bridge-per-circuit) artifact — cites `docs/note-unique-bridge-artifact.md` + frozen mechanism prereg SHA 8db4e8a7…. MANDATORY disclosure carried: §7 dry-pass previewed direction (ρ 0→+0.838); confirmatory quantifies effect already visible at calibration; pre-committed hypotheses were two-sided.
|
||||||
- BLIND HOLD (RQ3): RQ3 Results/Discussion AND authoritative Holm-7 LEFT AS HELD-BLIND placeholders — RQ3 battery still running (progress read by sidecar COUNT only = 3/90; NO rq3-run.json contents read). Lead RQ2-P1 NOT re-litigated (companion QUALIFIES; committed lead paper + frozen prereg f22331a72e… untouched). This stone: $0/offline, worktree-only. NEXT: HOLD for RQ3 — no new stones, no peeking at RQ3 data until `rq3-battery-results.json` lands.
|
- BLIND HOLD (RQ3): RQ3 Results/Discussion AND authoritative Holm-7 LEFT AS HELD-BLIND placeholders — RQ3 battery still running (progress read by sidecar COUNT only = 3/90; NO rq3-run.json contents read). Lead RQ2-P1 NOT re-litigated (companion QUALIFIES; committed lead paper + frozen prereg f22331a72e… untouched). This stone: $0/offline, worktree-only. NEXT: HOLD for RQ3 — no new stones, no peeking at RQ3 data until `rq3-battery-results.json` lands.
|
||||||
|
- 2026-07-22 RQ3 BATTERY COMPLETE (90/90 runs, live-docker-e2e, exited clean) → UN-BLIND + FROZEN analysis. RAW SEALED: `.../20260722T040640Z/confirmatory-data/SHA256SUMS` (battery-results.json `5b61e461…` + 90 rq3-run.json sidecars). New analyzer `analysis/rq3_confirm.py` (RQ3 analogue of rq2p3_confirm; frozen `stats`/`metrics` UNTOUCHED — run-level multi-arm bootstrap mirrors `two_sample_diff_ci`, 10k BCa, α=0.05). Sealed analysis `.../20260722T040640Z/analysis/rq3-confirmatory-analysis.json` (SHA `e09c66ef…`). Tests `tests/test_sor_rq3_confirm.py` 6 passed; full SOR suite 207 passed (no regression).
|
||||||
|
- RQ3 RESULT (effect+CI, never bare p): **RQ3-P1-perf FAIL/H0** — retention margin agent−max(static,random) = **−0.6pp**, BCa CI [−1.58pp, +0.39pp]; every selector heals ~all churn drops (~99% retention) so no ≥+10pp agent gain. **RQ3-P1-latency HOLDS budget** — added-latency(agent−min-baseline=random) = **−13.5ms**, CI [−52.1, +34.9]ms, upper ≤100ms (agent not slower). **RQ3-P2 FAIL/not-excluded** — rebuild-classifier AUC(agent vs pooled baseline, per-run mean-gap) = **0.587**, CI [0.458, 0.703], upper 0.703 > 0.60 → fingerprint NOT excluded (underpowered at n=30). **RQ3-P3 = H0** (P1-perf fails ∧ P2 fails). Honest null: the local agent selector neither beats baselines nor is certifiably non-classifiable.
|
||||||
|
- AUTHORITATIVE HOLM-7 (frozen family size=7; supersedes lead's conservative partial embedding, D3). RQ2-P3 slot carries the mechanism-corrected primary **H1-pooled Spearman p=0** (not the lead's degenerate p=1). **SURVIVORS: RQ1-P1 (r1×7), RQ2-P1 shrink (r2×6), RQ2-P3 mix (r3×5)** all holm_p=0. NON-survivors: RQ1-P2 (r4×4, holm_p=0.365), RQ3-P2 (r5×3, 0.511), RQ3-P1-perf (r6×2, 0.511), RQ3-P1-latency (r7×1, 0.511). Lead RQ1-P1 & RQ2-P1 survive regardless. Both prereg SHAs intact; $0/offline; worktree-only. NEXT: fill RQ3 half of companion paper (stone 2).
|
||||||
|
|||||||
@@ -0,0 +1,310 @@
|
|||||||
|
"""RQ3 confirmatory analysis — reads the SEALED live-docker battery, applies the
|
||||||
|
FROZEN prereg §6 plan, and computes the authoritative Holm-7 over the whole family.
|
||||||
|
|
||||||
|
This is the RQ3 analogue of ``rq2p3_confirm.py``: it does **not** re-specify anything.
|
||||||
|
It loads the immutable ``rq3-battery-results.json`` (the operator-gated live-docker
|
||||||
|
run) and the pinned params in ``docs/rq3-companion-run-brief.md`` and computes the three
|
||||||
|
frozen RQ3 confirmatory tests, each as an **effect + BCa 95% CI** (never a bare p; the p
|
||||||
|
is carried only to order the Holm family, per prereg §6):
|
||||||
|
|
||||||
|
* **RQ3-P1-perf** — throughput-retention margin: mean retention(agent) −
|
||||||
|
max(mean retention(static), mean retention(random)); perf holds iff the CI **lower**
|
||||||
|
bound ≥ +10 pp (frozen gate).
|
||||||
|
* **RQ3-P1-latency** — added-latency(agent) = median e2e latency(agent) −
|
||||||
|
min(median latency(static), median latency(random)) [the faster / min-latency baseline
|
||||||
|
arm, run-brief §3.2]; anonymity-latency budget holds iff the CI **upper** bound ≤ 100 ms.
|
||||||
|
* **RQ3-P2** — rebuild-classifier AUC separating the agent selector's per-run mean
|
||||||
|
inter-rebuild-gap signal from the pooled baseline selectors' signal (the fingerprint
|
||||||
|
question, §1(b)); anonymity holds iff the CI **upper** bound ≤ 0.60.
|
||||||
|
* **RQ3-P3** — logical AND: CONFIRM iff (P1-perf ∧ P1-latency) ∧ P2; else H0.
|
||||||
|
|
||||||
|
The BCa machinery is the frozen ``stats`` toolkit — this harness only *drives* it at the
|
||||||
|
run level (resample whole runs, the confirmatory grouping unit) via a generic multi-arm
|
||||||
|
bootstrap that mirrors ``stats.two_sample_diff_ci`` (independent per-arm resampling, a
|
||||||
|
combined leave-one-out jackknife, ``stats._bca_endpoints``). No frozen instrument is edited.
|
||||||
|
|
||||||
|
**Authoritative Holm-7.** Once all seven confirmatory p-values exist, this computes the
|
||||||
|
exact ``stats.holm_bonferroni`` step-down over the frozen size-7 family
|
||||||
|
{RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}. The four prior
|
||||||
|
p-values are read from the SEALED lead record (RQ1-P1, RQ1-P2, RQ2-P1) and the SEALED
|
||||||
|
RQ2-P3′ mechanism record (RQ2-P3 slot = its primary H1-pooled Spearman ρ — the direct
|
||||||
|
operationalization of the frozen single-slot "Spearman ρ between concentration and H";
|
||||||
|
the lead's degenerate as-instrumented RQ2-P3 is superseded by the mechanism-corrected
|
||||||
|
result, per operator decision D3). This Holm-7 is the authoritative final correction and
|
||||||
|
supersedes the lead paper's deliberately conservative *partial* embedding; both remain
|
||||||
|
valid, the partial never under-corrects.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import statistics
|
||||||
|
import sys
|
||||||
|
from collections import defaultdict
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Callable, Dict, List, Sequence
|
||||||
|
|
||||||
|
from cmd_chat.sor.analysis import stats
|
||||||
|
from cmd_chat.sor.analysis.metrics import rebuild_classifier_auc
|
||||||
|
|
||||||
|
# --- frozen family + gates -------------------------------------------------- #
|
||||||
|
FAMILY_SIZE = 7
|
||||||
|
FROZEN_FAMILY = (
|
||||||
|
"RQ1-P1", "RQ1-P2", "RQ2-P1", "RQ2-P3",
|
||||||
|
"RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2",
|
||||||
|
)
|
||||||
|
PERF_MARGIN_MIN_PP = 0.10 # RQ3-P1-perf gate: CI lower ≥ +10 pp
|
||||||
|
LATENCY_MAX_MS = 100.0 # RQ3-P1-latency gate: CI upper ≤ 100 ms
|
||||||
|
AUC_CEILING = 0.60 # RQ3-P2 gate: CI upper ≤ 0.60
|
||||||
|
|
||||||
|
# sealed prior records (immutable) that carry the 4 non-RQ3 family p-values
|
||||||
|
LEAD_RESULTS = "output/sor-confirmatory/20260720T060132Z/analysis/stage06-results.json"
|
||||||
|
RQ2P3_RESULTS = "output/sor-rq2p3-confirmatory/rq2p3-confirmatory-results.json"
|
||||||
|
|
||||||
|
|
||||||
|
def _seed(tag: str) -> int:
|
||||||
|
"""Deterministic, auditable per-test resampling seed."""
|
||||||
|
return int.from_bytes(hashlib.sha256(f"sor-rq3-confirm|{tag}".encode()).digest()[:8], "big")
|
||||||
|
|
||||||
|
|
||||||
|
# --- generic run-level multi-arm bootstrap (mirrors stats.two_sample_diff_ci) --- #
|
||||||
|
def _multi_arm_bootstrap(
|
||||||
|
arms: Dict[str, Sequence[float]],
|
||||||
|
statistic: Callable[[Dict[str, List[float]]], float],
|
||||||
|
*,
|
||||||
|
null: float,
|
||||||
|
seed: int,
|
||||||
|
n_resamples: int = stats.DEFAULT_RESAMPLES,
|
||||||
|
alpha: float = stats.DEFAULT_ALPHA,
|
||||||
|
):
|
||||||
|
"""Bootstrap CI + Holm-ordering p for a statistic over >=1 independent arms.
|
||||||
|
|
||||||
|
Each arm is resampled with replacement independently (the frozen two-sample rule
|
||||||
|
generalized to N arms); BCa uses a combined leave-one-out jackknife (each point
|
||||||
|
dropped from its OWN arm, others full), exactly as ``stats.two_sample_diff_ci``.
|
||||||
|
The p is ``stats.two_sided_bootstrap_p`` against ``null`` — carried only to order
|
||||||
|
the Holm family, never as a stand-alone decision.
|
||||||
|
"""
|
||||||
|
import random
|
||||||
|
|
||||||
|
names = list(arms)
|
||||||
|
data = {k: list(v) for k, v in arms.items()}
|
||||||
|
theta_hat = float(statistic(data))
|
||||||
|
rng = random.Random(seed)
|
||||||
|
thetas: List[float] = []
|
||||||
|
for _ in range(n_resamples):
|
||||||
|
sample = {k: [v[rng.randrange(len(v))] for _ in range(len(v))] for k, v in data.items()}
|
||||||
|
thetas.append(float(statistic(sample)))
|
||||||
|
thetas.sort()
|
||||||
|
|
||||||
|
# combined jackknife: drop point i from arm `name`, keep the others full
|
||||||
|
can_bca = all(len(v) > 1 for v in data.values())
|
||||||
|
if can_bca:
|
||||||
|
jack: List[float] = []
|
||||||
|
for name in names:
|
||||||
|
n = len(data[name])
|
||||||
|
for i in range(n):
|
||||||
|
d = dict(data)
|
||||||
|
d[name] = [data[name][j] for j in range(n) if j != i]
|
||||||
|
jack.append(float(statistic(d)))
|
||||||
|
q_lo, q_hi, used = stats._bca_endpoints(thetas, theta_hat, jack, alpha)
|
||||||
|
else:
|
||||||
|
q_lo, q_hi, used = alpha / 2.0, 1.0 - alpha / 2.0, "percentile"
|
||||||
|
|
||||||
|
ci = stats.CIResult(theta_hat, stats._percentile(thetas, q_lo),
|
||||||
|
stats._percentile(thetas, q_hi), alpha, n_resamples, used, seed)
|
||||||
|
p = stats.two_sided_bootstrap_p(thetas, null)
|
||||||
|
return ci, p
|
||||||
|
|
||||||
|
|
||||||
|
# --- battery loading -------------------------------------------------------- #
|
||||||
|
def load_battery(path: Path) -> Dict[str, Dict[str, List[float]]]:
|
||||||
|
"""Load the sealed battery into per-arm {retention, latency, gap_signal}.
|
||||||
|
|
||||||
|
``gap_signal`` is the per-run MEAN inter-rebuild gap — the confirmatory grouping
|
||||||
|
unit (same unit the frozen calibration gate used); runs with no gap are omitted
|
||||||
|
(no fabricated interval). Retention/latency are the per-run DVs as measured.
|
||||||
|
"""
|
||||||
|
doc = json.loads(Path(path).read_text())
|
||||||
|
arms: Dict[str, Dict[str, List[float]]] = {}
|
||||||
|
for _cid, cell in doc["cells"].items():
|
||||||
|
arms[cell["strategy"]] = {
|
||||||
|
"retention": list(cell["throughput_retention"]),
|
||||||
|
"latency": list(cell["added_latency_ms"]),
|
||||||
|
"gap_signal": [],
|
||||||
|
}
|
||||||
|
gap: Dict[str, List[float]] = defaultdict(list)
|
||||||
|
for run in doc["runs"]:
|
||||||
|
strat = run["cell_id"].split("selector=")[1].split("/")[0]
|
||||||
|
gaps = run.get("rebuild_gaps") or []
|
||||||
|
if gaps:
|
||||||
|
gap[strat].append(statistics.fmean(gaps))
|
||||||
|
for strat, sig in gap.items():
|
||||||
|
arms[strat]["gap_signal"] = sig
|
||||||
|
return arms
|
||||||
|
|
||||||
|
|
||||||
|
# --- the three frozen RQ3 tests --------------------------------------------- #
|
||||||
|
def _perf_margin(a: Dict[str, List[float]]) -> float:
|
||||||
|
return stats.mean(a["agent"]) - max(stats.mean(a["static"]), stats.mean(a["random"]))
|
||||||
|
|
||||||
|
|
||||||
|
def _added_latency(a: Dict[str, List[float]]) -> float:
|
||||||
|
# added-latency over the MIN-latency (faster) baseline arm, run-brief §3.2
|
||||||
|
return statistics.median(a["agent"]) - min(statistics.median(a["static"]), statistics.median(a["random"]))
|
||||||
|
|
||||||
|
|
||||||
|
def _p2_auc(a: Dict[str, List[float]]) -> float:
|
||||||
|
# fingerprint question: agent's per-run rebuild-gap signal vs the pooled baseline
|
||||||
|
return rebuild_classifier_auc(a["agent"], a["baseline"])
|
||||||
|
|
||||||
|
|
||||||
|
def analyze_rq3(arms: Dict[str, Dict[str, List[float]]], *, n_resamples: int = stats.DEFAULT_RESAMPLES) -> Dict:
|
||||||
|
ret = {k: arms[k]["retention"] for k in ("agent", "static", "random")}
|
||||||
|
lat = {k: arms[k]["latency"] for k in ("agent", "static", "random")}
|
||||||
|
pool_baseline = arms["static"]["gap_signal"] + arms["random"]["gap_signal"]
|
||||||
|
p2_arms = {"agent": arms["agent"]["gap_signal"], "baseline": pool_baseline}
|
||||||
|
|
||||||
|
perf_ci, perf_p = _multi_arm_bootstrap(ret, _perf_margin, null=0.0,
|
||||||
|
seed=_seed("perf"), n_resamples=n_resamples)
|
||||||
|
lat_ci, lat_p = _multi_arm_bootstrap(lat, _added_latency, null=0.0,
|
||||||
|
seed=_seed("latency"), n_resamples=n_resamples)
|
||||||
|
p2_ci, p2_p = _multi_arm_bootstrap(p2_arms, _p2_auc, null=0.5,
|
||||||
|
seed=_seed("p2-auc"), n_resamples=n_resamples)
|
||||||
|
|
||||||
|
perf_hold = perf_ci.lo >= PERF_MARGIN_MIN_PP # gate: CI lower ≥ +10 pp
|
||||||
|
lat_hold = lat_ci.hi <= LATENCY_MAX_MS # gate: CI upper ≤ 100 ms
|
||||||
|
p2_hold = p2_ci.hi <= AUC_CEILING # gate: CI upper ≤ 0.60
|
||||||
|
p1_hold = perf_hold and lat_hold
|
||||||
|
p3_confirm = p1_hold and p2_hold
|
||||||
|
|
||||||
|
which_lat_baseline = "random" if statistics.median(lat["random"]) <= statistics.median(lat["static"]) else "static"
|
||||||
|
return {
|
||||||
|
"RQ3-P1-perf": {
|
||||||
|
"effect": "throughput_retention_margin_agent_minus_max_baseline",
|
||||||
|
**perf_ci.as_dict(), "p_for_holm": perf_p,
|
||||||
|
"gate": f"CI lower ≥ +{PERF_MARGIN_MIN_PP}", "holds": bool(perf_hold),
|
||||||
|
"decision": "perf-gain" if perf_hold else "no-perf-gain",
|
||||||
|
},
|
||||||
|
"RQ3-P1-latency": {
|
||||||
|
"effect": "added_latency_ms_agent_minus_min_baseline",
|
||||||
|
"min_latency_baseline_arm": which_lat_baseline,
|
||||||
|
**lat_ci.as_dict(), "p_for_holm": lat_p,
|
||||||
|
"gate": f"CI upper ≤ {LATENCY_MAX_MS} ms", "holds": bool(lat_hold),
|
||||||
|
"decision": "within-latency-budget" if lat_hold else "over-latency-budget",
|
||||||
|
},
|
||||||
|
"RQ3-P2": {
|
||||||
|
"effect": "rebuild_classifier_auc_agent_vs_pooled_baseline",
|
||||||
|
"grouping_unit": "per-run mean inter-rebuild gap",
|
||||||
|
"n_agent_runs": len(p2_arms["agent"]), "n_baseline_runs": len(p2_arms["baseline"]),
|
||||||
|
**p2_ci.as_dict(), "p_for_holm": p2_p,
|
||||||
|
"gate": f"CI upper ≤ {AUC_CEILING}", "holds": bool(p2_hold),
|
||||||
|
"decision": "no-usable-fingerprint" if p2_hold else "fingerprint-not-excluded",
|
||||||
|
},
|
||||||
|
"RQ3-P3-joint": {
|
||||||
|
"rule": "CONFIRM iff (P1-perf ∧ P1-latency) ∧ P2",
|
||||||
|
"p1_holds": bool(p1_hold), "p2_holds": bool(p2_hold),
|
||||||
|
"confirm": bool(p3_confirm),
|
||||||
|
"decision": "agent-helps-without-fingerprint" if p3_confirm else "H0",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# --- authoritative Holm-7 --------------------------------------------------- #
|
||||||
|
def _load_prior_pvalues(lead_path: Path, rq2p3_path: Path) -> Dict[str, Dict]:
|
||||||
|
lead = json.loads(Path(lead_path).read_text())["confirmatory"]
|
||||||
|
rq2p3 = json.loads(Path(rq2p3_path).read_text())["results"]
|
||||||
|
h1 = rq2p3["H1_pooled_spearman"]
|
||||||
|
return {
|
||||||
|
"RQ1-P1": {"p": lead["RQ1-P1"]["p_for_holm"], "source": "sealed lead RQ1-P1"},
|
||||||
|
"RQ1-P2": {"p": lead["RQ1-P2"]["p_for_holm"], "source": "sealed lead RQ1-P2"},
|
||||||
|
"RQ2-P1": {"p": lead["RQ2-P1"]["p_for_holm"], "source": "sealed lead RQ2-P1 (shrink)"},
|
||||||
|
"RQ2-P3": {"p": h1["p_for_holm"],
|
||||||
|
"source": "sealed RQ2-P3′ H1-pooled Spearman ρ (mechanism-corrected, mix) "
|
||||||
|
"— supersedes the lead's degenerate as-instrumented RQ2-P3"},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def authoritative_holm7(rq3: Dict, lead_path: Path, rq2p3_path: Path) -> Dict:
|
||||||
|
prior = _load_prior_pvalues(lead_path, rq2p3_path)
|
||||||
|
pvals = {
|
||||||
|
"RQ1-P1": prior["RQ1-P1"]["p"],
|
||||||
|
"RQ1-P2": prior["RQ1-P2"]["p"],
|
||||||
|
"RQ2-P1": prior["RQ2-P1"]["p"],
|
||||||
|
"RQ2-P3": prior["RQ2-P3"]["p"],
|
||||||
|
"RQ3-P1-perf": rq3["RQ3-P1-perf"]["p_for_holm"],
|
||||||
|
"RQ3-P1-latency": rq3["RQ3-P1-latency"]["p_for_holm"],
|
||||||
|
"RQ3-P2": rq3["RQ3-P2"]["p_for_holm"],
|
||||||
|
}
|
||||||
|
assert set(pvals) == set(FROZEN_FAMILY), "family must be exactly the frozen size-7"
|
||||||
|
holm = stats.holm_bonferroni(pvals, family_size=FAMILY_SIZE)
|
||||||
|
rows = [{"name": h.name, "raw_p": h.p, "holm_p": h.p_adjusted,
|
||||||
|
"multiplier": h.multiplier, "rank": h.rank, "reject": h.reject} for h in holm]
|
||||||
|
return {
|
||||||
|
"family_size": FAMILY_SIZE,
|
||||||
|
"family": list(FROZEN_FAMILY),
|
||||||
|
"p_sources": {k: prior[k]["source"] for k in prior},
|
||||||
|
"rows": rows,
|
||||||
|
"survivors": [r["name"] for r in rows if r["reject"]],
|
||||||
|
"non_survivors": [r["name"] for r in rows if not r["reject"]],
|
||||||
|
"supersedes": "lead paper's conservative partial embedding (report-4 of family-of-7); "
|
||||||
|
"both valid, partial never under-corrects; RQ1-P1 and RQ2-P1 survive regardless",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def run(*, battery_path: Path, lead_path: Path, rq2p3_path: Path,
|
||||||
|
n_resamples: int = stats.DEFAULT_RESAMPLES) -> Dict:
|
||||||
|
arms = load_battery(battery_path)
|
||||||
|
rq3 = analyze_rq3(arms, n_resamples=n_resamples)
|
||||||
|
holm7 = authoritative_holm7(rq3, lead_path, rq2p3_path)
|
||||||
|
battery_sha = hashlib.sha256(Path(battery_path).read_bytes()).hexdigest()
|
||||||
|
return {
|
||||||
|
"schema": "sor-rq3-confirmatory/1",
|
||||||
|
"measured_from": "live-docker-e2e",
|
||||||
|
"battery_results_path": str(battery_path),
|
||||||
|
"battery_results_sha256": battery_sha,
|
||||||
|
"n_runs_per_arm": {k: len(arms[k]["retention"]) for k in arms},
|
||||||
|
"frozen_lead_prereg_sha256": "f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b",
|
||||||
|
"results": rq3,
|
||||||
|
"authoritative_holm7": holm7,
|
||||||
|
"honest_disclosure": (
|
||||||
|
"Every RQ3 test reports effect + BCa 95% CI; p is carried ONLY to order the Holm "
|
||||||
|
"family (prereg §6). Nulls are results: a selector that does not beat baselines, or "
|
||||||
|
"a rebuild pattern that is not certifiably non-classifiable, is the finding — not spun. "
|
||||||
|
"Reproducibility caveat (accepted, run-brief §2A): the agent (qwen2.5:3b via local "
|
||||||
|
"Ollama, temp 0) is reproducible via its committed decision-log + (seed, state-hash) "
|
||||||
|
"cache replay, NOT via independent model re-execution on other hardware."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _seal(out_dir: Path, report: Dict) -> Path:
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
path = out_dir / "rq3-confirmatory-analysis.json"
|
||||||
|
path.write_text(json.dumps(report, indent=2, sort_keys=True) + "\n", encoding="utf-8")
|
||||||
|
digest = hashlib.sha256(path.read_bytes()).hexdigest()
|
||||||
|
(out_dir / "SHA256SUMS").write_text(f"{digest} {path.name}\n", encoding="utf-8")
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv=None) -> int:
|
||||||
|
ap = argparse.ArgumentParser(prog="python -m cmd_chat.sor.analysis.rq3_confirm")
|
||||||
|
ap.add_argument("--battery", required=True)
|
||||||
|
ap.add_argument("--lead", default=LEAD_RESULTS)
|
||||||
|
ap.add_argument("--rq2p3", default=RQ2P3_RESULTS)
|
||||||
|
ap.add_argument("--out", default=None)
|
||||||
|
ap.add_argument("--n-resamples", type=int, default=stats.DEFAULT_RESAMPLES)
|
||||||
|
args = ap.parse_args(argv)
|
||||||
|
report = run(battery_path=Path(args.battery), lead_path=Path(args.lead),
|
||||||
|
rq2p3_path=Path(args.rq2p3), n_resamples=args.n_resamples)
|
||||||
|
if args.out:
|
||||||
|
path = _seal(Path(args.out), report)
|
||||||
|
print(f"sealed -> {path}", file=sys.stderr)
|
||||||
|
print(json.dumps(report, indent=2, sort_keys=True))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
e09c66efe00524ae3e5373bd20121e6e022e6885596f4fa9aabd51653c76929b rq3-confirmatory-analysis.json
|
||||||
@@ -0,0 +1,156 @@
|
|||||||
|
{
|
||||||
|
"authoritative_holm7": {
|
||||||
|
"family": [
|
||||||
|
"RQ1-P1",
|
||||||
|
"RQ1-P2",
|
||||||
|
"RQ2-P1",
|
||||||
|
"RQ2-P3",
|
||||||
|
"RQ3-P1-perf",
|
||||||
|
"RQ3-P1-latency",
|
||||||
|
"RQ3-P2"
|
||||||
|
],
|
||||||
|
"family_size": 7,
|
||||||
|
"non_survivors": [
|
||||||
|
"RQ1-P2",
|
||||||
|
"RQ3-P2",
|
||||||
|
"RQ3-P1-perf",
|
||||||
|
"RQ3-P1-latency"
|
||||||
|
],
|
||||||
|
"p_sources": {
|
||||||
|
"RQ1-P1": "sealed lead RQ1-P1",
|
||||||
|
"RQ1-P2": "sealed lead RQ1-P2",
|
||||||
|
"RQ2-P1": "sealed lead RQ2-P1 (shrink)",
|
||||||
|
"RQ2-P3": "sealed RQ2-P3\u2032 H1-pooled Spearman \u03c1 (mechanism-corrected, mix) \u2014 supersedes the lead's degenerate as-instrumented RQ2-P3"
|
||||||
|
},
|
||||||
|
"rows": [
|
||||||
|
{
|
||||||
|
"holm_p": 0.0,
|
||||||
|
"multiplier": 7,
|
||||||
|
"name": "RQ1-P1",
|
||||||
|
"rank": 1,
|
||||||
|
"raw_p": 0.0,
|
||||||
|
"reject": true
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.0,
|
||||||
|
"multiplier": 6,
|
||||||
|
"name": "RQ2-P1",
|
||||||
|
"rank": 2,
|
||||||
|
"raw_p": 0.0,
|
||||||
|
"reject": true
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.0,
|
||||||
|
"multiplier": 5,
|
||||||
|
"name": "RQ2-P3",
|
||||||
|
"rank": 3,
|
||||||
|
"raw_p": 0.0,
|
||||||
|
"reject": true
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.3648,
|
||||||
|
"multiplier": 4,
|
||||||
|
"name": "RQ1-P2",
|
||||||
|
"rank": 4,
|
||||||
|
"raw_p": 0.0912,
|
||||||
|
"reject": false
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.5112,
|
||||||
|
"multiplier": 3,
|
||||||
|
"name": "RQ3-P2",
|
||||||
|
"rank": 5,
|
||||||
|
"raw_p": 0.1704,
|
||||||
|
"reject": false
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.5112,
|
||||||
|
"multiplier": 2,
|
||||||
|
"name": "RQ3-P1-perf",
|
||||||
|
"rank": 6,
|
||||||
|
"raw_p": 0.177,
|
||||||
|
"reject": false
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"holm_p": 0.5112,
|
||||||
|
"multiplier": 1,
|
||||||
|
"name": "RQ3-P1-latency",
|
||||||
|
"rank": 7,
|
||||||
|
"raw_p": 0.497,
|
||||||
|
"reject": false
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"supersedes": "lead paper's conservative partial embedding (report-4 of family-of-7); both valid, partial never under-corrects; RQ1-P1 and RQ2-P1 survive regardless",
|
||||||
|
"survivors": [
|
||||||
|
"RQ1-P1",
|
||||||
|
"RQ2-P1",
|
||||||
|
"RQ2-P3"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"battery_results_path": "output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/rq3-battery-results.json",
|
||||||
|
"battery_results_sha256": "5b61e461004722985c1a7a9bc5fdfe855395df3180c71e6efdc3531b8ecf8039",
|
||||||
|
"frozen_lead_prereg_sha256": "f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b",
|
||||||
|
"honest_disclosure": "Every RQ3 test reports effect + BCa 95% CI; p is carried ONLY to order the Holm family (prereg \u00a76). Nulls are results: a selector that does not beat baselines, or a rebuild pattern that is not certifiably non-classifiable, is the finding \u2014 not spun. Reproducibility caveat (accepted, run-brief \u00a72A): the agent (qwen2.5:3b via local Ollama, temp 0) is reproducible via its committed decision-log + (seed, state-hash) cache replay, NOT via independent model re-execution on other hardware.",
|
||||||
|
"measured_from": "live-docker-e2e",
|
||||||
|
"n_runs_per_arm": {
|
||||||
|
"agent": 30,
|
||||||
|
"random": 30,
|
||||||
|
"static": 30
|
||||||
|
},
|
||||||
|
"results": {
|
||||||
|
"RQ3-P1-latency": {
|
||||||
|
"alpha": 0.05,
|
||||||
|
"ci_hi": 34.911185735836625,
|
||||||
|
"ci_lo": -52.14837333187461,
|
||||||
|
"decision": "within-latency-budget",
|
||||||
|
"effect": "added_latency_ms_agent_minus_min_baseline",
|
||||||
|
"gate": "CI upper \u2264 100.0 ms",
|
||||||
|
"holds": true,
|
||||||
|
"method": "bca",
|
||||||
|
"min_latency_baseline_arm": "random",
|
||||||
|
"n_resamples": 10000,
|
||||||
|
"p_for_holm": 0.497,
|
||||||
|
"point": -13.491572928614914,
|
||||||
|
"seed": 5867095940896968561
|
||||||
|
},
|
||||||
|
"RQ3-P1-perf": {
|
||||||
|
"alpha": 0.05,
|
||||||
|
"ci_hi": 0.003865947497819883,
|
||||||
|
"ci_lo": -0.015830325279388435,
|
||||||
|
"decision": "no-perf-gain",
|
||||||
|
"effect": "throughput_retention_margin_agent_minus_max_baseline",
|
||||||
|
"gate": "CI lower \u2265 +0.1",
|
||||||
|
"holds": false,
|
||||||
|
"method": "bca",
|
||||||
|
"n_resamples": 10000,
|
||||||
|
"p_for_holm": 0.177,
|
||||||
|
"point": -0.006312034914976117,
|
||||||
|
"seed": 16001687329924348899
|
||||||
|
},
|
||||||
|
"RQ3-P2": {
|
||||||
|
"alpha": 0.05,
|
||||||
|
"ci_hi": 0.7034362881655691,
|
||||||
|
"ci_lo": 0.45805555555555555,
|
||||||
|
"decision": "fingerprint-not-excluded",
|
||||||
|
"effect": "rebuild_classifier_auc_agent_vs_pooled_baseline",
|
||||||
|
"gate": "CI upper \u2264 0.6",
|
||||||
|
"grouping_unit": "per-run mean inter-rebuild gap",
|
||||||
|
"holds": false,
|
||||||
|
"method": "bca",
|
||||||
|
"n_agent_runs": 30,
|
||||||
|
"n_baseline_runs": 60,
|
||||||
|
"n_resamples": 10000,
|
||||||
|
"p_for_holm": 0.1704,
|
||||||
|
"point": 0.5869444444444445,
|
||||||
|
"seed": 9396385129919407505
|
||||||
|
},
|
||||||
|
"RQ3-P3-joint": {
|
||||||
|
"confirm": false,
|
||||||
|
"decision": "H0",
|
||||||
|
"p1_holds": false,
|
||||||
|
"p2_holds": false,
|
||||||
|
"rule": "CONFIRM iff (P1-perf \u2227 P1-latency) \u2227 P2"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"schema": "sor-rq3-confirmatory/1"
|
||||||
|
}
|
||||||
@@ -0,0 +1,91 @@
|
|||||||
|
5b61e461004722985c1a7a9bc5fdfe855395df3180c71e6efdc3531b8ecf8039 rq3-battery-results.json
|
||||||
|
110bc93cd92f43f679f108d34e16ca97af1a45076167597509caf431f539eb0d rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r0/rq3-run.json
|
||||||
|
80e238396a0c4359da81ac4a32d8da645c7bc045908d3c9e0ab6fbb47fd31c05 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r1/rq3-run.json
|
||||||
|
c1c41673ab235dde0c0e0ce968c0ca8b6876e37425486e5789ea8da2a9e2ae56 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r10/rq3-run.json
|
||||||
|
3e3609b9e78cb26bb5202e9f8f77acd4469a35ea2fc925a25c8f3c449c0598ee rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r11/rq3-run.json
|
||||||
|
add02b8c1fe94a27e7cfcfdd2017709a572fc635859ea7e9fa6a9aeb7752b179 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r12/rq3-run.json
|
||||||
|
a5fc0270f9e87a25ffbf3fbf41dc741fcb1aa981b7312249821c93d666786d19 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r13/rq3-run.json
|
||||||
|
72600ffe7378da5b04c8b7cc16c3618d11b1cce6ab3e166a5d58a76379bd3602 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r14/rq3-run.json
|
||||||
|
259543832d3d394379849b1c2736962119b290f8b3cad7aa05562bf5f18390cd rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r15/rq3-run.json
|
||||||
|
d78e5306cae58b7bc46f80244752090cac60102bd037d68d76fe2740b53ebd1e rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r16/rq3-run.json
|
||||||
|
9bffbb317df459bb7847dfc58c5462066ace4a4befad3ff5919f0986621673eb rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r17/rq3-run.json
|
||||||
|
d91c00e81cec24fa922a0fdcedc5c5745293d4522b4b2a5b4f3c7df6fe17a179 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r18/rq3-run.json
|
||||||
|
f293fe699c1aab60c9650eac031aedc4a605bb62c00cdb92794e2b768e2e8a50 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r19/rq3-run.json
|
||||||
|
ec661f5c52133cf814428bd954cea3d1f6118c35c36cb01640573530e1c35f26 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r2/rq3-run.json
|
||||||
|
ee1fa456439bdc1c8a8264252b11446a212fda555ac26688c885dbcf5d0d4b65 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r20/rq3-run.json
|
||||||
|
bb710e7fd11a50bc262665e857d591645520b47dc4ba9afbd15b770c7b36fa85 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r21/rq3-run.json
|
||||||
|
7c92a5ec490f8c40abd7d4d5d05effb00d5c106567426016a17a52cbbde32a28 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r22/rq3-run.json
|
||||||
|
cc57bded029cd9f7bb9bc61beef7a2eaecdad798a9872fa570986e10291b92a7 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r23/rq3-run.json
|
||||||
|
6d8318fa57e16597d500e6b29661843db532f09cb303e409051611befb8d03e5 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r24/rq3-run.json
|
||||||
|
95e2a0df7e320ad3aeec056eb8c6fe872789311b3fc2d85f3b0730fc1984c6e1 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r25/rq3-run.json
|
||||||
|
3a036309872aa31fcf150074097226da5bac6485f10786e25313825b983a3fc6 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r26/rq3-run.json
|
||||||
|
ad59b9aad30cd10be94c1bf0bd58716a3272fd2e75e1da58aaed5ecaca64bfd0 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r27/rq3-run.json
|
||||||
|
0b0c00ddc3db49aa0c85a1b3dc29292435cb53c3753446edd351205115d38f54 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r28/rq3-run.json
|
||||||
|
99e28987eaa6da592df7eddcc1134a0ee09bd833f79af54ebe3b6f2609cf256f rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r29/rq3-run.json
|
||||||
|
86cb773d2315ac118790289779337514e80006817c0c1ec07e0f8d99adaabf6c rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r3/rq3-run.json
|
||||||
|
08ad295c823da04a6ccbdd758d638c4aae4c7220e9cf4361c81672d800e4bef8 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r4/rq3-run.json
|
||||||
|
c7daf7fd45aad4c66570875a41aed8d7c69d0d1246f3291d7983b894d547a6e1 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r5/rq3-run.json
|
||||||
|
ae47e7fce7b8acbd36567e31ae1fa5879d5fa483c6d293e0ae588ebf46d925e7 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r6/rq3-run.json
|
||||||
|
ce234e3ed894d1c7ba2f6b650c50aec42a3bd423e894f422a64ff1b3a57b3376 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r7/rq3-run.json
|
||||||
|
723d66a199e09c13de8735b876db87145e5a3950aae9d749ef6abb91fa6f2af9 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r8/rq3-run.json
|
||||||
|
f30da953eb66d498d4e2be2607e8a1d47a4ac3f626f1cc891852c8ef07672919 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r9/rq3-run.json
|
||||||
|
f276456e1cfc914fcdb963f5d85e72a33d2858369547d4d11e8da739a2c315e8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r0/rq3-run.json
|
||||||
|
ebd567bb9b129d75d7e794b4f06fa1340263fe7c4abf2cc68546ab476e8e91b9 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r1/rq3-run.json
|
||||||
|
761a36d6ec339b84d465ffc9dcd88103e6b0130fb9b1ba4aebeb534a8bf581a4 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r10/rq3-run.json
|
||||||
|
2c663909db5624832bae8393648a1c50a4ac00ab580f4ae7a5f8bdbdce0ef3e2 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r11/rq3-run.json
|
||||||
|
5a663ef49d5860187fd2a7fc0453ae2f3e3e312ff1224cb75db57c3eb498711a rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r12/rq3-run.json
|
||||||
|
2af0bd978aa2b92842b5308ec262f508d3f244b9640cc7f16c0dd5f16c05f843 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r13/rq3-run.json
|
||||||
|
77331f8a9041755c77155da29b8176b4c27b674d741f23ee4bf213f84e008949 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r14/rq3-run.json
|
||||||
|
2f4b249f9afba300094bb2ce3b50265abe09350ce8959ec2b471ece8ffdfa6ae rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r15/rq3-run.json
|
||||||
|
0db6e249bf8b6ac0d6d255437a9f868d87641f66fc7fb93f7ca809674884bdb7 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r16/rq3-run.json
|
||||||
|
76aac9803ac18a5f6f8f0016ba3b7970fcff475a8475fc959834a4a88f1a0442 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r17/rq3-run.json
|
||||||
|
2a25407358b9ab36cb4997af9f232dcb1f0247015268278ec63cada29cfe6a39 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r18/rq3-run.json
|
||||||
|
e0f404e78b35404f789120a8b898e03a6ca1bad292bf1ac4efc68bd2004188f0 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r19/rq3-run.json
|
||||||
|
2515a88a3116a0d490b9d5484517479ed18af45436b33ffa31ee1293ee7fb6e8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r2/rq3-run.json
|
||||||
|
73593f9667e94ed16fdfa6b943d4b2cc7174cd62ca2cf6ec25f5bc4354d3ec62 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r20/rq3-run.json
|
||||||
|
9d89871c47cc12ac826dbeecab7d3943bff252248ecbbc5e88e444ec83f1fda1 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r21/rq3-run.json
|
||||||
|
d6753dc2097d22c67cc31fdfca225b336e8439a4975b4f1428879796c1d0410e rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r22/rq3-run.json
|
||||||
|
befe39c32b87e73169ec6c111ca79f15fbb24d4952b5dd817ce766399e10824e rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r23/rq3-run.json
|
||||||
|
cad75228e7ef404fbc54561d3f5b4969f441b047abaa0663ae1d51322f43acaa rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r24/rq3-run.json
|
||||||
|
2817c84be3af073fa471e0287a845d8147d01a1d0f8e800faa281ce02885afe4 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r25/rq3-run.json
|
||||||
|
3e7d841db693adb2fe61218588da3e269d31e337242241d0ac87115f516d221c rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r26/rq3-run.json
|
||||||
|
6a5b5d07bcdf363f4fb959135de6cc8b23c48e5a76888abd0fd2e67e80e5673f rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r27/rq3-run.json
|
||||||
|
44d512c8eeef83c410399303d482ef760fa490630b26a4cc5ca7294e7f324a14 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r28/rq3-run.json
|
||||||
|
6ea17d94630ccc117637b3a2ea0ddc8926f0085ad7cd5cea2aeba8de66997150 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r29/rq3-run.json
|
||||||
|
1d07bbd03e31d2717103d5d766d20c43730e3d65d1b95e551531f198744544cb rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r3/rq3-run.json
|
||||||
|
a22cd7af4000eea7529ba1e7dd257184f702d77fed44965a9a91c7944b65c2c9 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r4/rq3-run.json
|
||||||
|
7169dd2d4cf99c53537ce9fd5608f7dda118f645d01cbee5124721f3d03654cc rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r5/rq3-run.json
|
||||||
|
3c6352374e680178deb5f52f20839202cebf8536fe3b1afb7e0a8e475ae36c27 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r6/rq3-run.json
|
||||||
|
4cb9c0fcf183a53bb7dc428c12ba3482adae9d7554238e49520b4848b7599fb8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r7/rq3-run.json
|
||||||
|
763e75f8b465cf8b2b41c644b47f9e2c71a85995c569c2700e929feecd087c96 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r8/rq3-run.json
|
||||||
|
3bdc4dcf22de5f09f033b03bd16c01cb9c1cd4fc210ce9b57c8a0c65ec663e80 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r9/rq3-run.json
|
||||||
|
257326c8a34c602e028fca9593e8fe628a950eff63138f6eeeb425959112ed5a rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r0/rq3-run.json
|
||||||
|
1694018a795e64d07b5ff83f5ec512a0c6e38a4f210b3ea1d54b41281d54e21f rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r1/rq3-run.json
|
||||||
|
bdfb55a3b606ef441237ed0828d504dd885d6730b0de01b0d3cc50f8a5c05826 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r10/rq3-run.json
|
||||||
|
3a13e1de470c79965972a275d72e4c8bf8d1195e3a55a896520904de4d6c2e68 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r11/rq3-run.json
|
||||||
|
4449a0a53f7033cb7befd8520501b0f00b9201738e517c023576d7b861720935 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r12/rq3-run.json
|
||||||
|
45105408efac8d96b449dd3e1945f5727f9f0189a85d3e6bfa5a3c2db0b43882 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r13/rq3-run.json
|
||||||
|
07a9e31fec1807d5200978cbad6267966acd74d3eec640868e87bb778d4d4722 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r14/rq3-run.json
|
||||||
|
76830fb704f7439be2443fb2c051aeaeaaafddf32e35f209549bf74aab4ffe82 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r15/rq3-run.json
|
||||||
|
a4d9a9eb631a3c2c459252f79d0ac65c743440cef445c2545f7e08141130b34e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r16/rq3-run.json
|
||||||
|
7357f8e01e000bea1c27a0ccaed7d41f01a5e1f6f75569561a6e1901dc696442 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r17/rq3-run.json
|
||||||
|
4ee979e846426d382f6aea04e205cfe33a3478add09273956decf5aa6db28c37 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r18/rq3-run.json
|
||||||
|
19e6c0d747dd802e324ca6225a56e4a849e7db42f09df3b61c703aad4adb7082 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r19/rq3-run.json
|
||||||
|
cd2638c191fe0a73797f3e2eefbcfba8af9bca0759e6473c420e8af523d2c99e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r2/rq3-run.json
|
||||||
|
9f3730ef5c8feeed3c2578e8a30e471eeabc5c98d2a57514ad053d6b9328cb36 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r20/rq3-run.json
|
||||||
|
f947eec0f931023d0e6848809ee3484df191099236c50db46a32d3fcaecf88d8 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r21/rq3-run.json
|
||||||
|
9d5bc5dff8954fb6641ae9d2f9f83d659339867ac7395278c88d0e190991cd92 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r22/rq3-run.json
|
||||||
|
1cbeaf4904aff435ac72d8e002bd909ceeceeda2c176794f959b8c7264b021c0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r23/rq3-run.json
|
||||||
|
efaaad8ffd17cfe3c55e62d09f5b1dfd670c2474ea6219931df8e08eef2d83a0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r24/rq3-run.json
|
||||||
|
03c89c522e5e2cce9181c20f62edd434a734d8dc530629b36bfb6b5e58173c50 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r25/rq3-run.json
|
||||||
|
5c171990a45aca7adede446a2bb8b5cec5b72d454d5f4ae061b0883ec80eb5ab rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r26/rq3-run.json
|
||||||
|
4762af92e07ba3f8c7e6f0ffecc25af15469cbb3d15e151a9acb12619cf37e7e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r27/rq3-run.json
|
||||||
|
ec29c89ef5e42ec618ed848c23f9fa35709ce663053929214f05226a055ae2d1 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r28/rq3-run.json
|
||||||
|
68f3441aff941c9aabba81321e9352b966bf0883a1afb6b05172b4d5ea3c32d9 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r29/rq3-run.json
|
||||||
|
d8d6aaaa09b62f293750d8f46422ae27e504fd8863591f3f1c34f1a2135c6b0c rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r3/rq3-run.json
|
||||||
|
20e794c0082d6ec4ebdee3f8e659b5f6a19280a0706ef9910dc5af62329bbe32 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r4/rq3-run.json
|
||||||
|
5346c57b2b4822aa36906b4348d9be1c4c0c6504cd91676c49af5fa4c637769c rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r5/rq3-run.json
|
||||||
|
375089aa2ffa966668ad9063acb37e21f3ad176996bdd8f5a66f20d2bca7da53 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r6/rq3-run.json
|
||||||
|
8e5f07eb6b52125ccb13a58d4fd2316075a416670fe2391627c0954a42bf36f0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r7/rq3-run.json
|
||||||
|
e334641cd823338129c050a35a61d1d47a50976a6110301a42294a2e80b47c4a rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r8/rq3-run.json
|
||||||
|
e2c48513c108020c7e29a3107c7f8c5bfdb8bb8374120fc1827177d81fe772a1 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r9/rq3-run.json
|
||||||
+3415
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,128 @@
|
|||||||
|
"""RQ3 confirmatory-harness pins — structure, determinism, gate logic, Holm-7 wiring.
|
||||||
|
|
||||||
|
These exercise the RQ3 confirmatory analyzer's mechanics on a SMALL SYNTHETIC battery
|
||||||
|
(so the test owns its inputs and does not depend on the sealed live record): (a) the
|
||||||
|
battery loader recovers per-arm retention/latency and the per-run mean-gap signal; (b) the
|
||||||
|
three frozen gates read the right CI bound (perf: lower ≥ +10pp; latency: upper ≤ 100ms;
|
||||||
|
P2: upper ≤ 0.60) and RQ3-P3 is their logical AND; (c) the run-level multi-arm bootstrap is
|
||||||
|
byte-deterministic (same seed → identical CI) and BCa-well-formed; (d) the authoritative
|
||||||
|
Holm is over EXACTLY the frozen size-7 family with the RQ2-P3 slot carrying the
|
||||||
|
mechanism-corrected primary p, and the step-down/survivor logic matches stats.holm_bonferroni.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
|
||||||
|
from cmd_chat.sor.analysis import rq3_confirm as rc
|
||||||
|
from cmd_chat.sor.analysis import stats
|
||||||
|
|
||||||
|
|
||||||
|
def _synthetic_battery(tmp_path):
|
||||||
|
"""A 3-arm battery where agent clearly beats baselines on retention and is faster,
|
||||||
|
with distinct rebuild-gap signals — so the gate directions are unambiguous."""
|
||||||
|
def cell(strat, ret, lat):
|
||||||
|
return {"strategy": strat, "runs": len(ret),
|
||||||
|
"throughput_retention": ret, "added_latency_ms": lat}
|
||||||
|
cells = {
|
||||||
|
"a": cell("agent", [1.0] * 8, [100.0] * 8),
|
||||||
|
"s": cell("static", [0.5] * 8, [400.0] * 8),
|
||||||
|
"r": cell("random", [0.6] * 8, [300.0] * 8),
|
||||||
|
}
|
||||||
|
runs = []
|
||||||
|
# frozen classifier scores SHORTER gaps (more churn) as positive; give the agent
|
||||||
|
# shorter gaps than the baselines so it separates upward (AUC -> 1), the fingerprint
|
||||||
|
# direction the ≤0.60 ceiling is meant to catch.
|
||||||
|
for strat, gap in (("agent", 1.0), ("static", 3.0), ("random", 3.0)):
|
||||||
|
for i in range(8):
|
||||||
|
runs.append({"cell_id": f"RQ3/topo=1house/bridge=off/selector={strat}/churn=kp30s20",
|
||||||
|
"rebuild_gaps": [gap, gap, gap], "run_index": i})
|
||||||
|
doc = {"schema": "sor-rq3-battery-results/1", "cells": cells, "runs": runs}
|
||||||
|
p = tmp_path / "battery.json"
|
||||||
|
p.write_text(json.dumps(doc))
|
||||||
|
return p
|
||||||
|
|
||||||
|
|
||||||
|
# --- loader ----------------------------------------------------------------- #
|
||||||
|
def test_loader_recovers_per_arm_dvs_and_gap_signal(tmp_path):
|
||||||
|
arms = rc.load_battery(_synthetic_battery(tmp_path))
|
||||||
|
assert set(arms) == {"agent", "static", "random"}
|
||||||
|
assert arms["agent"]["retention"] == [1.0] * 8
|
||||||
|
assert arms["static"]["latency"] == [400.0] * 8
|
||||||
|
# per-run MEAN gap: three equal gaps -> that value per run, 8 runs
|
||||||
|
assert arms["agent"]["gap_signal"] == [1.0] * 8
|
||||||
|
assert arms["random"]["gap_signal"] == [3.0] * 8
|
||||||
|
|
||||||
|
|
||||||
|
# --- gate directions + RQ3-P3 AND ------------------------------------------- #
|
||||||
|
def test_gates_read_correct_ci_bounds_and_join(tmp_path):
|
||||||
|
arms = rc.load_battery(_synthetic_battery(tmp_path))
|
||||||
|
res = rc.analyze_rq3(arms, n_resamples=500)
|
||||||
|
perf, lat, p2, joint = (res["RQ3-P1-perf"], res["RQ3-P1-latency"],
|
||||||
|
res["RQ3-P2"], res["RQ3-P3-joint"])
|
||||||
|
# agent retention 1.0 vs max(0.5,0.6)=0.6 -> margin +0.4, CI lower well above +0.1
|
||||||
|
assert perf["point"] > 0.1 and perf["holds"] is True and perf["decision"] == "perf-gain"
|
||||||
|
# agent latency 100 vs min(400,300)=300 -> -200ms, upper bound << 100ms
|
||||||
|
assert lat["point"] < 0 and lat["holds"] is True
|
||||||
|
assert lat["min_latency_baseline_arm"] == "random"
|
||||||
|
# agent gap 1.0 vs baseline 3.0 -> fully separable upward AUC=1.0 -> upper > 0.60 -> fails
|
||||||
|
assert p2["point"] == 1.0 and p2["holds"] is False
|
||||||
|
# P3 = (perf ∧ latency) ∧ P2 ; P2 fails here -> H0
|
||||||
|
assert joint["p1_holds"] is True and joint["p2_holds"] is False
|
||||||
|
assert joint["confirm"] is False and joint["decision"] == "H0"
|
||||||
|
|
||||||
|
|
||||||
|
# --- determinism + BCa well-formed ------------------------------------------ #
|
||||||
|
def test_analysis_is_byte_deterministic(tmp_path):
|
||||||
|
b = _synthetic_battery(tmp_path)
|
||||||
|
a1 = rc.analyze_rq3(rc.load_battery(b), n_resamples=800)
|
||||||
|
a2 = rc.analyze_rq3(rc.load_battery(b), n_resamples=800)
|
||||||
|
assert a1 == a2
|
||||||
|
for t in ("RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"):
|
||||||
|
assert a1[t]["ci_lo"] <= a1[t]["point"] <= a1[t]["ci_hi"]
|
||||||
|
assert 0.0 <= a1[t]["p_for_holm"] <= 1.0
|
||||||
|
|
||||||
|
|
||||||
|
# --- multi-arm bootstrap matches the frozen two-sample rule ------------------ #
|
||||||
|
def test_multi_arm_bootstrap_matches_two_sample_diff_for_two_means(tmp_path):
|
||||||
|
a = [1.0, 2.0, 3.0, 4.0, 5.0]
|
||||||
|
b = [0.5, 1.5, 2.5, 3.5]
|
||||||
|
ci_gen, _ = rc._multi_arm_bootstrap(
|
||||||
|
{"x": a, "y": b}, lambda d: stats.mean(d["x"]) - stats.mean(d["y"]),
|
||||||
|
null=0.0, seed=123, n_resamples=1000)
|
||||||
|
ci_ref = stats.two_sample_diff_ci(a, b, stats.mean, seed=123, n_resamples=1000)
|
||||||
|
# same seed + same independent-arm resampling rule -> identical point + interval
|
||||||
|
assert abs(ci_gen.point - ci_ref.point) < 1e-12
|
||||||
|
assert abs(ci_gen.lo - ci_ref.lo) < 1e-9 and abs(ci_gen.hi - ci_ref.hi) < 1e-9
|
||||||
|
|
||||||
|
|
||||||
|
# --- authoritative Holm-7 --------------------------------------------------- #
|
||||||
|
def test_holm7_family_is_exactly_the_frozen_seven(tmp_path):
|
||||||
|
assert rc.FAMILY_SIZE == 7
|
||||||
|
assert set(rc.FROZEN_FAMILY) == {
|
||||||
|
"RQ1-P1", "RQ1-P2", "RQ2-P1", "RQ2-P3",
|
||||||
|
"RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_holm7_survivor_logic_with_three_zeros(tmp_path):
|
||||||
|
# three p=0 priors + non-significant RQ3 tests: only the three zeros survive Holm-7.
|
||||||
|
rq3 = {
|
||||||
|
"RQ3-P1-perf": {"p_for_holm": 0.18},
|
||||||
|
"RQ3-P1-latency": {"p_for_holm": 0.50},
|
||||||
|
"RQ3-P2": {"p_for_holm": 0.17},
|
||||||
|
}
|
||||||
|
# stub the two sealed prior records
|
||||||
|
lead = tmp_path / "lead.json"
|
||||||
|
lead.write_text(json.dumps({"confirmatory": {
|
||||||
|
"RQ1-P1": {"p_for_holm": 0.0}, "RQ1-P2": {"p_for_holm": 0.0912},
|
||||||
|
"RQ2-P1": {"p_for_holm": 0.0}}}))
|
||||||
|
rq2p3 = tmp_path / "rq2p3.json"
|
||||||
|
rq2p3.write_text(json.dumps({"results": {"H1_pooled_spearman": {"p_for_holm": 0.0}}}))
|
||||||
|
|
||||||
|
holm = rc.authoritative_holm7(rq3, lead, rq2p3)
|
||||||
|
assert holm["family_size"] == 7
|
||||||
|
assert holm["survivors"] == ["RQ1-P1", "RQ2-P1", "RQ2-P3"]
|
||||||
|
assert set(holm["non_survivors"]) == {"RQ1-P2", "RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"}
|
||||||
|
# RQ2-P3 slot must carry the mechanism-corrected primary p, not the lead degenerate 1.0
|
||||||
|
assert "mechanism-corrected" in holm["p_sources"]["RQ2-P3"]
|
||||||
|
# RQ1-P2 at rank 4 gets multiplier 4 -> 0.0912*4 = 0.3648, not rejected
|
||||||
|
row = {r["name"]: r for r in holm["rows"]}
|
||||||
|
assert row["RQ1-P2"]["multiplier"] == 4 and abs(row["RQ1-P2"]["holm_p"] - 0.3648) < 1e-9
|
||||||
Reference in New Issue
Block a user