analysis: seal RQ3 confirmatory battery + frozen analysis + authoritative Holm-7

RQ3 live-docker battery (90/90 runs) complete; un-blind and apply the frozen
prereg §6 plan. No re-specification.

Seal: SHA256SUMS over confirmatory-data/ (battery-results.json 5b61e461... + 90
rq3-run.json sidecars) — raw data immutable. New analyzer analysis/rq3_confirm.py
drives the frozen stats/metrics (untouched) via a run-level multi-arm bootstrap
mirroring two_sample_diff_ci (10k BCa, alpha=0.05). Effect+CI always, never bare p.

RQ3 result (honest null):
- RQ3-P1-perf: retention margin agent-max(static,random) = -0.6pp, CI
  [-1.58,+0.39]pp -> FAILS +10pp gate (all selectors heal ~all churn, ~99%).
- RQ3-P1-latency: added-latency(agent-min-baseline) = -13.5ms, CI [-52.1,+34.9],
  upper <= 100ms -> within budget (agent not slower).
- RQ3-P2: rebuild-classifier AUC(agent vs pooled baseline) = 0.587, CI
  [0.458,0.703], upper 0.703 > 0.60 -> fingerprint NOT excluded (underpowered).
- RQ3-P3 = H0 (P1-perf fails and P2 fails).

Authoritative Holm-7 over the frozen size-7 family (supersedes the lead's
conservative partial embedding; RQ2-P3 slot carries the mechanism-corrected
H1-pooled Spearman p=0, not the lead degenerate p=1). Survivors: RQ1-P1,
RQ2-P1 (shrink), RQ2-P3 (mix). Non-survivors: RQ1-P2, RQ3-P2, RQ3-P1-perf,
RQ3-P1-latency. Sealed: analysis/rq3-confirmatory-analysis.json (e09c66ef...).

Tests: tests/test_sor_rq3_confirm.py 6 passed; full SOR suite 207 passed.
Both prereg SHAs intact; $0/offline analysis; worktree only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
leetcrypt
2026-07-22 08:48:31 -07:00
parent e0d865d012
commit 47faee3a92
7 changed files with 4104 additions and 0 deletions
+3
View File
@@ -106,3 +106,6 @@ RQ2 RATIFIED (operator Andre, 2026-07-20, while blind) — observation-consisten
- RUNNING (do NOT trim): `executor.run_rq3_battery(live=True)` collecting the full frozen schedule — selector∈{static,random,agent} × pinned churn kp30/steps20, R=30 × C=50 = 4,500 real isolated-docker circuits (~12h est. at ~9.8s/circuit). Launch dir `output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/` (gitignored; force-add anchors on completion). Verified genuinely progressing: python PID 51087 alive + live 4-container circuit (client+3 hops) up on docker. **Completion-detection: `rq3-battery-results.json` written once at end; progress = count of `rq3-*/rq3-run.json` sidecars → 90 (3 cells × 30 runs).** Containment intact (isolated-docker only, self-traffic, $0 — local heuristic/random arms; no frontier spend). Both prereg SHAs intact; worktree-only. Analysis (RQ3-P1-perf/latency, RQ3-P2, Holm-7) runs AFTER completion on the sealed record. - RUNNING (do NOT trim): `executor.run_rq3_battery(live=True)` collecting the full frozen schedule — selector∈{static,random,agent} × pinned churn kp30/steps20, R=30 × C=50 = 4,500 real isolated-docker circuits (~12h est. at ~9.8s/circuit). Launch dir `output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/` (gitignored; force-add anchors on completion). Verified genuinely progressing: python PID 51087 alive + live 4-container circuit (client+3 hops) up on docker. **Completion-detection: `rq3-battery-results.json` written once at end; progress = count of `rq3-*/rq3-run.json` sidecars → 90 (3 cells × 30 runs).** Containment intact (isolated-docker only, self-traffic, $0 — local heuristic/random arms; no frontier spend). Both prereg SHAs intact; worktree-only. Analysis (RQ3-P1-perf/latency, RQ3-P2, Holm-7) runs AFTER completion on the sealed record.
- 2026-07-21 COMPANION PAPER — RQ2-P3 HALF FILLED POST-SEAL (D3 combined paper). `docs/stage-07-companion-methods.md`: §3 un-marked FREEZE-PENDING → FROZEN 2026-07-21 (SHA 8db4e8a7…); §5 Results + §6 Discussion for RQ2-P3 filled FROM THE SEALED RECORD ONLY (H1 ρ=+0.6244 CI[+0.5941,+0.6545]; H2 β=+0.7052 CI[+0.6195,+0.7903] n=270; H3 RESOLVED=MIX; Holm own {H1,H2} both reject — effect+CI, never bare p). Finding stated plainly: shared-pool concentration RAISES anonymity (mix), REFUTES naive funnel; framed as honest QUALIFICATION/CORRECTION of lead RQ2-P1 "shrink" as a unique-bridge (fresh-bridge-per-circuit) artifact — cites `docs/note-unique-bridge-artifact.md` + frozen mechanism prereg SHA 8db4e8a7…. MANDATORY disclosure carried: §7 dry-pass previewed direction (ρ 0→+0.838); confirmatory quantifies effect already visible at calibration; pre-committed hypotheses were two-sided. - 2026-07-21 COMPANION PAPER — RQ2-P3 HALF FILLED POST-SEAL (D3 combined paper). `docs/stage-07-companion-methods.md`: §3 un-marked FREEZE-PENDING → FROZEN 2026-07-21 (SHA 8db4e8a7…); §5 Results + §6 Discussion for RQ2-P3 filled FROM THE SEALED RECORD ONLY (H1 ρ=+0.6244 CI[+0.5941,+0.6545]; H2 β=+0.7052 CI[+0.6195,+0.7903] n=270; H3 RESOLVED=MIX; Holm own {H1,H2} both reject — effect+CI, never bare p). Finding stated plainly: shared-pool concentration RAISES anonymity (mix), REFUTES naive funnel; framed as honest QUALIFICATION/CORRECTION of lead RQ2-P1 "shrink" as a unique-bridge (fresh-bridge-per-circuit) artifact — cites `docs/note-unique-bridge-artifact.md` + frozen mechanism prereg SHA 8db4e8a7…. MANDATORY disclosure carried: §7 dry-pass previewed direction (ρ 0→+0.838); confirmatory quantifies effect already visible at calibration; pre-committed hypotheses were two-sided.
- BLIND HOLD (RQ3): RQ3 Results/Discussion AND authoritative Holm-7 LEFT AS HELD-BLIND placeholders — RQ3 battery still running (progress read by sidecar COUNT only = 3/90; NO rq3-run.json contents read). Lead RQ2-P1 NOT re-litigated (companion QUALIFIES; committed lead paper + frozen prereg f22331a72e… untouched). This stone: $0/offline, worktree-only. NEXT: HOLD for RQ3 — no new stones, no peeking at RQ3 data until `rq3-battery-results.json` lands. - BLIND HOLD (RQ3): RQ3 Results/Discussion AND authoritative Holm-7 LEFT AS HELD-BLIND placeholders — RQ3 battery still running (progress read by sidecar COUNT only = 3/90; NO rq3-run.json contents read). Lead RQ2-P1 NOT re-litigated (companion QUALIFIES; committed lead paper + frozen prereg f22331a72e… untouched). This stone: $0/offline, worktree-only. NEXT: HOLD for RQ3 — no new stones, no peeking at RQ3 data until `rq3-battery-results.json` lands.
- 2026-07-22 RQ3 BATTERY COMPLETE (90/90 runs, live-docker-e2e, exited clean) → UN-BLIND + FROZEN analysis. RAW SEALED: `.../20260722T040640Z/confirmatory-data/SHA256SUMS` (battery-results.json `5b61e461…` + 90 rq3-run.json sidecars). New analyzer `analysis/rq3_confirm.py` (RQ3 analogue of rq2p3_confirm; frozen `stats`/`metrics` UNTOUCHED — run-level multi-arm bootstrap mirrors `two_sample_diff_ci`, 10k BCa, α=0.05). Sealed analysis `.../20260722T040640Z/analysis/rq3-confirmatory-analysis.json` (SHA `e09c66ef…`). Tests `tests/test_sor_rq3_confirm.py` 6 passed; full SOR suite 207 passed (no regression).
- RQ3 RESULT (effect+CI, never bare p): **RQ3-P1-perf FAIL/H0** — retention margin agentmax(static,random) = **0.6pp**, BCa CI [1.58pp, +0.39pp]; every selector heals ~all churn drops (~99% retention) so no ≥+10pp agent gain. **RQ3-P1-latency HOLDS budget** — added-latency(agentmin-baseline=random) = **13.5ms**, CI [52.1, +34.9]ms, upper ≤100ms (agent not slower). **RQ3-P2 FAIL/not-excluded** — rebuild-classifier AUC(agent vs pooled baseline, per-run mean-gap) = **0.587**, CI [0.458, 0.703], upper 0.703 > 0.60 → fingerprint NOT excluded (underpowered at n=30). **RQ3-P3 = H0** (P1-perf fails ∧ P2 fails). Honest null: the local agent selector neither beats baselines nor is certifiably non-classifiable.
- AUTHORITATIVE HOLM-7 (frozen family size=7; supersedes lead's conservative partial embedding, D3). RQ2-P3 slot carries the mechanism-corrected primary **H1-pooled Spearman p=0** (not the lead's degenerate p=1). **SURVIVORS: RQ1-P1 (r1×7), RQ2-P1 shrink (r2×6), RQ2-P3 mix (r3×5)** all holm_p=0. NON-survivors: RQ1-P2 (r4×4, holm_p=0.365), RQ3-P2 (r5×3, 0.511), RQ3-P1-perf (r6×2, 0.511), RQ3-P1-latency (r7×1, 0.511). Lead RQ1-P1 & RQ2-P1 survive regardless. Both prereg SHAs intact; $0/offline; worktree-only. NEXT: fill RQ3 half of companion paper (stone 2).
+310
View File
@@ -0,0 +1,310 @@
"""RQ3 confirmatory analysis — reads the SEALED live-docker battery, applies the
FROZEN prereg §6 plan, and computes the authoritative Holm-7 over the whole family.
This is the RQ3 analogue of ``rq2p3_confirm.py``: it does **not** re-specify anything.
It loads the immutable ``rq3-battery-results.json`` (the operator-gated live-docker
run) and the pinned params in ``docs/rq3-companion-run-brief.md`` and computes the three
frozen RQ3 confirmatory tests, each as an **effect + BCa 95% CI** (never a bare p; the p
is carried only to order the Holm family, per prereg §6):
* **RQ3-P1-perf** — throughput-retention margin: mean retention(agent)
max(mean retention(static), mean retention(random)); perf holds iff the CI **lower**
bound ≥ +10 pp (frozen gate).
* **RQ3-P1-latency** — added-latency(agent) = median e2e latency(agent)
min(median latency(static), median latency(random)) [the faster / min-latency baseline
arm, run-brief §3.2]; anonymity-latency budget holds iff the CI **upper** bound ≤ 100 ms.
* **RQ3-P2** — rebuild-classifier AUC separating the agent selector's per-run mean
inter-rebuild-gap signal from the pooled baseline selectors' signal (the fingerprint
question, §1(b)); anonymity holds iff the CI **upper** bound ≤ 0.60.
* **RQ3-P3** — logical AND: CONFIRM iff (P1-perf ∧ P1-latency) ∧ P2; else H0.
The BCa machinery is the frozen ``stats`` toolkit — this harness only *drives* it at the
run level (resample whole runs, the confirmatory grouping unit) via a generic multi-arm
bootstrap that mirrors ``stats.two_sample_diff_ci`` (independent per-arm resampling, a
combined leave-one-out jackknife, ``stats._bca_endpoints``). No frozen instrument is edited.
**Authoritative Holm-7.** Once all seven confirmatory p-values exist, this computes the
exact ``stats.holm_bonferroni`` step-down over the frozen size-7 family
{RQ1-P1, RQ1-P2, RQ2-P1, RQ2-P3, RQ3-P1-perf, RQ3-P1-latency, RQ3-P2}. The four prior
p-values are read from the SEALED lead record (RQ1-P1, RQ1-P2, RQ2-P1) and the SEALED
RQ2-P3 mechanism record (RQ2-P3 slot = its primary H1-pooled Spearman ρ — the direct
operationalization of the frozen single-slot "Spearman ρ between concentration and H";
the lead's degenerate as-instrumented RQ2-P3 is superseded by the mechanism-corrected
result, per operator decision D3). This Holm-7 is the authoritative final correction and
supersedes the lead paper's deliberately conservative *partial* embedding; both remain
valid, the partial never under-corrects.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import statistics
import sys
from collections import defaultdict
from pathlib import Path
from typing import Callable, Dict, List, Sequence
from cmd_chat.sor.analysis import stats
from cmd_chat.sor.analysis.metrics import rebuild_classifier_auc
# --- frozen family + gates -------------------------------------------------- #
FAMILY_SIZE = 7
FROZEN_FAMILY = (
"RQ1-P1", "RQ1-P2", "RQ2-P1", "RQ2-P3",
"RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2",
)
PERF_MARGIN_MIN_PP = 0.10 # RQ3-P1-perf gate: CI lower ≥ +10 pp
LATENCY_MAX_MS = 100.0 # RQ3-P1-latency gate: CI upper ≤ 100 ms
AUC_CEILING = 0.60 # RQ3-P2 gate: CI upper ≤ 0.60
# sealed prior records (immutable) that carry the 4 non-RQ3 family p-values
LEAD_RESULTS = "output/sor-confirmatory/20260720T060132Z/analysis/stage06-results.json"
RQ2P3_RESULTS = "output/sor-rq2p3-confirmatory/rq2p3-confirmatory-results.json"
def _seed(tag: str) -> int:
"""Deterministic, auditable per-test resampling seed."""
return int.from_bytes(hashlib.sha256(f"sor-rq3-confirm|{tag}".encode()).digest()[:8], "big")
# --- generic run-level multi-arm bootstrap (mirrors stats.two_sample_diff_ci) --- #
def _multi_arm_bootstrap(
arms: Dict[str, Sequence[float]],
statistic: Callable[[Dict[str, List[float]]], float],
*,
null: float,
seed: int,
n_resamples: int = stats.DEFAULT_RESAMPLES,
alpha: float = stats.DEFAULT_ALPHA,
):
"""Bootstrap CI + Holm-ordering p for a statistic over >=1 independent arms.
Each arm is resampled with replacement independently (the frozen two-sample rule
generalized to N arms); BCa uses a combined leave-one-out jackknife (each point
dropped from its OWN arm, others full), exactly as ``stats.two_sample_diff_ci``.
The p is ``stats.two_sided_bootstrap_p`` against ``null`` — carried only to order
the Holm family, never as a stand-alone decision.
"""
import random
names = list(arms)
data = {k: list(v) for k, v in arms.items()}
theta_hat = float(statistic(data))
rng = random.Random(seed)
thetas: List[float] = []
for _ in range(n_resamples):
sample = {k: [v[rng.randrange(len(v))] for _ in range(len(v))] for k, v in data.items()}
thetas.append(float(statistic(sample)))
thetas.sort()
# combined jackknife: drop point i from arm `name`, keep the others full
can_bca = all(len(v) > 1 for v in data.values())
if can_bca:
jack: List[float] = []
for name in names:
n = len(data[name])
for i in range(n):
d = dict(data)
d[name] = [data[name][j] for j in range(n) if j != i]
jack.append(float(statistic(d)))
q_lo, q_hi, used = stats._bca_endpoints(thetas, theta_hat, jack, alpha)
else:
q_lo, q_hi, used = alpha / 2.0, 1.0 - alpha / 2.0, "percentile"
ci = stats.CIResult(theta_hat, stats._percentile(thetas, q_lo),
stats._percentile(thetas, q_hi), alpha, n_resamples, used, seed)
p = stats.two_sided_bootstrap_p(thetas, null)
return ci, p
# --- battery loading -------------------------------------------------------- #
def load_battery(path: Path) -> Dict[str, Dict[str, List[float]]]:
"""Load the sealed battery into per-arm {retention, latency, gap_signal}.
``gap_signal`` is the per-run MEAN inter-rebuild gap — the confirmatory grouping
unit (same unit the frozen calibration gate used); runs with no gap are omitted
(no fabricated interval). Retention/latency are the per-run DVs as measured.
"""
doc = json.loads(Path(path).read_text())
arms: Dict[str, Dict[str, List[float]]] = {}
for _cid, cell in doc["cells"].items():
arms[cell["strategy"]] = {
"retention": list(cell["throughput_retention"]),
"latency": list(cell["added_latency_ms"]),
"gap_signal": [],
}
gap: Dict[str, List[float]] = defaultdict(list)
for run in doc["runs"]:
strat = run["cell_id"].split("selector=")[1].split("/")[0]
gaps = run.get("rebuild_gaps") or []
if gaps:
gap[strat].append(statistics.fmean(gaps))
for strat, sig in gap.items():
arms[strat]["gap_signal"] = sig
return arms
# --- the three frozen RQ3 tests --------------------------------------------- #
def _perf_margin(a: Dict[str, List[float]]) -> float:
return stats.mean(a["agent"]) - max(stats.mean(a["static"]), stats.mean(a["random"]))
def _added_latency(a: Dict[str, List[float]]) -> float:
# added-latency over the MIN-latency (faster) baseline arm, run-brief §3.2
return statistics.median(a["agent"]) - min(statistics.median(a["static"]), statistics.median(a["random"]))
def _p2_auc(a: Dict[str, List[float]]) -> float:
# fingerprint question: agent's per-run rebuild-gap signal vs the pooled baseline
return rebuild_classifier_auc(a["agent"], a["baseline"])
def analyze_rq3(arms: Dict[str, Dict[str, List[float]]], *, n_resamples: int = stats.DEFAULT_RESAMPLES) -> Dict:
ret = {k: arms[k]["retention"] for k in ("agent", "static", "random")}
lat = {k: arms[k]["latency"] for k in ("agent", "static", "random")}
pool_baseline = arms["static"]["gap_signal"] + arms["random"]["gap_signal"]
p2_arms = {"agent": arms["agent"]["gap_signal"], "baseline": pool_baseline}
perf_ci, perf_p = _multi_arm_bootstrap(ret, _perf_margin, null=0.0,
seed=_seed("perf"), n_resamples=n_resamples)
lat_ci, lat_p = _multi_arm_bootstrap(lat, _added_latency, null=0.0,
seed=_seed("latency"), n_resamples=n_resamples)
p2_ci, p2_p = _multi_arm_bootstrap(p2_arms, _p2_auc, null=0.5,
seed=_seed("p2-auc"), n_resamples=n_resamples)
perf_hold = perf_ci.lo >= PERF_MARGIN_MIN_PP # gate: CI lower ≥ +10 pp
lat_hold = lat_ci.hi <= LATENCY_MAX_MS # gate: CI upper ≤ 100 ms
p2_hold = p2_ci.hi <= AUC_CEILING # gate: CI upper ≤ 0.60
p1_hold = perf_hold and lat_hold
p3_confirm = p1_hold and p2_hold
which_lat_baseline = "random" if statistics.median(lat["random"]) <= statistics.median(lat["static"]) else "static"
return {
"RQ3-P1-perf": {
"effect": "throughput_retention_margin_agent_minus_max_baseline",
**perf_ci.as_dict(), "p_for_holm": perf_p,
"gate": f"CI lower ≥ +{PERF_MARGIN_MIN_PP}", "holds": bool(perf_hold),
"decision": "perf-gain" if perf_hold else "no-perf-gain",
},
"RQ3-P1-latency": {
"effect": "added_latency_ms_agent_minus_min_baseline",
"min_latency_baseline_arm": which_lat_baseline,
**lat_ci.as_dict(), "p_for_holm": lat_p,
"gate": f"CI upper ≤ {LATENCY_MAX_MS} ms", "holds": bool(lat_hold),
"decision": "within-latency-budget" if lat_hold else "over-latency-budget",
},
"RQ3-P2": {
"effect": "rebuild_classifier_auc_agent_vs_pooled_baseline",
"grouping_unit": "per-run mean inter-rebuild gap",
"n_agent_runs": len(p2_arms["agent"]), "n_baseline_runs": len(p2_arms["baseline"]),
**p2_ci.as_dict(), "p_for_holm": p2_p,
"gate": f"CI upper ≤ {AUC_CEILING}", "holds": bool(p2_hold),
"decision": "no-usable-fingerprint" if p2_hold else "fingerprint-not-excluded",
},
"RQ3-P3-joint": {
"rule": "CONFIRM iff (P1-perf ∧ P1-latency) ∧ P2",
"p1_holds": bool(p1_hold), "p2_holds": bool(p2_hold),
"confirm": bool(p3_confirm),
"decision": "agent-helps-without-fingerprint" if p3_confirm else "H0",
},
}
# --- authoritative Holm-7 --------------------------------------------------- #
def _load_prior_pvalues(lead_path: Path, rq2p3_path: Path) -> Dict[str, Dict]:
lead = json.loads(Path(lead_path).read_text())["confirmatory"]
rq2p3 = json.loads(Path(rq2p3_path).read_text())["results"]
h1 = rq2p3["H1_pooled_spearman"]
return {
"RQ1-P1": {"p": lead["RQ1-P1"]["p_for_holm"], "source": "sealed lead RQ1-P1"},
"RQ1-P2": {"p": lead["RQ1-P2"]["p_for_holm"], "source": "sealed lead RQ1-P2"},
"RQ2-P1": {"p": lead["RQ2-P1"]["p_for_holm"], "source": "sealed lead RQ2-P1 (shrink)"},
"RQ2-P3": {"p": h1["p_for_holm"],
"source": "sealed RQ2-P3 H1-pooled Spearman ρ (mechanism-corrected, mix) "
"— supersedes the lead's degenerate as-instrumented RQ2-P3"},
}
def authoritative_holm7(rq3: Dict, lead_path: Path, rq2p3_path: Path) -> Dict:
prior = _load_prior_pvalues(lead_path, rq2p3_path)
pvals = {
"RQ1-P1": prior["RQ1-P1"]["p"],
"RQ1-P2": prior["RQ1-P2"]["p"],
"RQ2-P1": prior["RQ2-P1"]["p"],
"RQ2-P3": prior["RQ2-P3"]["p"],
"RQ3-P1-perf": rq3["RQ3-P1-perf"]["p_for_holm"],
"RQ3-P1-latency": rq3["RQ3-P1-latency"]["p_for_holm"],
"RQ3-P2": rq3["RQ3-P2"]["p_for_holm"],
}
assert set(pvals) == set(FROZEN_FAMILY), "family must be exactly the frozen size-7"
holm = stats.holm_bonferroni(pvals, family_size=FAMILY_SIZE)
rows = [{"name": h.name, "raw_p": h.p, "holm_p": h.p_adjusted,
"multiplier": h.multiplier, "rank": h.rank, "reject": h.reject} for h in holm]
return {
"family_size": FAMILY_SIZE,
"family": list(FROZEN_FAMILY),
"p_sources": {k: prior[k]["source"] for k in prior},
"rows": rows,
"survivors": [r["name"] for r in rows if r["reject"]],
"non_survivors": [r["name"] for r in rows if not r["reject"]],
"supersedes": "lead paper's conservative partial embedding (report-4 of family-of-7); "
"both valid, partial never under-corrects; RQ1-P1 and RQ2-P1 survive regardless",
}
def run(*, battery_path: Path, lead_path: Path, rq2p3_path: Path,
n_resamples: int = stats.DEFAULT_RESAMPLES) -> Dict:
arms = load_battery(battery_path)
rq3 = analyze_rq3(arms, n_resamples=n_resamples)
holm7 = authoritative_holm7(rq3, lead_path, rq2p3_path)
battery_sha = hashlib.sha256(Path(battery_path).read_bytes()).hexdigest()
return {
"schema": "sor-rq3-confirmatory/1",
"measured_from": "live-docker-e2e",
"battery_results_path": str(battery_path),
"battery_results_sha256": battery_sha,
"n_runs_per_arm": {k: len(arms[k]["retention"]) for k in arms},
"frozen_lead_prereg_sha256": "f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b",
"results": rq3,
"authoritative_holm7": holm7,
"honest_disclosure": (
"Every RQ3 test reports effect + BCa 95% CI; p is carried ONLY to order the Holm "
"family (prereg §6). Nulls are results: a selector that does not beat baselines, or "
"a rebuild pattern that is not certifiably non-classifiable, is the finding — not spun. "
"Reproducibility caveat (accepted, run-brief §2A): the agent (qwen2.5:3b via local "
"Ollama, temp 0) is reproducible via its committed decision-log + (seed, state-hash) "
"cache replay, NOT via independent model re-execution on other hardware."
),
}
def _seal(out_dir: Path, report: Dict) -> Path:
out_dir.mkdir(parents=True, exist_ok=True)
path = out_dir / "rq3-confirmatory-analysis.json"
path.write_text(json.dumps(report, indent=2, sort_keys=True) + "\n", encoding="utf-8")
digest = hashlib.sha256(path.read_bytes()).hexdigest()
(out_dir / "SHA256SUMS").write_text(f"{digest} {path.name}\n", encoding="utf-8")
return path
def main(argv=None) -> int:
ap = argparse.ArgumentParser(prog="python -m cmd_chat.sor.analysis.rq3_confirm")
ap.add_argument("--battery", required=True)
ap.add_argument("--lead", default=LEAD_RESULTS)
ap.add_argument("--rq2p3", default=RQ2P3_RESULTS)
ap.add_argument("--out", default=None)
ap.add_argument("--n-resamples", type=int, default=stats.DEFAULT_RESAMPLES)
args = ap.parse_args(argv)
report = run(battery_path=Path(args.battery), lead_path=Path(args.lead),
rq2p3_path=Path(args.rq2p3), n_resamples=args.n_resamples)
if args.out:
path = _seal(Path(args.out), report)
print(f"sealed -> {path}", file=sys.stderr)
print(json.dumps(report, indent=2, sort_keys=True))
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1 @@
e09c66efe00524ae3e5373bd20121e6e022e6885596f4fa9aabd51653c76929b rq3-confirmatory-analysis.json
@@ -0,0 +1,156 @@
{
"authoritative_holm7": {
"family": [
"RQ1-P1",
"RQ1-P2",
"RQ2-P1",
"RQ2-P3",
"RQ3-P1-perf",
"RQ3-P1-latency",
"RQ3-P2"
],
"family_size": 7,
"non_survivors": [
"RQ1-P2",
"RQ3-P2",
"RQ3-P1-perf",
"RQ3-P1-latency"
],
"p_sources": {
"RQ1-P1": "sealed lead RQ1-P1",
"RQ1-P2": "sealed lead RQ1-P2",
"RQ2-P1": "sealed lead RQ2-P1 (shrink)",
"RQ2-P3": "sealed RQ2-P3\u2032 H1-pooled Spearman \u03c1 (mechanism-corrected, mix) \u2014 supersedes the lead's degenerate as-instrumented RQ2-P3"
},
"rows": [
{
"holm_p": 0.0,
"multiplier": 7,
"name": "RQ1-P1",
"rank": 1,
"raw_p": 0.0,
"reject": true
},
{
"holm_p": 0.0,
"multiplier": 6,
"name": "RQ2-P1",
"rank": 2,
"raw_p": 0.0,
"reject": true
},
{
"holm_p": 0.0,
"multiplier": 5,
"name": "RQ2-P3",
"rank": 3,
"raw_p": 0.0,
"reject": true
},
{
"holm_p": 0.3648,
"multiplier": 4,
"name": "RQ1-P2",
"rank": 4,
"raw_p": 0.0912,
"reject": false
},
{
"holm_p": 0.5112,
"multiplier": 3,
"name": "RQ3-P2",
"rank": 5,
"raw_p": 0.1704,
"reject": false
},
{
"holm_p": 0.5112,
"multiplier": 2,
"name": "RQ3-P1-perf",
"rank": 6,
"raw_p": 0.177,
"reject": false
},
{
"holm_p": 0.5112,
"multiplier": 1,
"name": "RQ3-P1-latency",
"rank": 7,
"raw_p": 0.497,
"reject": false
}
],
"supersedes": "lead paper's conservative partial embedding (report-4 of family-of-7); both valid, partial never under-corrects; RQ1-P1 and RQ2-P1 survive regardless",
"survivors": [
"RQ1-P1",
"RQ2-P1",
"RQ2-P3"
]
},
"battery_results_path": "output/sor-rq3-confirmatory/20260722T040640Z/confirmatory-data/rq3-battery-results.json",
"battery_results_sha256": "5b61e461004722985c1a7a9bc5fdfe855395df3180c71e6efdc3531b8ecf8039",
"frozen_lead_prereg_sha256": "f22331a72e0d0ccf38b787e63acabbe9d666456ec76076787a6d545c3193425b",
"honest_disclosure": "Every RQ3 test reports effect + BCa 95% CI; p is carried ONLY to order the Holm family (prereg \u00a76). Nulls are results: a selector that does not beat baselines, or a rebuild pattern that is not certifiably non-classifiable, is the finding \u2014 not spun. Reproducibility caveat (accepted, run-brief \u00a72A): the agent (qwen2.5:3b via local Ollama, temp 0) is reproducible via its committed decision-log + (seed, state-hash) cache replay, NOT via independent model re-execution on other hardware.",
"measured_from": "live-docker-e2e",
"n_runs_per_arm": {
"agent": 30,
"random": 30,
"static": 30
},
"results": {
"RQ3-P1-latency": {
"alpha": 0.05,
"ci_hi": 34.911185735836625,
"ci_lo": -52.14837333187461,
"decision": "within-latency-budget",
"effect": "added_latency_ms_agent_minus_min_baseline",
"gate": "CI upper \u2264 100.0 ms",
"holds": true,
"method": "bca",
"min_latency_baseline_arm": "random",
"n_resamples": 10000,
"p_for_holm": 0.497,
"point": -13.491572928614914,
"seed": 5867095940896968561
},
"RQ3-P1-perf": {
"alpha": 0.05,
"ci_hi": 0.003865947497819883,
"ci_lo": -0.015830325279388435,
"decision": "no-perf-gain",
"effect": "throughput_retention_margin_agent_minus_max_baseline",
"gate": "CI lower \u2265 +0.1",
"holds": false,
"method": "bca",
"n_resamples": 10000,
"p_for_holm": 0.177,
"point": -0.006312034914976117,
"seed": 16001687329924348899
},
"RQ3-P2": {
"alpha": 0.05,
"ci_hi": 0.7034362881655691,
"ci_lo": 0.45805555555555555,
"decision": "fingerprint-not-excluded",
"effect": "rebuild_classifier_auc_agent_vs_pooled_baseline",
"gate": "CI upper \u2264 0.6",
"grouping_unit": "per-run mean inter-rebuild gap",
"holds": false,
"method": "bca",
"n_agent_runs": 30,
"n_baseline_runs": 60,
"n_resamples": 10000,
"p_for_holm": 0.1704,
"point": 0.5869444444444445,
"seed": 9396385129919407505
},
"RQ3-P3-joint": {
"confirm": false,
"decision": "H0",
"p1_holds": false,
"p2_holds": false,
"rule": "CONFIRM iff (P1-perf \u2227 P1-latency) \u2227 P2"
}
},
"schema": "sor-rq3-confirmatory/1"
}
@@ -0,0 +1,91 @@
5b61e461004722985c1a7a9bc5fdfe855395df3180c71e6efdc3531b8ecf8039 rq3-battery-results.json
110bc93cd92f43f679f108d34e16ca97af1a45076167597509caf431f539eb0d rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r0/rq3-run.json
80e238396a0c4359da81ac4a32d8da645c7bc045908d3c9e0ab6fbb47fd31c05 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r1/rq3-run.json
c1c41673ab235dde0c0e0ce968c0ca8b6876e37425486e5789ea8da2a9e2ae56 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r10/rq3-run.json
3e3609b9e78cb26bb5202e9f8f77acd4469a35ea2fc925a25c8f3c449c0598ee rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r11/rq3-run.json
add02b8c1fe94a27e7cfcfdd2017709a572fc635859ea7e9fa6a9aeb7752b179 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r12/rq3-run.json
a5fc0270f9e87a25ffbf3fbf41dc741fcb1aa981b7312249821c93d666786d19 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r13/rq3-run.json
72600ffe7378da5b04c8b7cc16c3618d11b1cce6ab3e166a5d58a76379bd3602 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r14/rq3-run.json
259543832d3d394379849b1c2736962119b290f8b3cad7aa05562bf5f18390cd rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r15/rq3-run.json
d78e5306cae58b7bc46f80244752090cac60102bd037d68d76fe2740b53ebd1e rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r16/rq3-run.json
9bffbb317df459bb7847dfc58c5462066ace4a4befad3ff5919f0986621673eb rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r17/rq3-run.json
d91c00e81cec24fa922a0fdcedc5c5745293d4522b4b2a5b4f3c7df6fe17a179 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r18/rq3-run.json
f293fe699c1aab60c9650eac031aedc4a605bb62c00cdb92794e2b768e2e8a50 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r19/rq3-run.json
ec661f5c52133cf814428bd954cea3d1f6118c35c36cb01640573530e1c35f26 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r2/rq3-run.json
ee1fa456439bdc1c8a8264252b11446a212fda555ac26688c885dbcf5d0d4b65 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r20/rq3-run.json
bb710e7fd11a50bc262665e857d591645520b47dc4ba9afbd15b770c7b36fa85 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r21/rq3-run.json
7c92a5ec490f8c40abd7d4d5d05effb00d5c106567426016a17a52cbbde32a28 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r22/rq3-run.json
cc57bded029cd9f7bb9bc61beef7a2eaecdad798a9872fa570986e10291b92a7 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r23/rq3-run.json
6d8318fa57e16597d500e6b29661843db532f09cb303e409051611befb8d03e5 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r24/rq3-run.json
95e2a0df7e320ad3aeec056eb8c6fe872789311b3fc2d85f3b0730fc1984c6e1 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r25/rq3-run.json
3a036309872aa31fcf150074097226da5bac6485f10786e25313825b983a3fc6 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r26/rq3-run.json
ad59b9aad30cd10be94c1bf0bd58716a3272fd2e75e1da58aaed5ecaca64bfd0 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r27/rq3-run.json
0b0c00ddc3db49aa0c85a1b3dc29292435cb53c3753446edd351205115d38f54 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r28/rq3-run.json
99e28987eaa6da592df7eddcc1134a0ee09bd833f79af54ebe3b6f2609cf256f rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r29/rq3-run.json
86cb773d2315ac118790289779337514e80006817c0c1ec07e0f8d99adaabf6c rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r3/rq3-run.json
08ad295c823da04a6ccbdd758d638c4aae4c7220e9cf4361c81672d800e4bef8 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r4/rq3-run.json
c7daf7fd45aad4c66570875a41aed8d7c69d0d1246f3291d7983b894d547a6e1 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r5/rq3-run.json
ae47e7fce7b8acbd36567e31ae1fa5879d5fa483c6d293e0ae588ebf46d925e7 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r6/rq3-run.json
ce234e3ed894d1c7ba2f6b650c50aec42a3bd423e894f422a64ff1b3a57b3376 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r7/rq3-run.json
723d66a199e09c13de8735b876db87145e5a3950aae9d749ef6abb91fa6f2af9 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r8/rq3-run.json
f30da953eb66d498d4e2be2607e8a1d47a4ac3f626f1cc891852c8ef07672919 rq3-RQ3_topo=1house_bridge=off_selector=agent_churn=kp30s20-r9/rq3-run.json
f276456e1cfc914fcdb963f5d85e72a33d2858369547d4d11e8da739a2c315e8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r0/rq3-run.json
ebd567bb9b129d75d7e794b4f06fa1340263fe7c4abf2cc68546ab476e8e91b9 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r1/rq3-run.json
761a36d6ec339b84d465ffc9dcd88103e6b0130fb9b1ba4aebeb534a8bf581a4 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r10/rq3-run.json
2c663909db5624832bae8393648a1c50a4ac00ab580f4ae7a5f8bdbdce0ef3e2 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r11/rq3-run.json
5a663ef49d5860187fd2a7fc0453ae2f3e3e312ff1224cb75db57c3eb498711a rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r12/rq3-run.json
2af0bd978aa2b92842b5308ec262f508d3f244b9640cc7f16c0dd5f16c05f843 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r13/rq3-run.json
77331f8a9041755c77155da29b8176b4c27b674d741f23ee4bf213f84e008949 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r14/rq3-run.json
2f4b249f9afba300094bb2ce3b50265abe09350ce8959ec2b471ece8ffdfa6ae rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r15/rq3-run.json
0db6e249bf8b6ac0d6d255437a9f868d87641f66fc7fb93f7ca809674884bdb7 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r16/rq3-run.json
76aac9803ac18a5f6f8f0016ba3b7970fcff475a8475fc959834a4a88f1a0442 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r17/rq3-run.json
2a25407358b9ab36cb4997af9f232dcb1f0247015268278ec63cada29cfe6a39 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r18/rq3-run.json
e0f404e78b35404f789120a8b898e03a6ca1bad292bf1ac4efc68bd2004188f0 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r19/rq3-run.json
2515a88a3116a0d490b9d5484517479ed18af45436b33ffa31ee1293ee7fb6e8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r2/rq3-run.json
73593f9667e94ed16fdfa6b943d4b2cc7174cd62ca2cf6ec25f5bc4354d3ec62 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r20/rq3-run.json
9d89871c47cc12ac826dbeecab7d3943bff252248ecbbc5e88e444ec83f1fda1 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r21/rq3-run.json
d6753dc2097d22c67cc31fdfca225b336e8439a4975b4f1428879796c1d0410e rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r22/rq3-run.json
befe39c32b87e73169ec6c111ca79f15fbb24d4952b5dd817ce766399e10824e rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r23/rq3-run.json
cad75228e7ef404fbc54561d3f5b4969f441b047abaa0663ae1d51322f43acaa rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r24/rq3-run.json
2817c84be3af073fa471e0287a845d8147d01a1d0f8e800faa281ce02885afe4 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r25/rq3-run.json
3e7d841db693adb2fe61218588da3e269d31e337242241d0ac87115f516d221c rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r26/rq3-run.json
6a5b5d07bcdf363f4fb959135de6cc8b23c48e5a76888abd0fd2e67e80e5673f rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r27/rq3-run.json
44d512c8eeef83c410399303d482ef760fa490630b26a4cc5ca7294e7f324a14 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r28/rq3-run.json
6ea17d94630ccc117637b3a2ea0ddc8926f0085ad7cd5cea2aeba8de66997150 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r29/rq3-run.json
1d07bbd03e31d2717103d5d766d20c43730e3d65d1b95e551531f198744544cb rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r3/rq3-run.json
a22cd7af4000eea7529ba1e7dd257184f702d77fed44965a9a91c7944b65c2c9 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r4/rq3-run.json
7169dd2d4cf99c53537ce9fd5608f7dda118f645d01cbee5124721f3d03654cc rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r5/rq3-run.json
3c6352374e680178deb5f52f20839202cebf8536fe3b1afb7e0a8e475ae36c27 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r6/rq3-run.json
4cb9c0fcf183a53bb7dc428c12ba3482adae9d7554238e49520b4848b7599fb8 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r7/rq3-run.json
763e75f8b465cf8b2b41c644b47f9e2c71a85995c569c2700e929feecd087c96 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r8/rq3-run.json
3bdc4dcf22de5f09f033b03bd16c01cb9c1cd4fc210ce9b57c8a0c65ec663e80 rq3-RQ3_topo=1house_bridge=off_selector=random_churn=kp30s20-r9/rq3-run.json
257326c8a34c602e028fca9593e8fe628a950eff63138f6eeeb425959112ed5a rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r0/rq3-run.json
1694018a795e64d07b5ff83f5ec512a0c6e38a4f210b3ea1d54b41281d54e21f rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r1/rq3-run.json
bdfb55a3b606ef441237ed0828d504dd885d6730b0de01b0d3cc50f8a5c05826 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r10/rq3-run.json
3a13e1de470c79965972a275d72e4c8bf8d1195e3a55a896520904de4d6c2e68 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r11/rq3-run.json
4449a0a53f7033cb7befd8520501b0f00b9201738e517c023576d7b861720935 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r12/rq3-run.json
45105408efac8d96b449dd3e1945f5727f9f0189a85d3e6bfa5a3c2db0b43882 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r13/rq3-run.json
07a9e31fec1807d5200978cbad6267966acd74d3eec640868e87bb778d4d4722 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r14/rq3-run.json
76830fb704f7439be2443fb2c051aeaeaaafddf32e35f209549bf74aab4ffe82 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r15/rq3-run.json
a4d9a9eb631a3c2c459252f79d0ac65c743440cef445c2545f7e08141130b34e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r16/rq3-run.json
7357f8e01e000bea1c27a0ccaed7d41f01a5e1f6f75569561a6e1901dc696442 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r17/rq3-run.json
4ee979e846426d382f6aea04e205cfe33a3478add09273956decf5aa6db28c37 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r18/rq3-run.json
19e6c0d747dd802e324ca6225a56e4a849e7db42f09df3b61c703aad4adb7082 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r19/rq3-run.json
cd2638c191fe0a73797f3e2eefbcfba8af9bca0759e6473c420e8af523d2c99e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r2/rq3-run.json
9f3730ef5c8feeed3c2578e8a30e471eeabc5c98d2a57514ad053d6b9328cb36 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r20/rq3-run.json
f947eec0f931023d0e6848809ee3484df191099236c50db46a32d3fcaecf88d8 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r21/rq3-run.json
9d5bc5dff8954fb6641ae9d2f9f83d659339867ac7395278c88d0e190991cd92 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r22/rq3-run.json
1cbeaf4904aff435ac72d8e002bd909ceeceeda2c176794f959b8c7264b021c0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r23/rq3-run.json
efaaad8ffd17cfe3c55e62d09f5b1dfd670c2474ea6219931df8e08eef2d83a0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r24/rq3-run.json
03c89c522e5e2cce9181c20f62edd434a734d8dc530629b36bfb6b5e58173c50 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r25/rq3-run.json
5c171990a45aca7adede446a2bb8b5cec5b72d454d5f4ae061b0883ec80eb5ab rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r26/rq3-run.json
4762af92e07ba3f8c7e6f0ffecc25af15469cbb3d15e151a9acb12619cf37e7e rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r27/rq3-run.json
ec29c89ef5e42ec618ed848c23f9fa35709ce663053929214f05226a055ae2d1 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r28/rq3-run.json
68f3441aff941c9aabba81321e9352b966bf0883a1afb6b05172b4d5ea3c32d9 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r29/rq3-run.json
d8d6aaaa09b62f293750d8f46422ae27e504fd8863591f3f1c34f1a2135c6b0c rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r3/rq3-run.json
20e794c0082d6ec4ebdee3f8e659b5f6a19280a0706ef9910dc5af62329bbe32 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r4/rq3-run.json
5346c57b2b4822aa36906b4348d9be1c4c0c6504cd91676c49af5fa4c637769c rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r5/rq3-run.json
375089aa2ffa966668ad9063acb37e21f3ad176996bdd8f5a66f20d2bca7da53 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r6/rq3-run.json
8e5f07eb6b52125ccb13a58d4fd2316075a416670fe2391627c0954a42bf36f0 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r7/rq3-run.json
e334641cd823338129c050a35a61d1d47a50976a6110301a42294a2e80b47c4a rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r8/rq3-run.json
e2c48513c108020c7e29a3107c7f8c5bfdb8bb8374120fc1827177d81fe772a1 rq3-RQ3_topo=1house_bridge=off_selector=static_churn=kp30s20-r9/rq3-run.json
+128
View File
@@ -0,0 +1,128 @@
"""RQ3 confirmatory-harness pins — structure, determinism, gate logic, Holm-7 wiring.
These exercise the RQ3 confirmatory analyzer's mechanics on a SMALL SYNTHETIC battery
(so the test owns its inputs and does not depend on the sealed live record): (a) the
battery loader recovers per-arm retention/latency and the per-run mean-gap signal; (b) the
three frozen gates read the right CI bound (perf: lower ≥ +10pp; latency: upper ≤ 100ms;
P2: upper ≤ 0.60) and RQ3-P3 is their logical AND; (c) the run-level multi-arm bootstrap is
byte-deterministic (same seed → identical CI) and BCa-well-formed; (d) the authoritative
Holm is over EXACTLY the frozen size-7 family with the RQ2-P3 slot carrying the
mechanism-corrected primary p, and the step-down/survivor logic matches stats.holm_bonferroni.
"""
import json
from cmd_chat.sor.analysis import rq3_confirm as rc
from cmd_chat.sor.analysis import stats
def _synthetic_battery(tmp_path):
"""A 3-arm battery where agent clearly beats baselines on retention and is faster,
with distinct rebuild-gap signals — so the gate directions are unambiguous."""
def cell(strat, ret, lat):
return {"strategy": strat, "runs": len(ret),
"throughput_retention": ret, "added_latency_ms": lat}
cells = {
"a": cell("agent", [1.0] * 8, [100.0] * 8),
"s": cell("static", [0.5] * 8, [400.0] * 8),
"r": cell("random", [0.6] * 8, [300.0] * 8),
}
runs = []
# frozen classifier scores SHORTER gaps (more churn) as positive; give the agent
# shorter gaps than the baselines so it separates upward (AUC -> 1), the fingerprint
# direction the ≤0.60 ceiling is meant to catch.
for strat, gap in (("agent", 1.0), ("static", 3.0), ("random", 3.0)):
for i in range(8):
runs.append({"cell_id": f"RQ3/topo=1house/bridge=off/selector={strat}/churn=kp30s20",
"rebuild_gaps": [gap, gap, gap], "run_index": i})
doc = {"schema": "sor-rq3-battery-results/1", "cells": cells, "runs": runs}
p = tmp_path / "battery.json"
p.write_text(json.dumps(doc))
return p
# --- loader ----------------------------------------------------------------- #
def test_loader_recovers_per_arm_dvs_and_gap_signal(tmp_path):
arms = rc.load_battery(_synthetic_battery(tmp_path))
assert set(arms) == {"agent", "static", "random"}
assert arms["agent"]["retention"] == [1.0] * 8
assert arms["static"]["latency"] == [400.0] * 8
# per-run MEAN gap: three equal gaps -> that value per run, 8 runs
assert arms["agent"]["gap_signal"] == [1.0] * 8
assert arms["random"]["gap_signal"] == [3.0] * 8
# --- gate directions + RQ3-P3 AND ------------------------------------------- #
def test_gates_read_correct_ci_bounds_and_join(tmp_path):
arms = rc.load_battery(_synthetic_battery(tmp_path))
res = rc.analyze_rq3(arms, n_resamples=500)
perf, lat, p2, joint = (res["RQ3-P1-perf"], res["RQ3-P1-latency"],
res["RQ3-P2"], res["RQ3-P3-joint"])
# agent retention 1.0 vs max(0.5,0.6)=0.6 -> margin +0.4, CI lower well above +0.1
assert perf["point"] > 0.1 and perf["holds"] is True and perf["decision"] == "perf-gain"
# agent latency 100 vs min(400,300)=300 -> -200ms, upper bound << 100ms
assert lat["point"] < 0 and lat["holds"] is True
assert lat["min_latency_baseline_arm"] == "random"
# agent gap 1.0 vs baseline 3.0 -> fully separable upward AUC=1.0 -> upper > 0.60 -> fails
assert p2["point"] == 1.0 and p2["holds"] is False
# P3 = (perf ∧ latency) ∧ P2 ; P2 fails here -> H0
assert joint["p1_holds"] is True and joint["p2_holds"] is False
assert joint["confirm"] is False and joint["decision"] == "H0"
# --- determinism + BCa well-formed ------------------------------------------ #
def test_analysis_is_byte_deterministic(tmp_path):
b = _synthetic_battery(tmp_path)
a1 = rc.analyze_rq3(rc.load_battery(b), n_resamples=800)
a2 = rc.analyze_rq3(rc.load_battery(b), n_resamples=800)
assert a1 == a2
for t in ("RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"):
assert a1[t]["ci_lo"] <= a1[t]["point"] <= a1[t]["ci_hi"]
assert 0.0 <= a1[t]["p_for_holm"] <= 1.0
# --- multi-arm bootstrap matches the frozen two-sample rule ------------------ #
def test_multi_arm_bootstrap_matches_two_sample_diff_for_two_means(tmp_path):
a = [1.0, 2.0, 3.0, 4.0, 5.0]
b = [0.5, 1.5, 2.5, 3.5]
ci_gen, _ = rc._multi_arm_bootstrap(
{"x": a, "y": b}, lambda d: stats.mean(d["x"]) - stats.mean(d["y"]),
null=0.0, seed=123, n_resamples=1000)
ci_ref = stats.two_sample_diff_ci(a, b, stats.mean, seed=123, n_resamples=1000)
# same seed + same independent-arm resampling rule -> identical point + interval
assert abs(ci_gen.point - ci_ref.point) < 1e-12
assert abs(ci_gen.lo - ci_ref.lo) < 1e-9 and abs(ci_gen.hi - ci_ref.hi) < 1e-9
# --- authoritative Holm-7 --------------------------------------------------- #
def test_holm7_family_is_exactly_the_frozen_seven(tmp_path):
assert rc.FAMILY_SIZE == 7
assert set(rc.FROZEN_FAMILY) == {
"RQ1-P1", "RQ1-P2", "RQ2-P1", "RQ2-P3",
"RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"}
def test_holm7_survivor_logic_with_three_zeros(tmp_path):
# three p=0 priors + non-significant RQ3 tests: only the three zeros survive Holm-7.
rq3 = {
"RQ3-P1-perf": {"p_for_holm": 0.18},
"RQ3-P1-latency": {"p_for_holm": 0.50},
"RQ3-P2": {"p_for_holm": 0.17},
}
# stub the two sealed prior records
lead = tmp_path / "lead.json"
lead.write_text(json.dumps({"confirmatory": {
"RQ1-P1": {"p_for_holm": 0.0}, "RQ1-P2": {"p_for_holm": 0.0912},
"RQ2-P1": {"p_for_holm": 0.0}}}))
rq2p3 = tmp_path / "rq2p3.json"
rq2p3.write_text(json.dumps({"results": {"H1_pooled_spearman": {"p_for_holm": 0.0}}}))
holm = rc.authoritative_holm7(rq3, lead, rq2p3)
assert holm["family_size"] == 7
assert holm["survivors"] == ["RQ1-P1", "RQ2-P1", "RQ2-P3"]
assert set(holm["non_survivors"]) == {"RQ1-P2", "RQ3-P1-perf", "RQ3-P1-latency", "RQ3-P2"}
# RQ2-P3 slot must carry the mechanism-corrected primary p, not the lead degenerate 1.0
assert "mechanism-corrected" in holm["p_sources"]["RQ2-P3"]
# RQ1-P2 at rank 4 gets multiplier 4 -> 0.0912*4 = 0.3648, not rejected
row = {r["name"]: r for r in holm["rows"]}
assert row["RQ1-P2"]["multiplier"] == 4 and abs(row["RQ1-P2"]["holm_p"] - 0.3648) < 1e-9