Defensive-measurement instrument scaffolds for the re-staged experiment, not
wired into the sealed pipeline:
- house_pool.py: distinct isolated container pools per house (own net + keys),
deterministic seed-derived node ids; refuses engine=="local" (containment).
- watchdog.py: pre-registered HALT / PAUSE-RESUME failsafe for load-bearing
node pools; deterministic, probe injected (no external target).
Both are DRAFT for a fresh pre-registration; neither touches the sealed battery.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Paper (stage-07 lead): add an Apparatus disclosure (§4.3) and a Node-distribution
limitation (§7) stating that all relay hops ran as isolated Docker containers on a
single engine host and the two phones were consenting endpoints, not forwarders.
Containment (engine != local) holds; cross-machine timing is out of scope / future work.
Add the self-contained study site (docs/paper-site/: build_site.py, index.html) built
straight from the sealed paper markdown, plus a generator + six 1080x1920 Instagram-story
teaser slides (make_story_slides.py, story/) using only sealed, audited numbers. Defensive
measurement framing throughout; the pre-registered nulls are reported, not spun.
3-persona integrity red-team of docs/stage-07-companion-methods.md (RQ2-P3
mechanism + RQ3), same discipline as the lead stage-08 review. Runs only on
already-sealed artifacts; computes no new statistic; edits no frozen prereg.
METHODOLOGIST: reproduced the authoritative Holm-7 step-down independently,
byte-for-byte vs the sealed record (survivors RQ1-P1 / RQ2-P1 / RQ2-P3;
non-survivors RQ1-P2 / RQ3-P2 / RQ3-P1-perf / RQ3-P1-latency). Confirmed the
supersedes-partial-embedding claim, the single-slot RQ2-P3 mechanism-corrected
p (family_size 7), run-level cluster-bootstrap validity, two-sided pre-commit.
DOMAIN SKEPTIC: mix-qualifies-shrink scoped as correction not overwrite;
calibration preview (dry-pass rho -> +0.838) disclosed and kept DISTINCT from
the confirmatory pooled +0.6244; RQ3 double-null unspun (perf = retention
ceiling / no headroom; P2 = n=30 power limitation, not a false all-clear).
REPRODUCIBILITY AUDITOR: every figure matches the sealed records (rq2p3
5fdcb379, rq3 analysis e09c66ef, rq3 battery 5b61e461); the 0.93 calibration
AUC (kp30-vs-kp5 regime) kept distinct from the 0.587 confirmatory
agent-vs-baseline AUC (no HARKing); Ollama temp-0 cross-hardware caveat
present; both prereg SHAs (f22331a72e, 8db4e8a7) recomputed intact; blinding
integrity = each track filled once, post-seal.
One prose-clarity defect found and fixed: §5 Holm-7 ranks 6-7 report the
step-down monotone-enforced 0.511 (bare products 0.354 / 0.497); added a
one-clause note so it cannot read as an arithmetic slip. ZERO numbers changed.
Worktree-only on feat/sor-consent-relay; containment intact; $0/offline.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace the HELD-BLIND RQ3 Results/Discussion + authoritative-Holm-7
placeholders in docs/stage-07-companion-methods.md with real numbers read
from the sealed record only (rq3-confirmatory-analysis.json e09c66ef…):
RQ3-P1-perf -0.6pp CI[-1.58,+0.39]pp -> H0 (no +10pp agent margin)
RQ3-P1-latency -13.5ms CI[-52.1,+34.9]ms -> within 100ms budget
RQ3-P2 AUC 0.587 CI[0.458,0.703] -> fingerprint NOT excluded
RQ3-P3 = H0 (perf fails AND P2 fails)
Authoritative Holm-7 (family size 7, supersedes lead partial embedding, D3):
survivors RQ1-P1 / RQ2-P1 / RQ2-P3; non-survivors RQ1-P2 / RQ3-P2 /
RQ3-P1-perf / RQ3-P1-latency. RQ2-P3 slot carries the mechanism-corrected
primary p, not the lead degenerate p=1.
Calibration disclosure carried: the 0.93 RQ3 calibration AUC is regime
discrimination (churn kp30 vs kp5), distinct from the 0.587 confirmatory
selector-vs-selector fingerprint — no HARKing, detectors frozen.
RQ2-P3 half not re-litigated (kept as filled at e0d865d). Nulls reported
openly as results. Frozen lead prereg f22331a72e… and RQ2-P3 prereg
8db4e8a7… both intact. $0/offline, worktree-only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
RQ3 live-docker battery (90/90 runs) complete; un-blind and apply the frozen
prereg §6 plan. No re-specification.
Seal: SHA256SUMS over confirmatory-data/ (battery-results.json 5b61e461... + 90
rq3-run.json sidecars) — raw data immutable. New analyzer analysis/rq3_confirm.py
drives the frozen stats/metrics (untouched) via a run-level multi-arm bootstrap
mirroring two_sample_diff_ci (10k BCa, alpha=0.05). Effect+CI always, never bare p.
RQ3 result (honest null):
- RQ3-P1-perf: retention margin agent-max(static,random) = -0.6pp, CI
[-1.58,+0.39]pp -> FAILS +10pp gate (all selectors heal ~all churn, ~99%).
- RQ3-P1-latency: added-latency(agent-min-baseline) = -13.5ms, CI [-52.1,+34.9],
upper <= 100ms -> within budget (agent not slower).
- RQ3-P2: rebuild-classifier AUC(agent vs pooled baseline) = 0.587, CI
[0.458,0.703], upper 0.703 > 0.60 -> fingerprint NOT excluded (underpowered).
- RQ3-P3 = H0 (P1-perf fails and P2 fails).
Authoritative Holm-7 over the frozen size-7 family (supersedes the lead's
conservative partial embedding; RQ2-P3 slot carries the mechanism-corrected
H1-pooled Spearman p=0, not the lead degenerate p=1). Survivors: RQ1-P1,
RQ2-P1 (shrink), RQ2-P3 (mix). Non-survivors: RQ1-P2, RQ3-P2, RQ3-P1-perf,
RQ3-P1-latency. Sealed: analysis/rq3-confirmatory-analysis.json (e09c66ef...).
Tests: tests/test_sor_rq3_confirm.py 6 passed; full SOR suite 207 passed.
Both prereg SHAs intact; $0/offline analysis; worktree only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Post-seal, un-blinded fill of the RQ2-P3 mechanism track in the combined
companion paper (D3), matching the lead paper's post-seal discipline. RQ3
remains fully HELD-BLIND (battery still running; progress read by sidecar
count only).
- §3: FREEZE-PENDING -> FROZEN 2026-07-21 (prereg SHA 8db4e8a7...)
- §5 Results (RQ2-P3): H1 rho=+0.6244 CI[+0.5941,+0.6545]; H2 slope=+0.7052
CI[+0.6195,+0.7903] (n=270, run-level cluster bootstrap); H3 RESOLVED=MIX;
Holm own {H1,H2} both reject. Effect+CI, never bare p.
- §6 Discussion: shared-pool concentration RAISES anonymity (mix), refuting
the naive funnel; honest qualification/correction of the lead RQ2-P1
"shrink" as a unique-bridge (fresh-bridge-per-circuit) artifact, per
note-unique-bridge-artifact.md + frozen mechanism prereg. Mandatory
disclosure: §7 dry-pass previewed direction (rho 0->+0.838); hypotheses
were pre-committed two-sided.
RQ3 Results/Discussion and the authoritative Holm-7 stay HELD-BLIND until
rq3-battery-results.json lands. Lead RQ2-P1 + frozen lead prereg untouched.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Operator-authorized (D2: GO on the live RQ3 battery, grid permitting). Grid
re-probed (real SSH + docker subprocess probes): full (3 reachable, 1 isolated
engine host, 0 down), sor-hop image present. Green preflight: both RQ3
calibration gates already pass and a 1x1x2 live rehearsal delivered a real
docker circuit with a measured added-latency.
New launcher rq3_confirmatory_run.py is the RQ3 analogue of confirmatory_run.py,
gated behind the same triple-lock before any live circuit: (1) operator token
SOR_CONFIRMATORY_GO=1, (2) frozen LEAD prereg SHA-256 verified (RQ3 is in the
frozen family-of-7), (3) assert_isolated(docker != local). The RQ3-P1-latency DV
is a real end-to-end measurement — nothing is fabricated.
Battery launched and confirmed progressing (python child alive + a live
4-container circuit up): executor.run_rq3_battery(live=True) collecting the full
frozen schedule (selector ∈ {static,random,agent} × pinned churn kp30/steps20,
R=30 × C=50 = 4,500 isolated-docker circuits), cells NOT trimmed. Completion is
detected by the once-written rq3-battery-results.json; progress by the per-run
rq3-run.json sidecars reaching 90 (3 cells × 30 runs). Confirmatory data seals
under output/ (gitignored) — anchors force-added on completion.
Containment load-bearing: isolated-docker only, self-generated fixture traffic,
$0 (local heuristic/random selector arms; no frontier-model spend). Both prereg
SHAs intact (lead f22331a72e…, RQ2-P3 8db4e8a7ac60…). Worktree-only on
feat/sor-consent-relay.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Operator-authorized (D1: freeze + run). New harness analysis/rq2p3_confirm.py
runs the frozen §4 battery entirely offline: all 9 sweep cells (B∈{2,4,8} ×
alpha∈{0,1,2}; B=50 anchor excluded) × R=30 × C=50 = 13,500 bridged circuits.
RQ2 entropy is analytic and pool bridge assignment is deterministic from the
seed, so no engine, no traffic, no grid, $0 — the same offline reconstruction
path confirm_load_rq2 uses for the lead RQ2. Seeds per frozen §6: run =
derive_seed(S0=20260719 ‖ cell_id ‖ run_index), per-circuit = the live
executor's real _circuit_seed rule (byte-identical). The frozen instruments
(stats.py, confirm.py, confirm_load_rq2.py, assembler.py) are UNCHANGED — the
only new code is seed enumeration, run-as-unit grouping, and an OLS-slope helper.
Frozen §8 analysis (effect + BCa 95% CI, never a bare p; 10,000 resamples;
α=0.05; run-level cluster bootstrap):
H1 pooled Spearman ρ(c_i,H_i) = +0.6244, CI [+0.5941, +0.6545] -> mix
H2 OLS dose-response slope = +0.7052, CI [+0.6195, +0.7903] -> mix (n=270)
H3 joint -> RESOLVED = mix (both exclude 0, agree in sign)
Holm over own family {H1-pooled, H2-slope}, size 2: both reject.
Honest disclosure (as commanded): the §7 dry pass already previewed this mix
(ρ 0→+0.838); the battery quantifies an effect already visible at calibration;
the two-sided pre-commitment stood. The +ρ MIX qualifies/corrects the lead
RQ2-P1 "shrink" headline as a unique-bridge artifact — a shared finite bridge
pool RAISES the per-circuit anonymity set, the opposite of the naive funnel.
That correction is the finding, reported openly. The lead RQ2-P1 result is NOT
re-litigated; any pool ΔH is exploratory (prereg §1).
Sealed immutably under output/ (gitignored → anchors force-added):
rq2p3-confirmatory-results.json (SHA-256 5fdcb379d8a2…) + SHA256SUMS;
byte-identical on re-run. Both prereg SHAs verified intact (lead f22331a72e…,
RQ2-P3 8db4e8a7ac60…). Tests: tests/test_sor_rq2p3_confirm.py 7 passed; full
SOR suite 201 passed, no regression. Defensive-measurement instrument only;
worktree-only on feat/sor-consent-relay.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Operator-authorized freeze sign-off (Andre delegated per-gate command
authority to the overseer 2026-07-21). Froze rq2p3-mechanism-prereg.md
AS-IS — no hypothesis or parameter changed; the §7 re-word already landed
at 7ab3466. §10 records FROZEN 2026-07-21, delegated approval, and the
locked levels B∈{2,4,8}, alpha∈{0,1.0,2.0}.
SHA convention mirrors the lead prereg: full-file sha256sum stored in a
sidecar (rq2p3-mechanism-prereg.sha256 = 8db4e8a7ac60…), not embedded
inline, to avoid the self-referential fixpoint. Lead prereg SHA
f22331a72e… re-verified intact (recomputed — unmoved).
Defensive-measurement instrument only; no confirmatory data collected in
this stone. Worktree-only on feat/sor-consent-relay.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Bounded companion-methods QA (blinding / citation-grounding / FREEZE-PENDING all
clean). Prose-only clarification, no number changed: mark the stray raw p ~ 0 as
the lead paper's already-published RQ1-P1/RQ2-P1 result so it cannot be misread as
a held companion figure. HOLDING — awaiting operator freeze (RQ2-P3 §10) / GO
(RQ3 live grid).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Draft the companion METHODS data-independent and blind, mirroring the lead-paper
SS4 discipline: methods drafted, all Results/Discussion HELD-BLIND. No confirmatory
run, no confirmatory number.
- Two standalone methods sections (paper structure left OPEN — operator editorial
call): (A) RQ2-P3 funnelling-mechanism study, header marked FREEZE-PENDING since
its prereg §10 is unsigned — pool instrument, B×alpha manipulated-IV design,
two-sided H1/H2/H3, run-level cluster bootstrap, re-worded §7 gate (cites
stage-05-rq2p3-gate-clarification.md); (B) RQ3 companion from the frozen prereg
§3/§4 — qwen2.5:3b + reproducibility caveat, churn kp=30/steps=20, frozen
RQ3-P1-perf/latency/RQ3-P2 gates, Holm-7-supersedes note.
- Intro/Related-Work deltas (unique-bridge/mix mechanism; churn-resilience +
rebuild-fingerprint motivation), reusing only grounded lead citations plus
Barton2025 from the frozen prereg's RQ3 DV.
- The only numbers written are already-produced CALIBRATION gate values, each
labelled calibration not result (RQ2-P3 dry preview mix; RQ3 churn-bites +
classifier calibration). Results/Discussion are HELD-BLIND placeholders.
HARD HOLDS: RQ2-P3 awaits §10 human freeze; RQ3 awaits operator-GO + live grid
(added-latency is live-only). Frozen lead prereg f22331a72e… untouched; no HARKing;
containment synthetic/offline, $0; worktree-only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Wire the RQ3 churn-resilience companion (severable follow-on; RQ3 frozen in its
own prereg) as SYNTHETIC/OFFLINE build + calibration only — no confirmatory
battery run.
- battery: enumerate_rq3_cells()/rq3_schedule() add selector in {static(control),
random, agent} at 1house/bridge-off under pinned churn kp=30/steps=20, kept
SEPARATE so the frozen 6-cell lead lattice stays byte-identical.
- executor: run_rq3_cell_run collects offline selector DVs (throughput-retention,
drops/rebuilds, rebuild-interval gaps); added-latency is live-only (offline path
records None, never fabricated); run_rq3_battery(live=False) hard-raises.
- analysis/rq3_calibration: DRY offline gate. Churn-bites (1589 drops/1576
rebuilds) + rebuild-classifier calibration green with real teeth — churned(kp30)
vs baseline(kp5) AUC 0.926 separable, baseline-vs-baseline null 0.518 blind.
Classifier scored on the PER-RUN mean inter-rebuild gap (the confirmatory
grouping unit); the frozen instrument (rebuild_interval_gaps,
rebuild_classifier_auc) is untouched, no fit to confirmatory data.
HARD HOLD: no RQ3 confirmatory battery (live added-latency is operator+grid-gated).
Lead prereg SHA f22331a72e… untouched; containment intact (synthetic/offline, $0
local Ollama, frontier arm inert); worktree-only.
Tests: tests/test_sor_rq3_wiring.py 9 passed; full SOR suite 194 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Operator ruling (Andre, 2026-07-21): the STOP was correct and the diagnosis
right — §7 items 1-2 did not fail because the instrument is broken, but because
they were worded under the naive-funnel prior. Under the ratified posterior the
pool is a MIX (shared bridge → shared exit_signature → larger anonymity set →
higher H). Corrected pre-freeze, logged transparently (no HARKing).
- rq2p3-mechanism-prereg.md §7: item 1 re-anchored on the FROZEN bridge-federated
branch (injective fresh-bridge → m_i=1 → H≈0, constant c=1/C — a pool draws with
replacement so it cannot reproduce the zero-variance degeneracy); item 2 B=1 →
c=1.0 AND high H (maximal mix, "low H" refuted by construction); added a §7 scope
note (the gate validates the INSTRUMENT and must NOT pre-assert the H-vs-c sign —
that stays the two-sided confirmatory question H1/H2). Confirmatory hypotheses
unchanged.
- docs/stage-05-rq2p3-gate-clarification.md: dated deviation-rationale log (cites
the dry-pass output, no number invented; records the two-sided direction-agnostic
pre-commitment is untouched; the mix ρ 0→+0.838 is an exploratory finding that
strengthens note-unique-bridge-artifact.md).
- rq2p3_calibration.py: item-1 teeth moved to frozen_branch_regression() on the
untouched bridge-federated branch; item-2 now asserts c=1.0 AND high H; added
all_pass aggregate. confirm.py/stats.py/confirm_load_rq2.py untouched.
Re-run: all four gate items PASS with real teeth (all_pass=True). Synthetic/DRY
only, no confirmatory record read. FREEZE STAYS A HUMAN GATE: §10 empty, no
confirmatory RQ2-P3 cell until operator records the SHA-256 + signs off. Lead
prereg SHA f22331a72e... untouched; worktree-only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Build the funnelling-mechanism instrument for the RQ2-P3 follow-up: a new
`bridge-federated-pool` topology drawing each circuit's bridge from a finite
willing-bridge pool (size B) under Zipf willingness skew (alpha), so bridge
concentration genuinely varies as a manipulated IV. Defensive measurement only:
synthetic seeds, no engine, no traffic, no confirmatory record read.
- assembler.py: add `bridge-federated-pool` branch + zipf_weights/weighted_draw
helpers; existing lead `bridge-federated` branch left untouched and
bit-reproducible (fresh bridge per seed).
- battery.py: enumerate_rq2p3_cells() = B{2,4,8} x alpha{0,1,2} + anchor B=50,
alpha=0 (separate fn; frozen 6-cell lead lattice unmutated).
- analysis/rq2p3_calibration.py: DRY synthetic-only §7 calibration gate; no
confirmatory data read. confirm_load_rq2.py/confirm.py/stats.py unchanged.
- tests/test_sor_rq2p3_pool.py: 9 synthetic tests (helpers + pool reuse/variance
+ lead branch reproducibility).
Calibration gate RAN: items 3 (monotonicity) & 4 (entropy) PASS; items 1 & 2 do
NOT match their naive-funnel wording because the ratified posterior yields a MIX
(rho>0 across the whole sweep incl anchor rho=+0.838; B=1 -> H=2.54 bits HIGH) --
the note-unique-bridge-artifact.md prediction, surfaced early. HARD HOLD: no
confirmatory battery, no prereg freeze; NEEDS-OPERATOR banner raised for the §7
item-1/2 re-word decision. Lead prereg SHA f22331a72e... untouched; worktree-only.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
RQ2-P3 mechanism study (NEW prereg, approved-path/params-locked, not yet frozen):
pool-based willing-bridge instrument, concentration swept as a manipulated IV
(B∈{2,4,8}×α∈{0,1,2}), two-sided/direction-agnostic, run-level cluster bootstrap.
RQ3 companion run-brief (execution-freeze for the already-frozen RQ3): agent=qwen2.5:3b
local Ollama, churn kill_prob=30/steps=20, baselines + Holm-7 completion pinned.
note-unique-bridge-artifact.md: on-record analysis that the lead RQ2-P1 shrink is plausibly
a fresh-bridge-per-circuit artifact (unique exit-signature -> singleton set -> H~0), and that
realistic bridge reuse (a mix) may raise anonymity -> the mechanism study may qualify the lead
RQ2 headline. No edit to the frozen prereg or committed lead paper.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Self red-team of the SS4-filled lead paper against the sealed SS3 results
(no new statistic, no raw-data touch, frozen prereg untouched). Documented in
docs/stage-08-adversarial-review.md.
PASS: number fidelity (every §5 figure matches stage06-results.json), Holm
arithmetic (family=7, multipliers 7/6/5/4, only RQ1-P1+RQ2-P1 survive @ .05),
spin/over-claim (no-leak framing for the below-chance AUC, honest-negative
reported with equal prominence, lab-scope disclaimers), blinding integrity
(§5-6 filled once post-seal), method faithfulness (bit-faithful --verify).
Defect fixed (prose only, ZERO numbers changed): RQ1-P1's marginally
below-chance AUC was incorrectly attributed to a bridge "padding stream
flattening" the profile. assembler.py sets padding_applied=(bridge==
"on+padding"), so the RQ1-P1 bridge-on NO-PAD arm carries no cover stream.
Corrected docs/stage-06-analysis.md and paper §5.2/§6 to state the offset is
an unexplained pooled-correlator artifact, explicitly not a padding effect.
Point estimate, CI, gate decision, and Holm outcome are unchanged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
One auditable, deterministic run of the four lead-paper §6 tests on the
frozen raw battery. Calibration gate PASS (linked AUC 1.0000, unlinked
0.5036 on §5 fixtures) → AUCs reportable.
- RQ1-P1 (bridge leak): AUC 0.4660 CI[0.4523,0.4798] = anomaly-below-chance.
NO measurable leak; Holm-sig but in the WRONG direction (not evidence of
linkability).
- RQ2-P1 (federation): ΔH −0.9587 bits CI[−1.056,−0.864] = SHRINK. Holm-sig
HONEST-NEGATIVE — federation reduces the anonymity set (opposite of RQ2's
motivation); reported with equal prominence per §6 two-sided framing.
- RQ1-P2 (padding): ΔAUC +0.0113 CI[−0.0025,+0.0234] padding-ineffective
(raw p 0.091, Holm adj-p 0.456, not rejected).
- RQ2-P3 (funnelling): Spearman ρ=0, zero-variance concentration =
as-instrumented degeneracy flagged in advance; inconclusive.
Holm family=7 report-4 (multipliers 7,6,5,4); only RQ1-P1 + RQ2-P1 survive.
Ran ONLY the pre-registered tests — no re-slicing, no post-hoc subgroups.
Method faithfulness: RQ1-P1/P2 CIs use a bit-faithful fast numpy bootstrap
(frozen O(n²)-per-fold jackknife intractable at n=75000); stage06_run.py
--verify proves both == frozen stats.bootstrap_ci bit-for-bit (point/lo/hi/
method within 1e-12, incl the average-rank tie path). RQ2-P1/P3 stay on the
frozen confirm.* paths (tractable). Frozen prereg SHA INTACT.
Artifacts: docs/stage-06-analysis.md (methods-faithful narrative, Holm table,
CONFIRMATORY/EXPLORATORY labels, nulls reported honestly) + force-added
output/…/analysis/stage06-results.json (regenerable via stage06_run).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Pairing ruling (docs/stage-05-rq1p2-pairing-clarification.md, RATIFIED-by-
derivation, made while BLIND): the frozen §6 L197-198 "paired" bootstrap pairs
the bridge-on (no-pad) and bridge-on+padding arms BY RUN INDEX — the only
balanced pairing the R=30 (§4 L111) interleaved (§3 L97) design supports; the
§4 L108 RQ1 unit is AUC over the (entry,exit) pair set, so the pairable unit is
the run, not the circuit. Prereg SHA untouched (f22331a72e…).
confirm_load.collect_rq1_p2_paired feeds one confirm.PairedCircuit per shared
run_index into the FROZEN confirm.rq1_p2_padding (§6 arm-level ΔAUC preserved);
per_run_delta_aucs exposes the auditable per-run ΔAUC_i; missing-arm indices are
unpaired (frozen R not redefined). Unit-tested on SYNTHETIC pcaps only
(prereg §2 blinding — unblocks CODE not RESULTS); full SOR suite 176 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Flip docs/stage-05-rq2-posterior-clarification.md PROPOSED->RATIFIED (operator
Andre, 2026-07-20, ratified while BLIND to RQ2 data — honest prereg-completion of
a gap the freeze left, not HARKing). Substance of the construction unchanged.
Add cmd_chat/sor/analysis/confirm_load_rq2.py: reconstructs each circuit's spec
OFFLINE from the persisted per_circuit_seeds via deterministic assemble(), derives
the observation-consistent anonymity set A_i (uniform / max-entropy -> [1]*m_i),
per-circuit Miller-Madow H_i, and the willing-bridge concentration series — the
exact inputs confirm.rq2_p1_delta_h (ΔH, two-sided) and rq2_p3_funnel (Spearman ρ)
consume. Grounded only in the ratified rule ([Serjantov2002]/[Diaz2002]).
BLINDING preserved (prereg §2): ratification unblocks CODE, not results. The
collect_* real-data entrypoints are BLIND-GATED and NOT run; the reconstruction is
unit-tested on SYNTHETIC specs/seeds only (test_sor_confirm_load_rq2.py, 6 passed;
full SOR suite 172 passed). Frozen prereg untouched; no RQ2 statistic computed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
docs/stage-07-paper-draft.md — Abstract skeleton, Intro (consent-gate novelty
G4 + RQ1 bridge-leak / RQ2 funnelling motivation), Related Work, full Methods
from the frozen prereg §2-§6, Limitations §7, References. Results/Discussion are
HELD-BLIND placeholders (battery still running, zero confirmatory numbers seen);
RQ2 subsections additionally gated on the pending posterior-construction
ratification. Frozen prereg untouched; no RQ2 statistic computed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Overseer ruled DIRECTION=A: fill the specification gap the frozen prereg
left on RQ2's per-circuit adversary sender-posterior, rather than
silently redefine the DV (B) or relabel a confirmatory test (C).
Proposes: adversary observes a circuit's exit-bridge/house and forms a
uniform posterior over the observation-consistent anonymity set A_i;
per-circuit H_i = Miller-Madow(log2 m_i), with S=2^H [Serjantov2002] and
d=H/log2 N [Diaz2002] as the cited normalizations. Makes RQ2-P1 (two-
sided ΔH) and RQ2-P3 (Spearman) computable OFFLINE from the immutable raw
records (per_circuit_seeds + deterministic assemble()).
Status PROPOSED — operator ratification pending. Pre-specified BLIND
(2026-07-20, battery running, zero RQ2 statistics computed) so this is
honest pre-registration completion, not HARKing. Frozen prereg untouched;
sign not presumed (honest-shrink reported with equal prominence).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
confirm_load.py reconstructs the RQ1 (entry,exit) circuit-pair set from
the REAL per-hop pcaps (bin_pcap_bytes -> score_matrix; diagonal=linked,
off-diagonal=unlinked) so confirm.rq1_p1_leak scores measured data — the
frozen §4 unit. classify_run reads persisted metrics.json;
collect_rq1_p1_pairs selects only bridge=on no-pad runs.
Plumbing unit-tested on SYNTHETIC pcaps only (prereg §2 blinding — no
inspecting intermediate confirmatory results before the battery
completes). RQ1-P2 nopad/pad pairing + all RQ2 held for the operator
ruling (NEEDS-OPERATOR: RQ2 per-circuit posterior underspecified).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Frozen §3/§4/§6 pin RQ2 DV as a per-circuit adversary sender-posterior
entropy (Miller-Madow, bootstrap over circuits; RQ2-P3 needs per-circuit
H). The committed executor measures a per-run realized-frequency entropy
instead — no per-circuit posterior, no per-circuit H. No formula for the
posterior construction exists in the frozen prereg or the instrument.
Flagged rather than invent a posterior model post-freeze (HARKing/
fabrication on a confirmatory DV). Battery left running (raw records
recoverable → RQ2 recomputable offline under any ruling). RQ1 proceeds.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Measures RQ1 bridge-correlation AUC + RQ2 Shannon entropy from REAL
per-hop pcaps of live isolated-docker circuits. run_battery(live=False)
hard-raises and the executor refuses (ExecutorError) any DV it did not
measure — the executor-side twin of the launcher's "never fabricate
cells" guard. Wired into confirmatory_run's triple-locked tokened GO
(executor.run_battery(live=True) behind operator token + verified frozen
prereg SHA + engine!=local + green preflight + full grid).
Containment intact: self-fixture bytes only, isolated engine only, no
external target. Frozen prereg untouched (SHA f22331a72e…).
Verify: test_sor_executor.py 7 + confirmatory_run + full SOR suite = 163
passed; preflight green (grid 3/3, READY); no confirmatory data collected.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Wire the live per-cell condition assembler (cmd_chat/sor/assembler.py): compose
the existing R1/R4/R5/R6 pieces into a condition-encoding CircuitSpec per frozen
§2 cell, so a cell_id maps to a genuinely distinct, ForwarderPlan-gated
(engine != local) isolated circuit rather than the plain 3-hop control:
- RQ1 bridge-on / on+padding insert a live bridge hop (+ R1 PADDING stream);
- RQ2 bridge-/directory-federated genuinely span >= 2 houses
(federation.select_federated_path, split-knowledge).
Plans only: opens no socket, moves no traffic, stands up no engine.
battery.assembler_dry_check validates on FIXTURES (6/6 cells distinct
fingerprints, all isolation-gated, same (cell,seed) reproduces, RQ1 bridge +
padding arms live, RQ2 federation >= 2 houses); write-once, kept out of any
confirmatory data dir. confirmatory_run preflight now runs it; the --operator-go
path no longer refuses cells as "not wired" — triple-lock + green preflight HOLD
on grid completion, then surface the operator's immutable data-run gate. No
confirmatory cell is ever fabricated.
Start-line instrument-validation gate (cmd_chat/sor/gate.py), grid + containment
pin (grid.py), and the guarded launcher (confirmatory_run.py) land alongside.
Holm: hold family_size = 7, report the 4 RQ1/RQ2 tests -> multipliers 7,6,5,4.
This is the pre-registration, not a deviation (prereg §6 [APPROVAL] size-7 family
+ D6 disclosure in the lead paper); docs/stage-05-holm-clarification.md marked
RATIFIED. RQ3-fold and alpha-split declined.
Containment intact; prereg untouched (SHA f22331a72e...); worktree-only. output/
gitignored (runtime artifacts, never source). Full SOR suite 156 passed. No
confirmatory data collected — the live data run remains the human gate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Enumerates the frozen §2 confirmatory cells (3 RQ1 bridge levels + 3 RQ2 topology
levels at their controls; bridge-off+padding recorded as a declared N/A, not run),
derives each run's seed by the frozen §4 rule (SHA256('<S0>|<cell_id>|<run_index>')
big-endian first 8 bytes -> u64, S0=20260719), and lays out the §2 schedule with
run order randomized within each cell and each RQ's control interleaved before and
after its treatments.
write_cell_plan emits the auditable cell x run plan artifact (R=30, C=50 -> 180
runs / 9000 circuits, seed rule, matched-N rule, full schedule). dry_pass exercises
the R2/R3 provenance pipeline on FIXTURES only (deterministic replay_and_seal, no
engine, no traffic) to prove schema-valid + checksummed provenance and that a seed
reproduces its circuit-build sequence. Collects no confirmatory data — that remains
the human gate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Encodes the four lead-paper confirmatory tests exactly per the frozen prereg §6,
calibrated on synthetic ground truth only (never re-fit to cell data):
- RQ1-P1 leak: correlation AUC + BCa CI; gate = CI excludes 0.5 (D2); 0.60 CI
lower bound = separate materiality label (material / weak-but-real), not the gate.
- RQ1-P2 padding: paired ΔAUC = AUC(no-pad) − AUC(pad) over circuits; effective
iff CI > 0.
- RQ2-P1 anonymity set: ΔH = H(federated) − H(single, matched N) with Miller-Madow
per-circuit entropy; two-sided grow / honest-shrink / inconclusive by CI sign.
- RQ2-P3 mechanism: Spearman ρ(top-3 bridge concentration, per-circuit H) + CI.
apply_holm corrects the reported RQ1/RQ2 subset against the full frozen family of
7 (family_size default 7) — never re-optimised to the 4 reported. Bootstrap
p-values order the Holm step-down only; every decision is a CI gate, never a bare p.
stats.py: bootstrap CIs can now return their resample distribution so the Holm
ordering p-value comes from the same resamples as the CI.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Turnkey implementation of the frozen prereg §6 analysis plan, written before any
confirmatory data exists (analysis-precedes-data, rigor-standards §Statistics).
Pure stdlib, no I/O, no engine, no traffic — calibrated on synthetic ground truth only:
- bootstrap_ci / two_sample_diff_ci: BCa 95% CIs (10k resamples default) with a
percentile fallback when bias/acceleration terms are degenerate; seeded and
reproducible (CIResult carries method+seed for the §6 three-seed spot-check).
- miller_madow_entropy_bits: plug-in Shannon entropy + Miller-Madow bias
correction (§3 estimator).
- spearman: rank correlation for RQ2-P3.
- holm_bonferroni: step-down multiplicity correction with explicit family_size so
a lead paper reporting a subset of the frozen 7-test family still corrects
against the full family (never re-optimised to the reported subset).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds the RQ3 agent-selector backend as a reproducible, offline measurement
decision — not relay/data-plane traffic. Containment is untouched: the model
query is a local call to localhost:11434.
- selector.py: SelectorPolicy ABC seam; static/random/agent(heuristic) built-ins;
make_policy() resolves (strategy, agent_backend); run_selection writes a
write-once selector.json provenance sidecar.
- agent_selector.py: OllamaAgentPolicy — reproducible confirmatory arm
(temp=0 + per-run seed, decisions cached keyed by (seed, state-hash),
deterministic heuristic fallback on any model/parse failure, model weights
digest pinned for the manifest). ClaudeExploratoryPolicy — EXPLORATORY-only
stub: makes no paid call, refuses to serve as a confirmatory backend.
- provenance.py: additive selector_backend field so an agent-selected run pins
the exact model id + weights digest.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
churn.py emits a seed-deterministic kill/spawn schedule (pure data — no real VM
spin/kill; the live fabric half stays gated by the containment law). selector.py
consumes the schedule and rebuilds a circuit whenever a kill drops one of its
hops, across static | random | agent strategies; the paid frontier-model agent
arm (GOAL envelope (c)) is human-gated and NOT wired — the agent strategy here is
a local stability heuristic that spends nothing. analysis/metrics.py aggregates
the four DV families (RQ1 correlation AUC, RQ2 entropy bits, RQ3 throughput
retention + rebuild-classifier AUC) into a schema-valid, write-once metrics.json.
Acceptance check green: under a fixed churn seed the selector rebuilds every
dropped circuit (every_drop_rebuilt, all strategies) and metrics.json is
produced. Python R7 selector suite 11 passed; full SOR suite 92 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
federation.py adds both federation modes as offline-verifiable logic (no socket,
no engine, no external target — this measures a trust model, it provides
anonymity to no one). directory-federation: a signature-gated HOUSE-PEER roster
(PeerRoster/build_peer_frame/parse_peer_frame, Ed25519-signed, forged/unsigned
rejected) merged into a pubkey->house Directory, with select_federated_path
drawing a seed-deterministic circuit that spans >=2 houses and refuses to
collapse to one — so no single house's logs hold every hop identity (RQ2 split
knowledge). bridge-member: BlindBridge holds no room key and relays only opaque
SOR tunnel bytes verbatim (chat/consent/unknown frames refused, plaintext never
read), emitting metadata-only bridge_forward R3 events. Acceptance check green:
roster signature-gated, path spans two houses, bridge is plaintext-blind. Python
federation suite 10 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
run_circuit_fixture stands up an isolated docker N-hop nested-SSH chain, pipes a
seed-deterministic self-generated payload through it, verifies end-to-end
delivery, captures + checksums a per-hop pcap, emits R3 events, and always tears
the circuit down. Containment is load-bearing: assert_isolated refuses
local/unknown engines up front, every hop is built through the ForwarderPlan
guard, only our own fixture bytes move between our own containers, and the runner
is wired for docker only (multipass e2e refused, not faked). Adds the sor-hop
fixture image (alpine + sshd + tcpdump, lab-relay only) and a docker-gated
acceptance test that skips where no daemon/image is present. Instrument-validation
gate item 1 GREEN: live 3-hop delivery verified, 3 distinct hops, pcaps checksum.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Land only the offline-validatable half of R7 (instrument-validation gate items
3 and 4): the detectors the gate calibrates on synthetic fixtures, never on
confirmatory-cell data. Pure stdlib functions over in-memory series — no pcaps,
no engine, no traffic, no VM fabric.
- analysis/detectors.py: shannon_entropy_bits (RQ2 anonymity-set entropy),
pearson/score_matrix/auc/linkage_auc, bridge_correlation_auc (RQ1
linkability scorer), and synthetic_bridge_fixture (seed-deterministic
known-linked / known-unlinked ground truth via the R1 SorRng).
Calibration green: entropy returns exactly log2(N) for N equiprobable senders
(gate item 4); a known-linked control pair scores AUC=1.0 and the unlinked
estimator is unbiased at chance (ensemble mean over 40 seeds = 0.498 ~ 0.5,
gate item 3). Detectors are calibrated on synthetic fixtures only — no fitting.
The traffic-moving R7 pieces (churn.py VM spin/kill, live selector rebuild loop,
metrics.json emission) are HELD for R4/R6 + a live grid and are absent here.
Python SOR suite 66 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Land only the containment-refusal half of R4: the guard every SOR forwarder
must pass before it does anything — assert engine != local, or refuse. It moves
no traffic, opens no socket, spawns no engine, and builds no SSH chain.
- forwarder.py: assert_isolated (raises on local and on any engine off the
manifest allow-list), isolation_prefix (docker/multipass exec argv, no host
path), ForwarderPlan (construction IS the gate — no plan for local/unknown/
no-container). Allow-list imported from provenance so forwarder and manifest
cannot disagree about what counts as isolated.
- There is no code path in this module that returns an exec prefix for the host;
containment is structural, not a runtime flag.
The traffic-moving half of R4 (gate item 1: 3-hop self-traffic e2e delivery +
per-hop pcaps across the grid) is deliberately NOT implemented — HELD for a live
isolated engine + grid. This guard is fully verifiable offline.
Acceptance (gate item 5) green: local/unknown/no-container refused; isolated
engine accepted with the correct isolation prefix. Python forwarder suite 12
passed; full SOR Python suite 49 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add the in-band, signature-gated consent handshake that governs whether a
node may be recruited into a measured SOR circuit, plus a per-recipient
sealed box so a hop credential decrypts only with the host's X25519 key.
- crypto.rs: seal_to_pubkey/open_sealed (ephemeral X25519 -> HKDF-SHA256 ->
Fernet AEAD), reusing the existing fernet crate (no new symmetric primitive).
- sor/consent.rs: signed ConsentRequest (Ed25519 persona), node_evaluate
(reject unless the signature verifies), CircuitBuilder that recruits a hop
only on an accept it can open; parse_sor_frame with never-panic proptests.
- net.rs: {"_sor":...} control frames are live-only and classified out-of-band
by the SOR layer (parse_sor), never surfaced as chat/app events.
- cmd_chat/sor/consent.py + bridge _sor case: bit-compatible Python mirror;
the agent only observes/classifies consent frames — it never auto-accepts
and stands up no forwarder (forwarding is R4, isolated-engine-only).
Acceptance check (roadmap R5) green both languages: host recruits a hop only
after an explicit accept; a reject leaves no entry; the credential opens only
with the host key (third party cannot); unsigned/forged request -> rejected.
Cross-language KDF parity pinned by a shared known-answer vector.
Rust 75 passed; Python SOR suite 37 passed.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds the tamper-evident event stream that makes every DV auditable. Each
measurable moment is appended as one JSON object per line to an append-only
output/sor-runs/<ts>/events.jsonl over the closed vocabulary {consent_request,
consent_accept, consent_reject, circuit_build, hop_add, bridge_forward,
churn_kill, churn_spawn, rebuild_start, rebuild_done}; on close the file is
SHA-256'd and the digest is sealed exactly once into manifest.json (the single
sanctioned None->hash completion of the manifest).
Records carry only metadata (fingerprints, byte counts, latencies, decisions),
never message plaintext, so the zero-knowledge relay property is preserved. The
deterministic replay_fixture_circuit emits a seed-driven event stream with no
engine, socket, or traffic (containment-safe) and is the basis of
instrument-validation gate item 6.
- cmd_chat/sor/events.py: EventLog (append-only, per-record schema),
replay_fixture_circuit, replay_and_seal.
- cmd_chat/sor/provenance.py: seal_manifest (seal-once events SHA + stop ts).
- tests: replay -> schema-valid JSONL whose SHA matches the manifest;
append-only (original prefix byte-identical after further appends);
seed-determinism modulo timestamps; seal-once immutability.
Live emit wiring into the _sor handlers lands with R4/R5.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
write_manifest() freezes at circuit-experiment start everything needed to
reproduce and audit a SOR run and writes an immutable
output/sor-runs/<ts>/manifest.json: the R1 --sor-seed, topology/selector/churn
schedule id, one persona fingerprint per participating node, worktree git SHA,
isolated engine kind + image digest, pip/cargo dependency freeze, and
start/stop timestamps. The events.jsonl SHA-256 field is reserved for R3.
Provenance only: no forwarder, no engine spawned. Defense-in-depth to keep
containment load-bearing, the schema refuses to record a non-isolated
("local") engine at run or node level; the binding assertion still lands with
the R4 forwarder.
- cmd_chat/sor/provenance.py: Node/RunManifest, node_fingerprint (bit-for-bit
mirror of persona.rs::fingerprint_of), git/deps capture, explicit
dependency-free schema validator.
- hh/src/persona.rs: fingerprint_of known-vector parity test.
- tests: R2 acceptance predicate (manifest present + schema-valid, non-empty
seed, >=1 fingerprint/node), immutability, containment rejection of local
engines, and cross-language fingerprint parity.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds the single stochasticity source for the SOR measurement instrument so
every later stochastic decision (path selection R4, churn/selector R7, padding
jitter R4) is driven by one --sor-seed and is byte-reproducible run-to-run —
the R1 acceptance check and instrument-validation gate item 2.
- hh/src/sor/mod.rs: SplitMix64 + SHA-256 domain-separated sub-streams +
unbiased Lemire bounded sampling -> deterministic select_path/bringup.
Pure bookkeeping: no forwarder, no socket, no engine. Isolated-engine
containment assertions land with the R4 forwarder.
- hh/src/main.rs: `sor-bringup --sor-seed` emits the deterministic
circuit-build sequence as JSON (observable reproducibility).
- cmd_chat/sor/config.py: bit-for-bit Python mirror so the seed means the
same thing on both sides.
- Tests (Rust proptest+unit, Python pytest): SplitMix64 known-answer vector,
same-seed determinism, seed divergence, bounded/no-panic, and a shared
cross-language parity vector asserted identically in both suites.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Integrate the laptop feature branch (multi-language capability benchmark,
model picker, and ESA-style pseudonymous file attribution) onto the current
church main (native AI harness, podman/vbox sandboxes, † theme). Only conflict
was the help/status line — kept the newer podman grammar + /help, spliced in
/export-signed, and aligned attribution strings to the † sigil.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adapt Princess_Pi's Encrypt-Share-Attribution scheme so shared files are
anonymous but provably attributable.
A. In-session: a persistent Ed25519 persona key (~/.config/hack-house/
persona_ed25519) signs every /send and /sendroom offer over the content
hash; receivers verify it and see the persona fingerprint. Optional
--attest <passphrase> attaches a revealable SHA-512(pass||sha256)
commitment. Additive JSON — wire-compatible with the Python client.
B. Portable: /export-signed <dir> packages a directory into Princess_Pi's
exact ESA 7z (fresh per-round Ed25519 sig over an inner 7z, SHA-512
checksums, bundled verify scripts). Builder embedded from
hh/tools/esa/esa_build.sh; verifiable with just bash+7z+ssh-keygen.
Tests: 45 pass (3 new persona). ESA archive build + verify-everything.sh +
passphrase reveal verified end-to-end.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add bench-lang.py + bench/ package: a third benchmark axis answering
"which open-source model is best for my workflow?" across Python,
JavaScript, Go, Rust and Bash.
- MultiPL-E (Go/Rust/JS/Bash) + original HumanEval (Python), loaded via
the HF datasets-server REST API with on-disk cache — no datasets/
pyarrow dependency.
- Completions go straight to Ollama /api/generate with raw=True so
instruct models continue the code instead of replying with prose.
- Code runs in rootless, network-less podman (safe default) with a
host-toolchain fallback; pass@1/pass@k via the HumanEval estimator.
- run/score separation: results persist to a scorecard JSON, then
`pick --workflow ops` re-ranks without re-running any model.
- Extensible: a new language is one Lang entry; a new workflow is one
block in workflows.json.
Also fix a --runs grant-persistence bug in bench-sandbox.py: the grant
leaked across runs, invalidating the L0-nogrant refusal test on runs 2+.
Each run now revokes the ACL and starts ungranted.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
bench-ai.py drives the /ai chat path end-to-end (SRP -> Fernet -> WebSocket
-> provider), measuring TTFT/total/tok-s per model, with a --direct provider
mode that isolates raw model throughput from event-loop contention.
bench-sandbox.py benchmarks the /ai <agent> !<task> sandbox code path by
playing the room owner on the zero-knowledge relay: it grants drive, sends
graded tasks (L0 no-grant refusal, file/script/logic/multistep, destructive
gating+confirm, blast-radius cap), captures the agent's injected _sbx:input
frames, and grades correctness behind the same destructive guard. Includes
REPL-aware exec replay (folds python3 sessions into heredocs), prompt-prefix
stripping, model-fail vs replay-limit tagging, --runs N averaging with
per-step timing, and auto-bumped timeouts for reasoning models.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
So the owner never has to name each model. Tracks an ai_agents set
(populated from `_ai` typing/stream frames and the "(ai) online" announce,
pruned on leave); `/grant ai` intersects it with the live roster and grants
all in one ACL broadcast. Help text gains a /grant ai row.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>