c5715ba2e3
bench-ai.py drives the /ai chat path end-to-end (SRP -> Fernet -> WebSocket -> provider), measuring TTFT/total/tok-s per model, with a --direct provider mode that isolates raw model throughput from event-loop contention. bench-sandbox.py benchmarks the /ai <agent> !<task> sandbox code path by playing the room owner on the zero-knowledge relay: it grants drive, sends graded tasks (L0 no-grant refusal, file/script/logic/multistep, destructive gating+confirm, blast-radius cap), captures the agent's injected _sbx:input frames, and grades correctness behind the same destructive guard. Includes REPL-aware exec replay (folds python3 sessions into heredocs), prompt-prefix stripping, model-fail vs replay-limit tagging, --runs N averaging with per-step timing, and auto-bumped timeouts for reasoning models. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>