test(ai): ground-truth benchmark for the native harness + first baseline
Add bench/ — a 4-category × easy/medium/hard task matrix (shell, code, git, multi) and a runner that drives the live TUI over tmux and grades each task by a `podman exec` verify snippet (exit 0 == PASS), never by the model's self-reported summary (which the weak CPU model fabricates). Tasks run in the agent's real cwd with bare filenames so the suite measures task completion, not the model's absolute-path discipline. Completion is detected off the viewport- independent `is thinking…` footer (the TUI is full-screen, so capture-pane scrollback is not chat history). First baseline (qwen2.5:3b): 2/12 PASS, high variance. Surfaces the next harness-addressable improvements — `<native>` tag leakage and bare tool-call-as-text — now measurable against this suite. Findings doc updated. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -31,6 +31,9 @@ err.log
|
||||
# Heavy / superseded demo-build kit (real demos live in video-toolkit)
|
||||
/docs/demo/
|
||||
|
||||
# Benchmark runtime results (commit a curated baseline with `git add -f`)
|
||||
/bench/results/
|
||||
|
||||
# Project scratch
|
||||
/i-try/
|
||||
test_rsa.py
|
||||
|
||||
Reference in New Issue
Block a user