Replay benchmark
Every case runs end to end: signed manifest, isolated sandbox, hidden tests, signed receipt. We publish where it refuses — and where it is still wrong.
By category
| Category | Right outcome |
|---|
Every case
| Case | What it tests | Expected | Observed | Match |
|---|
What changed after the first run
- Report forgery closed for black-box tests. With ordinary tests, code under test shares the pytest process and can forge the whole result (
shop/forge-exit, kept as a false PASS on purpose). Black-box packs run the candidate in a separate process with no access to the report channel; the same attack, and two stronger ones, now FAIL (shop-bb/*). - Three candidate replays. A FAIL needs every replay to fail, so a test that fails 30% of the time at random now blames correct code in about 2.7% of jobs instead of 9%.
- Limits broken in any replay = FAIL. A fork bomb that was contained in some replays used to read as “not reproducible”.
- Real repositories. Upstream bug fixes from toolz, boltons and more-itertools, each verified to fail before and pass after the fix. Small, well-tested libraries only — larger repositories come next.