Case study · shadow-mode pilot

12 real bugs, 3 fixers, 36 receipts.

We ran the full pilot process on bugs that maintainers of three open-source Python libraries actually fixed. Two AI agents got only the bug report. ResultBond judged every fix with hidden tests — and an independent check, the maintainers' own test suite, graded us.

97.2%agreed with the maintainers' tests36 verdicts
0good fixes blamedcorrect code never marked FAIL
1miss, explained belowa requirement the bug report never stated
36/36receipts verifiedEd25519, checkable at /verify
4.8smedian time to receiptthree replays per fix

How the pilot ran

  1. 1

    12 tasks from real history

    Each task is a commit where maintainers of toolz, boltons or more-itertools fixed a bug. The repository is frozen just before the fix.

  2. 2

    Hidden tests, proven first

    Black-box tests rewritten from the maintainers' own tests. The pilot tool refused to start until each hidden test failed on the buggy code — it rejected three of our first drafts.

  3. 3

    Three fixers per bug

    The maintainers' real fix, and two AI agents of different strength that saw only a one-paragraph bug report. 36 fixes in total.

  4. 4

    An independent grader

    Every fix was also run against the maintainers' complete test module from the fix commit — hundreds of tests ResultBond never used to decide. That is the "reviewer" our verdicts are compared with.

Every fix

Verdicts from the final round. 11 of agent A's and 10 of agent B's 12 fixes were accepted; every maintainer fix passed.

12/12Maintainers' fixfixes accepted
11/12Agent Afixes accepted
10/12Agent Bfixes accepted
TaskMaintainers' fixAgent AAgent B
boltons-guiderator-size
GUIDerator size limits are wrong
PASSPASSPASS
boltons-histogram-iqr
Histograms fail when the interquartile range is zero
PASSPASSPASS
boltons-pearson-zero
Stats.pearson_type fails on a zero denominator
PASSPASSPASS
boltons-relative-time-tz
relative_time() breaks on timezone-aware datetimes
PASSPASSPASS
mi-bucket-miss
bucket lookups invent keys
PASSFAILbroke existing testsFAILbroke existing tests
mi-ichunked-zero
ichunked() misbehaves for n=0
PASSPASSindependent check: FAILFAILbroke existing tests
mi-one-falsy-exception
one() and only() ignore falsy custom exceptions
PASSPASSPASS
mi-running-window
running_* window sizes are not validated
PASSPASSPASS
toolz-interpose-empty
interpose() raises on an empty sequence
PASSPASSPASS
toolz-reduceby-init
reduceby() should accept a callable init
PASSPASSPASS
toolz-topk-stable
topk() is not stable when keys tie
PASSPASSPASS
toolz-topk-tuple
topk() should return a tuple with its own implementation
PASSPASSPASS

What the pilot taught us

Limits

The final round as a client would see it: the pilot report, with every receipt one click from verification. Tasks, patches and both rounds' raw data are available to pilot partners on request.

Run this on your own bugs.

Four weeks, 30–50 real bugs, your agent of choice. Free in shadow mode.