Case study · shadow-mode pilot
12 real bugs, 3 fixers, 36 receipts.
We ran the full pilot process on bugs that maintainers of three open-source Python libraries actually fixed. Two AI agents got only the bug report. ResultBond judged every fix with hidden tests — and an independent check, the maintainers' own test suite, graded us.
How the pilot ran
- 1
12 tasks from real history
Each task is a commit where maintainers of toolz, boltons or more-itertools fixed a bug. The repository is frozen just before the fix.
- 2
Hidden tests, proven first
Black-box tests rewritten from the maintainers' own tests. The pilot tool refused to start until each hidden test failed on the buggy code — it rejected three of our first drafts.
- 3
Three fixers per bug
The maintainers' real fix, and two AI agents of different strength that saw only a one-paragraph bug report. 36 fixes in total.
- 4
An independent grader
Every fix was also run against the maintainers' complete test module from the fix commit — hundreds of tests ResultBond never used to decide. That is the "reviewer" our verdicts are compared with.
Every fix
Verdicts from the final round. 11 of agent A's and 10 of agent B's 12 fixes were accepted; every maintainer fix passed.
| Task | Maintainers' fix | Agent A | Agent B |
|---|---|---|---|
| boltons-guiderator-size GUIDerator size limits are wrong | PASS | PASS | PASS |
| boltons-histogram-iqr Histograms fail when the interquartile range is zero | PASS | PASS | PASS |
| boltons-pearson-zero Stats.pearson_type fails on a zero denominator | PASS | PASS | PASS |
| boltons-relative-time-tz relative_time() breaks on timezone-aware datetimes | PASS | PASS | PASS |
| mi-bucket-miss bucket lookups invent keys | PASS | FAILbroke existing tests | FAILbroke existing tests |
| mi-ichunked-zero ichunked() misbehaves for n=0 | PASS | PASSindependent check: FAIL | FAILbroke existing tests |
| mi-one-falsy-exception one() and only() ignore falsy custom exceptions | PASS | PASS | PASS |
| mi-running-window running_* window sizes are not validated | PASS | PASS | PASS |
| toolz-interpose-empty interpose() raises on an empty sequence | PASS | PASS | PASS |
| toolz-reduceby-init reduceby() should accept a callable init | PASS | PASS | PASS |
| toolz-topk-stable topk() is not stable when keys tie | PASS | PASS | PASS |
| toolz-topk-tuple topk() should return a tuple with its own implementation | PASS | PASS | PASS |
What the pilot taught us
- Round 1 missed two broken fixes. Both agents fixed the
bucketbug but quietly broke another behaviour: a key disappeared after being read. Our hand-written regression test didn't cover it, so we said PASS. Round 1 ended at 88.9% agreement with 4 false PASS: these two, plus the two explained in the third point. The independent check caught them. - We fixed the tool, not the numbers. Pilots now run the client's own test modules as a regression check inside the sealed worker (
--repo-tests). Round 2, same fixes: both broken fixes FAIL with “broke existing tests”, and one more regression surfaced in agent B'sichunkedfix. - One disagreement remains, and it's honest. For
ichunked, the maintainers also changed an error message that the bug report never mentioned. Agent A fixed what was asked; their test suite still wanted the new message. In a real pilot, that is a requirement to write down before the job — not a bug in the fix. - No good fix was ever blamed. Across both rounds, a correct fix was marked FAIL zero times.
Limits
- Small, well-tested pure-Python libraries. Larger codebases with services and databases come next.
- The agents were AI models we ran ourselves, not commercial products, and the bug reports were written by us from the maintainers' commit messages.
- Receipts are signed with demo keys: this was shadow mode, no money moved.
The final round as a client would see it: the pilot report, with every receipt one click from verification. Tasks, patches and both rounds' raw data are available to pilot partners on request.
Run this on your own bugs.
Four weeks, 30–50 real bugs, your agent of choice. Free in shadow mode.