Sample: the report a client gets at the end of a pilot, built from our open-source demo pilot (12 real bugs, two AI agents and the maintainers' own fixes). Here the reviewer is each project's upstream test suite; in a client pilot it is the client's engineer. The full case study.
ResultBond pilot report
Open-source demo pilot
36 runs · 12 tasks · 3 agents · shadow mode · 2 October 2026, 16:02 UTC
Recommendation: move to bonded modeEvery target is met.
97.2%agree with your reviewerstarget ≥ 95.0%MET
0correct fixes marked FAILtarget ≤ 0MET
100.0%runs got a PASS or FAILtarget ≥ 80.0%MET
4.8 smedian time to receipttarget ≤ 600 sMET
36/36receipts verifytarget allMET
36verdicts labelled by reviewerstarget ≥ 10MET
How each agent did
| Agent | Runs | PASS | FAIL | Held | Pass rate | Agreement with reviewer |
|---|---|---|---|---|---|---|
| agent A (Sonnet) | 12 | 11 | 1 | 0 | 91.7% | 91.7% of 12 labelled |
| agent B (Haiku) | 12 | 10 | 2 | 0 | 83.3% | 100.0% of 12 labelled |
| upstream maintainers | 12 | 12 | 0 | 0 | 100.0% | 100.0% of 12 labelled |
Where we disagreed
| Task | Run | Agent | What happened | ResultBond | Reviewer | Note |
|---|---|---|---|---|---|---|
| mi-ichunked-zero | 2 | agent A (Sonnet) | ResultBond passed a fix your reviewer rejected | PASSALL_CHECKS_PASS | FAIL | 1 of 619 upstream tests fail: test_negative |
Tasks
| Task | Bug | Hidden tests reviewed by | Runs |
|---|---|---|---|
| boltons-guiderator-size | GUIDerator size limits are wrong | — | 3 |
| boltons-histogram-iqr | Histograms fail when the interquartile range is zero | — | 3 |
| boltons-pearson-zero | Stats.pearson_type fails on a zero denominator | — | 3 |
| boltons-relative-time-tz | relative_time() breaks on timezone-aware datetimes | — | 3 |
| mi-bucket-miss | bucket lookups invent keys | — | 3 |
| mi-ichunked-zero | ichunked() misbehaves for n=0 | — | 3 |
| mi-one-falsy-exception | one() and only() ignore falsy custom exceptions | — | 3 |
| mi-running-window | running_* window sizes are not validated | — | 3 |
| toolz-interpose-empty | interpose() raises on an empty sequence | — | 3 |
| toolz-reduceby-init | reduceby() should accept a callable init | — | 3 |
| toolz-topk-stable | topk() is not stable when keys tie | — | 3 |
| toolz-topk-tuple | topk() should return a tuple with its own implementation | — | 3 |
Every run
| Task | Run | Agent | Verdict | Why | Time | Reviewer | Receipt |
|---|---|---|---|---|---|---|---|
| boltons-guiderator-size | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass | 4.2 s | PASS | check ↗boltons-guiderator-size__01.json |
| boltons-guiderator-size | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass | 4.3 s | PASS | check ↗boltons-guiderator-size__02.json |
| boltons-guiderator-size | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass | 4.3 s | PASS | check ↗boltons-guiderator-size__03.json |
| boltons-histogram-iqr | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass | 3.9 s | PASS | check ↗boltons-histogram-iqr__01.json |
| boltons-histogram-iqr | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass | 3.5 s | PASS | check ↗boltons-histogram-iqr__02.json |
| boltons-histogram-iqr | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass | 3.6 s | PASS | check ↗boltons-histogram-iqr__03.json |
| boltons-pearson-zero | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass | 3.6 s | PASS | check ↗boltons-pearson-zero__01.json |
| boltons-pearson-zero | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass | 3.4 s | PASS | check ↗boltons-pearson-zero__02.json |
| boltons-pearson-zero | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass | 3.4 s | PASS | check ↗boltons-pearson-zero__03.json |
| boltons-relative-time-tz | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass | 3.7 s | PASS | check ↗boltons-relative-time-tz__01.json |
| boltons-relative-time-tz | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass | 3.7 s | PASS | check ↗boltons-relative-time-tz__02.json |
| boltons-relative-time-tz | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass | 3.7 s | PASS | check ↗boltons-relative-time-tz__03.json |
| mi-bucket-miss | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 613 upstream tests in test_more.py pass | 80.5 s | PASS | check ↗mi-bucket-miss__01.json |
| mi-bucket-miss | 2 | agent A (Sonnet) | FAIL | the change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 613 upstream tests fail: test_list | 83.6 s | FAIL | check ↗mi-bucket-miss__02.json |
| mi-bucket-miss | 3 | agent B (Haiku) | FAIL | the change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 613 upstream tests fail: test_list | 85.9 s | FAIL | check ↗mi-bucket-miss__03.json |
| mi-ichunked-zero | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 619 upstream tests in test_more.py pass | 80.2 s | PASS | check ↗mi-ichunked-zero__01.json |
| mi-ichunked-zero | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: 1 of 619 upstream tests fail: test_negative | 83.3 s | FAIL | check ↗mi-ichunked-zero__02.json |
| mi-ichunked-zero | 3 | agent B (Haiku) | FAIL | the change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 619 upstream tests fail: test_negative | 83.0 s | FAIL | check ↗mi-ichunked-zero__03.json |
| mi-one-falsy-exception | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass | 79.6 s | PASS | check ↗mi-one-falsy-exception__01.json |
| mi-one-falsy-exception | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass | 83.3 s | PASS | check ↗mi-one-falsy-exception__02.json |
| mi-one-falsy-exception | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass | 79.0 s | PASS | check ↗mi-one-falsy-exception__03.json |
| mi-running-window | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass | 61.8 s | PASS | check ↗mi-running-window__01.json |
| mi-running-window | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass | 60.9 s | PASS | check ↗mi-running-window__02.json |
| mi-running-window | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass | 61.7 s | PASS | check ↗mi-running-window__03.json |
| toolz-interpose-empty | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass | 4.6 s | PASS | check ↗toolz-interpose-empty__01.json |
| toolz-interpose-empty | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass | 4.6 s | PASS | check ↗toolz-interpose-empty__02.json |
| toolz-interpose-empty | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass | 4.8 s | PASS | check ↗toolz-interpose-empty__03.json |
| toolz-reduceby-init | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass | 4.6 s | PASS | check ↗toolz-reduceby-init__01.json |
| toolz-reduceby-init | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass | 4.7 s | PASS | check ↗toolz-reduceby-init__02.json |
| toolz-reduceby-init | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass | 4.6 s | PASS | check ↗toolz-reduceby-init__03.json |
| toolz-topk-stable | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass | 4.7 s | PASS | check ↗toolz-topk-stable__01.json |
| toolz-topk-stable | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass | 4.8 s | PASS | check ↗toolz-topk-stable__02.json |
| toolz-topk-stable | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass | 4.9 s | PASS | check ↗toolz-topk-stable__03.json |
| toolz-topk-tuple | 1 | upstream maintainers | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass | 5.3 s | PASS | check ↗toolz-topk-tuple__01.json |
| toolz-topk-tuple | 2 | agent A (Sonnet) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass | 5.2 s | PASS | check ↗toolz-topk-tuple__02.json |
| toolz-topk-tuple | 3 | agent B (Haiku) | PASS | the fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass | 5.2 s | PASS | check ↗toolz-topk-tuple__03.json |