Sample: the report a client gets at the end of a pilot, built from our open-source demo pilot (12 real bugs, two AI agents and the maintainers' own fixes). Here the reviewer is each project's upstream test suite; in a client pilot it is the client's engineer. The full case study.

ResultBond pilot report

Open-source demo pilot

36 runs · 12 tasks · 3 agents · shadow mode · 2 October 2026, 16:02 UTC

Recommendation: move to bonded modeEvery target is met.
97.2%agree with your reviewerstarget ≥ 95.0%MET
0correct fixes marked FAILtarget ≤ 0MET
100.0%runs got a PASS or FAILtarget ≥ 80.0%MET
4.8 smedian time to receipttarget ≤ 600 sMET
36/36receipts verifytarget allMET
36verdicts labelled by reviewerstarget ≥ 10MET

How each agent did

AgentRunsPASSFAILHeldPass rateAgreement with reviewer
agent A (Sonnet)12111091.7%91.7% of 12 labelled
agent B (Haiku)12102083.3%100.0% of 12 labelled
upstream maintainers121200100.0%100.0% of 12 labelled

Where we disagreed

TaskRunAgentWhat happenedResultBondReviewerNote
mi-ichunked-zero2agent A (Sonnet)ResultBond passed a fix your reviewer rejectedPASSALL_CHECKS_PASSFAIL1 of 619 upstream tests fail: test_negative

Tasks

TaskBugHidden tests reviewed byRuns
boltons-guiderator-sizeGUIDerator size limits are wrong—3
boltons-histogram-iqrHistograms fail when the interquartile range is zero—3
boltons-pearson-zeroStats.pearson_type fails on a zero denominator—3
boltons-relative-time-tzrelative_time() breaks on timezone-aware datetimes—3
mi-bucket-missbucket lookups invent keys—3
mi-ichunked-zeroichunked() misbehaves for n=0—3
mi-one-falsy-exceptionone() and only() ignore falsy custom exceptions—3
mi-running-windowrunning_* window sizes are not validated—3
toolz-interpose-emptyinterpose() raises on an empty sequence—3
toolz-reduceby-initreduceby() should accept a callable init—3
toolz-topk-stabletopk() is not stable when keys tie—3
toolz-topk-tupletopk() should return a tuple with its own implementation—3

Every run

TaskRunAgentVerdictWhyTimeReviewerReceipt
boltons-guiderator-size1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass4.2 sPASScheck ↗boltons-guiderator-size__01.json
boltons-guiderator-size2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass4.3 sPASScheck ↗boltons-guiderator-size__02.json
boltons-guiderator-size3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_iterutils.py pass4.3 sPASScheck ↗boltons-guiderator-size__03.json
boltons-histogram-iqr1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass3.9 sPASScheck ↗boltons-histogram-iqr__01.json
boltons-histogram-iqr2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass3.5 sPASScheck ↗boltons-histogram-iqr__02.json
boltons-histogram-iqr3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 3 upstream tests in test_statsutils.py pass3.6 sPASScheck ↗boltons-histogram-iqr__03.json
boltons-pearson-zero1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass3.6 sPASScheck ↗boltons-pearson-zero__01.json
boltons-pearson-zero2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass3.4 sPASScheck ↗boltons-pearson-zero__02.json
boltons-pearson-zero3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 5 upstream tests in test_statsutils.py pass3.4 sPASScheck ↗boltons-pearson-zero__03.json
boltons-relative-time-tz1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass3.7 sPASScheck ↗boltons-relative-time-tz__01.json
boltons-relative-time-tz2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass3.7 sPASScheck ↗boltons-relative-time-tz__02.json
boltons-relative-time-tz3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 11 upstream tests in test_timeutils.py pass3.7 sPASScheck ↗boltons-relative-time-tz__03.json
mi-bucket-miss1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 613 upstream tests in test_more.py pass80.5 sPASScheck ↗mi-bucket-miss__01.json
mi-bucket-miss2agent A (Sonnet)FAILthe change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 613 upstream tests fail: test_list83.6 sFAILcheck ↗mi-bucket-miss__02.json
mi-bucket-miss3agent B (Haiku)FAILthe change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 613 upstream tests fail: test_list85.9 sFAILcheck ↗mi-bucket-miss__03.json
mi-ichunked-zero1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 619 upstream tests in test_more.py pass80.2 sPASScheck ↗mi-ichunked-zero__01.json
mi-ichunked-zero2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: 1 of 619 upstream tests fail: test_negative83.3 sFAILcheck ↗mi-ichunked-zero__02.json
mi-ichunked-zero3agent B (Haiku)FAILthe change broke behaviour that worked beforeREGRESSION_FAILUREreviewer: 1 of 619 upstream tests fail: test_negative83.0 sFAILcheck ↗mi-ichunked-zero__03.json
mi-one-falsy-exception1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass79.6 sPASScheck ↗mi-one-falsy-exception__01.json
mi-one-falsy-exception2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass83.3 sPASScheck ↗mi-one-falsy-exception__02.json
mi-one-falsy-exception3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 608 upstream tests in test_more.py pass79.0 sPASScheck ↗mi-one-falsy-exception__03.json
mi-running-window1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass61.8 sPASScheck ↗mi-running-window__01.json
mi-running-window2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass60.9 sPASScheck ↗mi-running-window__02.json
mi-running-window3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 148 upstream tests in test_recipes.py pass61.7 sPASScheck ↗mi-running-window__03.json
toolz-interpose-empty1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass4.6 sPASScheck ↗toolz-interpose-empty__01.json
toolz-interpose-empty2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass4.6 sPASScheck ↗toolz-interpose-empty__02.json
toolz-interpose-empty3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 51 upstream tests in test_itertoolz.py pass4.8 sPASScheck ↗toolz-interpose-empty__03.json
toolz-reduceby-init1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass4.6 sPASScheck ↗toolz-reduceby-init__01.json
toolz-reduceby-init2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass4.7 sPASScheck ↗toolz-reduceby-init__02.json
toolz-reduceby-init3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 40 upstream tests in test_itertoolz.py pass4.6 sPASScheck ↗toolz-reduceby-init__03.json
toolz-topk-stable1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass4.7 sPASScheck ↗toolz-topk-stable__01.json
toolz-topk-stable2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass4.8 sPASScheck ↗toolz-topk-stable__02.json
toolz-topk-stable3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 44 upstream tests in test_itertoolz.py pass4.9 sPASScheck ↗toolz-topk-stable__03.json
toolz-topk-tuple1upstream maintainersPASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass5.3 sPASScheck ↗toolz-topk-tuple__01.json
toolz-topk-tuple2agent A (Sonnet)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass5.2 sPASScheck ↗toolz-topk-tuple__02.json
toolz-topk-tuple3agent B (Haiku)PASSthe fix works: hidden and public tests pass, nothing brokeALL_CHECKS_PASSreviewer: all 43 upstream tests in test_itertoolz.py pass5.2 sPASScheck ↗toolz-topk-tuple__03.json