You have shipped a policy. Someone asks whether it is better than the last one. That is not one measurement, it is the difference between two measurements, and a difference needs far more evidence than either of the things being differenced. Here is the number nobody likes. To tell a seventy percent task from an eighty percent one at ninety-five percent confidence takes about two hundred and eighty seeds per policy. Per task. The field usually reports three to ten. Which means most reported rankings between close methods are not measurements, they are coin flips with a table around them. And a suite makes it worse, not better. Thirty-one tasks compared at ninety-five percent confidence each: the chance that every single ranking is sound is not ninety-five percent, it is zero point nine five to the thirty-first power, which is about twenty percent. Four times out of five, at least one row of your table is noise. You expect about one and a half false rankings from chance alone. None of this says do not build suites. It says a suite spends confidence to buy breadth, at an exchange rate you can compute -- and you should compute it before you draw a conclusion, not after someone questions one.