If a quantity is measured twice and the first measurement is , the second will usually be less so.
This is to the mean, and it follows from the fact that any measurement contains both signal and .
An reading is more likely than a middling one to have been pushed there by .
Noise does not persist, and so the next reading drifts back towards the centre.
Nothing has caused the improvement, and yet an explanation is almost always supplied.
The classic case is the flying instructor who praises a good landing and sees the next one .
He punishes a bad landing and sees the next one improve, and concludes that praise harms and criticism works.
The conclusion is false and the observation is accurate, which is what makes the error so .
Policy is affected more seriously than folklore.
A programme targeted at the worst-performing schools will record improvement whether or not it does anything.
The schools were selected because their results were , and results regress.
Any evaluation without a comparison group will therefore report a success it cannot claim.
The remedy is neither complicated nor expensive: select the schools, then which of them receives the programme.
Both groups regress by the same amount, and the difference between them is the effect.
Resistance to this design is rarely statistical and usually ethical or political.
Withholding a programme from schools that need it is difficult to defend in public.
It is also the only way to discover whether the programme is worth giving to anyone.
A measure adopted without that test will be on evidence that alone would have produced.