Measuring a teacher by the progress of their pupils sounds like the obvious approach.
attempt exactly that, comparing each pupil's result with a prediction from their earlier performance.
The statistical problems are well documented and have not prevented adoption.
for the same teacher vary substantially from one year to the next.
A teacher in the top fifth one year has a meaningful probability of appearing in the bottom fifth the next.
Class sizes of thirty produce estimates with wide enough to include most of the distribution.
Using such a number for or pay is therefore a decision made on .
Systems that used it that way produced predictable behaviour rather than better teaching.
Teachers competed for the classes likely to progress and avoided the pupils who would lower an average.
Some taught the test directly, which raises the measure and not the underlying ability.
The same data used differently has produced genuine improvement.
to a school and over three years, value-added figures are stable enough to identify departments that need support.
Used as information for a conversation rather than as a score for a person, the method survives its own error bars.
Observation of teaching is the necessary and is expensive, which is why numbers spread faster.
Every system that replaced judgement with a statistic did so because judgement is contestable and a number is not.