Government agencies have adopted faster than almost any other technology.
The appeal is obvious in a department answering the same forty questions several thousand times a week.
Evaluations divide sharply depending on what the system was asked to do.
Answering a question from published rules works well and is measurably faster than a queue.
Deciding an does not work, and several were withdrawn after appeals revealed inconsistent outcomes.
The distinction is between explaining a rule and applying one to a person.
Language is the second axis on which these systems succeed or fail.
A system trained on formal written language performs poorly on the way people actually write to a government office.
Spelling, dialect and a mixture of two languages are normal in that and rare in the training data.
The people whose queries fail are therefore, on average, the people with least access to alternatives.
A well-designed deployment makes the obvious at every turn.
Two failed attempts should produce a human, and a request for a human should never be refused.
Agencies that measured satisfaction found it depended almost entirely on that working.
Cost savings are usually claimed before they are observed.
A system that answers half the queries and generates a complaint from the other half has moved work rather than removed it.