Machine translation is excellent between languages with large parallel .
It is poor, and sometimes dangerous, for languages that lack them.
A model trained on little data produces fluent text with invented content.
is the trap, because a reader judges reliability by how natural the sounds.
Medical and legal settings have already recorded harm from exactly this failure.
The data gap is not a natural fact but a consequence of who has published online.
A language with millions of speakers can have very little written text in a digital form.
Community projects that collect and texts have therefore mattered more than model .
Recording speech is often easier than collecting writing, and speech models can be trained on it.
Ownership of that collected material is a live question in several communities.
A dataset gathered by volunteers can end up in a commercial model with no return to the speakers.
that permit research and restrict commercial use have become the common answer.
is the other weak point, since the standard scores were built for well-resourced pairs.
A high score can hide errors that a speaker would notice immediately.
Systems used in a clinic should be tested by speakers in a clinic, which is slow and is the only honest method.