C2 Đọc hiểu

Dữ liệu tổng hợp và giới hạn của ẩn danh

18câu8từ khoá6câu hỏi14phút đọc

A dataset from which names have been removed is described as , and the description is usually wrong.

Identity is carried not by names but by combinations of that are individually unremarkable.

A , a date of birth and a sex identify a large majority of a population uniquely.

None of those three is a name and no one of them identifies anybody.

proceeds by joining such a dataset to any other dataset sharing a few columns.

The second dataset need not be secret and is frequently a public register.

Published demonstrations have re-identified individuals in medical, transport and viewing records within days.

Institutions such data in good faith and in reliance on a legal test that the technique defeats.

is offered as the remedy and is a genuine improvement rather than a solution.

Publishing counts rather than records prevents the simplest attacks and permits subtler ones.

Two aggregate tables that differ by one person reveal that person's value in every column.

Statistical agencies have known this for decades and suppress small cells for exactly this reason.

is effective and it also removes the information most likely to be useful about small groups.

A minority too small to report is a minority about whom nothing can be published.

The trade-off between privacy and visibility falls on the groups least able to bear either loss.

Formal methods that add noise offer a way to state the trade-off rather than to escape it.

They produce a number describing how much privacy a costs, which no previous approach did.

Having the number does not decide what to do, and an argument conducted with one is a different argument.

Từ khoá