about
Semi-Supervised Knowledge Transfer for Deep Learning from Private Training Data (arxiv.org)
80 points by Katydid on Oct 25, 2016 | hide | past | pdf | 2 comments on HN

In plain words: Many separate models each learn from a different slice of private data, then stay hidden while a student trains on their noisy votes, so no single person's records shape the answer. It gave the best accuracy for a given privacy level on two image tasks.

Abstract · Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data

Some machine learning applications involve training data that is sensitive, such as the medical histories of patients in a clinical trial. A model may inadvertently and implicitly store some of its training data; careful analysis of the model may therefore reveal sensitive information. To address this problem, we demonstrate a generally applicable approach to providing strong privacy guarantees for training data: Private Aggregation of Teacher Ensembles (PATE). The approach combines, in a black-box fashion, multiple models trained with disjoint datasets, such as records from different subsets of users. Because they rely directly on sensitive data, these models are not published, but instead used as "teachers" for a "student" model. The student learns to predict an output chosen by noisy voting among all of the teachers, and cannot directly access an individual teacher or the underlying data or parameters. The student's privacy properties can be understood both intuitively (since no single teacher and thus no single dataset dictates the student's training) and formally, in terms of differential privacy. These properties hold even if an adversary can not only query the student but also inspect its internal workings. Compared with previous work, the approach imposes only weak assumptions on how teachers are trained: it applies to any model, including non-convex models like DNNs. We achieve state-of-the-art privacy/utility trade-offs on MNIST and SVHN thanks to an improved privacy analysis and semi-supervised learning.

Nicolas Papernot, Martín Abadi, Úlfar Erlingsson, Ian Goodfellow, Kunal Talwar
arXiv:1610.05755 · stat.ML, cs.CR, cs.LG · submitted Oct 18, 2016 · updated Mar 3, 2017
abstract · pdf · html · Accepted to ICLR 17 as an oral

add comment on HN

The title and abstract of the paper makes it seem like their approach can assure some form of absolute privacy. However, when you get down into the weeds of the article they really only have probabilistic guarantees. This seems like a problem because (1) most people will incorrectly assume the stronger form of privacy and (2) there is no objective way to determine the threshold required to be considered "private enough" for a given application. This is exactly the same as the issues I have with differential privacy.

In the end, I could see this being useful for protecting private corporate data where the concern is that the company does not want to lose the perceived value of their datasets just because they have released an external model using internal data. Theoretical guarantees that most data will be private should be good enough for this case. On the other hand, I would worry about using it on truly sensitive data (such as medical records) where even one compromised datum is of high concern.

It will be interesting to see how it plays out, but it's worth noting that there is a healthy amount of suspicion, some fear, and even a little bit of animosity toward google in the US health care industry. Opinions about surveillance and marketing (mis)uses aside, getting data can be a long and difficult process even for seasoned researchers in teaching hospitals. Personally, I don't fancy giving them my medical records any time soon.