about
Machine Unlearning (arxiv.org)
86 points by Tomte on Jan 28, 2020 | hide | past | pdf | 14 comments on HN

In plain words: To delete a user's data from a trained model, the usual fix is retraining everything from scratch. This splits the training data into pieces and trains separate models, so removing one point means rebuilding only its piece — up to 4.63 times faster.

Abstract

Once users have shared their data online, it is generally difficult for them to revoke access and ask for the data to be deleted. Machine learning (ML) exacerbates this problem because any model trained with said data may have memorized it, putting users at risk of a successful privacy attack exposing their information. Yet, having models unlearn is notoriously difficult. We introduce SISA training, a framework that expedites the unlearning process by strategically limiting the influence of a data point in the training procedure. While our framework is applicable to any learning algorithm, it is designed to achieve the largest improvements for stateful algorithms like stochastic gradient descent for deep neural networks. SISA training reduces the computational overhead associated with unlearning, even in the worst-case setting where unlearning requests are made uniformly across the training set. In some cases, the service provider may have a prior on the distribution of unlearning requests that will be issued by users. We may take this prior into account to partition and order data accordingly, and further decrease overhead from unlearning. Our evaluation spans several datasets from different domains, with corresponding motivations for unlearning. Under no distributional assumptions, for simple learning tasks, we observe that SISA training improves time to unlearn points from the Purchase dataset by 4.63x, and 2.45x for the SVHN dataset, over retraining from scratch. SISA training also provides a speed-up of 1.36x in retraining for complex learning tasks such as ImageNet classification; aided by transfer learning, this results in a small degradation in accuracy. Our work contributes to practical data governance in machine unlearning.

Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, Nicolas Papernot
arXiv:1912.03817 · cs.CR, cs.AI, cs.LG · submitted Dec 9, 2019 · updated Dec 15, 2020
abstract · pdf · html · Published in IEEE S&P 2021

add comment on HN
Also discussed: Aug 2022 (1 point, 0 comments) · Jan 2022 (9 points, 0 comments) · May 2021 (2 points, 0 comments) · Dec 2019 (1 point, 0 comments)

I object to this paper title. The idea of unlearning is already well-known, given there are benchmarks and previous results mentioned in the abstract. This paper introduces a new model and gets good performance but doesn't deserve to be named after the whole field.
I wonder: is this the usual case, that users behavior or other data is used as input for ML projects without their consent? Would they always have to opt out rather than preparing the data in a way that prevents from disregarding users' privacy? Or, assuming that testing for a relevant outcome invariably transgresses on users' privacy, I wonder if this kind of ML work isn't a bit unethical as a whole?
In many cases, the terms of service of a company include provisions that usage data of the products will be analyzed for research and experimentation purposes, which are often considered necessary for the health of the business (and thus even meet GDPR requirements for this data capture & use).

For example, a company couldn’t remain competitive and serve customers or continue existing if it can’t perform a/b testing on new features or changes, or look at descriptive statistics about which types of customers use which products. Creating statistical models to answer these questions or to have aspects of a product that personalize based on these data is a routine matter of business operation. Rightly or wrongly, the terms of service are usually enough to allow the business to use data this way, and often label it as critical for the operation of the business.

I don't see a problem of creating ML models for improving the same company's services, which I can't really even imagine to require much explicit customers' consent. If it's for the same company, the same way all clearance would go to researchers as it went to statisticians doing evaluations for companies in the past. If it's about optimizing your business, do you even need ML? Asking questions about statistical data has been done long before the current age of big-data statistical self-betrayal.

The way I see it is that people start to try building their businesses around the ideas of ML, basically ML as a service, the catch there is just that their ordering businesses data, which is really their customers' data will end up in the big mess of aggregated, weakly correlated data, from which they then try to derive their models that are supposed to make their money. At no point there, I as the customer of company A, can be sure if I'm correctly or incorrectly being correlated in those models. The need to delete me from these evaluations arises from my wish to protect not just my individuality from Brazil-like misinterpretations, but also to protect the companies asking the questions for their businesses, too.

I don't know about you, but to me this casts doubt on the utility of non-specific ML as an arbitrary interpretation of unspecific data that is as useless to me as it is to my competitors, seems just Jack shit, really. You wanna solve a problem? Go solve it by bringing the consumer and the producer closer together, that counts for any business out there, especially insurance and policy, and stop ramming another PC-driven layer of middle management ML between them.

I have such trouble trying to figure out which of these many algorithms that are released will end up having a significant impact on a 5 year horizon. But this I have a rare hunch about that it could be quite significant (at least the direction in which it's trying to push).
But if the point of Machine Learning is to generalise a given dataset, wouldn't a particular pattern (the one one wishes to forget) be (unintentionally) found given other unrelated and/or similar patterns?
I know it’s probably a stupid question but: why not just anonymising the data?
Anonymising data is surprisingly difficult. I'm wondering if there is the "falsehoods programmers believe about" about this, as there are for topics such as names, addresses, time, etc.?
Two quick thoughts: Some data is inherently non-anonymous, like face images. Data that is now anonymous may be de-anonymized in the future.
All human-produced data is inherently non-anonymous.

The only respectful mechanical way to use it is through information-theoretic guarantees, e.g. using it for zero-knowledge proofs and then burning the data.

De-anonymising is possible because of the data via direct or indirect correlations.
Sometimes the data is an image of your face and you don't want it to be used for any algorithm, regardless it has your name associated to it or not
> “ Machine learning (ML) exacerbates this problem because any model trained with said data may have memorized it,”

Yikes, that’s a really draconian scare tactic way to frame it. It clearly is meant to exacerbate misunderstandings of how statistical modeling actually works.