about
Dual Query: Practical Private Query Release for High Dimensional Data (arxiv.org)
7 points by user_235711 on Feb 12, 2014 | hide | past | pdf | discuss on HN

In plain words: It answers many questions about a dataset without exposing any one person, hiding the hard part in a math puzzle ordinary solvers can finish. On Netflix data with over 17,000 traits it handled millions of questions accurately, beating the previous best by orders of magnitude.

Abstract

We present a practical, differentially private algorithm for answering a large number of queries on high dimensional datasets. Like all algorithms for this task, ours necessarily has worst-case complexity exponential in the dimension of the data. However, our algorithm packages the computationally hard step into a concisely defined integer program, which can be solved non-privately using standard solvers. We prove accuracy and privacy theorems for our algorithm, and then demonstrate experimentally that our algorithm performs well in practice. For example, our algorithm can efficiently and accurately answer millions of queries on the Netflix dataset, which has over 17,000 attributes; this is an improvement on the state of the art by multiple orders of magnitude.

Marco Gaboardi, Emilio Jesús Gallego Arias, Justin Hsu, Aaron Roth, Zhiwei Steven Wu
arXiv:1402.1526 · cs.DS, cs.CR, cs.DB, cs.LG · submitted Feb 6, 2014 · updated Nov 19, 2015
abstract · pdf · html

add comment on HN