Anonymizing Unstructured Data

dc.creatorMotwani, Rajeev
dc.creatorNabar, Shubha U.
dc.date2008-10-31
dc.date2008-11-03
dc.date.accessioned2026-07-07T10:14:31Z
dc.date.available2026-07-07T10:14:31Z
dc.descriptionIn this paper we consider the problem of anonymizing datasets in which each individual is associated with a set of items that constitute private information about the individual. Illustrative datasets include market-basket datasets and search engine query logs. We formalize the notion of k-anonymity for set-valued data as a variant of the k-anonymity model for traditional relational datasets. We define an optimization problem that arises from this definition of anonymity and provide O(klogk) and O(1)-approximation algorithms for the same. We demonstrate applicability of our algorithms to the America Online query log dataset.
dc.description9 pages, 1 figure
dc.identifierhttps://arxiv.org/abs/0810.5582
dc.identifierhttp://arxiv.org/abs/0810.5582
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/172902
dc.subjectDatabases
dc.subjectData Structures and Algorithms
dc.titleAnonymizing Unstructured Data
dc.typetext

Files

Collections