Anonymizing Unstructured Data
| dc.creator | Motwani, Rajeev | |
| dc.creator | Nabar, Shubha U. | |
| dc.date | 2008-10-31 | |
| dc.date | 2008-11-03 | |
| dc.date.accessioned | 2026-07-07T10:14:31Z | |
| dc.date.available | 2026-07-07T10:14:31Z | |
| dc.description | In this paper we consider the problem of anonymizing datasets in which each individual is associated with a set of items that constitute private information about the individual. Illustrative datasets include market-basket datasets and search engine query logs. We formalize the notion of k-anonymity for set-valued data as a variant of the k-anonymity model for traditional relational datasets. We define an optimization problem that arises from this definition of anonymity and provide O(klogk) and O(1)-approximation algorithms for the same. We demonstrate applicability of our algorithms to the America Online query log dataset. | |
| dc.description | 9 pages, 1 figure | |
| dc.identifier | https://arxiv.org/abs/0810.5582 | |
| dc.identifier | http://arxiv.org/abs/0810.5582 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/172902 | |
| dc.subject | Databases | |
| dc.subject | Data Structures and Algorithms | |
| dc.title | Anonymizing Unstructured Data | |
| dc.type | text |