DocumentCode :
1762644
Title :
Right-Protected Data Publishing with Provable Distance-Based Mining
Author :
Zoumpoulis, Spyros I. ; Vlachos, Michail ; Freris, Nikolaos M. ; Lucchese, Claudio
Author_Institution :
Lab. for Inf. & Decision Syst., Massachusetts Inst. of Technol., Cambridge, MA, USA
Volume :
26
Issue :
8
fYear :
2014
fDate :
Aug. 2014
Firstpage :
2014
Lastpage :
2028
Abstract :
Protection of one´s intellectual property is a topic with important technological and legal facets. We provide mechanisms for establishing the ownership of a dataset consisting of multiple objects. The algorithms also preserve important properties of the dataset, which are important for mining operations, and so guarantee both right protection and utility preservation. We consider a right-protection scheme based on watermarking. Watermarking may distort the original distance graph. Our watermarking methodology preserves important distance relationships, such as: the Nearest Neighbors (NN) of each object and the Minimum Spanning Tree (MST) of the original dataset. This leads to preservation of any mining operation that depends on the ordering of distances between objects, such as NN-search and classification, as well as many visualization techniques. We prove fundamental lower and upper bounds on the distance between objects post-watermarking. In particular, we establish a restricted isometry property, i.e., tight bounds on the contraction/expansion of the original distances. We use this analysis to design fast algorithms for NN-preserving and MST-preserving watermarking that drastically prune the vast search space. We observe two orders of magnitude speedup over the exhaustive schemes, without any sacrifice in NN or MST preservation.
Keywords :
data mining; data visualisation; learning (artificial intelligence); trees (mathematics); watermarking; MST; MST-preserving watermarking; NN-preserving watermarking; dataset ownership; dataset property; intellectual property; minimum spanning tree; nearest neighbors; provable distance-based mining; restricted isometry property; right protection; right-protected data publishing; utility preservation; visualization techniques; watermarking methodology; Algorithm design and analysis; Correlation; Data mining; Frequency-domain analysis; Publishing; Vectors; Watermarking; Data and knowledge visualization; Data mining; Database Applications; Database Management; Information Technology and Systems; Watermarking; minimum spanning tree (MST); nearest neighbors (NN); restricted isometry property (RIP);
fLanguage :
English
Journal_Title :
Knowledge and Data Engineering, IEEE Transactions on
Publisher :
ieee
ISSN :
1041-4347
Type :
jour
DOI :
10.1109/TKDE.2013.90
Filename :
6529087
Link To Document :
بازگشت