Title :
Stochastic clustering for organizing distributed information sources
Author :
Shyu, Mei-Ling ; Chen, Shu-Ching ; Rubin, Stuart H.
Author_Institution :
Dept. of Electr. & Comput. Eng., Univ. of Miami, Coral Gables, FL, USA
Abstract :
The number of information sources and the volumes of data in these information sources have greatly increased, which may be attributed to the ever-increasing complexity of real-world applications. The enormous amount of information available in the information sources in a distributed information-providing environment has created a need to provide users with tools to effectively and efficiently navigate and retrieve information. Queries in such an environment often access information from multiple information sources. This may be attributed to navigational characteristics. Clusters provide a structure for organizing the large number of information sources for efficient browsing, searching, and retrieval. This paper presents a stochastically-based clustering mechanism, called the Markov model mediator (MMM), to group the information sources into a set of useful clusters. Each information source cluster groups those information sources that show similarities in their data access behavior. Information sources within the same cluster are expected to be able to provide most of the required information among themselves for user queries that are closely related with respect to a particular application. This can significantly improve system response time, query performance, and result in an overall improvement in decision support. Empirical studies on real databases are performed and the results demonstrate that our proposed mechanism leads to a better set of clusters in comparison with other clustering methods. This serves to illustrate the effectiveness of our proposed MMM mechanism.
Keywords :
Markov processes; data mining; distributed databases; information resources; information retrieval; pattern clustering; Markov model mediator; data access behavior; decision support; distributed information source; information navigation; information retrieval; query performance; stochastic clustering; system response time; Clustering methods; Data analysis; Data mining; Databases; Delay; Engineering management; Information retrieval; Navigation; Organizing; Stochastic processes; Artificial Intelligence; Cluster Analysis; Computer Communication Networks; Database Management Systems; Databases, Factual; Information Storage and Retrieval; Models, Statistical; Pattern Recognition, Automated; Stochastic Processes; Systems Integration; User-Computer Interface;
Journal_Title :
Systems, Man, and Cybernetics, Part B: Cybernetics, IEEE Transactions on
DOI :
10.1109/TSMCB.2004.833599