XCleaner: A new method for clustering XML documents by structure. With the vastly growing data resources on the Internet, XML is one of the most important standards for document management. Not only does it provide enhancements to document exchange and storage, but it is also helpful in a variety of information retrieval tasks. Document clustering is one of the most interesting research areas that utilize semi-structural nature of XML. In this paper, we put forward a new XML clustering algorithm that relies solely on document structure. We propose the use of maximal frequent subtrees and an operator called Satisfy/Violate to divide documents into groups. The algorithm is experimentally evaluated on real and synthetic data sets with promising results.
References in zbMATH (referenced in 1 article , 1 standard article )
Showing result 1 of 1.
- Brzeziński, Dariusz; Leśniewska, Anna; Morzy, Tadeusz; Piernik, Maciej: XCleaner: a new method for clustering XML documents by structure (2011)