MotifMiner: a general toolkit for efficiently identifying common substructures in molecules. Scientific research often involves examining structural relationships in molecules since scientists strongly believe in the causal relationship between structure and function. Traditionally, researchers have identified these patterns, or motifs, manually using biochemical expertise. However, with the massive influx of new biochemical data and the ability to gather data for very large molecules, there is great need for techniques that automatically and efficiently identify commonly occurring structural patterns in molecules. Previous automated substructure discovery approaches have each introduced variations of similar underlying techniques and have embedded domain knowledge. While doing so improves performance for the particular domain, this complicates extensibility to other domains. Also, they do not address scalability or noise, which is critical for certain structural domains like macromolecules. In this paper, we present MotifMiner, a general toolkit for automatically identifying common motifs in most any scientific molecular dataset. We describe both our application framework and services for identifying motifs, as well as demonstrate the flexibility of our system by analyzing several disparate domains, including protein, drug, and MD simulation datasets.

References in zbMATH (referenced in 7 articles )

Showing results 1 to 7 of 7.
Sorted by year (citations)

  1. Comin, Matteo; Verzotto, Davide: Filtering degenerate patterns with application to protein sequence analysis (2013)
  2. Shelokar, Prakash; Quirin, Arnaud; Cordón, Óscar: A multiobjective evolutionary programming framework for graph-based data mining (2013) ioport
  3. Assent, Ira; Krieger, Ralph; Glavic, Boris; Seidl, Thomas: Clustering multidimensional sequences in spatial and temporal databases (2008) ioport
  4. Marsolo, Keith; Parthasarathy, Srinivasan: On the use of structure and sequence-based features for protein classification and retrieval (2008) ioport
  5. Leung, Carson Kai-Sang; Khan, Quamrul I.; Li, Zhan; Hoque, Tariqul: CanTree: a canonical-order tree for incremental frequent-pattern mining (2007) ioport
  6. Leung, Carson Kai-Sang; Khan, Quamrul I.; Li, Zhan; Hoque, Tariqul: CanTree: A canonical-order tree for incremental frequent-pattern mining (2007) ioport
  7. Coatney, Matt; Parthasarathy, Srinivasan: MotifMiner: Efficient discovery of common substructures in biochemical molecules (2005) ioport