Abstract
Pattern discovery and subspace clustering play a central role in the biological domain, supporting for instance putative regulatory module discovery from omics data for both descriptive and predictive ends. In the presence of target variables (e.g. phenotypes), regulatory patterns should further satisfy delineate discriminative power properties, well-established in the presence of categorical outcomes, yet largely disregarded for numerical outcomes, such as risk profiles and quantitative phenotypes. DISA (Discriminative and Informative Subspace Assessment), a Python software package, is proposed to evaluate patterns in the presence of numerical outcomes using well-established measures together with a novel principle able to statistically assess the correlation gain of the subspace against the overall space. Results confirm the possibility to soundly extend discriminative criteria towards numerical outcomes without the drawbacks well-associated with discretization procedures. Results from four case studies confirm the validity and relevance of the proposed methods, further unveiling critical directions for research on biotechnology and biomedicine.Availability:DISA is freely available athttps://github.com/JupitersMight/DISAunder the MIT license.
Publisher
Public Library of Science (PLoS)
Reference43 articles.
1. Discriminative pattern mining and its applications in bioinformatics;X. Liu;Briefings In Bioinformatics,2015
2. Biclustering in data mining;S. Busygin;Computers & Operations Research,2008
3. Applications of frequent pattern mining;C. Aggarwal;Frequent Pattern Mining,2014
4. It is time to apply biclustering: a comprehensive review of biclustering applications in biological and biomedical data;J. Xie;Briefings In Bioinformatics,2019
5. Analyzing fibrous tissue pattern in fibrous dysplasia bone images using deep R-CNN networks for segmentation;A. Saranya;Soft Computing,2021
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献