Abstract
With innovations and advancements in analytical instruments and computer technology, omics studies based on statistical analysis, such as phytochemical omics, oilomics/lipidomics, proteomics, metabolomics, and glycomics, are increasingly popular in the areas of food chemistry and nutrition science. However, a remaining hurdle is the labor-intensive data process because learning coding skills and software operations are usually time-consuming for researchers without coding backgrounds. A MATLAB® coding basis and three-in-one integrated method, ‘Ana’, was created for data visualizations and statistical analysis in this work. The program loaded and analyzed an omics dataset from an Excel® file with 7 samples * 22 compounds as an example, and output six figures for three types of data visualization, including a 3D heatmap, heatmap hierarchical clustering analysis, and principal component analysis (PCA), in 18 s on a personal computer (PC) with a Windows 10 system and in 20 s on a Mac with a MacOS Monterey system. The code is rapid and efficient to print out high-quality figures up to 150 or 300 dpi. The output figures provide enough contrast to differentiate the omics dataset by both color code and bar size adjustments per their higher or lower values, allowing the figures to be qualified for publication and presentation purposes. It provides a rapid analysis method that would liberate researchers from labor-intensive and time-consuming manual or coding basis data analysis. A coding example with proper code annotations and completed user guidance is provided for undergraduate and postgraduate students to learn coding basis statistical data analysis and to help them utilize such techniques for their future research.
Funder
California Department of Food and Agriculture, 2020 Specialty Crop Block Grant Program
Subject
Paleontology,Space and Planetary Science,General Biochemistry, Genetics and Molecular Biology,Ecology, Evolution, Behavior and Systematics