Challenges of Big Data analysis-Reference-Cited by-同舟云学术

Challenges of Big Data analysis

Published:2014-02-05 Issue:2 Volume:1 Page:293-314
ISSN:2053-714X
Container-title:National Science Review
language:en
Short-container-title:

Author:

Fan Jianqing¹,Han Fang²,Liu Han¹

Affiliation:

1. Department of Operations Research and Financial Engineering, Princeton University, Princeton, NJ 08544, USA;

2. Department of Biostatistics, Johns Hopkins University, Baltimore, MD 21205, USA

Abstract

Abstract Big Data bring new opportunities to modern society and challenges to data scientists. On the one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data introduce unique computational and statistical challenges, including scalability and storage bottleneck, noise accumulation, spurious correlation, incidental endogeneity and measurement errors. These challenges are distinguished and require new computational and statistical paradigm. This paper gives overviews on the salient features of Big Data and how these features impact on paradigm change on statistical and computational methods as well as computing architectures. We also provide various new perspectives on the Big Data analysis and computation. In particular, we emphasize on the viability of the sparsest solution in high-confidence set and point out that exogenous assumptions in most statistical methods for Big Data cannot be validated due to incidental endogeneity. They can lead to wrong statistical inferences and consequently wrong scientific conclusions.

Publisher

Oxford University Press (OUP)

Subject

Multidisciplinary

Link

http://academic.oup.com/nsr/article-pdf/1/2/293/31565398/nwt032.pdf

Reference122 articles.

1. The case for cloud computing in genome informatics;Stein;Genome Biol,2010

2. High-dimensional data analysis: the curses and blessings of dimensionality;Donoho

3. Discussion on the paper ‘Sure independence screening for ultrahigh dimensional feature space’ by Fan and Lv;Bickel;J Roy Stat Soc B,2008

4. High dimensional classification using features annealed independence rules;Fan;Ann Stat,2008

5. Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes;Pittelkow;J Roy Stat Soc B,2008

Cited by 903 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Distributed optimal subsampling for quantile regression with massive data;Journal of Statistical Planning and Inference;2024-12

2. When we talk about Big Data, What do we really mean? Toward a more precise definition of Big Data;Frontiers in Big Data;2024-09-10

3. Special Economic Zones and Firms’ Trade Regime;Emerging Markets Finance and Trade;2024-09-05

4. A Systematic Review of Synthetic Data Generation Techniques Using Generative AI;Electronics;2024-09-04

5. Identify the most appropriate imputation method for handling missing values in clinical structured datasets: a systematic review;BMC Medical Research Methodology;2024-08-28