Using Intermediate Data of Map Reduce for Faster Execution-Reference-Cited by-同舟云学术

Using Intermediate Data of Map Reduce for Faster Execution

Published:2022-03-08 Issue: Volume:16 Page:20-26
ISSN:2074-1294
Container-title:International Journal of Computers and Communications
language:en
Short-container-title:

Author:

Prakash Shah Pratik¹,V. Pattabiraman¹

Affiliation:

1. School of Computing Science and Engineering VIT University – Chennai Campus Chennai, India

Abstract

Data of any kind structured, unstructured or semistructured is generated in large quantity around the globe in various domains. These datasets are stored on multiple nodes in a cluster. MapReduce framework has emerged as the most efficient technique and easy to use for parallel processing of distributed data. This paper proposes a new methodology for mapreduce framework workflow. The proposed methodology provides a way to process raw data in such a way that it requires less processing time to generate the required result. The methodology stores intermediate data which is generated between map and reduce phase and re-used as input to mapreduce. The paper presents methodology which focuses on improving the data reusability, scalability and efficiency of the mapreduce framework for large data analysis. MongoDB 2.4.2 is used to demonstrate the experimental work to show how we can store and reuse intermediate data as a part of mapreduce to improve the processing of large datasets.

Publisher

North Atlantic University Union (NAUN)

Subject

Literature and Literary Theory,History,Cultural Studies

Reference8 articles.

1. A. Espinosa, P. Hernandez, J.C. Moure, J. Protasio and A. Ripoll . “Analysis and improvement of map-reduce data distribution in mapping applications”. Published online: 8 June 2012 JSupercomput (2012) 62:1305-1317.

2. Diana Moise, Thi-Thu-Lan Trieu, Gabriel Antoniu and Luc Bougé. “Optimizing Intermediate Data Management in MapReduce Computations”. CloudCP „11 April 10, 2011 Salzburg, Austria ACM 978-1-4503-0727-7/11/04.

3. Iman Elghandour and Ashraf Aboulnaga. “ReStore: Reusing Results of MapReduce jobs”. The 38th International Conference on Very Large Data Bases, August 27th 31st 2012, Istanbul, Turkey. Proceedings of the VLDB Endowment, Vol. 5, No. 6.

4. J. Dean and S. Ghemawat. “MapReduce: Simplified data processing on large clusters”. In Proc. OSDI, pages 137-150. 2004.

5. Qiang Liu, Tim Todman, Wayne Luk and George A. Constantinides. “Automatic Optimisation of MapReduce Designs by Geometric Programming”. FPT 2009 IEEE.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. RA-MRS: A high efficient attribute reduction algorithm in big data;Journal of King Saud University - Computer and Information Sciences;2024-06