A Survey on Deep Learning for Software Engineering-Reference-Cited by-同舟云学术

A Survey on Deep Learning for Software Engineering

Published:2021-12-22 Issue: Volume: Page:
ISSN:0360-0300
Container-title:ACM Computing Surveys
language:en
Short-container-title:ACM Comput. Surv.

Author:

Yang Yanming¹^ORCID,Xia Xin²^ORCID,Lo David³^ORCID,Grundy John⁴^ORCID

Affiliation:

1. School of Computer Science and Technology, Zhejiang University, China

2. Software Engineering Application Technology Lab, Huawei, China

3. School of Information Systems, Singapore Management University, Singapore

4. Faculty of Information Technology, Monash University, Australia

Abstract

In 2006, Geoffrey Hinton proposed the concept of training “Deep Neural Networks (DNNs)” and an improved model training method to break the bottleneck of neural network development. More recently, the introduction of AlphaGo in 2016 demonstrated the powerful learning ability of deep learning and its enormous potential. Deep learning has been increasingly used to develop state-of-the-art software engineering (SE) research tools due to its ability to boost performance for various SE tasks. There are many factors, e.g., deep learning model selection, internal structure differences, and model optimization techniques, that may have an impact on the performance of DNNs applied in SE. Few works to date focus on summarizing, classifying, and analyzing the application of deep learning techniques in SE. To fill this gap, we performed a survey to analyze the relevant studies published since 2006. We first provide an example to illustrate how deep learning techniques are used in SE. We then conduct a background analysis (BA) of primary studies and present four research questions to describe the trend of DNNs used in SE (BA), summarize and classify different deep learning techniques (RQ1), analyze the data processing including data collection, data classification, data pre-processing, and data representation (RQ2). In RQ3, we depicted a range of key research topics using DNNs and investigated the relationships between DL-based model adoption and multiple factors (i.e., DL architectures, task types, problem types, and data types). We also summarized commonly used datasets for different SE tasks. In RQ4, we summarized the widely used optimization algorithms and provided important evaluation metrics for different problem types, including regression, classification, recommendation, and generation. Based on our findings, we present a set of current challenges remaining to be investigated and outline a proposed research road map highlighting key opportunities for future work.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science,Theoretical Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3505243

Reference13 articles.

1. Using Natural Language Processing to Automatically Detect Self-Admitted Technical Debt

2. Wei Fu and Tim Menzies. 2017. Easy over hard: A case study on deep learning. In FSE. 49–60. Wei Fu and Tim Menzies. 2017. Easy over hard: A case study on deep learning. In FSE. 49–60.

3. DECKARD: Scalable and Accurate Tree-Based Detection of Code Clones

Cited by 62 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Data preparation for Deep Learning based Code Smell Detection: A systematic literature review;Journal of Systems and Software;2024-10

2. Coverage-enhanced fault diagnosis for Deep Learning programs: A learning-based approach with hybrid metrics;Information and Software Technology;2024-09

3. Interpretable software estimation with graph neural networks and orthogonal array tunning method;Information Processing & Management;2024-09

4. A systematic review of machine learning methods in software testing;Applied Soft Computing;2024-09

5. Systematic Literature Review of Commercial Participation in Open Source Software;ACM Transactions on Software Engineering and Methodology;2024-08-30