Generating Question Titles for Stack Overflow from Mined Code Snippets-Reference-Cited by-同舟云学术

Generating Question Titles for Stack Overflow from Mined Code Snippets

Published:2020-10-31 Issue:4 Volume:29 Page:1-37
ISSN:1049-331X
Container-title:ACM Transactions on Software Engineering and Methodology
language:en
Short-container-title:ACM Trans. Softw. Eng. Methodol.

Author:

Gao Zhipeng¹,Xia Xin¹^ORCID,Grundy John¹,Lo David²,Li Yuan-Fang¹

Affiliation:

1. Monash University, Australia

2. Singapore Management University, Singapore

Abstract

Stack Overflow has been heavily used by software developers as a popular way to seek programming-related information from peers via the internet. The Stack Overflow community recommends users to provide the related code snippet when they are creating a question to help others better understand it and offer their help. Previous studies have shown that a significant number of these questions are of low-quality and not attractive to other potential experts in Stack Overflow. These poorly asked questions are less likely to receive useful answers and hinder the overall knowledge generation and sharing process. Considering one of the reasons for introducing low-quality questions in SO is that many developers may not be able to clarify and summarize the key problems behind their presented code snippets due to their lack of knowledge and terminology related to the problem, and/or their poor writing skills, in this study we propose an approach to assist developers in writing high-quality questions by automatically generating question titles for a code snippet using a deep sequence-to-sequence learning approach. Our approach is fully data-driven and uses an attention mechanism to perform better content selection, a copy mechanism to handle the rare-words problem and a coverage mechanism to eliminate word repetition problem. We evaluate our approach on Stack Overflow datasets over a variety of programming languages (e.g., Python, Java, Javascript, C# and SQL) and our experimental results show that our approach significantly outperforms several state-of-the-art baselines in both automatic and human evaluation. We have released our code and datasets to facilitate other researchers to verify their ideas and inspire the follow up work.

Publisher

Association for Computing Machinery (ACM)

Subject

Software

Link

https://dl.acm.org/doi/pdf/10.1145/3401026

Reference76 articles.

1. Why, when, and what: Analyzing Stack Overflow questions by topic, type, and code

2. Discovering value from community activity on focused question answering sites

3. The Good, the Bad and their Kins

Cited by 36 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Automatic smart contract comment generation via large language models and in-context learning;Information and Software Technology;2024-04

2. A survey on machine learning techniques applied to source code;Journal of Systems and Software;2024-03

3. AttSum: A Deep Attention-Based Summarization Model for Bug Report Title Generation;IEEE Transactions on Reliability;2023-12

4. Recommending Analogical APIs via Knowledge Graph Embedding;Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering;2023-11-30

5. Automatic recognizing relevant fragments of APIs using API references;Automated Software Engineering;2023-11-19