<scp>PyScribe</scp>–Learning to describe python code-Reference-Cited by-同舟云学术

PyScribe–Learning to describe python code

Published:2023-12-09 Issue: Volume: Page:
ISSN:0038-0644
Container-title:Software: Practice and Experience
language:en
Short-container-title:Softw Pract Exp

Author:

Guo Juncai¹^ORCID,Liu Jin¹²,Liu Xiao³,Wan Yao⁴,Zhao Yanjie⁵,Li Li⁶,Liu Kui⁷,Klein Jacques⁸^ORCID,Bissyandé Tegawendé F.⁸

Affiliation:

1. School of Computer Science Wuhan University Wuhan China

2. Key Laboratory of Network Assessment Technology Institute of Information Engineering, Chinese Academy of Sciences Beijing China

3. School of Information Technology Deakin University Burwood Melbourne Australia

4. School of Computer Science and Technology Huazhong University of Science and Technology Wuhan China

5. Faculty of Information Technology Monash University Clayton Victoria Australia

6. School of Software Beihang University Beijing China

7. Huawei Software Engineering Application Technology LabHangzhou China

8. SnT Centre University of Luxembourg Esch‐sur‐Alzette Luxembourg

Abstract

AbstractCode comment generation, which attempts to summarize the functionality of source code in textual descriptions, plays an important role in automatic software development research. Currently, several structural neural networks have been exploited to preserve the syntax structure of source code based on abstract syntax trees (ASTs). However, they can not well capture both the long‐distance and local relations between nodes while retaining the overall structural information of AST. To mitigate this problem, we present a prototype tool titled PyScribe, which extends the Transformer model to a new encoder‐decoder‐based framework. Particularly, the triplet position is designed and integrated into the node‐level and edge‐level structural features of AST for producing Python code comments automatically. This paper, to the best of our knowledge, makes the first effort to model the edges of AST as an explicit component for improved code representation. By specifying triplet positions for each node and edge, the overall structural information can be well preserved in the learning process. Moreover, the captured node and edge features go through a two‐stage decoding process to yield higher qualified comments. To evaluate the effectiveness of PyScribe, we resort to a large dataset of code‐comment pairs by mining Jupyter Notebooks from GitHub, for which we have made it publicly available to support further studies. The experimental results reveal that PyScribe is indeed effective, outperforming the state‐ofthe‐art by achieving an average BLEU score (i.e., av‐BLEU) of 0.28.

Funder

China Scholarship Council

National Natural Science Foundation of China

Publisher

Wiley

Subject

Software

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/spe.3291

Reference69 articles.

1. Self-Documenting Code

2. Comments are More Important than Code

3. Automatic documentation generation via source code summarization of method context

4. Towards automatically generating summary comments for Java methods

5. Automatically detecting and describing high level actions within methods