<scp>deGraphCS</scp> : Embedding Variable-based Flow Graph for Neural Code Search-Reference-Cited by-同舟云学术

deGraphCS : Embedding Variable-based Flow Graph for Neural Code Search

Published:2023-03-30 Issue:2 Volume:32 Page:1-27
ISSN:1049-331X
Container-title:ACM Transactions on Software Engineering and Methodology
language:en
Short-container-title:ACM Trans. Softw. Eng. Methodol.

Author:

Zeng Chen¹^ORCID,Yu Yue¹^ORCID,Li Shanshan¹^ORCID,Xia Xin²^ORCID,Wang Zhiming¹^ORCID,Geng Mingyang¹^ORCID,Bai Linxiao¹^ORCID,Dong Wei¹^ORCID,Liao Xiangke¹^ORCID

Affiliation:

1. School of Computer, National University of Defense Technology, Changsha, China

2. College of Computer Science and Technology, Zhejiang University, Hangzhou, China

Abstract

With the rapid increase of public code repositories, developers maintain a great desire to retrieve precise code snippets by using natural language. Despite existing deep learning-based approaches that provide end-to-end solutions (i.e., accept natural language as queries and show related code fragments), the performance of code search in the large-scale repositories is still low in accuracy because of the code representation (e.g., AST) and modeling (e.g., directly fusing features in the attention stage). In this paper, we propose a novel learnable de ep G raph for C ode S earch (called deGraphCS ) to transfer source code into variable-based flow graphs based on an intermediate representation technique, which can model code semantics more precisely than directly processing the code as text or using the syntax tree representation. Furthermore, we propose a graph optimization mechanism to refine the code representation and apply an improved gated graph neural network to model variable-based flow graphs. To evaluate the effectiveness of deGraphCS , we collect a large-scale dataset from GitHub containing 41,152 code snippets written in the C language and reproduce several typical deep code search methods for comparison. The experimental results show that deGraphCS can achieve state-of-the-art performance and accurately retrieve code snippets satisfying the needs of the users.

Funder

National Key R&D Program of China

National Natural Science Foundation of China

Publisher

Association for Computing Machinery (ACM)

Subject

Software

Link

https://dl.acm.org/doi/pdf/10.1145/3546066

Reference71 articles.

1. Neural machine translation by jointly learning to align and translate;Bahdanau Dzmitry;arXiv preprint arXiv:1409.0473,2014

2. Sourcerer

3. Tal Ben-Nun, Alice Shoshana Jakobovits, and Torsten Hoefler. 2018. Neural code comprehension: A learnable representation of code semantics. In Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.). Curran Associates, Inc., 3585–3597. http://papers.nips.cc/paper/7617-neural-code-comprehension-a-learnable-representation-of-code-semantics.pdf.

4. Joel Brandt, Mira Dontcheva, Marcos Weskamp, and Scott R. Klemmer. 2010. Example-centric programming: Integrating web search into the development environment. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 513–522.

5. BIKER: a tool for Bi-information source based API method recommendation

Cited by 19 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Specialized model initialization and architecture optimization for few-shot code search;Information and Software Technology;2025-01

2. Fine-grained vulnerability detection for medical sensor systems;Internet of Things;2024-12

3. On Representation Learning-based Methods for Effective, Efficient, and Scalable Code Retrieval;Neurocomputing;2024-10

4. SECON: Maintaining Semantic Consistency in Data Augmentation for Code Search;ACM Transactions on Information Systems;2024-08

5. An Empirical Study on Code Search Pre-trained Models: Academic Progresses vs. Industry Requirements;Proceedings of the 15th Asia-Pacific Symposium on Internetware;2024-07-24