Bin2vec: learning representations of binary executable programs for security tasks-Reference-Cited by-同舟云学术

Bin2vec: learning representations of binary executable programs for security tasks

Published:2021-07-01 Issue:1 Volume:4 Page:
ISSN:2523-3246
Container-title:Cybersecurity
language:en
Short-container-title:Cybersecur

Author:

Arakelyan Shushan,Arasteh Sima,Hauser Christophe,Kline Erik,Galstyan Aram

Abstract

AbstractTackling binary program analysis problems has traditionally implied manually defining rules and heuristics, a tedious and time consuming task for human analysts. In order to improve automation and scalability, we propose an alternative direction based on distributed representations of binary programs with applicability to a number of downstream tasks. We introduce Bin2vec, a new approach leveraging Graph Convolutional Networks (GCN) along with computational program graphs in order to learn a high dimensional representation of binary executable programs. We demonstrate the versatility of this approach by using our representations to solve two semantically different binary analysis tasks – functional algorithm classification and vulnerability discovery. We compare the proposed approach to our own strong baseline as well as published results, and demonstrate improvement over state-of-the-art methods for both tasks. We evaluated Bin2vec on 49191 binaries for the functional algorithm classification task, and on 30 different CWE-IDs including at least 100 CVE entries each for the vulnerability discovery task. We set a new state-of-the-art result by reducing the classification error by 40% compared to the source-code based inst2vec approach, while working on binary code. For almost every vulnerability class in our dataset, our prediction accuracy is over 80% (and over 90% in multiple classes).

Funder

USC Information Sciences Institute

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence,Computer Networks and Communications,Information Systems,Software

Link

https://link.springer.com/content/pdf/10.1186/s42400-021-00088-4.pdf

Reference55 articles.

1. Aafer, Y, Du W, Yin H (2013) Droidapiminer: Mining api-level features for robust malware detection in android. In: Zia TA, Zomaya AY, Varadharajan V, Mao ZM (eds)Security and Privacy in Communication Networks - 9th International ICST Conference, SecureComm 2013, Sydney, NSW, Australia, September 25-28, 2013, Revised Selected Papers, Springer, vol 127, 86–103. https://doi.org/10.1007/978-3-319-04283-1_6.

2. Abu-El-Haija, S, Perozzi B, Kapoor A, Alipourfard N, Lerman K, Harutyunyan H, Steeg GV, Galstyan A (2019) MixHop: Higher-order graph convolutional architectures via sparsified neighborhood mixing. PMLR Long Beach Calif USA Proc Mach Learn Res 97:21–29.

3. Allamanis, M, Barr ET, Devanbu PT, Sutton CA (2018) A survey of machine learning for big code and naturalness. ACM Comput Surv 51(4):81:1–81:37. https://doi.org/10.1145/3212695.

4. Andriesse, D, Chen X, van der Veen V, Slowinska A, Bos H (2016) An in-depth analysis of disassembly on full-scale x86/x64 binaries. In: USENIX In: USENIX Association, Austin. https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/andriesse.

5. Ben-Nun, T, Jakobovits AS, Hoefler T (2018) Neural code comprehension: A learnable representation of code semantics. In: Bengio S, Wallach HM, Larochelle H, Grauman K, Cesa-Bianchi N, Garnett R (eds)Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, 3-8 December, 2018, Montréal, Canada, 3589–3601. http://papers.nips.cc/paper/7617-neural-code-comprehension-a-learnable-representation-of-code-semantics.

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improving Malware Detection from Binary Control Flow Graphs Using Supervised Learning;2024 Intermountain Engineering, Technology and Computing (IETC);2024-05-13

2. VulANalyzeR: Explainable Binary Vulnerability Detection with Multi-task Learning and Attentional Graph Convolution;ACM Transactions on Privacy and Security;2023-04-14

3. Deep-Learning-Based Vulnerability Detection in Binary Executables;Foundations and Practice of Security;2023

4. HyperDbg: Reinventing Hardware-Assisted Debugging;Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security;2022-11-07

5. Automatically Detect Software Security Vulnerabilities Based on Natural Language Processing Techniques and Machine Learning Algorithms;Journal of ICT Research and Applications;2022-05-11