Seeing Is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability-Reference-Cited by-同舟云学术

Seeing Is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability

Published:2023-12-30 Issue:1 Volume:26 Page:41
ISSN:1099-4300
Container-title:Entropy
language:en
Short-container-title:Entropy

Author:

Liu Ziming¹,Gan Eric¹,Tegmark Max¹^ORCID

Affiliation:

1. Institute for Artificial Intelligence and Fundamental Interactions, Massachusetts Institute of Technology, Cambridge, MA 02139, USA

Abstract

We introduce Brain-Inspired Modular Training (BIMT), a method for making neural networks more modular and interpretable. Inspired by brains, BIMT embeds neurons in a geometric space and augments the loss function with a cost proportional to the length of each neuron connection. This is inspired by the idea of minimum connection cost in evolutionary biology, but we are the first the combine this idea with training neural networks with gradient descent for interpretability. We demonstrate that BIMT discovers useful modular neural networks for many simple tasks, revealing compositional structures in symbolic formulas, interpretable decision boundaries and features for classification, and mathematical structure in algorithmic datasets. Qualitatively, BIMT-trained networks have modules readily identifiable by the naked eye, but regularly trained networks seem much more complicated. Quantitatively, we use Newman’s method to compute the modularity of network graphs; BIMT achieves the highest modularity for all our test problems. A promising and ambitious future direction is to apply the proposed method to understand large models for vision, language, and science.

Funder

The Casey Family Foundation, the Foundational Questions Institute, the Rothberg Family Fund for Cognitive Science, the NSF Graduate Research Fellowship

IAIFI through NSF

Publisher

MDPI AG

Subject

General Physics and Astronomy

Link

https://www.mdpi.com/1099-4300/26/1/41/pdf

Reference44 articles.

1. Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. (2023, November 01). Zoom In: An Introduction to Circuits. Distill 2020. Available online: https://distill.pub/2020/circuits/zoom-in.

2. Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., and Chen, A. (2023, November 01). In-Context Learning and Induction Heads. Transform. Circuits Thread 2022. Available online: https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.

3. Michaud, E.J., Liu, Z., Girit, U., and Tegmark, M. (2023). The Quantization Model of Neural Scaling. arXiv.

4. Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., and Conerly, T. (2023, November 01). A Mathematical Framework for Transformer Circuits. Transform. Circuits Thread 2021. Available online: https://transformer-circuits.pub/2021/framework/index.html.

5. Wang, K.R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. (2023, January 1–5). Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small. Proceedings of the The Eleventh International Conference on Learning Representations, Kigali, Rwanda.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Machine Advisors: Integrating Large Language Models Into Democratic Assemblies;Social Epistemology;2024-08-18

2. Brain-Inspired Physics-Informed Neural Networks: Bare-Minimum Neural Architectures for PDE Solvers;Lecture Notes in Computer Science;2024