State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis-Reference-Cited by-同舟云学术

State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis

Published:2020-11-04 Issue:1 Volume:11 Page:
ISSN:2041-1723
Container-title:Nature Communications
language:en
Short-container-title:Nat Commun

Author:

Tetko Igor V.^ORCID,Karpov Pavel,Van Deursen Ruud^ORCID,Godin Guillaume^ORCID

Abstract

Abstract We investigated the effect of different training scenarios on predicting the (retro)synthesis of chemical compounds using text-like representation of chemical reactions (SMILES) and Natural Language Processing (NLP) neural network Transformer architecture. We showed that data augmentation, which is a powerful method used in image processing, eliminated the effect of data memorization by neural networks and improved their performance for prediction of new sequences. This effect was observed when augmentation was used simultaneously for input and the target data simultaneously. The top-5 accuracy was 84.8% for the prediction of the largest fragment (thus identifying principal transformation for classical retro-synthesis) for the USPTO-50k test dataset, and was achieved by a combination of SMILES augmentation and a beam search algorithm. The same approach provided significantly better results for the prediction of direct reactions from the single-step USPTO-MIT test set. Our model achieved 90.6% top-1 and 96.1% top-5 accuracy for its challenging mixed set and 97% top-5 accuracy for the USPTO-MIT separated set. It also significantly improved results for USPTO-full set single-step retrosynthesis for both top-1 and top-10 accuracies. The appearance frequency of the most abundantly generated SMILES was well correlated with the prediction outcome and can be used as a measure of the quality of reaction prediction.

Publisher

Springer Science and Business Media LLC

Subject

General Physics and Astronomy,General Biochemistry, Genetics and Molecular Biology,General Chemistry

Link

http://www.nature.com/articles/s41467-020-19266-y.pdf

Reference40 articles.

1. Corey, E. J. & Cheng, X.-M. The Logic of Chemical Synthesis. (John Wiley & Sons, New York, 1995).

2. Corey, E. J., Long, A. K. & Rubenstein, S. D. Computer-assisted analysis in organic synthesis. Science 228, 408–418 (1985).

3. Segler, M. H. S. & Waller, M. P. Neural-symbolic machine learning for retrosynthesis and reaction prediction. Chemistry 23, 5966–5971 (2017).

4. Coley, C. W., Barzilay, R., Jaakkola, T. S., Green, W. H. & Jensen, K. F. Prediction of organic reaction outcomes using machine learning. ACS Cent. Sci. 3, 434–443 (2017).

5. Segler, M. H. S., Preuss, M. & Waller, M. P. Planning chemical syntheses with deep neural networks and symbolic AI. Nature 555, 604–610 (2018).

Cited by 212 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Temporal graphs anomaly emergence detection: benchmarking for social media interactions;Applied Intelligence;2024-09-12

2. SAGB: self-attention with gate and BiGRU network for intrusion detection;Complex & Intelligent Systems;2024-09-09

3. Site-specific template generative approach for retrosynthetic planning;Nature Communications;2024-09-06

4. Reproducing Reaction Mechanisms with Machine‐Learning Models Trained on a Large‐Scale Mechanistic Dataset;Angewandte Chemie;2024-09-03

5. Reproducing Reaction Mechanisms with Machine‐Learning Models Trained on a Large‐Scale Mechanistic Dataset;Angewandte Chemie International Edition;2024-09-02