Scinovex
article Open AccessTop 1% cited

“Found in Translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models

Chemical Science · 2018 · Vol. 9(28) · pp. 6091–6098
Philippe SchwallerThéophile GaudinDávid LányiCostas BekasTeodoro Laino

Abstract

There is an intuitive analogy of an organic chemist's understanding of a compound and a language speaker's understanding of a word. Based on this analogy, it is possible to introduce the basic concepts and analyze potential impacts of linguistic analysis to the world of organic chemistry. In this work, we cast the reaction prediction task as a translation problem by introducing a template-free sequence-to-sequence model, trained end-to-end and fully data-driven. We propose a tokenization, which is arbitrarily extensible with reaction information. Using an attention-based model borrowed from human language translation, we improve the state-of-the-art solutions in reaction prediction on the top-1 accuracy by achieving 80.3% without relying on auxiliary knowledge, such as reaction templates or explicit atomic features. Also, a top-1 accuracy of 65.4% is reached on a larger and noisier dataset.

Machine Learning in Materials ScienceComputational Drug Discovery MethodsTopic ModelingSequence (biology)Translation (biology)ChemistryComputational biologyComputer scienceArtificial intelligenceBiologyBiochemistry
Citations
465
FWCI
21.70
field-weighted impact
References
34
Percentile
100%
vs. same field & year
Citations per year
References
Greedy function approximation: A gradient boosting machine.
The Annals of Statistics · 2001 · 27,794 citations
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
Neural‐Symbolic Machine Learning for Retrosynthesis and Reaction Prediction
Chemistry - A European Journal · 2017 · 608 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.