TY - JOUR
T1 - Electronic, redox, and optical property prediction of organic π-conjugated molecules through a hierarchy of machine learning approaches
AU - Bhat, Vinayak
AU - Sornberger, Parker
AU - Pokuri, Balaji Sesha Sarath
AU - Duke, Rebekah
AU - Ganapathysubramanian, Baskar
AU - Risko, Chad
N1 - Publisher Copyright:
© 2023 The Royal Society of Chemistry.
PY - 2022/11/17
Y1 - 2022/11/17
N2 - Accelerating the development of π-conjugated molecules for applications such as energy generation and storage, catalysis, sensing, pharmaceuticals, and (semi)conducting technologies requires rapid and accurate evaluation of the electronic, redox, or optical properties. While high-throughput computational screening has proven to be a tremendous aid in this regard, machine learning (ML) and other data-driven methods can further enable orders of magnitude reduction in time while at the same time providing dramatic increases in the chemical space that is explored. However, the lack of benchmark datasets containing the electronic, redox, and optical properties that characterize the diverse, known chemical space of organic π-conjugated molecules limits ML model development. Here, we present a curated dataset containing 25k molecules with density functional theory (DFT) and time-dependent DFT (TDDFT) evaluated properties that include frontier molecular orbitals, ionization energies, relaxation energies, and low-lying optical excitation energies. Using the dataset, we train a hierarchy of ML models, ranging from classical models such as ridge regression to sophisticated graph neural networks, with molecular SMILES representation as input. We observe that graph neural networks augmented with contextual information allow for significantly better predictions across a wide array of properties. Our best-performing models also provide an uncertainty quantification for the predictions. To democratize access to the data and trained models, an interactive web platform has been developed and deployed.
AB - Accelerating the development of π-conjugated molecules for applications such as energy generation and storage, catalysis, sensing, pharmaceuticals, and (semi)conducting technologies requires rapid and accurate evaluation of the electronic, redox, or optical properties. While high-throughput computational screening has proven to be a tremendous aid in this regard, machine learning (ML) and other data-driven methods can further enable orders of magnitude reduction in time while at the same time providing dramatic increases in the chemical space that is explored. However, the lack of benchmark datasets containing the electronic, redox, and optical properties that characterize the diverse, known chemical space of organic π-conjugated molecules limits ML model development. Here, we present a curated dataset containing 25k molecules with density functional theory (DFT) and time-dependent DFT (TDDFT) evaluated properties that include frontier molecular orbitals, ionization energies, relaxation energies, and low-lying optical excitation energies. Using the dataset, we train a hierarchy of ML models, ranging from classical models such as ridge regression to sophisticated graph neural networks, with molecular SMILES representation as input. We observe that graph neural networks augmented with contextual information allow for significantly better predictions across a wide array of properties. Our best-performing models also provide an uncertainty quantification for the predictions. To democratize access to the data and trained models, an interactive web platform has been developed and deployed.
UR - http://www.scopus.com/inward/record.url?scp=85144073944&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=85144073944&partnerID=8YFLogxK
U2 - 10.1039/d2sc04676h
DO - 10.1039/d2sc04676h
M3 - Article
AN - SCOPUS:85144073944
SN - 2041-6520
VL - 14
SP - 203
EP - 213
JO - Chemical Science
JF - Chemical Science
IS - 1
ER -