Skip to main content
Publications

All Publications

Propnet: A Knowledge Graph for Materials Science
Energy & Materials | February 5, 2020

TRI Author: Montoya, J. H.

All Authors: Mrdjenovich, D., Horton, M. K., Montoya, J. H., Legaspi, C. M., Dwaraknath, S., Tshitoyan, V., Jain A., Persson, K. A

Data-driven materials science is bolstered by the recent growth of online materials databases. However, the current informatics infrastructure has yet to unlock the full knowledge available within existing datasets or to explore connections between different materials science domains. Here, we present a streamlined system for codifying and connecting materials properties in an open-source Python framework: propnet. We demonstrate the capability of this framework to augment existing datasets of materials properties: by consecutively applying a network of physical relationships to calculate related information, propnet connects disparate domain knowledge. Beyond an immediate increase in available information, the results allow for the examination of correlations between sets of properties and guide the design of multifunctional materials. By emphasizing code extensibility and simplicity, we offer this software to the materials science community for general application to any experimental or computationally derived materials database.  Read More

Citations: Mrdjenovich, David, Matthew K. Horton, Joseph H. Montoya, Christian M. Legaspi, Shyam Dwaraknath, Vahe Tshitoyan, Anubhav Jain, and Kristin A. Persson. "propnet: A Knowledge Graph for Materials Science." Matter (2020).

 

Image
Propnet: A Knowledge Graph for Materials Science
Random Forest Machine Learning Models for Interpretable X‑Ray Absorption Near‑Edge Structure Spectrum‑Property Relationships
Energy & Materials | February 2, 2020

TRI Authors: Brian Rohr, Joseph Montoya, Santosh Suram, Linda Hung

All Authors: Steven Torrisi, Matthew Carbone, Brian Rohr, Joseph Montoya, Yang Ha, Junko Yano, Santosh Suram, Linda Hung

X-ray absorption spectroscopy (XAS) produces a wealth of information about the local structure of materials, but interpretation of spectra often relies on easily accessible trends and prior assumptions about the structure. Recently, researchers have demonstrated that machine learning models can automate this process to predict the coordinating environments of absorbing atoms from their XAS spectra. However, machine learning models are often difficult to interpret, making it challenging to determine when they are valid and whether they are consistent with physical theories. In this work, we present three main advances to the data-driven analysis of XAS spectra: we demonstrate the efficacy of random forests in solving two new property determination tasks (predicting Bader charge and mean nearest neighbor distance), we show that multiscale featurization can elucidate the regions and trends in spectra that encode various local properties, and we address the effect of normalization on model interpretability. The multiscale featurization transforms the spectrum into a vector of polynomial-fit features, and is contrasted with the commonly-used "pointwise" featurization that directly uses the entire spectrum as input. We find that across thousands of transition metal oxide spectra, the relative importance of features describing the curvature of the spectrum can be localized to individual energy ranges, and we can separate the importance of constant, linear, quadratic, and cubic trends, as well as the white line energy. This work has the potential to assist rigorous theoretical interpretations, expedite experimental data collection, and automate analysis of XAS spectra, thus accelerating discovery of new functional materials.   Read More

Citation: Torrisi, Steven, Matthew Carbone, Brian Rohr, Joseph H. Montoya, Yang Ha, Junko Yano, Santosh Suram, and Linda Hung. "Random Forest Machine Learning Models for Interpretable X-Ray Absorption Near-Edge Structure Spectrum-Property Relationships." chemRxiv preprint (2020). doi:10.26434/chemrxiv.11873691.v1

 

Image
random forest publication image
Benchmarking the acceleration of materials discovery by sequential learning
Energy & Materials | January 29, 2020

TRI Authors: Muratahan Aykol, Santosh K. Suram* All Authors: Brian Rohr, Helge S. Stein, Dan Guevarra, Yu Wang, Joel A. Haber, Muratahan Aykol, Santosh K. Suram* and John M. Gregoire* 

Sequential learning (SL) strategies, i.e. iteratively updating a machine learning model to guide experiments, have been proposed to significantly accelerate materials discovery and research. Applications on computational datasets and a handful of optimization experiments have demonstrated the promise of SL, motivating a quantitative evaluation of its ability to accelerate materials discovery, specifically in the case of physical experiments. The benchmarking effort in the present work quantifies the performance of SL algorithms with respect to a breadth of research goals: discovery of any “good” material, discovery of all “good” materials, and discovery of a model that accurately predicts the performance of new materials. To benchmark the effectiveness of different machine learning models against these goals, we use datasets in which the performance of all materials in the search space is known from high-throughput synthesis and electrochemistry experiments. Each dataset contains all pseudo-quaternary metal oxide combinations from a set of six elements (chemical space), the performance metric chosen is the electrocatalytic activity (overpotential) for the oxygen evolution reaction (OER). A diverse set of SL schemes is tested on four chemical spaces, each containing 2121 catalysts. The presented work suggests that research can be accelerated by up to a factor of 20 compared to random acquisition in specific scenarios. The results also show that certain choices of SL models are ill-suited for a given research goal resulting in substantial deceleration compared to random acquisition methods. The results provide quantitative guidance on how to tune an SL strategy for a given research goal and demonstrate the need for a new generation of materials-aware SL algorithms to further accelerate materials discovery. Read More Citation: Rohr, Brian, Helge S. Stein, Dan Guevarra, Yu Wang, Joel A. Haber, Muratahan Aykol, Santosh K. Suram, and John M. Gregoire. "Benchmarking the acceleration of materials discovery by sequential learning." Chemical Science 11, no. 10 (2020): 2696-2706.

 

Image
Benchmarking the acceleration of materials discovery by sequential learning
BEEP: A Python library for Battery Evaluation and Early Prediction
Energy & Materials | January 1, 2020

Battery evaluation and early prediction software package (BEEP) provides an open-source Python-based framework for the management and processing of high-throughput battery cycling data-streams. BEEPs features include file-system based organization of raw cycling data and metadata received from cell testing equipment, validation protocols that ensure the integrity of such data, parsing and structuring of data into Python-objects ready for analytics, featurization of structured cycling data to serve as input for machine-learning, and end-to-end examples that use processed data for anomaly detection and featurized data to train early-prediction models for cycle life. BEEP is developed in response to the software and expertise gap between cell-level battery testing and data-driven battery development. READ MORE

Image
beep article image
Efficient Pourbaix diagrams of many‑element compounds
Energy & Materials | August 31, 2019

TRI Author: Joseph Montoya

All Authors: Anjli Patel, Jens Nørskov, Kristin Persson, Joseph Montoya

Pourbaix diagrams have been used extensively to evaluate stability regions of materials subject to varying potential and pH conditions in aqueous environments. However, both recent advances in high-throughput material exploration and increasing complexity of materials of interest for electrochemical applications pose challenges for performing Pourbaix analysis on multidimensional systems. Specifically, current Pourbaix construction algorithms incur significant computational costs for systems consisting of four or more elemental components. Herein, we propose an alternative Pourbaix construction method that filters all potential combinations of species in a system to only those present on a compositional convex hull. By including axes representing the quantities of H+ and e− required to form a given phase, one can ensure every stable phase mixture is included in the Pourbaix diagram and reduce the computational time required to construct the resultant Pourbaix diagram by several orders of magnitude. This new Pourbaix algorithm has been incorporated into the pymatgen code and the Materials Project website, and it extends the ability to evaluate the Pourbaix stability of complex multicomponent systems.  Read More

Citation: Patel, Anjli M., Jens K. Nørskov, Kristin A. Persson, and Joseph H. Montoya. "Efficient Pourbaix diagrams of many-element compounds." Physical Chemistry Chemical Physics 21, no. 45 (2019): 25323-25327.

 

Image
Efficient Pourbaix diagrams of many‑element compounds
Network analysis of synthesizable materials discovery
Energy & Materials | May 1, 2019

TRI Authors: Muratahan Aykol,Linda Hung, Santosh Suram, Patrick Herring, Jens S. Hummelshoj

All Authors: Muratahan Aykol, Vinay I. Hegde, Linda Hung, Santosh Suram, Patrick Herring, Chris Wolverton, Jens S. Hummelshoj

Assessing the synthesizability of inorganic materials is a grand challenge for accelerating their discovery using computations. Synthesis of a material is a complex process that depends not only on its thermodynamic stability with respect to others, but also on factors from kinetics, to advances in synthesis techniques, to the availability of precursors. This complexity makes the development of a general theory or first-principles approach to synthesizability currently impractical. Here we show how an alternative pathway to predicting synthesizability emerges from the dynamics of the materials stability network: a scale-free network constructed by combining the convex free-energy surface of inorganic materials computed by high-throughput density functional theory and their experimental discovery timelines extracted from citations. The time-evolution of the underlying network properties allows us to use machine-learning to predict the likelihood that hypothetical, computer-generated materials will be amenable to successful experimental synthesis. 

Citation: Aykol, Muratahan, Vinay I. Hegde, Linda Hung, Santosh Suram, Patrick Herring, Chris Wolverton, and Jens S. Hummelshøj. "Network analysis of synthesizable materials discovery." Nature communications 10, no. 1 (2019): 1-7.

 

Image
TRI logo
CRYSTAL: a multi‑agent AI system for automated mapping of materials' crystal structures
Energy & Materials | April 24, 2019

TRI Author: Santosh Suram

All Authors: Carla P. Gomes, Junwen Bai, Yexiang Xue, Johan Björck, Brendan Rappazzo, Sebastian Ament, Richard Bernstein, Shufeng Kong, Santosh K. Suram, R. Bruce van Dover, and John M. Gregoire

We introduce CRYSTAL, a multi-agent AI system for crystal-structure phase mapping. CRYSTAL is the first system that can automatically generate a portfolio of physically meaningful phase diagrams for expert-user exploration and selection. CRYSTAL outperforms previous methods to solve the example Pd-Rh-Ta phase diagram, enabling the discovery of a mixed-intermetallic methanol oxidation electrocatalyst. The integration of multiple data-knowledge sources and learning and reasoning algorithms, combined with the exploitation of problem decompositions, relaxations, and parallelism, empowers AI to supersede human scientific data interpretation capabilities and enable otherwise inaccessible scientific discovery in materials science and beyond.  Read More

Citation: Gomes, Carla P., Junwen Bai, Yexiang Xue, Johan Björck, Brendan Rappazzo, Sebastian Ament, Richard Bernstein et al. "CRYSTAL: a multi-agent AI system for automated mapping of materials' crystal structures." MRS Communications 9, no. 2 (2019): 600-608.

 

Image
Crystal image
Genetic algorithms for computational materials discovery accelerated by machine learning
Energy & Materials | April 10, 2019

TRI Author: Jens Hummelshøj

All Authors: Paul C Jennings, Steen Lysgaard, Jens Strabo Hummelshøj, Tejs Vegge, Thomas Bligaard

Materials discovery is increasingly being impelled by machine learning methods that rely on pre-existing datasets. Where datasets are lacking, unbiased data generation can be achieved with genetic algorithms. Here a machine learning model is trained on-the-fly as a computationally inexpensive energy predictor before analyzing how to augment convergence in genetic algorithm-based approaches by using the model as a surrogate. This leads to a machine learning accelerated genetic algorithm combining robust qualities of the genetic algorithm with rapid machine learning. The approach is used to search for stable, compositionally variant, geometrically similar nanoparticle alloys to illustrate its capability for accelerated materials discovery, e.g., nanoalloy catalysts. The machine learning accelerated approach, in this case, yields a 50-fold reduction in the number of required energy calculations compared to a traditional “brute force” genetic algorithm. This makes searching through the space of all homotops and compositions of a binary alloy particle in a given structure feasible, using density functional theory calculations. Read More

Citation: Jennings, Paul C., Steen Lysgaard, Jens Strabo Hummelshøj, Tejs Vegge, and Thomas Bligaard. "Genetic algorithms for computational materials discovery accelerated by machine learning." npj Computational Materials 5, no. 1 (2019): 1-6.

 

Image
Genetic algorithms for computational materials discovery accelerated by machine learning
Data‑Driven Prediction of Battery Cycle Life Before Capacity Degradation
Energy & Materials | March 25, 2019

TRI Authors: Muratahan Aykol, Patrick K. Herring

All Authors: Kristen A. Severson, Peter M. Attia, Norman Jin, Nicholas Perkins, Benben Jiang, Zi Yang, Michael H. Chen, Muratahan Aykol, Patrick K. Herring, Dimitrios Fraggedakis, Martin Z. Bazant, Stephen J. Harris, William C. Chueh & Richard D. Braatz

Accurately predicting the lifetime of complex, nonlinear systems such as lithium-ion batteries is critical for accelerating technology development. However, diverse aging mechanisms, significant device variability and dynamic operating conditions have remained major challenges. We generate a comprehensive dataset consisting of 124 commercial lithium iron phosphate/graphite cells cycled under fast-charging conditions, with widely varying cycle lives ranging from 150 to 2,300 cycles. Using discharge voltage curves from early cycles yet to exhibit capacity degradation, we apply machine-learning tools to both predict and classify cells by cycle life. Our best models achieve 9.1% test error for quantitatively predicting cycle life using the first 100 cycles (exhibiting a median increase of 0.2% from initial capacity) and 4.9% test error using the first 5 cycles for classifying cycle life into two groups. This work highlights the promise of combining deliberate data generation with data-driven modelling to predict the behaviour of complex dynamical systems. Read More

Citation: Severson, Kristen A., Peter M. Attia, Norman Jin, Nicholas Perkins, Benben Jiang, Zi Yang, Michael H. Chen et al. "Data-driven prediction of battery cycle life before capacity degradation." Nature Energy 4, no. 5 (2019): 383-391.

Image
Data‑Driven Prediction of Battery Cycle Life Before Capacity Degradation
Machine Learning Accelerated Genetic Algorithms for Computational Materials Search
Energy & Materials | December 10, 2018

TRI Author: Jens Hummelshøj

All Authors: Steen Lysgaard, Paul C Jennings, Jens Strabo Hummelshøj, Thomas Bligaard, Tejs Vegge

A machine learning model is used as a surrogate fitness evaluator in a genetic algorithm (GA) optimization of the atomic distribution of Pt-Au nanoparticles. The machine learning accelerated genetic algorithm (MLaGA) yields a 50-fold reduction of required energy calculations compared to a traditional GA. Read More

Citation: Lysgaard, Steen, Paul C. Jennings, Jens Strabo Hummelshøj, Thomas Bligaard, and Tejs Vegge. "Machine Learning Accelerated Genetic Algorithms for Computational Materials Search." In ChemRxiv(2018).

 

Image
Machine Learning Accelerated Genetic Algorithms for Computational Materials Search