Featured Publications
All Publications
We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a question. In this work, we leverage the strong semantic reasoning capabilities of large vision-language models (VLMs) to efficiently explore and answer such questions. However, there are two main challenges when using VLMs in EQA: they do not have an internal memory for mapping the scene to be able to plan how to explore over time, and their confidence can be miscalibrated and can cause the robot to prematurely stop exploration or over-explore. We propose a method that first builds a semantic map of the scene based on depth information and via visual prompting of a VLM - leveraging its vast knowledge of relevant regions of the scene for exploration. Next, we use conformal prediction to calibrate the VLM's question answering confidence, allowing the robot to know when to stop exploration - leading to a more calibrated and efficient exploration strategy. To test our framework in simulation, we also contribute a new EQA dataset with diverse, realistic human-robot scenarios and scenes built upon the Habitat-Matterport 3D Research Dataset (HM3D). Both simulated and real robot experiments show our proposed approach improves the performance and efficiency over baselines that do no leverage VLM for exploration or do not calibrate its confidence. READ MORE
We present a 3D shape completion method that recovers the complete geometry of multiple objects in complex scenes from a single RGB-D image. Despite notable advancements in single object 3D shape completion, high-quality reconstructions in highly cluttered real-world multi-object scenes remains a challenge. To address this issue, we propose OctMAE, an architecture that leverages an Octree U-Net and a latent 3D MAE to achieve high-quality and near real-time multi-object shape completion through both local and global geometric reasoning. Because a naïve 3D MAE can be computationally intractable and memory intensive even in the latent space, we introduce a novel occlusion masking strategy and adopt 3D rotary embeddings, which significantly improves the runtime and shape completion quality. To generalize to a wide range of objects in diverse scenes, we create a large-scale photorealistic dataset, featuring a diverse set of 12K 3D object models from the Objaverse dataset which are rendered in multi-object scenes with physics-based positioning. Our method outperforms the current state-of-the-art on both synthetic and real-world datasets and demonstrates a strong zero-shot capability. READ MORE
The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DROID (Distributed Robot Interaction Dataset), a diverse robot manipulation dataset with 76k demonstration trajectories or 350 hours of interaction data, collected across 564 scenes and 84 tasks by 50 data collectors in North America, Asia, and Europe over the course of 12 months. We demonstrate that training with DROID leads to policies with higher performance and improved generalization ability. We open source the full dataset, policy learning code, and a detailed guide for reproducing our robot hardware setup. READ MORE
Solid polymer electrolytes hold significant promise as materials for next-generation batteries due to their superior safety performance, enhanced specific energy, and extended lifespans compared to liquid electrolytes. However, the material's low ionic conductivity impedes its commercialization, and the vast polymer space poses significant challenges for the screening and design. In this study, we assess the capabilities of generative artificial intelligence (AI) for the de novo design of polymer electrolytes. To optimize the generation, we compare different deep learning architectures, including both GPT-based and diffusion-based models, and benchmark the results with hyperparameter tuning. We further employ various evaluation metrics and full-atom molecular dynamics simulations to assess the performance of different generative model architectures and to validate the top candidates produced by each model. Out of only 45 candidates being tested, we discovered 17 polymers that achieve superior ionic conductivity better than any other polymers in our database, with some of them doubling the conductivity value. In addition, by adopting a pretraining and fine-tuning methodology, we significantly improve the efficacy of our generative models, achieving quicker convergence, enhanced performance with limited data, and greater diversity. Using the proposed method, we can easily generate a large number of novel, diverse, and valid polymers, with a chance of synthesizability, enabling us to identify promising candidates with markedly improved efficiency. READ MORE
Exploratory synthesis has been the main generator of new inorganic materials for decades. However, our Edisonian and bias-prone processes of synthetic exploration alone are no longer sufficient in an age that demands rapid advances in materials development. In this work, we demonstrate an end-to-end attempt towards systematic, computer-aided discovery and laboratory synthesis of inorganic crystalline compounds as a modern alternative to purely exploratory synthesis. Our approach initializes materials discovery campaigns by autonomously mapping the synthetic feasibility of a chemical system using density functional theory with AI feedback. Following expert-driven down-selection of newly generated phases, we use solid-state synthesis and in situ characterization via hot-stage X-ray diffraction in order to realize new ternary oxide phases experimentally. We applied this strategy in six ternary transition-metal oxide chemistries previously considered well-explored, one of which culminated in the discovery of two novel phases of calcium ruthenates. Detailed characterization using room temperature X-ray powder diffraction, 4D-STEM and SQUID measurements identifies the structure and composition and confirms distinct properties, including distinct defect concentrations, of one of the new phases formed in our experimental campaigns. While the discovery of a new material guided by AI and DFT theory represents a milestone, our procedure and results also highlight a number of critical gaps in the process that can inform future efforts towards the improvement of AI-coupled methodologies. READ MORE
The materials research community is increasingly using automation and artificial intelligence (AI) to accelerate research and development. A materials acceleration platform (MAP) typically encompasses several experimental techniques or instruments to establish a synthesis-characterization-evaluation workflow. With the advancement of workflow orchestration software and AI experiment design, the scope and complexity of MAPs are increasing, however each MAP typically operates as a standalone entity with dedicated experiment, compute, and database resources. The data from each MAP is thus siloed until subsequent efforts to integrate data into complex schema such as knowledge graphs. To lower the latency of data integration and establish an extensible community of MAPs, we must expand our automation efforts to include data handling that is decoupled from the resources of each MAP. Event-driven pipelines are well established in the computational community for building decoupled data processing systems. Such pipelines can be difficult to implement de novo due to their distributed nature and complex error handling. Fortunately, the broader computational science community has established a suite of cloud services that are well suited for this task. By leveraging cloud computing resources to establish event-driven data management, the MAP community can better realize the ideals of extensibility and interoperability in materials chemistry research. READ MORE
Lithium-ion batteries (LIBs) have attracted widespread attention as an efficient energy storage device on electric vehicles (EV) to achieve emission-free mobility. However, the performance of LIBs deteriorates with time and usage, and the state of health of used batteries are difficult to quantify. Having accurate estimations of a battery’s remaining life across different life stages would benefit maintenance, safety, and serve as a means of qualifying used batteries for second-life applications. Since the full history of a battery may not always be available in downstream applications, in this study, we demonstrate a deep learning framework that enables dynamic degradation rate prediction, including both short-term and long-term forecasting, while requiring only the most recent battery usage information. Specifically, our model takes a rolling window of current and voltage time-series inputs, and predicts the near-term and long-term capacity fade via a recurrent neural network. We exhaustively benchmark our model against a naive extrapolating model by evaluating the error on reconstructing the discharge capacity profile under different settings. We show that our model’s performance in accurately inferring the battery’s degradation profile is agnostic with respect to cell cycling history and its current state of health. This approach can provide a promising path towards evaluating battery health in running vehicles, enhance edge-computing battery diagnostics, and determine the state of health for used batteries with unknown cycling histories. READ MORE
Generative AI models have made significant progress in automating the creation of 3D shapes, which has the potential to transform car design. In engineering design and optimization, evaluating engineering metrics is crucial. To make generative models performance-aware and enable them to create high-performing designs, surrogate modeling of these metrics is necessary. However, the currently used representations of 3D shapes either require extensive computational resources to learn or suffer from significant information loss, which impairs their effectiveness in surrogate modeling. To address this issue, we propose a new 2D representation of 3D shapes. We develop a surrogate drag model based on this representation to verify its effectiveness in predicting 3D car drag. We construct a diverse dataset of 4,535 high-quality 3D car meshes labeled by drag coefficients computed from computational fluid dynamics simulations to train our model. Our experiments demonstrate that our model can accurately and efficiently evaluate drag coefficients with an R^2 value above 0.84 for various car categories. Our model is implemented using deep neural networks, making it compatible with recent AI image generation tools (such as Stable Diffusion) and a significant step towards the automatic generation of drag-optimized car designs. Moreover, we demonstrate a case study using the proposed surrogate model to guide a diffusion-based deep generative model for drag-optimized car body synthesis. We have made the dataset and code publicly available at https://decode.mit.edu/projects/dragprediction/. READ MORE
In this work, we introduce a polymer discovery platform designed to identify polymers with tailored properties efficiently, exemplified through the discovery of high-performance polymer electrolytes. The platform integrates three core components: a conditioned generative model, validation modules, and a feedback mechanism, creating a self-improving system for material innovation. To demonstrate the efficacy of this platform, it is used to identify polymer electrolyte materials with high ionic conductivity. A simple conditional generative model, based on the minGPT architecture, can effectively generate candidate polymers that exhibit a mean ionic conductivity that is significantly greater than those in the original training set. This approach, coupled with molecular dynamics simulations for validation and a specifically designed acquisition mechanism, allows the platform to refine its output iteratively. Notably, after the first iteration, we observed an increase in both the mean and the lower bound of the ionic conductivity of the new polymer candidates. The platform's effectiveness is underscored by the identification of 19 polymer repeating units, each displaying a computed ionic conductivity surpassing that of Polyethylene Oxide (PEO). The discovery of these polymers validates the platform's efficacy in identifying potential polymer materials. Acknowledging current limitations, future work will focus on enhancing modeling techniques, validation processes, and acquisition strategies, aiming for broader applicability in polymer science and machine learning. READ MORE
Electric vehicles (EVs) are generally considered more environmentally sustainable than internal combustion engine vehicles (ICEVs). Government and policy makers may want to incentivize multi-vehicle households who, if they purchase a new EV, would use their EV to replace a large portion of their ICEV mileage. Therefore, it is important to analyze how EV procurement affects annual EV mileage for different households. Given that many relevant data, especially experimental data, are often unavailable in the real world, we need causal analysis tools to answer this question. Additionally, our aim is to compare the expected EV mileage of different combinations of vehicles a household owns. Observing multiple combinations in an individual household is impossible since only one combination can exist, making causal inference challenging. In this paper, we construct a causal AI framework utilizing counterfactual reasoning methods to address this issue. READ MORE