Skip to main content
Publications

All Publications

Designing Rewards for Rewarding Designs: Demonstrating the Impact of Rewards on the Creative Design Process
Human-Centered AI | April 28, 2026

The creative design process involves transforming abstract goals into concrete outcomes through a series of decisions made under constraints. While such processes are commonly shaped by feedback like rewards, their impact on design decision making remains unclear. To better understand the role of rewards in the design process, we modeled a 3D parametric, goal-based chair design task as a Markov Decision Process. We tracked participants' decisions as they iteratively developed designs for an abstract design goal, and presented either a goal-aligned or goal-agnostic reward at every step. We tested the effect of these rewards on task behaviour and self-reported experience. With rewards, participants more thoroughly explored the design space, and maximised goal-aligned over goal-agnostic rewards while preserving diversity across designs. The nature of the goal also mattered, influencing participants' perception of the reward's usefulness. Building on these insights, we propose guidelines for designing effective feedback for design decision making.

Image
Design environment with all features
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
Human Interactive Driving | April 16, 2026

We propose a new approach for solving planning problems with a hierarchical structure, fusing reinforcement learning and MPC planning. Our formulation tightly and elegantly couples the two planning paradigms. It leverages reinforcement learning actions to inform the MPPI sampler, and adaptively aggregates MPPI samples to inform the value estimation. The resulting adaptive process leverages further MPPI exploration where value estimates are uncertain, and improves training robustness and the overall resulting policies. This results in a robust planning approach that can handle complex planning problems and easily adapts to different applications, as demonstrated over several domains, including race driving, modified Acrobot, and Lunar Lander with added obstacles. Our results in these domains show better data efficiency and overall performance in terms of both rewards and task success, with up to a 72% increase in success rate compared to existing approaches, as well as accelerated convergence (x2.1) compared to non-adaptive sampling.

Image
diagram of the combined approach
A Framework for Adapting In-Car Touchscreen Interfaces to Driver Behaviors, Perception, and Cognition
Human-Centered AI | April 13, 2026

Although in-car touchscreens expand interaction possibilities, they risk compromising driver safety and vigilance. We propose a data- and expert-informed framework for designing adaptive touchscreens that respond to a driver’s usage profile and cognitive state, maximizing usability while mitigating safety risks. First, in a driving simulator study, we find that cognitive load slows touchscreen button selections by 20% and produced shorter, more frequent off-road glances. We also find that enlarging buttons improves selection speeds by 0.3 seconds but at the cost of requiring more display pages. Next, these findings informed a co-design session with expert in-cabin designers, generating guidelines for adaptive interfaces that balance usability and safety. These guidelines form the basis of our Profile-State Adaptive (PSA) framework, which integrates driver profiles with cognitive states to guide interface adaptations. We then extend the framework to include a quantitative Time-Cost model as well as design patterns for adaptive layouts across usage profiles and cognitive demands.

Image
Conceptual diagram of the proposed Profile–State Adaptive (PSA) framework
Skill Modulates Coaching Language in Embodied Motor Learning
Human Interactive Driving | April 13, 2026

Effective coaching requires instruction to be tailored to each student’s skill level. While prior work has shown the importance of adaptive feedback, there is little quantitative evidence showing how coaching language itself changes as a function of skill level - particularly in real-time embodied domains. To address this gap, we collected a naturalistic dataset of an expert coach providing verbal instructions to students of varying skill in a high-performance driving simulation. Using linguistic and interaction measures, we show that, as skill improves, coaching language shifts from concrete, instructional feedback, towards evaluative and affective language. At the same time, while overall feedback amount decreases, language diversity increases, indicating a move toward denser, more expressive communication. Together, these findings provide us with a better understanding of how a human coach modulates language as a function of learner skill and provide empirical grounding for the design of adaptive and personalized AI coaching systems.

Image
Study Protocol Flow
SEGAR: Selective Enhancement for Generative Augmented Reality
Human Interactive Driving | March 25, 2026

Generative world models offer a compelling foundation for augmented-reality (AR) applications: by predicting future image sequences that incorporate deliberate visual edits, they enable temporally coherent, augmented future frames that can be computed ahead of time and cached, avoiding per-frame rendering from scratch in real time. In this work, we present SEGAR, a preliminary framework that combines a diffusion-based world model with a selective correction stage to support this vision. The world model generates augmented future frames with region-specific edits while preserving others, and the correction stage subsequently aligns safety-critical regions with real-world observations while preserving intended augmentations elsewhere. We demonstrate this pipeline in driving scenarios as a representative setting where semantic region structure is well defined and real-world feedback is readily available. We view this as an early step toward generative world models as practical AR infrastructure, where future frames can be generated, cached, and selectively corrected on demand.

Image
SEGAR system pipeline overview
ShaLa: Multimodal Shared Latent Generative Modelling
Human-Centered AI | March 14, 2026

This paper presents a novel generative framework for learning shared latent representations across multimodal data. Many advanced multimodal methods focus on capturing all combinations of modality-specific details across inputs, which can inadvertently obscure the high-level semantic concepts that are shared across modalities. Notably, Multimodal VAEs with low-dimensional latent variables are designed to capture shared representations, enabling various tasks such as joint multimodal synthesis and cross-modal inference. However, multimodal VAEs often struggle to design expressive joint variational posteriors and suffer from low-quality synthesis. In this work, ShaLa addresses these challenges by integrating a novel architectural inference model and a second-stage expressive diffusion prior, which not only facilitates effective inference of shared latent representation but also significantly improves the quality of downstream multimodal synthesis. We validate ShaLa extensively across multiple benchmarks, demonstrating superior coherence and synthesis quality compared to state-of-the-art multimodal VAEs. Furthermore, ShaLa scales to many more modalities while prior multimodal VAEs have fallen short in capturing the increasing complexity of the shared latent space.

Image
 Comparison of multimodal modeling paradigms
Short-Range Order and LixTM4−x Probability Maps for Disordered Rocksalt Cathodes
Energy & Materials | March 11, 2026

Short-range order (SRO) in the cation-disordered state is a controlling factor influencing the probability of finding tetrahedron clusters in disordered rocksalt (DRX) cathode materials. However, the prevalent  probability below the random limit across reported DRX compositions has not been systematically investigated, active strategies to surpass the random limit of  probability are lacking, and the fundamental ordering behavior on the face-centered cubic (FCC) lattice remains insufficiently explored. This research quantitatively examines pair SRO parameters and  probabilities via exhaustive Monte Carlo mapping across a simplified subset of the parameter space. The results indicate that, in the disordered state, the  probability is governed by the nearest neighbor (NN) pairwise SRO parameter, and that these quantities do not necessarily represent a simple attenuation of their corresponding low-temperature long-range order, particularly for the important cases of Layered and Spinel-like orderings. Strategies are proposed to mitigate or even reverse the lithium and transition metals mixing tendency of NN pair SRO to achieve  probabilities that exceed the random limit. This study advances the fundamental thermodynamic understanding of ordering behaviors, which can be generalized to any FCC system.

Image
graph from article
On the Strengths and Weaknesses of Data for Open-set Embodied Assistance
Human Interactive Driving | March 5, 2026

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models generalize to new users and new tasks. Diverse interactive data generation offers a promising avenue for providing data-efficient generalization capabilities for interactive embodied foundation models. In this paper, we investigate the generalization capabilities of a multimodal foundation model fine-tuned on diverse interactive assistance data in a synthetic domain. We explore generalization along two axes: a) assistance with unseen categories of user behavior and b) providing guidance in new configurations not encountered during training. We study a broad capability called Open-Set Corrective Assistance, in which the model needs to inspect lengthy user behavior and provide assistance through either corrective actions or language-based feedback. This task remains unsolved in prior work, which typically assumes closed corrective categories or relies on external planners, making it a challenging testbed for evaluating the limits of assistive data. To support this task, we generate synthetic assistive datasets in Overcooked and fine-tune a LLaMA-based model to evaluate generalization to novel tasks and user behaviors. Our approach provides key insights into the nature of assistive datasets required to enable open-set assistive intelligence. In particular, we show that performant models benefit from datasets that cover different aspects of assistance, including multimodal grounding, defect inference, and exposure to diverse scenarios.

Image
dataset example
Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
Human Interactive Driving | February 23, 2026

Generative models, like flows and diffusions, have recently emerged as popular and efficacious policy parameterizations in robotics. There has been much speculation as to the factors underlying their successes, ranging from capturing multi-modal action distribution to expressing more complex behaviors. In this work, we perform a comprehensive evaluation of popular generative control policies (GCPs) on common behavior cloning (BC) benchmarks. We find that GCPs do not owe their success to their ability to capture multi-modality or to express more complex observation-to-action mappings. Instead, we find that their advantage stems from iterative computation, as long as intermediate steps are supervised during training and this supervision is paired with a suitable level of stochasticity. As a validation of our findings, we show that a minimum iterative policy (MIP), a lightweight two-step regression-based policy, essentially matches the performance of flow GCPs, and often outperforms distilled shortcut models. Our results suggest that the distribution-fitting component of GCPs is less salient than commonly believed, and point toward new design spaces focusing solely on control performance.

Image
generative control policies graphs
Dynamic Association of Semantics and Parameter Estimates by Filtering
Human Interactive Driving | January 14, 2026

We propose a probabilistic semantic filtering framework in which parameters of a dynamical system are inferred and associated with a closed set of semantic classes in a map. We extend existing methods to a multi-parameter setting using a posterior that tightly couples semantics with the parameter likelihoods, and propose a filter to compute this posterior sequentially, subject to dynamics in the map's state. Using Bayesian moment matching, we show that the computational complexity of measurement updates scales linearly in the dimension of the parameter space. Finally, we demonstrate limitations of applying existing methods to a problem from the driving domain, and show that the proposed framework better captures time-varying parameter-to-semantic associations.

Image
semantic map image