I am a PhD student at NYU Courant, advised by Yann LeCun. My research focuses on world
models, planning, and representation learning, with the goal of building agents that can learn from
observation and interact with the world in an open-ended way.
I am currently especially interested in hierarchical world models and planning across multiple
temporal scales, where structure in the model can enable better long-horizon reasoning, more
efficient control, and stronger generalization from offline or reward-free data. More broadly, I am
interested in self-supervised learning for decision making, computer vision, and embodied
intelligence.
Learns a hierarchy of JEPA world models end-to-end, with each level predicting farther ahead in
its own latent space. Planning from coarse to fine improves long-horizon control while reducing
planning compute, raising Visual AntMaze success from 18% to 73%.
Introduces hierarchical planning across latent world models at multiple temporal scales, improving
long-horizon zero-shot control while reducing planning-time compute on robotics and simulated
control tasks.
This work introduces zero-shot planning using the Joint Embedding Prediction Architecture (JEPA)
trained from offline trajectories, demonstrating its strengths in generalization, trajectory
stitching, and data efficiency compared to traditional offline RL.
Introduces a JEPA-style latent dynamics model with a physics-inspired prober for real-time
quadrotor control, enabling accurate long-horizon prediction and robust zero-shot sim-to-real
transfer from automatically generated simulation data.
Presents an efficient probing benchmark to evaluate the fitness of unsupervised visual
representations for reinforcement learning (RL). Applied it to systematically improve pre-existing
SSL recipes for RL.
Showcases an industrial-scale end-to-end Automatic Speech Recognition model trained on 570k hours
of speech audio data using Noisy Student. It achieves competitive word error rates against larger
and more computationally expensive models.
Adapts the computer vision data augmentation technique MixUp to the natural language domain,
reducing calibration error of transformers for sentence classification by up to 50%.