Hi!

I'm Javier Lopetegui

PhD student at WILLOW, Inria Paris & ENS, advised by Cordelia Schmid.

I work on World-Action Models — video generation models adapted to jointly predict future observations and actions — as a paradigm for robot policy learning.

Javier Lopetegui

About me

I am a first-year PhD student in the WILLOW team at Inria Paris and École Normale Supérieure, supervised by Cordelia Schmid. I work on World-Action Models: large video diffusion models adapted to jointly predict future observations and actions, which makes them a promising paradigm for robot policy learning.

Before starting the PhD I completed the MVA master at ENS Paris-Saclay and an M1 in Artificial Intelligence at Université Paris-Saclay, supported by an Excellence Scholarship from the French Embassy in Cuba. I did my BSc in Computer Science at the University of Havana, where I also spent three years as a teaching assistant.

Along the way I worked on promptable medical image segmentation at Raidium, and on NLP for language varieties and biomedical text at Inria's ALMAnaCH team and at LISN.

News

  • Jul 2026 🇮🇹 Presented SA-WAM as a poster at the International Computer Vision Summer School (ICVSS 2026), Sicily.
  • Nov 2025 🎓 Started my PhD in the WILLOW team at Inria Paris / ENS, advised by Cordelia Schmid.
  • Oct 2025 ✅ Graduated from the MVA master at ENS Paris-Saclay.
  • Apr 2025 🩻 Joined Raidium as a research intern, working on text-promptable medical image segmentation.

Research

SA-WAM method diagram

Spatially Aware World Action Model via Geometric Latent Diffusion

Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid

Preprint, 2026

World Action Models (WAMs) leverage large-scale pretrained video diffusion models to jointly predict future observations and actions, yet the prevailing models operate exclusively on RGB and do not leverage 3D information. SA-WAM repurposes a pretrained video model for joint action, RGB and depth prediction within a single diffusion backbone: a nonlinear encoding maps the unbounded depth signal into the bounded input domain expected by the frozen VAE tokenizer, reusing the encoder without 3D-specific fine-tuning. It achieves state-of-the-art results on the RoboCasa and LIBERO-Plus benchmarks and outperforms strong baselines on a real UR5 arm.

Teaching

  • Introduction to Programming Student teaching assistant — practical sessions for 1st-year Computer Science students, Faculty of Mathematics and Computer Science, University of Havana
    Jan 2020 – Jul 2023