Yongyuan Cheryl Liang

I am a PhD student at UMD CS and was fortunate to work with Jianwei Yang and Huazhe Xu. I received my B.S. degree in Mathematics from Sun Yat-sen University.

I have interned and conducted research at NVIDIA, Adobe, and Microsoft Research.

If you’re interested in my research, potential collaborations, or simply want to catch up, feel free to drop me an email.

I'm currently on the full-time job market.

profile photo

Email  /  Google Scholar  /  Github /  Twitter

Research Highlight

I study how to build large agentic models that can generalize across tasks and environments while interacting with the virtual and physical world over time. My research focuses on three capabilities needed to move beyond one-shot perception-to-action mappings toward adaptive, closed-loop agents:

Memory Maintain and update task-relevant information across interactions, viewpoints, and open-ended environments.
Interaction memory: TraceVLA; Spatial memory: MomaGraph, SAW-Bench, Spatial Memory in Frontier Models
Reasoning Learn a structured understanding of the world and reason prospectively about the consequences of actions.
Cross-model reasoning: ROVER; With 3D modality: Lemon; As training objectives: Magma
Planning Compose decisions over long horizons and adapt them as new observations and execution feedback emerge.
Action planning: MomaGraph; Task planning: Plan with Vibe Agents

Together, these capabilities are deeply intertwined within a shared generalist agentic model, where memory, reasoning, and planning jointly shape decisions that are grounded through flexible action interfaces, while interaction feedback continuously updates all three.

Research overview: generalist agentic models interacting with the physical world

Aug' 26  

We drop DanceOPD and DiffusionOPSD.

May' 26  

SAW-Bench won the Best Paper Award Runner-Up at CVPR 2026 WMAS.

Apr' 26  

SAW-Bench to appear in ICML 2026 as Spotlight (2%).

Apr' 26  

Drop a new blog post about Vibe Agents Step Into the Real World.

Mar' 26  

We release a new blog post about Spatial Memory in Frontier Models.

Feb' 26  

Three papers to appear in CVPR 2026 (2 main track and 1 findings).

Jan' 26  

MomaGraph selected as Oral presentation (1%) in ICLR 2026.

Jan' 26  

One papers to appear in ICRA 2026.

Jan' 26  

Two papers to appear in ICLR 2026 (ROVER and MomaGraph).

Sept' 25  

One paper to appear in NeurIPS 2025.

Feb' 25  

Magma to appear in CVPR 2025.

Jan' 25  

Two papers to appear in ICLR 2025.

Jan' 25  

Start to update Awesome-Generalist-Agents.

Sept' 24  

Make-An-Agent to appear in NeurIPS 2024.

June' 24  

ACE selected as Oral presentation (1%) in ICML 2024.

May' 24  

Two papers to appear in ICML 2024.

Jan' 24  

Three papers to appear in ICLR 2024, including two spotlights and one poster.


Selected Publications and Preprints
Filter by: show selected / show all by date / Generalist Agentic Model / Reinforcement Learning / Other Topics

* denotes Equal Contributions and Project Lead; † indicates Equal Advising.

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning
Yuanchen Ju*, Yongyuan Liang*, Yen-Jen Wang*, Gireesh Nandiraju, Yuanliang Ju, Seungjae Lee, Qiao Gu, Elvis Hsieh, Furong Huang†, Koushil Sreenath†

ICLR, 2026 Oral, Top 1%
Project Page  /  Paper  /  Code /  Benchmark /  Twitter

Anticipatory Planning for Multimodal Agents
Yongyuan Liang, Shijie Zhou, Yu Gu, Hao Tan, Gang Wu, Franck Dernoncourt, Jihyung Kil, Ryan A. Rossi, Ruiyi Zhang

CVPR Findings, 2026
Paper  /  Twitter

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
Yongyuan Liang*, Wei Chow*, Feng Li, Ziqiao Ma, Xiyao Wang, Jiageng Mao, Jiuhai Chen, Jiatao Gu, Yue Wang†, Furong Huang†

ICLR, 2026
Project Page  /  Paper  /  Code /  Benchmark /  Twitter

Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
Yongyuan Liang, Xiyao Wang, Yuanchen Ju, Jianwei Yang, Furong Huang

arXiv, 2025
Spotlight Talks at CVPR Workshop CVinW, 2025
Project Page  /  Paper  /  Code /  Models & Datasets /  Twitter

Magma: A Foundation Model for Multimodal AI Agents
Magma Team

CVPR, 2025
Project Page  /  Paper  /  Code /  Models & Datasets /  Twitter

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
Ruijie Zheng*, Yongyuan Liang*, Shuaiyi Huang, Jianfeng Gao, Hal Daumé III, Andrey Kolobov, Furong Huang, Jianwei Yang

ICLR, 2025
Oral Talks at ICLR Workshop GenBot, 2025
Project Page  /  Paper  /  Code /  Models /  Twitter

Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion
Yongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang, Furong Huang, Huazhe Xu

NeurIPS, 2024
Oral Talks at NeurIPS Workshop AFM, 2024
Project Page  /  Paper  /  Code /  Models & Dataset /  Twitter

ACE: Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
Tianying Ji*, Yongyuan Liang*, Yan Zeng, Yu Luo, Guowei Xu, Jiawei Guo, Ruijie Zheng, Furong Huang, Fuchun Sun, Huazhe Xu

ICML, 2024 Oral, Top 1%
Project Page  /  Paper  /  Code /  Twitter

Blogs
Mar 2026

What We Talk About When We Talk About Spatial Memory in Frontier Models

Spatial Awareness, World Modeling

Apr 2026

Vibes Meet Gravity: AI Agents Step Into the Real World

Vibe Agents, Agentic Multimodal Models

Coming soon

Avocado: Multi-Objective Alignment of Language Models

Alignment Steering, interpretability


Professional Service

Conference Program Committee: ICML(2022-2025), NeurIPS(2021-2025), ICLR(2021-2026)

Workshop Program Committee: FMDM at NeurIPS 2023, Bi-Align at ICLR 2025, CVinW at CVPR 2025


Misc

If my name is a bit tricky to pronounce for you, I’d love to go by Cheryl [ˈʃerəl].

Classic INTJ-A.

I've been playing the violin🎻 for over 15 years and served as Principal First Violin in the university orchestra. I also play the piano for more than 10 years.

I enjoy reading Japanese and Western literature. Here's some of my reading notes.

Been a fan of Novak Djokovic since 2012.

My Erdős number = 4.

life in photos 🍊



© Yongyuan Liang. All rights reserved for content and custom design.
Base template by Jon Barron.