Yuhao Zhang (张钰浩)

I am a Master student at Shanghai Jiao Tong University, Department of Computer Science, working in ScaleLab under the supervision of Prof. Yao Mu.

My research focuses on Embodied AI, with a particular interest in Embodied World Models and Vision-Language-Action (VLA) models.

Email  /  Scholar  /  GitHub

profile photo

News

  • [2026.07] Our paper R3DP was accepted to ECCV 2026 as Oral.
  • [2026.06] Selected as Outstanding Camper at Shanghai Innovation Institute (上海创智学院优秀营员).
  • [2026.06] Two papers (R3DP and M2Tok) accepted to ECCV 2026.
  • [2026.02] Our paper MM-ACT was accepted to CVPR 2026.
  • [2026.01] Our paper SimpleVLA-RL was accepted to ICLR 2026.

Research

I'm interested in embodied intelligence, world models, and VLA. Representative papers are highlighted. * denotes equal contribution.

R3DP R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Yuhao Zhang*, Wanxi Dong, Yue Shi, Yi Liang, Jingnan Gao, Qiaochu Yang, Yaxing Lyu, Zhixuan Liang, Yibin Liu, Congsheng Xu, etc.
European Conference on Computer Vision (ECCV), 2026   (Oral)
project page / arXiv / code

A real-time 3D-aware manipulation policy that couples VGGT-based priors with a lightweight temporal predictor and multi-view fuser.

Implicit Drifting Policy Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry
Zemin Yang, Yaoyu He, Yiming Zhong, Yuhao Zhang, Xinge Zhu, Yao Mu, Qingqiu Huang, Yuexin Ma
arXiv preprint, 2026
arXiv

A one-step imitation learning framework using conditional expert geometry to enforce action-manifold constraints without explicit drifting-field estimation.

M2Tok M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models
Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu
European Conference on Computer Vision (ECCV), 2026

A multi-head multi-codebook discrete action tokenization scheme for scaling VLA models on continuous control.

MM-ACT MM-ACT: Learn from Multimodal Parallel Generation to Act
Haotian Liang, Xinyi Chen, Bin Wang, Mingkang Chen, Yitian Liu, Yuhao Zhang, Zanxin Chen, Tianshuo Yang, Yilun Chen, Jiangmiao Pang, etc.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
arXiv / code

A unified VLA model unifying text, image, and action generation in a shared token space; 96.3% on LIBERO.

SimpleVLA-RL SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Haozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang*, Zhaohui Yang, Kaiyan Zhang, Xuekai Zhu, Yuchen Zhang, Tianxing Chen, Ganqu Cui, etc.
International Conference on Learning Representations (ICLR), 2026
arXiv / code

A simple, scalable RL recipe for training vision-language-action models.

RoboTwin Challenge Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop
Tianxing Chen, Kaixuan Wang, Zhaohui Yang, Yuhao Zhang*, Zanxin Chen, Baijun Chen, Wanxi Dong, Ziyuan Liu, Dong Chen, Tianshuo Yang, etc.
CVPR 2025 MEIS Workshop (Challenge Report), 2025
arXiv / challenge page

A large-scale challenge on generalizable bimanual manipulation across 17 tasks in simulation and the real world.

Education

  • 2024.09 – present: M.S. in Computer Science, Shanghai Jiao Tong University
  • 2020.09 – 2024.06: B.E. in Software Engineering, Shandong University

Awards

  • 2026: Outstanding Camper, Shanghai Innovation Institute (上海创智学院优秀营员)
  • 2026: ECCV 2026 Oral Presentation — R3DP

Template borrowed from Jon Barron.