GAMES Webinar 2026 – 417期(构建物理智能基座:从第一人称数据到可拓展预训练模型)|洪方舟(Ropedia), 徐英豪(香港科技大学)

GAMES Webinar 2026 – 417期(构建物理智能基座:从第一人称数据到可拓展预训练模型)

报告嘉宾:洪方舟 Ropedia

报告时间:2026年10月08日 晚上20:00-20:30(北京时间)

报告题目:

物理智能的第一人称数据基座

报告摘要:

第一人称人类数据记录了人类与世界发生交互的准确物理过程,是机器人学习移动与操作任务的重要数据来源,并且因其容易规模化的特点,而获得了社区越来越多的关注。在Ropedia,我们致力于构建第一人称人类数据的基础设施,构建了包括硬件、平台、算法管线、与模型在内的一整套数据闭环。我们最近发布了我们的第二代第一人称数据采集头环Ropedia HOMIE Gen2,具有更高的硬件参数规格与可扩展性,能够取得更高的数据标注精度。此外,我们也通过一系列模型研究,验证了我们第一人称数据的有效性,尤其是第一人称世界模型Xperience-0,通过我们的数据训练,展现出了在移动操作与跨本体等方面的新能力。

讲者简介:

洪方舟博士是Ropedia的联合创始人兼CTO,正带领团队为物理智能(Physical AI)构建数据基础设施,让机器真正理解并适应物理世界。他主导打造的旗舰多模态数据集Xperience-10M,在Hugging Face上下载量突破270万次、登顶平台周热门榜前三。他于新加坡南洋理工大学获得博士学位(MMLab@NTU / S-Lab,师从刘子纬教授),本科毕业于清华大学软件学院。他在三维生成、三维重建与具身感知方向发表了一系列工作,在CVPR、ICLR、SIGGRAPH等顶级会议发表并入选Oral/Spotlight。他曾于Meta Reality Labs实习参与第一人称多模态模型研究。他曾获得Google博士生奖学金、China3DV学术新锐奖。

讲者主页:https://hongfz16.github.io/

报告嘉宾:徐英豪 香港科技大学

报告时间:2026年10月08日 晚上20:30-21:00(北京时间)

报告题目:

Embodied foundational model with scalable video-action pretraining

报告摘要:

Most robot foundation models inherit their knowledge from vision-language pretraining, yet robot control is fundamentally about how the physical world evolves under actions. This talk argues that video-action pretraining offers a complementary and scalable foundation for robot learning: by learning to predict scene dynamics, a model acquires manipulation priors from large-scale robot and human videos and translates them into executable actions. We systematically study how to design effective proxy tasks for scalable pretraining along three dimensions: the latent space, the autoregressive pretraining objective, and in-context learning. We further show that, with asynchronous inference and efficient deployment, the pretrained model enables real-time closed-loop control on real robots. Together, these results suggest that video-action pretraining provides a scalable path toward general-purpose robot foundation models.

讲者简介:
Yinghao Xu is an Assistant Professor in the Department of Computer Science and Engineering at the Hong Kong University of Science and Technology (HKUST). Previously, he was a Staff Research Scientist at RobbyAnt, working on world models and embodied AI. Before that, he was a postdoctoral researcher at Stanford University. His research lies at the intersection of 3D computer vision, generative AI, and embodied AI, with a recent focus on building world models that unify 3D reconstruction, world simulation, and embodied action. His work has been selected as a Best Paper Award candidate at ECCV 2026 and CVPR 2021. He was the recipient of the Yunfan Award at WAIC 2024 and was nominated for the Snap Fellowship in 2022.

讲者主页:https://justimyhxu.github.io/


主持人简介:

马月昕,上海科技大学长聘副教授、博士生导师,博士毕业于香港大学。主要研究方向为三维视觉、具身智能、自动驾驶。共发表相关领域顶会或顶刊论文100余篇,其中一作与通讯论文50余篇,包括TPAMI、CVPR、ICCV、SIGGRAPH等,谷歌学术引用近1万次。参与指导的论文获MICCAI 2024唯一最佳论文奖,ACM MM 2024最佳论文候选。担任TVCG、Visual Intelligence、RA-L编委,曾获上海市海外高层次人才、China 3DV 2025年度优秀青年学者、入选全球前2%顶尖科学家榜单等。

You may also like...