GAMES Webinar 2026 – 417期(构建物理智能基座:从第一人称数据到可拓展预训练模型)|洪方舟(Ropedia), 徐英豪(香港科技大学)
GAMES Webinar 2026 – 417期(构建物理智能基座:从第一人称数据到可拓展预训练模型)
报告嘉宾:洪方舟 Ropedia
报告时间:2026年10月08日 晚上20:00-20:30(北京时间)
报告题目:
报告摘要:
第一人称人类数据记录了人类与世界发生交互的准确物理过程,是机器人学习移动与操作任务的重要数据来源,并且因其容易规模化的特点,而获得了社区越来越多的关注。在Ropedia,我们致力于构建第一人称人类数据的基础设施,构建了包括硬件、平台、算法管线、与模型在内的一整套数据闭环。我们最近发布了我们的第二代第一人称数据采集头环Ropedia HOMIE Gen2,具有更高的硬件参数规格与可扩展性,能够取得更高的数据标注精度。此外,我们也通过一系列模型研究,验证了我们第一人称数据的有效性,尤其是第一人称世界模型Xperience-0,通过我们的数据训练,展现出了在移动操作与跨本体等方面的新能力。
洪方舟博士是Ropedia的联合创始人兼CTO,正带领团队为物理智能(Physical AI)构建数据基础设施,让机器真正理解并适应物理世界。他主导打造的旗舰多模态数据集Xperience-10M,在Hugging Face上下载量突破270万次、登顶平台周热门榜前三。他于新加坡南洋理工大学获得博士学位(MMLab@NTU / S-Lab,师从刘子纬教授),本科毕业于清华大学软件学院。他在三维生成、三维重建与具身感知方向发表了一系列工作,在CVPR、ICLR、SIGGRAPH等顶级会议发表并入选Oral/Spotlight。他曾于Meta Reality Labs实习参与第一人称多模态模型研究。他曾获得Google博士生奖学金、China3DV学术新锐奖。
报告嘉宾:徐英豪 香港科技大学
报告时间:2026年10月08日 晚上20:30-21:00(北京时间)
报告题目:
报告摘要:
Most robot foundation models inherit their knowledge from vision-language pretraining, yet robot control is fundamentally about how the physical world evolves under actions. This talk argues that video-action pretraining offers a complementary and scalable foundation for robot learning: by learning to predict scene dynamics, a model acquires manipulation priors from large-scale robot and human videos and translates them into executable actions. We systematically study how to design effective proxy tasks for scalable pretraining along three dimensions: the latent space, the autoregressive pretraining objective, and in-context learning. We further show that, with asynchronous inference and efficient deployment, the pretrained model enables real-time closed-loop control on real robots. Together, these results suggest that video-action pretraining provides a scalable path toward general-purpose robot foundation models.
讲者主页:https://justimyhxu.github.io/
主持人简介: