报告摘要:生成式AI在学习高维数据底层分布方面已取得重大进展,这为构建通用AI系统奠定了基础。 在本次报告中,我将分享我们针对虚拟世界内容所开发的大规模生成模型的工作,包括图像、视频和3D内容的实例。 更进一步,我将介绍近期在具身视频基础模型方面的研究,该模型能够充分利用“数据金字塔”, 借助互联网规模的视频数据实现双臂操作的强大泛化能力,展现了构建具身基础模型的广阔前景。
讲者简介:朱军,清华大学计算机系Bosch AI教授、清华大学人工智能研究院副院长、 IEEE/AAAI Fellow,曾任卡内基梅隆大学兼职教授。主要从事机器学习研究,担任国际著名期刊IEEE TPAMI的副主编, 担任ICML、NeurIPS、ICLR等资深领域主席和最佳论文评审委员等。获中国青年科技奖、中国科协求是杰出青年奖、 陈嘉庚青年科技奖、科学探索奖、ICLR国际会议杰出论文奖等。研制首个全面对标Sora的Vidu视频大模型, 以及开源的概率编程库“珠算”、深度强化学习库“天授”等,通过技术转化孵化瑞莱智慧和生数科技。
报告摘要:Recovering the 3D scene geometry is an important step towards spatial intelligence. As we continuously explore our world, real-time reconstructing such 3D world becomes the basic requirement for perceiving the 3D world. However, current 3D reconstruction method is usually very slow. To solve this issue, we propose LongSplat, an online real-time 3D Gaussian reconstruction framework designed for long-sequence image input. Although we can reconstruct the holistic scene, it is still not enough as current 3D reconstruction methods could not infer the unseen parts of the scene, separate each object in the scene, and understand the spatial relation between different objects. To solve these issues, we propose the CUPID, a generative 3D reconstruction pipeline, which that accurately infers the camera pose, 3D shape, and texture from a single image. At last, we introduce All-Angles Bench, a comprehensive benchmark to evaluate MLLMs’ multi-view understanding. Our evaluation of 27 representative models across over 2,100 annotated multi-view question-answer pairs in the six tasks, we reveal significant limitations in geometric consistency and cross-view correspondence, particularly in cross-view identification and camera pose estimation. These findings highlight the need for domain-specific training to achieve human-level spatial intelligence with MLLMs.
讲者简介:Shenghua Gao is an Associate Professor, and the assistant director of School of Computing and Data Science, the University of Hong Kong. He also serves as the associate director for the Theory and System Center, Shenzhen Loop Area Institute. His research interests include 3D reconstruction, image and video understanding and generation, 3D generation, AI4Science, etc. He has served as an area chair for multiple conferences (CVPR, NeurIPS, ICCV, ACM MM, ECCV, etc.). He also served as an associate editor for IEEE TPAMI and TMM.
报告摘要:近年来,随着大模型的快速发展,三维重建与生成技术也取得了显著进展,而且两者技术的结合既可以提升重建的鲁棒性和完整度也能提升生成的质量和时空一致性,已成为一个重要发展趋势。本次报告主要介绍我们课题组近年来在三维场景的高效重建与生成方面的研究工作,并展示相关的应用。例如,根据输入的一张或多张稀疏视图,可以自动生成多段时空一致的长视频序列,结合三维重建与3D高斯溅射技术可以进一步生成高质量的可供用户自由漫游的三维场景。
讲者简介:章国锋,浙江大学求是特聘教授,博士生导师,国家杰出青年科学基金获得者。主要从事三维视觉、增强现实与空间智能方面的研究,尤其在SLAM、三维重建和生成方面取得了一系列重要成果,开源了一系列相关系统和算法的源代码,是OpenXRLab扩展现实开源平台的主要发起人。曾获2011年全国优秀博士学位论文奖、2020年浙江省技术发明奖一等奖(排名第4)、2021年浙江省自然科学奖一等奖(排名第2)以及国际顶级会议ISMAR 2020唯一最佳论文奖。担任国际顶级期刊IJCV编委,以及《Virtual Reality & Intelligent Hardware》、《计算机辅助设计与图形学学报》和《中国图象图形学报》等期刊编委,中国图象图形学学会虚拟现实专委会副主任、增强现实核心技术产业联盟副理事长、浙江省人工智能学会增强现实分会副会长。
报告摘要:三维场景是空间智能的基础底座,机器人通过深度相机及/或RGB相机来感知和理解三维场景。实现机器人自主式的三维场景重建需要解决两个基本问题:路径规划和视点规划。我们从场景中三维物体的局部信息、相互关系等信息来实现机器人的自主式漫游、扫描、重建与物体理解的任务。另一方面,生成式AI的快速发展能够使得我们通过简单的文本与图像等输入方式快速生成三维场景。在这个报告中,我们将介绍基于多模态数据来实现三维场景的重建与生成的系列探索和思考
讲者简介:刘利刚,中国科学技术大学教授,国家自然基金委“杰出青年”项目获得者。从事计算机图形学及CAD/CAE方向研究。曾获中国计算机图形学杰出奖,Siggraph Asia首届时间检验奖 (Test-of-Time Award)等奖项。任中国工业与应用数学学会几何设计与计算专业委员会 (CSIAM GDC) 主任、国际几何建模与处理(GMP)协会指导委员会委员、国际实体建模协会(SMA)执行委员会委员、亚洲图形学协会(Asiagraphics)副主席。