Events
CMAI seminar: World Models - Spatial Physical Generative AI
Centre for Multimodal AICentre for Multimodal AI seminar
Speaker: Prof Tae-Kyun Kim (KAIST, currently at Samsung AI Cambridge)
Title: World Models: Spatial Physical Generative AI
When: Thursday, 15th October, 2pm - 3pm
Where: Room 701, Graduate Centre
Abstract:
Recent advances in generative AI have produced impressive language, image, and video models, yet these systems often lack a grounded understanding of 3D space, motion, interaction, and physical causality. This talk explores a path beyond large language models toward spatial and physical artificial general intelligence through physics-grounded world models. It presents our recent work on physically realistic digital humans, deformable hand–object interactions, and simulation-guided video generation, including MPMAvatar, PhysHanDI, and 3DPhysVideo. By integrating differentiable physical simulation, 3D scene reconstruction, and generative video models, these approaches improve physical consistency, controllability, and generalization across objects, materials, and interactions. The talk also introduces the Goldberg Machine Test as a grand challenge for evaluating compositional and long-horizon physical intelligence. Together, these developments suggest that combining learned representations with physical reasoning is essential for building world models capable of supporting embodied AI and, ultimately, spatial physical AGI.
Bio:
Tae-Kyun (T-K) Kim is Professor and the director of Computer Vision and Learning Lab at School of Computing, KAIST since 2020, and has been an adjunct reader of Imperial College London (ICL), UK for 2020-2024. He led Computer Vision and Learning Lab at Imperial College during 2010-2020. He obtained his PhD from Univ. of Cambridge in 2008 and Junior Research Fellowship (governing body) of Sidney Sussex College, Univ. of Cambridge during 2007-2010. His BSc and MSc are from KAIST in 1998 and 2000, he worked at Samsung AIT for 2000-2004 (military duty). His research interests primarily lie in machine (deep) learning for 3D computer vision, generative AI and AI with Physics, including: articulated 3D hand/body reconstruction, face analysis and recognition, 6D object pose estimation, activity recognition, object detection/tracking, active robot vision, which lead to novel active and interactive visual sensing. He has co-authored over 120 academic papers in top-tier conferences and journals in the field, and has co-organised series of HANDS workshops and 6D Object Pose workshops (in conjunction with CVPR/ICCV/ECCV) since 2015 to 2020. He was the general chair of BMVC17 in London, the program co-chair of BMVC23, and is Associate Editor of IEEE Trans on PAMI, Pattern Recognition Journal, Image and Vision Computing Journal. He regularly serves as an (Senior) Area Chair for top-tier vision/ML conferences. He received KUKA best service robotics paper award at ICRA 2014, and 2016 best paper award by the ASCE Journal of Computing, and the best paper finalist at CVPR 2020, and his co-authored algorithm for face image representation is an international standard of MPEG-7 ISO/IEC.
| Contact: | Changjae Oh |
| Email: | c.oh@qmul.ac.uk |
Updated by: Emmanouil Benetos

