LI Kai李锴
李锴LI KaiMultimodal LLMs · Video World Models · 3D Spatial Understanding多模态大模型 · 视频世界模型 · 三维空间理解
My research centres on large multimodal models and how they acquire spatial and temporal understanding. I am currently working on memory in video world models, alongside reinforcement learning for sparse-view 3D scene reconstruction. Earlier, during my PhD, I proposed the Offset Token, which brought off-nadir photogrammetry into the perceptual scope of vision foundation models.我的研究以多模态大模型为核心,关注模型如何获得空间与时间层面的理解能力。目前正在研究视频世界模型(Video World Model)中的记忆机制,并同时探索用强化学习实现稀疏视角三维场景重建。博士期间,我提出了偏移量词元(Offset Token)的概念,将偏移摄影带入视觉大模型的理解视角。
Browse my research directions浏览我的研究方向
Current Focus当前重心
Teaching a video world model to remember让视频世界模型学会记忆
Video world models learn to predict how a scene evolves, but they largely lack a persistent notion of what they have already seen. My current work asks how such a model should remember: how to retain scene state over long horizons, stay consistent when the camera revisits a place, and keep memory affordable as the rollout grows. This is the temporal counterpart to the spatial understanding I built during my PhD.视频世界模型能够预测场景如何演化,却普遍缺少对「已经看到过什么」的持久表征。 我目前的研究关注这类模型应当如何记忆:如何在长时序上保持场景状态、 如何在镜头重新回到同一位置时维持一致、以及如何在推演变长时控制记忆开销。 这与我博士期间构建的空间理解能力,构成时间维度上的呼应。
Research Line研究主线
Multimodal LLMs meet 3D scene understanding多模态大模型与 3D 场景理解
After accumulating substantial experience in 3D spatial understanding, this line extends my work towards today's frontier: letting large multimodal models reason about 3D structure directly, without camera calibration or dense views — and, building on that foundation, using reinforcement learning to keep sparse-view indoor reconstruction consistent across viewpoints.在积累了丰富的三维空间理解经验之后,这一部分把工作拓展到当下最前沿的方向: 让多模态大模型在没有相机参数、没有密集视图的条件下直接推理三维结构;并在此基础上, 尝试用强化学习让稀疏视角下的室内场景重建保持跨视角一致。
Research Line研究主线
Offset Token: from a concept to a system偏移量词元:从一个概念到一套体系
My PhD research targets building extraction under off-nadir satellite photography. I first proposed the concept of the Offset Token, which brings off-nadir photography into the perceptual scope of large vision models. Around this concept I built a complete loop — vectorised building extraction, positional repair between historical vector maps and updated imagery, efficient optimisation of the offset token, and full decoupling of offset learning.博士期间的研究方向为遥感卫星偏移摄影视角下的建筑物提取。研究主体上, 首先提出了偏移量词元(Offset Token)的概念,将偏移摄影带入大模型的理解视角; 并围绕所提出的偏移量词元概念展开更加细腻的研究,实现了偏移摄影下矢量化的建筑物提取、 历史矢量地图与更新影像之间的位置修复、偏移量词元的高效优化计算、偏移量学习解耦等内容, 完成了「概念提出 → 具象化应用 → 任务解耦与独立 → 底层理论降维」的闭环。
Concept概念提出
Offset Token enters the foundation-model view偏移量词元进入大模型视角
Application具象应用
Vector footprints & historical-label alignment矢量化提取与历史标签对齐
Decoupling任务解耦
Offset learning becomes an independent task偏移量学习独立成任务
Reduction理论降维
7-DoF formulation, far cheaper inference7 维充分描述,大幅降低算力
News近期动态
Recent updates近期动态
- 2026.08 🚀 Started working on memory mechanisms for video world models at the Tencent Omni Research Team.🚀 开始在腾讯 Omni 研究团队从事视频世界模型记忆机制的研究。
- 2026.08 📄 Remote-Sensing City Layout Extraction with MLLM is on arXiv — first-authored by Zigan Zhou, the master's student I mentor.📄 Remote-Sensing City Layout Extraction with MLLM 已挂 arXiv —— 由我指导的硕士生 Zigan Zhou 担任第一作者。
- 2025.10 🎉 One paper accepted by IEEE JSTARS.🎉 一篇论文被 IEEE JSTARS 接收。
- 2025.10 👍 Honoured to join the Tencent Omni Research Team again.👍 很荣幸再次加入腾讯 Omni 研究团队。
- 2025.09 🎉 One paper accepted by IEEE TGRS.🎉 一篇论文被 IEEE TGRS 接收。
- 2025.08 🎉 One paper accepted by IEEE TGRS.🎉 一篇论文被 IEEE TGRS 接收。
- 2025.08 👏 Nominated for the IEEE Excellence in Technical Communication Student Prize Award by IEEE GRSS in 2025 — the only Chinese nominee of the year.👏 获 IEEE GRSS 提名 2025 年度 IEEE Excellence in Technical Communication Student Prize Award(年度唯一华人)。
- 2025.07 👏 OBM was selected for the 3MT Final Round (Top 10) at IGARSS 2025.👏 OBM 入选 IGARSS 2025 三分钟论文竞赛决赛(Top 10)。
- 2025.07 🎉 Our PolyFootNet is accepted by IEEE TGRS.🎉 我们的 PolyFootNet 被 IEEE TGRS 接收。
- 2025.04 🎉 One paper accepted by WWW Companion.🎉 一篇论文被 WWW Companion 接收。
- 2025.04 👏 Granted an IGARSS 2025 Travel Grant. See you in Brisbane!👏 获得 IGARSS 2025 Travel Grant 资助,布里斯班见!
- 2024.10 🎉 Our OBM is accepted by IEEE TGRS.🎉 我们的 OBM 被 IEEE TGRS 接收。
- 2024.09 🎉 One paper accepted by ISPRS Journal of Photogrammetry and Remote Sensing.🎉 一篇论文被 ISPRS Journal of Photogrammetry and Remote Sensing 接收。
Background教育与经历
Where I come from教育、经历与教学
Education教育背景
-
2024 – Present
Ph.D. Candidate in Data Science博士研究生,数据科学
City University of Hong Kong (CityU) · School of Data Science香港城市大学(CityU)· 数据科学学院
-
2021 – Present
Ph.D. Candidate (direct entry), Signal and Information Processing博士研究生(本科直博),信号与信息处理
University of Chinese Academy of Sciences (UCAS) · Aerospace Information Research Institute中国科学院大学(UCAS)· 空天信息创新研究院
-
Until 2021
B.Eng. in Space Information and Digital Technology工学学士,空间信息与数字技术
University of Electronic Science and Technology of China (UESTC)电子科技大学(UESTC)
Experience研究与实习经历
-
2025.10 – Present
Research Intern, Omni Research Team研究实习生,Omni 研究团队
Tencent腾讯
-
2021 – Present
Research Assistant, Off-Nadir Remote Sensing Group研究助理,偏移摄影遥感方向
Aerospace Information Research Institute, Chinese Academy of Sciences中国科学院空天信息创新研究院
-
2021.05 – 2021.09
Computer Vision Algorithm Intern计算机视觉算法实习生
Tencent腾讯
Mentorship研究生指导
Jiajun ZhangJiajun Zhang
M.Sc., City University of Hong Kong · co-mentored with Prof. Xiangyu Zhao香港城市大学硕士研究生 · 与赵翔宇教授联合指导
Zigan ZhouZigan Zhou
M.Sc., City University of Hong Kong · co-mentored with Prof. Xiangyu Zhao香港城市大学硕士研究生 · 与赵翔宇教授联合指导
Weizhi ChenWeizhi Chen
University of Chinese Academy of Sciences (AIRCAS) · co-mentored with Prof. Jingbo Chen中国科学院大学(空天信息创新研究院)· 与陈静波教授联合指导
Teaching教学经历
Teaching Assistant, Optimization Methods助教,最优化方法
Prof. Zhenjiang Zhao · University of Chinese Academy of Sciences赵振江 教授 · 中国科学院大学
Teaching Assistant, Academic English助教,学术英语
Prof. James · University of Chinese Academy of SciencesJames 教授 · 中国科学院大学
Get in touch联系我
Let's talk about off-nadir vision and 3D scenes欢迎交流偏移摄影与三维场景理解
Open to faculty, postdoctoral, research scientist and algorithm engineer positions from 2026.2026 年起开放教职、博士后、研究科学家与算法工程师岗位机会。