Skip to content
PhD Candidate · UCAS × CityU Hong Kong中国科学院大学 × 香港城市大学 联合培养博士生Seeking 2026 Positions2026 求职中

LI Kai李锴

李锴LI Kai

Multimodal LLMs · Video World Models · 3D Spatial Understanding多模态大模型 · 视频世界模型 · 三维空间理解

My research centres on large multimodal models and how they acquire spatial and temporal understanding. I am currently working on memory in video world models, alongside reinforcement learning for sparse-view 3D scene reconstruction. Earlier, during my PhD, I proposed the Offset Token, which brought off-nadir photogrammetry into the perceptual scope of vision foundation models.我的研究以多模态大模型为核心,关注模型如何获得空间与时间层面的理解能力。目前正在研究视频世界模型(Video World Model)中的记忆机制,并同时探索用强化学习实现稀疏视角三维场景重建。博士期间,我提出了偏移量词元(Offset Token)的概念,将偏移摄影带入视觉大模型的理解视角。

LI Kai
19
Publications发表论文
152
Citations论文引用
7
h-indexh 指数

Current Focus当前重心

Teaching a video world model to remember让视频世界模型学会记忆

Video world models learn to predict how a scene evolves, but they largely lack a persistent notion of what they have already seen. My current work asks how such a model should remember: how to retain scene state over long horizons, stay consistent when the camera revisits a place, and keep memory affordable as the rollout grows. This is the temporal counterpart to the spatial understanding I built during my PhD.视频世界模型能够预测场景如何演化,却普遍缺少对「已经看到过什么」的持久表征。 我目前的研究关注这类模型应当如何记忆:如何在长时序上保持场景状态、 如何在镜头重新回到同一位置时维持一致、以及如何在推演变长时控制记忆开销。 这与我博士期间构建的空间理解能力,构成时间维度上的呼应。

Research Line研究主线

Multimodal LLMs meet 3D scene understanding多模态大模型与 3D 场景理解

After accumulating substantial experience in 3D spatial understanding, this line extends my work towards today's frontier: letting large multimodal models reason about 3D structure directly, without camera calibration or dense views — and, building on that foundation, using reinforcement learning to keep sparse-view indoor reconstruction consistent across viewpoints.在积累了丰富的三维空间理解经验之后,这一部分把工作拓展到当下最前沿的方向: 让多模态大模型在没有相机参数、没有密集视图的条件下直接推理三维结构;并在此基础上, 尝试用强化学习让稀疏视角下的室内场景重建保持跨视角一致。

Research Line研究主线

Offset Token: from a concept to a system偏移量词元:从一个概念到一套体系

My PhD research targets building extraction under off-nadir satellite photography. I first proposed the concept of the Offset Token, which brings off-nadir photography into the perceptual scope of large vision models. Around this concept I built a complete loop — vectorised building extraction, positional repair between historical vector maps and updated imagery, efficient optimisation of the offset token, and full decoupling of offset learning.博士期间的研究方向为遥感卫星偏移摄影视角下的建筑物提取。研究主体上, 首先提出了偏移量词元(Offset Token)的概念,将偏移摄影带入大模型的理解视角; 并围绕所提出的偏移量词元概念展开更加细腻的研究,实现了偏移摄影下矢量化的建筑物提取、 历史矢量地图与更新影像之间的位置修复、偏移量词元的高效优化计算、偏移量学习解耦等内容, 完成了「概念提出 → 具象化应用 → 任务解耦与独立 → 底层理论降维」的闭环。

Concept概念提出

Offset Token enters the foundation-model view偏移量词元进入大模型视角

Application具象应用

Vector footprints & historical-label alignment矢量化提取与历史标签对齐

Decoupling任务解耦

Offset learning becomes an independent task偏移量学习独立成任务

Reduction理论降维

7-DoF formulation, far cheaper inference7 维充分描述,大幅降低算力

News近期动态

Recent updates近期动态

  • 2026.08 🚀 Started working on memory mechanisms for video world models at the Tencent Omni Research Team.🚀 开始在腾讯 Omni 研究团队从事视频世界模型记忆机制的研究。
  • 2026.08 📄 Remote-Sensing City Layout Extraction with MLLM is on arXiv — first-authored by Zigan Zhou, the master's student I mentor.📄 Remote-Sensing City Layout Extraction with MLLM 已挂 arXiv —— 由我指导的硕士生 Zigan Zhou 担任第一作者。
  • 2025.10 🎉 One paper accepted by IEEE JSTARS.🎉 一篇论文被 IEEE JSTARS 接收。
  • 2025.10 👍 Honoured to join the Tencent Omni Research Team again.👍 很荣幸再次加入腾讯 Omni 研究团队
  • 2025.09 🎉 One paper accepted by IEEE TGRS.🎉 一篇论文被 IEEE TGRS 接收。
  • 2025.08 🎉 One paper accepted by IEEE TGRS.🎉 一篇论文被 IEEE TGRS 接收。

Background教育与经历

Where I come from教育、经历与教学

Education教育背景

  • 2024 – Present

    Ph.D. Candidate in Data Science博士研究生,数据科学

    City University of Hong Kong (CityU) · School of Data Science香港城市大学(CityU)· 数据科学学院

  • 2021 – Present

    Ph.D. Candidate (direct entry), Signal and Information Processing博士研究生(本科直博),信号与信息处理

    University of Chinese Academy of Sciences (UCAS) · Aerospace Information Research Institute中国科学院大学(UCAS)· 空天信息创新研究院

  • Until 2021

    B.Eng. in Space Information and Digital Technology工学学士,空间信息与数字技术

    University of Electronic Science and Technology of China (UESTC)电子科技大学(UESTC)

Experience研究与实习经历

  • 2025.10 – Present

    Research Intern, Omni Research Team研究实习生,Omni 研究团队

    Tencent腾讯

  • 2021 – Present

    Research Assistant, Off-Nadir Remote Sensing Group研究助理,偏移摄影遥感方向

    Aerospace Information Research Institute, Chinese Academy of Sciences中国科学院空天信息创新研究院

  • 2021.05 – 2021.09

    Computer Vision Algorithm Intern计算机视觉算法实习生

    Tencent腾讯

Mentorship研究生指导

  • Jiajun ZhangJiajun Zhang

    M.Sc., City University of Hong Kong · co-mentored with Prof. Xiangyu Zhao香港城市大学硕士研究生 · 与赵翔宇教授联合指导

  • Zigan ZhouZigan Zhou

    M.Sc., City University of Hong Kong · co-mentored with Prof. Xiangyu Zhao香港城市大学硕士研究生 · 与赵翔宇教授联合指导

  • Weizhi ChenWeizhi Chen

    University of Chinese Academy of Sciences (AIRCAS) · co-mentored with Prof. Jingbo Chen中国科学院大学(空天信息创新研究院)· 与陈静波教授联合指导

Teaching教学经历

  • Teaching Assistant, Optimization Methods助教,最优化方法

    Prof. Zhenjiang Zhao · University of Chinese Academy of Sciences赵振江 教授 · 中国科学院大学

  • Teaching Assistant, Academic English助教,学术英语

    Prof. James · University of Chinese Academy of SciencesJames 教授 · 中国科学院大学

Get in touch联系我

Let's talk about off-nadir vision and 3D scenes欢迎交流偏移摄影与三维场景理解

Open to faculty, postdoctoral, research scientist and algorithm engineer positions from 2026.2026 年起开放教职、博士后、研究科学家与算法工程师岗位机会。

Beijing, China / Hong Kong SAR中国北京 / 中国香港 UCAS · AIRCAS · CityU AML Lab中国科学院大学 · 空天信息创新研究院 · 香港城市大学 AML 实验室WeChat: kaili37