Full interactive site完整互动版网站

LI Kai李锴PhD Candidate · UCAS × CityU Hong Kong中国科学院大学 × 香港城市大学 联合培养博士生

Multimodal LLMs · Video World Models · 3D Spatial Understanding多模态大模型 · 视频世界模型 · 三维空间理解

My research centres on large multimodal models and how they acquire spatial and temporal understanding. I am currently working on memory in video world models, alongside reinforcement learning for sparse-view 3D scene reconstruction. Earlier, during my PhD, I proposed the Offset Token, which brought off-nadir photogrammetry into the perceptual scope of vision foundation models.我的研究以多模态大模型为核心,关注模型如何获得空间与时间层面的理解能力。目前正在研究视频世界模型(Video World Model)中的记忆机制,并同时探索用强化学习实现稀疏视角三维场景重建。博士期间,我提出了偏移量词元(Offset Token)的概念,将偏移摄影带入视觉大模型的理解视角。

Publications发表论文

Building & Spatial Structure Understanding with Offset Tokens基于偏移量词元的建筑与空间结构理解

[1] Prompt-Driven Building Footprint Extraction in Aerial Images with Offset-Building Model K. Li, Y. Deng, Y. Kong, D. Liu, J. Chen, Y. Meng, J. Ma, C. Wang· IEEE TGRS, 2024

[2] PolyFootNet: Extracting Polygonal Building Footprints in Off-Nadir Remote Sensing Images K. Li, Y. Deng, J. Chen, Y. Meng, Z. Xi, J. Ma, C. Wang, M. Wang, X. Zhao· IEEE TGRS, 2025

[3] DragOSM: Extract Building Roofs and Footprints from Aerial Images by Aligning Historical Labels K. Li, X. Weng, Y. Deng, Y. Meng, C. Pang, G.-S. Xia, X. Zhao· arXiv, 2025

[4] ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction K. Li, Y. Deng, L. Deng, Z. Xi, C. Wang, J. Zhang, Y. Ji, Y. Meng, X. Zhao· arXiv, 2026

[5] LODEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery K. Li, et al.· 2025

Multimodal LLMs for 3D Scene Understanding, with Reinforcement Learning多模态大模型驱动的 3D 场景理解与强化学习

[6] Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs K. Li, L. Jiang, Z. Li, J. Dong, J. Zhang, Y. Yin, R. Zhang, K. Yan, X. Huang, et al.· arXiv, 2026

· Reinforcement-Learning-Driven, Sparse-View-Consistent Indoor Scene Reconstruction基于强化学习的稀疏视角一致性室内场景重建 On Working研究中

Current Focus: Memory in Video World Models当前重心:视频世界模型中的记忆机制

· Memory Mechanisms for Video World Models视频世界模型的记忆机制 On Working研究中

Other Collaborative Publications其他合作论文

[7] Remote-Sensing City Layout Extraction with MLLM Z. Zhou, K. Li, Y. Deng. arXiv preprint arXiv:2608.16484, 2026.

[8] ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control H. Wang, Z. Li, Y. Zhang, Z. He, L. Jiang, K. Li, Y. Zhao, L. Fan, W. Hou, T. Liang, et al. arXiv preprint arXiv:2605.09869, 2026.

[9] DGTRSD and DGTRSCLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision-Language Foundation Model for Alignment W. Chen, Y. Deng, W. Jin, J. Chen, J. Chen, Y. Feng, Z. Xi, D. Liu, K. Li, Y. Meng. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS), 2025.

[10] IRSAMap: Towards Large-Scale, High-Resolution Land Cover Map Vectorization Y. Meng, L. Deng, Z. Xi, J. Chen, J. Chen, A. Yue, D. Liu, K. Li, C. Wang, K. Li, et al. IEEE Transactions on Geoscience and Remote Sensing, 2025.

[11] PCP: A Prompt-Based Cartographic-Level Polygonal Vector Extraction Framework for Remote Sensing Images C. Wang, Z. Xi, D. Liu, Y. Feng, Y. Deng, K. Li, J. Chen, J. Chen, Y. Meng. IEEE Transactions on Geoscience and Remote Sensing, 2025.

[12] LRSCLIP: A Vision-Language Foundation Model for Aligning Remote Sensing Image with Longer Text W. Chen, J. Chen, Y. Deng, J. Chen, Y. Feng, Z. Xi, D. Liu, K. Li, Y. Meng. arXiv preprint arXiv:2503.19311, 2025.

[13] SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion J. Ma, S. Wang, J. Yang, J. Hu, J. Liang, G. Lin, K. Li, Y. Meng. arXiv preprint arXiv:2502.11515, 2025.

[14] CusMer: Multimodal Intent Recognition in Customer Service via Data Augment and LLM Merge Z. Li, B. Wu, Y. Zhang, X. Li, K. Li, W. Chen. Companion Proceedings of the ACM Web Conference (WWW Companion), 2025, pp. 3058-3062.

[15] SAMPolyBuild: Adapting the Segment Anything Model for polygonal building extraction C. Wang, J. Chen, Y. Meng, Y. Deng, K. Li, Y. Kong. ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 218: 707-720.

[16] 基于边界曲线的多因素机场出租车管理系统模型 程荟璇, 贺益鑫, 李锴, 覃思义. 实验科学与技术, 2022, 20(5): 40-44.

[17] Land Price Assessment Based on Deep Neural Network A. Hou, J. Liu, Y. Tao, S. Jiang, K. Li, Z. Zheng, J. Xia, Y. He, M. Zhu, G. Zhou, et al. IGARSS 2019, Yokohama, Japan.

[18] Urban Functional Regions Discovering Based on Deep Learning F. Mou, R. Kong, K. Li, Z. Zheng, J. Xia, Y. He, M. Zhu, G. Zhou, H. Zhang, Z. Liu, et al. IGARSS 2019, Yokohama, Japan, pp. 9462-9465.

[19] A WordNet-Based Geospatial Web Services Search Method Supporting Quality of Service Constraints L. Jiang, K. Li, D. Liu, Y. Zhou, Z. Zheng, F. Huang. IGARSS 2019, Yokohama, Japan, pp. 863-866.

Education教育背景

Experience研究与实习经历

Mentorship研究生指导

Teaching教学经历

Awards & Honours获奖与荣誉

PhD Studies博士阶段

Undergraduate Studies本科阶段

Get in touch联系我

UCAS · AIRCAS · CityU AML Lab中国科学院大学 · 空天信息创新研究院 · 香港城市大学 AML 实验室
Beijing, China / Hong Kong SAR中国北京 / 中国香港
likai211@mails.ucas.ac.cn
WeChat: kaili37