Skip to content

Research & Publications研究与论文

Six chapters, one thread六个章节,一条主线

My work is organised as a narrative rather than a flat list: Part I builds the Offset Token system for off-nadir building extraction; Part II extends 3D spatial understanding to vision-language models. Expand any chapter for the figure, abstract, highlights and resources.我把研究成果组织成一条叙事线索,而不是一份平铺的清单:第一部分构建偏移摄影下建筑物提取的偏移量词元体系;第二部分把三维空间理解拓展到视觉语言大模型。展开任一章节即可查看示意图、摘要、研究亮点与资源链接。

Google ScholarGoogle Scholar

PhD research main line博士生涯研究主线

Building & Spatial Structure Understanding with Offset Tokens基于偏移量词元的建筑与空间结构理解

My PhD research targets building extraction under off-nadir satellite photography. I first proposed the concept of the Offset Token, which brings off-nadir photography into the perceptual scope of large vision models. Around this concept I built a complete loop — vectorised building extraction, positional repair between historical vector maps and updated imagery, efficient optimisation of the offset token, and full decoupling of offset learning.博士期间的研究方向为遥感卫星偏移摄影视角下的建筑物提取。研究主体上, 首先提出了偏移量词元(Offset Token)的概念,将偏移摄影带入大模型的理解视角; 并围绕所提出的偏移量词元概念展开更加细腻的研究,实现了偏移摄影下矢量化的建筑物提取、 历史矢量地图与更新影像之间的位置修复、偏移量词元的高效优化计算、偏移量学习解耦等内容, 完成了「概念提出 → 具象化应用 → 任务解耦与独立 → 底层理论降维」的闭环。

CHAPTER章节 01 Background & the birth of the Offset Token背景与 Offset Token 概念的提出 Extends what a ViT token can mean in vision, and introduces the Offset Token as the carrier of spatial displacement.拓展 Transformer / ViT 词元在视觉中的内涵,首次引出偏移量词元作为空间位移的承载者。 OBMIEEE TGRS 2024
OBM Figure coming soon示意图待补充 OBM — method overview

Prompt-Driven Building Footprint Extraction in Aerial Images with Offset-Building Model

K. Li, Y. Deng, Y. Kong, D. Liu, J. Chen, Y. Meng, J. Ma, C. Wang

IEEE Transactions on Geoscience and Remote SensingIEEE Transactions on Geoscience and Remote Sensing2024Published已发表

In off-nadir aerial imagery a building roof is projected away from its real footprint, so the footprint is frequently occluded and directly invisible. OBM reformulates footprint extraction as a promptable task: alongside the usual detection and segmentation queries, the Transformer decoder receives an additional Offset Token that regresses the roof-to-footprint displacement. This is the first work to bring off-nadir displacement modelling into the promptable foundation-model paradigm.在偏移摄影的航空影像中,建筑物屋顶被投影到真实足迹之外,足迹往往被遮挡、无法直接观测。 OBM 将足迹提取重构为可提示(Promptable)任务:在 Transformer 解码器中, 除常规的检测与分割查询之外,引入额外的偏移量词元(Offset Token), 直接回归屋顶到足迹的位移。这是首次把偏移摄影的位移建模引入可提示大模型范式的工作。

CHAPTER章节 02 First application: vector-level building extractionOffset Token 概念的初步应用:矢量级建筑物提取 Applies the new token to polygonal footprint extraction, validating its power on complex 3D structures.将新提出的词元概念应用于矢量级别的建筑物提取,验证其在复杂三维结构中的有效性与潜力。 PolyFootNetIEEE TGRS 2025
PolyFootNet Figure coming soon示意图待补充 PolyFootNet — method overview

PolyFootNet: Extracting Polygonal Building Footprints in Off-Nadir Remote Sensing Images

K. Li, Y. Deng, J. Chen, Y. Meng, Z. Xi, J. Ma, C. Wang, M. Wang, X. Zhao

IEEE Transactions on Geoscience and Remote SensingIEEE Transactions on Geoscience and Remote Sensing2025Published已发表

Raster masks are not what GIS pipelines consume — maps are vectors. PolyFootNet carries the Offset Token from pixel space into polygon space: roof polygons are predicted and then translated by learned offsets into footprint polygons, producing directly usable vector products without a hand-crafted raster-to-vector post-processing chain.栅格掩膜并非 GIS 生产流程真正需要的产物——地图本质上是矢量的。PolyFootNet 把偏移量词元从像素空间带入矢量多边形空间:先预测屋顶多边形, 再由学习到的偏移量整体平移得到足迹多边形,直接产出可用的矢量成果, 免去人工设计的栅格转矢量后处理链路。

CHAPTER章节 03 From Offset Token to Alignment Token: interactive alignment从 Offset Token 到 Alignment Token:具象化与交互对齐 Instantiates the Offset Token as an Alignment Token, repairing the spatial gap between historical OSM labels and new imagery through an interactive drag mechanism.将偏移量词元具象为对齐词元(Alignment Token),通过交互式「拖拽」机制修复历史 OSM 标签与更新影像之间的空间偏移。 DragOSMarXiv 2025
DragOSM Figure coming soon示意图待补充 DragOSM — method overview

DragOSM: Extract Building Roofs and Footprints from Aerial Images by Aligning Historical Labels

K. Li, X. Weng, Y. Deng, Y. Meng, C. Pang, G.-S. Xia, X. Zhao

arXiv preprint arXiv:2509.17951arXiv 预印本 arXiv:2509.179512025Preprint预印本

Enormous amounts of historical vector labels (e.g. OpenStreetMap) already exist, but they no longer line up with freshly captured off-nadir imagery. DragOSM treats the stale label as a prompt to be dragged: an Alignment Token learns where the historical polygon should move so that roof and footprint are both recovered. Instead of re-annotating from scratch, existing map assets are repaired and reused.世界上已经积累了海量历史矢量标签(如 OpenStreetMap),但它们与新采集的偏移摄影影像 不再对齐。DragOSM 将过期标签视为可被拖拽的提示:由对齐词元 (Alignment Token)学习历史多边形应当移动到何处,从而同时恢复屋顶与足迹。 这使得已有地图资产可被修复复用,而无需从零重新标注。

CHAPTER章节 04 Decoupling offset learning & building a benchmark偏移量学习的解耦、独立模块化与基准构建 Strips offset learning out of semantic segmentation entirely, defines it as the standalone RFOV extraction task, and releases a benchmark plus baseline.将偏移量学习从语义分割中彻底剥离,定义为独立的 RFOV 提取任务,并给出基准数据集与基线模型。 DragRoofarXiv 2026
DragRoof Figure coming soon示意图待补充 DragRoof — method overview

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

K. Li, Y. Deng, L. Deng, Z. Xi, C. Wang, J. Zhang, Y. Ji, Y. Meng, X. Zhao

arXiv preprint arXiv:2607.25210arXiv 预印本 arXiv:2607.252102026Preprint预印本

Offset estimation had always ridden along with segmentation. This work makes it first-class: the Roof-to-Footprint Offset Vector (RFOV) becomes an independent extraction task, the dragging process is modelled as an ordinary differential equation, and the community gets the ObliCity benchmark together with the DragRoof baseline model.以往偏移量估计总是搭载在分割任务之上。本工作将其提升为一等问题:把 屋顶到足迹偏移向量(RFOV)定义为独立的提取任务, 以常微分方程(ODE)建模模拟拖拽过程,并构建了 ObliCity 基准数据集与 DragRoof 基线模型。

CHAPTER章节 05 Low-dimensional, efficient Offset Tokens底层机制优化与降维推导:低维高效的偏移标识 Derives the underlying mathematics from a 3D-vision standpoint: the ideal offset mapping is fully described by 7 variables, which collapses terminal compute cost.从三维视觉(3DV)角度深入剖析偏移量词元的底层数学推导:理想的 offset 映射可由 7 维变量充分描述,从而大幅降低终端计算复杂度。 LODEOTPreprint预印本
LODEOT Figure coming soon示意图待补充 LODEOT — method overview

LODEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery

K. Li, et al.

2025Preprint预印本

Why should a displacement need a high-dimensional token at all? Viewed through 3D vision, the ideal roof-to-footprint mapping of a whole scene is sufficiently described by only 7 variables. LODEOT exploits this to pair high-dimensional queries (detection / segmentation) with a low-dimensional offset token, cutting inference cost on the terminal side while keeping accuracy.位移信息真的需要高维词元来承载吗?从三维视觉的角度看,整幅场景理想的 屋顶到足迹映射仅需7 个变量 即可充分描述。LODEOT 据此 将高维查询(检测 / 分割)与低维偏移标识高效结合, 在保持精度的同时显著降低终端推理开销。

Spatial understanding空间维度

Multimodal LLMs for 3D Scene Understanding, with Reinforcement Learning多模态大模型驱动的 3D 场景理解与强化学习

After accumulating substantial experience in 3D spatial understanding, this line extends my work towards today's frontier: letting large multimodal models reason about 3D structure directly, without camera calibration or dense views — and, building on that foundation, using reinforcement learning to keep sparse-view indoor reconstruction consistent across viewpoints.在积累了丰富的三维空间理解经验之后,这一部分把工作拓展到当下最前沿的方向: 让多模态大模型在没有相机参数、没有密集视图的条件下直接推理三维结构;并在此基础上, 尝试用强化学习让稀疏视角下的室内场景重建保持跨视角一致。

CHAPTER章节 06 Sparse-view 3D reconstruction via executable scene programs基于可执行场景程序的稀疏视图 3D 场景重建 Drops the reliance on camera parameters and dense views: a VLM infers an executable scene program straight from sparse RGB observations.摆脱对相机参数与密集视图的依赖:由视觉语言大模型直接从稀疏 RGB 观测推导可执行的场景程序。 ScenixarXiv 2026
Scenix Figure coming soon示意图待补充 Scenix — method overview

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

K. Li, L. Jiang, Z. Li, J. Dong, J. Zhang, Y. Yin, R. Zhang, K. Yan, X. Huang, et al.

arXiv preprint arXiv:2608.07012arXiv 预印本 arXiv:2608.070122026Preprint预印本

Classical reconstruction wants calibrated cameras and many views. Scenix instead asks a vision-language model (LayoutVLM) to read a handful of RGB images and emit an executable scene program — jointly predicting room layout, object attributes and semantic descriptions, which are then instantiated into a fully editable 3D scene.传统重建依赖标定过的相机与密集视图。Scenix 转而让视觉语言大模型(LayoutVLM) 读取少量 RGB 图像,直接输出可执行的场景程序(Scene Program)—— 联合预测房间结构、物体属性与语义描述,并最终实例化为完全可编辑的 三维场景

Reinforcement-Learning-Driven, Sparse-View-Consistent Indoor Scene Reconstruction基于强化学习的稀疏视角一致性室内场景重建 Extending Scene Programs with reinforcement learning for cross-view consistency在场景程序基础上引入强化学习,实现稀疏视角下的一致性重建 On Working研究中

Early-stage project building on the Scene-Program idea behind Scenix: using reinforcement learning objectives so that indoor reconstructions from sparse, uncalibrated views stay geometrically and semantically consistent across viewpoints.

早期阶段项目,延续 Scenix 中场景程序(Scene Program)的思路:引入 强化学习目标,使稀疏、无标定视角下的室内场景重建在几何与语义上 保持跨视角一致。

Ongoing research在研方向

Current Focus: Memory in Video World Models当前重心:视频世界模型中的记忆机制

Video world models learn to predict how a scene evolves, but they largely lack a persistent notion of what they have already seen. My current work asks how such a model should remember: how to retain scene state over long horizons, stay consistent when the camera revisits a place, and keep memory affordable as the rollout grows. This is the temporal counterpart to the spatial understanding I built during my PhD.视频世界模型能够预测场景如何演化,却普遍缺少对「已经看到过什么」的持久表征。 我目前的研究关注这类模型应当如何记忆:如何在长时序上保持场景状态、 如何在镜头重新回到同一位置时维持一致、以及如何在推演变长时控制记忆开销。 这与我博士期间构建的空间理解能力,构成时间维度上的呼应。

Memory Mechanisms for Video World Models视频世界模型的记忆机制 How should a video world model remember what it has already seen?视频世界模型应当如何记住它已经看到过的内容? On Working研究中

Ongoing research on memory in video world models: giving a generative video model a persistent representation of scene state so that long-horizon rollouts stay self-consistent, revisited viewpoints agree with what was generated before, and the cost of remembering does not grow unchecked with sequence length.

正在进行的视频世界模型记忆机制研究:为生成式视频模型引入对场景状态的 持久表征,使长时序推演保持自洽、镜头回到旧视角时与先前生成的内容一致, 并让记忆开销不随序列长度无约束增长。

Other Collaborative Publications其他合作论文

2026 Remote-Sensing City Layout Extraction with MLLM Z. Zhou, K. Li, Y. Deng. arXiv preprint arXiv:2608.16484, 2026. · arXiv 2026 ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control H. Wang, Z. Li, Y. Zhang, Z. He, L. Jiang, K. Li, Y. Zhao, L. Fan, W. Hou, T. Liang, et al. arXiv preprint arXiv:2605.09869, 2026. · arXiv 2025 DGTRSD and DGTRSCLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision-Language Foundation Model for Alignment W. Chen, Y. Deng, W. Jin, J. Chen, J. Chen, Y. Feng, Z. Xi, D. Liu, K. Li, Y. Meng. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS), 2025. · IEEE JSTARS 2025 IRSAMap: Towards Large-Scale, High-Resolution Land Cover Map Vectorization Y. Meng, L. Deng, Z. Xi, J. Chen, J. Chen, A. Yue, D. Liu, K. Li, C. Wang, K. Li, et al. IEEE Transactions on Geoscience and Remote Sensing, 2025. · IEEE TGRS 2025 PCP: A Prompt-Based Cartographic-Level Polygonal Vector Extraction Framework for Remote Sensing Images C. Wang, Z. Xi, D. Liu, Y. Feng, Y. Deng, K. Li, J. Chen, J. Chen, Y. Meng. IEEE Transactions on Geoscience and Remote Sensing, 2025. · IEEE TGRS 2025 LRSCLIP: A Vision-Language Foundation Model for Aligning Remote Sensing Image with Longer Text W. Chen, J. Chen, Y. Deng, J. Chen, Y. Feng, Z. Xi, D. Liu, K. Li, Y. Meng. arXiv preprint arXiv:2503.19311, 2025. · arXiv 2025 SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion J. Ma, S. Wang, J. Yang, J. Hu, J. Liang, G. Lin, K. Li, Y. Meng. arXiv preprint arXiv:2502.11515, 2025. · arXiv 2025 CusMer: Multimodal Intent Recognition in Customer Service via Data Augment and LLM Merge Z. Li, B. Wu, Y. Zhang, X. Li, K. Li, W. Chen. Companion Proceedings of the ACM Web Conference (WWW Companion), 2025, pp. 3058-3062. · WWW Companion 2024 SAMPolyBuild: Adapting the Segment Anything Model for polygonal building extraction C. Wang, J. Chen, Y. Meng, Y. Deng, K. Li, Y. Kong. ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 218: 707-720. · ISPRS J. P&RS 2022 基于边界曲线的多因素机场出租车管理系统模型 程荟璇, 贺益鑫, 李锴, 覃思义. 实验科学与技术, 2022, 20(5): 40-44. · 实验科学与技术 2019 Land Price Assessment Based on Deep Neural Network A. Hou, J. Liu, Y. Tao, S. Jiang, K. Li, Z. Zheng, J. Xia, Y. He, M. Zhu, G. Zhou, et al. IGARSS 2019, Yokohama, Japan. · IGARSS 2019 Urban Functional Regions Discovering Based on Deep Learning F. Mou, R. Kong, K. Li, Z. Zheng, J. Xia, Y. He, M. Zhu, G. Zhou, H. Zhang, Z. Liu, et al. IGARSS 2019, Yokohama, Japan, pp. 9462-9465. · IGARSS 2019 A WordNet-Based Geospatial Web Services Search Method Supporting Quality of Service Constraints L. Jiang, K. Li, D. Liu, Y. Zhou, Z. Zheng, F. Huang. IGARSS 2019, Yokohama, Japan, pp. 863-866. · IGARSS