“欢迎访问北京理工大学学报自然科学中文版网站!”
加入收藏 | 地图 | 联系方式      
 
基于大模型双代理协同的无人机视觉语言导航方法
Dual-Agent Collaborative Aerial Vision-and-Language Navigation Based on Large Vision-Language Model
投稿时间:2026-07-31  修订日期:2026-08-16
DOI:
中文关键词:  无人机  视觉语言导航  多模态大模型  零样本推理
English Keywords:UAV? vision-and-language navigation? vision-language model? zero-shot inference
基金项目:
作者单位邮编
詹昭焕 深圳北理莫斯科大学工程系 518172
敖艺廊 北京理工大学机电学院 518172
余丽莎* 深圳信息职业技术大学 518172
周麦琪 深圳北理莫斯科大学 518172
晏楚阳 深圳北理莫斯科大学 518172
赵劲驰 深圳北理莫斯科大学 518172
卢昭羽 深圳北理莫斯科大学 518172
摘要点击次数: 110
全文下载次数: 0
中文摘要:
      视觉语言大模型具备较强的零样本语义理解能力,为降低无人机视觉语言导航对任务特定训练数据的依赖提供了新的技术路径。然而,直接利用大模型从城市俯视地图预测导航目标仍面临两方面问题:一方面,无人机航向在世界坐标系中表示,而视觉观测基于图像坐标系,现有方法缺少对二者转换关系的显式建模,容易产生方向理解偏差;另一方面,导航器通常直接预测目标像素坐标,易出现坐标越界、语义匹配但定位偏移等问题。针对上述问题,提出一种面向CityNav任务的零样本双代理协同导航方法。该方法将高层导航推理解耦为感知代理与规划代理:感知代理通过方向感知投影,将世界坐标系下的无人机航向映射至图像坐标系,从而为方向关系推理提供统一的坐标基准。在此基础上,该代理结合局部区域细化与候选节点管理,将目标相关的视觉证据、空间关系及地标信息组织为结构化候选集合;规划代理在候选集合上综合多源证据进行受约束的离散选择,将开放式坐标生成转化为候选节点决策。选定节点经确定性坐标变换后被转换为世界坐标,局部控制器据此生成无人机动作。整个导航过程无需进行任务特定微调。CityNav数据集上的实验结果表明,该方法在未见测试集上成功率达到22.03%,理想成功率达到44.16%,较先进方法分别提升0.83与8.78个百分点,验证了基于大模型双代理协同的导航方法的有效性。
English Summary:
      Large vision-language models (VLMs) possess strong zero-shot semantic understanding capabilities, offering a promising approach to reducing the dependence of UAV vision-language navigation on task-specific training data. However, directly employing VLMs to predict navigation goals from urban overhead maps presents two major challenges. First, UAV headings are represented in the world coordinate system, whereas visual observations are defined in the image coordinate system. Existing methods lack explicit modeling of the transformation between these two coordinate systems, which can lead to errors in directional understanding. Second, navigators typically predict target pixel coordinates directly, making them susceptible to out-of-bounds predictions and localization errors despite correct semantic matching. To address these issues, we propose a zero-shot dual-agent collaborative navigation method for the CityNav task. The method decomposes high-level navigation reasoning into a perception agent and a planning agent. Through direction-aware projection, the perception agent maps the UAV heading from the world coordinate system to the image coordinate system, thereby establishing a unified coordinate reference for directional reasoning. Building upon this projection, the agent combines local-region refinement with candidate-node management to organize target-relevant visual evidence, spatial relationships, and landmark information into a structured candidate set. The planning agent then integrates multi-source evidence to perform constrained discrete selection over the candidate set, reformulating open-ended coordinate generation as candidate-node selection. The selected node is converted into world coordinates through a deterministic coordinate transformation, based on which a local controller generates UAV actions. The entire navigation process requires no task-specific fine-tuning. Experimental results on the CityNav dataset show that the proposed method achieves a success rate of 22.03% and an oracle success rate of 44.16% on the unseen test split, outperforming the state-of-the-art method by 0.83 and 8.78 percentage points, respectively. These results demonstrate the effectiveness of the proposed LVLM-based dual-agent collaborative navigation method.
  查看/发表评论  下载PDF阅读器

您是第18045049位访问者  今日共有 1588访问者
版权所有:北京理工大学学术期刊办公室
主管单位:中华人民共和国工业和信息化部 主办单位:北京理工大学 地址:北京海淀区中关村南大街5号
技术支持:北京勤云科技发展有限公司