智能网联汽车基于逆强化学习的轨迹规划优化机制研究

Research on Inverse Reinforcement Learning-Based Trajectory Planning Optimization Mechanism for Autonomous Connected Vehicles

  • 摘要: 针对当前轨迹规划策略存在实时性差、优化目标权重系数难以标定、模仿学习方法可解释性差等问题,提出了基于最大熵原则的逆强化学习方法,通过学习经验驾驶员驾驶轨迹的内在优化机制,从而规划出符合人类驾驶经验的整体最优的换道专家轨迹,为解决轨迹规划方法的实时性问题和可解释性问题奠定了理论基础. 以一般风险场景和高风险场景为应用案例,通过Matlab/Simulink 仿真验证了所提逆强化学习方法实现轨迹规划的可行性与有效性.

     

    Abstract: Trajectory planning is one of the most significant technologies of autonomous connected vehicles. However, there are some problems in existing trajectory planning strategy, for example, weak real-time ability, difficult to calibrate weighting coefficients of optimization objectives and the poor interpretability for direct imitation learning method in the trajectory planning strategy. Therefore, an inverse reinforcement learning (IRL) method was proposed based on the maximum entropy principle in this paper. Learning the underlying optimization mechanism of driving trajectories from experienced drivers, the planning of lane-changing expert trajectories was achieved aligning with the human driving experience, laying a theoretical foundation for solving the real-time and interpretability problems of trajectory planning methods. Finally, taking general risk scenarios and high-risk scenarios as application cases respectively, the feasibility and effectiveness of the proposed trajectory planning method were validated through Matlab/Simulink simulations.

     

/

返回文章
返回