面向端到端目标检测神经网络的高效硬件加速系统设计

Efficient Hardware Acceleration System Design for End-to-End Object Detection Neural Network

  • 摘要: 针对神经网络目标检测系统在硬件资源受限与功耗敏感的边缘计算设备中应用的问题,提出了一种基于现场可编程门阵列(FPGA)实现的YOLOv3-Tiny神经网络目标检测硬件加速系统. 利用网络结构重组、层间融合与动态数值量化,缩减YOLOv3-Tiny网络规模. 基于通道并行与权值驻留硬件加速算法、紧密流水线处理流程与硬件运算单元复用,提升硬件资源利用效率. 所设计的端到端目标检测加速系统被部署在UltraScale+ XCZU9EG FPGA上,达到了96.6 GOPS的吞吐量与17.3 FPS的检测帧率,功耗为4.12 W,并具有0.32 GOPS/DSP与2.68 GOPS/kLUT的硬件资源利用效率. 在保持高效准确目标检测能力的同时,硬件资源利用效率优于其他已有的YOLOv3-Tiny目标检测硬件加速器.

     

    Abstract: To solve the problem of limited hardware resources and sensitive power consumption in the application of neural network object detection system for edge computing devices, a YOLOv3-Tiny neural network object detection hardware acceleration system was proposed based on field programmable gate array (FPGA). The scale of YOLOv3-Tiny network was reduced by using network structure reorganization, inter layer fusion and dynamic numerical quantization. Based on channel parallel and weight resident hardware acceleration algorithm, tight pipeline processing flow and hardware operation unit reuse, the utilization efficiency of hardware resources was improved. The designed end-to-end object detection acceleration system was deployed on UltraScale+ XCZU9EG FPGA. The result shows that it can achieve 96.6 GOPS throughput, 17.3 FPS detection frame rate and 4.12 W power consumption. The hardware resource utilization efficiency is 0.32 GOPS/DSP and 2.68 GOPS/kLUT. Maintaining efficient and accurate object detection capability, the utilization efficiency of hardware resources is better than other existing YOLOv3-Tiny object detection hardware accelerators.

     

/

返回文章
返回