Lightweight Deep Learning Framework for UAV Object Detection Using Adaptive Multi-Scale Feature Fusion

Authors

  • Yifei Gao School of Electrical and Control Engineering, Liaoning Technical University, Xingcheng, Huludao, Liaoning 125105, China.
  • Shengzhi Dong College of Intelligent Science and Engineering, Liaoning Technical University, Xingcheng, Huludao, Liaoning 125105, China.

DOI:

https://doi.org/10.56979/1102/2026/1658

Keywords:

UAV object detection, VisDrone, lightweight detection, small objects, multi-scale feature fusion, adaptive feature fusion, YOLO26n

Abstract

UAV imagery contains many small and densely distributed objects whose apparent scale changes sharply with altitude and viewpoint. This study defines an architecture framework for lightweight UAV object detection that combines a nano-scale one-stage detector, a high-resolution P2 branch, and an Adaptive Multi-Scale Feature Fusion (AMFF) module operating across P2-P5 features. AMFF aligns neighbouring feature maps spatially and by channel, estimates image-dependent scale weights from pooled descriptors, normalizes the weights with softmax, and applies residual weighted fusion. The evaluation protocol is anchored to the public VisDrone2019-DET configuration, comprising 6,471 training images, 548 validation images, and 1,610 test-dev images across ten classes. Source-backed architecture analysis shows that the standard YOLO26n configuration contains 2,572,280 parameters and 6.1 GFLOPs, whereas the P2 variant contains 2,662,400 parameters and 9.5 GFLOPs. Thus, adding P2 increases parameter count by only 3.50% but increases GFLOPs by 55.74%, showing that high-resolution detection is parameter-light but not compute-neutral. The supplied experimental package does not contain executed AMFF checkpoints, prediction files, or hardware timing logs. Accordingly, mAP@0.50:0.95, precision, recall, AMFF accuracy gain, model size after export, and inference latency are not fabricated or inferred. The current evidence establishes the reproducible architecture and controlled evaluation design, while the effectiveness of adaptive fusion requires validation through the defined baseline-to-AMFF experiments on the stated VisDrone splits.

Downloads

Published

2026-09-01

How to Cite

Yifei Gao, & Shengzhi Dong. (2026). Lightweight Deep Learning Framework for UAV Object Detection Using Adaptive Multi-Scale Feature Fusion. Journal of Computing & Biomedical Informatics, 11(02). https://doi.org/10.56979/1102/2026/1658