Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks

Sai Likhith Karri; Ansh Saxena

arXiv:2510.25797·cs.CV·October 31, 2025

Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks

Sai Likhith Karri, Ansh Saxena

PDF

TL;DR

This paper explores the enhancement of underwater object detection by integrating spatio-temporal modeling and spatial attention mechanisms into YOLOv5, demonstrating significant improvements in detection accuracy in complex marine environments.

Contribution

The study introduces a novel combination of T-YOLOv5 with CBAM, showing how spatial attention and temporal modeling together improve underwater object detection performance.

Findings

01

T-YOLOv5 outperforms standard YOLOv5 in accuracy.

02

Adding CBAM further improves detection in challenging scenarios.

03

Models show superior generalization in dynamic marine environments.

Abstract

This study examines the effectiveness of spatio-temporal modeling and the integration of spatial attention mechanisms in deep learning models for underwater object detection. Specifically, in the first phase, the performance of temporal-enhanced YOLOv5 variant T-YOLOv5 is evaluated, in comparison with the standard YOLOv5. For the second phase, an augmented version of T-YOLOv5 is developed, through the addition of a Convolutional Block Attention Module (CBAM). By examining the effectiveness of the already pre-existing YOLOv5 and T-YOLOv5 models and of the newly developed T-YOLOv5 with CBAM. With CBAM, the research highlights how temporal modeling improves detection accuracy in dynamic marine environments, particularly under conditions of sudden movements, partial occlusions, and gradual motion. The testing results showed that YOLOv5 achieved a mAP@50-95 of 0.563, while T-YOLOv5 and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.