Fine-Tuning Panoptic FPN with ResNet-50 for Maritime Obstacle Detection on the LaRS Dataset

Authors

  • Istifa Shania Putri Telkom University, Indonesia Author
  • Sugih Ahmad Fauzan Bandung Institute of Technology, Indonesia Author
  • Mega Fitri Yani Telkom University, Indonesia Author
  • Cindy Muhdiantini Telkom University, Indonesia Author

DOI:

https://doi.org/10.5281/zenodo.20768168

Keywords:

Panoptic Segmentation, Fine-Tunning, Mask R-CNN, Feature Pyramid Networks, Maritime Obstacle Detection, Detection LaRS Dataset

Abstract

Maritime obstacle detection is a critical challenge for Unmanned Surface Vehicles (USVs) operating in complex and dynamic environments. This study investigates the effectiveness of fine-tuning Panoptic FPN, a Mask R-CNN-based architecture augmented with Feature Pyramid Networks, for panoptic segmentation on the LaRS (Lake, River, Seas) dataset. Unlike prior work that explored model comparisons broadly, this research focuses specifically on the impact of hyperparameter tuning and backbone selection on maritime panoptic segmentation performance. Through systematic ablation studies, we demonstrate that adjusting the learning rate to 0.002 and the gamma decay factor to 0.2 yields significant improvements. Our fine-tuned Panoptic FPN with a ResNet-50 backbone achieves a Panoptic Quality (PQ) of 45.31%, surpassing the previous state-of-the-art Mask2Former Swin-B (41.7%) by 3.61 percentage points. Notably, ResNet-50 outperforms the deeper ResNet-101 backbone (36.47% PQ), suggesting that heavier architectures may overfit on domain-specific maritime datasets. Furthermore, Panoptic FPN requires only 8 hours of training compared to approximately 2 days for Mask2Former Swin-L, demonstrating superior computational efficiency. These findings highlight that targeted fine-tuning of lightweight architectures can outperform larger transformer-based models in maritime panoptic segmentation tasks.

Downloads

Download data is not yet available.

References

Sánchez-Beaskoetxea, J., Basterretxea-Iribar, I., Sotés, I., & Maruri Machado, M. de las M. (2021). Human error in marine accidents: Is the crew normally to blame? Maritime Transport Research, 2, 100016. https://doi.org/10.1016/j.martra.2021.100016

Jin, J., Liu, D., Li, F., Dai, Y., Li, L., & Ma, Y. (2024). Wide baseline stereovision based obstacle detection for unmanned surface vehicles. Signal, Image and Video Processing. https://doi.org/10.1007/s11760-024-03098-0

Žust, L., Perš, J., & Kristan, M. (2023). LaRS: A diverse panoptic maritime obstacle detection dataset and benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 20247–20257). https://doi.org/10.1109/ICCV51070.2023.01857

Jaus, A., Yang, K., & Stiefelhagen, R. (2023). Panoramic panoptic segmentation: Insights into surrounding parsing for mobile agents via unsupervised contrastive learning. IEEE Transactions on Intelligent Transportation Systems, 24(4), 4438–4453. https://doi.org/10.1109/TITS.2022.3232897

Dou, Y., et al. (2024). PanopticUAV: Panoptic segmentation of UAV images for marine environment monitoring. Computer Modeling in Engineering & Sciences, 138(1), 1001–1014. https://doi.org/10.32604/cmes.2023.027764

Cheng, B., et al. (2020). Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12472–12482). https://doi.org/10.1109/CVPR42600.2020.01249

Kirillov, A., Girshick, R., He, K., & Dollár, P. (2019). Panoptic feature pyramid networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6392–6401). https://doi.org/10.1109/CVPR.2019.00656

Qiao, D., Liu, G., Li, W., Lyu, T., & Zhang, J. (2022). Automated full scene parsing for marine ASVs using monocular vision. Journal of Intelligent & Robotic Systems, 104(2), 37. https://doi.org/10.1007/s10846-021-01543-7

Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1280–1289). https://doi.org/10.1109/CVPR52688.2022.00135

Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2016). Feature pyramid networks for object detection. arXiv:1612.03144. https://arxiv.org/abs/1612.03144

Liu, Z., et al. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. arXiv:2103.14030. https://arxiv.org/abs/2103.14030

Li, X., & Chen, D. (2022). A survey on deep learning-based panoptic segmentation. Digital Signal Processing, 120, 103283. https://doi.org/10.1016/j.dsp.2021.103283

Gengwei, Z., Gao, Y., Xu, H., Zhang, H., Li, Z., & Liang, X. (2021). Ada-Segment: Automated multi-loss adaptation for panoptic segmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 35, 3333–3341. https://doi.org/10.1609/aaai.v35i4.16445

Milioto, A., Behley, J., McCool, C., & Stachniss, C. (2020). LiDAR panoptic segmentation for autonomous driving. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 8505–8512). https://doi.org/10.1109/IROS45743.2020.9340837

Zendel, O., Schörghuber, M., Rainer, B., Murschitz, M., & Beleznai, C. (2022). Unifying panoptic segmentation for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 21319–21328). https://doi.org/10.1109/CVPR52688.2022.02066

Wang, H., Zhu, Y., Adam, H., Yuille, A., & Chen, L.-C. (2021). MaX-DeepLab: End-to-end panoptic segmentation with mask transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5459–5470). https://doi.org/10.1109/CVPR46437.2021.00542

Ashrafi, S. S., Shokouhi, S. B., & Ayatollahi, A. (2023). Still image action recognition based on interactions between joints and objects. Multimedia Tools and Applications, 82(17), 25945–25971. https://doi.org/10.1007/s11042-023-14350-z

Elharrouss, O., Al-Maadeed, S. A., Subramanian, N., Ottakath, N., Almaadeed, N., & Himeur, Y. (2021). Panoptic segmentation: A review. arXiv:2111.10250. https://arxiv.org/abs/2111.10250

Sellat, Q., Bisoy, S. K., & Priyadarshini, R. (2022). Semantic segmentation for self-driving cars using deep learning: A survey. In Cognitive Big Data Intelligence with a Metaheuristic Approach (pp. 211–238). Academic Press. https://doi.org/10.1016/B978-0-323-85117-6.00002-9

Darbyshire, M., Sklar, E., & Parsons, S. (2023). Hierarchical Mask2Former: Panoptic segmentation of crops, weeds and leaves. In Proceedings of the IEEE International Conference on Robotics and Biomimetics (ROBIO).

Sun, H., et al. (2023). Brain tumor image segmentation based on improved FPN. BMC Medical Imaging, 23(1), 172. https://doi.org/10.1186/s12880-023-01131-1

Bovcon, B., Muhovič, J., Vranac, D., Mozetič, D., Perš, J., & Kristan, M. (2022). MODS: A USV-oriented object detection and obstacle segmentation benchmark. IEEE Transactions on Intelligent Transportation Systems, 23(8), 13403–13418. https://doi.org/10.1109/TITS.2021.3124192

Downloads

Published

02-07-2026

Issue

Section

Articles

How to Cite

Fine-Tuning Panoptic FPN with ResNet-50 for Maritime Obstacle Detection on the LaRS Dataset. (2026). SITEKNIK: Sistem Informasi, Teknik Dan Teknologi Terapan, 3(3), 108-115. https://doi.org/10.5281/zenodo.20768168

Share

Similar Articles

1-10 of 13

You may also start an advanced similarity search for this article.