Mix-layers semantic extraction and multi-scale aggregation transformer for semantic segmentation

Tianping Li; Xiaolong Yang; Zhenyi Zhang; Zhaotong Cui; Zhou Maoxia

doi:10.1007/s40747-024-01650-6

Complex & Intelligent Systems (Nov 2024)

Mix-layers semantic extraction and multi-scale aggregation transformer for semantic segmentation

Tianping Li,
Xiaolong Yang,
Zhenyi Zhang,
Zhaotong Cui,
Zhou Maoxia

Affiliations

Tianping Li: School of Physics and Electronics, Shandong Normal University
Xiaolong Yang: School of Physics and Electronics, Shandong Normal University
Zhenyi Zhang: School of Physics and Electronics, Shandong Normal University
Zhaotong Cui: School of Physics and Electronics, Shandong Normal University
Zhou Maoxia: School of Physics and Electronics, Shandong Normal University

DOI: https://doi.org/10.1007/s40747-024-01650-6
Journal volume & issue: Vol. 11, no. 1
pp. 1 – 15

Abstract

Read online

Abstract Recently, a number of vision transformer models for semantic segmentation have been proposed, with the majority of these achieving impressive results. However, they lack the ability to exploit the intrinsic position and channel features of the image and are less capable of multi-scale feature fusion. This paper presents a semantic segmentation method that successfully combines attention and multiscale representation, thereby enhancing performance and efficiency. This represents a significant advancement in the field. Multi-layers semantic extraction and multi-scale aggregation transformer decoder (MEMAFormer) is proposed, which consists of two components: mix-layers dual channel semantic extraction module (MDCE) and semantic aggregation pyramid pooling module (SAPPM). The MDCE incorporates a multi-layers cross attention module (MCAM) and an efficient channel attention module (ECAM). In MCAM, horizontal connections between encoder and decoder stages are employed as feature queries for the attention module. The hierarchical feature maps derived from different encoder and decoder stages are integrated into key and value. To address long-term dependencies, ECAM selectively emphasizes interdependent channel feature maps by integrating relevant features across all channels. The adaptability of the feature maps is reduced by pyramid pooling, which reduces the amount of computation without compromising performance. SAPPM is comprised of several distinct pooled kernels that extract context with a deeper flow of information, forming a multi-scale feature by integrating various feature sizes. The MEMAFormer-B0 model demonstrates superior performance compared to SegFormer-B0, exhibiting gains of 4.8%, 4.0% and 3.5% on the ADE20K, Cityscapes and COCO-stuff datasets, respectively.

Published in Complex & Intelligent Systems

ISSN: 2199-4536 (Print); 2198-6053 (Online)
Publisher: Springer
Country of publisher: Switzerland
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science; Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: https://www.springer.com/journal/40747

About the journal

Abstract

Keywords