Tamed Warping Network for High-Resolution Semantic Video Segmentation

Songyuan Li; Junyi Feng; Xi Li

doi:10.3390/app131810102

Applied Sciences (Sep 2023)

Tamed Warping Network for High-Resolution Semantic Video Segmentation

Songyuan Li,
Junyi Feng,
Xi Li

Affiliations

Songyuan Li: College of Computer Science, Zhejiang University, Hangzhou 310027, China
Junyi Feng: College of Computer Science, Zhejiang University, Hangzhou 310027, China
Xi Li: College of Computer Science, Zhejiang University, Hangzhou 310027, China

DOI: https://doi.org/10.3390/app131810102
Journal volume & issue: Vol. 13, no. 18
p. 10102

Abstract

Read online

Recent approaches for fast semantic video segmentation have reduced redundancy by warping feature maps across adjacent frames, greatly speeding up the inference phase. However, the accuracy drops seriously owing to the errors incurred by warping. In this paper, we propose a novel framework and design a simple and effective correction stage after warping. Specifically, we build a non-key-frame CNN, fusing warped context features with current spatial details. Based on the feature fusion, our context feature rectification (CFR) module learns the model’s difference from a per-frame model to correct the warped features. Furthermore, our residual-guided attention (RGA) module utilizes the residual maps in the compressed domain to help CRF focus on error-prone regions. Results on Cityscapes show that the accuracy significantly increases from 67.3% to 71.6%, and the speed edges down from 65.5 FPS to 61.8 FPS at a resolution of 1024×2048. For non-rigid categories, e.g., “human” and “object”, the improvements are even higher than 18 percentage points.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords