ZTE Communications ›› 2026, Vol. 24 ›› Issue (2): 71-82.DOI: 10.12142/ZTECOM.202602009

• Industry-Academia Co-Research • Previous Articles     Next Articles

ACTap: Integrating Attention and Convolution for Network Modality Recognition

Ling Zihan1, Zhang Tianwei1, Ning Yuwei1, Cao Yang1(), Shen Can2   

  1. 1.Huazhong University of Science and Technology, Wuhan 430074, China
    2.ZTE Corporation, Shenzhen 518057, China
  • Received:2025-08-14 Online:2026-06-25 Published:2026-06-16
  • About author:Ling Zihan received his BE degree from the School of Microelectronics and Communication Engineering, Chongqing University, China in 2023. He is currently pursuing the MS degree at the School of Electronic Information and Communications, Huazhong University of Science and Technology, China. His current research interests include real-time VR video transmission, network modal analysis, and adaptive bitrate algorithms.
    Zhang Tianwei received his MS degree in wireless communication systems from the University of Sheffield, UK in 2020. He is currently pursuing the PhD degree at the School of Electronic Information and Communications, Huazhong University of Science and Technology, China. His current research interests include real-time video transmission, video processing, and adaptive streaming.
    Ning Yuwei is a third-year undergraduate student at Huazhong University of Science and Technology, China, supervised by Prof. Cao Yang. His research interests include deep learning, time series processing, and computer vision.
    Cao Yang (ycao@hust.edu.cn) is currently an associate professor at the School of Electronic Information and Communications, Huazhong University of Science and Technology, China. From 2011 to 2013, he was a visiting scholar at the School of Electrical, Computer, and Energy Engineering, Arizona State University, USA. His research interests include immersive video transmission and edge/distributed computing. He has coauthored 50 papers in refereed IEEE journals and conferences. He received the CHINACOM Best Paper Award in 2010 and the Microsoft Research Fellowship in 2011.
    Shen Can received his PhD degree in engineering in 1997 and has since been engaged in the research and development of audio and video technologies at ZTE Corporation. He has received the Second Prize of the National Science and Technology Progress Award, published 20 papers, and holds over 70 patents.
  • Supported by:
    This work was supported in part by ZTE Industry?University?Institute Cooperation Funds under Grant No. IA20230728008 and the National Natural Science Foundation of China (NSFC) under Grant No. 62271224.

Abstract:

With the sustained growth of live video streaming, the demand for high-quality video services for mobile users across diverse network environments is increasing rapidly. In this paper, we define the network fluctuation characteristics in different environments as network modality. To comprehensively investigate network modality across different environments, we construct a network modality dataset by collecting multi-dimensional network metrics from various real-world scenarios and suggest that network modality exhibits separability. Therefore, network modality recognition, which aims to distinguish the scenarios where a user is located based on network modality sequences, is feasible and can be formulated as a multivariate time series (MTS) classification problem. To address this problem, we propose a novel neural network (NN)-based classification model called ACTap. Specifically, the model first integrates a two-stage attention (TSA) mechanism and a convolutional neural network (CNN) to extract features from network modality sequences. Then, it filters out noisy feature representations to learn discriminative class prototypes, and finally recognizes network modality based on the distance between their feature representations and class prototypes. Experimental results validate the separability of network modality and show that ACTap outperforms four benchmark models in terms of classification accuracy on the network modality dataset.

Key words: network modality, multivariate time series classification, feature fusion, class prototype learning