ZTE Communications ›› 2026, Vol. 24 ›› Issue (2): 71-82.DOI: 10.12142/ZTECOM.202602009
• Industry-Academia Co-Research • Previous Articles Next Articles
Ling Zihan1, Zhang Tianwei1, Ning Yuwei1, Cao Yang1(
), Shen Can2
Received:2025-08-14
Online:2026-06-25
Published:2026-06-16
About author:Ling Zihan received his BE degree from the School of Microelectronics and Communication Engineering, Chongqing University, China in 2023. He is currently pursuing the MS degree at the School of Electronic Information and Communications, Huazhong University of Science and Technology, China. His current research interests include real-time VR video transmission, network modal analysis, and adaptive bitrate algorithms.Supported by:Ling Zihan, Zhang Tianwei, Ning Yuwei, Cao Yang, Shen Can. ACTap: Integrating Attention and Convolution for Network Modality Recognition[J]. ZTE Communications, 2026, 24(2): 71-82.
Add to citation manager EndNote|Ris|BibTeX
URL: https://zte.magtechjournal.com/EN/10.12142/ZTECOM.202602009
| Dataset | Sampling Time Interval/s | Multi‑Dimensional Modality Information | Multiple Scenes | Full Pipe | Protocol |
|---|---|---|---|---|---|
| Ours | 0.1 | √ | √ | √ | UDP/TCP |
| Raca et al.[ | 1 | √ | √ | × | TCP |
| Yang et al.[ | unknown | × | √ | × | UDP/TCP |
| Van et al.[ | 1 | × | √ | × | TCP |
| Pan et al.[ | 1 | √ | × | × | UDP/TCP |
Kousias et al.[ | millisecond-level | √ | √ | × | TCP |
Table 1 Comparison of network datasets
| Dataset | Sampling Time Interval/s | Multi‑Dimensional Modality Information | Multiple Scenes | Full Pipe | Protocol |
|---|---|---|---|---|---|
| Ours | 0.1 | √ | √ | √ | UDP/TCP |
| Raca et al.[ | 1 | √ | √ | × | TCP |
| Yang et al.[ | unknown | × | √ | × | UDP/TCP |
| Van et al.[ | 1 | × | √ | × | TCP |
| Pan et al.[ | 1 | √ | × | × | UDP/TCP |
Kousias et al.[ | millisecond-level | √ | √ | × | TCP |
| Symbol | Description |
|---|---|
| The network modality sequence | |
| The corresponding class label of the network modality sequence | |
| The length of the network modality sequence | |
| The feature representation of the network modality sequence | |
| The length of the feature representation | |
| The | |
| The length of segment | |
| The | |
| The length of the feature vector | |
| The input of the TSA layer | |
The outputs of the cross-time stage and cross-dimension stage, respectively | |
The extracted features of the TSA layer and CNN layer, respectively | |
The set of all representations and the set of selected representations, respectively | |
| The attention weight for class | |
| The prototype of class | |
The distance between network modality sequence class prototype | |
The probability of network modality sequence class |
Table 2 Summary of notations
| Symbol | Description |
|---|---|
| The network modality sequence | |
| The corresponding class label of the network modality sequence | |
| The length of the network modality sequence | |
| The feature representation of the network modality sequence | |
| The length of the feature representation | |
| The | |
| The length of segment | |
| The | |
| The length of the feature vector | |
| The input of the TSA layer | |
The outputs of the cross-time stage and cross-dimension stage, respectively | |
The extracted features of the TSA layer and CNN layer, respectively | |
The set of all representations and the set of selected representations, respectively | |
| The attention weight for class | |
| The prototype of class | |
The distance between network modality sequence class prototype | |
The probability of network modality sequence class |
| Model | Accuracy |
|---|---|
| ACTap | 0.902 3 |
| TapNet | 0.671 6 |
| CNNTap | 0.850 0 |
| AttentionTap | 0.813 3 |
| ACMLP | 0.847 9 |
Table 3 Accuracy performance of models
| Model | Accuracy |
|---|---|
| ACTap | 0.902 3 |
| TapNet | 0.671 6 |
| CNNTap | 0.850 0 |
| AttentionTap | 0.813 3 |
| ACMLP | 0.847 9 |
| [1] | Carlucci G, De Cicco L, Holmer S, et al. Analysis and design of the google congestion control for web real-time communication (WebRTC) [C]//Proceedings of the 7th International Conference on Multimedia Systems. ACM, 2016: 1–12. DOI: 10.1145/2910017.2910605 |
| [2] | Mao H Z, Netravali R, Alizadeh M. Neural adaptive video streaming with Pensieve [C]//Proceedings of the Conference of the ACM Special Interest Group on Data Communication. ACM, 2017: 197–210. DOI: 10.1145/3098822.3098843 |
| [3] | Zhang H H, Zhou A F, Lu J M, et al. OnRL: improving mobile video telephony via online reinforcement learning [C]//Proceedings of the 26th Annual International Conference on Mobile Computing and Networking. ACM, 2020: 1–14. DOI: 10.1145/3372224.3419186 |
| [4] | Raca D, Leahy D, Sreenan C J, et al. Beyond throughput, the next generation: a 5G dataset with channel and context metrics [C]//Proceedings of the 11th ACM Multimedia Systems Conference. ACM, 2020: 303–308. DOI: 10.1145/3339825.3394938 |
| [5] | Yang L, Yuan M X, Wang W, et al. Apps on the move: a fine-grained analysis of usage behavior of mobile apps [C]//Proceedings of the 35th Annual IEEE International Conference on Computer Communications. IEEE, 2016: 1–9. DOI: 10.1109/INFOCOM.2016.7524464 |
| [6] | van der Hooft J, Petrangeli S, Wauters T, et al. HTTP/2-based adaptive streaming of HEVC video over 4G/LTE networks [J]. IEEE communications letters, 2016, 20(11): 2177–2180. DOI: 10.1109/LCOMM.2016.2601087 |
| [7] | Pan Y Y, Li R H, Xu C R. The first 5G-LTE comparative study in extreme mobility [J]. Proceedings of the ACM on measurement and analysis of computing systems, 2022, 6(1): 1–22. DOI: 10.1145/3508040 |
| [8] | Kousias K, Rajiullah M, Caso G, et al. A large-scale dataset of 4G, NB-IoT, and 5G non-standalone network measurements [J]. IEEE communications magazine, 2024, 62(5): 44–49. DOI: 10.1109/MCOM.011.2200707 |
| [9] | Lu J G, Zheng Q F. Ultra-lightweight face animation method for ultra-low bitrate video conferencing [J]. ZTE communications, 2023, 21(1): 64–71. DOI: 10.12142/ZTECOM.202301008 |
| [10] | Xie L, Zhang X G, Huang C, et al. Markov based rate adaption approach for live streaming over HTTP/2 [J]. ZTE communications, 2018, 16(2): 37–41. DOI: 10.3969/j.issn.1673-5188.2018.02.007 |
| [11] | HoloWan APP [EB/OL]. [2024-08-01]. |
| [12] | ZHANG Y H, YAN J C. Crossformer: transformer utilizing cross-dimension dependency for multivariate time series forecasting [C/OL]//The Eleventh International Conference on Learning Representations. ICLR, 2022. |
| [13] | Zhang X C, Gao Y F, Lin J, et al. TapNet: multivariate time series classification with attentional prototypical network [J]. Proceedings of the AAAI conference on artificial intelligence, 2020, 34(4): 6845–6852. DOI: 10.1609/aaai.v34i04.6165 |
| [14] | Li G Z, Choi B, Xu J L, et al. ShapeNet: a shapelet-neural network approach for multivariate time series classification [J]. Proceedings of the AAAI conference on artificial intelligence, 2021, 35(9): 8375–8383. DOI: 10.1609/aaai.v35i9.17018 |
| [15] | Gao G, Gao Q T, Yang X, et al. A reinforcement learning-informed pattern mining framework for multivariate time series classification [C]//31st International Joint Conference on Artificial Intelligence. IJCAI, 2022: 2994–3000. DOI: 10.24963/ijcai.2022/415 |
| [16] | Khan M, Wang H Z, Ngueilbaye A, et al. End-to-end multivariate time series classification via hybrid deep learning architectures [J]. Personal and ubiquitous computing, 2023, 27(2): 177–191. DOI: 10.1007/s00779-020-01447-7 |
| [17] | Zerveas G, Jayaraman S, Patel D, et al. A transformer-based framework for multivariate time series representation learning [C]//Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. ACM, 2021: 2114–2124. DOI: 10.1145/3447548.3467401 |
| [18] | Zuo R D, Li G Z, Choi B, et al. SVP-T: a shape-level variable-position transformer for multivariate time series classification [J]. Proceedings of the AAAI conference on artificial intelligence, 2023, 37(9): 11497–11505. DOI: 10.1609/aaai.v37i9.26359 |
| [19] | Gao N Z, Yu Y F, Hua X H. A content-aware bitrate selection method using multi-step prediction for 360-degree video streaming [J]. ZTE communications, 2022, 20(4): 96–109. DOI: 10.12142/ZTECOM.202204012 |
| [20] | Li J X, Xu Y T, Cao Y, et al. Utility-driven joint caching and bitrate allocation for real-time immersive videos [J]. IEEE journal of selected topics in signal processing, 2023, 17(5): 1106–1118. DOI: 10.1109/JSTSP.2023.3295597 |
| [21] | Xu Y T, Du J H, Wang J H, et al. Panonut360: a head and eye tracking dataset for panoramic video [C]//Proceedings of the 15th ACM Multimedia Systems Conference. ACM, 2024: 319–325. DOI: 10.1145/3625468.3652176 |
| [22] | Du D Z, Su B, Wei Z W. Preformer: predictive transformer with multi-scale segment-wise correlations for long-term time series forecasting [C]//2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1–5. DOI: 10.1109/ICASSP49357.2023.10096881 |
| [23] | Zhou H Y, Zhang S H, Peng J Q, et al. Informer: beyond efficient transformer for long sequence time-series forecasting [J]. Proceedings of the AAAI conference on artificial intelligence, 2021, 35(12): 11106–11115. DOI: 10.1609/aaai.v35i12.17325 |
| [24] | Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16×16 words: transformers for image recognition at scale [PP/OL]. arXiv (2021-06-03) [2024-08-01]. |
| [25] | Vaswani A, Shazeer N, Parmar N, et al. 2017. Attention is all you need [C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS, 2017: 6000–6010. DOI:10.48550/arXiv.1706.03762 |
| [26] | Ioffe S, Szegedy C. Batch normalization: accelerating deep network training by reducing internal covariate shift [C]//International conference on machine learning. PMLR, 2015: 448-456. DOI: 10.48550/arXiv.1502.03167 |
| [27] | Xu B, Wang N Y, Chen T Q, et al. Empirical evaluation of rectified activations in convolutional network [PP/OL]. V2. arXiv (2015-11-27). [2024-06-07]. |
| [28] | MacQueen J. Some methods for classification and analysis of multivariate observations [C]//Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics. University of California Press, 1967: 281–297 |
| [29] | Kingma D P, Ba J. Adam: a method for stochastic optimization [PP/OL]. V9. arXiv (2017-01-30) [2024-06-07]. |
| [30] | Aliyun [EB/OL]. [2024-05-19]. |
| [31] | Cardwell N, Cheng Y, Gunn C S, et al. BBR: congestion-based congestion control [J]. Communications of the ACM, 2017, 60(2): 58–66. DOI: 10.1145/3009824 |
| [32] | Pytorch [EB/OL]. [2024-05-19]. |
| [33] | Bagnall A, Dau H A, Lines J, et al. The UEA multivariate time series classification archive, 2018 [PP/OL]. V1. arXiv (2018-10-31)[2024-04-21]. |
| [34] | van der Maaten L, Hinton G. Visualizing data using t-SNE [J/OL]. Journal of machine learning research, 2008, 9(86): 2579-2605. |
| [1] | Zheng Wangguandong, Lu Ping, Deng Fangwei, Huang Shijun, Xia Siyu. Steel Surface Anomaly Detection Using 3D Depth and 2D RGB Features [J]. ZTE Communications, 2026, 24(1): 81-87. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||