ZTE Communications ›› 2026, Vol. 24 ›› Issue (2): 26-32.DOI: 10.12142/ZTECOM.202602004
• Special Topic • Previous Articles Next Articles
Chen Yu1, Li Fan2, Wu Jie3, Gao Weipeng1(
), Ouyang Ye1
Received:2026-01-15
Online:2026-06-16
Published:2026-06-16
About author:Chen Yu received his ME degree in computer science from Southeast University, China in 2017. He currently serves as an AI algorithm engineer at the AI Lab of AsiaInfo Technologies. He holds 3 patents. His research interests include LLMs, reinforcement learning for agents, RAG, multi-agent systems, and the application of AI in telecommunications and enterprise intelligence.Chen Yu, Li Fan, Wu Jie, Gao Weipeng, Ouyang Ye. Training Optimization for Complex Reasoning Tasks in ACN: Dynamic Batch- Aware Advantage Weighting for Agentic RAG[J]. ZTE Communications, 2026, 24(2): 26-32.
Add to citation manager EndNote|Ris|BibTeX
URL: https://zte.magtechjournal.com/EN/10.12142/ZTECOM.202602004
| Base Model | Method | HotpotQA | NQ | 2Wiki | Avg | Improvement over GRPO |
|---|---|---|---|---|---|---|
| Qwen2.5-3B | GRPO | 0.23 | 0.49 | 0.15 | 0.290 | - |
| DAPO | 0.25 | 0.49 | 0.16 | 0.300 | +3.4% | |
| DB-AW (Full) | 0.31 | 0.53 | 0.19 | 0.340 | +17.2% | |
| Qwen2.5-7B | GRPO | 0.25 | 0.52 | 0.17 | 0.313 | - |
| DAPO | 0.27 | 0.52 | 0.18 | 0.323 | +3.2% | |
| DB-AW (Full) | 0.33 | 0.57 | 0.21 | 0.370 | +18.2% | |
| LLaMA3.2-3B | GRPO | 0.21 | 0.45 | 0.13 | 0.263 | - |
| DAPO | 0.22 | 0.45 | 0.14 | 0.270 | +2.7% | |
| DB-AW (Full) | 0.26 | 0.49 | 0.16 | 0.303 | +15.2% |
Table 1 Main results: cross-model generalization (2 000 samples, validation accuracy)
| Base Model | Method | HotpotQA | NQ | 2Wiki | Avg | Improvement over GRPO |
|---|---|---|---|---|---|---|
| Qwen2.5-3B | GRPO | 0.23 | 0.49 | 0.15 | 0.290 | - |
| DAPO | 0.25 | 0.49 | 0.16 | 0.300 | +3.4% | |
| DB-AW (Full) | 0.31 | 0.53 | 0.19 | 0.340 | +17.2% | |
| Qwen2.5-7B | GRPO | 0.25 | 0.52 | 0.17 | 0.313 | - |
| DAPO | 0.27 | 0.52 | 0.18 | 0.323 | +3.2% | |
| DB-AW (Full) | 0.33 | 0.57 | 0.21 | 0.370 | +18.2% | |
| LLaMA3.2-3B | GRPO | 0.21 | 0.45 | 0.13 | 0.263 | - |
| DAPO | 0.22 | 0.45 | 0.14 | 0.270 | +2.7% | |
| DB-AW (Full) | 0.26 | 0.49 | 0.16 | 0.303 | +15.2% |
| Configuration | HotpotQA | NQ | 2Wiki | Avg |
|---|---|---|---|---|
| GRPO (baseline) | 0.23 | 0.49 | 0.15 | 0.290 |
| DB-AW (Filtering Only) | 0.26 | 0.50 | 0.16 | 0.307 (+5.9%) |
| DB-AW (Weighting Only) | 0.29 | 0.52 | 0.18 | 0.330 (+13.8%) |
| DB-AW (Full) | 0.31 | 0.53 | 0.19 | 0.340 (+17.2%) |
Table 2 Ablation study on Qwen2.5-3B-Instruct (2 000 samples, validation accuracy)
| Configuration | HotpotQA | NQ | 2Wiki | Avg |
|---|---|---|---|---|
| GRPO (baseline) | 0.23 | 0.49 | 0.15 | 0.290 |
| DB-AW (Filtering Only) | 0.26 | 0.50 | 0.16 | 0.307 (+5.9%) |
| DB-AW (Weighting Only) | 0.29 | 0.52 | 0.18 | 0.330 (+13.8%) |
| DB-AW (Full) | 0.31 | 0.53 | 0.19 | 0.340 (+17.2%) |
| [1] | Mobile China. AI-Agent communication network white paper [R]. 2025 |
| [2] | Singh A, Ehtesham A, Kumar S, et al. Agentic retrieval-augmented generation: a survey on agentic RAG [PP/OL]. arXiv (2025-04-01) [2026-02-06]. |
| [3] | Yang Z L, Qi P, Zhang S Z, et al. HotpotQA: a dataset for diverse, explainable multi-hop question answering [PP/OL]. arXiv (2018-09-25) [2026-02-06]. |
| [4] | Ho X, Nguyen A K, Sugawara S, et al. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps [C]//The 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, 2020: 6609–6625. DOI: 10.18653/v1/2020.coling-main.580 |
| [5] | Trivedi H, Balasubramanian N, Khot T, et al. MuSiQue: multihop questions via single-hop question composition [J]. Transactions of the association for computational linguistics, 2022, 10: 539–554. DOI: 10.1162/tacl_a_00475 |
| [6] | Kwiatkowski T, Palomaki J, Redfield O, et al. Natural questions: a benchmark for question answering research [J]. Transactions of the association for computational linguistics, 2019, 7: 453–466. DOI: 10.1162/tacl_a_00276 |
| [7] | Yao S Y, Zhao J, Yu D, et al. ReAct: synergizing reasoning and acting in language models [PP/OL]. arXiv (2022-10-06) [2026-02-06]. |
| [8] | Asai A, Wu Z Q, Wang Y Z, et al. Self-RAG: learning to retrieve, generate, and critique through self- reflection [PP/OL]. arXiv (2023-10-17) [2026-02-06]. |
| [9] | Shinn N, Cassano F, Berman E, et al. Reflexion: language agents with verbal reinforcement learning [PP/OL]. arXiv (2023-03-20) [2026-02-06]. |
| [10] | Li X X, Dong G T, Jin J J, et al. Search-o1: agentic search-enhanced large reasoning models [PP/OL]. arXiv (2025-01-09) [2026-02-06]. |
| [11] | Shao Z H, Wang P Y, Zhu Q H, et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models [PP/OL]. arXiv (2024-02-05) [2026-02-06]. |
| [12] | Jin B W, Zeng H S, Yue Z R, et al. Search-R1: training LLMs to reason and leverage search engines with reinforcement learning [PP/OL]. arXiv (2025-03-12) [2026-02-06]. |
| [13] | Song H T, Jiang J H, Min Y Q, et al. R1-searcher: a novel two-stage outcome-based RL approach for search-enhanced large reasoning models [PP/OL]. arXiv (2025-03-12) [2026-02-06]. |
| [14] | Song Y, Ramaneti K, Sheikh Z, et al. Agent data protocol: unifying datasets for diverse, effective fine-tuning of LLM agents [PP/OL]. arXiv (2025-10-28) [2026-02-06]. |
| [15] | Eldeeb E, Alves H. Offline multi-agent reinforcement learning for 6G communications: fundamentals, applications and future directions [PP/OL]. arXiv (2026-01-01) [2026-02-06]. |
| [16] | Christianos F, Papoudakis G, Rahman M A, et al. Scaling multi-agent reinforcement learning with selective parameter sharing [C]//The 38th International Conference on Machine Learning. PMLR, 2021: 1989–1998 |
| [17] | Shrivastava A, Gupta A, Girshick R. Training region-based object detectors with online hard example mining [C]//Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2016: 761–769. DOI: 10.1109/cvpr.2016.89 |
| [18] | Soviany P, Ionescu R T, Rota P, et al. Curriculum learning: a survey [J]. International journal of computer vision, 2022, 130(6): 1526–1565. DOI: 10.1007/s11263-022-01611-x |
| [19] | Yu Q Y, Zhang Z, Zhu R F, et al. DAPO: an open-source LLM reinforcement learning system at scale [PP/OL]. arXiv (2025-03-18) [2026-02-06]. |
| [20] | Ahmadian A, Cremer C, Gallé M, et al. Back to basics: revisiting reinforce style optimization for learning from human feedback in LLMs [PP/OL]. arXiv (2024-02-26) [2026-02-06]. |
| [21] | Hu J, Liu J K, Xu H T, et al. REINFORCE++: an efficient RLHF algorithm with robustness to both prompt and reward models [PP/OL]. arXiv (2025-11-10) [2026-02-06]. |
| [22] | Rafailov R, Sharma A, Mitchell E, et al. Direct preference optimization: your language model is secretly a reward model [PP/OL]. arXiv (2024-07-29) [2026-02-06]. |
| [23] | Meng Y, Xia M Z, Chen D Q. SimPO: simple preference optimization with a reference-free reward [PP/OL]. arXiv (2024-05-23) [2026-02-06]. |
| [24] | Schulman J, Wolskie F, Dhariwal P, et al. Proximal policy optimization algorithms [PP/OL]. arXiv (2017-07-20) [2026-02-06]. |
| [25] | Cui G Q, Yuan L F, Wang Z F, et al. Process reinforcement through implicit rewards [PP/OL]. arXiv (2025-02-03) [2026-02-06]. |
| [26] | Zhang E C, Yan X G, Lin W, et al. Learning like humans: advancing LLM reasoning capabilities via adaptive difficulty curriculum learning and expert-guided self-reformulation [C]//The Conference on Empirical Methods in Natural Language Processing. ACL, 2025: 6630–6644. DOI: 10.18653/v1/2025.emnlp-main.336 |
| [1] | Zhang Xiaotian, Xiao Han, Wang Dan, Huang Zhenglei, Xu Changqiao. Toward AI-Agent-Native 6G Networks: A Survey on Protocols, Multimodal Coordination, and ISCC-Driven Dynamic Networking [J]. ZTE Communications, 2026, 24(2): 16-25. |
| [2] | FAN Kaiqing, YAO Yuze, GAO Ning, LI Xiao, JIN Shi. RIS Enabled Simultaneous Transmission and Key Generation with PPO: Exploring Security Boundary of RIS Phase Shift [J]. ZTE Communications, 2025, 23(1): 11-17. |
| [3] | WEI Zhiqing, ZHANG Yongji, JI Danna, LI Chenfei. Sensing and Communication Integrated Fast Neighbor Discovery for UAV Networks [J]. ZTE Communications, 2024, 22(3): 69-82. |
| [4] | SHEN Jiahao, JIANG Ke, TAN Xiaoyang. Boundary Data Augmentation for Offline Reinforcement Learning [J]. ZTE Communications, 2023, 21(3): 29-36. |
| [5] | REN Min, XU Renyu, ZHU Ting. Double Deep Q-Network Decoder Based on EEG Brain-Computer Interface [J]. ZTE Communications, 2023, 21(3): 3-10. |
| [6] | FENG Bingyi, FENG Mingxiao, WANG Minrui, ZHOU Wengang, LI Houqiang. Multi-Agent Hierarchical Graph Attention Reinforcement Learning for Grid-Aware Energy Management [J]. ZTE Communications, 2023, 21(3): 11-21. |
| [7] | YU Junpeng, CHEN Yiyu. A Practical Reinforcement Learning Framework for Automatic Radar Detection [J]. ZTE Communications, 2023, 21(3): 22-28. |
| [8] | YOU Qian, XU Qian, YANG Xin, ZHANG Tao, CHEN Ming. RIS-Assisted UAV-D2D Communications Exploiting Deep Reinforcement Learning [J]. ZTE Communications, 2023, 21(2): 61-69. |
| [9] | JIA Haonan, HE Zhenqing, TAN Wanlong, RUI Hua, LIN Wei. Distributed Multi-Cell Multi-User MISO Downlink Beamforming via Deep Reinforcement Learning [J]. ZTE Communications, 2022, 20(4): 69-77. |
| [10] | JI Hong, ZHANG Tianxiang, ZHANG Kai, WANG Wanyuan, WU Weiwei. Efficient Network Slicing with Dynamic Resource Allocation [J]. ZTE Communications, 2021, 19(1): 11-19. |
| [11] | Stephen ANOKYE, Mohammed SEID, SUN Guolin. A Survey on Machine Learning Based Proactive Caching [J]. ZTE Communications, 2019, 17(4): 46-55. |
| [12] | DONG Shaokang, CHEN Jiarui, LIU Yong, BAO Tianyi, GAO Yang. Reinforcement Learning from Algorithm Model to Industry Innovation: A Foundation Stone of Future Artificial Intelligence [J]. ZTE Communications, 2019, 17(3): 31-41. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||