| [1] |
Mobile China. AI-Agent communication network white paper [R]. 2025
|
| [2] |
Singh A, Ehtesham A, Kumar S, et al. Agentic retrieval-augmented generation: a survey on agentic RAG [PP/OL]. arXiv (2025-04-01) [2026-02-06].
|
| [3] |
Yang Z L, Qi P, Zhang S Z, et al. HotpotQA: a dataset for diverse, explainable multi-hop question answering [PP/OL]. arXiv (2018-09-25) [2026-02-06].
|
| [4] |
Ho X, Nguyen A K, Sugawara S, et al. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps [C]//The 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, 2020: 6609–6625. DOI: 10.18653/v1/2020.coling-main.580
|
| [5] |
Trivedi H, Balasubramanian N, Khot T, et al. MuSiQue: multihop questions via single-hop question composition [J]. Transactions of the association for computational linguistics, 2022, 10: 539–554. DOI: 10.1162/tacl_a_00475
|
| [6] |
Kwiatkowski T, Palomaki J, Redfield O, et al. Natural questions: a benchmark for question answering research [J]. Transactions of the association for computational linguistics, 2019, 7: 453–466. DOI: 10.1162/tacl_a_00276
|
| [7] |
Yao S Y, Zhao J, Yu D, et al. ReAct: synergizing reasoning and acting in language models [PP/OL]. arXiv (2022-10-06) [2026-02-06].
|
| [8] |
Asai A, Wu Z Q, Wang Y Z, et al. Self-RAG: learning to retrieve, generate, and critique through self- reflection [PP/OL]. arXiv (2023-10-17) [2026-02-06].
|
| [9] |
Shinn N, Cassano F, Berman E, et al. Reflexion: language agents with verbal reinforcement learning [PP/OL]. arXiv (2023-03-20) [2026-02-06].
|
| [10] |
Li X X, Dong G T, Jin J J, et al. Search-o1: agentic search-enhanced large reasoning models [PP/OL]. arXiv (2025-01-09) [2026-02-06].
|
| [11] |
Shao Z H, Wang P Y, Zhu Q H, et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models [PP/OL]. arXiv (2024-02-05) [2026-02-06].
|
| [12] |
Jin B W, Zeng H S, Yue Z R, et al. Search-R1: training LLMs to reason and leverage search engines with reinforcement learning [PP/OL]. arXiv (2025-03-12) [2026-02-06].
|
| [13] |
Song H T, Jiang J H, Min Y Q, et al. R1-searcher: a novel two-stage outcome-based RL approach for search-enhanced large reasoning models [PP/OL]. arXiv (2025-03-12) [2026-02-06].
|
| [14] |
Song Y, Ramaneti K, Sheikh Z, et al. Agent data protocol: unifying datasets for diverse, effective fine-tuning of LLM agents [PP/OL]. arXiv (2025-10-28) [2026-02-06].
|
| [15] |
Eldeeb E, Alves H. Offline multi-agent reinforcement learning for 6G communications: fundamentals, applications and future directions [PP/OL]. arXiv (2026-01-01) [2026-02-06].
|
| [16] |
Christianos F, Papoudakis G, Rahman M A, et al. Scaling multi-agent reinforcement learning with selective parameter sharing [C]//The 38th International Conference on Machine Learning. PMLR, 2021: 1989–1998
|
| [17] |
Shrivastava A, Gupta A, Girshick R. Training region-based object detectors with online hard example mining [C]//Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2016: 761–769. DOI: 10.1109/cvpr.2016.89
|
| [18] |
Soviany P, Ionescu R T, Rota P, et al. Curriculum learning: a survey [J]. International journal of computer vision, 2022, 130(6): 1526–1565. DOI: 10.1007/s11263-022-01611-x
|
| [19] |
Yu Q Y, Zhang Z, Zhu R F, et al. DAPO: an open-source LLM reinforcement learning system at scale [PP/OL]. arXiv (2025-03-18) [2026-02-06].
|
| [20] |
Ahmadian A, Cremer C, Gallé M, et al. Back to basics: revisiting reinforce style optimization for learning from human feedback in LLMs [PP/OL]. arXiv (2024-02-26) [2026-02-06].
|
| [21] |
Hu J, Liu J K, Xu H T, et al. REINFORCE++: an efficient RLHF algorithm with robustness to both prompt and reward models [PP/OL]. arXiv (2025-11-10) [2026-02-06].
|
| [22] |
Rafailov R, Sharma A, Mitchell E, et al. Direct preference optimization: your language model is secretly a reward model [PP/OL]. arXiv (2024-07-29) [2026-02-06].
|
| [23] |
Meng Y, Xia M Z, Chen D Q. SimPO: simple preference optimization with a reference-free reward [PP/OL]. arXiv (2024-05-23) [2026-02-06].
|
| [24] |
Schulman J, Wolskie F, Dhariwal P, et al. Proximal policy optimization algorithms [PP/OL]. arXiv (2017-07-20) [2026-02-06].
|
| [25] |
Cui G Q, Yuan L F, Wang Z F, et al. Process reinforcement through implicit rewards [PP/OL]. arXiv (2025-02-03) [2026-02-06].
|
| [26] |
Zhang E C, Yan X G, Lin W, et al. Learning like humans: advancing LLM reasoning capabilities via adaptive difficulty curriculum learning and expert-guided self-reformulation [C]//The Conference on Empirical Methods in Natural Language Processing. ACL, 2025: 6630–6644. DOI: 10.18653/v1/2025.emnlp-main.336
|