This review examines evaluation of tool-using language agents. The organizing question is how benchmarks can distinguish planning, tool selection, argument construction, execution, and recovery failures. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to treating task completion as evidence that intermediate actions were safe and authorized. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in agentic services that act on external tools.
- Bran, A. M., Cox, S., Schilter, O., Baldassari, C., White, A. D., & Schwaller, P. (2024). Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5), 525-535. https://doi.org/10.1038/s42256-024-00832-8 DOI
- Chen, X., Xiang, J., Lu, S., Liu, Y., He, M., & Shi, D. (2025). Evaluating large language models and agents in healthcare: key challenges in clinical applications. Intelligent Medicine, 5(2), 151-163. https://doi.org/10.1016/j.imed.2025.03.002 DOI
- Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., & Li, Y. (2024). Large language models empowered agent-based modeling and simulation: a survey and perspectives. Humanities and Social Sciences Communications, 11(1). https://doi.org/10.1057/s41599-024-03611-3 DOI
- Huang, W., Abbeel, P., Pathak, D., & Mordatch, I. (2022). Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2201.07207 DOI
- Jiang, F., Peng, Y., Dong, L., Wang, K., Yang, K., Pan, C., Niyato, D., & Dobre, O. A. (2024). Large Language Model Enhanced Multi-Agent Systems for 6G Communications. IEEE Wireless Communications, 31(6), 48-55. https://doi.org/10.1109/mwc.016.2300600 DOI
- Lomuscio, A., Qu, H., & Raimondi, F. (2015). MCMAS: an open-source model checker for the verification of multi-agent systems. International Journal on Software Tools for Technology Transfer, 19(1), 9-30. https://doi.org/10.1007/s10009-015-0378-x DOI
- Mehandru, N., Miao, B. Y., Almaraz, E. R., Sushil, M., Butte, A. J., & Alaa, A. (2024). Evaluating large language models as agents in the clinic. npj Digital Medicine, 7(1), 84. https://doi.org/10.1038/s41746-024-01083-y DOI
- Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A., & Lenert, L. A. (2025). The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review. Journal of the American Medical Informatics Association, 32(6), 1071-1086. https://doi.org/10.1093/jamia/ocaf063 DOI
- Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2302.04761 DOI
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2303.11366 DOI
- Journal
- Journal of Algorithmic Discovery and Applied AI
- Volume
- 1 (2026)
- Article number
- jadai20260003
- License
- CC BY 4.0