UTD24 Research Publishing
Journal of Algorithmic Discovery and Applied AI

Evaluating Tool-Using Language Agents in Open-Ended Environments

Read & download PDF
Abstract

This review examines evaluation of tool-using language agents. The organizing question is how benchmarks can distinguish planning, tool selection, argument construction, execution, and recovery failures. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to treating task completion as evidence that intermediate actions were safe and authorized. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in agentic services that act on external tools.

Keywords
language agentstool useevaluationbenchmarkssafety
References
  1. Bran, A. M., Cox, S., Schilter, O., Baldassari, C., White, A. D., & Schwaller, P. (2024). Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5), 525-535. https://doi.org/10.1038/s42256-024-00832-8 DOI
  2. Chen, X., Xiang, J., Lu, S., Liu, Y., He, M., & Shi, D. (2025). Evaluating large language models and agents in healthcare: key challenges in clinical applications. Intelligent Medicine, 5(2), 151-163. https://doi.org/10.1016/j.imed.2025.03.002 DOI
  3. Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., & Li, Y. (2024). Large language models empowered agent-based modeling and simulation: a survey and perspectives. Humanities and Social Sciences Communications, 11(1). https://doi.org/10.1057/s41599-024-03611-3 DOI
  4. Huang, W., Abbeel, P., Pathak, D., & Mordatch, I. (2022). Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2201.07207 DOI
  5. Jiang, F., Peng, Y., Dong, L., Wang, K., Yang, K., Pan, C., Niyato, D., & Dobre, O. A. (2024). Large Language Model Enhanced Multi-Agent Systems for 6G Communications. IEEE Wireless Communications, 31(6), 48-55. https://doi.org/10.1109/mwc.016.2300600 DOI
  6. Lomuscio, A., Qu, H., & Raimondi, F. (2015). MCMAS: an open-source model checker for the verification of multi-agent systems. International Journal on Software Tools for Technology Transfer, 19(1), 9-30. https://doi.org/10.1007/s10009-015-0378-x DOI
  7. Mehandru, N., Miao, B. Y., Almaraz, E. R., Sushil, M., Butte, A. J., & Alaa, A. (2024). Evaluating large language models as agents in the clinic. npj Digital Medicine, 7(1), 84. https://doi.org/10.1038/s41746-024-01083-y DOI
  8. Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A., & Lenert, L. A. (2025). The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review. Journal of the American Medical Informatics Association, 32(6), 1071-1086. https://doi.org/10.1093/jamia/ocaf063 DOI
  9. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2302.04761 DOI
  10. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2303.11366 DOI
Publication details
Journal
Journal of Algorithmic Discovery and Applied AI
Volume
1 (2026)
Article number
jadai20260003
License
CC BY 4.0