This review examines reproducible evaluation of multimodal models. The organizing question is which tests separate perception, grounding, reasoning, calibration, and answer-generation errors. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to allowing a single benchmark score to conceal contamination and modality-specific failure. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in multimodal systems in science, industry, and public services.
- Alcaraz, J. M. L., Bouma, H., & Strodthoff, N. (2025). Enhancing clinical decision support with physiological waveforms — A multimodal benchmark in emergency care. Computers in Biology and Medicine, 192(Pt A), 110196. https://doi.org/10.1016/j.compbiomed.2025.110196 DOI
- Chai, W., & Wang, G. (2022). Deep Vision Multimodal Learning: Methodology, Benchmark, and Trend. Applied Sciences, 12(13), 6588. https://doi.org/10.3390/app12136588 DOI
- Doris, A. C., Grandi, D., Tomich, R., Alam, M. F., Ataei, M., Cheong, H., & Ahmed, F. (2024). DesignQA: A Multimodal Benchmark for Evaluating Large Language Models’ Understanding of Engineering Documentation. Journal of Computing and Information Science in Engineering, 25(2). https://doi.org/10.1115/1.4067333 DOI
- Hu, J., Liu, R., Hong, D., Camero, A., Yao, J., Schneider, M., Kurz, F., Segl, K., & Zhu, X. X. (2023). MDAS: a new multimodal benchmark dataset for remote sensing. Earth system science data, 15(1), 113-131. https://doi.org/10.5194/essd-15-113-2023 DOI
- Li, H., Wang, Z., Wang, J., Wang, Y., Lau, A. K. H., & Qu, H. (2024). CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2409.19058 DOI
- Liu, A. A., Xu, N., Nie, W. Z., Su, Y. T., Wong, Y., & Kankanhalli, M. (2016). Benchmarking a Multimodal and Multiview and Interactive Dataset for Human Action Recognition. IEEE Transactions on Cybernetics, 47(7), 1781-1794. https://doi.org/10.1109/tcyb.2016.2582918 DOI
- Panetta, K., Rajendran, R., Ramesh, A., Rao, S., & Agaian, S. (2021). Tufts Dental Database: A Multimodal Panoramic X-Ray Dataset for Benchmarking Diagnostic Systems. IEEE Journal of Biomedical and Health Informatics, 26(4), 1650-1659. https://doi.org/10.1109/jbhi.2021.3117575 DOI
- Qu, B. Y., Liang, J. J., Wang, Z. Y., Chen, Q., & Suganthan, P. N. (2015). Novel benchmark functions for continuous multimodal optimization with comparative results. Swarm and Evolutionary Computation, 26, 23-34. https://doi.org/10.1016/j.swevo.2015.07.003 DOI
- Ren, H., Sun, L., Guo, J., & Han, C. (2022). A Dataset and Benchmark for Multimodal Biometric Recognition Based on Fingerprint and Finger Vein. IEEE Transactions on Information Forensics and Security, 17, 2030-2043. https://doi.org/10.1109/tifs.2022.3175599 DOI
- Sharma, V., Goyal, P., Lin, K., Thattai, G., Gao, Q., & Sukhatme, G. S. (2022). CH-MARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2208.13626 DOI
- Journal
- Journal of Algorithmic Discovery and Applied AI
- Volume
- 1 (2026)
- Article number
- jadai20260005
- License
- CC BY 4.0