UTD24 Research Publishing
Journal of Algorithmic Discovery and Applied AI

Small Language Models for Resource-Constrained Applied AI

Read & download PDF
Abstract

This review examines small language models for resource-constrained applied artificial intelligence. The organizing question is which compression and adaptation choices preserve utility, privacy, and calibration under tight compute budgets. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to reporting parameter reduction without measuring end-to-end quality and energy. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in edge, mobile, and embedded AI deployment.

Keywords
small language modelsefficient AIquantizationdistillationedge inference
References
  1. Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., Schirrmeister, R. T., & Hutter, F. (2025). Accurate predictions on small data with a tabular foundation model. Nature, 637(8045), 319-326. https://doi.org/10.1038/s41586-024-08328-6 DOI
  2. Kulkarni, R. C. (2026). Energy-Efficient AI Inference at the Edge: Optimizing Semiconductor Hardware for Small Language Models. International Journal of AI BigData Computational and Management Studies, 7, 202-218. https://doi.org/10.63282/3050-9416.ijaibdcms-v7i1p132 DOI
  3. Lin, Z., Qu, G., Chen, Q., Chen, X., Chen, Z., & Huang, K. (2025). Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities. IEEE Communications Magazine, 63(9), 52-59. https://doi.org/10.1109/mcom.001.2400764 DOI
  4. Menghani, G. (2023). Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better. ACM Computing Surveys, 55(12), 1-37. https://doi.org/10.1145/3578938 DOI
  5. Pujari, M., Goel, A., & Pakina, A. K. (2024). Efficient TinyML Architectures for On-Device Small Language Models: Privacy-Preserving Inference at the Edge. International Journal Science and Technology, 3(3), 67-75. https://doi.org/10.56127/ijst.v3i3.1958 DOI
  6. Ramazzotti, D., Caravagna, G., Loohuis, L. O., Graudenzi, A., Korsunsky, I., Mauri, G., Antoniotti, M., & Mishra, B. (2015). CAPRI: efficient inference of cancer progression models from cross-sectional data. Bioinformatics, 31(18), 3016-3026. https://doi.org/10.1093/bioinformatics/btv296 DOI
  7. Sledzieski, S., Kshirsagar, M., Baek, M., Dodhia, R., Ferres, J. L., & Berger, B. (2024). Democratizing protein language models with parameter-efficient fine-tuning. Proceedings of the National Academy of Sciences, 121(26), e2405840121. https://doi.org/10.1073/pnas.2405840121 DOI
  8. Wang, G., Chen, Y., An, P., Hong, H., Hu, J., & Huang, T. (2023). UAV-YOLOv8: A Small-Object-Detection Model Based on Improved YOLOv8 for UAV Aerial Photography Scenarios. Sensors, 23(16), 7190. https://doi.org/10.3390/s23167190 DOI
  9. Wang, R., Gao, Z., Zhang, L., Yue, S., & Gao, Z. (2025). Empowering large language models to edge intelligence: A survey of edge efficient LLMs and techniques. Computer Science Review, 57, 100755. https://doi.org/10.1016/j.cosrev.2025.100755 DOI
  10. Wang, T., Ren, Z., Ding, Y., Fang, Z., Sun, Z., MacDonald, M. L., Sweet, R. A., Wang, J., & Chen, W. (2016). FastGGM: An Efficient Algorithm for the Inference of Gaussian Graphical Model in Biological Networks. PLoS Computational Biology, 12(2), e1004755. https://doi.org/10.1371/journal.pcbi.1004755 DOI
Publication details
Journal
Journal of Algorithmic Discovery and Applied AI
Volume
1 (2026)
Article number
jadai20260002
License
CC BY 4.0