OMLS-Bench: a multi-level benchmark for engineering LLMS
Abstract
This paper presents the OMLS-Benchmark (Open Multi-Level Skills Benchmark) assessment system, a two–stage framework for the comprehensive assessment of large language models in software engineering tasks. The aim of the proposed approach is to overcome the limitations of existing techniques that either measure narrow subtasks or apply a single level of complexity and do not capture the dynamics of engineering skill and the interactive nature of diagnostic reasoning. The described system covers nine domains (Back-end, Front-end, Mobile, DevOps, Data Analysis, Machine Learning, Big Data, IoT, embedded systems) and five levels of complexity, which allows you to stratify quality by domains and levels. A two–stage procedure is proposed: Stage I standardized tasks with multiple choice and a strict response format, Stage II scenario tasks with step-by-step checklists and an independent judge model. The Tier Accuracy and Domain Accuracy metrics have been formalized, the OPS integral indicator has been introduced; the variables of formula (1) have been disclosed. Artifacts are published for reproducibility: JSON schemas of tasks, Russianlanguage templates of projects and an evaluation script. eval_mc.py with a description of the input/output parameters. Experiments show: heterogeneity of quality between domains; decreased results when switching from tests with fixed options to scenario tasks; detailed diagnostic reports on outstanding checklist items. The OMLS-Bench can serve as a practical tool for comparing LLMs in engineering tasks and as a basis for purposefully fine-tuning models to specific areas. The initial large-scale assessment of ten modern large language models revealed a clear stratification of results by complexity levels and by area: larger-scale models demonstrated high accuracy in widely represented web-oriented areas, while specialized areas (mobile development, embedded systems) showed significantly worse performance. These observations highlight the importance of both the size of the model and the variety of subject data in training. OMLS-Bench provides a reproducible and extensible evaluation tool that can serve as a basis for the development of more reliable and domain-specific assistant engineer models. In the future, it is planned to develop the interactive phase, increase the realism of scenarios and finalize control checklists to bring testing closer to professional practice.
About the Authors
S. A. YarushevRussian Federation
Sergey A. Yarushev, Ph.D. in Engineering, Director of the Research Center
Moscow
A. O. Anurov
Russian Federation
Alexandr O. Anurov, research assistant
Moscow
G. G. Bulgakov
Russian Federation
Gennadii G. Bulgakov, postgraduate student
Moscow
References
1. Ануров А. О., Булгаков Г. Г., Ярушев С. А. OMLS-Bench [Электронный ресурс]. – URL: https://github.com/Blgkff/OMLS-BENCH (дата обращения: 14.06.2025).
2. D. Hendrycks et al., “Measuring Massive Multitask Language Understanding,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021. [Online]. Available: https://arxiv.org/abs/2009.03300
3. T. B. Brown et al., “Language Models are Few-Shot Learners,” Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, pp. 1877–1901, 2020. [Online]. Available: https://arxiv.org/abs/2005.14165
4. A. Wang et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in *Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. 9th Int. Jt. Conf. Nat. Lang. Process. (EMNLP-IJCNLP)*, pp. 353–361, 2019. [Online]. Available: https://arxiv.org/abs/1804.07461
5. A. Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” in Proc. 3rd Workshop Eval. Compar. NLP Syst., 2020. [Online]. Available: https://arxiv.org/abs/1905.00537
6. S. Lu et al., “CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2021. [Online]. Available: https://arxiv.org/abs/2102.04664
7. M. Chen et al., “Evaluating Large Language Models Trained on Code,” arXiv preprint arXiv:2107.03374, 2021. [Online]. Available: https://arxiv.org/abs/2107.03374
8. M. Mitchell and D. C. Krakauer, “The Debate Over Understanding in AI’s Large Language Models,” arXiv preprint arXiv:2210.13966, 2022. [Online]. Available: https://arxiv.org/abs/2210.13966
9. D. Hendrycks et al., “Aligning AI With Shared Human Values,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021. [Online]. Available: https://arxiv.org/abs/2008.02275
10. M. Suzgun et al., “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them,” in Find. Assoc. Comput. Linguist.: ACL, pp. 4670–4685, 2022. [Online]. Available: https://arxiv.org/abs/2210.09261
11. M. Kazemi et al., “BIG-Bench Extra Hard,” arXiv preprint arXiv:2502.19187, 2025. [Online]. Available: https://arxiv.org/abs/2502.19187
12. H. Li et al., “CMMLU: Measuring Massive Multitask Language Understanding in Chinese,” in Find. Assoc. Comput. Linguist.: ACL, pp. 6543–6558, 2024. [Online]. Available: https://aclanthology.org/2024.findings-acl.671
13. Y. Wang et al., “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark,” arXiv preprint arXiv:2406.01574, 2024. [Online]. Available: https://arxiv.org/abs/2406.01574
14. L. Austin et al., “Program Synthesis with Large Language Models,” in NeurIPS Workshop Mach. Learn. Syst., 2021. [Online]. Available: https://arxiv.org/abs/2108.07732
15. K. Cobbe et al., “Training Verifiers to Solve Math Word Problems,” Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 34, 2021. [Online]. Available: https://arxiv.org/abs/2110.14168
16. A. Radford et al., “Language Models are Unsupervised Multitask Learners,” OpenAI Tech. Rep., 2019. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
Review
For citations:
Yarushev S.A., Anurov A.O., Bulgakov G.G. OMLS-Bench: a multi-level benchmark for engineering LLMS. Intelligent transport. 2025;(3(35)):96-112. (In Russ.)
JATS XML




