Preview

Intelligent transport

Advanced search

OMLS-Bench: a multi-level benchmark for engineering LLMS

Abstract

This paper presents the OMLS-Benchmark (Open Multi-Level Skills Benchmark) assessment system, a two–stage framework for the comprehensive assessment of large language models in software engineering tasks. The aim of the proposed approach is to overcome the limitations of existing techniques that either measure narrow subtasks or apply a single level of complexity and do not capture the dynamics of engineering skill and the interactive nature of diagnostic reasoning. The described system covers nine domains (Back-end, Front-end, Mobile, DevOps, Data Analysis, Machine Learning, Big Data, IoT, embedded systems) and five levels of complexity, which allows you to stratify quality by domains and levels. A two–stage procedure is proposed: Stage I standardized tasks with multiple choice and a strict response format, Stage II scenario tasks with step-by-step checklists and an independent judge model. The Tier Accuracy and Domain Accuracy metrics have been formalized, the OPS integral indicator has been introduced; the variables of formula (1) have been disclosed. Artifacts are published for reproducibility: JSON schemas of tasks, Russianlanguage templates of projects and an evaluation script. eval_mc.py with a description of the input/output parameters. Experiments show: heterogeneity of quality between domains; decreased results when switching from tests with fixed options to scenario tasks; detailed diagnostic reports on outstanding checklist items. The OMLS-Bench can serve as a practical tool for comparing LLMs in engineering tasks and as a basis for purposefully fine-tuning models to specific areas. The initial large-scale assessment of ten modern large language models revealed a clear stratification of results by complexity levels and by area: larger-scale models demonstrated high accuracy in widely represented web-oriented areas, while specialized areas (mobile development, embedded systems) showed significantly worse performance. These observations highlight the importance of both the size of the model and the variety of subject data in training. OMLS-Bench provides a reproducible and extensible evaluation tool that can serve as a basis for the development of more reliable and domain-specific assistant engineer models. In the future, it is planned to develop the interactive phase, increase the realism of scenarios and finalize control checklists to bring testing closer to professional practice.

About the Authors

S. A. Yarushev
Federal Research Center «Computer Science and Control» of the Russian Academy of Sciences
Russian Federation

Sergey A. Yarushev, Ph.D. in Engineering, Director of the Research Center

Moscow



A. O. Anurov
Plekhanov Russian University of Economics
Russian Federation

Alexandr O. Anurov, research assistant

Moscow



G. G. Bulgakov
Plekhanov Russian University of Economics
Russian Federation

Gennadii G. Bulgakov, postgraduate student

Moscow



References

1. Ануров А. О., Булгаков Г. Г., Ярушев С. А. OMLS-Bench [Электронный ресурс]. – URL: https://github.com/Blgkff/OMLS-BENCH (дата обращения: 14.06.2025).

2. D. Hendrycks et al., “Measuring Massive Multitask Language Understanding,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021. [Online]. Available: https://arxiv.org/abs/2009.03300

3. T. B. Brown et al., “Language Models are Few-Shot Learners,” Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, pp. 1877–1901, 2020. [Online]. Available: https://arxiv.org/abs/2005.14165

4. A. Wang et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in *Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. 9th Int. Jt. Conf. Nat. Lang. Process. (EMNLP-IJCNLP)*, pp. 353–361, 2019. [Online]. Available: https://arxiv.org/abs/1804.07461

5. A. Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” in Proc. 3rd Workshop Eval. Compar. NLP Syst., 2020. [Online]. Available: https://arxiv.org/abs/1905.00537

6. S. Lu et al., “CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2021. [Online]. Available: https://arxiv.org/abs/2102.04664

7. M. Chen et al., “Evaluating Large Language Models Trained on Code,” arXiv preprint arXiv:2107.03374, 2021. [Online]. Available: https://arxiv.org/abs/2107.03374

8. M. Mitchell and D. C. Krakauer, “The Debate Over Understanding in AI’s Large Language Models,” arXiv preprint arXiv:2210.13966, 2022. [Online]. Available: https://arxiv.org/abs/2210.13966

9. D. Hendrycks et al., “Aligning AI With Shared Human Values,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021. [Online]. Available: https://arxiv.org/abs/2008.02275

10. M. Suzgun et al., “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them,” in Find. Assoc. Comput. Linguist.: ACL, pp. 4670–4685, 2022. [Online]. Available: https://arxiv.org/abs/2210.09261

11. M. Kazemi et al., “BIG-Bench Extra Hard,” arXiv preprint arXiv:2502.19187, 2025. [Online]. Available: https://arxiv.org/abs/2502.19187

12. H. Li et al., “CMMLU: Measuring Massive Multitask Language Understanding in Chinese,” in Find. Assoc. Comput. Linguist.: ACL, pp. 6543–6558, 2024. [Online]. Available: https://aclanthology.org/2024.findings-acl.671

13. Y. Wang et al., “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark,” arXiv preprint arXiv:2406.01574, 2024. [Online]. Available: https://arxiv.org/abs/2406.01574

14. L. Austin et al., “Program Synthesis with Large Language Models,” in NeurIPS Workshop Mach. Learn. Syst., 2021. [Online]. Available: https://arxiv.org/abs/2108.07732

15. K. Cobbe et al., “Training Verifiers to Solve Math Word Problems,” Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 34, 2021. [Online]. Available: https://arxiv.org/abs/2110.14168

16. A. Radford et al., “Language Models are Unsupervised Multitask Learners,” OpenAI Tech. Rep., 2019. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf


Review

For citations:


Yarushev S.A., Anurov A.O., Bulgakov G.G. OMLS-Bench: a multi-level benchmark for engineering LLMS. Intelligent transport. 2025;(3(35)):96-112. (In Russ.)

Views: 28

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 3033-6007 (Online)