500 rub
Journal Highly available systems №3 for 2026 г.
Article in number:
VEXL domain-specific language for automated verification of practical assignments with text-based answers
Type of article: scientific article
DOI: https://doi.org/10.18127/j20729472-202603-03
UDC: 004.434
Authors:

B.S. Ksemidov1

1 Federal Research Center “Computer Science and Control” of the Russian Academy of Sciences (Moscow, Russia)
1 sokboriswork@yandex.com

Abstract:

Problem Statement. In the context of the digitalization of education, automating the assessment of practical assignments that require writing program code or providing open-ended text answers is becoming increasingly relevant. However, existing tools (such as GIFT, IMS QTI, Gherkin) are primarily oriented toward classical test formats and lack capabilities for code verification or semantic text analysis, which places a heavy burden on instructors who must grade solutions manually.

Goal. To develop the VEXL (Verification & Exercise Language) domain-specific language (DSL), which implements a hybrid determi­nistic-stochastic approach for the automated verification of practical assignment solutions using large language models (LLMs).

Results. The VEXL domain-specific language has been developed, based on evaluating solutions using families of positive and negative verifiers. An integration mechanism for large language models based on the "LLM-as-a-judge" paradigm has been implemented for the automated verification of text responses. An experimental evaluation on a dataset of 84 student responses demonstrated the high accuracy of the proposed approach: Precision reached 98.3%, Recall was 96.7%, and the F1-score was 97.5%. Furthermore, a threefold reduction in the average grading time per assignment was shown (from 15 to 5 minutes) compared to manual grading.

Practical Significance. The developed VEXL language and its interpreter software serve as a tool for automating grading and providing feedback in educational courses. Its simple syntax allows instructors without deep programming skills to create complex verification scenarios, facilitating the scaling of educational programs and significantly reducing the routine workload on teaching staff.

Pages: 30-36
For citation

Ksemidov B.S. VEXL domain-specific language for automated verification of practical assignments with text-based answers // Highly Available Systems. 2026. V. 22. № 3. P. 30−36. DOI: https://doi.org/10.18127/j20729472-202603-03

References
  1. Saviny`x I., Zhuravleva E. Formaty` dlya obmena testovy`mi zadaniyami: AIKEN i GIFT. Vestnik Marijskogo gosudarstvennogo universiteta. 2009. № 3. S. 99–100. (in Russian).
  2. Sary`kov E. Sozdanie fajlov s testovy`mi voprosami formata GIFT i ix importirovanie v LMS Moodle. Innovacionny`e obrazovatel`ny`e texnologii v sisteme «Shkola-vuz». 2016. S. 62–70. (in Russian).
  3. Boussakuk M. i dr. Designing and developing e-assessment delivery system under IMS QTI ver. 2.2 specification. International Journal of Emerging Technologies in Learning (iJET). International Journal of Emerging Technology in Learning, 2021. T. 16. № 1. S. 219–233.
  4. Volkov A.A., Otbetkina T.A., Vidmanova A.N. Innovacionny`j metod povy`sheniya e`ffektivnosti obrazovatel`nogo processa s primeneniem texnologii LLM (na primere ChatGPT). Informatika, vy`chislitel`naya texnika i upravlenie. 2023. (in Russian).
  5. Nikitina A.S. II-instrumenty` i bol`shie yazy`kovy`e modeli v obrazovatel`noj praktike. Vestnik molody`x uchyony`x i specialistov Samarskogo universiteta. 2025. № 1 (26). S. 141–148. (in Russian).
  6. Ferreira M. i dr. Acceptance test generation with large language models: An industrial case study. 2025 IEEE/ACM International Conference on Automation of Software Test (AST). IEEE. 2025. S. 1–11.
  7. Karpurapu S. i dr. Comprehensive evaluation and insights into the use of large language models in the automation of behavior-driven development acceptance test formulation. IEEE Access. IEEE. 2024. T. 12. S. 58715–58721.
  8. Quispe J., Cordova H., Wong L. Approach for Generating and Executing Acceptance Tests Based on Computer Vision and GPT-4. 2026 8th International Conference on Software Engineering and Computer Science (CSECS). IEEE. 2026. S. 1–6.
  9. Ameniczkij A.V. i dr. Prichiny`, e`ticheskie problemy` i profilaktika gallyucinacii LLM. INTELLEKT 3. 2024. S. 12. (in Russian).
  10. Mernik M., Heering J., Sloane A.M. When and how to develop domain-specific languages. ACM computing surveys (CSUR). ACM New York, NY, USA, 2005. T. 37. № 4. S. 316–344.
  11. Kaur A. i dr. A literature review on device-to-device data exchange formats for iot applications. Journal of Intelligent Systems and Computing. 2020. T. 1. № 1. S. 1–10.
Date of receipt: 30.07.2026
Approved after review: 11.08.2026
Accepted for publication: 31.08.2026