500 rub
Journal Neurocomputers №4 for 2026 г.
Article in number:
Architecture and algorithm of a deterministic pipeline for structuring JSON and XML data using LLM
Type of article: scientific article
DOI: https://doi.org/10.18127/j19998554-202604-04
UDC: 004.85:311.2
Authors:

N.N. Oltyan1
1 Financial University under the Government of the Russian Federation (Moscow, Russia)

1 nikitaoltyan@mail.ru

Abstract:

This paper presents a theoretical study that formalizes the concept and algorithm of a deterministic pipeline for transforming un-structured text into JSON/XML structures using large language models (LLMs). The proposed architecture integrates grammar-constrained decoding, schema validation, strict typing, and repair mechanisms, thereby ensuring syntactic and semantic correctness, reproducibility, and predictability of the results. Correctness invariants have been defined, and an algorithm that eliminates non-determinism in generation has been formulated. The proposed solution provides a foundation for reliable and controllable application of LLMs in data structuring systems.

Pages: 47-53
For citation

Oltyan N.N. Architecture and algorithm of a deterministic pipeline for structuring JSON and XML data using LLM // Neurocomputers. 2026. V. 28. № 4. P. 47–53. DOI: https://doi.org/10.18127/j19998554-202604-04

References
  1. Oltyan N.N. Evolyutsiya metodov izvlecheniya i strukturirovaniya dannykh iz teksta v JSON i XML. Nejrokomp'yutery. 2025. № 6. S. 37–49. (in Russian)
  2. Beurer-Kellner L., Fischer M., Vechev M. Guiding LLMs the right way: Fast, non-invasive constrained generation. arXiv preprint. arXiv:2403.06988. 2024.
  3. Geng S., Cooper H., Moskal M. et al. JSONSchemaBench: A rigorous benchmark of structured outputs for language models. arXiv preprint. arXiv:2501.10868. 2025.
  4. Geng S., Josifoski M., Peyrard M. et al. Grammar-constrained decoding for structured NLP tasks without finetuning. arXiv preprint. arXiv:2305.13971. 2023.
  5. Korn F., Saha B., Srivastava D. et al. On repairing structural problems in semi-structured data. Proceedings of the VLDB Endowment. 2013. V. 6. № 9. P. 601–612.
  6. Shen Z., Wang D.Y.-B., Mishra S.S. et al. SLOT: Structuring the output of large language models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025. P. 472–491.
  7. Viotti J.C., Mior M.J. Blaze: Compiling JSON schema for 10x faster validation. arXiv preprint. arXiv:2503.02770. 2025.
  8. Yao Y., Mao S., Zhang N. et al. Schema-aware reference as prompt improves data-efficient knowledge graph construction. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2023. P. 911–921.
Date of receipt: 09.02.2026
Approved after review: 27.02.2026
Accepted for publication: 29.06.2026