500 rub
Journal Highly available systems №3 for 2026 г.
Article in number:
Research of the architecture of the distributed computing system MPI-GRID based on real-time operating systems
Type of article: scientific article
DOI: https://doi.org/10.18127/j20729472-202603-07
UDC: 004.252
Authors:

Yu.P. Titov1, N.S. Andrejanov2

1 FIC «Informatics and Management» RAS (Moscow, Russia)
1 Moscow Aviation Institute (Moscow, Russia)
2 Federal State Unitary Enterprise Scientific and Production Enterprise "Gamma" (Moscow, Russia)
1 kalengul@mail.ru, 2 nik00789030127799@bk.ru

Abstract:

Problem statement. The modern stage of computing technology development is characterized by the rapid growth in the number of autonomous real-time devices, including robotic complexes, unmanned aerial vehicles, industrial controllers, and wearable devices. These systems collectively possess significant aggregate computing potential. However, existing distributed computing architectures GRID and MPI have fundamental limitations when applied in such heterogeneous real-time environments. GRID systems, while supporting heterogeneous and geographically distributed resources, do not provide the necessary level of parallelism for interacting processes and lack deterministic timing guarantees. MPI clusters, conversely, offer powerful parallelism mechanisms but require hardware homogeneity and predictable low-latency networks, which are unattainable in heterogeneous environments with variable network delays. This creates a critical gap between available computational resources and the ability to utilize them effectively for applications requiring both real-time coordination and parallel computations.

Objective. The primary objective is the development of a conceptual architecture for a hybrid distributed computing system, designated MPI-GRID, built upon real-time operating systems. This architecture aims to unite heterogeneous autonomous devices into a unified computing environment capable of simultaneously supporting coordinated real-time actions and resource-intensive parallel computations characteristic of GRID technologies.

Methods. The proposed architecture comprises a global scheduler on a dedicated server and node agents deployed on each computing node running FreeRTOS. The global scheduler employs dynamic node classification based on a novel MPI-fitness score integrating Worst-Case Execution Time (WCET) of MPI operations, network latency, interrupt jitter, and deadline compliance. Nodes are categorized into three classes (MPI-class, hybrid class, GRID-class) to optimize task placement according to determinism and performance requirements. Communication is organized through pre-scheduled synchronization epochs, during which local schedulers suspend computational tasks and prioritize network exchanges. The node agent is implemented as a set of cooperating FreeRTOS tasks with fixed priorities (Task_Sync priority 6, Task_Comm priority 5, Task_Executor priority 4, Task_Monitor priority 3) communicating through queues and event groups. Time synchronization across nodes is achieved using the Precision Time Protocol (PTP) with sub-10 μs accuracy. For optimal task distribution among nodes considering their dynamically changing characteristics, the global scheduler incorporates an Ant Colony Optimization (ACO) modification adapted for heterogeneous real-time environments.

Results. A three-level classification of computing nodes (MPI-class, hybrid class, GRID-class) has been implemented, ensuring optimal task distribution considering their requirements for determinism and performance. The node agent, implemented as a set of cooperating tasks with fixed priorities, introduces overhead of less than 6% CPU and about 10 KB RAM, confirming the possibility of its use even on resource-constrained devices. Time synchronization based on the PTP protocol with accuracy not worse than 10 μs is ensured. A fault recovery mechanism has been developed including node failure detection (15-second timeout), task redistribution, and state recovery. Markov model analysis showed that with at least 3 MPI-class nodes and cold reservation, the probability of failure-free operation over 24 hours exceeds 0.999, and the system availability coefficient reaches 0.9999. Application of the ACO method for task distribution improved computational resource utilization efficiency by 15-25% compared to classical greedy algorithms while maintaining acceptable planning time.

Practical significance. The proposed concept enables creation of fault-tolerant distributed computing systems from heterogeneous real-time devices, ensuring predictable execution of critical operations and efficient use of idle resources. The achieved availability coefficient of 0.9999 (less than 52 minutes downtime annually) and sub-500 ms recovery time satisfy stringent requirements for mission-critical systems in transport dispatching, electrical grid management, and industrial automation. The system architecture naturally supports graceful degradation upon node failures through dynamic task redistribution without operator intervention. Potential applications include swarm robotics for search and rescue operations, distributed sensor networks for environmental monitoring, real-time data processing in smart manufacturing, and coordinated control systems for autonomous vehicle fleets. The open-source FreeRTOS-based implementation facilitates adoption across a wide range of embedded platforms.

Pages: 71-84
For citation

Titov Yu.P., Andrejanov N.S. Research of the architecture of the distributed computing system MPI-GRID based on real-time operating systems // Highly Available Systems. 2026. V. 22. № 3. P. 71−84. DOI: https://doi.org/10.18127/j20729472-202603-07

References
  1. Stankovic J.A. Research Directions for the Internet of Things. IEEE Internet of Things Journal. 2014. V. 1. № 1. P. 3–9. DOI: 10.1109/JIOT.2014.2312291
  2. Cisco Systems. Cisco Annual Internet Report (2018–2023). 2020.
  3. Kumar V., Grama A., Gupta A., Karypis G. Introduction to Parallel Computing: Design and Analysis of Algorithms. Benjamin-Cummings, 1994.
  4. Anderson D.P., Cobb J., Korpela E. SETI@home: An Experiment in Public-Resource Computing. Communications of the ACM. 2002. V. 45. № 11. P. 56–61. DOI: 10.1145/581571.581573
  5. Abdelzaher T., Stankovic J.A., Lu C. Feedback Performance Control in Software Services. IEEE Control Systems Magazine. 2003. V. 23. № 3. P. 74–90. DOI: 10.1109/MCS.2003.1200251
  6. Tindell K., Burns A., Wellings A.J. Analysis of Hard Real-Time Communications. Real-Time Systems. 1995. V. 9. № 2. P. 147–171. DOI: 10.1007/BF01088856
  7. Stankovic J.A. When Sensor and Actuator Networks Cover the World. ETRI Journal. 2008. V. 30. № 5. P. 627–633. DOI: 10.4218/etrij.08.1308.0181
  8. Kim J.E., Abdelzaher T., Sha L. Sporadic Decision-Centric Data Scheduling with Normally-Off Sensors. IEEE Real-Time Systems Symposium. 2016. P. 135–146. DOI: 10.1109/RTSS.2016.022
  9. Eidson J., Lee K. IEEE 1588 Standard for a Precision Clock Synchronization Protocol for Networked Measurement and Control Systems. IEEE Sensors. 2002. DOI: 10.1109/ICSENS.2002.1037261
  10. Macenski S., Foote T., Gerkey B. Robot Operating System 2: Design, Architecture, and Uses In The Wild. Science Robotics. 2022. V. 7. № 66. DOI: 10.1126/scirobotics.abm6074
  11. Avizienis A., Laprie J.-C., Randell B., Landwehr C. Basic Concepts and Taxonomy of Dependable and Secure Computing. IEEE Transactions on Dependable and Secure Computing. 2004. V. 1. № 1. P. 11–33. DOI: 10.1109/TDSC.2004.2
  12. Burns A., Wellings A. Real-Time Systems and Programming Languages. Addison-Wesley, 2009.
  13. Foster I., Kesselman C. The Grid 2: Blueprint for a New Computing Infrastructure. Morgan Kaufmann, 2004.
  14. Buyya R. et al. Grid Computing: Making the Global Infrastructure a Reality. Wiley, 2003. DOI: 10.1002/0470867167
  15. Evdokimov A.A., Voronoy S.M. Upravlenie parallel'nymi zadaniyami v GRID [Parallel Job Management in GRID]. Informatika i komp'yuternye tekhnologii. 2009. (in Russian).
  16. Dongarra J., Meuer H.W., Strohmaier E. Top500 Supercomputer Sites: Performance Development for the Last Decade. Concurrency and Computation: Practice and Experience. 2005. DOI: 10.1002/cpe.934
  17. Gropp W., Lusk E., Skjellum A. Using MPI: Portable Parallel Programming with the Message-Passing Interface. MIT Press, 1999.
  18. Snir M. et al. MPI: The Complete Reference. MIT Press, 1998.
  19. Kalia A., Kaminsky M., Anderson T. Design Guidelines for High Performance RDMA Systems. USENIX Annual Technical Conference. 2016. P. 437–450.
  20. Sha L. et al. Real-Time Systems: Scheduling, Analysis, and Verification. Wiley, 2004.
  21. Liu C.L., Layland J.W. Scheduling Algorithms for Multiprogramming in a Hard-Real-Time Environment. Journal of the ACM. 1973. V. 20. № 1. P. 46–61. DOI: 10.1145/321738.321743
  22. Mills D.L. Network Time Protocol (NTP): A Brief Tutorial. RFC 1305. 1992.
  23. Wilhelm R., Engblom J., Ermedahl A. The Worst-Case Execution-Time Problem – Overview of Methods and Survey of Tools. ACM Transactions on Embedded Computing Systems. 2008. V. 7. № 3. Art. 36. DOI: 10.1145/1347375.1347389
  24. Thakur R., Rabenseifner R., Gropp W. Optimization of Collective Communication Operations in MPICH. International Journal of High Performance Computing Applications. 2005. V. 19. № 1. P. 49–66. DOI: 10.1177/1094342005051521
  25. Paxson V. End-to-End Internet Packet Dynamics. IEEE/ACM Transactions on Networking. 1999. V. 7. № 3. P. 277–292. DOI: 10.1109/90.779192
  26. Ramamritham K., Stankovic J.A. Scheduling Algorithms and Operating Systems Support for Real-Time Systems. Proceedings of the IEEE. 1994. V. 82. № 1. P. 55–67. DOI: 10.1109/5.259426
  27. Burdonov I.B., Kosachev A.S., Ponomarenko V.N. Operatsionnye sistemy real'nogo vremeni [Real-Time Operating Systems]. Preprint Instituta sistemnogo programmirovaniya RAN. 2006. № 14. (in Russian).
  28. Fundamentals of GRID Technologies [Electronic resource]. URL: http://book.itep.ru/4/7/grid.htm (in Russian).
  29. FreeRTOS Reference Manual. Real Time Engineers Ltd., 2023.
  30. Bernstein D. Containers and Cloud: From LXC to Docker and Beyond. IEEE Cloud Computing. 2014. V. 1. № 3. P. 20–26. DOI: 10.1109/MCC.2014.51
  31. IEEE Standard for a Precision Clock Synchronization Protocol for Networked Measurement and Control Systems. IEEE Std 1588-2008. 2008. DOI: 10.1109/IEEESTD.2008.4579760
Date of receipt: 10.08.2026
Approved after review: 21.08.2026
Accepted for publication: 31.08.2026