Exploring Agent Behavior and Performance in Shortened Simulations via Multiagent Reinforcement Learning

Authors

  • Stanislav Safranek Department of Information Technologies University of Hradec Kralove Hradec Kralove, Czech Republic
  • Brian Kirk New Mexico Institute of Mining and Technology New Mexico, United States of America

DOI:

https://doi.org/10.11113/ijic.v16n1-2.683

Keywords:

Agent-based Simulation, Artificial Neural Networks (ANN), Multiagent Reinforcement Learning, Pathfinding Optimization, Reinforcement Learning Stability, Statistical Performance Analysis

Abstract

This study explores the potential of Multiagent Reinforcement Learning (MARL) for autonomous navigation in a discrete two-dimensional environment. We design and implement an agent that learns an optimal path through a 5×5 grid via repeated interactions and a reward-based mechanism. Over 500 training episodes, we examine the convergence speed of the learned policy, the stability of agent behavior, and the success rate in reaching the goal. Our approach combines artificial neural networks with a multiagent framework, enabling decentralized decision making and scalable adaptation. We discuss critical factors affecting learning stability, including reward function design and network architecture, and outline avenues for extending the methodology to more complex, real-time tasks. The findings demonstrate MARL’s promise in solving navigation problems efficiently and provide concrete recommendations for tuning training parameters and network structures to enhance performance and robustness.

References

Bellman, R., & Dreyfus, S. (2010). Dynamic programming (Princeton Landmarks in Mathematics ed., with a new introduction). Princeton University Press.

Bușoniu, L., Babuška, R., & De Schutter, B. (2008). A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2), 156–172. https://doi.org/10.1109/TSMCC.2007.913919.

Chopra, A. K., Artikis, A., Bentahar, J., Colombetti, M., Dignum, F., Fornara, N., Jones, A. J. I., Singh, M. P., & Yolum, P. (2013). Research directions in agent communication. ACM Transactions on Intelligent Systems and Technology, 4(2), 1–23. https://doi.org/10.1145/2438653.2438655.

Farhan, M., Gohre, B., & Junprung, E. (2020). Reinforcement learning in analogic simulation models: A guiding example using Pathmind. In Proceedings of the 2020 Winter Simulation Conference (WSC). https://doi.org/10.1109/WSC48552.2020.9383916.

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

Murphy, K. P. (2013). Machine learning: A probabilistic perspective. MIT Press.

Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML) (pp. 807–814).

Olfati-Saber, R., Fax, J. A., & Murray, R. M. (2007). Consensus and cooperation in networked multi-agent systems. Proceedings of the IEEE, 95(1), 215–233. https://doi.org/10.1109/JPROC.2006.887293.

Onken, D., Nurbekyan, L., Li, X., Fung, S. W., Osher, S., & Ruthotto, L. (2023). A neural network approach for high-dimensional optimal control applied to multiagent path finding. IEEE Transactions on Control Systems Technology, 31(1), 235–251. https://doi.org/10.1109/TCST.2022.3172872.

Pan, Y., Ji, W., Lam, H.-K., & Cao, L. (2024). An improved predefined-time adaptive neural control approach for nonlinear multiagent systems. IEEE Transactions on Automation Science and Engineering, 1–10. https://doi.org/10.1109/TASE.2023.3324397.

Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. https://doi.org/10.1038/323533a0.

Russell, S. J., Norvig, P., & Chang, M.-W. (2022). Artificial intelligence: A modern approach (4th ed.). Pearson.

Sandholm, T., & Lesser, V. R. (1995). Issues in automated negotiation and electronic commerce: Extending the contract net framework. In Proceedings of the International Conference on Multi-Agent Systems (ICMAS) (pp. 12–14).

Schulman, J., Chen, X., & Abbeel, P. (2017). Equivalence between policy gradients and soft Q-learning (Version 4). arXiv. https://doi.org/10.48550/arXiv.1704.06440

Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1), 1929–1958.

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Wooldridge, M. J. (2009). An introduction to multiagent systems (2nd ed.). John Wiley & Sons.

Wu, W., Liu, J., Li, F., Zhang, Y., & Hu, Z. (2023). Prescribed settling time adaptive neural network consensus control of multiagent systems with unknown time-varying input dead-zone. Mathematics, 11(4), 988. https://doi.org/10.3390/math11040988,

Zhang, C., & Wu, J. (2021). Global adaptive consensus control for multiagent systems with predefined accuracy. Complexity, 2021, 1–19. https://doi.org/10.1155/2021/3396482

Zhang, Y., Niu, B., Zhang, J.-M., & Wang, X.-M. (2021). Predefined-time adaptive neural-network-based consensus tracking control for nonlinear multiagent systems with zero tracking error. IEEE Access, 9, 132205–132214. https://doi.org/10.1109/ACCESS.2021.3115118.

Downloads

Published

2026-07-28

How to Cite

Exploring Agent Behavior and Performance in Shortened Simulations via Multiagent Reinforcement Learning. (2026). International Journal of Innovative Computing, 16(1-2), 123-131. https://doi.org/10.11113/ijic.v16n1-2.683