A*-guided deep reinforcement learning for single- and dual-agent navigation : cross-algorithm evaluation and transformer-based policies : a thesis presented in partial fulfilment of the requirements for the degree of Master of Information Science in Computer Science at Massey University, Auckland, New Zealand
Loading...
Date
DOI
Open Access Location
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Massey University
Rights
The author
Abstract
Integrating classical path planning with deep reinforcement learning (DRL) offers a promising approach for navigation in obstacle-dense environments, where purely learned policies often suffer from sparse rewards, inefficient exploration, unstable con vergence, and collision-prone behaviour. This thesis investigates the integration of the A-star (A*) path-planning algorithm with deep and multi-agent reinforcement learning (MARL), with a particular focus on reproducibility, cross-algorithm evaluation, policy architecture design, and the transition from single-agent to dual-agent navigation. The study builds on previous A*-guided DRL research by reproducing and extending its core ideas within the Vectorized Multi-Agent Simulator (VMAS) and Benchmarking Multi-Agent Reinforcement Learning (BenchMARL) frameworks. The thesis makes four main contributions. First, it establishes a reproducible experimental pipeline for evaluating A*-guided reinforcement learning in navigation tasks. Second, it examines how A*-generated waypoint guidance affects different reinforcement learning algorithms, including Multi-Agent Proximal Policy Optimization (MAPPO), Multi-Agent Deep Deterministic Policy Gradient (MADDPG), and the Q value Mixing Network (QMIX). Third, it evaluates the use of structured waypoint information through a Task-Adaptive Waypoint-Aware Transformer configuration. Fourth, it investigates how A*-guided learning and waypoint-aware policies behave when extended from controlled single-agent navigation to dual-agent navigation. In the proposed framework, A* is used to generate waypoint paths that support reinforcement learning through reward shaping and, in some configurations, through explicit waypoint observation features. For the multilayer perceptron (MLP) baseline and Standard Transformer configurations, A*-derived guidance is used during training and removed during evaluation. For the Task-Adaptive Waypoint-Aware Transformer configurations, A*-based reward shaping is disabled during evaluation, but A*-derived waypoint observation features may remain part of the policy input. The baseline re production shows positive descriptive differences in final-window evaluation reward and the collision-penalty proxy in controlled single-agent settings, particularly in compact obstacle-dense environments. However, the results also show that A* guidance does not necessarily improve sample efficiency or convergence speed and introduces additional computational overhead. Extending the framework to MAPPO, MADDPG, and QMIX reveals that the effects of A* guidance are algorithm-dependent, with different learning paradigms responding differently to planning-derived reward signals. The architectural evaluation shows that a Standard Transformer Policy Baseline does not automatically outperform the Fully Connected MLP Policy Baseline. This indicates that simply replacing an MLP with a Transformer is insufficient for improving navigation performance. In contrast, the Task-Adaptive Waypoint-Aware Trans former shows positive descriptive trends in selected final performance metrics, including higher average-reward and final-window evaluation-reward values in some configurations. However, the five-seed statistical analysis does not establish statistically significant improvements after Holm-Bonferroni correction. The dual-agent experiments reveal an important limitation of directly transferring single-agent A*-guided strategies to multi-agent navigation. While A* guidance reduces collisions in single-agent settings, independently generated A* paths can create inter agent path conflicts in dual-agent scenarios. This produces a qualitative shift in the role of A* guidance: its collision-reduction benefits do not automatically transfer to multi-agent settings, although it may still support generalization and task completion in some cases. The results also show that scaling waypoint-aware Transformer policies to dual-agent navigation increases training difficulty and produces algorithm-dependent effects. MADDPG benefits most clearly from the waypoint-aware architecture, QMIX shows mixed effects, and MAPPO exhibits instability under the tested configuration. Overall, this thesis demonstrates that classical planning signals can provide useful structure for reinforcement learning-based navigation, but their benefits depend strongly on the learning algorithm, policy architecture, evaluation setting, and number of agents. The findings clarify both the potential and the limitations of A*-guided DRL and MARL. They suggest that waypoint-aware Transformer policies are a promising direction for structured navigation learning, while also highlighting the need for larger scale statistical validation, improved multi-agent coordination mechanisms, and more robust approaches for resolving conflicts between independently planned paths.
