The rapid growth of urban air mobility (UAM) presents a fundamental challenge: how do we safely and efficiently manage dozens — potentially hundreds — of autonomous eVTOL aircraft arriving at and departing from vertiports in dense urban environments? Traditional air traffic management approaches, designed for human-controlled aircraft in open airspace, are ill-suited for the high-density, low-altitude operations that characterize UAM. Vertiports, the designated takeoff and landing zones for eVTOLs, become critical bottlenecks where scheduling conflicts, resource allocation, and safety constraints converge into a complex decision-making problem.

Deep Reinforcement Learning (DRL) offers a compelling framework for addressing this challenge. By modeling vertiport operations as a Markov Decision Process (MDP), we can train intelligent agents to make real-time decisions about aircraft sequencing, gate assignment, and departure scheduling. In our research, we employed Proximal Policy Optimization (PPO), a policy gradient method that has demonstrated remarkable stability and sample efficiency across a range of continuous control tasks. The agent observes the current state of the vertiport — including queued aircraft, weather conditions, pad availability, and fuel constraints — and learns a policy that maximizes throughput while maintaining safety margins. Unlike rule-based systems, the DRL agent adapts to novel situations and can handle the stochastic nature of arrival patterns.

Our multi-agent extension places independent PPO agents at each vertiport within a network, enabling decentralized coordination. Using a shared reward signal that penalizes delays and safety violations, the agents learn to cooperate without explicit communication protocols. Simulation results across a modeled urban corridor with six vertiports showed a 34% improvement in throughput compared to first-come-first-served scheduling, with a 61% reduction in safety margin violations. The agents discovered emergent strategies such as holding patterns and dynamic pad reallocation that were not explicitly programmed, demonstrating the creative problem-solving capacity of DRL in safety-critical systems.

Looking ahead, integrating DRL-based scheduling into the broader U-space framework will be essential for scaling urban air mobility. U-space, Europe’s vision for drone traffic management, envisions a highly automated system where service providers manage airspace dynamically. Our research suggests that DRL agents could serve as the decision engine within U-space service providers, handling the real-time optimization that human operators cannot perform at the required speed. Future work will focus on transfer learning across vertiport topologies, robustness to adversarial scenarios, and the critical step of bridging the sim-to-real gap through hardware-in-the-loop testing.