Reinforcement Learning with Complex-Valued uRNN Memories Achieves Significant Gains in Partially Observable Environments
A new approach introduces three versions of unitary recurrent networks (uRNNs) with complex-valued memories as drop-in replacements for recurrent PPO architectures in reinforcement learning. These methods show substantial improvements on tasks requiring long-term memory in partially observable settings.
Researchers have proposed three versions of unitary recurrent networks (uRNNs) that leverage complex-valued hidden states for reinforcement learning in partially observable environments. These uRNNs are designed as drop-in replacements for recurrent architectures in Proximal Policy Optimization (PPO).

The key innovation is representing recurrence and hidden states as complex vectors, where the phase component allows for richer, more flexible memory representations. The norm-preserving property of uRNNs supports information propagation over extended sequences, improving the agent's ability to recall long-term dependencies.
Empirical results show that these uRNN-based architectures achieve significant performance gains—up to 2 to 3 times the reward—over baseline recurrent methods in memory-intensive tasks like rocksample and Craftax. The work also explores how phase information in the complex state can be preserved and leveraged for phase-aware policies, drawing analogies to quantum state measurement.
- uRNNs can replace standard recurrent PPO layers with no major architectural changes.
- Significant gains are observed in settings where memory and long-term planning are crucial.
- Phase-preserving policy approaches are experimentally validated.
