Logo PTI Logo FedCSIS

Position Papers of the 21st Conference on Computer Science and Intelligence Systems

Annals of Computer Science and Information Systems, Volume 48

StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning

DOI: http://dx.doi.org/10.15439/2026F6447

Citation: Ivo Nowak (). StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems. ACSIS, Vol. 48, pages 143–148.

Full text

Abstract. Reinforcement learning is typically treated as a uniform optimization process that does not explicitly exploit global structure, leading to slow and inefficient information propagation. In contrast, dynamic programming achieves high efficiency by leveraging structured propagation of value information. In this paper, we argue that such structure is implicitly encoded in its learning dynamics. Focusing on distributional reinforcement learning, we show that the temporal evolution of return variance reveals a propagation process across the state space. We introduce a temporal learning indicator t*(s) that captures when a state is most strongly affected by propagated updates. Empirically, this signal induces an ordering over states consistent with dynamic programming--style propagation, where information spreads from terminal regions outward. Based on this insight, we propose StructRL, a framework that exploits this emergent structure to guide exploration and sampling. In a gridworld setting, this leads to substantially faster and more stable convergence compared to a standard baseline. Our results support a new perspective: reinforcement learning can be understood as a structured propagation process, and learning dynamics themselves provide actionable information to control it. This opens a pathway toward more sample- and energy-efficient reinforcement learning.

References

  1. M. G. Bellemare, W. Dabney, and R. Munos, “A Distributional Perspective on Reinforcement Learning,” in Proceedings of the 34th International Conference on Machine Learning, PMLR, vol. 70, pp. 449–458, 2017. https://dx.doi.org/10.48550/arXiv.1707.06887.
  2. W. Dabney, G. Ostrovski, D. Silver, and R. Munos, “Implicit Quantile Networks for Distributional Reinforcement Learning,” in Proceedings of the 35th International Conference on Machine Learning, PMLR, vol. 80, pp. 1096–1105, 2018. https://dx.doi.org/10.48550/arXiv.1806.06923.
  3. T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized Experience Replay,” in International Conference on Learning Representations, 2016. https://dx.doi.org/10.48550/arXiv.1511.05952.
  4. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018. ISBN: 978-0-262-039246.
  5. D. P. Bertsekas, Dynamic Programming and Optimal Control, Athena Scientific, 1996. ISBN: 978-1-886529-04-0.
  6. V. Aziz, I. Nowak, and E. M. T. Hendrix, “Distributional Reinforcement Learning via the Cramér Distance,” arXiv preprint https://arxiv.org/abs/2605.08104, 2026. https://dx.doi.org/10.48550/arXiv.2605.08104.
  7. D. Silver and R. S. Sutton, “Welcome to the Era of Experience,” Google DeepMind, 2025. Available: https: //storage.googleapis.com/deepmind-media/Era-of-Experience/ The%20Era%20of%20Experience%20Paper.pdf.
  8. A. W. Moore and C. G. Atkeson, “Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time,” Machine Learning, vol. 13, no. 1, pp. 103–130, 1993. https://dx.doi.org/10.1007/BF00993104.
  9. M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying Count-Based Exploration and Intrinsic Motivation,” in Advances in Neural Information Processing Systems, vol. 29, pp. 1471–1479, 2016. https://dx.doi.org/10.48550/arXiv.1606.01868.
  10. D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-Driven Exploration by Self-Supervised Prediction,” in Proceedings of the 34th International Conference on Machine Learning, PMLR, vol. 70, pp. 2778–2787, 2017. https://dx.doi.org/10.48550/arXiv.1705.05363.
  11. R. S. Sutton, D. Precup, and S. Singh, “Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning,” Artificial Intelligence, vol. 112, no. 1– 2, pp. 181–211, 1999. https://dx.doi.org/10.1016/S0004-3702(99)00052-1.
  12. T. G. Dietterich, “Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition,” Journal of Artificial Intelligence Research, vol. 13, pp. 227–303, 2000. https://dx.doi.org/10.1613/jair.639.