Logo PTI Logo FedCSIS

Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 47

Outbound Traffic Control in Call/Contact Center Systems Using the TD3 Algorithm

, , ,

DOI: http://dx.doi.org/10.15439/2026F4509

Citation: Marcin Kozłowski, , ,

Full text

Abstract. Outbound traffic control in Call Center systems requires simultaneous consideration of efficient utilization of human resources and compliance with stringent quality requirements, particularly those related to acceptable call abandonment levels. This paper proposes an adaptive approach based on reinforcement learning, in which the control task is formulated as a continuous control problem. The operational policy is derived using the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, characterized by a stable learning process and deterministic action selection. A simulation environment and a reward function were developed to incorporate key operational performance indicators, such as the consultant occupancy rate (OR) and the abandonment rate (AR). The effectiveness of the proposed solution was evaluated using historical data from a real Call Center system, demonstrating improved resource utilization while maintaining a high level of service quality.

References

  1. M. Płaza, Ł. Pawlik, and R. S. Deniziak, “Call transcription methodology for contact center systems,” IEEE Access, vol. 9, pp. 110975–110988, 2021, https://dx.doi.org/10.1109/ACCESS.2021.3102502.
  2. Ł. Pawlik, M. Płaza, R. S. Deniziak, and E. Boksa, “A method for improving bot effectiveness by recognising implicit customer intent in contact centre conversations,” Speech Communication, vol. 143, pp. 33–45, 2022.
  3. M. Płaza, R. Kazała, Z. Koruba, M. Kozłowski, M. Lucińska, K. Sitek, and J. Spyrka, “Emotion recognition method for call/contact centre systems,” Applied Sciences, vol. 12, no. 21, Art. no. 10951, 2022.
  4. M. Ibrahim and N. B. Shroff, “Dynamic resource allocation for cloud-based call centers using reinforcement learning,” IEEE Transactions on Network and Service Management, vol. 16, no. 3, pp. 1122–1135, 2019.
  5. L. Feinberg, A. Mandelbaum, and B. Shalev-Shwartz, “Online learning for call center workforce management,” Operations Research, vol. 66, no. 6, pp. 1674–1691, 2018.
  6. S. Asghari, S. Yousefi, and M. S. Pishvaee, “A data-driven optimization approach for call center performance improvement,” European Journal of Operational Research, vol. 275, no. 2, pp. 676–690, 2019.
  7. G. Koole, Call Center Optimization. MG Books, 2013.
  8. N. Gans, G. Koole, and A. Mandelbaum, “Telephone call centers: Tutorial, review, and research prospects,” Manufacturing & Service Operations Management, vol. 5, no. 2, pp. 79–141, 2003.
  9. A. N. Avramidis and P. L’Ecuyer, “Modeling and simulation of call centers,” in Proc. Winter Simulation Conf., 2005, pp. 144–152.
  10. W. M. N. Dilini, A. Karunaratne, and D. N. Ranasinghe, “Methods for controlling the call placement ratio for outbound dialing of a call center,” in Proc. Statistical Concepts and Methods for the Modern World Conf., Colombo, 2011, pp. 1–12.
  11. N. Anisimov, N. Korolev, and H. Ristock, “Modeling and simulation of a pacing engine for proactive campaigns in contact center environment,” in Proc. Spring Simulation Multiconference, 2008, pp. 249–255.
  12. P. J. de Freitas Filho, G. F. da Cruz, R. Seara, and G. Steinmann, “Using simulation to predict market behavior for outbound call centers,” in Proc. Winter Simulation Conf., 2007, pp. 2247–2251.
  13. S. Fourati and S. Tabbane, “Optimization of a predictive dialing algorithm,” in Proc. Advanced Int. Conf. Telecommunications, 2010, pp. 212–218.
  14. P. M. Amaral and M. M. Vital, “Predictive dialer intensity optimization using genetic algorithms,” Int. J. Machine Learning and Computing, vol. 4, no. 3, pp. 286–291, 2014.
  15. Aurus, “Predictive vs. progressive outbound dialers test results,” 2023. [Online]. Available: https://aurus5.com/blog/cisco/predictive-vs-progressive-outbound-dialers-test-results/ [Accessed: Jun. 20, 2026].
  16. L. Qi, S. Ma, and K. Liu, “Research on predictive dialing system based on distributed call center,” in Proc. 4th Int. Conf. Software Engineering Research, Management and Applications (SERA’06), 2006, pp. 194–201, https://dx.doi.org/10.1109/SERA.2006.58.
  17. V. B. Iversen, Teletraffic Engineering Handbook. Technical University of Denmark (DTU), 2015.
  18. M. Zhang, Y. Sheng, N. Tian, W. Liu, H. Wang, L. Zhu, and Q. Xu, “Research and application of traffic forecasting in customer service center based on ARIMA model and LSTM neural network model,” J. Phys.: Conf. Ser., vol. 1881, Art. no. 032063, 2021.
  19. M. Płaza and Ł. Pawlik, “Influence of the contact center systems development on key performance indicators,” IEEE Access, vol. 9, pp. 44580–44591, 2021, https://dx.doi.org/10.1109/ACCESS.2021.3066801.
  20. Ofcom, “Statement of policy on persistent misuse of electronic communications networks or services,” 2006 (rev. 2016).
  21. Federal Communications Commission, “47 C.F.R. § 64.1200: Delivery restrictions,” Electronic Code of Federal Regulations, 2024.
  22. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. MIT Press, 2018.
  23. C. Szepesvári, Algorithms for Reinforcement Learning. Morgan & Claypool, 2010.
  24. C. J. C. H. Watkins and P. Dayan, “Q-learning,” Machine Learning, vol. 8, pp. 279–292, 1992.
  25. V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015.
  26. M. Hu, “Problems with continuous action space,” in The Art of Reinforcement Learning. Apress, 2023.
  27. T. Degris, M. White, and R. S. Sutton, “Off-policy actor-critic,” in Proc. 29th Int. Conf. Machine Learning (ICML), 2012, pp. 179–186.
  28. T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in Proc. 4th Int. Conf. Learning Representations (ICLR), 2016. [Online]. Available: https://arxiv.org/abs/1509.02971
  29. S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor–critic methods,” in Proc. 35th Int. Conf. Machine Learning (ICML), 2018, pp. 1587–1596.
  30. T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proc. 35th Int. Conf. Machine Learning (ICML), 2018, pp. 1861–1870.
  31. S. Liu, “An evaluation of DDPG, TD3, SAC, and PPO: Deep reinforcement learning algorithms for controlling continuous systems,” in Proc. 2023 Int. Conf. Data Science, Advanced Algorithm and Intelligent Computing (DAI 2023), Atlantis Press, 2024.
  32. D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proc. 31st Int. Conf. Machine Learning (ICML), 2014, pp. 387–395.
  33. I. Smirnov and S. Gu, “RLBenchNet: The right network for the right reinforcement learning task,” arXiv preprint https://arxiv.org/abs/2505.15040, 2025.
  34. Stable-Baselines3 Contributors, “Stable-Baselines3 documentation,” 2021. [Online]. Available: https://stable-baselines3.readthedocs.io [Accessed: Jun. 20, 2026].
  35. Z. Ben Hazem, F. Saidi, N. Guler, and A. H. Altaif, “A hybrid reinforcement learning framework combining TD3 and PID control for robust trajectory tracking of a 5-DOF robotic arm,” Automation, vol. 6, no. 4, Art. no. 56, 2025.
  36. P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proc. AAAI Conf. Artificial Intelligence, vol. 32, no. 1, 2018.