Logo PTI Logo FedCSIS

Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 47

Voicemail Detection in Telephone Recordings for Call/Contact Center Systems Using Hierarchical Spectral Windowing

, ,

DOI: http://dx.doi.org/10.15439/2026F7609

Citation: Michał Zawadzki, ,

Full text

Abstract. This paper proposes a custom system for automatic voicemail detection designed for production-grade call/contact center environments. The detection problem is formulated as a binary classification task, where the objective is to distinguish live speech from a recorded voicemail message. The proposed system is based on a two-level hierarchical audio signal windowing structure combined with dominant spectral component extraction using the Discrete Fourier Transform. The extracted feature representation serves as input to a multilayer perceptron performing binary classification. The system was designed for telephone recordings sampled at 8 kHz. Experiments were conducted on a dataset of 7321 authentic telephone recordings from a production contact center. A systematic search over filtering threshold parameters yielded 56 experiments. The best accuracy of 93.35\% was achieved using a small window threshold of 1.5 and a large window threshold of 0.1. The results confirmed the effectiveness of the proposed approach and highlighted the role of filtering parameters in classification accuracy.

References

  1. C. Sánchez, S. Maldonado and C. Vairetti, Improving debt collection via contact center information: A predictive analytics framework. Decision Support Systems, 2022. doi.org/10.1016/j.dss.2022.113812.
  2. M.H. Rim, K.C. Thomas, J. Chandramouli, S.A. Barrus and N.A. Nickman, Implementation and quality assessment of a pharmacy services call center for outpatient pharmacies and specialty pharmacy services in an academic health system. Am J Health Syst Pharm. 2018. https://dx.doi.org/10.2146/ajhp170319.
  3. A. Graves, A. Mohamed and G. Hinton, Speech recognition with deep recurrent neural networks, 2013 IEEE International Conference on Acoustics. Speech and Signal Processing, Vancouver, BC, Canada, 2013, pp. 6645-6649. https://dx.doi.org/10.1109/ICASSP.2013.6638947.
  4. T. Noga, The use of chatbots and voicebots by public institutions in the communication process with clients. Zeszyty Naukowe. Organizacja i Zarządzanie/Politechnika Śląska, 2023. https://dx.doi.org/10.29119/1641-3466.2023.174.6.
  5. M. Płaza, R. Kazała, Z. Koruba, M. Kozłowski, M. Lucińska, K. Sitek, and J. Spyrka, Emotion Recognition Method for Call/Contact Centre Systems. Applied Sciences, vol. 12, no. 21, p. 10951, 2022. doi.org/10.3390/app122110951.
  6. M. Bojanić, V. Delić and A. Karpov, Call Redistribution for a Call Center Based on Speech Emotion Recognition. Applied Sciences. 2020; 10(13):4653. doi.org/10.3390/app10134653.
  7. K. Poczęta, M. Płaza, T. Michno, M. Krechowicz, and M. Zawadzki, A multi-label text message classification method designed for applications in call/contact centre systems, Applied Soft Computing, vol. 145, pp. 1–20, 2023. doi.org/10.1016/j.asoc.2023.110562.
  8. Z. Ahanin, M.A. Ismail, N.S.S. Singh and A. AL-Ashmorim, Hybrid Feature Extraction for Multi-Label Emotion Classification in English Text Messages. Sustainability. 2023; 15(16):12539. doi.org/10.3390/su151612539.
  9. M. Kozłowski, M. Płaza, M. Lucińska and M. Kowalczyk, Metoda predykcji ruchu wychodzącego w systemach klasy call/contact center wykorzystująca algorytm uczenia ze wzmocnieniem, Przegląd Telekomunikacyjny - Wiadomości Telekomunikacyjne, vol. 4, pp. 348–351, 2025. https://dx.doi.org/10.15199/59.2025.4.78.
  10. M. Płaza and Ł. Pawlik, Influence of the Contact Center Systems Development on Key Performance Indicators, IEEE Access, vol. 9, pp. 44580–44591, 2021. https://dx.doi.org/10.1109/ACCESS.2021.3066801.
  11. C.S.C Correia, Optimizing Call-Center Operations with Reinforcement Learning. MS thesis. Universidade do Porto (Portugal), 2020.
  12. E. Minaj, Optimization of answering machine detection application in asterisk. 2020. lutpub.lut.fi/handle/10024/161698.
  13. A. Kopylov, O. Seredin, A. Filin and B. Tyshkevich, Detection of interactive voice response (IVR) in phone call records. International Journal of Speech Technology 23.4 2020: 907-915. doi.org/10.1007/s10772-020-09754-3.
  14. K. Altwlkany, S. Delalić, E. Selmanović, A. Alihodžić and I. Lovrić, A Recurrent Neural Network Approach to the Answering Machine Detection Problem. 2024 47th MIPRO ICT and Electronics Convention (MIPRO). IEEE, 2024. https://dx.doi.org/10.1109/MIPRO60963.2024.10569812.
  15. T. Joseph, K. Tyagi and R. Kumbhare, Quantitative analysis of DTMF tone detection using DFT, FFT and Goertzel algorithm. 2019 Global Conference for Advancement in Technology (GCAT). IEEE, 2019. https://dx.doi.org/10.1109/GCAT47503.2019.8978284.
  16. V. Balambica, N. Hamadamen, A. Karthikayen and M. Praveen, Digital signal processing dual tone multifrequency detector. 2023. https://dx.doi.org/10.37896/YMER22.02/98.
  17. P. Sysel and P. Rajmic, Goertzel algorithm generalized to non-integer multiples of fundamental frequency. EURASIP J. Adv. Signal Process. 2012, 56 (2012). https://doi.org/10.1186/1687-6180-2012-56.
  18. D. Pereira and R. Oliveira, Detection of Abnormal SIP Signaling Patterns: A Deep Learning Comparison. Computers 2022. doi.org/10.3390/computers11020027.
  19. H. Luo and Y. Zhang, Beep tone detection within RTP streams based on TK energy operator and DESA2 algorithm, 2011 International Conference on Electronics, Communications and Control (ICECC), Ningbo, China, 2011, pp. 766-769, https://dx.doi.org/10.1109/ICECC.2011.6066454.
  20. https://www.rfc-editor.org/rfc/rfc7044.txt.
  21. https://datatracker.ietf.org/doc/html/rfc4458.
  22. J.K. Deters, S. Janus, J.A.L. Silva, H.J. Wörtche and S.U. Zuidema, Sensor-based agitation prediction in institutionalized people with dementia A systematic review, Pervasive and Mobile Computing, 2024. doi.org/10.1016/j.pmcj.2024.101876.
  23. B.T. Atmaja and M. Akagi, On The Differences Between Song and Speech Emotion Recognition: Effect of Feature Sets, Feature Types, and Classifiers, 2020 IEEE REGION 10 CONFERENCE (TENCON), Osaka, Japan, 2020, pp. 968-972, https://dx.doi.org/10.1109/TENCON50793.2020.9293852.
  24. S. Lachenani, H. Kheddar and M. Ouldzmirli, Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning, 2024 International Conference on Telecommunications and Intelligent Systems (ICTIS), Djelfa, Algeria, 2024, pp. 1-6, https://dx.doi.org/10.1109/ICTIS62692.2024.10894266.
  25. N. Anisimov, M. Olkhovskyr and K. Kishinsky, Answering machines detection in contact centers using convolutional neural networks. url: www.fiztech-usa.net/anisimov/papers/answering-machines-detection_6.pdf.
  26. B. Jiang and J.Yang, Preferred frame length for the short-time magnitude spectrum on speech intelligibility and speech quality, 2011 8th International Conference on Information, Communications & Signal Processing, Singapore, 2011, pp. 1-3, https://dx.doi.org/10.1109/ICICS.2011.6174266.
  27. G. Ma, P. Hu, J. Kang, S. Huang and H. Huang, Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognition, 2022. doi.org/10.21437/Interspeech.2021-964.
  28. A.B. Abdusalomov, F. Safarov, M. Rakhimov, B. Turaev and T.K. Whangbo, Improved Feature Parameter Extraction from Speech Signals Using Machine Learning Algorithm. Sensors 2022, 22, 8122. doi.org/10.3390/s22218122.
  29. M. Pereira, S. Chapaneri and D. Jayaswal, Analysis of windowing techniques for speech emotion recognition, 2016 International Conference on Information Communication and Embedded Systems (ICICES), Chennai, India, 2016, pp. 1-6, https://dx.doi.org/10.1109/ICICES.2016.7518859.
  30. J.G. Proakis and D.G. Manolakis, Digital Signal Processing: Principles, Algorithms, and Applications, Prentice-hall international INC. 1996. ISBN-10: 0133942899.