Research Article | ![]()
Error Signal Based Method for Real Time Pitch Extraction from Noisy Speech
Author(s): Mohammad Forhad Hossain1, Nargis Parvi 2, Md. Saifur Rahman3, Pintu Chandra Paul4, and Moinur Rahman5*
Published In : International Journal of Electrical and Electronics Research (IJEER) Volume 14, Issue 3
Publisher : FOREX Publication
Published : 30 September 2026
e-ISSN : 2347-470X
Page(s) : 842-851
Abstract
Pitch estimation efficiency always gets affected by two major phenomena. The first one is additive noise which by nature are not correlated with the source speech signal. The second one is the vocal tract effect which creates unnecessary peaks within the signal waveform, resulting inaccurate detection of the pitch value. This paper proposes a Modified Weighted Autocorrelation Function (MWAF) executed in the Linear Predictive Coding (LPC) residual domain, instead of directly executing over the noisy source speech signal. Converting the noisy source speech signal to LP residual helps eliminating the vocal tract effects and emphasizing the necessary effects of glottal excitation. The Autocorrelation Function (ACF) and the Circular Average Magnitude Difference Function (CAMDF) are directly applied on this residual. The proposed MWAF is then calculated by multiplying the ACF with the reciprocal of the CAMDF, which enhances the periodic components as well as reduces the noise induced false peaks. Experimental results were evaluated on the KEELE and NTT databases across White, Babble, Train and HF Channel noise conditions. The result shows that our proposed algorithm reduces the Gross Pitch Error (GPE) by an average of 40.34% compared to baseline methods such as BaNa, PEFAC and WAF. Furthermore, the algorithm achieves an overall computational gain of 53.44% compared to BaNa, PEFAC and WAF, offering a stable solution for pitch tracking in degraded acoustic conditions.
Keywords: Pitch, Linear Predictive Coding, Modified Weighted Autocorrelation, LPC Residual.
Mohammad Forhad Hossain , Department of ICT, Comilla University, Bangladesh; Email: forhad-ict-cou@stud.cou.ac.bd
Nargis Parvin , Associate Professor, Department of CSE, Bangladesh Army International University of Science and Technology, Bangladesh; Email: nargis.cse@baiust.ac.bd
Md. Saifur Rahman, Professor, Department of ICT, Comilla University, Bangladesh; Email: saifurice@cou.ac.bd
Pintu Chandra Paul, Assistant Professor, Department of ICT, Comilla University, Bangladesh; Email: pintu@cou.ac.bd
Moinur Rahman , Lecturer, Department of ICT, Comilla University, Bangladesh; Email: moinur.rahman@cou.ac.bd
-
[1] Vary, P. and Martin, R., “Digital Speech Transmission: Enhancement, Coding and Error Concealment,” John Wiley & Sons, New York, 2006, https://doi.org/10.1002/0470031743.
-
[2] Hess W. “Pitch determination of speech signals: algorithms and devices,” in Springer Science & Business Media, Dec 6 2012, https://doi.org/10.1007/978-3-642-81926-1
-
[3] M. S. Rahman, “Pitch extraction for speech signals in noisy environments,” P.hD. Dissertation, Department of Mathematics, Electronics, and Informatics, Saitama University, Saitama, Japan, 2020. [Online].Available:https://sucra.repo.nii.ac.jp/record/19377/files/GD0001258.pdf.
-
[4] D. Wang, C. Yu and J. H. L. Hansen, "Robust Harmonic Features for Classification-Based Pitch Estimation," in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 5, pp. 952-964, May 2017, doi: 10.1109/TASLP.2017.2667879.
-
[5] S. S. Upadhya, "Pitch detection in time and frequency domain," in International Conference on Communication, Information & Computing Technology (ICCICT), Mumbai, India, 2012, pp. 1-5, doi: 10.1109/ICCICT.2012.6398150.
-
[6] L. R. Rabiner, "On the use of autocorrelation analysis for pitch detection", IEEE Transaction on Acoustics, Speech, Signal Processing, Vol. ASSP-25, No. 1, pp. 24–33, 1977, doi: 10.1109/TASSP.1977.1162905.
-
[7] M. Ross, H. Shaffer, A. Cohen, R. Freudberg and H. Manley, "Average magnitude difference function pitch extractor," IEEE Transactions on Acoustics, Speech, and Signal Processing, Vol. 22, No. 5, pp. 353-362, 1974, doi: 10.1109/TASSP.1974.1162598.
-
[8] R. Chakraborty, D. Sengupta, and S. Sinha, "Pitch tracking of acoustic signals based on average squared mean difference function," Signal, image and video processing, Vol. 3, No. 4, pp. 319–327, 2009, doi: 10.1007/s11760-008-0072-5.
-
[9] T. Shimamura and H. Kobayashi, "Weighted autocorrelation for pitch extraction of noisy speech,” IEEE Transactions on Speech and Audio Processing, Vol. 9, No. 7, pp. 727-730, 2001, doi: 10.1109/89.952490
-
[10] A. De Cheveigne and H. Kawahara, "Yin, a fundamental frequency estimator for speech and music," The Journal of the Acoustical Society of America, Vol. 111, No. 4, pp. 1917–1930, 2002, doi: 10.1121/1.1458024.
-
[11] S. Ahmadi and A. S. Spanias, "Cepstrum-based pitch detection using a new statistical v/uv classification algorithm," IEEE Transactions on Speech and Audio Processing, Vol. 7, No. 3, pp. 333–338, 1999, doi: 10.1109/89.759042.
-
[12] Kunieda N, Shimamura T, Suzuki J, "Pitch extraction by using autocorrelation function on the log spectrum," Electronics and Communications in Japan, Part 3, Vol. 83, No.1, pp. 90–98, 2000, doi: 10.1002/(SICI)1520-6440(200001)83.
-
[13] Hasan MAFMR, Rahman MS, Shimamura T. "Windowless autocorrelation-based cepstrum method for pitch extraction of noisy speech,” Journal of Signal Processing, Vol. 16, No. 3, pp. 231-239, 2012, doi: 10.2299/jsp.16.231.
-
[14] S. Gonzalez and M. Brookes, "PEFAC - A Pitch Estimation Algorithm Robust to High Levels of Noise," IEEE/ACM Transactions on Audio, Speech, and Language Processing, Vol. 22, No. 2, pp. 518-530, 2014, doi: 10.1109/TASLP.2013.2295918.
-
[15] N. Yang, H. Ba, W. Cai, I. Demirkol and W. Heinzelman, "BaNa: A Noise Resilient Fundamental Frequency Detection Algorithm for Speech and Music," IEEE/ACM Transactions on Audio, Speech, and Language Processing, Vol. 22, No. 12, pp. 1833- 1848, 2014, doi: 10.1109/TASLP.2014.2352453.
-
[16] J. W. Kim, J. Salamon, P. Li and J. P. Bello, "Crepe: A Convolutional Representation for Pitch Estimation," 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 2018, pp. 161-165, doi: 10.1109/ICASSP.2018.8461329.
-
[17] Zhang, Y., Wang, H., Wang, D. "Densely-connected Convolutional Recurrent Network for Fundamental Frequency Estimation in Noisy Speech," Proc. Interspeech 2022, pp. 401-405, doi: 10.21437/Interspeech.2022-11156
-
[18] H. Schröter, T. Rosenkranz, A.N. Escalante-B, and A. Maier, "LACOPE: Latency-Constrained Pitch Estimation for Speech Enhancement," Proc. Interspeech 2021, pp. 656-660, 2021, doi: 10.21437/Interspeech.2021-633
-
[19] B. Li and X. Zhang, "A Pitch Estimation Algorithm for Speech in Complex Noise Environments Based on the Radon Transform," in IEEE Access, vol. 11, pp. 9876-9889, 2023, doi: 10.1109/ACCESS.2023.3240181.
-
[20] V. Këpuska and H. Elharati, "RNN-Based F0 Estimation Method with Attention Mechanism," MDPI Information, vol. 16, no. 12, 2024
-
[21] T. T. Nway and T. Shimamura, "Weighted Autocorrelation with Convolutional Neural Network for Noisy Speech Pitch Estimation," 2024 IEEE 13th Global Conference on Consumer Electronics (GCCE), Kitakyushu, Japan, 2024, pp. 510-511, doi: 10.1109/GCCE62371.2024.10760946
-
[22] K. Subramani, J. -M. Valin, J. Büthe, P. Smaragdis and M. Goodwin, "Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11851-11855, doi: 10.1109/ICASSP48485.2024.10447962.
-
[23] C. Y. Jeong, Y. Song, S. Shin, and M. Kim, "Efficient pitch‐estimation network for edge devices," ETRI Journal, vol. 47, no. 1, pp. 112-122, 2025, doi: 10.4218/etrij.2023-0430.
-
[24] Rabiner, Lawrence R. "Theory and applications of digital speech processing." (2011).
-
[25] Fant, G. "Acoustic theory of speech production. the Hague, the Netherlands: Mouton & Co.", 1960, pp: 582-585.
-
[26] J. Makhoul, "Linear prediction: A tutorial review," in Proceedings of the IEEE, vol. 63, no. 4, pp. 561-580, April 1975, doi: 10.1109/PROC.1975.9792.
-
[27] Plante F, Meyer G, Ainsworth W, "A fundamental frequency extraction reference database," Proceedings of the Eurospeech, pp. 837–840, 1995, doi: 10.21437/Eurospeech.1995-191.
-
[28] 20 Countries Language Database, NTT Advanced Technology Corp., Jpn, (1988)
-
[29] Varga, A., Steeneken, HJ., Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems, Speech Communication vol. 12, no. 3, pp. 247–251, 1993.
-
[30] Itahashi, S.: Creating speech copora for speech science and technology, IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences vol. 74, no. 7, pp. 1906 1910, 1991
-
[31] Rabiner, L., Cheng, M., Rosenberg, A., McGonegal, C., A comparative performance study of several pitch detection algorithms, IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 24, no. 5, pp. 399-418, 1976.

I. J. of Electrical & Electronics Research