(2) Dessi Puji Lestari
(3) * Aulia Rahmawati
*corresponding author
AbstractThe use of future context in acoustic modeling seems to give an impact on system performance such us Bidirectional Long Short-Term Memory (BLSTM). It has been used as an acoustic model on Speech Recognition System for Quran recitation and show better result than Hidden Markov Model - Gaussian mixture model (HMM-GMM) with average Word Error Rate (WER) value 4.6%. but, the architectural complexity of BLSTM make the latency during decoding process is high. To reduce the latency, Minimal Gated Recurrent Unit with Temporal Convolution (mGRUIPTC) acoustic model was used. Text data such as transcription, lexicon, and corpus used in training are represented at phone level to handle phone level detection. The transcription is generated using modified QScript to handle reciting rules in detail. In the test, the system can reduce decode process latency by up to 11 seconds with Phone Error Rate (PER) difference of up to 1.46% compared to BLSTM. However, our model still needs to be trained with more data to detect error recitation better.
Keywordsacoustic modeling; future context; gated recurrent unit; quran recitation; speech recognition
|
DOIhttps://doi.org/10.29099/ijair.v10i1.1613 |
Article metrics10.29099/ijair.v10i1.1613 Abstract views : 153 | PDF views : 13 |
Cite |
Full Text Download
|
References
M. Fatkhurrohman, “Penerapan Metode Muraja’ah dalam meningkatkan Kwalitas Hafalan Al-Qur’an Siswa Kelas VII A di SMPAL-Muayyad Surakarta Tahun Pelajaran 2018/2019,” Institut Agama Islam Negeri Surakarta, 2019.
F. Thirafi and D. P. Lestari, “Hybrid HMM-BLSTM-Based Acoustic Modeling for Automatic Speech Recognition on Quran Recitation,” in Proceedings of the 2018 International Conference on Asian Language Processing, IALP 2018, Institute of Electrical and Electronics Engineers Inc., Jan. 2019, pp. 203–208. doi: 10.1109/IALP.2018.8629184.
J. Li, X. Wang, Y. Zhao, and Y. Li, “Gated Recurrent Unit Based Acoustic Modeling with Future Context,” in INTERSPEECH, May 2018, pp. 1788–1792. Accessed: Jun. 07, 2020. [Online]. Available: http://arxiv.org/abs/1805.07024
W. M. Muhammad, R. Muhammad, A. Muhammad, and A. M. Martinez-Enriquez, “Voice content matching system for quran readers,” in Proceedings of Special Session - 9th Mexican International Conference on Artificial Intelligence: Advances in Artificial Intelligence and Applications, MICAI 2010, 2010, pp. 148–153. doi: 10.1109/MICAI.2010.11.
R. Yuwan and D. P. Lestari, “Automatic Extraction Phonetically Rich and Balanced Verses for Speaker-Dependent Quranic Speech Recognition System,” in International Conference of the Pacific Association for Computational Linguistics. Springer, Springer International Publishing, 2015.
M. Mohri, F. Pereira, and M. Riley, “SPEECH RECOGNITION WITH WEIGHTED FINITE-STATE TRANSDUCERS.”
D. Povey et al., “The Kaldi Speech Recognition Toolkit.” Accessed: Jul. 30, 2020. [Online]. Available: http://kaldi.sf.net/
D. Povey et al., “Purely sequence-trained neural networks for ASR based on lattice-free MMI.”
O. M. Strand and A. Egeberg, “Cepstral mean and variance normalization in the model domain,” 2004. Accessed: Dec. 18, 2020. [Online]. Available: http://www.isca-speech.org/archive
S. Geirhofer, “Feature Reduction with Linear Discriminant Analysis and its Performance on Phoneme Recognition,” 2004.
M. J. F. Gales, “Maximum Likelihood Linear Transformations for HMM-Based Speech Recognition,” Comput. Speech Lang., vol. 12, no. 2, pp. 75–98, 1998.
S. Matsoukas, R. Schwartz, H. Jin, and L. Nguyen, “Practical Implementations of Speaker-Adaptive Training,” in DARPA Speech Recognition Workshop, 1997. [Online]. Available: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.43.481&rep=rep1&type=pdf
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker adaptation of neural network acoustic models using i-vectors,” in 2013 IEEE Workshop on Automatic Speech Recognition and Understanding, ASRU 2013 - Proceedings, 2013, pp. 55–59. doi: 10.1109/ASRU.2013.6707705.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
________________________________________________________
The International Journal of Artificial Intelligence Research
Organized by: Prodi Teknik Informatika Fakultas Teknologi Bisnis dan Sains
Published by: Universitas Dharma Wacana
Jl. Kenanga No. 03 Mulyojati 16C Metro Barat Kota Metro Lampung
Email: jurnal.ijair@gmail.com

This work is licensed under Creative Commons Attribution-ShareAlike 4.0 International License.













Download