References

tuzsut

Труды учебных заведений связи

Proceedings of Telecommunication Universities

1813-324X2712-8830

СПбГУТ

10.31854/1813-324X-2023-9-5-35-42

tuzsut-513

Research Article

ЭЛЕКТРОНИКА, ФОТОНИКА, ПРИБОРОСТРОЕНИЕ И СВЯЗЬ

ELECTRONICS, PHOTONICS, INSTRUMENTATION AND COMMUNICATIONS

Разработка и исследование системы автоматического распознавания цифр йеменского диалекта арабской речи с использованием нейронных сетей

Development and Research of a System for Automatic Recognition of the Digits Yemeni Dialect of Arabic Speech Using Neural Networks

https://orcid.org/0009-0006-1723-2782

Радан

Н.Х.А.

Radan

N.H.

аспирант кафедры информационных систем Тверского государственного технического университета

naeem.radan@gmail.com

https://orcid.org/0000-0003-1119-2610

Сидоров

К. В.

Sidorov

кандидат технических наук, доцент, доцент кафедры автоматизации технологических процессов Тверского государственного технического университета

bmisidorov@mail.ru

Тверской государственный технический университетРоссияTver State Technical UniversityRussian Federation

2023

14112023

953542

2023

Радан Н., Сидоров К.В.

Radan N., Sidorov K.

Данная работа распространяется под лицензией Creative Commons Attribution 4.0.

This work is licensed under a Creative Commons Attribution 4.0 License.

https://tuzs.sut.ru/jour/article/view/513

В статье описаны результаты исследований по разработке и тестированию системы автоматического распознавания речи (САРР) на арабских цифрах с помощью искусственных нейронных сетей. Для проведения исследований использовались звукозаписи (речевые сигналы) арабского йеменского диалекта, записанные в Республике Йемен. САРР представляет собой изолированную систему распознавания целых слов, она реализована в двух режимах: «дикторозависимая система» (дикторы при обучении и тестировании системы используются одни и те же) и «дикторонезависимая система» (дикторы, используемые для обучения системы, отличаются от тех, которые применяются для ее тестирования). В процессе распознавания речевой сигнал очищается от шумов с помощью фильтров, далее сигнал предварительно локализуется, обрабатывается и анализируется окном Хэмминга (применяется алгоритм временного выравнивания для компенсации различий в произношении). Информативные признаки извлекаются из речевого сигнала с использованием мел-частотных кепстральных коэффициентов. Разработанная САРР обеспечивает высокую точность распознавания арабских цифр йеменского диалекта – 96,2 % (для дикторозависимой системы) и 98,8 % (для дикторонезависимой системы).

The article describes the results of research on the development and testing of an automatic speech recognition system (SAR) in Arabic numerals using artificial neural networks. Sound recordings (speech signals) of the Arabic Yemeni dialect recorded in the Republic of Yemen were used for the research. SAR is an isolated system of recognition of whole words, it is implemented in two modes: "speaker-dependent system" (the same speakers are used for training and testing the system) and "speaker-independent system" (the speakers used for training the system differ from those used for testing it). In the process of speech recognition, the speech signal is cleared of noise using filters, then the signal is pre-localized, processed and analyzed by the Hamming window (a time alignment algorithm is used to compensate for differences in pronunciation). Informative features are extracted from the speech signal using mel-frequency cepstral coefficients. The developed SAR provides high accuracy of the recognition of Arabic numerals of the Yemeni dialect – 96.2 % (for a speaker-dependent system) and 98.8 % (for a speaker-independent system).

нейронные сетираспознавание речийеменский диалект

neural networksspeech recognitionYemeni dialect

References1

Al-Zabibi M. An acoustic-phonetic approach in automatic Arabic speech recognition. Loughborough University. Doctoral Thesis. 1990. URL: https://hdl.handle.net/2134/6949 (Accessed 02.10.2023)

Al-Zabibi M. An acoustic-phonetic approach in automatic Arabic speech recognition. Loughborough University. Doctoral Thesis. 1990. URL: https://hdl.handle.net/2134/6949 [Accessed 02.10.2023]

Alkhouli M. Alaswaat Alaghawaiyah // Daar Alfalah, Jordan. 1990 (in Arabic)

Alkhouli M. Alaswaat Alaghawaiyah. Daar Alfalah, Jordan. 1990 (in Arabic)

Deller J., Hansen J., Proakis J. Discrete-Time Processing of Speech Signal. 1993. DOI:10.1109/9780470544402

Elshafei M. Toward an Arabic Text-to-Speech System // The Arabian Journal for Scince and Engineering. 1991. Vol. 16. Iss. 4B. PP. 565‒583.

Elshafei M. Toward an Arabic Text-to-Speech System. The Arabian Journal for Science and Engineering. 1991;16(4B):565‒583.

Hagos E. Implementation of an Isolated Word Recognition System. M.Sc. Thesis. King Fahd University of Petroleum & Minerals Dhahran, Saudi Arabia. 1985.

Abdulla W.H., Abdul-Karim M.A.H. Real-time spoken Arabic digit recognizer // International Journal of Electronics. 1985. Vol. 59. Iss. 5. PP. 645–648. DOI:10.1080/00207218508920741

Abdulla W.H., Abdul-Karim M.A.H. Real-time spoken Arabic digit recognizer. International Journal of Electronics. 1985; 59(5):645–648. DOI:10.1080/00207218508920741

Alotaibi Y.A. Investigating spoken Arabic digits in speech recognition setting. Information Sciences. 2005. Vol. 173. Iss. 1-3. PP. 115–139. DOI:10.1016/j.ins.2004.07.008

Alotaibi Y.A. Investigating spoken Arabic digits in speech recognition setting. Information Sciences. 2005;173(1-3):115–139. DOI:10.1016/j.ins.2004.07.008

Alotaibi Y.A. High performance Arabic digits recognizer using neural networks // Proceedings of the International Joint Conference on Neural Networks (Portland, USA, 20‒24 July 2003). IEEE, 2003. DOI:10.1109/ijcnn.2003.1223444

Alotaibi Y.A. High performance Arabic digits recognizer using neural networks. Proceedings of the International Joint Conference on Neural Networks, 20‒24 July 2003, Portland, USA. IEEE; 2003. DOI:10.1109/ijcnn.2003.1223444

Alotaibi Y.A. Analyzing Arabic digit recognizer errors using spectrograms // Proceedings 7th International Conference on Signal Processing (ICSP, Beijing, China, 31 August 2004 ‒ 04 September 2004). IEEE, 2004. DOI:10.1109/icosp.2004.1452746

Alotaibi Y.A. Analyzing Arabic digit recognizer errors using spectrograms // Proceedings 7th International Conference on Signal Processing, ICSP, 31 August 2004 ‒ 04 September 2004, Beijing, China. IEEE; 2004. DOI:10.1109/icosp.2004.1452746

Hassine M., Boussaid L., Massaoud H. Tunisian Dialect Recognition Based on Hybrid Techniques // International Arab Journal of Information Technology. 2018. Vol. 15. No. 1. PP. 58–65.

Hassine M., Boussaid L., Massaoud H. Tunisian Dialect Recognition Based on Hybrid Techniques. International Arab Journal of Information Technology. 2018;15(1):58–65.

Аль-Дайбани А.М.С. Исследование методов и разработка алгоритмов обработки сигналов для систем автоматического распознавания телефонной речи в республике Йемен. Дис. … канд. техн. наук. Владимир: Владимирский государственный университет имени Александра Григорьевича и Николая Григорьевича Столетовых, 2019. 150 с.

Al-Daibani A.M.S. Research of methods and development of algorithms of signal processing for systems of automatic recognition of telephone speech in the Republic of Yemen. PhD Thesis. Vladimir: Vladimir State University named after Alex-ander Grigorievich and Nikolai Grigorievich Stoletov Publ.; 2019. 150 p.

Радан Н.Х. Системы автоматического распознавания арабской речи и йеменского диалекта // Научно-аналитический журнал «Вестник Санкт-Петербургского университета государственной противопожарной службы МЧС России». 2023. № 2. С. 194–212.

Radan N. Automatic speech recognition systems for Arabic speech and Yemeni dialect. Bulletin of St. Petersburg Uni-versity of the State Fire Service of the Ministry of Emergency Situations of Russia. 2023;2:194–212.

The authors declare that there are no conflicts of interest present.