BengaliLCP: A Dataset for Lexical Complexity Prediction in the Bengali Texts

Nabila Ayman; Md. Akram Hossain; Abdul Aziz; Rokan Uddin Faruqui; Abu Nowshed Chy

BengaliLCP: A Dataset for Lexical Complexity Prediction in the Bengali Texts

Nabila Ayman, Md. Akram Hossain, Abdul Aziz, Rokan Uddin Faruqui, Abu Nowshed Chy

Abstract

Encountering intricate or ambiguous terms within a sentence produces distress for the reader during comprehension. Lexical Complexity Prediction (LCP) deals with predicting the complexity score of a word or a phrase considering its context. This task poses several challenges including ambiguity, context sensitivity, and subjectivity in perceiving complexity. Despite having 300 million native speakers and ranking as the seventh most spoken language in the world, Bengali falls behind in the research on lexical complexity when compared to other languages. To bridge this gap, we introduce the first annotated Bengali dataset, that assists in performing the task of LCP in this language. Besides, we propose a transformer-based deep neural approach with a pairwise multi-head attention mechanism and LSTM model to predict the lexical complexity of Bengali tokens. The outcomes demonstrate that the proposed neural approach outperformed the existing state-of-the-art models for the Bengali language.

Anthology ID:: 2024.lrec-main.200
Volume:: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Month:: May
Year:: 2024
Address:: Torino, Italia
Editors:: Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
Venues:: LREC | COLING
SIG:
Publisher:: ELRA and ICCL
Note:
Pages:: 2227–2237
Language:
URL:: https://aclanthology.org/2024.lrec-main.200
DOI:
Bibkey:
Cite (ACL):: Nabila Ayman, Md. Akram Hossain, Abdul Aziz, Rokan Uddin Faruqui, and Abu Nowshed Chy. 2024. BengaliLCP: A Dataset for Lexical Complexity Prediction in the Bengali Texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2227–2237, Torino, Italia. ELRA and ICCL.
Cite (Informal):: BengaliLCP: A Dataset for Lexical Complexity Prediction in the Bengali Texts (Ayman et al., LREC-COLING 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.lrec-main.200.pdf

PDF Cite Search