UMBCLU at SemEval-2024 Task 1: Semantic Textual Relatedness with and without machine translation

Shubhashis Roy Dipta, Sai Vallurupalli


Abstract
The aim of SemEval-2024 Task 1, “Semantic Textual Relatedness for African and Asian Languages” is to develop models for identifying semantic textual relatedness (STR) between two sentences using multiple languages (14 African and Asian languages) and settings (supervised, unsupervised, and cross-lingual). Large language models (LLMs) have shown impressive performance on several natural language understanding tasks such as multilingual machine translation (MMT), semantic similarity (STS), and encoding sentence embeddings. Using a combination of LLMs that perform well on these tasks, we developed two STR models, TranSem and FineSem, for the supervised and cross-lingual settings. We explore the effectiveness of several training methods and the usefulness of machine translation. We find that direct fine-tuning on the task is comparable to using sentence embeddings and translating to English leads to better performance for some languages. In the supervised setting, our model performance is better than the official baseline for 3 languages with the remaining 4 performing on par. In the cross-lingual setting, our model performance is better than the baseline for 3 languages (leading to 1st place for Africaans and 2nd place for Indonesian), is on par for 2 languages and performs poorly on the remaining 7 languages.
Anthology ID:
2024.semeval-1.195
Volume:
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
Month:
June
Year:
2024
Address:
Mexico City, Mexico
Editors:
Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, Aiala Rosá
Venue:
SemEval
SIG:
SIGLEX
Publisher:
Association for Computational Linguistics
Note:
Pages:
1351–1357
Language:
URL:
https://aclanthology.org/2024.semeval-1.195
DOI:
10.18653/v1/2024.semeval-1.195
Bibkey:
Cite (ACL):
Shubhashis Roy Dipta and Sai Vallurupalli. 2024. UMBCLU at SemEval-2024 Task 1: Semantic Textual Relatedness with and without machine translation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), pages 1351–1357, Mexico City, Mexico. Association for Computational Linguistics.
Cite (Informal):
UMBCLU at SemEval-2024 Task 1: Semantic Textual Relatedness with and without machine translation (Roy Dipta & Vallurupalli, SemEval 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.semeval-1.195.pdf
Supplementary material:
 2024.semeval-1.195.SupplementaryMaterial.txt