AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction

Bhavya Bhavya, Shradha Sehgal, Jinjun Xiong, ChengXiang Zhai


Abstract
Textual analogies that make comparisons between two concepts are often used for explaining complex ideas, creative writing, and scientific discovery. In this paper, we propose and study a new task, called Analogy Detection and Extraction (AnaDE), which includes three synergistic sub-tasks: 1) detecting documents containing analogies, 2) extracting text segments that make up the analogy, and 3) identifying the (source and target) concepts being compared. To facilitate the study of this new task, we create a benchmark dataset by scraping Metamia.com and investigate the performances of state-of-the-art models on all sub-tasks to establish the first-generation benchmark results for this new task. We find that the Longformer model achieves the best performance on all the three sub-tasks demonstrating its effectiveness for handling long texts. Moreover, smaller models fine-tuned on our dataset perform better than non-finetuned ChatGPT, suggesting high task difficulty. Overall, the models achieve a high performance on documents detection suggesting that it could be used to develop applications like analogy search engines. Further, there is a large room for improvement on the segment and concept extraction tasks.
Anthology ID:
2024.eacl-long.103
Volume:
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:
March
Year:
2024
Address:
St. Julian’s, Malta
Editors:
Yvette Graham, Matthew Purver
Venue:
EACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
1723–1737
Language:
URL:
https://aclanthology.org/2024.eacl-long.103
DOI:
Bibkey:
Cite (ACL):
Bhavya Bhavya, Shradha Sehgal, Jinjun Xiong, and ChengXiang Zhai. 2024. AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1723–1737, St. Julian’s, Malta. Association for Computational Linguistics.
Cite (Informal):
AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction (Bhavya et al., EACL 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.eacl-long.103.pdf
Video:
 https://aclanthology.org/2024.eacl-long.103.mp4