On the Impact of Knowledge-based Linguistic Annotations in the Quality of Scientific Embeddings
Andres Garcia-Silva, Ronald Denaux, Jose Manuel Gomez-Perez

TL;DR
This paper investigates how incorporating explicit linguistic annotations, such as lexical, grammatical, and semantic features, affects the quality of scientific word embeddings across different evaluation tasks.
Contribution
It provides a comprehensive analysis of the impact of various linguistic annotations on embedding quality, highlighting their benefits in scientific text representations.
Findings
Linguistic annotations generally improve embedding evaluation scores.
The impact of annotations varies depending on the specific task.
Combining multiple annotations can enhance embedding quality.
Abstract
In essence, embedding algorithms work by optimizing the distance between a word and its usual context in order to generate an embedding space that encodes the distributional representation of words. In addition to single words or word pieces, other features which result from the linguistic analysis of text, including lexical, grammatical and semantic information, can be used to improve the quality of embedding spaces. However, until now we did not have a precise understanding of the impact that such individual annotations and their possible combinations may have in the quality of the embeddings. In this paper, we conduct a comprehensive study on the use of explicit linguistic annotations to generate embeddings from a scientific corpus and quantify their impact in the resulting representations. Our results show how the effect of such annotations in the embeddings varies depending on the…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
