Acronym Disambiguation: A Domain Independent Approach

Aditya Thakker; Suhail Barot; Sudhir Bagul

arXiv:1711.09271·cs.CL·December 19, 2017·2 cites

Acronym Disambiguation: A Domain Independent Approach

Aditya Thakker, Suhail Barot, Sudhir Bagul

PDF

Open Access 1 Repo

TL;DR

This paper introduces a domain-independent system for acronym disambiguation that leverages Wikipedia and AcronymsFinder.com to gather context and uses Doc2Vec embeddings to achieve over 90% accuracy in selecting correct expansions.

Contribution

It presents a novel, general approach for acronym disambiguation using context retrieval and paragraph embeddings, applicable across domains.

Findings

01

Achieved 90.9% accuracy in disambiguation

02

Built a dataset with 707 acronyms and 14,876 disambiguations

03

Demonstrated effectiveness of Doc2Vec in context scoring

Abstract

Acronyms are omnipresent. They usually express information that is repetitive and well known. But acronyms can also be ambiguous because there can be multiple expansions for the same acronym. In this paper, we propose a general system for acronym disambiguation that can work on any acronym given some context information. We present methods for retrieving all the possible expansions of an acronym from Wikipedia and AcronymsFinder.com. We propose to use these expansions to collect all possible contexts in which these acronyms are used and then score them using a paragraph embedding technique called Doc2Vec. This method collectively led to achieving an accuracy of 90.9% in selecting the correct expansion for given acronym, on a dataset we scraped from Wikipedia with 707 distinct acronyms and 14,876 disambiguations.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

adityathakker/AcronymExpansion
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBiomedical Text Mining and Ontologies · Natural Language Processing Techniques · Advanced Text Analysis Techniques