XLEnt: Mining a Large Cross-lingual Entity Dataset with   Lexical-Semantic-Phonetic Word Alignment

Ahmed El-Kishky; Adithya Renduchintala; James Cross; Francisco; Guzm\'an; Philipp Koehn

arXiv:2104.08597·cs.CL·September 13, 2021

XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment

Ahmed El-Kishky, Adithya Renduchintala, James Cross, Francisco, Guzm\'an, Philipp Koehn

PDF

Open Access

TL;DR

XLEnt introduces LSP-Align, a novel method for automatically mining a large-scale cross-lingual entity dataset from web data, significantly aiding multilingual NLP tasks by providing extensive entity pairs across 120 languages.

Contribution

The paper presents LSP-Align, a new technique that outperforms existing methods in extracting cross-lingual entity pairs and releases a large multilingual entity dataset for NLP research.

Findings

01

Extracted 164 million cross-lingual entity pairs

02

Outperforms baseline methods in entity pair extraction

03

Provides a resource for 120 languages aligned with English

Abstract

Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification. While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing. To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data. We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English. We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Speech and dialogue systems