Scalable Cross-lingual Document Similarity through Language-specific   Concept Hierarchies

Carlos Badenes-Olmedo; Jose-Luis Redondo Garc\'ia; Oscar Corcho

arXiv:2101.03026·cs.CL·January 11, 2021

Scalable Cross-lingual Document Similarity through Language-specific Concept Hierarchies

Carlos Badenes-Olmedo, Jose-Luis Redondo Garc\'ia, Oscar Corcho

PDF

1 Repo

TL;DR

This paper introduces an unsupervised, scalable method for cross-lingual document similarity that leverages language-specific concept hierarchies without needing parallel corpora or translation resources.

Contribution

It proposes a novel unsupervised algorithm that annotates topics with cross-lingual labels using independently-trained models, enabling scalable multi-lingual document comparison.

Findings

01

Effective classification and sorting of documents across English, Spanish, and French.

02

No need for parallel or comparable corpora or translation resources.

03

Promising results in multi-lingual document similarity tasks.

Abstract

With the ongoing growth in number of digital articles in a wider set of languages and the expanding use of different languages, we need annotation methods that enable browsing multi-lingual corpora. Multilingual probabilistic topic models have recently emerged as a group of semi-supervised machine learning models that can be used to perform thematic explorations on collections of texts in multiple languages. However, these approaches require theme-aligned training data to create a language-independent space. This constraint limits the amount of scenarios that this technique can offer solutions to train and makes it difficult to scale up to situations where a huge collection of multi-lingual documents are required during the training phase. This paper presents an unsupervised document similarity algorithm that does not require parallel or comparable corpora, or any other type of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

cbadenes/crosslingual-semantic-similarity
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.