Don't Just Scratch the Surface: Enhancing Word Representations for   Korean with Hanja

Kang Min Yoo; Taeuk Kim; Sang-goo Lee

arXiv:1908.09282·cs.CL·November 1, 2019

Don't Just Scratch the Surface: Enhancing Word Representations for Korean with Hanja

Kang Min Yoo, Taeuk Kim, Sang-goo Lee

PDF

3 Repos

TL;DR

This paper introduces a method to improve Korean word representations by incorporating Hanja through cross-lingual transfer learning, enhancing their quality for various NLP tasks.

Contribution

It presents a novel approach leveraging Hanja and cross-lingual transfer learning to enhance Korean word embeddings, validated through intrinsic and downstream evaluations.

Findings

01

Improved performance on word analogy and similarity tests.

02

Enhanced results on Korean news headline generation.

03

Effective use of Hanja for cross-lingual transfer in Korean NLP.

Abstract

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that Hanja is closely related to Chinese. We evaluate the intrinsic quality of representations learned through our approach using the word analogy and similarity tests. In addition, we demonstrate their effectiveness on several downstream tasks, including a novel Korean news headline generation task.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.