Similarity-Based Approaches to Natural Language Processing

Lillian Lee (Cornell University)

arXiv:cmp-lg/9708011·cmp-lg·February 3, 2008·76 cites

Similarity-Based Approaches to Natural Language Processing

Lillian Lee (Cornell University)

PDF

Open Access

TL;DR

This thesis introduces two similarity-based methods for natural language processing, including hierarchical clustering and nearest-neighbor models, demonstrating significant improvements in word sense disambiguation, event prediction, and speech recognition.

Contribution

It presents novel hierarchical clustering and nearest-neighbor approaches tailored for sparse data problems in NLP, outperforming standard techniques.

Findings

01

Superior performance in word sense disambiguation

02

Over 20% perplexity reduction in low-frequency event prediction

03

Significant speech recognition error-rate improvements

Abstract

This thesis presents two similarity-based approaches to sparse data problems. The first approach is to build soft, hierarchical clusters: soft, because each event belongs to each cluster with some probability; hierarchical, because cluster centroids are iteratively split to model finer distinctions. Our second approach is a nearest-neighbor approach: instead of calculating a centroid for each class, as in the hierarchical clustering approach, we in essence build a cluster around each word. We compare several such nearest-neighbor approaches on a word sense disambiguation task and find that as a whole, their performance is far superior to that of standard methods. In another set of experiments, we show that using estimation techniques based on the nearest-neighbor model enables us to achieve perplexity reductions of more than 20 percent over standard techniques in the prediction of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques · Music and Audio Processing