Effects of Stop Words Elimination for Arabic Information Retrieval: A   Comparative Study

Ibrahim Abu El-Khair

arXiv:1702.01925·cs.CL·February 8, 2017·22 cites

Effects of Stop Words Elimination for Arabic Information Retrieval: A Comparative Study

Ibrahim Abu El-Khair

PDF

Open Access

TL;DR

This study evaluates the impact of different stop words lists and weighting schemes on Arabic information retrieval effectiveness, finding that a general stoplist combined with BM25 weighting yields the best results.

Contribution

It compares three stop words lists and three weighting schemes, demonstrating the effectiveness of combining linguistic and statistical approaches for Arabic IR.

Findings

01

General stoplist outperforms other lists

02

BM25 weighting scheme yields best performance

03

Stoplists improve retrieval effectiveness

Abstract

The effectiveness of three stop words lists for Arabic Information Retrieval---General Stoplist, Corpus-Based Stoplist, Combined Stoplist ---were investigated in this study. Three popular weighting schemes were examined: the inverse document frequency weight, probabilistic weighting, and statistical language modelling. The Idea is to combine the statistical approaches with linguistic approaches to reach an optimal performance, and compare their effect on retrieval. The LDC (Linguistic Data Consortium) Arabic Newswire data set was used with the Lemur Toolkit. The Best Match weighting scheme used in the Okapi retrieval system had the best overall performance of the three weighting algorithms used in the study, stoplists improved retrieval effectiveness especially when used with the BM25 weight. The overall performance of a general stoplist was better than the other two lists.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsInformation Retrieval and Search Behavior · Text and Document Classification Technologies