Wild Card Queries for Searching Resources on the Web

Davood Rafiei; Haobin Li

arXiv:0908.2588·cs.DB·August 19, 2009·2 cites

Wild Card Queries for Searching Resources on the Web

Davood Rafiei, Haobin Li

PDF

Open Access

TL;DR

This paper introduces a flexible, domain-independent wild card query framework for extracting facts from natural language texts, addressing challenges in query expansion and result ranking to improve retrieval accuracy.

Contribution

It presents a novel wild card query mechanism for resource retrieval, along with analysis and evaluation of query expansion and ranking strategies.

Findings

01

The framework effectively retrieves facts with high precision.

02

Query expansion can introduce false positives, requiring careful ranking.

03

Ranking strategies improve the accuracy of retrieved results.

Abstract

We propose a domain-independent framework for searching and retrieving facts and relationships within natural language text sources. In this framework, an extraction task over a text collection is expressed as a query that combines text fragments with wild cards, and the query result is a set of facts in the form of unary, binary and general $n$ -ary tuples. A significance of our querying mechanism is that, despite being both simple and declarative, it can be applied to a wide range of extraction tasks. A problem in querying natural language text though is that a user-specified query may not retrieve enough exact matches. Unlike term queries which can be relaxed by removing some of the terms (as is done in search engines), removing terms from a wild card query without ruining its meaning is more challenging. Also, any query expansion has the potential to introduce false positives. In…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsWeb Data Mining and Analysis · Natural Language Processing Techniques · Topic Modeling