Initial Experiences Re-Exporting Duplicate and Similarity Computation   with an OAI-PMH aggregator

Terry L. Harrison; Aravind Elango; Johan Bollen; Michael Nelson

arXiv:cs/0401001·cs.DL·May 23, 2007·5 cites

Initial Experiences Re-Exporting Duplicate and Similarity Computation with an OAI-PMH aggregator

Terry L. Harrison, Aravind Elango, Johan Bollen, Michael Nelson

PDF

Open Access

TL;DR

This paper presents initial experiences with re-exporting similarity computations of metadata records via an OAI-PMH aggregator, aiding in duplicate detection and metadata quality improvement.

Contribution

It introduces an implementation that uses the <about> container to re-export similarity results, enabling better metadata management across harvests.

Findings

01

Successfully computed similarities for 3751 records

02

Detected duplicates and metadata errors effectively

03

Demonstrated usefulness for service providers

Abstract

The proliferation of the Open Archive Initiative Protocol for Metadata Harvesting (OAI-PMH) has resulted in the creation of a large number of service providers, all harvesting from either data providers or aggregators. If data were available regarding the similarity of metadata records, service providers could track redundant records across harvests from multiple sources as well as provide additional end-user services. Due to the large number of metadata formats and the diverse mapping strategies employed by data providers, similarity calculation requirements necessitate the use of information retrieval strategies. We describe an OAI-PMH aggregator implementation that uses the optional ``<about>'' container to re-export the results of similarity calculations. Metadata records (3751) were harvested from a NASA data provider and similarities for the records were computed. The results were…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Quality and Management · Advanced Data Storage Technologies · Advanced Database Systems and Queries