PTHash: Revisiting FCH Minimal Perfect Hashing

Giulio Ermanno Pibiri; Roberto Trani

arXiv:2104.10402·cs.DS·February 8, 2022

PTHash: Revisiting FCH Minimal Perfect Hashing

Giulio Ermanno Pibiri, Roberto Trani

PDF

2 Repos

TL;DR

This paper revisits the FCH minimal perfect hashing algorithm, improving its scalability and space efficiency while maintaining fast lookup times, and demonstrates its competitiveness through extensive experiments.

Contribution

The authors present an improved FCH-based algorithm that scales well, reduces space, and offers faster lookup times compared to previous methods.

Findings

01

Achieves 2-4x faster lookup times

02

Reduces space consumption to be competitive with state-of-the-art

03

Scales efficiently to large key sets

Abstract

Given a set $S$ of $n$ distinct keys, a function $f$ that bijectively maps the keys of $S$ into the range ${0, \dots, n - 1}$ is called a minimal perfect hash function for $S$ . Algorithms that find such functions when $n$ is large and retain constant evaluation time are of practical interest; for instance, search engines and databases typically use minimal perfect hash functions to quickly assign identifiers to static sets of variable-length keys such as strings. The challenge is to design an algorithm which is efficient in three different aspects: time to find $f$ (construction time), time to evaluate $f$ on a key of $S$ (lookup time), and space of representation for $f$ . Several algorithms have been proposed to trade-off between these aspects. In 1992, Fox, Chen, and Heath (FCH) presented an algorithm at SIGIR providing very fast lookup evaluation. However, the approach received little…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.