Large and Small Deviations for Statistical Sequence Matching

Lin Zhou; Qianyun Wang; Jingjing Wang; Lin Bai; Alfred O. Hero

arXiv:2407.02816·cs.IT·July 4, 2024

Large and Small Deviations for Statistical Sequence Matching

Lin Zhou, Qianyun Wang, Jingjing Wang, Lin Bai, Alfred O. Hero

PDF

Open Access

TL;DR

This paper analyzes the statistical sequence matching problem, providing theoretical performance guarantees for the generalized likelihood ratio test (GLRT) and demonstrating its optimality in large and small deviations regimes, with extensions to unknown match counts.

Contribution

It offers a comprehensive theoretical analysis of GLRT for sequence matching, including optimality proofs and extensions to unknown match scenarios, improving upon prior results.

Findings

01

Explicit characterization of mismatch and false reject tradeoffs

02

Proof of GLRT optimality under Neyman-Pearson criterion

03

Strengthening of previous results for single-sequence databases

Abstract

We revisit the problem of statistical sequence matching between two databases of sequences initiated by Unnikrishnan (TIT 2015) and derive theoretical performance guarantees for the generalized likelihood ratio test (GLRT). We first consider the case where the number of matched pairs of sequences between the databases is known. In this case, the task is to accurately find the matched pairs of sequences among all possible matches between the sequences in the two databases. We analyze the performance of the GLRT by Unnikrishnan and explicitly characterize the tradeoff between the mismatch and false reject probabilities under each hypothesis in both large and small deviations regimes. Furthermore, we demonstrate the optimality of Unnikrishnan's GLRT test under the generalized Neyman-Person criterion for both regimes and illustrate our theoretical results via numerical examples.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTime Series Analysis and Forecasting