Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

Hatef Otroshi Shahreza; Anjith George; S\'ebastien Marcel

arXiv:2601.15406·cs.CV·January 23, 2026

Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

Hatef Otroshi Shahreza, Anjith George, S\'ebastien Marcel

PDF

Open Access

TL;DR

This paper systematically evaluates multimodal large language models for heterogeneous face recognition across various spectral modalities, revealing significant performance gaps compared to classical systems and highlighting current limitations.

Contribution

It provides a comprehensive benchmark of state-of-the-art MLLMs for cross-modality face recognition, emphasizing their limitations and the need for rigorous biometric evaluation.

Findings

01

MLLMs underperform classical face recognition in cross-spectral scenarios.

02

Significant performance gaps exist between MLLMs and traditional methods.

03

Current MLLMs are limited for practical heterogeneous face recognition applications.

Abstract

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we conduct a systematic evaluation of state-of-the-art MLLMs for heterogeneous face recognition (HFR), where enrollment and probe images are from different sensing modalities, including visual (VIS), near infrared (NIR), short-wave infrared (SWIR), and thermal camera. We benchmark multiple open-source MLLMs across several cross-modality scenarios, including VIS-NIR, VIS-SWIR, and VIS-THERMAL face recognition. The recognition performance of MLLMs is evaluated using biometric protocols and based on different metrics, including Acquire Rate, Equal Error Rate (EER), and True Accept Rate (TAR). Our results reveal substantial performance gaps between MLLMs and classical face…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsFace recognition and analysis · Face and Expression Recognition · Speech and Audio Processing