jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers

Florian H\"onicke; Michael G\"unther; Andreas Koukounas; Mohammad Kalim Akram; Scott Martens; Saba Sturua; Han Xiao

arXiv:2605.08384·cs.CL·May 13, 2026

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers

Florian H\"onicke, Michael G\"unther, Andreas Koukounas, Mohammad Kalim Akram, Scott Martens, Saba Sturua, Han Xiao

PDF

26 Models

TL;DR

GELATO introduces a multimodal embedding model that efficiently encodes text, images, audio, and video into a unified semantic space by freezing backbone models and training only minimal connecting components.

Contribution

The paper presents GELATO, a novel multimodal embedding approach that extends existing models with minimal training, achieving state-of-the-art performance across multiple modalities.

Findings

01

GELATO produces competitive results with larger models.

02

Training only 0.35% of total weights reduces computational cost.

03

The model maintains text embedding quality identical to prior models.

Abstract

In this work, we introduce GELATO (Geometry-preserving Embeddings via Locked Aligned TOwers), a novel approach to multimodal embedding models. We build on the VLM-style architecture, in which non-text encoders are adapted to produce input for a language model, which in turn generates embeddings for all varieties of input. We present the result: the jina-embeddings-v5-omni suite, a pair of models that encode text, image, audio, and video input into a single semantic embedding space. GELATO extends the two Jina Embeddings v5 Text models to support additional modality by adding encoders for images and audio. The backbone text embedding models and the added non-text modality encoders remain frozen. We only trained the connecting components, representing 0.35% of the total weights of the joint model. Training is therefore much more efficient than full-parameter retraining. Additionally, the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.