MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

Tim Strohmeyer; Lucas Morin; Gerhard Ingmar Meijer; Val\'ery Weber; Ahmed Nassar; Peter Staar

arXiv:2603.28550·cs.CV·March 31, 2026

MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

Tim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer, Val\'ery Weber, Ahmed Nassar, Peter Staar

PDF

1 Repo

TL;DR

MarkushGrapher-2 is an end-to-end multimodal system that accurately recognizes complex chemical Markush structures from documents by integrating OCR, vision, and layout encoding, and is supported by a large dataset and benchmark.

Contribution

It introduces a novel multimodal recognition approach for Markush structures, including a new dataset, benchmark, and a two-stage training strategy for improved accuracy.

Findings

01

Outperforms state-of-the-art models in Markush structure recognition.

02

Effectively fuses text, image, and layout information for chemical structure extraction.

03

Maintains strong performance in molecule recognition tasks.

Abstract

Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to recognize molecules represented either in figures or in text independently. However, methods for recognizing chemical structures from multimodal descriptions (Markush structures) lag behind in precision and cannot be used for automatic large-scale processing. In this work, we present MarkushGrapher-2, an end-to-end approach for the multimodal recognition of chemical structures in documents. First, our method employs a dedicated OCR model to extract text from chemical images. Second, the text, image, and layout information are jointly encoded through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. Finally, the resulting encodings are effectively fused through a two-stage…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

null
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.