PDFMathTranslate: Scientific Document Translation Preserving Layouts
Rongxin Ouyang, Chang Chu, Zhikuang Xin, Xiangyao Ma

TL;DR
PDFMathTranslate is an open-source tool that translates scientific documents while maintaining their original layouts, significantly improving the accessibility and dissemination of scientific knowledge across language barriers.
Contribution
It introduces a novel approach combining large language models and layout detection to accurately translate scientific documents with preserved formatting.
Findings
Achieved high translation precision and layout preservation.
Open-sourced with over 222k downloads.
Enhances accessibility of scientific literature across languages.
Abstract
Language barriers in scientific documents hinder the diffusion and development of science and technologies. However, prior efforts in translating such documents largely overlooked the information in layouts. To bridge the gap, we introduce PDFMathTranslate, the world's first open-source software for translating scientific documents while preserving layouts. Leveraging the most recent advances in large language models and precise layout detection, we contribute to the community with key improvements in precision, flexibility, and efficiency. The work has been open-sourced at https://github.com/byaidu/pdfmathtranslate with more than 222k downloads.
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsBiomedical Text Mining and Ontologies · Natural Language Processing Techniques · Mathematics, Computing, and Information Processing
