Transfer Learning in Biomedical Natural Language Processing: An   Evaluation of BERT and ELMo on Ten Benchmarking Datasets

Yifan Peng; Shankai Yan; Zhiyong Lu

arXiv:1906.05474·cs.CL·June 19, 2019·71 cites

Transfer Learning in Biomedical Natural Language Processing: An Evaluation of BERT and ELMo on Ten Benchmarking Datasets

Yifan Peng, Shankai Yan, Zhiyong Lu

PDF

Open Access 4 Repos 1 Models

TL;DR

This paper introduces the BLUE benchmark for biomedical NLP, evaluates BERT and ELMo models on it, and finds that domain-specific BERT models outperform others across various biomedical and clinical tasks.

Contribution

The paper creates a new benchmark for biomedical NLP and systematically evaluates popular language models, highlighting the effectiveness of domain-specific pre-training.

Findings

01

BERT models trained on biomedical data outperform general models.

02

The BLUE benchmark covers diverse biomedical and clinical NLP tasks.

03

Pre-trained models and datasets are publicly available for research.

Abstract

Inspired by the success of the General Language Understanding Evaluation benchmark, we introduce the Biomedical Language Understanding Evaluation (BLUE) benchmark to facilitate research in the development of pre-training language representations in the biomedicine domain. The benchmark consists of five tasks with ten datasets that cover both biomedical and clinical texts with different dataset sizes and difficulties. We also evaluate several baselines based on BERT and ELMo and find that the BERT model pre-trained on PubMed abstracts and MIMIC-III clinical notes achieves the best results. We make the datasets, pre-trained models, and codes publicly available at https://github.com/ncbi-nlp/BLUE_Benchmark.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Models

🤗
tsantos/PathologyBERT
model· 831 dl· ♡ 7
831 dl♡ 7

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Biomedical Text Mining and Ontologies