# Concept annotation in the CRAFT corpus

**Authors:** Michael Bada, Miriam Eckert, Donald Evans, Kristin Garcia, Krista Shipley, Dmitry Sitnikov, William A Baumgartner, K Bretonnel Cohen, Karin Verspoor, Judith A Blake, Lawrence E Hunter

PMC · DOI: 10.1186/1471-2105-13-161 · BMC Bioinformatics · 2012-07-09

## TL;DR

The CRAFT Corpus is a large, annotated biomedical text collection that helps improve automated methods for identifying concepts in scientific literature.

## Contribution

The paper introduces a richly annotated corpus with concept annotations from multiple biomedical ontologies, supporting NLP research.

## Key findings

- The CRAFT Corpus includes annotations for over 560,000 tokens in its initial release.
- It features concept annotations from nine biomedical ontologies and terminologies.
- The corpus is freely available and supports diverse biomedical disciplines.

## Abstract

Manually annotated corpora are critical for the training and evaluation of automated methods to identify concepts in biomedical text.

This paper presents the concept annotations of the Colorado Richly Annotated Full-Text (CRAFT) Corpus, a collection of 97 full-length, open-access biomedical journal articles that have been annotated both semantically and syntactically to serve as a research resource for the biomedical natural-language-processing (NLP) community. CRAFT identifies all mentions of nearly all concepts from nine prominent biomedical ontologies and terminologies: the Cell Type Ontology, the Chemical Entities of Biological Interest ontology, the NCBI Taxonomy, the Protein Ontology, the Sequence Ontology, the entries of the Entrez Gene database, and the three subontologies of the Gene Ontology. The first public release includes the annotations for 67 of the 97 articles, reserving two sets of 15 articles for future text-mining competitions (after which these too will be released). Concept annotations were created based on a single set of guidelines, which has enabled us to achieve consistently high interannotator agreement.

As the initial 67-article release contains more than 560,000 tokens (and the full set more than 790,000 tokens), our corpus is among the largest gold-standard annotated biomedical corpora. Unlike most others, the journal articles that comprise the corpus are drawn from diverse biomedical disciplines and are marked up in their entirety. Additionally, with a concept-annotation count of nearly 100,000 in the 67-article subset (and more than 140,000 in the full collection), the scale of conceptual markup is also among the largest of comparable corpora. The concept annotations of the CRAFT Corpus have the potential to significantly advance biomedical text mining by providing a high-quality gold standard for NLP systems. The corpus, annotation guidelines, and other associated resources are freely available at http://bionlp-corpora.sourceforge.net/CRAFT/index.shtml.

## Full-text entities

- **Genes:** Grid2ip (glutamate receptor, ionotropic, delta 2 (Grid2) interacting protein 1) [NCBI Gene 170935], TPP1 (tripeptidyl peptidase 1) [NCBI Gene 1200] {aka CLN2, GIG1, LPIC, SCAR7, TPP-1}, Suox (sulfite oxidase) [NCBI Gene 211389] {aka SO}, Ntf3 (neurotrophin 3) [NCBI Gene 18205] {aka HDNF, NGF-2, Nt3, Ntf-3}, Ces1e (carboxylesterase 1E) [NCBI Gene 13897] {aka Eg, Es-22, Es22, egasyn}
- **Diseases:** CL (MESH:D002971), diabetes (MESH:D003920), MF (MESH:C567116), SO (MESH:D010855), MGD (MESH:D004482), cancer (MESH:D009369)
- **Chemicals:** dinucleotides (MESH:D015226), dipeptides (MESH:D004151), oligopeptides (MESH:D009842), cation (MESH:D002412), oligonucleotides (MESH:D009841), glutamate (MESH:D018698), peptides (MESH:D010455), serine (MESH:D012694), IAA (-)
- **Species:** Xenopus laevis (African clawed frog, species) [taxon 8355], Mus musculus (house mouse, species) [taxon 10090], Rattus (rat, genus) [taxon 10114], Rattus norvegicus (brown rat, species) [taxon 10116], Homo sapiens (human, species) [taxon 9606]

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/PMC3476437/full.md

## Figures

2 figures with captions in the complete paper: https://tomesphere.com/paper/PMC3476437/full.md

## References

79 references — full list in the complete paper: https://tomesphere.com/paper/PMC3476437/full.md

---
Source: https://tomesphere.com/paper/PMC3476437