# PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

**Authors:** John D. Westbrook, Jasmine Y. Young, Chenghua Shao, Zukang Feng, Vladimir Guranovic, Catherine L. Lawson, Brinda Vallat, Paul D. Adams, John M Berrisford, Gerard Bricogne, Kay Diederichs, Robbie P. Joosten, Peter Keller, Nigel W. Moriarty, Oleg V. Sobolev, Sameer Velankar, Clemens Vonrhein, David G. Waterman, Genji Kurisu, Helen M. Berman, Stephen K. Burley, Ezra Peisach

PMC · DOI: 10.1016/j.jmb.2022.167599 · Journal of molecular biology · 2023-06-26

## TL;DR

PDBx/mmCIF is a foundational data standard for structural biology that supports diverse data types and accelerates scientific discovery through semantic tools and FAIR data principles.

## Contribution

The paper introduces the architecture and community-driven tools of PDBx/mmCIF, emphasizing its role in enabling FAIR data delivery in structural biology.

## Key findings

- PDBx/mmCIF is used across multiple platforms for storing and sharing biological macromolecule structures and computational models.
- Community-driven development has made PDBx/mmCIF a semantically rich and extensible framework for structural biology data.
- The PDBx/mmCIF ecosystem supports FAIR data principles, benefiting millions of users globally.

## Abstract

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArc-hive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

## Full-text entities

- **Diseases:** BIRD (MESH:D053591), CCD (MESH:C566443), DDL (MESH:D007806), PDB (MESH:D011488), SARS-Cov-2 (MESH:D000086382)
- **Species:** Human immunodeficiency virus 1 (no rank) [taxon 11676], Zika virus (no rank) [taxon 64320], Homo sapiens (human, species) [taxon 9606]

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/PMC10292674/full.md

## Figures

5 figures with captions in the complete paper: https://tomesphere.com/paper/PMC10292674/full.md

## References

68 references — full list in the complete paper: https://tomesphere.com/paper/PMC10292674/full.md

---
Source: https://tomesphere.com/paper/PMC10292674