# Linkage-based ortholog refinement in bacterial pangenomes with CLARC

**Authors:** Indra González Ojeda, Samantha G Palace, Pamela P Martinez, Taj Azarian, Lindsay R Grant, Laura L Hammitt, William P Hanage, Marc Lipsitch

PMC · DOI: 10.1093/nar/gkaf488 · 2025-06-20

## TL;DR

CLARC improves bacterial pangenome analysis by refining gene groupings, reducing overestimation of accessory genes and enhancing evolutionary predictions.

## Contribution

CLARC introduces a novel method for refining ortholog groups using functional and linkage data, reducing accessory gene overestimation.

## Key findings

- CLARC reduced accessory gene estimates by over 30% in Streptococcus pneumoniae.
- The method improves evolutionary predictions based on accessory gene frequencies.
- CLARC is broadly applicable across different bacterial species.

## Abstract

Bacterial genomes exhibit significant variation in gene content and sequence identity. Pangenome analyses explore this diversity by classifying genes into core and accessory clusters of orthologous groups (COGs). However, strict sequence identity cutoffs can misclassify divergent alleles as different genes, inflating accessory gene counts. CLARC (Connected Linkage and Alignment Redefinition of COGs) (https://github.com/IndraGonz/CLARC) improves pangenome analyses by condensing accessory COGs using functional annotation and linkage information. Through this approach, orthologous groups are consolidated into more practical units of selection. Analyzing 8000+ Streptococcus pneumoniae genomes, CLARC reduced accessory gene estimates by >30% and improved evolutionary predictions based on accessory gene frequencies. CLARC is effective across different bacterial species, making it a broadly applicable tool for comparative genomics. By refining COG definitions, CLARC offers critical insights into bacterial evolution, aiding genetic studies across diverse populations.

Graphical Abstract

## Linked entities

- **Species:** Streptococcus pneumoniae (taxon 1313)

## Full-text entities

- **Chemicals:** CLARC (-)
- **Species:** Streptococcus pneumoniae (species) [taxon 1313]

## Figures

8 figures with captions in the complete paper: https://tomesphere.com/paper/PMC12204703/full.md

---
Source: https://tomesphere.com/paper/PMC12204703