# Exploiting redundancy in large materials datasets for efficient machine learning with less data

**Authors:** Kangming Li, Daniel Persaud, Kamal Choudhary, Brian DeCost, Michael Greenwood, Jason Hattrick-Simpers

PMC · DOI: 10.1038/s41467-023-42992-y · Nature Communications · 2023-11-10

## TL;DR

This paper shows that most materials data used for machine learning is redundant, and removing up to 95% of it doesn't hurt prediction accuracy.

## Contribution

The study reveals high redundancy in materials datasets and introduces uncertainty-based active learning to build smaller, effective datasets.

## Key findings

- Up to 95% of data in materials datasets can be removed without affecting in-distribution prediction performance.
- Redundant data does not improve performance on out-of-distribution samples.
- Active learning can create smaller datasets that maintain prediction accuracy and robustness.

## Abstract

Extensive efforts to gather materials data have largely overlooked potential data redundancy. In this study, we present evidence of a significant degree of redundancy across multiple large datasets for various material properties, by revealing that up to 95% of data can be safely removed from machine learning training with little impact on in-distribution prediction performance. The redundant data is related to over-represented material types and does not mitigate the severe performance degradation on out-of-distribution samples. In addition, we show that uncertainty-based active learning algorithms can construct much smaller but equally informative datasets. We discuss the effectiveness of informative data in improving prediction performance and robustness and provide insights into efficient data acquisition and machine learning training. This work challenges the “bigger is better” mentality and calls for attention to the information richness of materials data rather than a narrow emphasis on data volume.

Big data is crucial for machine learning, but the redundancies in the datasets are rarely studied. Here the authors reveal significant redundancy in large materials datasets, showing that up to 95% of data can be removed without impacting prediction accuracy.

## Full-text entities

- **Genes:** LIM2 (lens intrinsic membrane protein 2) [NCBI Gene 3982] {aka CTRCT19, MP17, MP19}
- **Cell lines:** MP — Homo sapiens (Human), Induced pluripotent stem cell (CVCL_B5NJ)

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/PMC10638383/full.md

## Figures

6 figures with captions in the complete paper: https://tomesphere.com/paper/PMC10638383/full.md

## References

50 references — full list in the complete paper: https://tomesphere.com/paper/PMC10638383/full.md

---
Source: https://tomesphere.com/paper/PMC10638383