Fusing heterogeneous data sets

Yipeng Song

arXiv:1908.09653·q-bio.GN·August 27, 2019

Fusing heterogeneous data sets

Yipeng Song

PDF

Open Access

TL;DR

This paper presents statistical methods for fusing heterogeneous biological data sets that differ in data type and measurement scale, improving integration in systems biology research.

Contribution

It introduces novel statistical techniques specifically designed to handle heterogeneity in data types and scales in biological data fusion.

Findings

01

Methods outperform existing approaches in simulations

02

Effective integration demonstrated on real biological data

03

Enhanced understanding of biological systems through data fusion

Abstract

In systems biology, it is common to measure biochemical entities at different levels of the same biological system. One of the central problems for the data fusion of such data sets is the heterogeneity of the data. This thesis discusses two types of heterogeneity. The first one is the type of data, such as metabolomics, proteomics and RNAseq data in genomics. These different omics data reflect the properties of the studied biological system from different perspectives. The second one is the type of scale, which indicates the measurements obtained at different scales, such as binary, ordinal, interval and ratio-scaled variables. In this thesis, we developed several statistical methods capable to fuse data sets of these two types of heterogeneity. The advantages of the proposed methods in comparison with other approaches are assessed using comprehensive simulations as well as the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsGene expression and cancer classification · Metabolomics and Mass Spectrometry Studies · Statistical Methods and Inference