Cross Validation for Correlated Data in Regression and Classification   Models, with Applications to Deep Learning

Oren Yuval; Saharon Rosset

arXiv:2502.14808·stat.ME·March 14, 2025

Cross Validation for Correlated Data in Regression and Classification Models, with Applications to Deep Learning

Oren Yuval, Saharon Rosset

PDF

Open Access 1 Repo

TL;DR

This paper introduces a bias-corrected cross-validation method for correlated data in regression and classification, applicable to deep learning, improving model evaluation accuracy in complex data scenarios.

Contribution

It extends bias correction in cross-validation to any model and performance criterion, addressing correlated data issues beyond linear models.

Findings

01

Bias-corrected CV outperforms standard CV in correlated data scenarios.

02

Method applicable to deep neural networks and various prediction criteria.

03

Validated on synthetic and real-world datasets with complex correlation structures.

Abstract

We present a methodology for model evaluation and selection where the sampling mechanism violates the i.i.d. assumption. Our methodology involves a formulation of the bias between the standard Cross-Validation (CV) estimator and the mean generalization error, denoted by $w_{c v}$ , and practical data-based procedures to estimate this term. This concept was introduced in the literature only in the context of a linear model with squared error loss as the criterion for prediction performance. Our proposed bias-corrected CV estimator, $CV_{c} = CV + w_{c v}$ , can be applied to any learning model, including deep neural networks, and to a wide class of criteria for prediction performance in regression and classification tasks. We demonstrate the applicability of the proposed methodology in various scenarios where the data contains complex correlation structures (such as clustered and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

OrenYuv/CVc-for-General-Models
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications