Estimating regression errors without ground truth values

Henri Tiittanen; Emilia Oikarinen; Andreas Henelius; Kai Puolam\"aki

arXiv:1910.04069·stat.ML·October 10, 2019·1 cites

Estimating regression errors without ground truth values

Henri Tiittanen, Emilia Oikarinen, Andreas Henelius, Kai Puolam\"aki

PDF

Open Access

TL;DR

This paper introduces a framework for estimating the generalization error of regression models without needing ground truth, aiding in detecting overfitting and concept drift in real-world datasets.

Contribution

The paper presents a novel, theoretically derived framework for estimating regression errors without ground truth, applicable across various regression models.

Findings

01

Framework performs robustly in real-world datasets

02

Effective in detecting concept drift

03

Applicable to any family of regression functions

Abstract

Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications we do not know the true value of the outcome variable being predicted outside the training data, i.e., the ground truth is unknown. It is hence not straightforward to directly observe when the estimate from a model potentially is wrong, due to phenomena such as overfitting and concept drift. In this paper we present an efficient framework for estimating the generalization error of regression functions, applicable to any family of regression functions when the ground truth is unknown. We present a theoretical derivation of the framework and empirically evaluate its strengths and limitations. We find that it performs robustly and is useful for detecting concept drift in datasets in several real-world domains.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Stream Mining Techniques · Anomaly Detection Techniques and Applications · Stock Market Forecasting Methods