The Proper Care and Feeding of CAMELS: How Limited Training Data Affects   Streamflow Prediction

Martin Gauch; Juliane Mai; Jimmy Lin

arXiv:1911.07249·cs.LG·November 23, 2020

The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction

Martin Gauch, Juliane Mai, Jimmy Lin

PDF

1 Repo

TL;DR

This study investigates how limited training data impacts the accuracy of tree- and LSTM-based streamflow prediction models, revealing that LSTMs outperform trees with larger datasets, informing data collection strategies.

Contribution

It systematically evaluates the effects of training data size and input length on model accuracy, providing insights into optimal data usage for streamflow prediction models.

Findings

01

LSTMs outperform trees with larger training datasets.

02

Both models perform similarly with small datasets.

03

Additional training data improves prediction accuracy.

Abstract

Accurate streamflow prediction largely relies on historical meteorological records and streamflow measurements. For many regions, however, such data are only scarcely available. Facing this problem, many studies simply trained their machine learning models on the region's available data, leaving possible repercussions of this strategy unclear. In this study, we evaluate the sensitivity of tree- and LSTM-based models to limited training data, both in terms of geographic diversity and different time spans. We feed the models meteorological observations disseminated with the CAMELS dataset, and individually restrict the training period length, number of training basins, and input sequence length. We quantify how additional training data improve predictions and how many previous days of forcings we should feed the models to obtain best predictions for each training set size. Further, our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

gauchm/ealstm_regional_modeling
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.