Assessing the Performance of a Long Short-Term Memory Algorithm in the Dataset with Missing Values-Reference-Cited by-同舟云学术

Assessing the Performance of a Long Short-Term Memory Algorithm in the Dataset with Missing Values

Published:2022-12-31 Issue:12 Volume:44 Page:636-642
ISSN:1225-5025
Container-title:Journal of Korean Society of Environmental Engineers
language:en
Short-container-title:J Korean Soc Environ Eng

Author:

Park Hyun-Geoun^ORCID,Suh Sang-Ik^ORCID,Jo Gyeong Cheol^ORCID,Jang Jinuk^ORCID,Ki Seo Jin^ORCID

Abstract

This study was conducted to assess the performance of a long short-term memory algorithm (LSTM), which was suitable for time series prediction, in the multivariate dataset with missing values. The full dataset for the adopted LSTM model was prepared by running a popular watershed model Hydrological Simulation Program-Fortran (HSPF) in the upper Nam River Basin for 3 years from 2016 to 2018, excluding a one-year warm-up period, on a daily time step. The accuracy of prediction for the LSTM model was evaluated in response to various interpolation methods as well as changes in the number of missing values (for dependent variables) and independent variables (containing a fixed number of missing values for either single or multiple variables). Note that the entire dataset is divided into training and test datasets at a ratio of 7:3. Results showed that different interpolation methods resulted in a considerable variation in performance of the LSTM model. Out of them, StructTS and RPART were selected as the best imputation methods recovering missing values for discharge and total phosphorus, respectively. The prediction error of the LSTM model increased gradually with increasing the number of missing values from 300 to 700. The LSTM model, however, appeared to maintain its performance fairly well even in data sets with a large amount of missing values as long as adequate interpolation methods were adopted for each dependent variable. The performance of the LSTM model degraded further as the number of independent variables containing the fixed number of missing values increased from 1 to 7. We believe that the proposed methodology can be used not only to reconstruct missing values in a real-time monitoring dataset with excellent performance, but also to improve the accuracy of prediction for (time series) deep learning models.

Funder

Rural Development Administration

Publisher

Korean Society of Environmental Engineering

Subject

General Medicine

Link

http://jksee.or.kr/upload/pdf/KSEE-2022-44-12-636.pdf

Reference16 articles.

1. Estimation of cyanobacteria pigments in the main rivers of South Korea using spatial attention convolutional neural network with hyperspectral imagery

2. Retrieval of water quality parameters from hyperspectral images using a hybrid feedback deep factorization machine model

3. Drought prediction till 2100 under RCP 8.5 climate change scenarios for Korea

4. Long-term relationship between air and water temperatures in Lake Paldang, South Korea

5. Soft detection of 5-day BOD with sparse matrix in city harbor water using deep learning techniques