Testing the Tests: What Are the Impacts of Incorrect Assumptions When Applying Confidence Intervals or Hypothesis Tests to Compare Competing Forecasts?-Reference-Cited by-同舟云学术

Testing the Tests: What Are the Impacts of Incorrect Assumptions When Applying Confidence Intervals or Hypothesis Tests to Compare Competing Forecasts?

Published:2018-06 Issue:6 Volume:146 Page:1685-1703
ISSN:0027-0644
Container-title:Monthly Weather Review
language:en
Short-container-title:Mon. Wea. Rev.

Author:

Gilleland Eric¹,Hering Amanda S.¹,Fowler Tressa L.¹,Brown Barbara G.¹

Affiliation:

1. Research Applications Laboratory, National Center for Atmospheric Research, Boulder, Colorado

Abstract

Which of two competing continuous forecasts is better? This question is often asked in forecast verification, as well as climate model evaluation. Traditional statistical tests seem to be well suited to the task of providing an answer. However, most such tests do not account for some of the special underlying circumstances that are prevalent in this domain. For example, model output is seldom independent in time, and the models being compared are geared to predicting the same state of the atmosphere, and thus they could be contemporaneously correlated with each other. These types of violations of the assumptions of independence required for most statistical tests can greatly impact the accuracy and power of these tests. Here, this effect is examined on simulated series for many common testing procedures, including two-sample and paired t and normal approximation z tests, the z test with a first-order variance inflation factor applied, and the newer Hering–Genton (HG) test, as well as several bootstrap methods. While it is known how most of these tests will behave in the face of temporal dependence, it is less clear how contemporaneous correlation will affect them. Moreover, it is worthwhile knowing just how badly the tests can fail so that if they are applied, reasonable conclusions can be drawn. It is found that the HG test is the most robust to both temporal dependence and contemporaneous correlation, as well as the specific type and strength of temporal dependence. Bootstrap procedures that account for temporal dependence stand up well to contemporaneous correlation and temporal dependence, but require large sample sizes to be accurate.

Funder

National Science Foundation

Publisher

American Meteorological Society

Subject

Atmospheric Science

Link

http://journals.ametsoc.org/doi/pdf/10.1175/MWR-D-17-0295.1

Reference52 articles.

1. Brockwell, P. J., and R. A. Davis, 2010: Introduction to Time Series and Forecasting. 2nd ed. Springer, 437 pp.

2. Test inversion bootstrap confidence intervals

3. Clark, T. E., and M. W. McCracken, 2013: Advances in forecast evaluation. Handbook of Economic Forecasting, G. Elliott and A. Timmermann, Eds., Vol. 2, Elsevier, 1107–1201.