A study on the impact of fatigue on human raters when scoring speaking responses-Reference-Cited by-同舟云学术

A study on the impact of fatigue on human raters when scoring speaking responses

Published:2014-05-06 Issue:4 Volume:31 Page:479-499
ISSN:0265-5322
Container-title:Language Testing
language:en
Short-container-title:Language Testing

Author:

Ling Guangming¹,Mollaun Pamela¹,Xi Xiaoming¹

Affiliation:

1. Educational Testing Service, USA

Abstract

The scoring of constructed responses may introduce construct-irrelevant factors to a test score and affect its validity and fairness. Fatigue is one of the factors that could negatively affect human performance in general, yet little is known about its effects on a human rater’s scoring quality on constructed responses. In this study, we compared the scoring quality of 72 raters under four shift conditions differing on the shift length (total scoring time in a day) and session length (time continuously spent on a task). About 14,000 audio responses to four TOEFL iBT speaking tasks were scored, including 5446 validity responses that have pre-assigned “true” scores used to measure scoring accuracy. Our results suggest that the overall scoring accuracy is high for the TOEFL iBT Speaking Test, but varying levels of rating accuracy and consistency exist across shift conditions. The raters working the shorter shifts or shorter sessions on average maintain greater rating productivity, accuracy, and consistency than those working longer shifts or sessions do. The raters working the 6-hour shift with three 2-hour sessions outperform those under other shift conditions in both rating accuracy and consistency.

Publisher

SAGE Publications

Subject

Linguistics and Language,Social Sciences (miscellaneous),Language and Linguistics

Link

http://journals.sagepub.com/doi/pdf/10.1177/0265532214530699

Reference28 articles.

1. Perceived fatigue during physical work: an experimental evaluation of a fatigue inventory

2. Rater reliability and "judgmental fatigue."

Cited by 15 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Advancing Consistency in Education: A Reliability Analysis of the Clinical Reasoning Assessment Tool;Journal of Physical Therapy Education;2024-08-08

2. Short answer scoring with GPT-4;Proceedings of the Eleventh ACM Conference on Learning @ Scale;2024-07-09

3. Assessing pronunciation using dictation tools;Journal of Second Language Pronunciation;2024-06-06

4. Inter-rater reliability of the Silverman and Andersen index-a measure of respiratory distress in preterm infants;PLOS ONE;2023-06-30

5. Modeling Rating Order Effects Under Item Response Theory Models for Rater-Mediated Assessments;Applied Psychological Measurement;2023-05-13