Abstract
Context An ability to predict calving difficulty could help farmers make better farm-management decisions, thereby improving dairy farm profitability and welfare. Aims This study aimed to predict calving difficulty in Iranian dairy herds using machine-learning (ML) algorithms and to evaluate sampling methods to deal with imbalanced datasets. Methods For this purpose, the history records of cows that calved between 2011 and 2021 on two commercial dairy farms were used. Using WEKA software, four commonly used ML algorithms, namely naïve Bayes, random forest, decision trees, and logistic regression, were applied to the dataset. The calving difficulty was considered as a binary trait with 0, normal or unassisted calving, and 1, difficult calving, i.e. receiving any help during parturition from farm personnel involvement to surgical intervention. The average rate of difficult calving was 18.7%, representing an imbalanced dataset. Therefore, down-sampling and cost-sensitive techniques were implemented to tackle this problem. Different models were evaluated on the basis of F-measure and the area under the curve. Key results The results showed that sampling techniques improved the predictive model (P = 0.07, and P = 0.03, for down-sampling and cost-sensitive techniques respectively). F-measure ranged from 0.387 (decision tree) to 0.426 (logistic regression) with the balanced dataset. However, when applied to the original imbalanced dataset, naïve Bayes had the best performance of up to 0.388 in terms of F-measure. Conclusions Overall, sampling techniques improved the prediction model compared with original imbalanced dataset. Although prediction models performed worse than expected (due to an imbalanced dataset, and missing values), the implementation of ML algorithms can still lead to an effective method of predicting calving difficulty. Implications This research indicated the capability of ML algorithms to predict the incidence of calving difficulty within a balanced dataset, but that more explanatory variables (e.g. genetic information) are required to improve the prediction based on an unbalanced original dataset.
Subject
Animal Science and Zoology,Food Science
Reference54 articles.
1. Influence of calving ease on in-line milk lactose and other milk components.;Animals,2021
2. Prevalence, risk factors and consequent effect of dystocia in Holstein dairy cows in Iran.;Asian–Australasian Journal of Animal Sciences,2012
3. Baaken D, Hess S (2021) Forecasting regional milk production quantity: a comparison of regression models and machine learning. In ‘2021 Conference, Virtual 315117’, 17–31 August 2021. (International Association of Agricultural Economists)
4. Evaluation measures for models assessment over imbalanced data sets.;Journal of Information Engineering and Applications,2013
5. Boakari YL, Ali HE-S (2021) Management to prevent dystocia. In ‘Bovine reproduction’. (Ed. RM Hopper) pp. 590–596. (John Wiley & Sons, Inc.)
Cited by
2 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献