ABSTRACT
Background and Aim: Lameness is a major health, welfare, and production concern in dairy cattle, whereas conventional visual locomotion scoring is intermittent and susceptible to observer-dependent variation. Markerless computer vision may enable objective and repeatable locomotion monitoring under routine farm conditions. This study aimed to develop a 38-keypoint whole-body pose-estimation framework for automated dairy cattle lameness assessment and evaluate its agreement with blinded Sprecher locomotion scoring.
Materials and Methods: A You Only Look Once version 8 (YOLOv8)-pose framework was developed using 12,000 annotated images obtained from 449 Holstein cows. Thirty-eight anatomical keypoints represented the head, spine, pelvis, forelimbs, hindlimbs, and hooves. Four model and input resolution configurations were evaluated. Joint-angle time series, coefficients of variation, and whole-body asymmetry features were derived from detected keypoints. For independent clinical validation, 120 cattle videos, obtained from animals with no overlap at the animal level with the 449-cow model-development cohort, were evaluated blindly by three assessors using the five-point Sprecher scale. Inter-observer reliability, consensus scores, rank associations, group differences, and exploratory receiver operating characteristic analyses were determined.
Results: YOLOv8l-pose at 2560 × 2560 pixels showed the best performance, achieving pose mean average precision at an intersection over union threshold of 0.5 of 83.11%, mean average precision across thresholds of 0.5–0.95 of 37.39%, and recall of 91.74%. All three assessors assigned identical Sprecher scores to 87.5% of videos. Pairwise quadratic-weighted Cohen κ ranged from 0.843 to 0.926, Fleiss κ was 0.779, and intraclass correlation coefficients were 0.883 and 0.958 for single and average measurements, respectively. The continuous artificial intelligence (AI) score correlated significantly with consensus Sprecher scores (Spearman ρ = 0.234, p = 0.010), and score distributions differed among consensus groups (Kruskal–Wallis H = 13.09, p = 0.0014). Exploratory discrimination of consensus Sprecher scores ≥3, using a threshold internally derived from and evaluated within the same stratified validation cohort (not externally validated), yielded an area under the curve of 0.875; at an AI threshold of 2.9, sensitivity and specificity were 62.5% and 98.2%, respectively.
Conclusion: The 38-keypoint whole-body model demonstrates technical feasibility as an anatomically comprehensive, sensor-free approach for quantifying dairy cattle locomotion. Because the clinical validation cohort was deliberately stratified and diagnostic thresholds were internally derived, its agreement with independent visual assessment should be interpreted as evidence of potential decision-support utility rather than established clinical performance. Prospective external validation and independent calibration of diagnostic thresholds in an unselected cohort are warranted before clinical or commercial deployment.
Keywords: artificial intelligence, bovine lameness, computer vision, dairy cattle, locomotion assessment, pose-estimation, Sprecher score, whole-body keypoints.