Skip to main navigation Skip to search Skip to main content

Toward Multimodal AI for Livestock: Vision-Language Modelling for Non-Invasive Cattle Behaviour Analysis in Smart Barns

Research output: Contribution to conferencePaperpeer-review

Abstract

Cattle behaviour monitoring is an important area in precision livestock farming, as it directly influences animal welfare, reproductive health, and early disease detection. Traditional manual approaches, such as manual observation or vision-only systems, are labourintensive and often struggle to capture diverse behaviours under complex barn settings. In this work, we presented a non-invasive framework for analysing cattle behaviours using a vision-language model. The framework is based on BLIP (Bootstrapping Language-Image Pretraining), which was trained on our cattle behaviour dataset. By jointly leveraging visual features from a Vision Transformer and textual embeddings, BLIP enables accurate classification of common behaviours, including lying, standing, resting, eating, drinking and contraction. The model generates both individual and group-level cattle activities for individual image frames before predictions are temporally smoothed across video sequences. The results highlight the potential of vision-language modelling as a reliable, informative, and non-invasive solution for livestock monitoring, achieving an accuracy of 92.4 % for behaviour classification and a BLEU-4 score of 0.64 for caption generation. This illustrated its contribution to improved welfare assessment and the advancement of innovative barn technologies.
Original languageEnglish
Pages0279-0288
Number of pages10
DOIs
Publication statusPublished - 22 Oct 2025
EventIEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON) -
Duration: 22 Oct 202524 Oct 2025

Conference

ConferenceIEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON)
Period22/10/2524/10/25

Fingerprint

Dive into the research topics of 'Toward Multimodal AI for Livestock: Vision-Language Modelling for Non-Invasive Cattle Behaviour Analysis in Smart Barns'. Together they form a unique fingerprint.

Cite this