Artificial intelligence has quietly taken over the modern pig barn. Cameras hang from ceilings, watching around the clock as deep learning models count animals, track their movements, detect tail-biting outbreaks, and even predict when a sow is about to give birth. What began as a niche research curiosity has grown into a substantial scientific field, and a new systematic review published in Smart Agricultural Technology has now mapped that landscape in unprecedented detail. The study, led by Hassan-Roland Nasser with Claudia Kasper and Vladimir Živković, screened more than a thousand publications and subjected 131 of them to rigorous quality appraisal, producing the first pig-focused systematic synthesis of deep learning based two-dimensional computer vision in precision livestock farming.
The technical foundations of the field rest on convolutional neural networks, the deep learning architectures that transformed computer vision by learning hierarchical feature representations directly from image data. Early layers of these networks capture simple patterns such as edges and textures, while deeper layers encode abstract concepts like objects and behaviors. Most systems deployed in pig production follow a modular design: a pretrained backbone network, often a ResNet, VGG, or EfficientNet model originally trained on large-scale datasets such as ImageNet, serves as a feature extractor, while task-specific prediction heads perform the actual work of detection, classification, or segmentation. This transfer learning approach is crucial in agriculture, where annotated data are expensive and scarce. More recently, transformer-based architectures such as Vision Transformers have begun to appear, prized for their ability to capture long-range spatial dependencies in the crowded, visually chaotic environment of a commercial pen.
The review organizes the field’s outputs into what the authors call visiotypes: algorithmically extracted visual variables that serve as proxies for an animal’s biological state. These range from perception-level outputs like detection, tracking, and counting to higher-level inferences such as behavior recognition, estrus detection, and weight estimation. Behavior recognition emerged as the single most frequently reported visiotype, appearing in 41 studies, followed by posture recognition with 26 and object detection with 21. The dominant application is activity monitoring, addressed by 74 studies, followed by welfare assessment with 39. Breeding and reproduction, productivity assessment, and health status applications trail considerably behind, suggesting that the field has concentrated on what cameras see most easily rather than on what matters most economically.
On the modeling side, the YOLO family of detectors dominates, alongside Faster R-CNN and SSD, prized for their balance of accuracy, inference speed, and robustness to the occlusion and variable illumination that plague barn environments. For pixel-level analysis, Mask R-CNN and its variants are the workhorses, enabling researchers to separate overlapping animals and extract specific body regions such as heads, tails, and ears. Individual identification is typically handled by CNN classifiers trained on pig faces or heads, with lightweight architectures frequently exceeding 95 percent accuracy. For behaviors that unfold over time, such as aggression or farrowing, researchers pair frame-level feature extractors with temporal models, most commonly Long Short-Term Memory networks, though three-dimensional CNNs and spatiotemporal graph convolutional networks are gaining ground.
Reported performance figures are strikingly high. Median precision across classification studies stands at 95.2 percent, median accuracy at 95.3 percent, and median mAP50 for detection and segmentation at 95.0 percent. Tracking systems report a median MOTA of 93.5. Yet the review’s most sobering finding concerns how these numbers were obtained. Only 17 of the 131 studies, a mere 13 percent, reported external or cross-farm validation, and 93 studies drew all their data from a single source. The authors judged 45 studies, roughly a third of the total, to carry a high risk of data leakage, most often because continuous video was split randomly at the frame level, placing near-duplicate images in both training and test sets. The glittering benchmark scores, in other words, were typically earned under within-dataset, single-site conditions that may say little about real-world performance.
The environmental challenges documented across the studies reinforce this concern. Occlusion was by far the most frequently cited obstacle, mentioned 80 times, followed by lighting variability with 39 mentions and behavioral complexity with 36. Overlapping animals and pen-environment constraints each drew 23 mentions. These are precisely the conditions that define commercial pig housing: dozens of visually similar animals jostling in dim, dusty pens under artificial light. Pose estimation studies carried the highest leakage risk of all, at 60 percent, a consequence of the extreme labor cost of keypoint annotation, which forces researchers to work with small datasets drawn from few animals and to sample densely from continuous footage.
Reproducibility emerges as another structural weakness. Only 0.7 percent of the reviewed papers provided both a public dataset and source code, about 5.2 percent shared code alone, and 7.5 percent released datasets. The overwhelming majority offered neither, limiting independent validation, benchmarking, and cumulative progress. Annotation practices varied widely in granularity and behavioral definitions, often without reported inter-annotator agreement, further complicating comparison across studies. The authors argue that openly accessible, collaborative, multi-site benchmark datasets with harmonized behavioral ontologies are the enabling foundation on which everything else in the field depends.
There are also notable gaps in what the field chooses to study. Adult growing-finishing pigs are markedly over-represented, while piglets and sows receive far less attention, despite the direct relevance of early life stages and reproduction to piglet survival and farm economics. Postural behaviors dominate the behavioral literature, likely because they are easy to annotate and visually distinct, while aggressive and damaging behaviors, though highly relevant to welfare, remain underrepresented. Methodologically, transformer-based architectures account for less than 15 percent of reported models, and not a single surveyed study employed vision-language foundation models, leaving a conceptual rather than sensor-level gap that the authors identify as a promising frontier for zero-shot behavioral interpretation and flexible, query-based monitoring.
The review’s authors frame progress as three interdependent layers: enabling conditions such as shared benchmarks and leakage-aware validation, methodological advances including transformers and self-supervised learning, and a translation layer covering real-time edge deployment, maintenance, and cost-benefit analysis. Indicators of practical feasibility remain scarce, with inference speed reported by only a minority of studies at a median of roughly 29 frames per second, and computational complexity, embedded-hardware testing, and economic analysis reported even less often. Until the field confronts its validation weaknesses and builds the shared infrastructure to compare methods honestly, the authors conclude, technically impressive prototypes will keep falling short of the robust, field-ready systems that precision livestock farming ultimately demands.
Subject of Research: Deep learning-based 2D computer vision applications for pig monitoring in precision livestock farming
Article Title: Deep learning based 2D computer vision in precision livestock farming for pigs: A systematic literature review and future research directions
Article References: Nasser, H.-R., Kasper, C., & Živković, V. (2026). Deep learning based 2D computer vision in precision livestock farming for pigs: A systematic literature review and future research directions. Smart Agricultural Technology, 15, Article 102426. https://doi.org/10.1016/j.atech.2026.102426
Image Credits: AI Generated
DOI: 10.1016/j.atech.2026.102426
Keywords: deep learning, computer vision, precision livestock farming, pig welfare, behavior recognition, object detection, YOLO, systematic review, animal monitoring, data leakage, transformers, agricultural technology
Cite Scienmag News
Alan Morgan. (October 3, 2026). AI Cameras Are Watching Pigs Like Never Before, But a Massive New Review Reveals the Field’s Hidden Flaws. Scienmag. https://scienmag.com/ai-cameras-are-watching-pigs-like-never-before-but-a-massive-new-review-reveals-the-fields-hidden-flaws/
Alan Morgan. "AI Cameras Are Watching Pigs Like Never Before, But a Massive New Review Reveals the Field’s Hidden Flaws." Scienmag, 3 October 2026, https://scienmag.com/ai-cameras-are-watching-pigs-like-never-before-but-a-massive-new-review-reveals-the-fields-hidden-flaws/. Accessed 3 October 2026.
Alan Morgan. "AI Cameras Are Watching Pigs Like Never Before, But a Massive New Review Reveals the Field’s Hidden Flaws." Scienmag. October 3, 2026. https://scienmag.com/ai-cameras-are-watching-pigs-like-never-before-but-a-massive-new-review-reveals-the-fields-hidden-flaws/

