Sunday, September 20, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play

September 20, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play

AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play

AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Mutual eye contact is one of the most powerful signals in human social life. It conveys attention, intimacy, trust, and emotional state, and its absence early in development has long been recognized as a clinically meaningful marker of atypical social development, most notably in autism spectrum disorder. Yet the tools clinicians use to measure it remain stubbornly low-tech: trained observers watch recorded sessions and manually code when eye contact begins and ends, a process that is slow, expensive, and vulnerable to human bias and error. Now, a team of researchers from the University of Sydney, the Chinese University of Hong Kong, Shanghai Jiao Tong University, and collaborating institutions has unveiled an artificial intelligence framework that can automatically detect clinically defined eye contact between children and assessors during real, unscripted free play sessions—potentially transforming how social engagement is quantified in clinical assessments.

The study, published in the journal Cognitive Computation, introduces a multi-modal, multi-view deep learning framework called the Multi-Modal Social Behavior Recognition (MSBR) system. Unlike eye-tracking glasses or constrained single-camera setups, the framework relies on four synchronized, wall-mounted high-definition cameras capturing complementary views of a clinical assessment room. Children—103 diagnosed with autism and 33 typically developing controls, all aged 3 to 12 years—completed semi-structured 60-minute social interaction sessions that included a five-minute free play task. During free play, children could choose toys, move around the room, and interact spontaneously with an assessor and, at times, their caregivers. Crucially, no wearable devices were attached to the children, avoiding the calibration drift, facial obstruction, and sensory discomfort that plague head-mounted eye trackers, particularly in children with sensory sensitivities.

Defining what counts as eye contact was a careful clinical exercise. For this study, an eye contact event was defined as the child directing visual attention toward the eyes of the co-located assessor while the assessor simultaneously maintained eye contact with the child. The researchers stress that this is a video-observed behavioral event, not a measurement of gaze vectors, ocular fixation, or gaze angle. Two trained coders with domain knowledge annotated the start and end times of each mutual eye contact event using the MATLAB Video Labeler, achieving moderate to high inter-rater reliability, with mean Cohen’s Kappa values of 0.72 in the typically developing group and 0.84 in the autism group. Ambiguous cases—such as distinguishing genuine mutual eye contact from general face-looking—were reviewed and resolved with lead clinical researchers according to strict coding criteria.

The scale of the annotated dataset reflects the rarity of the target behavior. Across 136 videos totaling 733 minutes of recording, mutual eye contact events were identified in 77 autistic children and 25 typically developing children, totaling just 567 seconds of cumulative eye contact. The framework’s developers trimmed the videos into clips labeled either “Eye Contact” or “Other,” yielding 1,726 eye contact clips and 2,435 non-eye-contact clips. Each clip was then decoded into two complementary visual representations: RGB frames capturing static appearance, and optical flow images capturing motion patterns between consecutive frames. This dual representation feeds the framework’s two branches—a spatial-domain branch processing single RGB frames and a temporal-domain branch processing sequences of optical flow images—both built on the ResNet-50 convolutional backbone within a Temporal Segment Network architecture.

Training the model on such imbalanced data required careful engineering. Because most of the free play time did not involve eye contact, the researchers used randomly sampled, fixed-length sequences from long clips, increasing exposure to positive eye contact events while preventing the model from being dominated by lengthy “Other” segments. Data augmentation techniques, including MultiScaleCrop for the RGB branch and RandomResizedCrop for the optical flow branch, were applied only to training data, with all images resized to 224 by 224 pixels and horizontally flipped with 50 percent probability. The model was trained in PyTorch on two NVIDIA GeForce RTX 2080Ti GPUs using Adam optimization with a cosine annealing learning rate scheduler, for up to 300 epochs with categorical cross-entropy loss.

A methodological hallmark of the study is its rigorous participant-independent evaluation. The researchers used five-fold stratified cross-validation with participant-level partitioning: within each fold, 60 percent of unique participants were allocated to training, 20 percent to validation, and 20 percent to held-out testing, ensuring that no participant’s data appeared in more than one partition. Autistic and typically developing children were stratified across partitions to maintain group representation. Hyperparameters were tuned exclusively on validation sets, and held-out test data were reserved for final evaluation, providing a robust measure of how the model performs on children it has never seen.

The results were striking. A fused single-view model achieved F1 scores of 0.85 for eye contact behavior and 0.90 for non-eye-contact behavior. But the multi-view fusion of predictions from all four cameras pushed performance to an F1 score of 0.92 for eye contact and 0.94 for “Other” behaviors, with a Top-1 accuracy of 0.94 for the best fused configuration. Paired fold-level comparisons confirmed that multi-view prediction significantly outperformed single-view prediction across all modalities and diagnostic groups, with mean paired F1 improvements of 0.065 for eye contact and 0.044 for non-eye-contact behaviors and very large fold-level effect sizes. Notably, no statistically significant performance differences emerged between the autism and typically developing groups, and the spatial-domain RGB models outperformed temporal optical flow models—an outcome the authors attribute to the transient, small-amplitude nature of eye contact movements, which makes motion features difficult to extract.

Interpretability received dedicated attention. Using Gradient-weighted Class Activation Mapping, or Grad-CAM, the team generated heatmaps showing which image regions drove the model’s classifications. In the deeper convolutional layers, activation concentrated around faces, upper bodies, and interaction-relevant areas such as toys being manipulated—precisely the visual information clinicians and trained coders rely on when distinguishing mutual eye contact from its absence. The authors caution, however, that these visualizations are qualitative and do not demonstrate that the model estimates gaze vectors directly; they indicate associations, not causal explanations of the model’s decision process.

The study’s limitations are candidly acknowledged. The framework was trained and evaluated within a specific four-camera clinical setup and a particular age range, and external validation across independent sites has not yet been performed. The system recognizes video-observed behavioral events rather than measuring physiological gaze, so future work should validate these clinically defined behaviors against eye-tracking or gaze-estimation methods. The substantial manual annotation burden remains a challenge, motivating future development of weakly supervised or action localization approaches. Privacy, secure data governance, and deployment feasibility also demand attention, though the models’ relatively modest footprint—about 23.5 million parameters and 89.7 MiB of FP32 storage per modality—suggests that deployment on GPU-enabled workstations or edge-computing platforms may be feasible, and privacy-preserving frameworks using de-identified features offer a promising direction.

Even with these caveats, the implications are considerable. An objective, video-derived behavioral marker of mutual eye contact could relieve clinicians of hours of manual video coding, reduce subjective bias, and enable scalable, quantitative assessment of social engagement consistent with established clinical criteria. Such a tool could support more detailed behavioral characterization in autism research and inform future clinical decision-support systems, while preserving the ecological richness of spontaneous multi-person interaction—toy play, variable body orientation, free movement, and unscripted engagement—that constrained laboratory paradigms sacrifice. The researchers envision extending the framework to additional behavioral modalities, including gesture dynamics, body movement, speech features, and audio-visual interaction patterns, moving toward a comprehensive digital characterization of children’s social behavior. If validated across independent clinical sites and integrated with complementary gaze-measurement technologies, multi-view AI frameworks of this kind could become a routine component of developmental assessment, turning ordinary clinical video into rigorous, quantifiable evidence about how children connect with the people around them.

Subject of Research: Automated multi-view deep learning recognition of clinically defined eye contact between children and assessors during free play autism assessments

Article Title: Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors

Article References: Sun, C., Guastella, A. J., Ouyang, W., Zhou, L., Thapa, R., Thomas, E. E., Zhao, H., & McEwan, A. (2026). Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors. Cognitive Computation, 18(1), Article 110. https://doi.org/10.1007/s12559-026-10659-7

Image Credits: AI Generated

DOI: 10.1007/s12559-026-10659-7

Keywords: eye contact, autism spectrum disorder, deep learning, multi-view video, free play assessment, social behavior recognition, computer vision, Grad-CAM, behavioral coding, clinical assessment, optical flow, ResNet-50

Cite Scienmag News

Blake Davidson. (September 20, 2026). AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play. Scienmag. https://scienmag.com/ai-learns-to-spot-eye-contact-between-autistic-children-and-clinicians-during-free-play/

Blake Davidson. "AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play." Scienmag, 20 September 2026, https://scienmag.com/ai-learns-to-spot-eye-contact-between-autistic-children-and-clinicians-during-free-play/. Accessed 20 September 2026.

Blake Davidson. "AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play." Scienmag. September 20, 2026. https://scienmag.com/ai-learns-to-spot-eye-contact-between-autistic-children-and-clinicians-during-free-play/

Tags: advanced AI frameworks for social behavior analysisAI tools for measuring eye contact in childrenAI-based eye contact detection in autism assessmentautism spectrum disorderautomated social engagement analysis through deep learningautomatic coding of eye contact during free playbehavioral codingclinical assessmentcomputer visioncomputer vision in autism diagnosisdeep learningenhancing clinical autism assessments with AI technologyeye contactfree play assessmentGrad-CAMmachine learning for autism social behavior markersmulti-modal video analysis for autism researchmulti-view videomulti-view video analysis for autism diagnosticsoptical flowreal-time eye contact recognition in clinical settingsResNet-50social behavior recognitionunobtrusive assessment of social interactions in autism
Share26Tweet16
Previous Post

Oral Semaglutide Shows Real-World Heart Protection in Diabetes Patients with Cardiovascular Disease

Next Post

Action Research Helps Hospitals Build Care Pathways That Bend Without Breaking

Related Posts

Microplastics Amplify the Deadly Toll of Ozone and Heat on Bumblebees
Technology and Engineering

Microplastics Amplify the Deadly Toll of Ozone and Heat on Bumblebees

September 20, 2026
A Century of Electrochemistry Reshapes How Metals Are Made
Technology and Engineering

A Century of Electrochemistry Reshapes How Metals Are Made

September 20, 2026
Extrusion Doubles Strength of Heat-Resistant Aluminum-Cerium-Magnesium Alloy
Technology and Engineering

Extrusion Doubles Strength of Heat-Resistant Aluminum-Cerium-Magnesium Alloy

September 20, 2026
Nanoparticle Electrode and Machine Learning Team Up to Catch Toxic Lead and Cadmium in Water
Technology and Engineering

Nanoparticle Electrode and Machine Learning Team Up to Catch Toxic Lead and Cadmium in Water

September 20, 2026
Simple Fix Guarantees Connected Graphs for Robust Text Clustering
Technology and Engineering

Simple Fix Guarantees Connected Graphs for Robust Text Clustering

September 20, 2026
Ring Shear Tests Reveal How Concrete and Cement-Treated Soil Grip Together in Composite Piles
Technology and Engineering

Ring Shear Tests Reveal How Concrete and Cement-Treated Soil Grip Together in Composite Piles

September 20, 2026
Next Post
Action Research Helps Hospitals Build Care Pathways That Bend Without Breaking

Action Research Helps Hospitals Build Care Pathways That Bend Without Breaking

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Microplastics Amplify the Deadly Toll of Ozone and Heat on Bumblebees
  • Experts Reach Consensus on 85 Essential Topics for Eating Disorder Training
  • Hidden Enzyme KMT9 Helps Prostate Tumors Evade Immune Attack
  • Action Research Helps Hospitals Build Care Pathways That Bend Without Breaking

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading