<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-assisted early autism detection &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-assisted-early-autism-detection/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 09:40:51 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-assisted early autism detection &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Reads Parent Interviews to Screen Toddlers for Autism With Near-Expert Accuracy</title>
		<link>https://scienmag.com/ai-reads-parent-interviews-to-screen-toddlers-for-autism-with-near-expert-accuracy/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:40:51 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI-assisted early autism detection]]></category>
		<category><![CDATA[AI-based autism screening for toddlers]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[autism screening]]></category>
		<category><![CDATA[autism spectrum disorder]]></category>
		<category><![CDATA[automated scoring]]></category>
		<category><![CDATA[behavioral development screening for toddlers]]></category>
		<category><![CDATA[BMC Psychiatry]]></category>
		<category><![CDATA[C-BeDevel-I]]></category>
		<category><![CDATA[case-control study]]></category>
		<category><![CDATA[clinical agreement]]></category>
		<category><![CDATA[clinical validation of AI screening methods]]></category>
		<category><![CDATA[conversational AI in pediatric diagnosis]]></category>
		<category><![CDATA[diagnostic accuracy]]></category>
		<category><![CDATA[early childhood development]]></category>
		<category><![CDATA[language model clinical assessment]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in healthcare]]></category>
		<category><![CDATA[machine learning accuracy in early childhood screening]]></category>
		<category><![CDATA[near-expert accuracy AI autism screening tools]]></category>
		<category><![CDATA[parent interview analysis for autism detection]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[structured parent interviews for autism]]></category>
		<category><![CDATA[translation of clinical conversations into diagnostic scores]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234526</guid>

					<description><![CDATA[Researchers report that large language models can score physician-administered parent interviews for autism screening in young children with near-expert agreement and high diagnostic discrimination, while cautioning that prospective real-world trials are still needed.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has taken another step into the clinic, and this time the target is one of the most consequential questions in early childhood medicine: does this toddler have autism? A team of researchers from Nanjing Brain Hospital, the Affiliated Brain Hospital of Nanjing Medical University, Seoul National University Bundang Hospital, and Xuzhou Medical University has shown that large language models, the same class of AI systems behind conversational chatbots, can score a structured parent interview for autism screening with accuracy that approaches, and in some measures nearly matches, that of trained human experts. The study, published in BMC Psychiatry, is among the first to rigorously test whether an AI system can convert the messy, narrative-rich transcript of a real clinical conversation into a reliable screening score for young children.</p>
<p>The tool at the center of the study is the C-BeDevel-I, the Chinese version of the Behavior Development Screening for Toddlers administered as an interview. Unlike brief parent checklists such as the M-CHAT, which rely on yes-or-no questionnaire items, this is a second-level screening instrument: a physician sits down with a parent or primary caregiver and conducts a structured but conversational interview about the child&#8217;s development, social communication, and restricted and repetitive behaviors. The richness of that format is both its strength and its bottleneck. Narrative answers capture nuances that checkboxes miss, but scoring them demands trained professionals, time, and consistency. The researchers asked a deceptively simple question: could a large language model read the transcript of such an interview and produce a score that a clinician would trust?</p>
<p>To answer it, the team assembled three retrospective analytical datasets. The first, Cohort A, included 518 children aged 9 to 42 months and was used to establish the internal consistency and diagnostic validity of the C-BeDevel-I itself as a second-level screening parent interview for autism spectrum disorder. The second, Cohort B, was a carefully selected case-control sample of 256 distinct children, 130 with autism spectrum disorder and 126 with typical development, whose interview recordings were transcribed and then scored twice by three different AI systems: Gemini-2.5, GPT-5-mini, and Qwen-Plus. The third, Cohort C, served as an independent external test. It comprised 60 children, 30 with autism and 30 with typical development, enrolled at the Department of Child Health of Lianyungang First People&#8217;s Hospital, a separate center where the clinical diagnosis had been established before any AI scoring took place.</p>
<p>The technical pipeline behind the experiment is as important as the results. Interview audio was converted to text using automated speech recognition, and the resulting transcripts were passed to the language models through an application programming interface. The workflow incorporated retrieval-augmented generation, a technique that grounds the model&#8217;s judgments in reference material, such as the scoring manual, retrieved at the moment of inference rather than relying solely on what the model absorbed during training. Chain-of-thought prompting was used to encourage the models to reason step by step through each interview item before committing to a score, mirroring the way a trained clinician works through a rating scale. Each model produced a total score for every interview, and the first complete output was treated as the primary result, while the second, independent call assessed how stable the model&#8217;s judgments were when the same transcript was scored again.</p>
<p>The headline numbers are striking. In Cohort B, the area under the receiver operating characteristic curve, a standard measure of how well a test separates cases from controls, reached 0.9961 for Gemini-2.5, with a 95 percent stratified-bootstrap confidence interval of 0.9916 to 0.9991. GPT-5-mini achieved 0.9861 and Qwen-Plus 0.9770. After Holm correction of paired DeLong tests, only the difference between Gemini-2.5 and Qwen-Plus remained statistically significant, an AUC difference of 0.0191 with an adjusted p-value of 0.012, while the other pairwise comparisons fell just short of significance. In practical terms, all three models separated autistic children from typically developing children in this selected case-control sample with near-ceiling discrimination, with Gemini-2.5 statistically ahead of only one competitor.</p>
<p>Discrimination alone, however, is not the same as agreement with human experts, and the researchers measured both. Gemini-2.5&#8217;s scores showed a concordance correlation coefficient of 0.9631 with expert total scores, a mean absolute error of just 0.824 points, and a mean AI-minus-expert difference of +0.184 points, indicating negligible systematic bias. The limits of agreement ranged from −2.235 to 2.602 points on the scale, meaning that in the vast majority of cases the machine landed within a few points of the clinician. When the same transcript was scored twice, the intraclass correlation coefficient for repeated outputs was 0.9875, an unusually high figure for a generative system whose outputs are not deterministic. Categorized against the diagnostic cutoff, the model achieved a sensitivity of 0.9692 and a specificity of 0.9206, correctly flagging nearly all autistic children while wrongly labeling fewer than one in twelve typically developing children as at risk.</p>
<p>The external validation in Cohort C provided the most sobering and arguably most informative test. At the independent hospital, Gemini-2.5 correctly classified 53 of 60 children, yielding a sensitivity of 0.9333, a specificity of 0.8333, and an overall accuracy of 0.8833. Score concordance with experts remained high at 0.9387, and the mean absolute error rose only slightly to 0.875 points. Specificity, the ability to correctly clear children who do not have autism, dropped compared with the internal cohort, a pattern consistent with the well-known performance gap between selected case-control samples and messier real-world populations. The authors are explicit that these findings demonstrate feasibility and support further clinician-supervised evaluation, not yet real-world clinical effectiveness.</p>
<p>That caution matters because of how the study was designed. The cohorts were retrospective and case-control by construction, deliberately balancing autistic children against typically developing peers. Real screening clinics see a far more heterogeneous stream of children, including many with developmental delay, language disorders, ADHD, and other neurodevelopmental conditions that can mimic or co-occur with autism. A screening tool that performs brilliantly at distinguishing autism from typical development may perform less impressively at distinguishing autism from these look-alike conditions, and the authors themselves call for future prospective studies using larger, consecutive, and clinically heterogeneous cohorts with neurodevelopmental comparators. The diagnostic reference in each cohort was a consensus DSM-5 clinical assessment, which is a strong standard, but the children whose interviews fed the AI were not a random sample of the screening population.</p>
<p>Even with those caveats, the implications are considerable. Second-level autism screening is a bottleneck in many health systems: trained clinicians are scarce, waitlists are long, and the earlier a child is identified, the earlier intervention can begin, with well-documented benefits for developmental outcomes. A workflow in which a physician conducts the interview and an AI drafts the scoring, subject to clinician review, could compress the time between interview and actionable result from days or weeks to minutes. The study also addressed the ethical infrastructure that such a workflow demands: it was conducted under the Declaration of Helsinki with central ethics approval from Nanjing Brain Hospital, written informed consent from all parents and guardians explicitly covered audio recording, automated transcription, and the transmission of de-identified transcript text to commercial AI providers strictly for research evaluation, and the authors declared no competing interests. The work was funded by the National Natural Science Foundation of China and the Nanjing Science and Technology Development Plan.</p>
<p>The deeper significance of the study may lie less in the specific scores than in what it demonstrates about the maturing relationship between language models and clinical measurement. Scoring a narrative interview is precisely the kind of task that defeated earlier generations of machine learning, which needed rigid input formats and labeled features. Modern large language models, augmented with retrieval and structured reasoning prompts, can now operate on the raw substance of a clinical conversation, and they can do so with quantified repeatability, agreement statistics, and confidence intervals rather than anecdotes. The path from a retrospective case-control demonstration to a deployed screening aid runs through prospective trials, regulatory scrutiny, and careful attention to bias across languages, cultures, and clinical populations. But the study offers a concrete, statistically disciplined proof of concept that the AI can meet the expert halfway, and for families waiting on an autism assessment, that halfway point may one day mark the difference between early answers and long silence.</p>
<p><strong>Subject of Research:</strong> Large language model-assisted scoring of parent interviews for early autism screening in young children</p>
<p><strong>Article Title:</strong> Large language model–assisted scoring of physician-administered parent interviews for autism screening in young children: a retrospective diagnostic case-control and agreement study</p>
<p><strong>Article References:</strong> Zhu, J., Yue, Z., Han, Y., Yoo, H. J., Bong, G., Chen, H., Dai, Y., Qiu, N., Yin, S., Ma, Y., Wang, S., &amp; Ke, X. (2026). Large language model–assisted scoring of physician-administered parent interviews for autism screening in young children: a retrospective diagnostic case-control and agreement study. <em>BMC Psychiatry</em>. <a href="https://doi.org/10.1186/s12888-026-08686-7" rel="noopener noreferrer">https://doi.org/10.1186/s12888-026-08686-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12888-026-08686-7" rel="noopener noreferrer">10.1186/s12888-026-08686-7</a></p>
<p><strong>Keywords:</strong> autism spectrum disorder, large language models, artificial intelligence, autism screening, C-BeDevel-I, diagnostic accuracy, retrieval-augmented generation, automated scoring, early childhood development, BMC Psychiatry, clinical agreement, case-control study</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234526</post-id>	</item>
	</channel>
</rss>
