<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>statistical analysis with AI &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/statistical-analysis-with-ai/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 02:09:43 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>statistical analysis with AI &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Joins the Psychometrics Lab: LLMs Help Build a New Emotion Regulation Scale</title>
		<link>https://scienmag.com/ai-joins-the-psychometrics-lab-llms-help-build-a-new-emotion-regulation-scale/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 02:09:43 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI in mental health research]]></category>
		<category><![CDATA[AI-assisted psychometric scale development]]></category>
		<category><![CDATA[AI-generated questionnaire items]]></category>
		<category><![CDATA[anxiety]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[bias screening in scale development]]></category>
		<category><![CDATA[BMC Psychology]]></category>
		<category><![CDATA[Depression]]></category>
		<category><![CDATA[emotion regulation]]></category>
		<category><![CDATA[emotion regulation measurement tools]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[human-AI collaboration in psychometrics]]></category>
		<category><![CDATA[innovative methods in emotion regulation assessment]]></category>
		<category><![CDATA[integration of generative AI in psychological research]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in psychology]]></category>
		<category><![CDATA[Mental health]]></category>
		<category><![CDATA[psychological assessment]]></category>
		<category><![CDATA[psychometric evaluation of emotion regulation scales]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[scale development]]></category>
		<category><![CDATA[statistical analysis with AI]]></category>
		<category><![CDATA[translation of psychological assessment tools]]></category>
		<category><![CDATA[well-being]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260826</guid>

					<description><![CDATA[Researchers at King Abdulaziz University used a multi-model large language framework with human oversight to develop and preliminarily validate a 17-item Emotion Regulation Scale in 743 adults.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has written poetry, passed bar exams, and drafted computer code, but one of its most delicate tests may now be in the psychology laboratory. Researchers at King Abdulaziz University in Jeddah, Saudi Arabia, have reported the first psychometric evaluation of a new Emotion Regulation Scale (ERS) developed with the assistance of multiple large language models, or LLMs, working under continuous human supervision. The study, published in BMC Psychology, describes how generative AI systems were woven into nearly every stage of scale construction, from drafting candidate items to screening them for bias, translating them, and helping interpret the statistical structure that emerged from real-world data. The result is a 17-item questionnaire that shows promising, if preliminary, evidence of measuring how people manage their emotions across cognitive, behavioral, and contextual domains.</p>
<p>Emotion regulation, the process by which people influence which emotions they have, when they have them, and how they experience and express them, is one of the most heavily studied constructs in mental health science. Decades of research, much of it built on James Gross&#8217;s influential Process Model, have linked difficulties in regulating emotion to depression, anxiety, and diminished well-being. Yet the authors argue that existing instruments each capture only part of the picture. Some focus narrowly on specific strategies such as cognitive reappraisal or expressive suppression, while others measure regulatory difficulties rather than the broader repertoire of regulatory processes. The ERS was designed to close that gap by covering cognitive, behavioral, and contextual dimensions of regulation within a single, compact instrument.</p>
<p>What makes the study notable is not just the new scale but the framework behind it. Rather than delegating the work to a single chatbot, the researchers deployed a multi-model pipeline in which several LLMs and AI-based tools each contributed at different stages. One model might generate an initial pool of items grounded in the theoretical framework, while others were used to refine wording, flag potentially biased or ambiguous phrasing, and assist with translation. After the statistical analyses were complete, the models also helped interpret the emerging factor structure. At every step, the researchers emphasize, human judgment remained in charge: the AI proposals were reviewed, edited, and validated by the investigators before anything entered the instrument. The initial pool of 25 items was progressively refined down to the final 17-item scale.</p>
<p>The empirical test involved 743 adults recruited through convenience sampling, a cross-sectional design typical of early psychometric work. Crucially, the sample was randomly split into two non-overlapping subsamples, a methodological safeguard that allows exploratory and confirmatory analyses to be conducted on independent data. The first half of the sample was used for exploratory factor analysis, the statistical technique that searches for hidden dimensions underlying patterns of item responses. The second half was reserved for confirmatory factor analysis, which tests whether a hypothesized structure fits new data. This split-sample approach is considered good practice in scale development because it prevents researchers from polishing a model on the same data used to discover it.</p>
<p>The exploratory analysis revealed a preliminary four-factor structure that together explained 50.85 percent of the variance in responses, a respectable figure for a newly developed instrument. The four factors appear to map onto distinct facets of emotion regulation, spanning adaptive strategies and maladaptive cognitive patterns. In the confirmatory stage, the researchers tested a post hoc refined 17-item model and found fit indices that met conventional benchmarks: the ratio of chi-square to degrees of freedom was 1.95, comfortably below the threshold of concern, while the Comparative Fit Index reached 0.934 and the Tucker-Lewis Index 0.918, both approaching or exceeding the 0.90 standard. The root mean square error of approximation came in at 0.050, sitting right at the boundary of what methodologists typically consider acceptable fit.</p>
<p>Reliability, the consistency with which the scale measures its intended constructs, was acceptable for an instrument at this early stage. Cronbach&#8217;s alpha values ranged from 0.68 to 0.74 across the subscales, and McDonald&#8217;s omega, a more modern reliability coefficient, ranged from 0.68 to 0.75. These figures fall short of the 0.80 or higher often desired for mature scales, but the authors are candid that such values are not unusual for short subscales in initial development, and the numbers provide a baseline for refinement in future studies. The modest reliability also reflects the trade-off inherent in brief instruments, which sacrifice some internal consistency for speed and ease of administration.</p>
<p>Validity evidence followed two complementary paths. Convergent validity examines whether the new scale correlates with established measures of related constructs in the expected directions. Here, the adaptive strategies subscale correlated positively with cognitive reappraisal, with coefficients ranging from 0.522 to 0.608, while maladaptive cognitions correlated with expressive suppression at 0.400. Concurrent validity then linked the ERS to mental health outcomes: maladaptive regulation was associated with depression at 0.432 and anxiety at 0.462, and negatively with well-being and quality of life, while adaptive strategies showed the opposite pattern, correlating negatively with depression and anxiety and positively with well-being and quality of life. Known-groups validity added a final layer, with analyses of variance showing significant differences in maladaptive cognition scores across levels of depression and anxiety, with F statistics of 39.32 and 64.49 respectively, both highly significant.</p>
<p>The implications extend beyond emotion research. Scale development has long been a slow, labor-intensive craft, often requiring months of item drafting, expert review, and pilot testing before a single data point is collected. If LLMs can reliably accelerate item generation and bias screening while humans retain editorial control, the pipeline from construct definition to validated instrument could shorten considerably, particularly for researchers in under-resourced settings or for languages where few validated measures exist. The multi-model approach also offers a hedge against the known weaknesses of any single AI system, since flaws introduced by one model can be caught by another or by human reviewers. The authors stress that the framework is assistive rather than autonomous, and that established psychometric procedures, not the AI, remain the backbone of the science.</p>
<p>Caution, however, is warranted on several fronts. The study relied on convenience sampling, which limits how far the findings generalize beyond the recruited population, and the authors themselves call for further research with diverse samples to evaluate the scale&#8217;s psychometric properties, generalizability, and broader utility. The reliability coefficients, while acceptable for early-stage work, leave room for improvement, and the four-factor structure is described as preliminary pending replication. The role of AI in the process also raises questions that the field is only beginning to grapple with, including how to document and audit the contributions of generative models and how to ensure that AI-drafted items do not subtly encode cultural or linguistic biases. The researchers&#8217; decision to keep human oversight explicit at every stage offers one template, but standards for AI-assisted measurement are still being written.</p>
<p>Even with those caveats, the study marks a genuine milestone: a peer-reviewed psychological instrument whose creation was shaped end to end by a structured collaboration between multiple AI systems and human scientists, validated against real data from hundreds of respondents. As generative models grow more capable, the boundary between computational tools and scientific instruments is likely to blur further, and this work offers an early, carefully documented glimpse of what that collaboration can look like in practice. For now, the 17-item ERS stands as both a candidate measure of a foundational mental health construct and a proof of concept that the psychometric toolkit of the future may include a large language model at the bench, always with a human hand on the controls.</p>
<p><strong>Subject of Research:</strong> AI-assisted psychometric development of an Emotion Regulation Scale using large language models</p>
<p><strong>Article Title:</strong> A multi-model LLM framework for AI-assisted psychological scale development: initial psychometric evaluation of the Emotion Regulation Scale (ERS)</p>
<p><strong>Article References:</strong> Almohammadi, I. A., &amp; Alzahrani, M. A. (2026). A multi-model LLM framework for AI-assisted psychological scale development: initial psychometric evaluation of the Emotion Regulation Scale (ERS). <em>BMC Psychology</em>. <a href="https://doi.org/10.1186/s40359-026-05759-w" rel="noopener noreferrer">https://doi.org/10.1186/s40359-026-05759-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40359-026-05759-w" rel="noopener noreferrer">10.1186/s40359-026-05759-w</a></p>
<p><strong>Keywords:</strong> artificial intelligence, large language models, emotion regulation, psychometrics, scale development, psychological assessment, mental health, factor analysis, BMC Psychology, depression, anxiety, well-being</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260826</post-id>	</item>
	</channel>
</rss>
