<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>language models in clinical psychology &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/language-models-in-clinical-psychology/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 10:21:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>language models in clinical psychology &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Can Write Your Mental Health Quiz, But It Cannot Prove It Works</title>
		<link>https://scienmag.com/ai-can-write-your-mental-health-quiz-but-it-cannot-prove-it-works/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 10:21:26 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI and subjective data analysis]]></category>
		<category><![CDATA[AI in psychological measurement]]></category>
		<category><![CDATA[algorithmic bias]]></category>
		<category><![CDATA[automation bias]]></category>
		<category><![CDATA[clinical narrative analysis with AI]]></category>
		<category><![CDATA[digital mental health]]></category>
		<category><![CDATA[ethical considerations in AI mental health tools]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI in mental health assessments]]></category>
		<category><![CDATA[generative psychometrics challenges]]></category>
		<category><![CDATA[item generation]]></category>
		<category><![CDATA[language models in clinical psychology]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[limitations of AI in mental health diagnosis]]></category>
		<category><![CDATA[measurement invariance]]></category>
		<category><![CDATA[measurement validity]]></category>
		<category><![CDATA[PLOS Mental Health]]></category>
		<category><![CDATA[psychological assessment]]></category>
		<category><![CDATA[psychological assessment validity standards]]></category>
		<category><![CDATA[psychometric rigor in AI applications]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[rapid AI assessment tools and validation]]></category>
		<category><![CDATA[synthetic data]]></category>
		<category><![CDATA[validity of AI-driven psychological tests]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247062</guid>

					<description><![CDATA[A new PLOS Mental Health review proposes a lifecycle framework for using generative AI in psychological measurement while warning that fluency, synthetic respondents, and algorithmic scores are no substitute for validated psychometric evidence.]]></description>
										<content:encoded><![CDATA[<p>Generative artificial intelligence is quietly infiltrating one of psychology&#8217;s most guarded territories: the measurement of the human mind. A new review published in PLOS Mental Health argues that while large language models can draft questionnaire items, classify patient narratives, and extract scores from clinical notes at unprecedented speed, none of that fluency constitutes valid measurement. The paper, led by David Villarreal-Zegarra of Universidad Continental in Peru, lays out a lifecycle framework designed to let researchers harness generative AI without quietly dismantling a century of psychometric rigor. Its central warning is blunt: linguistic fluency is not validity, and acceleration is not validation.</p>
<p>The stakes are higher than they might appear. Psychological assessment increasingly depends on language, narrative, and context, precisely the raw material that large language models and large multimodal models process best. An emerging vision sometimes called generative psychometrics proposes using these models to organize unstructured subjective data and produce structured psychological characterizations while preserving quantitative rigor. But the review stresses that a model&#8217;s ability to produce coherent, clinically plausible language says nothing about whether its outputs measure anything at all. Under established frameworks such as the Standards for Educational and Psychological Testing and the COSMIN taxonomy, validity is an evidentiary claim about how scores are interpreted and used, and that burden does not shrink because a machine wrote the items.</p>
<p>The authors draw a crucial terminological line by distinguishing four levels of evidence: raw data, candidate indicators, algorithmic scores, and validated psychometric measures. Digital traces from smartphones, chatbots, ecological momentary assessment, wearables, social media, and clinical notes should not be treated as psychometric measures by default, they argue. They remain raw data or candidate indicators until their construct interpretation and intended use are theoretically specified, technically verified, empirically calibrated, and validated with human data. Only then can a digital or AI-assisted system earn the label of a psychometric measure. This framing directly challenges the growing habit of treating any quantifiable signal, from heart-rate variability to keyboard dynamics, as if quantification alone conferred measurement status.</p>
<p>To organize the field, the review maps generative AI onto a seven-stage lifecycle of measurement development: construct definition, generation of items or signals, evaluation of content and data quality, piloting and calibration, validation, fairness and measurement invariance, and finally scoring, interpretation, and documentation. At each stage, AI can assist but no stage should be delegated entirely to it. Construct definition remains the anchor: language models can summarize literature and propose domains, but their conceptualizations are probabilistic syntheses of training data rather than theory-driven definitions, and they risk reinforcing dominant-language or culturally narrow perspectives. In mental health, where neighboring constructs such as distress, depression, burnout, and loneliness overlap semantically while differing clinically, that imprecision is dangerous.</p>
<p>The empirical evidence so far supports a conservative reading. Studies of AI-generated items, including ChatGPT-produced concept inventory items in physics education and machine-authored personality items, show that such material can achieve acceptable psychometric properties only after careful prompt engineering, expert review, selection, and testing with real respondents. Even then, problems of ambiguity, redundancy, and weak construct alignment persist. The review also highlights the V3 framework for biometric monitoring technologies, which separates verification, analytical validation, and clinical validation, and notes that generative AI cannot compensate for unverified devices, poorly validated algorithms, or unexamined missingness in sensor-derived data.</p>
<p>One of the most provocative findings concerns synthetic respondents. Researchers have begun using large language models as simulated survey participants to pre-screen items before costly piloting. In one study, six large language models were evaluated as respondents within an item response theory framework, and some model-derived item parameters correlated highly with human-calibrated ones. But the models&#8217; ability distributions were markedly narrower than human distributions, failing to reproduce real population variability. Synthetic responses may help flag obviously poor items or prioritize candidates for pilot testing, the authors conclude, but they cannot replace calibration in human samples when the goal is estimating symptom severity, prevalence, clinical cut-offs, or treatment response.</p>
<p>The review also introduces a conceptual distinction that may shape the field for years: GenAI for psychometrics, in which generative models support measurement development and scoring, versus psychometrics for GenAI, in which psychometric methods are turned on the models themselves. When a large language model becomes an active component of a measurement pipeline, eliciting narratives, classifying chatbot turns, scoring open-ended responses, or acting as an automated judge, its stability, bias, prompt sensitivity, and behavioral consistency directly affect the validity of the resulting scores. Recent preprint evidence underscores the concern: an LLM-native psychometric instrument found that stable model self-reports did not reliably predict observed model behavior across 25 models, suggesting a model can appear internally consistent while still missing the construct human observers care about.</p>
<p>The catalog of failure modes is extensive. Construct drift can occur when generated items gradually shift the target construct while remaining superficially coherent. Face validity without structural validity arises when AI-generated items sound right but prove factorially unstable or poorly related to external criteria. Algorithmic bias can emerge because training corpora encode and reproduce biases tied to gender, race, age, language, disability, culture, and socioeconomic position, making measurement invariance and differential item functioning non-optional analyses rather than afterthoughts. Prompt sensitivity and model drift mean that outputs can change with a provider&#8217;s silent update, undermining reproducibility across time and sites. Automation bias tempts clinicians and researchers to over-trust machine outputs, and data-security failures involving sensitive mental health information threaten confidentiality, scoring integrity, and public trust.</p>
<p>There is also a subtler philosophical risk: the flattening of subjectivity. Mental health constructs are inherently subjective, contextual, and culturally mediated. When open-text responses, clinical narratives, or expert annotations are aggregated into single labels, consensus scores, or embeddings, meaningful disagreement and individual variability can be obscured. The authors argue that even technically sophisticated AI-assisted systems may produce psychometrically limited scores if they do not preserve, model, or report the uncertainty and plurality of human psychological experience, an emerging concern that remains far from standard practice in annotation and evaluation pipelines.</p>
<p>The paper closes with a five-priority research agenda: treating fairness and measurement invariance as first-order requirements; establishing longitudinal reproducibility, prompt stability, and recertification rules as models drift; defining the boundary conditions under which LLM respondents and synthetic data help or harm calibration; building psychometric benchmarks for evaluating model behavior; and clarifying validity standards for multimodal measurement combining text, speech, wearables, and clinical notes. The authors also demand radical transparency, calling for reporting of exact model names, versions, prompts, inference settings, human edits, and provenance of every AI-generated element, in line with guidelines such as TRIPOD-LLM and CONSORT-AI. Their conclusion is deliberately conservative: generative AI should augment, not replace, psychometric science. It does not reduce the need for psychometrics; it makes psychometrics more central, more demanding, and more consequential than ever.</p>
<p><strong>Subject of Research:</strong> Psychometric applications and risks of generative artificial intelligence in digital mental health measurement</p>
<p><strong>Article Title:</strong> Psychometric applications of generative artificial intelligence: Lifecycle, risks, and research agenda</p>
<p><strong>Article References:</strong> Villarreal-Zegarra, D., Paredes-Gonzales, Y., &amp; García-Serna, J. (2026). Psychometric applications of generative artificial intelligence: Lifecycle, risks, and research agenda. <em>PLOS Mental Health, 3</em>(9), e0000702. <a href="https://doi.org/10.1371/journal.pmen.0000702" rel="noopener noreferrer">https://doi.org/10.1371/journal.pmen.0000702</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pmen.0000702" rel="noopener noreferrer">10.1371/journal.pmen.0000702</a></p>
<p><strong>Keywords:</strong> generative AI, psychometrics, large language models, digital mental health, measurement validity, psychological assessment, algorithmic bias, measurement invariance, synthetic data, item generation, automation bias, PLOS Mental Health</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247062</post-id>	</item>
	</channel>
</rss>
