<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>factor analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/factor-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 15:50:35 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>factor analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Italian Nurses Get a New Tool to Predict Who Will Quit</title>
		<link>https://scienmag.com/italian-nurses-get-a-new-tool-to-predict-who-will-quit/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 15:50:35 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Anticipated Turnover Scale]]></category>
		<category><![CDATA[Anticipated Turnover Scale validation]]></category>
		<category><![CDATA[burnout]]></category>
		<category><![CDATA[content validity]]></category>
		<category><![CDATA[Cronbach's alpha]]></category>
		<category><![CDATA[cross-cultural adaptation]]></category>
		<category><![CDATA[early warning tools for nurse retention]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[health workforce shortage]]></category>
		<category><![CDATA[healthcare organizational management]]></category>
		<category><![CDATA[healthcare staffing challenges]]></category>
		<category><![CDATA[hospital management]]></category>
		<category><![CDATA[hospital staff retention strategies]]></category>
		<category><![CDATA[impact of nurse attrition on patient care]]></category>
		<category><![CDATA[Italian healthcare workforce]]></category>
		<category><![CDATA[nurse job satisfaction assessment]]></category>
		<category><![CDATA[nurse retention]]></category>
		<category><![CDATA[nurse turnover prediction]]></category>
		<category><![CDATA[nursing]]></category>
		<category><![CDATA[nursing profession workforce planning]]></category>
		<category><![CDATA[nursing workforce shortages]]></category>
		<category><![CDATA[predicting nurse resignations]]></category>
		<category><![CDATA[psychometric validation]]></category>
		<category><![CDATA[turnover intention]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=230658</guid>

					<description><![CDATA[Italian researchers have translated and preliminarily validated the Anticipated Turnover Scale, giving hospital managers a short questionnaire that can flag nurses' intention to quit before they resign.]]></description>
										<content:encoded><![CDATA[<p>Nursing is a profession under strain almost everywhere on Earth, and the numbers behind that strain are sobering. The World Health Organization and the International Council of Nurses estimate that by 2030 the world will be short roughly nine million nurses, a gap driven by aging populations, rising chronic disease, and working conditions that push experienced clinicians out of hospitals faster than they can be replaced. When a nurse leaves, the cost is not merely administrative. Losing seasoned professionals erodes intellectual capital, reduces productivity, and has been linked to poorer patient satisfaction, compromised staff safety, longer hospital stays, and lower quality of care. Organizations must then spend heavily to recruit, hire, and train replacements, while the nurses who remain often face heavier workloads, diminished job satisfaction, and weakened team cohesion. Against this backdrop, a research team in Italy has taken a step that could give hospital managers an early warning system: a carefully translated and preliminarily validated Italian version of the Anticipated Turnover Scale, a short questionnaire designed to measure whether nurses are thinking about walking out the door.</p>
<p>The Anticipated Turnover Scale, or ATS, was developed by Jan R. Atwood and Ada Sue Hinshaw in 1983 and has long been prized for its simplicity and reliability. The instrument consists of twelve items rated on a Likert scale, asking nurses how strongly they agree or disagree with statements about their intention to leave their current position, whether by moving to another hospital or transferring within the same organization. In its original form, responses ranged from 1, strongly disagree, to 7, strongly agree, producing total scores between 12 and 84, with higher scores signaling a greater intention to leave. The scale deliberately mixes positively and negatively worded items to blunt response bias, and it takes only about five minutes to complete. In the original validation, Cronbach&#8217;s alpha, a standard measure of internal consistency, was 0.84, and factor analysis revealed two main factors explaining 54.9 percent of the variance. A later study of newly graduated nurses using a five-point format achieved an alpha of 0.90. Until now, however, no reliable Italian-language version existed.</p>
<p>That gap mattered because intention to leave is widely regarded as the single best early indicator of actual turnover. The cognitive and emotional process of considering, planning, and deciding to quit precedes resignation, and research consistently shows a strong association between stated intention and eventual departure. Known drivers include job dissatisfaction, burnout, stress, work-life conflict, low pay, weak professional engagement, patient safety concerns, and leadership styles, including toxic leadership that breeds organizational mistrust. The COVID-19 pandemic intensified the problem dramatically, with roughly one-third of nurses reporting that they considered leaving the profession because of burnout and psychological distress. Generational patterns add another layer: Millennials change jobs more readily than Baby Boomers or Generation X, though across all generations, workload and staffing shortages dominate the reasons for leaving, while leadership support and salary anchor retention. Giving nurse managers a validated tool to detect these intentions early is therefore essential for planning retention strategies before resignations cascade.</p>
<p>The Italian team, publishing in Nursing Open, began by securing official authorization from Atwood and Hinshaw themselves, including permission to adopt a five-point Likert format instead of the original seven-point scale. The researchers chose the simplified format to improve usability for respondents, consistent with its successful use in prior studies. Translation and cultural adaptation then followed a rigorous four-step protocol aligned with internationally recommended procedures. Four translators with deliberately different profiles produced independent translations: a certified translator with no healthcare background, a non-healthcare translator with bicultural Italian-American experience, a nurse translator, and a nurse translator with bicultural Italian-English experience. Each was instructed to translate the instrument and its instructions while noting any difficulties. A bilingual committee, including a Director of Health Professions, a Departmental Nursing Manager, a Nursing Coordinator, a Labor Sociologist, a Psychologist, and an Associate Professor of Nursing, then reviewed the versions using a decentralized approach to preserve meaning in both languages.</p>
<p>The third and fourth steps closed the loop on linguistic fidelity. Two independent native English-speaking translators, one with and one without a healthcare background, back-translated the Italian version into English without ever seeing the original instrument or its translations. The research team then compared the two back-translated English versions, discussed discrepancies, and proposed a final version that the original author confirmed. With the translated scale in hand, the team moved to preliminary validation in a real-world setting: a public university hospital in northeastern Italy. All hospital nurses were invited to participate, with freelance nurses and those in managerial positions excluded. A self-administered questionnaire was distributed by email, and to minimize response bias the ATS was administered first, before any socio-demographic or occupational questions. Data collection ran from October 31 to November 20, 2023, through an anonymous online platform. Of 300 questionnaires distributed, 172 were returned and 168 were deemed valid for analysis.</p>
<p>The psychometric evaluation examined three pillars: face validity, content validity, and the internal structure of the scale. A panel of six experts judged that the translated instrument clearly and appropriately captured the phenomenon of nursing turnover, confirming its face validity without modifications and highlighting its potential for both research and practice. Content validity told a more nuanced story. Experts rated each item on a four-point relevance scale, with scores converted into binary relevant or not relevant categories and analyzed with a binomial test. Items 2 and 12 achieved unanimous relevance ratings with statistically significant agreement, while items 1, 5, 6, and 11 showed high though non-significant agreement. But items 7 and 9 were judged not relevant by five of the six evaluators, yielding item-level content validity index values of just 0.17. The scale-level index stood at 0.68, rising to 0.78 when the two problematic items were excluded. The authors caution that with only six evaluators, a single disagreement can substantially shift these indexes, so the estimates deserve careful interpretation.</p>
<p>The internal structure was explored through Principal Axis Factoring with Promax rotation, and the data proved highly suitable for factor analysis, with a Kaiser-Meyer-Olkin value of 0.891 and a significant Bartlett&#8217;s test of sphericity. For the full twelve-item version, the analysis extracted three factors explaining 57.82 percent of the variance. The dominant first factor, accounting for 44.59 percent, appeared to reflect a general attitude toward staying in or leaving the current job, primarily capturing internal turnover, meaning movement within the same organization. The second factor seemed tied to external turnover, the inclination to leave the organization entirely, with item 11, which expresses serious doubts about remaining, loading at a full 1.0. The third factor was associated almost exclusively with item 9, which conveys uncertainty about how long to stay in the current role. When items 7 and 9 were removed, a second factor analysis on the remaining ten items produced two factors explaining 59.16 percent of the variance, mirroring the internal-versus-external distinction.</p>
<p>Reliability results strengthened the case for the shortened version. The full twelve-item Italian scale achieved a Cronbach&#8217;s alpha of 0.828, already indicating good internal consistency, but excluding items 7 and 9 pushed the coefficient to 0.901, a level considered excellent. The first factor extracted from the twelve-item analysis showed an alpha of 0.912, while the first factor of the ten-item solution reached 0.809. These values sit comfortably within the range reported across four decades of ATS research: the original 1983 study reported 0.84, a Portuguese validation in Lisbon found 0.87 for twelve items rising to 0.91 after item exclusions, and a meta-analysis by Barlow and Zangaro documented alphas between 0.85 and 0.94 with a mean of 0.89. The Italian coefficients are slightly lower but largely acceptable, and the improvement after removing items 7 and 9 echoes the Portuguese findings, suggesting those items may pose cultural or linguistic difficulties beyond Italy alone.</p>
<p>Comparing validations across languages and forty years of healthcare change requires caution, and the authors are explicit about the limits of their work. The original development involved 1,525 professionals from fifteen rural Arizona hospitals, while the Portuguese study of 259 nurses is more comparable in scale. The Italian factor structure differs from previous validations even though it explains a similar proportion of variance, and no confirmatory factor analysis was performed, which would be needed to verify the observed structures in an independent sample. Criterion validity could not be assessed because no Italian gold standard for turnover measurement exists, and establishing a clinical cut-off score was infeasible due to privacy and timing constraints. Temporal stability, or test-retest reliability, remains untested, and the six-member expert panel may have been too small to fully capture content validity.</p>
<p>Even with these caveats, the implications for nursing management are concrete. A validated Italian instrument for measuring intention to leave would let nurse managers monitor turnover risk in real time and design targeted, personalized interventions to boost job satisfaction, engagement, and well-being before resignations occur. The authors conclude that the ten-item version, with its stronger content validity and internal consistency, is closest to practical use, though the scale is not yet ready for routine application. Future research should confirm the factor structure with independent samples, establish test-retest reliability, and link scores to actual turnover to define a meaningful threshold. In the meantime, the study adds an important piece to the global puzzle of nurse retention, offering Italian healthcare systems a scientifically grounded way to hear, earlier and more clearly, the quiet signal that precedes a resignation letter.</p>
<p><strong>Subject of Research:</strong> Italian translation and preliminary psychometric validation of the Anticipated Turnover Scale for measuring nurses&#x27; intention to leave</p>
<p><strong>Article Title:</strong> Evaluation of Nurses&#x27; Intent to Leave: Italian Translation, Cultural Adaptation, and Initial Psychometric Assessment of the Anticipated Turnover Scale</p>
<p><strong>Article References:</strong> Cignola, S., Bembich, S., &amp; Sanson, G. (2026). Evaluation of Nurses&#x27; Intent to Leave: Italian Translation, Cultural Adaptation, and Initial Psychometric Assessment of the Anticipated Turnover Scale. <em>Nursing Open, 13</em>(10), Article e70848. <a href="https://doi.org/10.1002/nop2.70848" rel="noopener noreferrer">https://doi.org/10.1002/nop2.70848</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1002/nop2.70848" rel="noopener noreferrer">10.1002/nop2.70848</a></p>
<p><strong>Keywords:</strong> nursing, turnover intention, Anticipated Turnover Scale, psychometric validation, cross-cultural adaptation, nurse retention, health workforce shortage, Cronbach&#x27;s alpha, factor analysis, content validity, burnout, hospital management</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">230658</post-id>	</item>
		<item>
		<title>New Turkish Version of ARFID Questionnaire Proves Accurate in Adults</title>
		<link>https://scienmag.com/new-turkish-version-of-arfid-questionnaire-proves-accurate-in-adults/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 11:44:11 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[adult ARFID assessment]]></category>
		<category><![CDATA[ARFID]]></category>
		<category><![CDATA[ARFID assessment]]></category>
		<category><![CDATA[ARFID symptoms and diagnosis]]></category>
		<category><![CDATA[development of culturally adapted mental health assessments]]></category>
		<category><![CDATA[eating disorder research in Turkey]]></category>
		<category><![CDATA[eating disorder screening tools]]></category>
		<category><![CDATA[eating disorders]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[food neophobia]]></category>
		<category><![CDATA[impact of ARFID on daily functioning]]></category>
		<category><![CDATA[mental health assessment]]></category>
		<category><![CDATA[non-English ARFID instruments]]></category>
		<category><![CDATA[nutritional deficiencies in ARFID]]></category>
		<category><![CDATA[PARDI-AR-Q]]></category>
		<category><![CDATA[PARDI-AR-Q validation study]]></category>
		<category><![CDATA[psychometric validation]]></category>
		<category><![CDATA[questionnaire]]></category>
		<category><![CDATA[screening tool]]></category>
		<category><![CDATA[test-retest reliability]]></category>
		<category><![CDATA[Turkish adaptation of ARFID questionnaire]]></category>
		<category><![CDATA[Turkish adults]]></category>
		<category><![CDATA[validated ARFID diagnostic tool]]></category>
		<category><![CDATA[validity]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227603</guid>

					<description><![CDATA[A new study validates the Turkish version of the PARDI-AR-Q, the first non-English adaptation of a brief self-report questionnaire for assessing ARFID symptoms in adults.]]></description>
										<content:encoded><![CDATA[<p>Avoidant/Restrictive Food Intake Disorder, better known as ARFID, has long lived in the shadow of its more famous eating disorder cousins, anorexia nervosa and bulimia nervosa. Unlike those conditions, ARFID is not driven by concerns about body shape or weight. Instead, people with ARFID avoid or restrict food for very different reasons: the texture, smell, or appearance of certain foods may be intolerable, their interest in eating may be strikingly low, or they may genuinely fear that eating will lead to choking or vomiting. The consequences can be serious, including nutritional deficiencies, weight loss, dependence on nutritional supplements, and profound interference with daily life. Although the diagnosis was formally introduced into the Diagnostic and Statistical Manual of Mental Disorders in 2013, most assessment tools remain available only in English, leaving clinicians in much of the world without validated instruments. A new study published in the Journal of Eating Disorders now changes that picture for Turkish-speaking adults, reporting the first non-English adaptation of a brief self-report questionnaire designed to capture the core dimensions of ARFID.</p>
<p>The instrument at the center of the study is the Pica, ARFID and Rumination Disorder Interview–ARFID Questionnaire, abbreviated as PARDI-AR-Q. It grew out of the PARDI, a structured clinician interview developed by an international team that includes Rachel Bryant-Waugh of King&#8217;s College London and the South London and Maudsley NHS Foundation Trust, one of the co-authors of the new paper. The questionnaire version condenses that interview into a self-report format, asking respondents about sensory-based avoidance of food, lack of interest in eating or appetite, fear of aversive consequences of eating, and the impact of eating problems on everyday functioning. Because ARFID presentations differ so markedly from person to person, capturing these distinct dimensions separately is critical. A person who avoids fruit because of its texture needs a different clinical conversation from someone who eats reluctantly because they fear vomiting, and a good screening tool must be able to distinguish between them.</p>
<p>Translating a psychological questionnaire is far more involved than swapping words between languages. Concepts such as food neophobia, sensory sensitivity, and appetite must resonate with the everyday vocabulary of the target population, and the response scales must behave in the same way statistically. To establish whether the Turkish PARDI-AR-Q met these standards, a research team led by Makbule Esen Öksüzoğlu of Kastamonu University, together with colleagues from several Turkish hospitals and universities, recruited 1,343 university students aged 18 to 40 from three Turkish universities. The sample was 53.4 percent female, with a mean age of 23.6 years. Participants completed the Turkish PARDI-AR-Q alongside a battery of comparison measures: the Nine-Item ARFID Screen, known as NIAS, the Food Neophobia Scale, the short form of the Eating Disorder Examination Questionnaire, and the Depression Anxiety Stress Scale-21. A subset of 138 participants filled out the PARDI-AR-Q a second time roughly four weeks later, allowing the researchers to test whether scores remained stable over time.</p>
<p>The statistical workup was thorough. The team examined internal consistency, the degree to which items within a scale measure the same underlying construct; test–retest reliability, the stability of scores across weeks; exploratory and confirmatory factor analyses, which test whether the questionnaire&#8217;s structure matches the theoretical dimensions of ARFID; measurement invariance across sex, age, ARFID-risk status, and weight status; and both convergent and criterion-referenced validity. Each of these analyses answers a different question about the instrument, and together they provide a comprehensive portrait of how the Turkish version behaves in practice. The study received ethics approval from Uşak University and was conducted in accordance with the Declaration of Helsinki, with written informed consent from all participants.</p>
<p>The headline finding is that the Turkish PARDI-AR-Q performed remarkably well. Exploratory factor analysis supported a four-factor structure mirroring the intended dimensions of the instrument, and confirmatory factor analysis showed excellent fit for a correlated four-factor model, with a comparative fit index of 0.990, a Tucker-Lewis index of 0.985, a root mean square error of approximation of 0.049, and a standardized root mean square residual of just 0.017. For readers unfamiliar with these statistics, values close to 1.0 on the first two indices and low values on the latter two indicate that the observed data align almost perfectly with the proposed model. In plain terms, the questionnaire measures four distinct but related facets of ARFID exactly as designed, and it does so in Turkish as cleanly as psychometricians could hope.</p>
<p>Reliability figures were equally impressive. Internal consistency, measured by Cronbach&#8217;s alpha, reached 0.95 for the total score, which is excellent, and ranged from 0.86 to 0.92 across the subscales, spanning good to excellent. The four-week test–retest reliability, expressed as intraclass correlation coefficients, fell between 0.921 and 0.947, meaning that a person&#8217;s scores changed very little over a month in the absence of any intervention. This stability matters enormously for research: if a screening tool gives wildly different results from one month to the next, it cannot be used to track symptoms or evaluate treatment. The Turkish PARDI-AR-Q passes that test with room to spare.</p>
<p>Validity evidence came from several directions. The PARDI-AR-Q subscales correlated most strongly with the corresponding subscales of the NIAS, with correlation coefficients ranging from 0.663 to 0.760, exactly the pattern expected if both instruments tap the same constructs. Participants who scored above the established NIAS cut-off for ARFID risk scored significantly higher on the PARDI-AR-Q, providing criterion-referenced support. The researchers also demonstrated scalar measurement invariance across sex, age group, ARFID-risk status, and weight status, which means the questionnaire functions comparably for women and men and for younger and older adults, allowing meaningful score comparisons across these groups. On the descriptive side, endorsement of diagnostic and risk-related items was generally low to moderate: 19.8 percent of participants reported perceiving an eating problem themselves, 18.8 percent said a professional had identified a nutritional deficiency, and 17.9 percent reported that others had noticed an eating problem. These figures suggest that eating-related concerns are far from rare in young adult populations, even if few meet full diagnostic criteria.</p>
<p>Why does this matter beyond Turkey? Psychometric validation studies can seem dry, but they are the quiet infrastructure of mental health science. Without validated tools in local languages, entire populations are effectively invisible to epidemiological research, and clinicians must rely on informal judgment or instruments that were never tested for their patients. ARFID in particular has been under-recognized in adults, with much of the literature focused on children and adolescents. The new study is explicitly the first to validate the PARDI-AR-Q outside English-speaking populations, making the Turkish version a template for adaptations elsewhere. The authors note that the questionnaire can now serve as a short, practical screening instrument for Turkish-speaking adults who may be experiencing ARFID symptoms and who might benefit from further assessment or support, opening the door to prevalence studies and clinical trials in a population of more than 80 million people.</p>
<p>The researchers are candid about the limits of their work. The sample consisted entirely of university students, a group that is young, educated, and not representative of the general adult population, let alone of clinical populations. Future studies, they write, should test the questionnaire in people with diagnosed eating disorders and in more diverse groups, including adults outside higher education. A screening tool validated in students may behave differently in older adults, in people with chronic medical conditions, or in psychiatric inpatients. There is also the inherent limitation of self-report: ARFID can involve limited insight, and some individuals may underreport symptoms, which is why the questionnaire is positioned as a screening and research tool rather than a substitute for a full clinical interview such as the original PARDI.</p>
<p>Even with those caveats, the study represents a meaningful step forward for a disorder that has struggled for recognition. ARFID is not picky eating, and it is not a phase that people simply grow out of; it is a condition that can compromise nutrition, growth, and quality of life across the lifespan. Giving clinicians and researchers in Turkey a rigorously validated, quick-to-administer measure of its core symptoms makes it possible to find the people who have been suffering quietly, often mislabeled as merely fussy or anxious around food. As awareness of ARFID grows worldwide, studies like this one quietly ensure that the science of measurement keeps pace, so that no one&#8217;s eating difficulties go unseen simply because of the language they speak.</p>
<p><strong>Subject of Research:</strong> Validation of the Turkish PARDI-AR-Q questionnaire for assessing ARFID symptoms in adults</p>
<p><strong>Article Title:</strong> Validity and reliability of the pica, ARFID and rumination disorder interview- ARFID questionnaire (PARDI-AR-Q) in Turkish adults</p>
<p><strong>Article References:</strong> Esen Öksüzoğlu, M., Günal Okumuş, H., Kaşak, M., Çelik, Y. S., Öğütlü, H., &amp; Bryant-Waugh, R. (2026). Validity and reliability of the pica, ARFID and rumination disorder interview- ARFID questionnaire (PARDI-AR-Q) in Turkish adults. <em>Journal of Eating Disorders</em>. <a href="https://doi.org/10.1186/s40337-026-01750-3" rel="noopener noreferrer">https://doi.org/10.1186/s40337-026-01750-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40337-026-01750-3" rel="noopener noreferrer">10.1186/s40337-026-01750-3</a></p>
<p><strong>Keywords:</strong> ARFID, PARDI-AR-Q, eating disorders, psychometric validation, questionnaire, Turkish adults, test-retest reliability, factor analysis, screening tool, food neophobia, mental health assessment, Validity</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227603</post-id>	</item>
		<item>
		<title>New Driving Self-Esteem Questionnaire Measures How Good You Feel Behind the Wheel</title>
		<link>https://scienmag.com/new-driving-self-esteem-questionnaire-measures-how-good-you-feel-behind-the-wheel/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 09:55:14 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[aggressive driving]]></category>
		<category><![CDATA[Current Psychology]]></category>
		<category><![CDATA[development of driving self-esteem questionnaire]]></category>
		<category><![CDATA[driver confidence assessment]]></category>
		<category><![CDATA[driving anger]]></category>
		<category><![CDATA[driving competence self-assessment]]></category>
		<category><![CDATA[driving self-esteem]]></category>
		<category><![CDATA[DSEQ]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[impact of self-esteem on driving behavior]]></category>
		<category><![CDATA[measurement invariance]]></category>
		<category><![CDATA[measuring driving self-worth]]></category>
		<category><![CDATA[psychological tools for driver confidence]]></category>
		<category><![CDATA[psychometric measurement of driving ability]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[questionnaire validation]]></category>
		<category><![CDATA[road safety]]></category>
		<category><![CDATA[self-esteem]]></category>
		<category><![CDATA[self-esteem in traffic psychology]]></category>
		<category><![CDATA[self-evaluation in driving]]></category>
		<category><![CDATA[self-perception of driving skills]]></category>
		<category><![CDATA[specialized self-esteem scales for drivers]]></category>
		<category><![CDATA[traffic psychology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226999</guid>

					<description><![CDATA[Researchers have developed and validated the Driving Self-Esteem Questionnaire, a unidimensional, gender-invariant scale showing that driving-specific self-esteem predicts anger on the road better than global self-esteem in key situations.]]></description>
										<content:encoded><![CDATA[<p>Ask any driver what they think of their own abilities behind the wheel and you will almost certainly get a confident answer, often a very confident one. Yet psychologists have long known that how people feel about themselves in general does not always predict how they feel about themselves in specific situations. A person may hold a healthy overall sense of self-worth while quietly dreading multi-lane roundabouts, or carry modest general self-esteem while believing they are among the best drivers on the road. A research team led by David Herrero-Fernández of the Universidad Europea del Atlántico in Spain, working with colleagues in Romania and Mexico, has now built and rigorously tested the first dedicated instrument to capture this elusive construct: the Driving Self-Esteem Questionnaire, or DSEQ. The work, published in Current Psychology, offers traffic researchers a psychometrically sound tool for measuring how drivers evaluate their own competence on the road.</p>
<p>The theoretical motivation behind the study rests on a distinction that has shaped self-esteem research for decades. Global self-esteem, famously measured by Morris Rosenberg&#8217;s 1965 scale, reflects a person&#8217;s overall evaluation of their own worth. But since the 1970s, researchers such as Richard Shavelson and Herbert Marsh have argued that self-concept is hierarchical and multidimensional: general feelings of self-worth sit atop a structure of domain-specific evaluations covering academic ability, physical appearance, social relationships and more. Meta-analytic work, including a 2022 synthesis of longitudinal studies by Lea Dapp, Sandra Krauss and Ulrich Orth, has shown that these domain-specific feelings can feed upward into global self-esteem, and that the domains are meaningfully different from one another. Driving, the Spanish-led team argues, is a domain where this specificity matters enormously, because driving self-worth is entangled with anger, risk-taking and aggression in ways that general self-esteem only partly captures.</p>
<p>The evidence for that entanglement is substantial. Global self-esteem has repeatedly been shown to predict anger and aggressive behaviour both on and off the road, and studies of driving anger, going back to Jerry Deffenbacher&#8217;s Driving Anger Scale in 1994, have documented how hostile emotions behind the wheel translate into dangerous conduct. Earlier work by Herrero-Fernández and colleagues had already suggested that self-esteem acts as a distal predictor of trait driving anger, and a companion study published in 2026 in Transportation Research Part F pointed to driving self-esteem specifically as a moderator of the link between general anger and anger on the road. What was missing was a validated way to measure that driving-specific self-evaluation directly, rather than inferring it from general scales. The DSEQ was designed to fill exactly that gap.</p>
<p>To build the questionnaire, the researchers generated an item pool covering the everyday texture of driving self-worth: feeling like a good or skilled driver, staying calm on uphill starts, handling overtaking on two-way roads, feeling insecure every time one gets behind the wheel, or preferring that a licensed passenger take over because they would do it better. The final published item set, presented bilingually in Spanish and English, includes statements such as &#8216;I am a good driver&#8217;, &#8216;I feel fearful every time I approach a large roundabout or one with many lanes&#8217; and &#8216;Nowadays, I would easily pass the practical driving test.&#8217; The team then administered the items to 530 Spanish drivers with an average age of 30.55 years, of whom 67.4 percent were women, and subjected the responses to a battery of modern psychometric techniques.</p>
<p>The analytical approach was deliberately conservative. Because questionnaire responses are ordinal rather than continuous, the team used polychoric correlations rather than ordinary Pearson coefficients in their exploratory factor analyses, and applied parallel analysis to decide how many factors the data genuinely supported. The initial results hinted at a multidimensional structure, but the researchers traced this pattern to item wording effects, the well-documented tendency of positively and negatively phrased statements to cluster separately regardless of content, a phenomenon also seen in the Rosenberg Self-Esteem Scale itself. Once wording artifacts were accounted for, the data supported a unidimensional structure: a single underlying dimension of driving self-esteem. The team reinforced this conclusion with dedicated unidimensionality indices, including UniCo, ECV and MIREAL, statistics developed by Pere Ferrando and Urbano Lorenzo-Seva to test whether a scale truly measures one thing rather than several bundled together.</p>
<p>The final model showed adequate fit and high reliability, meaning the items consistently tap the same construct. Crucially, the researchers also tested measurement invariance across gender, confirming that the questionnaire functions equivalently for men and women at the configural, metric and scalar levels. This matters because gender differences in self-esteem are among the most replicated findings in social psychology, and any comparison between male and female drivers would be meaningless if the scale itself behaved differently in each group. With invariance established, researchers can now compare driving self-esteem across genders with confidence that any observed differences reflect real psychological differences rather than measurement artifacts.</p>
<p>Validity evidence came from relationships with external variables. As expected, driving self-esteem correlated positively with global self-esteem, confirming that the new scale is related to, but distinct from, the general construct. More strikingly, it correlated negatively with several dimensions of driving anger, the emotional responses measured by Deffenbacher&#8217;s framework, including anger at discourtesy, slow driving and other common provocations. Drivers who feel competent and secure behind the wheel appear to be less easily provoked into anger on the road, a pattern consistent with theories linking threatened self-worth to hostile reactions.</p>
<p>The incremental validity results were more nuanced. In hierarchical regression analyses, the team tested whether the DSEQ explained additional variance in driving anger outcomes beyond what global self-esteem already accounted for. The answer was yes, but selectively: the questionnaire added small but meaningful increases in explained variance for anger related to discourtesy and slow driving, while showing no incremental contribution for other dimensions. The authors interpret this as evidence of modest, domain-specific incremental validity, which is precisely what a domain-specific instrument should show. A driving-specific measure should not outperform general self-esteem everywhere; it should add predictive power exactly where the driving context is most psychologically relevant.</p>
<p>The practical implications extend well beyond the laboratory. Traffic psychology has long grappled with the paradox that overconfident drivers can be as dangerous as insecure ones, with studies showing that high implicit self-esteem predicts risky behaviours such as dangerous mobile phone use while driving. A validated measure of driving self-esteem gives intervention designers a target: programmes aimed at recalibrating drivers&#8217; self-evaluations, whether inflated or deflated, can now assess whether they actually shift the construct they intend to change. The scale also opens the door to studying how driving self-esteem develops over the lifespan, how it relates to self-rated driving ability in older adults, and whether it moderates the well-documented pathway from driving anger to aggressive behaviour. The researchers note that the English version of the items is a direct translation of the original Spanish, and they advise any team adapting the DSEQ into other languages to follow the International Test Commission&#8217;s guidelines for test translation and adaptation.</p>
<p>For a field that has catalogued scales and inventories extensively, a 2026 review in Transportation Research Part F documented just how crowded the measurement landscape has become, the arrival of a new instrument demands justification. The DSEQ earns it by targeting a construct that general measures cannot reach: the specific self-worth a driver carries into the cabin. With a unidimensional structure, strong reliability, gender invariance and demonstrated incremental validity in the domains where driving anger bites hardest, the questionnaire gives researchers a precise instrument for one of the most consequential self-evaluations people make. Given that road traffic injuries remain a leading cause of death worldwide, understanding how drivers feel about themselves behind the wheel, and how those feelings fuel or defuse anger in traffic, is far more than an academic exercise. It is a step toward predicting, and ultimately preventing, the hostile encounters that turn ordinary journeys into dangerous ones.</p>
<p><strong>Subject of Research:</strong> Development and validation of a self-report questionnaire measuring driving-specific self-esteem in Spanish drivers</p>
<p><strong>Article Title:</strong> Development and validation of the driving self-esteem questionnaire</p>
<p><strong>Article References:</strong> Herrero-Fernández, D., Bogdan-Ganea, S. R., Martín-Ayala, J. L., &amp; Vistorte, A. O. R. (2026). Development and validation of the driving self-esteem questionnaire. <em>Current Psychology, 45</em>(19), Article 1562. <a href="https://doi.org/10.1007/s12144-026-10124-6" rel="noopener noreferrer">https://doi.org/10.1007/s12144-026-10124-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12144-026-10124-6" rel="noopener noreferrer">10.1007/s12144-026-10124-6</a></p>
<p><strong>Keywords:</strong> driving self-esteem, DSEQ, psychometrics, driving anger, self-esteem, traffic psychology, questionnaire validation, measurement invariance, factor analysis, aggressive driving, Current Psychology, road safety</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226999</post-id>	</item>
		<item>
		<title>New Burnout Scale for Autism Service Providers Passes Its First Big Test</title>
		<link>https://scienmag.com/new-burnout-scale-for-autism-service-providers-passes-its-first-big-test/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:12:25 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[applied behavior analysis]]></category>
		<category><![CDATA[autism]]></category>
		<category><![CDATA[Autism service provider burnout]]></category>
		<category><![CDATA[behavioral health providers]]></category>
		<category><![CDATA[burnout]]></category>
		<category><![CDATA[burnout among behavior analysts and direct support staff]]></category>
		<category><![CDATA[burnout assessment for developmental disabilities]]></category>
		<category><![CDATA[burnout research in developmental disabilities]]></category>
		<category><![CDATA[developmental disabilities]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[impact of caseloads and ethical strains]]></category>
		<category><![CDATA[improving workforce well-being in autism services]]></category>
		<category><![CDATA[Maslach Burnout Inventory]]></category>
		<category><![CDATA[measuring burnout in autism workforce]]></category>
		<category><![CDATA[Mental health]]></category>
		<category><![CDATA[Occupational Stress]]></category>
		<category><![CDATA[occupational stress in autism care]]></category>
		<category><![CDATA[occupational stressors in autism support roles]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[scale validation]]></category>
		<category><![CDATA[tailored burnout measurement tools]]></category>
		<category><![CDATA[validation of BADDS scale]]></category>
		<category><![CDATA[workforce retention]]></category>
		<category><![CDATA[workplace conditions in autism services]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225574</guid>

					<description><![CDATA[A 566-provider study distilled a 39-item burnout assessment for developmental disability settings into six validated factors that track with established burnout measures.]]></description>
										<content:encoded><![CDATA[<p>Burnout has quietly become one of the most pressing problems in autism services. The clinicians, behavior analysts, and direct support staff who deliver care to people with autism and related developmental disabilities face a distinctive cocktail of occupational stressors: physical aggression risk, heavy caseloads, demanding family interactions, thin supervision, and ethical strains that few other professions encounter in quite the same combination. Yet the field has long relied on generic burnout instruments that measure the symptoms of exhaustion without identifying the workplace conditions that produce them. A new study published in the Journal of Autism and Developmental Disorders takes a major step toward closing that gap, reporting the factor structure and validity of the Burnout Assessment for Developmental Disability Settings, or BADDS, a self-report scale purpose-built for this workforce.</p>
<p>The research team, led by Summer Bottini of Marcus Autism Center and Emory University School of Medicine, together with colleagues including Laura Johnson, Scott Gillespie, Mindy Scheithauer, Alexandra Hardee, and Lawrence Scahill, surveyed 566 autism service providers online. Participants completed the BADDS alongside three well-established comparison instruments: the Maslach Burnout Inventory Human Services Survey, the standard measure of burnout symptoms in helping professions; the Areas of Worklife Survey, which captures organizational conditions tied to burnout; and the Patient Health Questionnaire-9, a widely used screen for depressive symptoms. The design allowed the researchers to test whether the new scale behaves the way a valid measure of workplace-driven burnout should.</p>
<p>The path to a finished instrument was anything but smooth, and the study is refreshingly candid about that. An earlier confirmatory factor analysis, based on the scale&#8217;s initial qualitative development work, showed poor fit, meaning the original item structure did not hang together statistically the way the developers had hoped. An expert panel stepped in and cut 17 redundant items, leaving 56. The researchers then ran an exploratory factor analysis on those remaining items, a statistical technique that lets the data itself reveal how questions cluster into coherent dimensions. That analysis trimmed the scale further, down to 39 items organized into six distinct factors.</p>
<p>Those six factors read like a field guide to the pressures of developmental disability work. Physical Safety captures concerns about aggression and injury risk, an occupational hazard that distinguishes this workforce from most healthcare settings. Training and Supervision reflects the adequacy of mentorship and skill development, long identified in the literature as a buffer against burnout. Stakeholder Interactions addresses the complex web of relationships with families, schools, and other professionals. Professional Development covers career growth and advancement, Stress Management addresses coping resources and workload, and Ethics and Values probes alignment between personal principles and organizational practice. Together they sketch a portrait of a profession strained not only by the intensity of the work itself but by the systems surrounding it.</p>
<p>The psychometric numbers are solid. Internal reliability, measured by how consistently items within each subscale correlate with one another, ranged from 0.75 to 0.92 across the six subscales, comfortably within accepted standards for research and applied instruments. Values above 0.70 are generally considered acceptable, and the upper reaches of the BADDS range approach the levels seen in gold-standard clinical measures. This means that providers answering the scale are responding to coherent, stable dimensions of their work environment rather than to a scattering of loosely related complaints.</p>
<p>Validity evidence came from the correlations with the comparison measures. Correlations between BADDS subscales and the Maslach Burnout Inventory and Areas of Worklife Survey subscales ranged from 0.04 to 0.60. That pattern is exactly what a well-designed stressor measure should show: strong enough relationships with symptom measures to demonstrate convergent validity, but not so strong as to suggest the BADDS is merely duplicating existing instruments. The low end of the range also makes sense, since some workplace stressors, such as ethical concerns, would not be expected to track tightly with every symptom dimension. The scale is measuring something related to burnout but not identical to it, which is precisely the point.</p>
<p>Perhaps the most practically important finding concerns dose and response. The researchers found that elevations across BADDS subscales were associated with incremental increases in Maslach Burnout Inventory scores. In plain terms, the more a provider reports problems in a given workplace domain, the more burnout symptoms they report, in a graded fashion. This dose-response pattern is a hallmark of a meaningful measure and supports the idea that the BADDS is capturing workplace conditions that actually matter for worker wellbeing, not just noise. It also suggests the scale could serve as a diagnostic tool for organizations, pinpointing which domains are most elevated in a given clinic or agency before those problems translate into exhaustion, depression, and turnover.</p>
<p>The stakes for this kind of measurement are high. The autism services workforce has been documented in prior research to experience elevated burnout, particularly among early-career board certified behavior analysts with low collegial support, and turnover among behavior technicians and analysts disrupts continuity of care for children and families who depend on consistent intervention. The World Health Organization&#8217;s inclusion of burnout in the ICD-11 classification has sharpened attention on occupational burnout as a legitimate workplace phenomenon, and the field of behavior analysis has increasingly turned to organizational behavior management frameworks to address it. What has been missing is an instrument that speaks the specific language of this workforce, and the BADDS is designed to fill that role.</p>
<p>The study&#8217;s methodology reflects the realities of modern survey research. The team used strategies to detect insincere respondents, a growing concern in online research where bots and inattentive participants can contaminate data. The work was supported by the Emory University Pediatric Biostatistics Core and the Marcus Autism Center Clinical Innovation Fund, and the authors report no competing interests. The scale&#8217;s development followed a two-stage logic common in contemporary instrument design: qualitative exploration to generate items grounded in the lived experience of providers, followed by quantitative refinement to test and trim the resulting item pool against real data.</p>
<p>The authors are careful to note that additional research is warranted to confirm the psychometrics and utility of the BADDS, and that caution is appropriate. This study establishes the scale&#8217;s internal structure and its relationships with established measures in a single large sample; future work will need to test it across different settings, track providers over time to see whether BADDS scores predict later burnout and turnover, and evaluate whether interventions guided by BADDS profiles actually improve working conditions and retention. Still, the arrival of a validated, domain-specific burnout assessment gives autism service organizations something they have never had before: a way to measure the specific pressures wearing down their staff, and a roadmap for fixing them before the workforce burns out entirely.</p>
<p><strong>Subject of Research:</strong> Psychometric validation of a burnout assessment scale for behavioral health providers in autism and developmental disability settings</p>
<p><strong>Article Title:</strong> Factor Structure and Validity of the Burnout Assessment for Developmental Disability Settings (BADDS)</p>
<p><strong>Article References:</strong> Bottini, S., Johnson, L., Gillespie, S., Scheithauer, M., Hardee, A., &amp; Scahill, L. (2026). Factor Structure and Validity of the Burnout Assessment for Developmental Disability Settings (BADDS). <em>Journal of Autism and Developmental Disorders</em>. <a href="https://doi.org/10.1007/s10803-026-07553-4" rel="noopener noreferrer">https://doi.org/10.1007/s10803-026-07553-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10803-026-07553-4" rel="noopener noreferrer">10.1007/s10803-026-07553-4</a></p>
<p><strong>Keywords:</strong> burnout, autism, developmental disabilities, psychometrics, factor analysis, behavioral health providers, Maslach Burnout Inventory, occupational stress, workforce retention, scale validation, applied behavior analysis, mental health</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225574</post-id>	</item>
		<item>
		<title>Bowel Dysfunction After Rectal Cancer Surgery Splits Into Three Distinct Symptom Domains, Study Finds</title>
		<link>https://scienmag.com/bowel-dysfunction-after-rectal-cancer-surgery-splits-into-three-distinct-symptom-domains-study-finds/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 19:24:24 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[bowel dysfunction]]></category>
		<category><![CDATA[bowel dysfunction after rectal cancer surgery]]></category>
		<category><![CDATA[clustering of bowel movements]]></category>
		<category><![CDATA[exploratory factor analysis in gastrointestinal studies]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[fecal incontinence]]></category>
		<category><![CDATA[fecal incontinence in rectal surgery patients]]></category>
		<category><![CDATA[impact of sphincter-preserving surgery]]></category>
		<category><![CDATA[LARS score]]></category>
		<category><![CDATA[LARS symptom domains]]></category>
		<category><![CDATA[low anterior resection syndrome]]></category>
		<category><![CDATA[mixed-effects models]]></category>
		<category><![CDATA[patient-reported outcome measures for bowel function]]></category>
		<category><![CDATA[patient-reported outcomes]]></category>
		<category><![CDATA[pelvic floor]]></category>
		<category><![CDATA[postoperative bowel management]]></category>
		<category><![CDATA[radiotherapy]]></category>
		<category><![CDATA[Rectal cancer surgery]]></category>
		<category><![CDATA[rectal tumor resection complications]]></category>
		<category><![CDATA[risk factors for bowel dysfunction]]></category>
		<category><![CDATA[treatment approaches for LARS]]></category>
		<category><![CDATA[urgency]]></category>
		<category><![CDATA[Wexner score]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223586</guid>

					<description><![CDATA[A Japanese factor-analytic study of 320 rectal surgery patients shows that postoperative bowel dysfunction comprises three distinct symptom domains with different risk factors, challenging the use of composite LARS and Wexner scores alone.]]></description>
										<content:encoded><![CDATA[<p>For the hundreds of thousands of people who undergo sphincter-preserving surgery for rectal cancer each year, saving the sphincter muscle is only half the battle. Many survivors are left with a stubborn cluster of bowel problems known as low anterior resection syndrome, or LARS, a condition that can bring urgency, incontinence, frequent stools and a phenomenon called clustering, in which several bowel movements arrive in rapid succession. Now, a new study from Japan suggests that this syndrome is not one disorder at all, but several distinct conditions hiding inside a single score, each with its own risk factors and, potentially, its own treatment.</p>
<p>Researchers at Fukuoka University Hospital analyzed data from 320 patients who underwent low anterior resection for rectal tumors between 2016 and 2025. Their findings, published in Annals of Gastroenterological Surgery, used a statistical technique called exploratory factor analysis to dissect the two most widely used patient-reported measures of postoperative bowel function: the LARS score and the Cleveland Clinic Florida Fecal Incontinence Score, better known as the Wexner score. What emerged was a three-part structure that the authors say should change how clinicians think about, measure and treat bowel dysfunction after rectal surgery.</p>
<p>The logic behind the study is straightforward but powerful. The LARS score bundles five symptoms into a single number, while the Wexner score focuses mainly on incontinence. Clinicians have long noticed that the two instruments sometimes disagree, giving conflicting impressions of the same patient. The Japanese team suspected that this discordance arises because the scores are actually measuring different underlying dimensions of bowel dysfunction, dimensions that a composite total quietly blends together. To test that idea, they went beneath the total scores and examined the individual symptom items themselves.</p>
<p>First, the researchers mapped how the symptoms correlated with one another. Some pairs were tightly linked: Wexner liquid incontinence and LARS liquid stool incontinence correlated strongly, with a Spearman coefficient of 0.71, and urgency tracked closely with clustering at 0.59. Other symptoms, notably bowel frequency, floated loosely, correlating weakly with nearly everything else. When the team reorganized the correlation matrix using hierarchical clustering, related symptoms visibly grouped together, hinting that the syndrome had natural fault lines running through it.</p>
<p>Formal factor analysis confirmed the hunch. Using maximum likelihood extraction with Quartimin rotation on pooled observations from 6, 12 and 24 months after surgery, the analysis retained three factors that together explained 64.7 percent of the total variance. The first factor gathered the incontinence-related items, including liquid and solid stool incontinence, pad use and lifestyle alteration. The second was defined almost entirely by urgency and clustering, with a smaller contribution from bowel frequency. The third captured gas incontinence from both instruments, pointing to a separate domain of impaired gas control that neither score was designed to isolate on its own.</p>
<p>The real payoff came when the team asked which clinical factors drove each domain. Using multivariable mixed-effects models that accounted for repeated measurements in the same patients over time, they found that the same risk factors did not affect all domains equally. Very low anterior resection, in which the stapled anastomosis is created inside the anal canal itself, and a transanal surgical approach strongly worsened the incontinence domain, with beta coefficients of 0.437 and 0.860 respectively, but had much weaker effects on urgency and clustering. This pattern fits the anatomy: an ultra-low anastomosis and transanal dissection can directly injure the sphincter complex, undermining the mechanical barrier against leakage rather than the reflexes governing urgency.</p>
<p>Radiotherapy told a different story. Preoperative radiation worsened all three domains, suggesting it damages pelvic tissue in a broader, less selective way. Preoperative chemotherapy raised incontinence severity as well. Meanwhile, the age findings were the most striking and perhaps the most surprising. Compared with patients under 50, those aged 50 to 70 had greater incontinence-domain severity, consistent with an age-related decline in pelvic floor reserve. Yet patients aged 70 and older actually reported less urgency and clustering than the youngest group, with a negative beta coefficient of -0.257. The authors interpret this inversion as evidence that the two domains arise from different biology: incontinence from structural, anatomical vulnerability, and urgency with clustering from functional abnormalities, such as heightened visceral sensitivity and colonic hypermotility, that resemble diarrhea-predominant irritable bowel syndrome and tend to fade with age.</p>
<p>That distinction has immediate therapeutic implications. Patients whose suffering is dominated by urgency and clustering, often younger people, may respond to drugs that calm colonic hypermotility; the authors point to emerging evidence that 5-HT3 receptor antagonists such as ramosetron can relieve urgency and frequent stools in selected LARS patients. Patients with incontinence-dominant symptoms, by contrast, are more likely to benefit from physical rehabilitation of the pelvic floor and sphincter, including biofeedback and pelvic floor muscle training. A single composite score, the study argues, cannot tell these patients apart, because two people with identical LARS totals may occupy entirely different symptom worlds.</p>
<p>The study has important caveats, which the authors lay out candidly. It was conducted at a single center, and the three-factor structure was derived exploratively without confirmation in an independent cohort, so the domains remain hypothesis-generating rather than established. The LARS items were reweighted onto an ordinal severity scale before analysis, since the original score&#8217;s non-linear item weights were built for classification, not structural modeling. Questionnaire response fell to 66.9 percent at 24 months, and missing data were handled under a missing-at-random assumption without formal sensitivity analysis, meaning late symptom severity may have been underestimated. The analysis also excluded obstructed defecation and quality-of-life measures, which international consensus definitions consider integral to the full LARS construct, and the classification of very low anterior resection relied on the operating surgeon&#8217;s intraoperative judgment rather than a standardized anatomical measurement.</p>
<p>Even with those limitations, the message is hard to ignore. Postoperative bowel dysfunction, long summarized by a single number on a questionnaire, appears to be a family of related but separable conditions, each with its own anatomy, epidemiology and treatment logic. A domain-based approach, the authors conclude, offers a more precise framework for individualized functional assessment and could guide the design of future intervention trials, whether testing antihypertensive-style motility drugs for young patients with urgency-dominant disease or targeted pelvic floor rehabilitation for those with structural incontinence. For rectal cancer survivors, the difference between one blurred score and three clear domains may ultimately be the difference between generic reassurance and treatment that actually fits the problem.</p>
<p><strong>Subject of Research:</strong> Symptom domain structure of postoperative bowel dysfunction after sphincter-preserving rectal cancer surgery</p>
<p><strong>Article Title:</strong> Exploring Symptom Domains of Postoperative Bowel Dysfunction by Integrating LARS and Wexner Scores: A Longitudinal Factor‐Analytic Study</p>
<p><strong>Article References:</strong> Matsumoto, Y., Takeshita, I., Shiokawa, K., Sahara, K., Munechika, T., Nagata, K., Nagano, H., Takahashi, H., &amp; Hasegawa, S. (2026). Exploring Symptom Domains of Postoperative Bowel Dysfunction by Integrating LARS and Wexner Scores: A Longitudinal Factor‐Analytic Study. <em>Annals of Gastroenterological Surgery</em>, Article ags3.70289. <a href="https://doi.org/10.1002/ags3.70289" rel="noopener noreferrer">https://doi.org/10.1002/ags3.70289</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1002/ags3.70289" rel="noopener noreferrer">10.1002/ags3.70289</a></p>
<p><strong>Keywords:</strong> low anterior resection syndrome, LARS score, Wexner score, rectal cancer surgery, fecal incontinence, factor analysis, bowel dysfunction, urgency, pelvic floor, radiotherapy, patient-reported outcomes, mixed-effects models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223586</post-id>	</item>
		<item>
		<title>Muscle Symptoms May Anchor a Hidden Web of Body-Wide Dysregulation in ME/CFS</title>
		<link>https://scienmag.com/muscle-symptoms-may-anchor-a-hidden-web-of-body-wide-dysregulation-in-me-cfs/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:27:50 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[age-related symptom variation in ME/CFS]]></category>
		<category><![CDATA[autonomic dysfunction]]></category>
		<category><![CDATA[body-wide physiological dysregulation]]></category>
		<category><![CDATA[chronic fatigue syndrome]]></category>
		<category><![CDATA[chronic fatigue syndrome diagnostic biomarkers]]></category>
		<category><![CDATA[cross-sectional ME/CFS symptom analysis]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[gender differences in ME/CFS symptomatology]]></category>
		<category><![CDATA[Journal of Translational Medicine]]></category>
		<category><![CDATA[ME/CFS]]></category>
		<category><![CDATA[ME/CFS muscle symptoms]]></category>
		<category><![CDATA[menopausal status]]></category>
		<category><![CDATA[multisystem symptoms in ME/CFS]]></category>
		<category><![CDATA[muscle pain and autonomic dysfunction]]></category>
		<category><![CDATA[muscle problems and breathing difficulties]]></category>
		<category><![CDATA[muscle symptoms]]></category>
		<category><![CDATA[physiological dysregulation]]></category>
		<category><![CDATA[physiological mechanisms underlying ME/CFS]]></category>
		<category><![CDATA[post-exertional malaise]]></category>
		<category><![CDATA[sex differences]]></category>
		<category><![CDATA[structural equation modeling]]></category>
		<category><![CDATA[symptom architecture in ME/CFS]]></category>
		<category><![CDATA[symptom phenotyping]]></category>
		<category><![CDATA[systemic inflammation in ME/CFS]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223254</guid>

					<description><![CDATA[A new analysis of 736 ME/CFS patients finds that muscle symptoms are statistically embedded in a latent physiological dysregulation factor spanning breathing, cardiovascular, thermoregulatory, visual, and flu-like symptoms, with exploratory differences by sex and menopausal-status proxy group.]]></description>
										<content:encoded><![CDATA[<p>Muscle problems are among the most disabling features of myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS), a condition that afflicts millions worldwide and remains stubbornly resistant to clear-cut diagnostic biomarkers. A new exploratory study published in the Journal of Translational Medicine suggests that these muscle symptoms may not be an isolated complaint but rather a visible surface expression of a deeper, shared physiological dysregulation that also drives breathing difficulties, cardiovascular symptoms, thermoregulatory problems, visual disturbances, and flu-like feelings. Drawing on cross-sectional data from 736 individuals with physician-diagnosed ME/CFS enrolled in the APAV-ME/CFS registry, Lotte Habermann-Horstmeier of the Villingen Institute of Public Health and Lukas M. Horstmeier of the Institute for Medical Biometry and Statistics at University Hospital Freiburg set out to map where muscle symptoms sit within the disease&#8217;s broader symptom architecture, and whether that architecture looks different in women versus men and across age-defined menopausal-status proxy groups.</p>
<p>The analytical strategy was deliberately layered. First, the researchers used multivariable logistic regression to ask which of 14 symptom domains independently predicted the presence of muscle problems. In the full model, three symptoms emerged as significant independent predictors: breathing problems carried an odds ratio of 2.49 (95 percent confidence interval 1.42 to 4.38, p = 0.001), flu-like symptoms an odds ratio of 1.97 (1.14 to 3.42, p = 0.015), and temperature-regulation disorders an odds ratio of 1.97 (1.10 to 3.51, p = 0.021). When the model was restricted to a reduced set of physiological symptoms, cardiovascular symptoms additionally reached significance with an odds ratio of 1.87 (1.09 to 3.24, p = 0.024). These odds ratios indicate that patients reporting each of these symptoms had roughly double the odds of also reporting muscle problems compared with patients who did not, after accounting for the other symptoms in the model.</p>
<p>Correlation analysis reinforced the impression that muscle symptoms travel with a distinct physiological cluster. Tetrachoric correlations, which estimate the latent association between binary symptom variables, showed substantial relationships between muscle problems and breathing difficulties (rho = 0.52), cardiovascular symptoms (rho = 0.49), and temperature-regulation disorders (rho = 0.49), with somewhat weaker but still notable links to visual disturbances (rho = 0.39) and flu-like symptoms (rho = 0.39). Values in this range suggest that the co-occurrence of these symptoms is far from random and points toward a common underlying driver rather than a coincidental clustering of unrelated complaints.</p>
<p>The centerpiece of the study, however, was its latent-variable modeling. Using exploratory factor analysis followed by structural equation modeling (SEM), the authors tested whether the physiological symptoms could be explained by a single latent factor, an unobserved variable hypothesized to generate the observed symptom pattern. The model fit was strong by conventional standards: the root mean square error of approximation (RMSEA) was 0.046, well below the 0.06 threshold typically considered acceptable; the comparative fit index (CFI) was 0.971 and the Tucker-Lewis index (TLI) was 0.952, both close to the 0.95 benchmark; and the standardized root mean square residual (SRMR) was 0.026, far under the 0.08 cutoff. Together these indices indicate that a single latent physiological dysregulation factor provides a parsimonious and statistically credible account of how these symptoms co-occur in the sample.</p>
<p>Not every symptom behaved identically within this structure. Gastrointestinal complaints loaded substantially on the latent factor even though they showed no independent association with muscle problems in the multivariable regression, suggesting they belong to the broader physiological pattern but are not tightly coupled to the muscle-specific manifestation. Urogenital symptoms also displayed a smaller but significant loading on the factor alongside substantial item-specific variance, indicating that they are partly woven into the shared dysregulation and partly driven by their own distinct processes. This kind of decomposition, separating shared from symptom-specific variance, is precisely what latent-variable approaches are designed to deliver, and it hints that ME/CFS pathology may involve both a common multisystem thread and parallel, partially independent symptom channels.</p>
<p>The sex-stratified analyses produced a striking descriptive asymmetry. In women, all five physiological symptoms, breathing, cardiovascular, thermoregulatory, visual, and flu-like, independently predicted muscle problems. In men, only breathing problems and temperature-regulation disorders remained significant. Yet when the authors formally tested the overall symptom-by-sex interaction, the result did not reach statistical significance, meaning the data do not establish that the symptom associations genuinely differ between the sexes. One exception appeared at the individual symptom level: the association between breathing-related symptoms and muscle problems showed a statistically significant interaction with sex, hinting that this particular link may be stronger in one sex than the other. The authors urge caution here, noting that the overall interaction test was non-significant and the confidence interval for the interaction estimate was wide, leaving the finding exploratory at best.</p>
<p>Menopausal status, approximated by age-defined proxy groups, added another descriptive layer. Women classified as premenopausal by this proxy showed significant associations between muscle problems and cardiovascular symptoms, visual disturbances, and temperature-regulation disorders, whereas women classified as postmenopausal showed significant associations with breathing difficulties and flu-like symptoms. Flu-like symptoms were independently associated with muscle problems in the postmenopausal subgroup but not in the premenopausal one. Once again, however, the formal test told a more conservative story: the overall menopausal-by-symptom interaction was not statistically significant (Wald chi-square of 5.96 on 5 degrees of freedom, p = 0.310). The authors therefore frame these subgroup contrasts as exploratory differences in association rather than evidence of a menopause-specific effect, a distinction that matters enormously for how the findings should be cited and built upon.</p>
<p>Stability across disease duration offered further reassurance about the latent factor&#8217;s measurement properties. The specified physiological factor showed statistical comparability across cross-sectional disease-duration groups, meaning there was no overall evidence that the factor means something different in patients early versus late in their illness. This kind of measurement invariance is a prerequisite for meaningful comparison across patient subgroups and lends weight to the idea that the dysregulation factor is a stable feature of the disease rather than an artifact of illness stage, survivorship bias, or shifting symptom reporting over time.</p>
<p>What could the latent factor represent biologically? The authors are careful not to overclaim, but they note that the findings are compatible with the hypothesis that the symptom-dysregulation factor reflects a chronic post-exertional malaise-associated multisystem response. The symptom cluster it captures, spanning respiratory, cardiovascular, thermoregulatory, visual, and flu-like domains, overlaps substantially with the territory of autonomic dysfunction, vascular dysregulation, immune activation, and metabolic disturbance, all of which have been implicated in ME/CFS by prior physiological studies. Whether the latent factor corresponds to genuinely interacting mechanisms across these systems is a question the present cross-sectional design cannot answer, and the authors explicitly call for longitudinal and biomarker-based studies to test it.</p>
<p>The practical implications could be considerable. If muscle symptoms indeed serve as a candidate integrative symptom within a measurable latent dysregulation factor, then symptom-based phenotyping may help stratify patients for clinical trials, a persistent challenge in a disease as heterogeneous as ME/CFS. Better stratification could, in turn, accelerate the development of mechanism-based therapeutic approaches by ensuring that patients sharing a common physiological signature are studied together. The study was conducted in 2022 under Declaration of Helsinki protocols approved by the Ethics Committee of Furtwangen University, Germany, and was partially funded by the Ministry of Social Affairs, Health and Integration using state funds approved by the Baden-Württemberg State Parliament, with the subsequent analyses unfunded. As an exploratory, cross-sectional analysis of registry data, it cannot establish causality or direction, and its subgroup findings demand replication. But by quantifying how tightly muscle symptoms are woven into a broader physiological web, and by providing a statistically well-fitting latent structure that future biomarker studies can interrogate, the work offers ME/CFS researchers a concrete, testable framework for one of medicine&#8217;s most perplexing chronic illnesses.</p>
<p><strong>Subject of Research:</strong> Latent physiological symptom-dysregulation and muscle symptoms in ME/CFS analyzed by sex and menopausal-status proxy group</p>
<p><strong>Article Title:</strong> Muscle symptoms and a latent physiological symptom-dysregulation factor in ME/CFS: exploratory analyses by sex and age-defined menopausal-status proxy group</p>
<p><strong>Article References:</strong> Habermann-Horstmeier, L., &amp; Horstmeier, L. M. (2026). Muscle symptoms and a latent physiological symptom-dysregulation factor in ME/CFS: exploratory analyses by sex and age-defined menopausal-status proxy group. <em>Journal of Translational Medicine</em>. <a href="https://doi.org/10.1186/s12967-026-09018-9" rel="noopener noreferrer">https://doi.org/10.1186/s12967-026-09018-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12967-026-09018-9" rel="noopener noreferrer">10.1186/s12967-026-09018-9</a></p>
<p><strong>Keywords:</strong> ME/CFS, chronic fatigue syndrome, muscle symptoms, physiological dysregulation, post-exertional malaise, structural equation modeling, factor analysis, sex differences, menopausal status, autonomic dysfunction, symptom phenotyping, Journal of Translational Medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223254</post-id>	</item>
		<item>
		<title>The Standard Test for Female Sexual Function Holds Up in Pregnancy—But Only in Part</title>
		<link>https://scienmag.com/the-standard-test-for-female-sexual-function-holds-up-in-pregnancy-but-only-in-part/</link>
		
		<dc:creator><![CDATA[Harold Sullivan]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 22:27:50 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[Archives of Sexual Behavior]]></category>
		<category><![CDATA[clinical implications of FSFI in pregnancy]]></category>
		<category><![CDATA[development of reliable tools for assessing female sexual health]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[Female sexual function assessment during pregnancy]]></category>
		<category><![CDATA[Female Sexual Function Index]]></category>
		<category><![CDATA[impact of childbirth on sexual function]]></category>
		<category><![CDATA[item response theory]]></category>
		<category><![CDATA[measurement invariance]]></category>
		<category><![CDATA[measurement of desire and arousal during pregnancy]]></category>
		<category><![CDATA[pain and orgasm issues in pregnant women]]></category>
		<category><![CDATA[perinatal health]]></category>
		<category><![CDATA[postpartum]]></category>
		<category><![CDATA[postpartum sexual health]]></category>
		<category><![CDATA[Pregnancy]]></category>
		<category><![CDATA[pregnancy-related sexual dysfunction]]></category>
		<category><![CDATA[psychometric validation]]></category>
		<category><![CDATA[psychometric validation of sexual health questionnaires]]></category>
		<category><![CDATA[sexual dysfunction]]></category>
		<category><![CDATA[sexual function]]></category>
		<category><![CDATA[sexual health research in perinatal populations]]></category>
		<category><![CDATA[sexual problems in perinatal period]]></category>
		<category><![CDATA[validation of Female Sexual Function Index in pregnant women]]></category>
		<category><![CDATA[validity theory]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216721</guid>

					<description><![CDATA[A new validation study finds that the six domain scores of the Female Sexual Function Index are meaningful for pregnant and postpartum individuals, but the widely used total score is not supported in these populations.]]></description>
										<content:encoded><![CDATA[<p>Sexual problems are strikingly common during pregnancy and after childbirth, with prevalence estimates in the perinatal period ranging from 36 to 88 percent. Low desire, difficulties with arousal and lubrication, pain during intercourse, and orgasm problems are all frequently reported by expectant and new mothers. Yet despite decades of research relying on a single questionnaire to measure these experiences—the Female Sexual Function Index, or FSFI—almost nobody had rigorously tested whether the instrument actually works in pregnant and postpartum populations. A new validation study published in Archives of Sexual Behavior by Pablo Santos-Iglesias of Cape Breton University, Natalie O. Rosen of Dalhousie University, and Samantha J. Dawson of the University of British Columbia set out to answer that question, and its findings carry important implications for both researchers and clinicians.</p>
<p>The FSFI has been a workhorse of sexual medicine since its introduction in 2000. The 19-item self-report questionnaire assesses six domains of sexual function over the previous four weeks: desire, arousal, lubrication, orgasm, satisfaction, and pain. Respondents rate each item on a scale, and the domain scores are typically summed into a total score, with a widely used clinical cutoff distinguishing women with and without sexual dysfunction. The instrument was developed and validated in general community and clinical samples of adult women, however—not in people navigating the profound hormonal, anatomical, and relational upheavals of pregnancy and the postpartum period. Whether its scores mean the same thing in those populations had remained largely an open question.</p>
<p>To address it, the research team adopted a contemporary, argument-based approach to validity, guided by modern frameworks that treat validity not as a property of the test itself but of the interpretations and decisions built on its scores. Rather than asking simply whether the FSFI is valid, the researchers examined whether specific claims about what FSFI scores mean in perinatal samples are supported by evidence. This matters because perinatal sexual function is not simply typical sexual function in a different context. Hormonal shifts, tissue changes from delivery, breastfeeding, fatigue, and the psychological transition to parenthood all reshape sexual response in ways that a questionnaire designed for other populations may or may not capture.</p>
<p>The centerpiece of the investigation was a series of factor analyses, statistical techniques that test whether the pattern of people&#8217;s answers matches the structure the questionnaire assumes. The results offered a clear verdict on one front: the six-domain structure of the FSFI—desire, arousal, lubrication, orgasm, satisfaction, and pain—was supported in sexually active pregnant and postpartum individuals. In other words, the questionnaire&#8217;s underlying architecture holds together in this population, and each domain appears to tap a distinct, coherent aspect of sexual function. That finding gives researchers license to continue using and interpreting the six domain scores with reasonable confidence in perinatal research.</p>
<p>The same could not be said for the total score. Models that assumed a single overarching dimension of sexual function underlying all the items—essentially, the assumption baked into summing everything into one number—were not supported by the data. This is a technically significant result. When a unidimensional model fails, it means the domains are not interchangeable reflections of one common construct, and a total score risks blending genuinely different phenomena into a single, potentially misleading figure. A person might score low overall because of pain alone, for example, while their desire, arousal, and satisfaction remain intact—clinically meaningful distinctions that a total score would obscure. The authors&#8217; conclusion on this point is unambiguous: the use of an FSFI total score is not supported in pregnant and postpartum samples.</p>
<p>Item response theory analyses added a second layer of nuance. These models examine how well each individual item distinguishes between people at different levels of the underlying trait. The findings showed that FSFI items discriminated well at low-to-average levels of sexual function but performed less effectively at higher levels. In practical terms, the questionnaire is most informative for pregnant and postpartum individuals who have poorer sexual function or who are experiencing sexual difficulties. For people functioning well, the items become less precise, meaning small differences among high-functioning individuals may not be reliably captured. This has a silver lining for clinical and research applications focused on identifying and understanding sexual problems, but it cautions against over-interpreting fine-grained differences among people reporting generally healthy sexual function.</p>
<p>The researchers also examined how FSFI domain scores relate to external variables, a standard strategy for testing whether a measure behaves the way theory says it should. With a few exceptions, this evidence indicated that the domain scores are sensitive to real individual differences in sexual function and can distinguish between individuals with and without distressing sexual difficulties. That is, the domains correlate with other measures in sensible ways and show the kind of clinical sensitivity that a useful assessment tool requires. Two domains, however, showed weaker performance: sexual arousal and sexual satisfaction. The authors advise that results from these two domains, like the total score, be interpreted with particular caution in perinatal samples until further evidence accumulates.</p>
<p>The stakes of this psychometric housekeeping are higher than they might appear. Perinatal sexual dysfunction is associated with depression symptoms, relationship dissatisfaction, and distress for both members of a couple, and researchers have increasingly documented trajectories of sexual well-being across the transition to parenthood. Interventions are being developed to support couples&#8217; sexual health during this period, and clinical conversations about resuming intercourse after childbirth depend on accurate assessment. If the field&#8217;s dominant measurement tool were quietly mismeasuring these constructs, studies could reach distorted conclusions, clinical cutoffs could misclassify people, and patients&#8217; experiences could be misunderstood. This study provides a foundation of evidence showing which parts of the instrument can be trusted in this population and which cannot.</p>
<p>The work also exemplifies a broader shift in how psychological and sexual health measures are evaluated. Older validation studies often reported a handful of statistics and declared an instrument valid or invalid. Contemporary validity theory, by contrast, builds an integrated argument, weighing evidence from internal structure, item-level performance, and relations to external variables to evaluate specific claims about score interpretation. By applying this framework to the FSFI, the researchers modeled a more rigorous standard for sexuality research, where measurement quality has historically received less scrutiny than in other areas of psychology. The study&#8217;s data and analysis code were also made publicly available through the Open Science Framework, supporting transparency and reuse.</p>
<p>For now, the practical guidance is clear. Researchers studying pregnancy and the postpartum period can continue to use the FSFI&#8217;s six domain scores as meaningful indicators of the underlying constructs, and the measure appears well suited to identifying individuals with sexual difficulties. But the familiar habit of reporting a single total FSFI score should end in perinatal research, and the arousal and satisfaction domains warrant cautious interpretation. The authors note that further research is needed to address the limitations identified, including the weaker performance at higher levels of function. Until that evidence arrives, this study offers the field something it has long lacked: a clear, evidence-based map of what one of sexual medicine&#8217;s most widely used questionnaires can—and cannot—reliably tell us about sexual health during one of life&#8217;s most transformative transitions.</p>
<p><strong>Subject of Research:</strong> Validation of the Female Sexual Function Index for measuring sexual function in pregnant and postpartum individuals</p>
<p><strong>Article Title:</strong> A Validation Study of the Female Sexual Function Index for Use in Pregnant and Postpartum Samples</p>
<p><strong>Article References:</strong> Santos-Iglesias, P., Rosen, N. O., &amp; Dawson, S. J. (2026). A Validation Study of the Female Sexual Function Index for Use in Pregnant and Postpartum Samples. <em>Archives of Sexual Behavior, 55</em>(6), 2523-2541. <a href="https://doi.org/10.1007/s10508-026-03519-w" rel="noopener noreferrer">https://doi.org/10.1007/s10508-026-03519-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10508-026-03519-w" rel="noopener noreferrer">10.1007/s10508-026-03519-w</a></p>
<p><strong>Keywords:</strong> Female Sexual Function Index, sexual function, pregnancy, postpartum, psychometric validation, factor analysis, item response theory, sexual dysfunction, perinatal health, validity theory, Archives of Sexual Behavior, measurement invariance</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216721</post-id>	</item>
		<item>
		<title>Hidden Wording Effects Distort Psychology Scores, But a New Statistical Fix Emerges</title>
		<link>https://scienmag.com/hidden-wording-effects-distort-psychology-scores-but-a-new-statistical-fix-emerges/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 00:39:29 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[confirmatory factor analysis]]></category>
		<category><![CDATA[dimensionality]]></category>
		<category><![CDATA[enhancing validity of psychological scales]]></category>
		<category><![CDATA[Exploratory Graph Analysis]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[impact of question phrasing on survey responses]]></category>
		<category><![CDATA[improving psychological assessment accuracy]]></category>
		<category><![CDATA[influence of positive and negative item wording]]></category>
		<category><![CDATA[method variance]]></category>
		<category><![CDATA[parallel analysis]]></category>
		<category><![CDATA[psychological measurement]]></category>
		<category><![CDATA[psychometric method variance]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[questionnaire bias correction methods]]></category>
		<category><![CDATA[random intercept item factor analysis]]></category>
		<category><![CDATA[randomized intercept item factor analysis]]></category>
		<category><![CDATA[research on survey response distortions]]></category>
		<category><![CDATA[Short Grit Scale]]></category>
		<category><![CDATA[statistical techniques in psychometrics]]></category>
		<category><![CDATA[structural validity]]></category>
		<category><![CDATA[survey design bias]]></category>
		<category><![CDATA[survey methodology]]></category>
		<category><![CDATA[wording effect in psychological questionnaires]]></category>
		<category><![CDATA[wording effects]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215687</guid>

					<description><![CDATA[A new study in Behavior Research Methods shows that random intercept item factor analysis can separate genuine trait variance from wording artifacts in mixed-worded psychological scales, using the Short Grit Scale as a test case.]]></description>
										<content:encoded><![CDATA[<p>Every year, millions of people fill out psychological questionnaires designed to measure everything from grit and self-esteem to anxiety and burnout. These instruments shape hiring decisions, clinical diagnoses, and entire research literatures. Yet a subtle flaw has haunted survey design for decades: when questionnaires mix positively and negatively worded items, respondents often produce answers that reflect the phrasing of the questions rather than the trait being measured. A new study published in Behavior Research Methods by Palmira Faraci and Giuliana Nasonte of the Psychometrics Laboratory at the University Kore of Enna puts a sophisticated statistical tool, random intercept item factor analysis, to the test, and the results suggest it can cleanly separate genuine personality signals from the noise created by question wording.</p>
<p>The problem, known in psychometrics as method variance or the wording effect, arises because scales are frequently built with a deliberate balance of positively and negatively keyed items. The idea sounds sensible: reversing the direction of some questions should discourage automatic agreement and force respondents to read carefully. But the strategy backfires in a measurable way. Items that share the same wording direction tend to correlate with one another for reasons that have nothing to do with the underlying construct. When researchers run a factor analysis on such data, the statistical machinery often obligingly splits the scale into two clusters, one of positively worded items and one of negatively worded items, rather than the single dimension the scale was designed to capture.</p>
<p>This artificial splitting, sometimes called spurious bidimensionality, has real consequences. A researcher who concludes that a questionnaire measures two distinct traits may publish misleading findings, build theories on phantom subfactors, or compute reliability estimates that are inflated or deflated for the wrong reasons. The grit scale, a wildly popular measure of perseverance and passion for long-term goals, has been a particular battleground for this debate. Since its development by Angela Duckworth and colleagues, researchers have argued endlessly about whether grit truly has two facets, perseverance of effort and consistency of interest, or whether the apparent two-factor structure is simply an artifact of the negatively worded items embedded in the short version of the scale.</p>
<p>Faraci and Nasonte tackled this question with random intercept item factor analysis, or RIIFA, a modeling approach originally introduced by Albert Maydeu-Olivares and Donna Coffman in 2006. The core insight of RIIFA is elegant: it adds a random intercept factor to the standard factor model, a latent variable on which every item in the scale loads regardless of its wording direction. This intercept factor absorbs the shared response tendency that cuts across all items, including acquiescence, the general willingness to agree with statements, and other response styles. Once that pervasive method variance is siphoned off into its own latent dimension, the remaining substantive factor can be interpreted as the true trait, purified of the contamination introduced by how the questions happen to be phrased.</p>
<p>To evaluate whether RIIFA delivers on this promise, the researchers conducted two studies using two independent samples from the United Kingdom. The first study analyzed data from 977 participants, and the second from 496. Their test case was the Short Grit Scale, known as Grit-S, an eight-item instrument that mixes positively and negatively worded questions. The choice was strategic: the Grit-S is short, widely used, and famously prone to the wording-effect problem, making it an ideal proving ground for a method intended to restore structural validity to mixed-worded scales.</p>
<p>The first study focused on dimensionality assessment, the task of determining how many latent factors actually underlie a set of items. The researchers compared traditional factor retention techniques, including parallel analysis, a classic method that compares observed eigenvalues against those from random data, with RIIFA-based counterparts that incorporate the random intercept factor. They also employed exploratory graph analysis, a newer network psychometrics approach that treats items as nodes in a network and uses community detection algorithms, borrowed originally from network science methods like the fast unfolding of communities, to identify clusters of tightly connected items. The results were striking. Traditional retention methods consistently overestimated the number of factors, suggesting the Grit-S contained two dimensions when the theoretical expectation was one. In contrast, the RIIFA-based techniques produced unidimensional solutions, and bootstrap analyses, which resample the data thousands of times to gauge stability, showed that these solutions were considerably more stable than those obtained without controlling for the wording effect.</p>
<p>The second study shifted from exploration to confirmation, using confirmatory factor analysis to formally test competing models of the scale&#8217;s structure. The researchers fit models with and without a random intercept factor and compared them on both fit and parsimony, the principle that simpler explanations should be preferred unless complexity earns its keep. The model incorporating the random intercept factor achieved the best balance between the two. Its fit statistics were excellent by conventional standards: a root mean square error of approximation of .048, with a confidence interval spanning .022 to .072, a comparative fit index of .984, a Tucker-Lewis index of .974, and a standardized root mean square residual of .027. Values close to .95 or above on the fit indices and below about .06 or .08 on the error measures are typically considered indicative of a well-fitting model, and this model cleared every bar comfortably.</p>
<p>Reliability told a similarly encouraging story. The RIIFA model yielded a hierarchical omega of .84, a coefficient derived from bifactor-style measurement models that estimates how well the total scale score reflects a single dominant common factor after accounting for the variance absorbed by subsidiary dimensions. In practical terms, this means that once the method variance was reallocated to the random intercept factor, the substantive grit factor explained enough common variance to support meaningful interpretation of total scores. The researchers interpret this as evidence that RIIFA does not merely hide the problem; it actively redistributes explained variance between the substantive factor and the method factor, mitigating the artificial bidimensionality that plagues conventional analyses and enhancing the interpretability of the latent structure.</p>
<p>The implications reach well beyond the grit scale. Mixed-worded instruments are ubiquitous across psychology, appearing in measures of self-esteem, core self-evaluations, need for cognition, perceived stress, learning burnout, and countless clinical and organizational questionnaires. For each of these, the same dilemma recurs: researchers must decide whether an emerging second factor represents a genuine substantive distinction or a methodological artifact. Historically, that judgment has been made with tools, such as standard parallel analysis and conventional confirmatory factor analysis, that have no built-in mechanism for separating content from phrasing. The findings of this study suggest that RIIFA offers a principled alternative, and the authors explicitly recommend its application in cases where wording effects threaten the validity of psychometric measurement.</p>
<p>The study also connects to a broader movement in quantitative psychology toward modeling response styles and careless responding rather than simply hoping they wash out in large samples. Recent work has documented how even a few inconsistent respondents can confound the structure of personality survey data, how acquiescence distorts exploratory factor analysis, and how attention checks and response-style detection methods carry their own complications. Faraci and Nasonte&#8217;s contribution fits squarely into this research program, demonstrating on real data that a model explicitly designed to partition substantive and method variance can outperform traditional approaches. Importantly, the authors have made their data openly available through the Open Science Framework, and all of the R and Mplus code used in the analyses is accessible as well, lowering the barrier for other researchers to adopt the technique. For a field that depends so heavily on self-report questionnaires, a validated method for ensuring that scores reflect traits rather than question phrasing is not a technical nicety. It is a foundation for the credibility of the measurements themselves, and this study provides compelling empirical grounds for adding random intercept item factor analysis to the standard psychometric toolkit.</p>
<p><strong>Subject of Research:</strong> Using random intercept item factor analysis to separate substantive variance from wording-effect method variance in mixed-worded psychological scales</p>
<p><strong>Article Title:</strong> Disentangling substantive and method variance in mixed-worded scales: An empirical application of the random intercept item factor analysis (RIIFA)</p>
<p><strong>Article References:</strong> Faraci, P., &amp; Nasonte, G. (2026). Disentangling substantive and method variance in mixed-worded scales: An empirical application of the random intercept item factor analysis (RIIFA). <em>Behavior Research Methods, 58</em>(11), Article 299. <a href="https://doi.org/10.3758/s13428-026-03163-1" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03163-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03163-1" rel="noopener noreferrer">10.3758/s13428-026-03163-1</a></p>
<p><strong>Keywords:</strong> psychometrics, random intercept item factor analysis, wording effects, method variance, Short Grit Scale, factor analysis, exploratory graph analysis, parallel analysis, dimensionality, structural validity, confirmatory factor analysis, survey methodology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215687</post-id>	</item>
		<item>
		<title>New Turkish Questionnaire Captures the Full Spectrum of Emotional Eating</title>
		<link>https://scienmag.com/new-turkish-questionnaire-captures-the-full-spectrum-of-emotional-eating/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:27:09 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[comprehensive emotional eating scale]]></category>
		<category><![CDATA[cultural adaptation of eating behavior questionnaires]]></category>
		<category><![CDATA[eating behavior and emotional states]]></category>
		<category><![CDATA[eating disorders]]></category>
		<category><![CDATA[emotion regulation]]></category>
		<category><![CDATA[emotional eating]]></category>
		<category><![CDATA[emotional eating in response to negative emotions]]></category>
		<category><![CDATA[emotional eating in response to positive emotions]]></category>
		<category><![CDATA[emotional eating measurement]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[impact of emotions on eating patterns]]></category>
		<category><![CDATA[Journal of Eating Disorders]]></category>
		<category><![CDATA[measurement invariance]]></category>
		<category><![CDATA[multidimensional emotional eating assessment]]></category>
		<category><![CDATA[overeating]]></category>
		<category><![CDATA[psychological assessment of emotional eating]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[research on emotional eating in Turkey]]></category>
		<category><![CDATA[scale validation]]></category>
		<category><![CDATA[test-retest reliability]]></category>
		<category><![CDATA[Turkish adults]]></category>
		<category><![CDATA[Turkish emotional eating questionnaire]]></category>
		<category><![CDATA[undereating]]></category>
		<category><![CDATA[validation of emotional eating tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213307</guid>

					<description><![CDATA[Turkish researchers have validated the Comprehensive Emotional Eating Scale in 897 adults, producing the first Turkish questionnaire to measure overeating and undereating in response to both positive and negative emotions.]]></description>
										<content:encoded><![CDATA[<p>When people reach for ice cream after a breakup or lose their appetite entirely before a job interview, they are exhibiting what psychologists call emotional eating. For decades, the scientific literature has treated this behavior as a single phenomenon, usually defined as eating more in response to negative feelings. A new study published in the Journal of Eating Disorders argues that this picture is far too narrow, and it offers Turkish researchers the first validated tool capable of measuring emotional eating in all of its complexity. The work, led by Beyza Katırcıoğlu and Murat Baş of Acibadem Mehmet Ali Aydinlar University in Istanbul, translated and tested the Comprehensive Emotional Eating Scale, or CEES, in a large sample of Turkish adults, and the results suggest the instrument performs remarkably well.</p>
<p>The conceptual foundation of the CEES is deceptively simple but represents a significant departure from earlier measures. Emotional eating, in this framework, is a multidimensional construct with four distinct dimensions: overeating in response to positive emotions, overeating in response to negative emotions, undereating in response to positive emotions, and undereating in response to negative emotions. Most existing questionnaires capture only the negative-overeating quadrant, which means researchers have historically been blind to people who celebrate with food, who suppress their appetite when distressed, or who cannot eat when they are happy. The Turkish validation study is the first to bring this complete four-dimensional framework to a Turkish-speaking population, filling what the authors describe as a genuine gap in the available assessment instruments.</p>
<p>To test the scale, the research team recruited 897 healthy adults between the ages of 18 and 65. Methodological studies of this kind require careful statistical architecture, and the investigators split their sample into two demographically matched subsamples of 449 and 448 participants. The first subsample was used for exploratory factor analysis, a technique that allows the underlying structure of a questionnaire to emerge from the data rather than being imposed in advance. The second subsample was reserved for confirmatory factor analysis, which tests whether the structure discovered in the first half holds up in an independent group of respondents. This two-step design is considered best practice in psychometrics because it guards against the common pitfall of building and validating a model on the same data.</p>
<p>The exploratory analysis delivered a clean verdict. The Turkish CEES reproduced the expected four-factor structure, corresponding to positive overeating, positive undereating, negative overeating, and negative undereating, and these four factors together accounted for 70.0 percent of the total variance in responses. In questionnaire research, explaining seven-tenths of the variance is an unusually strong result, indicating that the items cluster tightly around the intended dimensions. Internal consistency, which reflects how well the items within each subscale measure the same underlying construct, was excellent across the board, with Cronbach&#8217;s alpha values ranging from 0.948 to 0.962. Values above 0.90 are typically regarded as excellent, and these figures place the Turkish CEES among the most internally consistent instruments in the emotional eating literature.</p>
<p>Reliability over time was assessed by asking a subset of participants to complete the questionnaire a second time, two weeks after their first administration. The intraclass correlation coefficients for the four subscales ranged from 0.80 to 0.84, and the total score achieved an intraclass correlation of 0.91, indicating that people&#8217;s answers remained stable across the two-week interval. This test-retest stability matters because emotional eating is conceptualized as a relatively enduring behavioral tendency rather than a fleeting mood-dependent state. A questionnaire whose results swing wildly from one fortnight to the next would be of limited use for tracking individuals in research studies or, eventually, in clinical practice.</p>
<p>The confirmatory stage of the study involved a sophisticated comparison of competing statistical models. The researchers tested a standard confirmatory factor model, an exploratory structural equation model, and a bifactor exploratory structural equation model, the latter allowing for both a general emotional eating factor and specific subscale factors simultaneously. The bifactor model that incorporated three additional specific factors for semantically overlapping items demonstrated the best fit to the data, with a comparative fit index of 0.996, a Tucker-Lewis index of 0.994, a root mean square error of approximation of 0.089, and a standardized root mean square residual of 0.043. Fit indices close to or above 0.95 for CFI and TLI, and below 0.08 for SRMR, are conventionally interpreted as indicating excellent model fit. Notably, the bifactor solution supported the four subscale factors alongside only a weak general factor, which carries an important implication: emotional eating in this framework is best understood as four related but distinct patterns rather than a single unified tendency, and researchers should analyze the subscales separately rather than relying on a total score alone.</p>
<p>A further question in cross-cultural psychometrics is whether a questionnaire functions equivalently across different groups. If an instrument measures something different in men than in women, or in younger versus older adults, comparisons between those groups become meaningless. The Turkish CEES team therefore tested measurement invariance across gender, age, marital status, and body mass index categories. At the subscale level, changes in fit indices were generally consistent with metric and scalar invariance across sex, marital status, and age, meaning the scale measures the same constructs in the same way across these groups. The authors do, however, urge caution in two respects: the baseline models showed limited absolute fit, and the comparisons across body mass index categories were less straightforward, so BMI-related group comparisons should be interpreted with care until further evidence accumulates.</p>
<p>Convergent validity, the demonstration that a new scale correlates with other measures in theoretically expected ways, was established through correlations with four established instruments: the 13-item Eating Disorder Examination Questionnaire, the eight-item Depression Anxiety and Stress Scale, the 16-item Difficulties in Emotion Regulation Scale, and the 14-item Perceived Stress Scale. The Turkish CEES scores exhibited the theoretically consistent associations with eating disorder psychopathology, emotional dysregulation, stress, and psychological distress. In other words, people who reported more emotional eating of a given type also tended to report more eating-related problems, greater difficulty managing their emotions, and higher levels of stress and psychological distress, exactly the pattern that theories of emotional eating would predict.</p>
<p>The practical significance of this validation extends well beyond the statistics. Türkiye has a growing community of researchers working on obesity, disordered eating, and the psychological drivers of dietary behavior, but until now they have lacked a Turkish-language instrument that captures the full positive and negative, overeating and undereating spectrum. With the CEES now available in Turkish, studies can distinguish, for example, between someone who eats excessively when anxious and someone who stops eating when celebrating, two behaviors with potentially very different metabolic and psychological consequences. The four-dimensional profile could also help clinicians and dietitians understand why some patients gain weight under stress while others lose it, a distinction that single-dimension scales have obscured.</p>
<p>The authors are careful to note the limits of their work. The validation was conducted in healthy adults, and the clinical utility of the Turkish CEES in patient populations, such as individuals with diagnosed eating disorders or obesity undergoing treatment, requires further validation before the scale can be used confidently in clinical settings. The study also received no specific external funding, and the authors declare no competing interests. Even with those caveats, the study represents a substantial methodological contribution: a rigorously tested, four-dimensional, psychometrically strong Turkish measure of emotional eating, validated in nearly nine hundred adults with state-of-the-art factor analytic techniques. For a field that has long relied on instruments capturing only a fraction of the emotional eating phenomenon, the arrival of a comprehensive Turkish-language tool is likely to reshape how researchers in Türkiye study the tangled relationship between feelings and food.</p>
<p><strong>Subject of Research:</strong> Validation of the Turkish version of the Comprehensive Emotional Eating Scale, a four-dimensional measure of emotional eating</p>
<p><strong>Article Title:</strong> Psychometric properties of the Turkish version of the comprehensive emotional eating scale</p>
<p><strong>Article References:</strong> Psychometric properties of the Turkish version of the comprehensive emotional eating scale. (n.d.). <a href="https://doi.org/10.1186/s40337-026-01773-w" rel="noopener noreferrer">https://doi.org/10.1186/s40337-026-01773-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40337-026-01773-w" rel="noopener noreferrer">10.1186/s40337-026-01773-w</a></p>
<p><strong>Keywords:</strong> emotional eating, psychometrics, scale validation, overeating, undereating, eating disorders, emotion regulation, factor analysis, Turkish adults, measurement invariance, test-retest reliability, Journal of Eating Disorders</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213307</post-id>	</item>
		<item>
		<title>New Knee Surgery Scorecard Captures What Matters to Chinese Patients</title>
		<link>https://scienmag.com/new-knee-surgery-scorecard-captures-what-matters-to-chinese-patients/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 22:49:30 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[China]]></category>
		<category><![CDATA[Chinese knee surgery recovery assessment]]></category>
		<category><![CDATA[Cronbach's alpha]]></category>
		<category><![CDATA[cross-cultural adaptation]]></category>
		<category><![CDATA[culturally tailored knee arthroplasty questionnaires]]></category>
		<category><![CDATA[development of Chinese-specific orthopedic outcome scores]]></category>
		<category><![CDATA[factor analysis]]></category>
		<category><![CDATA[functional recovery after knee replacement in China]]></category>
		<category><![CDATA[impact of culturally relevant assessment tools in orthopedic care]]></category>
		<category><![CDATA[importance of patient perspectives in knee surgery outcomes]]></category>
		<category><![CDATA[improvements in postoperative quality of life measurement]]></category>
		<category><![CDATA[knee osteoarthritis]]></category>
		<category><![CDATA[Knee replacement patient-reported outcome measures]]></category>
		<category><![CDATA[measuring kneeling and squatting ability post-knee surgery]]></category>
		<category><![CDATA[orthopedic surgery]]></category>
		<category><![CDATA[patient-centered knee surgery evaluation tools]]></category>
		<category><![CDATA[patient-reported outcome measures]]></category>
		<category><![CDATA[PROMs in Chinese joint surgery]]></category>
		<category><![CDATA[psychometric validation]]></category>
		<category><![CDATA[Quality of Life]]></category>
		<category><![CDATA[responsiveness]]></category>
		<category><![CDATA[total knee arthroplasty]]></category>
		<category><![CDATA[validation of Chinese knee surgery PROMs]]></category>
		<category><![CDATA[WOMAC]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=210962</guid>

					<description><![CDATA[Researchers in China have developed and validated a new patient-reported outcome measure for knee replacement surgery built around culturally relevant activities such as squatting, Tai Chi, and rising from low seating.]]></description>
										<content:encoded><![CDATA[<p>When surgeons replace a worn-out knee joint, the technical success of the operation is only half the story. What patients really want to know is whether they can kneel to pray, squat to tend a vegetable plot, or join the nightly square-dancing crowd in the park without wincing. A new study published in The Lancet Regional Health – Western Pacific argues that the questionnaires used worldwide to measure recovery after total knee arthroplasty often miss exactly those moments, particularly for patients in China. The research, led by Chao Xu, Xiaofeng Chang, and colleagues at hospitals in Xi&#8217;an, Chengdu, and Lanzhou, introduces and rigorously tests a home-grown measurement tool built from the ground up with Chinese patients, rather than translated from Western originals.</p>
<p>The instrument, called the Chinese Total Knee Arthroplasty Patient-reported Outcome Measurement, or CTP, belongs to a class of questionnaires known as patient-reported outcome measures, or PROMs. Instead of relying on clinicians to grade range of motion on an X-ray or flexion angle with a goniometer, PROMs ask patients directly about pain, function, mood, and quality of life. Proponents argue this approach captures the outcomes that matter most, improves communication between doctor and patient, and detects meaningful improvement that physician-scored measures can overlook. Yet the field has a persistent blind spot: most widely used knee PROMs, including the Oxford Knee Score, the new Knee Society Score, and the WOMAC index, were developed in Western populations and only later adapted for Chinese speakers, often with limited attention to whether the underlying activities and social situations are actually relevant to the people answering them.</p>
<p>The stakes are enormous. Symptomatic knee osteoarthritis affects roughly 14.6 percent of adults in China, and the number of total knee replacement procedures grew at an annual rate of 27.43 percent between 2012 and 2019. Despite this surge, between 10 and 20 percent of patients report dissatisfaction after surgery, a figure that suggests recovery is not simply a matter of a technically flawless implant. The research team&#8217;s earlier searches of PubMed, Embase, Scopus, and major Chinese databases found no previously validated, de novo, multidomain PROM built specifically for Chinese knee replacement patients — a gap the new study set out to fill.</p>
<p>The project unfolded in stages. In prior qualitative work, the investigators built an initial pool of 135 candidate items through patient interviews, concept elicitation with a multidisciplinary panel of 17 joint surgeons, orthopedic nurses, rehabilitation therapists, statisticians, and psychologists, two rounds of Delphi expert consultation, and cognitive debriefing with patients. That process produced a six-domain, 35-item preliminary questionnaire. The new study then put the instrument through quantitative refinement across three sequential, independent cohorts totaling 1,748 participants recruited between June 2023 and June 2024, with twelve-month follow-up completed by June 2025. No patient appeared in more than one cohort, a design choice that strengthens the validity of each analytic phase.</p>
<p>The item-reduction statistics tell a striking story about cultural fit. Questionnaires only work if the answers spread out; if nearly everyone gives the same response, the item carries almost no information. Three items asking about preoperative expectations showed ceiling concentrations of 94.7, 84.7, and 56.7 percent, meaning the vast majority of patients gave the top answer, and the entire expectations domain was dropped. A jogging item floored at 71.7 percent ceiling, while questions about embarrassment when others see the knee produced floor effects. At the same time, activities central to Chinese daily life — squatting, rising from low seating without armrests, mild field labour, climbing stairs in buildings without elevators, and community exercise such as Tai Chi or square dancing — loaded strongly onto the scale and stayed in. The result is a final 22-item clinical questionnaire with ten bilateral symptom items, eight function items, three quality-of-life items, and a single global satisfaction question.</p>
<p>For the formal psychometric testing, the team concentrated on a 16-item operated-knee core score covering symptoms, function, and quality of life, deliberately scoring only the replaced joint to avoid contamination from the other arthritic knee. The reliability figures were impressive. Cronbach&#8217;s alpha, a measure of internal consistency that reflects how coherently items hang together, reached 0.864 for the total core score, with subscale values between 0.809 and 0.942. When 150 participants filled out the questionnaire twice, two weeks apart and before any surgery, the intraclass correlation coefficient was 0.836, indicating that scores are stable over time rather than bouncing with mood or measurement noise.</p>
<p>Validity results were more nuanced. Exploratory factor analysis on a random half of the final cohort supported a clean three-factor structure matching Symptoms, Function, and Quality of Life, with all items loading above 0.40 and the three factors together explaining 72.2 percent of variance. But confirmatory factor analysis on the other half told a partially different story: absolute fit indices such as the SRMR (0.05), RMSEA (0.04), and the chi-square ratio (2.13) met prespecified thresholds, while the comparative fit index (0.92) and non-normed fit index (0.91) fell just short of the strict 0.95 criterion. The authors are candid about this, describing the structure as receiving partial rather than uniform support that requires independent replication — an unusually honest framing in a literature where fit indices are sometimes selectively reported.</p>
<p>The team also benchmarked the new measure against WOMAC, the most entrenched Western instrument. Convergent validity was strong where constructs genuinely overlap: the CTP function domain correlated with WOMAC function at r = 0.698, and CTP symptoms tracked WOMAC pain at r = 0.605, exactly the pattern expected if both instruments are measuring the same underlying phenomena. The weak correlation (r = 0.355) between CTP quality of life and WOMAC stiffness was not a failure but a confirmation that the two domains are conceptually distinct — stiffness is not a proxy for life satisfaction. Most importantly, in the responsiveness analysis, the CTP total score dropped from 33.52 at baseline to 21.81 one year after surgery, an effect size of 1.34 and a standardised response mean of 1.66, both comfortably in the large-effect range. WOMAC produced even larger standardised changes, but the authors caution that effect-size magnitude depends on an instrument&#8217;s content and score distribution and should not be read as a head-to-head superiority contest. What the CTP offers is complementary coverage of symptoms, psychosocial impact, and culturally salient activities that WOMAC simply does not ask about.</p>
<p>The study&#8217;s limitations are clearly acknowledged. All three hospitals were tertiary centres in western China, so the findings establish multicentre evidence within that region rather than nationwide generalisability; dialect, occupation, healthcare access, and community habits vary considerably across the country. No recognised Chinese gold-standard knee PROM existed, making criterion validity impossible to assess, and the quality-of-life domain still awaits validation against independent generic health or mental-health measures. Future work, the authors write, should test measurement invariance across diverse regions, establish anchor-based interpretability thresholds so that clinicians know what a given score change actually means, and evaluate whether the CTP adds predictive information beyond established instruments.</p>
<p>Even so, the implications reach well beyond orthopedics. Roughly half of the participants reported primary education or less, and nearly three-quarters relied on rural cooperative medical insurance — populations routinely underrepresented in instrument development. The CTP demonstrates that a psychometrically sound measure can be built from patient interviews upward, tested in more than 1,700 real patients across multiple centres, and refined with transparent statistical criteria rather than ad hoc judgment. It takes patients an average of under five minutes to complete. As knee replacement volumes continue their steep climb in China and across Asia&#8217;s aging societies, the study makes a forceful case that outcome measurement should be grounded in the lives patients actually lead — whether that means a walk around a Western suburb or a deep squat in a courtyard garden. The questionnaire does not replace WOMAC and its peers, but it may finally ask the right questions of the people answering them.</p>
<p><strong>Subject of Research:</strong> Development and multicentre psychometric validation of a culturally tailored patient-reported outcome measure for Chinese total knee arthroplasty patients</p>
<p><strong>Article Title:</strong> Quantitative refinement and multicentre psychometric validation of the Chinese total knee arthroplasty patient-reported outcome measure: a prospective cohort study with repeated measures</p>
<p><strong>Article References:</strong> Xu, C., Chang, X., Ye, F., Wu, H., Chen, G., Zhang, R., Yao, S., Shang, L., &amp; Ma, J. (2026). Quantitative refinement and multicentre psychometric validation of the Chinese total knee arthroplasty patient-reported outcome measure: a prospective cohort study with repeated measures. <em>The Lancet Regional Health &#8211; Western Pacific, 74</em>, Article 101989. <a href="https://doi.org/10.1016/j.lanwpc.2026.101989" rel="noopener noreferrer">https://doi.org/10.1016/j.lanwpc.2026.101989</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.lanwpc.2026.101989" rel="noopener noreferrer">10.1016/j.lanwpc.2026.101989</a></p>
<p><strong>Keywords:</strong> total knee arthroplasty, patient-reported outcome measures, psychometric validation, knee osteoarthritis, Cronbach&#x27;s alpha, WOMAC, cross-cultural adaptation, responsiveness, factor analysis, China, orthopedic surgery, quality of life</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">210962</post-id>	</item>
	</channel>
</rss>
