<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Structured observation tools &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/structured-observation-tools/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 13:26:07 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Structured observation tools &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Finnish Work Performance Test Passes Its First Big Statistical Stress Check</title>
		<link>https://scienmag.com/finnish-work-performance-test-passes-its-first-big-statistical-stress-check/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 13:26:07 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Assessment of Work Performance (AWP)]]></category>
		<category><![CDATA[assessment tools]]></category>
		<category><![CDATA[AWP-FI]]></category>
		<category><![CDATA[construct validity]]></category>
		<category><![CDATA[differential item functioning]]></category>
		<category><![CDATA[Finland]]></category>
		<category><![CDATA[Finnish occupational therapy studies]]></category>
		<category><![CDATA[Finnish work performance test]]></category>
		<category><![CDATA[Job ability evaluation]]></category>
		<category><![CDATA[Long-term sick leave recovery]]></category>
		<category><![CDATA[Model of Human Occupation]]></category>
		<category><![CDATA[Occupational health assessments]]></category>
		<category><![CDATA[occupational therapy]]></category>
		<category><![CDATA[Occupational therapy evaluation]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[Rasch analysis]]></category>
		<category><![CDATA[rehabilitation assessment tools]]></category>
		<category><![CDATA[Return to work assessment]]></category>
		<category><![CDATA[Structured observation tools]]></category>
		<category><![CDATA[vocational rehabilitation]]></category>
		<category><![CDATA[work ability]]></category>
		<category><![CDATA[Work ability assessment]]></category>
		<category><![CDATA[work performance]]></category>
		<category><![CDATA[Work performance measurement]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227979</guid>

					<description><![CDATA[A new Rasch analysis of the Finnish version of the Assessment of Work Performance finds the observational tool largely valid and reliable, while revealing targeting and multidimensionality issues that warrant further testing.]]></description>
										<content:encoded><![CDATA[<p>When someone on long-term sick leave is trying to return to work, the stakes of a single assessment can be enormous. Benefits, rehabilitation plans, and even the course of a person&#8217;s career can hinge on how accurately a clinician judges their ability to perform a job. Yet measuring work ability is notoriously difficult: it is a complex, multifaceted phenomenon, and the tools used to capture it are not always as reliable as the decisions built upon them. A new study published in the Scandinavian Journal of Occupational Therapy now offers the first rigorous statistical examination of the Finnish version of the Assessment of Work Performance, a structured observation tool designed to show how well a person actually works, not just how well they describe working.</p>
<p>The Assessment of Work Performance, or AWP, was originally developed in Sweden and has now been translated into ten languages, with Finnish being the most recent addition. Unlike interview-based instruments, the AWP relies on direct observation. An occupational therapist watches a client carry out a real or simulated work task and rates fourteen distinct observable skills on a four-point scale, from incompetent to competent performance. These skills fall into three domains: motor skills such as mobility, coordination, and strength; process skills such as time management, planning, and adaptation; and communication and interaction skills such as social contact and information exchange. Crucially, the instrument is not tied to any particular diagnosis. Because it measures the interaction between person and environment, it can be applied to clients with mental health conditions, musculoskeletal problems, or any other work-related challenge, in genuine workplaces or constructed assessment settings.</p>
<p>In the new study, a team led by Jennie Nyman of Metropolia University of Applied Sciences, together with colleagues at Linköping University in Sweden, put the Finnish translation, the AWP-FI, through a demanding statistical technique known as Rasch analysis. Seventeen occupational therapists across Finland, all trained in a one-day session and two follow-up meetings, collected ninety-four assessments as part of their routine clinical work. The clients ranged from 18 to 60 years of age, and a majority of those with reported diagnoses had mental or behavioural disorders. Most assessments were conducted through direct observation, and the overwhelming majority, 84 percent, took place in simulated settings, with tasks such as book binding, household chores, and arranging and indexing magazines standing in for real jobs.</p>
<p>Rasch analysis is a form of modern psychometrics that converts ordinal rating-scale data into interval-level measurements. The method estimates two quantities simultaneously: the ability of each person and the difficulty of each item, placing both on a common scale. If the data behave as the model expects, the instrument can be said to measure a single coherent underlying trait, in this case work performance, in a way that is fair and comparable across individuals. The researchers structured their evaluation around three questions drawn from established psychometric guidance: whether the scale&#8217;s difficulty range matched the abilities of the sample, whether a reliable measurement scale had been constructed, and whether the individuals in the sample were being measured reliably and meaningfully.</p>
<p>The headline result was encouraging. The AWP-FI showed an overall fit to the Rasch model, with a non-significant chi-square test indicating that observed responses aligned well with model expectations. Individual items performed within acceptable statistical limits, and, importantly, no item behaved differently for different groups. Tests for differential item functioning, which detect whether an item is biased, for example by being harder for women than for men at the same underlying ability level, came back clean across gender, across direct versus participatory observation, and across simulated versus genuine work tasks. This stability suggests the Finnish instrument can be used confidently with men and women alike, in real workplaces or assessment laboratories, and whether the therapist watches from the side or joins the work activity as a participant.</p>
<p>The analysis was not without complications. Five of the fourteen items showed disordered thresholds, meaning that adjacent response categories were not being used in a consistent, ordered way. In practice, the middle rating of the four-point scale was rarely the most probable response for those items, likely because the observed clients clustered at similar performance levels and the full range of the scale was seldom used. The researchers collapsed the affected categories and reran the analysis, a standard corrective procedure, but they caution that any permanent change to the rating scale should wait for confirmation in larger and more varied samples.</p>
<p>Two deeper issues also emerged. First, targeting was suboptimal: the mean person location of 1.6 sat well above the item mean of zero, and the item difficulty range, spanning roughly from minus 1.3 to 0.9 on the logit scale, was narrow. In plain terms, the clients assessed were generally more capable than the hardest items could measure, leaving the scale without items at the demanding end of the work-performance continuum. The authors attribute this largely to the tasks chosen: most assessments used simulated, non-work-specific activities with relatively low demands. Their clinical recommendation is that assessors think carefully about task selection, since the nature and demands of the chosen task directly shape what the assessment can reveal. Adding harder items could widen the measure, but the authors urge caution, because new items must remain faithful to the instrument&#8217;s theoretical foundation in the Model of Human Occupation, and expanding item counts carries its own psychometric risks.</p>
<p>Second, the analysis flagged local dependency among items and indications of multidimensionality. Twenty-four of ninety-one residual correlations exceeded the statistical cut-off, and these correlations clustered neatly within the three AWP domains, suggesting that items within a domain share information beyond what the single underlying trait of work performance can explain. When the researchers compared person estimates derived from different domains, between 11 and 17 percent of respondents fell outside the range expected for a truly unidimensional scale, well above the 5 percent benchmark. Notably, this finding diverges from an earlier evaluation of the Swedish original, which found no multidimensionality. The authors suggest the clinical impact may be minimal, since the domain structure reflects the deliberate theoretical architecture of the Model of Human Occupation, but they recommend that future studies with larger samples apply multidimensional or testlet-based Rasch approaches to confirm the pattern.</p>
<p>On the reliability front, the news was solid. The person-separation index reached 0.83, comfortably above the 0.8 benchmark, indicating the scale can distinguish roughly three clinically meaningful groups of clients. Only four participants, about 4 percent, showed erratic response patterns outside acceptable fit limits, meaning nearly everyone was measured in a statistically coherent way. Taken together, the results provide an initial validation of the AWP-FI and support its use as a valid and reliable observational tool for assessing work performance in Finland, where the Finnish Institute for Health and Welfare has explicitly called for more reliable and valid vocational assessment practices. Used alongside the interview-based Worker Role Interview, whose Finnish version passed a similar evaluation in 2023, the AWP-FI can contribute to the kind of holistic, multi-source assessment that sickness insurance and vocational rehabilitation systems increasingly demand.</p>
<p>The study&#8217;s limitations are candidly acknowledged. The sample of ninety-four assessments is modest, the DIF analysis suffered from very unequal group sizes, with simulated tasks dominating at 84 percent, which raises the risk that genuine item bias could have gone undetected, and the narrow spread of client performance likely constrained several of the analyses. The authors therefore frame this as a first step rather than a final verdict, recommending further testing with larger and more diverse samples to resolve the local dependency and targeting issues. For clinicians, the practical message is immediate and actionable: the AWP-FI works, but its usefulness depends on choosing work tasks demanding enough to reveal the full range of a client&#8217;s capability, and on remembering that the fourteen observed skills are organised into three theoretically distinct domains that should inform, not obscure, the interpretation of results.</p>
<p><strong>Subject of Research:</strong> Psychometric validation of the Finnish version of the Assessment of Work Performance using Rasch analysis</p>
<p><strong>Article Title:</strong> Psychometric evaluation of the Finnish version of the Assessment of Work Performance (AWP-FI)</p>
<p><strong>Article References:</strong> Nyman, J., Sandqvist, J., Pihlava, J., Ekbladh, E., &amp; Yngve, M. (2025). Psychometric evaluation of the Finnish version of the Assessment of Work Performance (AWP-FI). <em>Scandinavian Journal of Occupational Therapy, 32</em>(1), Article 2555182. <a href="https://doi.org/10.1080/11038128.2025.2555182" rel="noopener noreferrer">https://doi.org/10.1080/11038128.2025.2555182</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1080/11038128.2025.2555182" rel="noopener noreferrer">10.1080/11038128.2025.2555182</a></p>
<p><strong>Keywords:</strong> psychometrics, Rasch analysis, occupational therapy, work ability, vocational rehabilitation, AWP-FI, construct validity, differential item functioning, Model of Human Occupation, Finland, assessment tools, work performance</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227979</post-id>	</item>
	</channel>
</rss>
