Personality psychology has long rested on the idea that human character can be captured by five broad domains: openness to experience, conscientiousness, extraversion, agreeableness, and neuroticism. But researchers have increasingly recognized that these domains are too coarse for many scientific purposes. Within each of the Big Five sit two finer-grained aspects, yielding ten building blocks of personality that often predict behavior more precisely than the broad domains themselves. Measuring all ten usually required long questionnaires, which limited their use in clinics, large surveys, and studies where participants’ time is scarce. A new study published in Current Psychology reports that a compact 40-item questionnaire designed to capture all ten aspects now works just as well in Italian as it does in English, offering one of the most rigorous cross-cultural tests of a short personality measure to date.
The instrument in question is the Big Five Aspect Scales short version, known as the BFAS-40. The original Big Five Aspect Scales, developed by Colin DeYoung and colleagues, grew out of the observation that each Big Five domain splits into two aspects: for example, extraversion divides into enthusiasm and assertiveness, while neuroticism splits into withdrawal and volatility. These aspects sit between the broad domains and the even narrower facets measured by instruments such as the NEO Personality Inventory, and they have proven useful for linking personality to psychopathology, political ideology, and well-being. The BFAS-40, developed by Gallagher, Stevenor, Samo, and McAbee, condensed the original longer scales into just four items per aspect, aiming to preserve psychometric strength while cutting completion time dramatically.
Translating a questionnaire is not simply a matter of swapping words. Personality items often rely on idioms, subtle emotional vocabulary, and response conventions that do not map neatly across languages. Italian researchers have a long history of adapting personality measures, going back to factor-analytic work on the NEO inventories in the 1990s and 2000s, but each new instrument must be validated on its own terms. A team led by Chloe Lau of the University of Toronto and Francesco Bruno of Universitas Mercatorum in Rome, together with Francesca Chiesi of the University of Florence and Lena Quilty of the Centre for Addiction and Mental Health, took on that task for the BFAS-40, testing the translated scale in a community sample of 662 Italian-speaking adults.
The methodological heart of the study is its use of multidimensional item response theory, or MIRT, a modern statistical framework that models how each individual item functions across levels of the underlying trait. Unlike classical test theory, which focuses on total scores, item response theory estimates parameters for every question: how well it discriminates between people with different trait levels, and where its response thresholds fall along the trait continuum. Because the ten personality aspects are correlated and each item is written to tap one aspect within a broader domain, the researchers fit confirmatory multidimensional models rather than treating each aspect in isolation. This approach allowed them to test whether the hypothesized two-aspect structure actually held within each of the five Big Five domains in the Italian data.
The results were encouraging. Across all ten aspects, the confirmatory MIRT models showed acceptable to strong factor loadings and discrimination parameters, with response thresholds that were well ordered, meaning the answer options functioned as intended along the trait scale. Bayesian fit statistics told a consistent story: standardized root mean square residual values ranged from 0.045 to 0.063, root mean square error of approximation values from 0.032 to 0.045, and comparative and Tucker-Lewis fit indices ranged from 0.950 to 0.975. In conventional terms, these numbers indicate acceptable to strong model fit, supporting the structural validity of the two-factor solutions within each domain. In plain language, the Italian items grouped themselves into the same ten aspects that theory predicted, without needing to be reshuffled or discarded.
Reliability, the consistency with which the items measure each aspect, was assessed using Bayesian methods rather than traditional point estimates alone. The resulting estimates ranged from low to excellent across the ten aspects, a spread that is not unusual for very short scales. With only four items per aspect, some aspects inevitably show weaker internal consistency than others, and the authors are transparent about this variability. The Bayesian framework is particularly well suited to this kind of evaluation because it produces full distributions of the reliability estimates rather than single numbers, giving researchers a more honest picture of the precision of the measurement in a sample of this size.
Validity was then tested against external criteria. The Italian BFAS-40 was administered alongside the Big Five Inventory short form, allowing the researchers to check whether the new scale’s domain scores converged with an established measure of the same constructs. They also examined associations with indices of well-being and with measures of depression, anxiety, and stress drawn from the DASS-21. The pattern of correlations supported both convergent and discriminant validity: aspects related to negative emotionality, such as withdrawal and volatility, correlated with distress measures in the expected directions, while aspects tied to positive engagement aligned with well-being scores. This network of expected relationships suggests the translated items are measuring the same psychological substance as their English originals, not merely producing superficially similar responses.
Perhaps the most internationally significant part of the study is its cross-cultural comparison. The researchers compared item functioning between the Italian sample and a Canadian sample of 347 participants, using differential item functioning analysis, a technique that detects whether an item behaves differently for people from different groups after accounting for their trait levels. Of the 40 items, 39 showed no evidence of differential functioning between Italian and Canadian respondents. Only a single Agreeableness item was flagged. This near-perfect record is remarkable for a translated personality measure and suggests that the Italian adaptation achieved what psychometricians call measurement equivalence, the property that allows scores to be compared meaningfully across cultures rather than within each country alone.
The one flagged item is worth noting rather than dismissing. Agreeableness, which splits into compassion and politeness, is the aspect most closely tied to social values and interpersonal norms, domains where cultural expectations differ most between countries. An item that functions differently across cultures does not necessarily indicate a translation error; it may reflect genuine cultural variation in how agreeable behavior is expressed or interpreted. For researchers planning cross-cultural studies, the practical implication is that 39 of the 40 Italian items can be used in direct comparison with Canadian data, while the flagged Agreeableness item should be interpreted with caution in such analyses.
The study arrives at a moment when demand for brief, validated personality measures is accelerating. Short scales are increasingly used in electronic health records, large online panels, and studies that follow thousands of people over time, where every additional minute of questionnaire length translates into dropout. Recent work has extended the BFAS-40 to German, French, and Italian samples, and the present study adds a level of statistical scrutiny, multidimensional item response modeling with Bayesian estimation and formal cross-cultural DIF testing, that goes beyond typical translation checks. For Italian-speaking researchers and clinicians, the BFAS-40 now offers a theoretically grounded, efficient tool that captures the ten aspects of the Big Five in a fraction of the time of the original instruments. For the broader field, it demonstrates that careful psychometric engineering can make even a nuanced hierarchical model of personality travel across languages without losing its shape, a small but meaningful step toward personality science that speaks with one measurement language worldwide.
Subject of Research: Cross-cultural psychometric validation of the Italian BFAS-40 personality questionnaire using multidimensional item response theory
Article Title: The Italian adaptation of the big five aspects scale short version (BFAS-40): a multidimensional item response theory model and cross-cultural validation study
Article References: Lau, C., Bruno, F., Chiesi, F., & Quilty, L. C. (2026). The Italian adaptation of the big five aspects scale short version (BFAS-40): a multidimensional item response theory model and cross-cultural validation study. Current Psychology, 45(18), Article 1508. https://doi.org/10.1007/s12144-026-09982-x
Image Credits: AI Generated
DOI: 10.1007/s12144-026-09982-x
Keywords: personality psychology, Big Five, BFAS-40, psychometrics, item response theory, cross-cultural validation, Italian adaptation, reliability, validity, differential item functioning, Bayesian analysis, psychological assessment
Cite Scienmag News
Glenn Wilkins. (October 9, 2026). A 40-Question Personality Test Passes Its Italian Exam With Flying Colors. Scienmag. https://scienmag.com/a-40-question-personality-test-passes-its-italian-exam-with-flying-colors/
Glenn Wilkins. "A 40-Question Personality Test Passes Its Italian Exam With Flying Colors." Scienmag, 9 October 2026, https://scienmag.com/a-40-question-personality-test-passes-its-italian-exam-with-flying-colors/. Accessed 9 October 2026.
Glenn Wilkins. "A 40-Question Personality Test Passes Its Italian Exam With Flying Colors." Scienmag. October 9, 2026. https://scienmag.com/a-40-question-personality-test-passes-its-italian-exam-with-flying-colors/

