Recognizing a familiar face is usually an effortless part of social life. A colleague may be identified across a busy office, a friend may be spotted in a crowd, and the face of someone recently met can often be recalled later without conscious effort. For people with developmental prosopagnosia, however, these ordinary experiences can remain persistently difficult. The condition, often known as “face blindness,” describes a lifelong difficulty recognizing familiar faces that occurs despite normal vision and the absence of a brain injury. Because many people adapt to the condition or remain unaware that it has a name, developmental prosopagnosia is frequently undiagnosed. A new study from Japan has now delivered the most comprehensive evaluation to date of one of the field’s most widely used screening instruments, the 20-item Prosopagnosia Index, or PI20.
Researchers led by Associate Professor Masaki Mori of the Japan Advanced Institute of Science and Technology, working with former graduate student Midori Sugiyama of Keio University, examined how effectively each PI20 question measures difficulties with face recognition. Their findings, published in Behavior Research Methods on August 12, 2026, support the questionnaire’s overall value as an accessible screening tool while revealing that not every item contributes equally to the measurement. The analysis could help researchers and clinicians use the PI20 more appropriately and guide the development of a revised questionnaire that identifies developmental prosopagnosia with greater precision. The work is especially significant because self-report tools are often the first practical step for people seeking an explanation for experiences that may have affected their social and professional lives for years.
The PI20 asks respondents to rate their agreement with 20 statements describing everyday face-recognition experiences. These include situations such as recognizing familiar people outside their usual surroundings, remembering a person’s face after an introduction, and identifying someone encountered previously. Respondents select one of five response categories, ranging from strongly disagree to strongly agree. Items written in opposite directions are scored in reverse so that the total score consistently reflects the same underlying trait: greater difficulty recognizing familiar faces. The questionnaire is not a diagnostic test by itself, but it can indicate whether a person’s experiences warrant further assessment using behavioral face-recognition tests and professional evaluation.
To investigate the instrument in detail, the researchers analyzed responses from 599 Japanese adults using two complementary statistical frameworks. Classical test theory evaluates the questionnaire as a whole, focusing on properties such as internal consistency, reliability, and relationships among scores. In practical terms, it asks whether the items work together coherently and whether the resulting score is sufficiently stable to support interpretation. Item response theory, or IRT, examines the behavior of individual questions across different levels of the measured trait. This approach can estimate how strongly an item discriminates between people with slightly different degrees of face-recognition difficulty and identify the range of ability or impairment in which a question provides the most information.
The researchers used a unidimensional graded response model, an IRT model designed for ordered response categories such as the five-point agreement scale used by the PI20. In the model, a latent variable represented by theta, or θ, describes a person’s position along the continuum of face-recognition difficulty. Each item is represented by characteristic curves showing the probability of selecting each response category as θ changes. Well-functioning items display an orderly progression: as face-recognition difficulty increases, the most likely response shifts step by step from stronger disagreement toward stronger agreement. When curves are well separated, the question can distinguish effectively between respondents. If the curves overlap heavily or fail to progress in an orderly way, the item contributes less precise information.
Taken together, the classical and IRT analyses found that the PI20 has excellent reliability and validity as a screening measure. Most questions showed the expected stepwise movement across response categories as respondents’ levels of face-recognition difficulty increased. These items were able to differentiate people more effectively, particularly when they described recognizable real-world situations involving familiar faces. The results suggest that the PI20 captures a meaningful and relatively coherent psychological construct rather than simply reflecting general social discomfort, poor memory, or dissatisfaction with vision. This does not mean that every person with a high score has developmental prosopagnosia, but it indicates that the questionnaire can provide a useful first signal for identifying individuals who may benefit from more detailed testing.
The analysis also exposed two items that performed comparatively poorly. One concerns recognizing people with distinctive facial features. Highly unusual faces may be memorable even for individuals who struggle with ordinary face recognition, meaning that the question may not measure the specific impairment associated with developmental prosopagnosia. A distinctive scar, unusually shaped feature, or other salient characteristic can provide a non-face cue that allows a person to identify someone without constructing a detailed representation of the face itself. In measurement terms, responses to this item may therefore be influenced by the memorability of unusual features rather than by the respondent’s general ability to recognize familiar faces.
Another weaker item asks about recognizing one’s own face in photographs. Although self-face recognition can be relevant to some aspects of visual identity processing, many people with developmental prosopagnosia do not report substantial difficulty identifying themselves. Recognition of one’s own face may also rely on contextual information, hairstyle, clothing, body shape, or familiarity built through repeated exposure. As a result, the item may have limited power to separate people with different degrees of developmental face-recognition difficulty. The researchers’ findings do not imply that these questions are meaningless in every context, but they suggest that they may be less suitable than other PI20 items for screening specifically for developmental prosopagnosia.
The study’s implications extend beyond the current questionnaire. The researchers have made one of the largest PI20 datasets publicly available along with the accompanying analysis code, creating an open resource for researchers investigating face recognition across languages and cultures. Such access can make it easier to test whether the questionnaire operates equivalently in different populations, a crucial issue in psychological measurement. Translation alone does not guarantee that an item has the same meaning or difficulty in another culture; everyday expectations about recognizing people, social encounters, and self-images may vary substantially. Open data and reproducible code can help investigators compare item performance, detect cultural differences, and evaluate revised versions rather than relying only on total scores.
The findings provide a stronger scientific foundation for a questionnaire already used widely in research and public-facing screening. They also illustrate why a high reliability coefficient is not enough to establish that every question is useful: a scale can be consistent overall while containing individual items that add little information or measure an unintended feature. By combining whole-test evaluation with item-level modeling, Mori and Sugiyama have shown where the PI20 performs well and where it can be refined. Their work may ultimately support earlier recognition of developmental prosopagnosia, more accurate research estimates, and better-targeted assessments for people whose face-recognition difficulties have long been misunderstood. The study was supported by the Waseda University Grant for Special Research Projects 2025, and the authors report no relevant financial or non-financial conflicts of interest.
Subject of Research: People
Article Title: Item analyses of the 20-item prosopagnosia index using classical test theory and item response theory
News Publication Date: 12-Aug-2026
Web References: https://doi.org/10.3758/s13428-026-03145-3
References: Mori, Masaki, and Midori Sugiyama. “Item analyses of the 20-item prosopagnosia index using classical test theory and item response theory.” Behavior Research Methods. DOI: 10.3758/s13428-026-03145-3
Image Credits: Dr. Masaki Mori, Japan Advanced Institute of Science and Technology (JAIST), Japan
Keywords: Developmental prosopagnosia, face blindness, face recognition, PI20, Prosopagnosia Index, item response theory, classical test theory, psychological assessment, screening questionnaire, cognitive science, vision science, psychometrics

