Large language models may be able to do more than generate personality questionnaires: they may already contain a working statistical model of how people tend to think and respond. In a study published August 6 in iScience, researchers from the Hebrew University of Jerusalem used GPT-4 to transform source texts into psychological surveys, then tested whether the resulting questionnaires could measure meaningful personality patterns. Their most striking finding was that ChatGPT predicted how people would answer the surveys before any participants had completed them.
The researchers began with a simple but unconventional question: could an artificial intelligence system read almost any text describing human characteristics and convert it into a structured personality assessment? To test the idea, they supplied GPT-4 with passages from the Diagnostic and Statistical Manual of Mental Disorders, or DSM-5, a clinical reference used to describe and diagnose mental disorders. They also provided an astrology textbook as a deliberately contrasting source, because astrology assigns personality traits to zodiac signs but does not rest on the same empirical foundations as modern psychological science.
GPT-4 processed the texts and generated questionnaires containing statements that participants could rate on a five-point scale, ranging from strong disagreement to strong agreement. From the DSM-5 material, for example, the model created items based on descriptions associated with paranoid personality disorder, including whether a person often suspects other people’s motives or finds it easy to trust them. The astrology-based questionnaire used similar statements, but linked them to personality descriptions attributed to astrological signs and elements. In both cases, the language model served as a translator between unstructured written material and measurable psychological variables.
This process resembles a form of automated instrument construction. A personality questionnaire is not simply a list of interesting statements; its items must be sufficiently clear, consistent, and related to the underlying trait being studied. Researchers therefore examined whether responses to the AI-generated questions formed coherent clusters. If several items are intended to reflect a shared psychological dimension, participants’ answers should show statistical relationships with one another. This internal consistency is a fundamental property of reliable measurement and helps distinguish a meaningful scale from a collection of disconnected claims.
The study included 600 participants, who completed both AI-generated questionnaires and the Big Five Inventory, widely regarded as one of the most extensively validated frameworks for assessing personality. The DSM-5-based survey produced coherent patterns among traits that commonly occur together in real-world psychological data. Characteristics such as avoidance and dependency, for example, showed relationships that broadly resembled those observed in established personality research. The response structure also mirrored patterns found in the Big Five Inventory, providing evidence that the model had extracted psychologically relevant information rather than merely producing plausible-sounding questions.
The astrology-based questionnaire produced a different result. Although its individual statements could still contain recognizable descriptions of human behavior, the broader groups of traits showed weak internal consistency. Characteristics associated with the same astrological category did not reliably move together in participants’ responses. In statistical terms, the survey lacked the coherent latent dimensions expected from a scientifically grounded personality model. The result suggests that an AI system can extract useful psychological language from a non-scientific source without validating the source’s larger explanatory system.
That distinction became particularly important when the researchers examined whether the questionnaires could predict life outcomes. Both the DSM-5- and astrology-derived surveys showed associations with measures related to depression, anxiety, and well-being at levels comparable to the Big Five questionnaire. The finding does not mean that astrology has been scientifically confirmed. Instead, it indicates that personality-relevant descriptions embedded in an astrology text may still capture general behavioral and emotional patterns that have psychological signal. In effect, the model may have separated useful descriptions of people from the unsupported framework used to organize them.
The study’s most unexpected test took place before participants answered the questions. Using the source material and the generated questionnaires, GPT-4 was asked to forecast the population-level results, including average responses to individual items and correlations between questions. The model’s predictions reportedly matched the participants’ response patterns with high accuracy. It did not need to know how any particular individual would answer; rather, it anticipated broad statistical regularities, such as which statements people would tend to endorse together and which would remain largely unrelated.
This ability may arise from the way large language models are trained. Systems such as GPT-4 learn from enormous collections of language gathered from websites, books, and social media. Because people routinely express emotions, motives, habits, fears, and social judgments through language, the training data may contain repeated traces of how personality traits are described and how they tend to co-occur. The model is not explicitly taught a psychological theory in the manner of a student or clinician, but its internal representations may encode associations that resemble population-level psychological knowledge. The authors describe this as a possible natural byproduct of learning human language.
The findings could make questionnaire development faster and more flexible, particularly for exploratory research, but they do not eliminate the need for human validation. A survey generated by an AI system can inherit cultural assumptions, reproduce biases in its training data, or confuse vivid language with a measurable trait. The researchers also caution that the results may be weaker in languages and cultures underrepresented in large language-model training data. Further studies will need to test whether the same patterns appear across countries, languages, age groups, and clinical populations. For now, the work suggests that generative AI is not merely producing psychological content on demand; it may also be capable of recognizing—and forecasting—the statistical structure of human personality.
Subject of Research: People
Article Title: Generating and analyzing personality questionnaires using large language models
News Publication Date: 6-Aug-2026
Web References: https://doi.org/10.1016/j.isci.2026.116909; http://www.cell.com/iscience
References: Monsa et al., iScience, DOI: 10.1016/j.isci.2026.116909
Image Credits: Hagar Segev
Keywords: Generative AI, large language models, GPT-4, machine learning, deep learning, personality tests, psychological assessment, clinical psychology, social psychology, Big Five Inventory, DSM-5, astrology, personality research

