For more than four decades, psychological research on how people read emotions from faces has depended on photographs of human actors. From the classic Pictures of Facial Affect compiled by Paul Ekman and Wallace Friesen in 1976 to modern datasets such as the Karolinska Directed Emotional Faces and the Amsterdam Dynamic Facial Expression Set, scientists have relied on posed expressions captured in controlled studio conditions. Now a team at the University of Haifa has demonstrated a rigorous new alternative: facial expression stimuli created entirely with generative artificial intelligence, validated so thoroughly that more than 2,000 study participants could not reliably tell them apart from real photographs.
The research, published in Behavior Research Methods, describes both a method for producing photorealistic AI-generated emotional expressions and a comprehensive validation framework for testing whether such images can stand in for photographs in behavioral experiments. The work, led by Shlomo Hareli and Shlomo David, addresses a long-standing tension in emotion research between experimental control and ecological validity. Traditional photographic datasets are constrained by the actors available, the expressions they can convincingly produce, and the ethical complications of using images of real people. Computational alternatives such as 3D morphable models and FACS-based animation tools offer more control but require expensive specialized software, months of technical training, and often produce faces that fall squarely into the uncanny valley.
Generative AI, the researchers argue, relaxes this trade-off. The method is built on the open-source platform ComfyUI and the Flux.1-dev diffusion model, enhanced with Low-Rank Adaptation, or LoRA, fine-tuning. The authors trained custom LoRA models for each combination of emotion and gender using 15 images per combination drawn, with permission, from the validated FACES database of real facial expressions. Each model was trained for 16 epochs on a single consumer-grade laptop GPU, requiring roughly six hours per emotion-gender pair, with a fixed random seed to guarantee that other laboratories can reproduce the results exactly. Prompt engineering then controlled body weight, clothing, lighting, background, and gaze direction, with weighted attention values emphasizing specific features during generation.
Crucially, the pipeline does not stop at generation. Every image was screened by Py-Feat, a Python library for facial expression analysis, using a strict criterion that the intended emotion had to reach at least a 95 percent classification likelihood before the image was accepted. Images that failed were iteratively refined with a node-based expression editor until they passed. The final Study 1 dataset contained 72 images, evenly balanced across gender, three body weight categories, and three emotional expressions, with a mean intended-emotion confidence of 98.38 percent. All workflows, models, datasets, and analysis scripts were released openly on the Open Science Framework, a move the authors say is central to making the method genuinely reproducible.
The first human validation study recruited 606 US participants through Prolific, each viewing a single image. The results confirmed the manipulations across the board: happy faces were rated significantly happier than neutral ones, which were rated happier than sad ones, and perceived weight rose in clear steps from the thin through the average to the higher-weight categories. Perhaps most strikingly, when asked whether the person in the photo was real or AI-generated, on a scale where 3 meant they could not decide, no condition was clearly identified as artificial. Even the least convincing cell, higher-weight male faces, scored only around 3.25, barely past the uncertainty midpoint.
The age judgments revealed an unexpected form of validation. Happy faces were perceived as younger than sad or neutral ones, and perceived age increased systematically with body weight, patterns that replicate well-documented findings from research on real human photographs. Rather than revealing artificial confounds in the stimuli, these effects suggest the AI generation process captured realistic correlations between facial morphology, body weight, and perceived age, enhancing rather than undermining the ecological validity of the images.
Study 2 extended the framework in two important ways. First, it added anger as a fourth emotion. Second, it adopted a same-character design, in which a single AI-generated identity appears across all four emotional expressions, eliminating potential confounds from differing facial features between conditions. This posed a technical challenge: maintaining identity while altering expression. The team solved it with a cyclic workflow combining an expression editor, LoRA-based inpainting with the identity-preserving PuLID technique, and a face-swapping node, validated computationally with FaceShape software that confirmed 100 percent within-character similarity, benchmarked against the ADFES database. A total of 96 images depicting 24 distinct characters passed both identity and emotion checks, and 736 new participants confirmed that all four emotions and all weight categories were perceived as intended.
A third study, Study 2b, pushed validation beyond simple emotion recognition. With 731 additional participants rating the same 96 images, the researchers examined valence, intensity, naturalness, authenticity, and non-focal emotions. The stimuli behaved like real-expression photographs on every dimension: happiness was rated positive, sadness and anger negative, and neutral faces affectively neutral yet not expressively empty, with measurable intensity. The ordering of naturalness and authenticity ratings, happiness and neutrality highest, anger lowest, mirrored the genuineness hierarchy previously established for posed expressions in real databases. Male faces were rated as angrier than female ones, consistent with known gender stereotypes in emotion perception, and body weight influenced sadness and fear ratings, showing that the stimuli carry meaningful social-category information rather than functioning as abstract emotion displays.
The authors are careful about the limits of their achievement. Photorealism as judged by human observers does not mean AI-generated and real faces are equivalent in every respect; computational algorithms can still detect subtle artifacts invisible to people, and the current framework addresses only static images, whereas dynamic expressions are known to aid emotion recognition. The pipeline also still demands considerable technical skill, and the authors call for future template interfaces to make it accessible without workflow-level expertise. Nevertheless, they conclude that AI-generated stimuli can now serve as a scientifically validated, ethically attractive alternative to photographs, particularly for research on stigmatized characteristics such as body weight, and that the dual computational and human validation framework they established provides a template any laboratory can adapt as generative technology continues to evolve.
Subject of Research: Creation and validation of photorealistic AI-generated facial expression stimuli for emotion perception research
Article Title: Creating and validating photorealistic AI-generated facial expression stimuli for emotion research
Article References: Hareli, S., & David, S. (2026). Creating and validating photorealistic AI-generated facial expression stimuli for emotion research. Behavior Research Methods, 58(11), Article 296. https://doi.org/10.3758/s13428-026-03156-0
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03156-0
Keywords: artificial intelligence, facial expressions, emotion perception, stimulus validation, body weight, generative AI, LoRA, ComfyUI, Py-Feat, behavioral research, photorealism, psychology methods
Cite Scienmag News
Glenn Wilkins. (September 22, 2026). AI-Generated Faces Pass the Realness Test in New Emotion Research Toolset. Scienmag. https://scienmag.com/ai-generated-faces-pass-the-realness-test-in-new-emotion-research-toolset/
Glenn Wilkins. "AI-Generated Faces Pass the Realness Test in New Emotion Research Toolset." Scienmag, 22 September 2026, https://scienmag.com/ai-generated-faces-pass-the-realness-test-in-new-emotion-research-toolset/. Accessed 22 September 2026.
Glenn Wilkins. "AI-Generated Faces Pass the Realness Test in New Emotion Research Toolset." Scienmag. September 22, 2026. https://scienmag.com/ai-generated-faces-pass-the-realness-test-in-new-emotion-research-toolset/








