Cognitive scientists have long faced an uncomfortable trade-off. To test formal models of how people judge, categorize, and remember, researchers need stimuli whose properties can be precisely measured and fed into computational models. The easiest way to guarantee that control is to invent artificial materials: geometric shapes that vary in color and size, fictitious insects with binary features, or cartoon characters whose attributes are assigned by the experimenter. These materials yield elegant, replicable experiments, but they raise a nagging question about external validity. When a model of categorization succeeds with imaginary bugs, does the same success await when people classify real foods, animals, or nations? A new open-access resource now aims to close that gap by giving the research community three large, curated sets of real-world stimuli, complete with the psychological “feature spaces” needed to drive computational models of cognition.
The resource, published in the journal Behavior Research Methods, was created by David Izydorczyk and Arndt Bröder of the University of Mannheim. It consists of 80 representative items from each of three everyday domains — foods, mammals, and countries — along with more than 280,000 pairwise similarity judgments collected from 1,798 participants recruited through the online platform Prolific. From those judgments, the authors derived multidimensional scaling (MDS) representations: 14 psychological dimensions for foods and ten each for mammals and countries. Crucially, these low-dimensional spaces reconstructed the original similarity data with high accuracy, achieving correlations of at least .93 in every domain. In a proof-of-concept study, the same dimensions also explained substantial variance in participants’ numerical judgments about the items, demonstrating that they are not merely statistical artifacts but functionally relevant summaries of how people actually think about these objects.
The methodological logic behind the project rests on a long-standing insight from cognitive psychology: computational models of judgment and memory need inputs. Exemplar models of categorization, for instance, compute the similarity between a new object and stored examples; lens models of judgment assume that people infer a hidden criterion, such as an object’s value, from probabilistic cues available in the environment. In every case, the model must be told what features each stimulus possesses. With artificial stimuli, those features are simply defined by the experimenter — this shape is red, that bug has wings. With real-world stimuli, no such predefined cue structure exists, so researchers have historically hand-picked cues based on intuition. In the classic city-size task of Gigerenzer and Goldstein, for example, the presence of an airport or a major-league soccer team served as binary cues for inferring population. But nothing guarantees that experimenter-chosen cues match the representations people actually use.
Multidimensional scaling offers an elegant, data-driven alternative. The technique, pioneered analytically in the early 1960s by Roger Shepard and Joseph Kruskal, takes a matrix of pairwise similarity ratings and represents each stimulus as a point in a multidimensional space, such that items judged similar end up near each other and dissimilar items end up far apart. The resulting dimensions are then interpreted as the psychological properties along which the stimuli vary — sweetness for foods, size for mammals, wealth for countries. Because the features are extracted from human judgments rather than imposed by the researcher, they are candidates for what the literature calls “functionally relevant” features: the properties that actually guide cognitive processing in a given task. Nosofsky and colleagues previously applied this strategy to sets of rocks and minerals, showing that MDS-derived dimensions could fuel computational models of categorization and recognition memory, and more recently Izydorczyk and Bröder used the approach to model how people learn to estimate the maximum flight speeds of bird species.
The new work extends this data-driven philosophy to three domains chosen with deliberate care. The authors applied three selection criteria. First, the domains needed high everyday relevance, so that even laypeople bring meaningful knowledge to the laboratory and stay motivated across long rating sessions. Second, each domain had to support many possible judgment and decision criteria: foods can be evaluated for caloric content, sugar levels, or carbon footprint; countries can be assessed on population, gross domestic product, or life expectancy. Third, the team wanted diversity in the kind of domain itself — one natural taxonomic category (mammals), one purely cultural-political domain (countries), and one mixed domain (foods), the latter selected partly for its relevance to public health and nutritional literacy. Where massive databases like THINGS, which catalogs more than 1,850 object concepts, aim for breadth across the object universe, this project aims for depth: 80 well-chosen items per domain, dense enough to support subtle within-domain comparisons that sparser databases cannot capture.
The stimulus curation reflects the same balance of control and realism. Food photographs were taken under standardized conditions with an iPhone 13 Pro, every item plated on an identical white dish with fork and knife for scale, spanning staples like rice, apples, and chicken as well as processed items such as cheesecake and cherry jam. Country stimuli consisted of names, sampled randomly from a Wikipedia list with balanced representation across six continental regions. Mammal images — from guinea pigs to African elephants and bottlenosed dolphins — were cropped from their original backgrounds and placed on a uniform white canvas in Adobe Photoshop to eliminate visual confounds. All materials, images, and data are freely available on the Open Science Framework under a CC-BY-SA-NC-4.0 license for non-commercial research use.
Collecting the similarity data required a substantial logistics effort. With 80 items per domain, there are 3,160 possible pairs, and asking any single participant to rate them all would be prohibitively tedious. The team therefore used a balanced incomplete block design in which each participant rated only 158 pairs — about five percent of the total — while the pooled sample covered every pair multiple times: each pair received ratings from roughly 30 participants in the food domain, 26 in the country domain, and 33 in the mammal domain. Participants rated similarity on a seven-point scale, with instructions kept deliberately vague to avoid steering them toward purely perceptual or purely semantic features, allowing the most psychologically salient dimensions to emerge on their own. Attention checks, self-report items on cheating and data quality, and correlations between individual ratings and group averages were used to prune the dataset; after exclusions, the final samples comprised 607 participants for foods, 520 for countries, and 671 for mammals.
The reliability of the aggregated data proved excellent. Following established practice, the authors estimated reliability as the mean split-half correlation corrected with the Spearman–Brown formula across 1,000 random splits, yielding values around .94 for the food domain, with similarly high estimates in the other two. Descriptively, average normalized similarity was highest for countries (mean .33 on a 0–1 scale), reflecting perhaps the shared conceptual framing of nations, compared with foods (.17) and mammals (.21), where ratings spanned nearly the full range from identical-looking twins to maximally dissimilar pairs. Participants’ open-ended descriptions of the features they used — sweet versus savory, wild versus domesticated, rich versus poor — provided a sanity check that the emerged dimensions align with conscious intuitions.
The proof-of-concept study addressed the question that matters most for modelers: do these similarity-derived dimensions predict behavior on tasks other than similarity rating? After a brief training phase, additional participants judged a target criterion in each domain — numerical estimates of domain-specific quantities. When the MDS dimensions were used as predictors in regression-style analyses, they explained substantial variance in those judgments. This functional relevance is what distinguishes the resource from alternatives. In recent years, many researchers have turned to deep neural networks, extracting feature vectors from pre-trained image or language models as cheap, scalable stimulus representations. Those embeddings have real advantages — they cost nothing to generate for vast stimulus sets — but they carry the fingerprints of their training task, and their transferability degrades as tasks diverge. Human similarity ratings, by contrast, directly reflect the mental representations of the very species under study.
For the field, the significance of the resource lies less in novelty of method than in curation and access. The authors are explicit that they introduce no new scaling or feature-extraction technique; the contribution is a well-validated, ready-to-use database that lowers the barrier for running model-based studies with realistic materials. A researcher testing an exemplar model of categorization, a multiple-cue probability model of learning, or a global-memory model of recognition can now download the stimuli, images, similarity matrices, and dimension coordinates and begin experimenting immediately — across tasks as varied as estimating calories, classifying mammals, or judging national wealth. The three domains differ, too, in how much they rely on perceptual similarity versus conceptual reasoning, allowing investigators to probe when each mode dominates. In an era when psychology faces increasing pressure to demonstrate that laboratory findings generalize to the world outside, tools like this one make ecological validity a practical option rather than an aspiration.
Cite Scienmag News
Glenn Wilkins. (September 7, 2026). New dataset aids modeling of judgment, categorization, memory across three domains. Scienmag. https://scienmag.com/new-dataset-aids-modeling-of-judgment-categorization-memory-across-three-domains/
Glenn Wilkins. "New dataset aids modeling of judgment, categorization, memory across three domains." Scienmag, 7 September 2026, https://scienmag.com/new-dataset-aids-modeling-of-judgment-categorization-memory-across-three-domains/. Accessed 7 September 2026.
Glenn Wilkins. "New dataset aids modeling of judgment, categorization, memory across three domains." Scienmag. September 7, 2026. https://scienmag.com/new-dataset-aids-modeling-of-judgment-categorization-memory-across-three-domains/

