Every political argument runs on something deeper than policy: a claim about what is ultimately good. Security versus freedom, tradition versus change, the collective versus the individual — these abstract values form the invisible architecture of democratic debate, yet they have been notoriously hard to measure in the wild. Now an international consortium coordinated by the European Commission’s Joint Research Centre has unveiled the largest expert-annotated dataset of human values in political text ever assembled. Described in the journal Behavior Research Methods, the ValuesML collection comprises 2,648 texts and 74,231 sentences drawn from news articles and election manifestos in nine languages: Bulgarian, Dutch, English, French, German, Greek, Hebrew, Italian and Turkish. Crucially, each value-laden passage was identified and interpreted by trained values experts, and each carries a second, unusual layer of information — whether the value is being celebrated as attained or invoked as under threat. The researchers say the resource could reshape how scientists study political polarization and how engineers build artificial intelligence that grasps what humans actually care about.
The dataset is anchored in the refined theory of human values developed by psychologist Shalom Schwartz and colleagues. In that framework, values are broad, trans-situational motivational goals — desirable end states such as power, achievement, hedonism, stimulation, self-direction, universalism, benevolence, tradition, conformity and security — that steer judgment and behavior across wildly different contexts. These ten basic values are not an unstructured grab-bag. They are arranged in a quasi-circumplex, a circular map in which neighboring values are motivationally compatible and opposing values collide, summarized by two axes: self-transcendence versus self-enhancement, and openness to change versus conservation. A 2012 refinement carved the circle into 19 finer values, splitting security into personal and societal varieties, and these categories have proven stable across hundreds of samples from more than 80 countries. The team chose Schwartz’s model over the rival Moral Foundations Theory because it extends beyond moral intuitions such as care and loyalty to goals like power and conformity, and because its circular geometry yields precise, testable predictions about which values should co-occur in text.
Measuring such abstractions in flowing prose has long defeated simpler tools. The dominant approach — dictionary methods that count words from predefined value lists — struggles in political language, where meaning is notoriously context-dependent. The word “party” can signal a celebration or a political organization, and “freedom of” versus “freedom to” invokes opposing motivational worlds depending on what follows. Transformer-based neural models, the architectures behind modern language technology, read context far better and have been shown to outperform dictionaries at detecting moral content in political communication. Yet the authors argue that such classifiers are typically trained on task-specific labels with opaque provenance, and that corpora generated or labeled by large language models still depend on gold-standard human annotation for calibration and evaluation. Earlier value datasets also leaned on crowdsourced workers judging short argument snippets, largely in English, and therefore failed to cover the genres where citizens actually encounter politics — news reporting and election manifestos at scale.
Building the new corpus was an exercise in disciplined scale. The project ran from May 2023 to April 2024 and drew on two veins of political text. The first was news: for each of seven EU languages, the team sampled 32,500 articles per language from the Europe Media Monitor, an automated system that scans thousands of outlets in more than 50 languages, retained pieces published between January 2019 and March 2023 on topics such as immigration, health, defense, the economy, environment and education, and filtered down to roughly 600 articles per language, mixing mainstream sources with outlets flagged by fact-checkers for spreading disinformation. The second vein was programmatic: excerpts from manifestos of 70 parties that contested recent parliamentary elections in twelve countries, including Austria, Germany, France, Greece, Israel, Turkey, the United Kingdom and the United States, covering themes from renewable energy and immigration to the minimum wage and national defense. After cleaning and screening for “value-laden” content — a signal-driven strategy, since prior work suggests only about a third of natural texts contain identifiable value expression — the final corpus held 2,354 news articles and 294 manifesto excerpts.
Annotation was treated as a science in itself. The campaign recruited 66 expert annotators, 21 curators and nine language leads, drawn from social psychology, political science, sociology, economics and data science, via an open call and snowball sampling through research networks. All annotators trained on a shared guide to Schwartz’s 19 refined values, then progressed through a flashcard program of escalating difficulty, from unambiguous single-value sentences to context-dependent level-four cases, before a joint calibration exercise in which every annotator coded the same political speech: Ursula von der Leyen’s 2022 State of the European Union address. The actual annotation ran on the INCEpTION collaborative platform, where experts highlighted value-expressive spans up to sentence length and attached value and attainment labels; multiple values could be marked in overlapping spans, but annotators were trained to name a dominant value when interpretations competed. In total, 43,248 text spans were coded, 82 percent of texts received two independent expert annotations, and language leads met weekly with the core team to resolve thorny cases and fold rulings back into iteratively updated guidelines.
The dataset’s most distinctive feature is its second annotation layer: attainment. A value can be framed as attained, as when a policy is praised for “ensuring freedom and democracy,” or as constrained, invoked as under threat, as when the same policy is attacked for “threatening our safety.” Pilot testing showed this attained-versus-constrained distinction was more informative than simple positive or negative framing, and every annotation carries it. The resulting distributions are telling. Across nearly all values, attained framings dominated, consistent with the theory’s picture of values as desirable end states. The conspicuous exception is security, which was more often annotated as constrained — a quantitative fingerprint of threat-based political rhetoric, in which speakers mobilize audiences not by celebrating safety but by insisting that it is slipping away. The team also tested whether annotators’ own personal value priorities colored their judgments, finding only weak and unsystematic associations with the labels they produced.
Because value identification is inherently interpretive, the team measured agreement with unusual care, computing Cohen’s kappa for annotator pairs and Krippendorff’s alpha across each language group on spans with at least 50 percent word overlap. Across all annotations, alpha reached 0.546 for the ten basic values and 0.486 for the 19 refined values — moderate agreement, below the commonly cited 0.667 benchmark, although the Greek and French groups exceeded it at 0.696 and 0.685. The authors argue these scores reflect the inherent interpretive complexity of the task rather than careless coding, in line with recent perspectivist approaches to subjective annotation. A structured curation phase then reconciled disagreements: annotations on which two experts independently agreed were accepted automatically, while contested cases went to expert curators who adjudicated against the guidelines and documented precedents, discarding unsupported labels. A final cross-lingual audit clustered semantically similar annotated spans via English translations and sent divergent labels to language leads for joint review, tightening comparability across languages without erasing legitimate cultural nuance.
The curated corpus, released openly as the Touché24-ValueEval benchmark for an international shared task on value detection, already sketches a coherent portrait of political rhetoric. Societal security emerged as the most frequently annotated value, followed by achievement and rule-oriented conformity, while hedonism, humility and tradition barely registered — evidence that political discourse is organized around collective and institutional concerns rather than personal gratification. Manifestos proved richer in value expression than news articles, with self-transcendence values such as universalism concentrated in party programs, suggesting politicians strategically foreground collective ideals in manifestos while journalism distributes value language more thinly. Aggregated by ideology, the profiles track theory: liberal parties peak on self-direction and achievement; nationalist and radical-right parties on power, face and rule-oriented conformity; social democratic and left-wing parties on universalism and personal security; green parties on universalism-nature; and conservative parties on tradition.
External checks suggest the annotations capture something real. When the team computed conditional co-occurrence ratios — how much more often two values share a sentence than chance predicts — adjacent values on the Schwartz circle kept company as theory demands, while opposites shunned each other: power co-occurred with hedonism at a ratio of just 0.21, and universalism rarely traveled with achievement. One anomaly, an elevated hedonism-tradition pairing of 2.41, dissolved on inspection into references to traditional festivities and national celebrations, a reminder that concrete instantiations of values can bend the idealized motivational circle. Most strikingly, a winning algorithm from the shared task, trained on this dataset, scanned 2020 manifestos from seven countries and reproduced recognizable ideological value profiles for entire party families, from agrarian to radical right — tentative evidence that aggregated annotations carry genuine political signal.
The resource, openly available on Zenodo with aligned English translations for every sentence, arrives at a charged moment for both democracy and machine learning. It offers a benchmark for building transparent, theory-grounded models that detect value framing at scale — a capability relevant to monitoring polarization, designing public communication and aligning large language models with the values societies actually hold. The authors are candid about limits: the corpus deliberately oversamples value-laden text rather than mirroring the full news landscape, expert curation may smooth away legitimate interpretive variation, and domains such as sport, art and personal relationships are underrepresented. They advise treating the annotations as probabilistic, context-dependent indicators and triangulating them with survey measures such as the European Social Survey. Even so, the consortium argues, the dataset lays a foundation for measuring, modeling and debating the motivational core of politics — a shared map, in nine languages, of what democracies hold dear and what they fear losing.
Cite Scienmag News
Glenn Wilkins. (August 30, 2026). New multilingual dataset detects values in news and political manifestos. Scienmag. https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/
Glenn Wilkins. "New multilingual dataset detects values in news and political manifestos." Scienmag, 30 August 2026, https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/. Accessed 30 August 2026.
Glenn Wilkins. "New multilingual dataset detects values in news and political manifestos." Scienmag. August 30, 2026. https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/

