Sunday, August 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Psychology & Psychiatry

New multilingual dataset detects values in news and political manifestos

August 30, 2026
in Psychology & Psychiatry
Glenn Wilkins
By Glenn Wilkins Scienmag Editorial Profile - Clinical Psychology
Reading Time: 6 mins read
0
New multilingual dataset detects values in news and political manifestos

New multilingual dataset detects values in news and political manifestos

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every political argument runs on something deeper than policy: a claim about what is ultimately good. Security versus freedom, tradition versus change, the collective versus the individual — these abstract values form the invisible architecture of democratic debate, yet they have been notoriously hard to measure in the wild. Now an international consortium coordinated by the European Commission’s Joint Research Centre has unveiled the largest expert-annotated dataset of human values in political text ever assembled. Described in the journal Behavior Research Methods, the ValuesML collection comprises 2,648 texts and 74,231 sentences drawn from news articles and election manifestos in nine languages: Bulgarian, Dutch, English, French, German, Greek, Hebrew, Italian and Turkish. Crucially, each value-laden passage was identified and interpreted by trained values experts, and each carries a second, unusual layer of information — whether the value is being celebrated as attained or invoked as under threat. The researchers say the resource could reshape how scientists study political polarization and how engineers build artificial intelligence that grasps what humans actually care about.

The dataset is anchored in the refined theory of human values developed by psychologist Shalom Schwartz and colleagues. In that framework, values are broad, trans-situational motivational goals — desirable end states such as power, achievement, hedonism, stimulation, self-direction, universalism, benevolence, tradition, conformity and security — that steer judgment and behavior across wildly different contexts. These ten basic values are not an unstructured grab-bag. They are arranged in a quasi-circumplex, a circular map in which neighboring values are motivationally compatible and opposing values collide, summarized by two axes: self-transcendence versus self-enhancement, and openness to change versus conservation. A 2012 refinement carved the circle into 19 finer values, splitting security into personal and societal varieties, and these categories have proven stable across hundreds of samples from more than 80 countries. The team chose Schwartz’s model over the rival Moral Foundations Theory because it extends beyond moral intuitions such as care and loyalty to goals like power and conformity, and because its circular geometry yields precise, testable predictions about which values should co-occur in text.

Measuring such abstractions in flowing prose has long defeated simpler tools. The dominant approach — dictionary methods that count words from predefined value lists — struggles in political language, where meaning is notoriously context-dependent. The word “party” can signal a celebration or a political organization, and “freedom of” versus “freedom to” invokes opposing motivational worlds depending on what follows. Transformer-based neural models, the architectures behind modern language technology, read context far better and have been shown to outperform dictionaries at detecting moral content in political communication. Yet the authors argue that such classifiers are typically trained on task-specific labels with opaque provenance, and that corpora generated or labeled by large language models still depend on gold-standard human annotation for calibration and evaluation. Earlier value datasets also leaned on crowdsourced workers judging short argument snippets, largely in English, and therefore failed to cover the genres where citizens actually encounter politics — news reporting and election manifestos at scale.

Building the new corpus was an exercise in disciplined scale. The project ran from May 2023 to April 2024 and drew on two veins of political text. The first was news: for each of seven EU languages, the team sampled 32,500 articles per language from the Europe Media Monitor, an automated system that scans thousands of outlets in more than 50 languages, retained pieces published between January 2019 and March 2023 on topics such as immigration, health, defense, the economy, environment and education, and filtered down to roughly 600 articles per language, mixing mainstream sources with outlets flagged by fact-checkers for spreading disinformation. The second vein was programmatic: excerpts from manifestos of 70 parties that contested recent parliamentary elections in twelve countries, including Austria, Germany, France, Greece, Israel, Turkey, the United Kingdom and the United States, covering themes from renewable energy and immigration to the minimum wage and national defense. After cleaning and screening for “value-laden” content — a signal-driven strategy, since prior work suggests only about a third of natural texts contain identifiable value expression — the final corpus held 2,354 news articles and 294 manifesto excerpts.

Annotation was treated as a science in itself. The campaign recruited 66 expert annotators, 21 curators and nine language leads, drawn from social psychology, political science, sociology, economics and data science, via an open call and snowball sampling through research networks. All annotators trained on a shared guide to Schwartz’s 19 refined values, then progressed through a flashcard program of escalating difficulty, from unambiguous single-value sentences to context-dependent level-four cases, before a joint calibration exercise in which every annotator coded the same political speech: Ursula von der Leyen’s 2022 State of the European Union address. The actual annotation ran on the INCEpTION collaborative platform, where experts highlighted value-expressive spans up to sentence length and attached value and attainment labels; multiple values could be marked in overlapping spans, but annotators were trained to name a dominant value when interpretations competed. In total, 43,248 text spans were coded, 82 percent of texts received two independent expert annotations, and language leads met weekly with the core team to resolve thorny cases and fold rulings back into iteratively updated guidelines.

The dataset’s most distinctive feature is its second annotation layer: attainment. A value can be framed as attained, as when a policy is praised for “ensuring freedom and democracy,” or as constrained, invoked as under threat, as when the same policy is attacked for “threatening our safety.” Pilot testing showed this attained-versus-constrained distinction was more informative than simple positive or negative framing, and every annotation carries it. The resulting distributions are telling. Across nearly all values, attained framings dominated, consistent with the theory’s picture of values as desirable end states. The conspicuous exception is security, which was more often annotated as constrained — a quantitative fingerprint of threat-based political rhetoric, in which speakers mobilize audiences not by celebrating safety but by insisting that it is slipping away. The team also tested whether annotators’ own personal value priorities colored their judgments, finding only weak and unsystematic associations with the labels they produced.

Because value identification is inherently interpretive, the team measured agreement with unusual care, computing Cohen’s kappa for annotator pairs and Krippendorff’s alpha across each language group on spans with at least 50 percent word overlap. Across all annotations, alpha reached 0.546 for the ten basic values and 0.486 for the 19 refined values — moderate agreement, below the commonly cited 0.667 benchmark, although the Greek and French groups exceeded it at 0.696 and 0.685. The authors argue these scores reflect the inherent interpretive complexity of the task rather than careless coding, in line with recent perspectivist approaches to subjective annotation. A structured curation phase then reconciled disagreements: annotations on which two experts independently agreed were accepted automatically, while contested cases went to expert curators who adjudicated against the guidelines and documented precedents, discarding unsupported labels. A final cross-lingual audit clustered semantically similar annotated spans via English translations and sent divergent labels to language leads for joint review, tightening comparability across languages without erasing legitimate cultural nuance.

The curated corpus, released openly as the Touché24-ValueEval benchmark for an international shared task on value detection, already sketches a coherent portrait of political rhetoric. Societal security emerged as the most frequently annotated value, followed by achievement and rule-oriented conformity, while hedonism, humility and tradition barely registered — evidence that political discourse is organized around collective and institutional concerns rather than personal gratification. Manifestos proved richer in value expression than news articles, with self-transcendence values such as universalism concentrated in party programs, suggesting politicians strategically foreground collective ideals in manifestos while journalism distributes value language more thinly. Aggregated by ideology, the profiles track theory: liberal parties peak on self-direction and achievement; nationalist and radical-right parties on power, face and rule-oriented conformity; social democratic and left-wing parties on universalism and personal security; green parties on universalism-nature; and conservative parties on tradition.

External checks suggest the annotations capture something real. When the team computed conditional co-occurrence ratios — how much more often two values share a sentence than chance predicts — adjacent values on the Schwartz circle kept company as theory demands, while opposites shunned each other: power co-occurred with hedonism at a ratio of just 0.21, and universalism rarely traveled with achievement. One anomaly, an elevated hedonism-tradition pairing of 2.41, dissolved on inspection into references to traditional festivities and national celebrations, a reminder that concrete instantiations of values can bend the idealized motivational circle. Most strikingly, a winning algorithm from the shared task, trained on this dataset, scanned 2020 manifestos from seven countries and reproduced recognizable ideological value profiles for entire party families, from agrarian to radical right — tentative evidence that aggregated annotations carry genuine political signal.

The resource, openly available on Zenodo with aligned English translations for every sentence, arrives at a charged moment for both democracy and machine learning. It offers a benchmark for building transparent, theory-grounded models that detect value framing at scale — a capability relevant to monitoring polarization, designing public communication and aligning large language models with the values societies actually hold. The authors are candid about limits: the corpus deliberately oversamples value-laden text rather than mirroring the full news landscape, expert curation may smooth away legitimate interpretive variation, and domains such as sport, art and personal relationships are underrepresented. They advise treating the annotations as probabilistic, context-dependent indicators and triangulating them with survey measures such as the European Social Survey. Even so, the consortium argues, the dataset lays a foundation for measuring, modeling and debating the motivational core of politics — a shared map, in nine languages, of what democracies hold dear and what they fear losing.

Subject of Research: Expert annotation of human value expression in political news articles and party manifestos across nine languages, grounded in Schwartz’s refined theory of human values

Subject of Research: Psychology & Psychiatry

Article Title: ValuesML: A new multilingual dataset for values detection in news and political manifestos

Article References: Scharfbillig, M., Reitis-Münstermann, T., Stefanovitch, N., Kiesel, J., Brock, P. S., Sneddon, J., Cartier, E., Ardag, M., Arieli, S., Daniel, E., Dobewall, H., Karl, J., Krasteva, A., Oeschger, T. P., Petasis, G., Russo, L., Seddone, A., Tamo-Larrieux, A., van Herk, H., ... Ye, S. (2026). ValuesML: A new multilingual dataset for values detection in news and political manifestos. Behavior Research Methods, 58(10), Article 276. https://doi.org/10.3758/s13428-026-03092-z

Image Credits: AI Generated

DOI: 10.3758/s13428-026-03092-z

Keywords: human values, value expression, political discourse, cross-cultural research, multilingual dataset, annotation, natural language processing, political manifestos, news media, machine learning, inter-annotator agreement, computational social science

Cite Scienmag News

Glenn Wilkins. (August 30, 2026). New multilingual dataset detects values in news and political manifestos. Scienmag. https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/

Glenn Wilkins. "New multilingual dataset detects values in news and political manifestos." Scienmag, 30 August 2026, https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/. Accessed 30 August 2026.

Glenn Wilkins. "New multilingual dataset detects values in news and political manifestos." Scienmag. August 30, 2026. https://scienmag.com/new-multilingual-dataset-detects-values-in-news-and-political-manifestos/

Tags: AI understanding of human valuesAI understanding of political valuescomparative analysis of political texts across languagescross-lingual political text datasetsdataset for AI understanding of democratic debatesdetecting values as celebrated or threateneddetecting values in election manifestosexpert-annotated political datasetshuman values detection in news and manifestoshuman values in news and manifestosimpact of values on democratic debatesinternational political text datasetsmachine learning for political value identificationmeasuring abstract political valuesmultilingual dataset for political text analysismultilingual natural language processing for politicsmultilingual political text analysispolitical argument underlying valuespolitical polarization research toolspolitical values detectionSchwartz human values theory applicationSchwartz theory of human values in political discoursevalues annotation in political discoursevalues-based political polarization research
Share26Tweet16
Previous Post

Full-Stack Designs Bring Intelligence to Brain-Computer Interfaces

Next Post

Employees See Promise in Workplace Mindfulness, but Doubts Remain, Study Finds

Related Posts

Employees See Promise in Workplace Mindfulness, but Doubts Remain, Study Finds
Psychology & Psychiatry

Employees See Promise in Workplace Mindfulness, but Doubts Remain, Study Finds

August 30, 2026
New validated scale offers brief measure of adolescent sexual wellbeing
Psychology & Psychiatry

New validated scale offers brief measure of adolescent sexual wellbeing

August 30, 2026
What an ADHD diagnosis means for seemingly well-functioning women
Psychology & Psychiatry

What an ADHD diagnosis means for seemingly well-functioning women

August 29, 2026
Mouse model reveals new insights into brain dynamics in Angelman syndrome
Psychology & Psychiatry

Mouse model reveals new insights into brain dynamics in Angelman syndrome

August 29, 2026
Study reveals what drives moral distress in medical science students
Psychology & Psychiatry

Study reveals what drives moral distress in medical science students

August 29, 2026
New framework guides mental health integration across all policy sectors
Psychology & Psychiatry

New framework guides mental health integration across all policy sectors

August 29, 2026
Next Post
Employees See Promise in Workplace Mindfulness, but Doubts Remain, Study Finds

Employees See Promise in Workplace Mindfulness, but Doubts Remain, Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • How SHH-Wnt crosstalk, DNA methylation, and miRNAs drive uterine fibroids
  • Mineral phase of iron nanoparticles dictates toxicity to lettuce germination and growth
  • Immune-on-chip systems recreate human immunity for immunotherapy, vaccines, and autoimmune modeling
  • Smartphone sensors and Raspberry Pi devices combine to detect cosmic rays

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading