Wednesday, September 23, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Teaching AI to Think Before Flagging Hateful and Propagandistic Memes

September 23, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 4 mins read
0
Teaching AI to Think Before Flagging Hateful and Propagandistic Memes

Teaching AI to Think Before Flagging Hateful and Propagandistic Memes

Teaching AI to Think Before Flagging Hateful and Propagandistic Memes

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Memes have become one of the most pervasive modes of communication on social media, blending images and text with humor, irony, and cultural references. While often harmless, this format can be exploited to spread hate speech, disinformation, and propaganda, and the very humor that makes memes shareable can trivialize toxic content and normalize hostile views. Detecting harmful memes automatically is notoriously difficult because their meaning rarely lives in the image or the caption alone; it emerges from the interaction between the two, often through implicit stereotypes or satirical framing that a naive classifier will miss entirely.

A new study published in Machine Learning with Applications addresses this challenge with a reasoning-centric training methodology for multimodal large language models, or MLLMs. Led by Mohamed Bayan Kmainasi of Qatar Computing Research Institute and colleagues including Mucahid Kutlu, Ali Ezzat Shahroor, Abul Hasnat, and Firoj Alam, the work demonstrates that reinforcement learning with chain-of-thought supervision can push explainable meme moderation to state-of-the-art performance on both English hateful memes and Arabic propagandistic memes. The research represents, according to the authors, the first systematic study of group relative policy optimization, known as GRPO, in multimodal reasoning under cross-lingual, fine-grained, and self-training settings.

The team’s central insight is that thinking-based MLLMs, which generate explicit intermediate reasoning steps before committing to an answer, are ideally suited to memes because their meaning depends on image-text interaction rather than unimodal cues. But whether the reasoning capabilities of such models, typically honed on mathematics and code generation, transfer to subjective, culturally situated tasks like meme moderation remained an open question. Standard supervised fine-tuning alone provides limited control over the balance between prediction correctness and rationale faithfulness, since cross-entropy loss treats all output tokens equally regardless of their functional role.

To resolve this, the researchers designed a three-stage training pipeline. The first stage is a supervised fine-tuning warm-up that aligns the model with gold labels, natural language explanations, and distilled reasoning traces produced by GPT-4.1. The second stage applies GRPO with a composite reward function that jointly optimizes classification correctness, output-format compliance, reference-based explanation similarity measured by METEOR, explanation length, and a novel thinking-length reward called Rthink. The third stage, self-training GRPO or ST-GRPO, extends the approach to unlabeled data using consensus-based pseudo-labels derived from the model’s own majority-vote predictions.

The composite reward is carefully weighted so that the two binary objectives, label correctness and format compliance, each receive 0.35 and together dominate the three auxiliary components, which contribute 0.30 combined. The thinking-length reward is particularly significant. Without it, the authors observed a consistent reward-hacking pattern: the model learned to compress or empty its reasoning traces while still collecting high reward, a length-shortening bias that was especially pronounced on the harder Arabic dataset. Rthink penalizes only reasoning traces shorter than a minimum threshold of 150 words, discouraging degenerate outputs without incentivizing verbosity.

The evaluation spans two distinct tasks and two languages. The English benchmark is the Facebook Hateful Memes dataset, containing roughly 11,000 memes where classification requires joint multimodal understanding. The Arabic benchmark, ArMeme, contains about 5,700 memes with four labels covering propaganda, not-propaganda, not-meme, and other. Because no fine-grained propaganda annotations existed for ArMeme, the team built a dual-annotator pipeline using GPT-4.1 and Llama-4-Scout as independent labelers of 23 propaganda techniques, consolidated by Gemini-3-Pro as an arbiter. Human validation on 584 memes showed the consolidated annotations aligned better with human reference labels than either single-model source.

The results are striking. On the Hateful Memes benchmark, the best supervised GRPO configuration with thinking-length regularization achieved 82.0 percent accuracy and 0.80 macro-F1, outperforming prior reported results and strong sequence-classification baselines such as Qwen3-VL-8B-Instruct and Gemma-3-12B-IT. On ArMeme, self-training GRPO reached 0.612 macro-F1, improving over previous work by 7.6 points and over the original ArMeme benchmark by 6.1 points. Notably, unimodal baselines lagged far behind: the best text-only model on ArMeme reached only 0.509 macro-F1, and image-only models averaged just 0.267, confirming that cross-modal reasoning captures signals that neither modality provides alone.

The self-training stage showed a telling asymmetry. On ArMeme, where unlabeled data was collected from the same social media sources as the labeled set, ST-GRPO improved macro-F1 by 1.5 points with gains concentrated in minority classes. On the English dataset, it slightly degraded performance, a result the authors attribute to distribution mismatch between the unlabeled pool and the benchmark, and to majority-vote bias amplification in binary classification, where small prediction biases can produce near-unanimous consensus that reinforces rather than corrects class skew. The finding suggests that consensus-based pseudo-labeling requires both distribution alignment and sufficient label-space diversity to provide reliable supervision.

Beyond raw accuracy, the model produces natural language explanations alongside its predictions, a property that sequence classifiers lack. Modality ablations showed the model genuinely depends on both inputs, with image removal hurting the English task most and OCR text removal crippling Arabic propaganda detection. An LLM-as-judge evaluation using GPT-4.1 and Gemini-2.5-Pro found the trained models substantially outperformed the zero-shot backbone on grounding, correctness, and usefulness for moderation, approaching the quality of human-verified reference explanations. The team has released all code, data extensions, prompting templates, and evaluation resources on GitHub.

The implications extend well beyond memes. The study shows that fine-grained supervision and distilled chain-of-thought rationales complement reinforcement-learning-based optimization, that multi-LLM annotation pipelines can scale fine-grained labeling to previously unlabeled domains, and that even a simple thresholded reasoning-length penalty can stabilize RL training against reward hacking in multimodal settings. For content moderation at scale, where subjective interpretation of culturally embedded imagery is the daily reality, the work offers a concrete template for building systems that not only flag harmful content but also articulate why, enabling meaningful human review rather than opaque automated verdicts.

Subject of Research: Reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes

Article Title: Adapting reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes

Article References: Kmainasi, M. B., Kutlu, M., Shahroor, A. E., Hasnat, A., & Alam, F. (2026). Adapting reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes. Machine Learning with Applications, 26, Article 101003. https://doi.org/10.1016/j.mlwa.2026.101003

Image Credits: AI Generated

DOI: 10.1016/j.mlwa.2026.101003

Keywords: reinforcement learning, chain-of-thought, multimodal large language models, hateful memes, propaganda detection, content moderation, GRPO, self-training, explainable AI, Arabic memes, reward hacking, fine-grained annotation

Cite Scienmag News

Denise Maddox. (September 23, 2026). Teaching AI to Think Before Flagging Hateful and Propagandistic Memes. Scienmag. https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/

Denise Maddox. "Teaching AI to Think Before Flagging Hateful and Propagandistic Memes." Scienmag, 23 September 2026, https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/. Accessed 23 September 2026.

Denise Maddox. "Teaching AI to Think Before Flagging Hateful and Propagandistic Memes." Scienmag. September 23, 2026. https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/

Tags: AI-driven social media content filteringArabic memeschain-of-thoughtchain-of-thought supervision in AI modelschallenges in automatic harmful content recognitioncontent moderationcross-lingual meme analysisexplainable AIexplainable AI for social media moderationfine-grained annotationgroup relative policy optimization in AIGRPOHate speech detection in memeshateful memesmultilingual meme moderation techniquesmultimodal large language modelsmultimodal large language models for content moderationpropaganda and disinformation detection in memespropaganda detectionreasoning-based AI training for harmful content identificationreinforcement learningreinforcement learning in hate speech detectionreward hackingself-training
Share26Tweet16
Previous Post

AI Turns School Websites Into a Measuring Stick for Digital Transformation

Next Post

Scientists Stack Drought and Flood Tolerance Genes Into Beloved Rice Variety

Related Posts

Machine Learning Reveals a Surprising Concentration Paradox in UK Student Mobility
Technology and Engineering

Machine Learning Reveals a Surprising Concentration Paradox in UK Student Mobility

September 23, 2026
Lab-Grown Heart Models Move Closer to Predicting Human Cardiac Disease
Technology and Engineering

Lab-Grown Heart Models Move Closer to Predicting Human Cardiac Disease

September 23, 2026
Open-Source Solar-Powered Aeroponic Tower Grows Food Off the Grid for $720
Technology and Engineering

Open-Source Solar-Powered Aeroponic Tower Grows Food Off the Grid for $720

September 23, 2026
AI Framework Cortex Maps Career Goals to Prerequisite-Ready Learning Paths
Technology and Engineering

AI Framework Cortex Maps Career Goals to Prerequisite-Ready Learning Paths

September 23, 2026
Plant-Based Polymer Electrolyte Boosts Magnesium Supercapacitor Performance
Technology and Engineering

Plant-Based Polymer Electrolyte Boosts Magnesium Supercapacitor Performance

September 23, 2026
Attention Weights Turned Into Powerful New Explanations for AI Transformers
Technology and Engineering

Attention Weights Turned Into Powerful New Explanations for AI Transformers

September 23, 2026
Next Post
Scientists Stack Drought and Flood Tolerance Genes Into Beloved Rice Variety

Scientists Stack Drought and Flood Tolerance Genes Into Beloved Rice Variety

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Chemical Tag on Histones Governs Rice Immunity Against Viruses and Insect Pests
  • How a chance discovery of 2,500 files revealed the birth of plastic surgery
  • Healthcare Workers in Japan Felt Less Happy Even After COVID Emergency Ended
  • Feeling Connected Helps Adults with Type 1 Diabetes Manage Their Condition Better

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading