When AI Says No: How Language Models Became the World’s Newest Moral Authorities
A new study of 100,000 tweets shows that LLaMA 3.2 refuses about 10 percent of paraphrasing requests, enforcing a narrow, ...
A new study of 100,000 tweets shows that LLaMA 3.2 refuses about 10 percent of paraphrasing requests, enforcing a narrow, ...
A new fairness-regularized framework called CoR-Hate retrieves real counterfactual examples from corpus data to reduce identity bias in hate-speech detection ...
A new study systematically tests early, late, and hybrid fusion strategies for combining image and text in AI systems that ...
Researchers have developed a reinforcement learning method that trains multimodal AI models to detect hateful and propagandistic memes in English ...
Researchers have unveiled a Profit-Driven Simulation framework that tests AI toxicity detection models in realistic, revenue-sensitive social media environments and ...
© 2025 Scienmag - Science Magazine
© 2025 Scienmag - Science Magazine