<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>large language model biases &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/large-language-model-biases/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 13 Apr 2026 17:06:41 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>large language model biases &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Switch-Driven Prompts Expose ChatGPT’s Gender Bias</title>
		<link>https://scienmag.com/switch-driven-prompts-expose-chatgpts-gender-bias/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 13 Apr 2026 17:06:41 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI and feminist ideology]]></category>
		<category><![CDATA[AI and LGBTQ+ advocacy]]></category>
		<category><![CDATA[AI and socioeconomic controversies]]></category>
		<category><![CDATA[AI ethical challenges in gender]]></category>
		<category><![CDATA[AI ideological frameworks]]></category>
		<category><![CDATA[AI response to gender issues]]></category>
		<category><![CDATA[AI stereotype reinforcement]]></category>
		<category><![CDATA[algorithmic bias in AI]]></category>
		<category><![CDATA[ChatGPT gender bias]]></category>
		<category><![CDATA[generative AI and gender equality]]></category>
		<category><![CDATA[large language model biases]]></category>
		<category><![CDATA[OpenAI gender policy]]></category>
		<guid isPermaLink="false">https://scienmag.com/switch-driven-prompts-expose-chatgpts-gender-bias/</guid>

					<description><![CDATA[In the rapidly evolving landscape of generative artificial intelligence, recent research sheds light on the intricate dynamics of ChatGPT’s engagement with gender equality discourses. A study conducted by Song, Liang, and Zhao (2026) unpacks how this flagship large language model (LLM) simultaneously demonstrates a perceptive commitment to gender equality while paradoxically reproducing entrenched gender biases, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the rapidly evolving landscape of generative artificial intelligence, recent research sheds light on the intricate dynamics of ChatGPT’s engagement with gender equality discourses. A study conducted by Song, Liang, and Zhao (2026) unpacks how this flagship large language model (LLM) simultaneously demonstrates a perceptive commitment to gender equality while paradoxically reproducing entrenched gender biases, highlighting a critical gap in AI design and implementation. Their findings reveal that ChatGPT’s seemingly progressive, sometimes educative responses to straightforward queries on gender issues are not merely coincidental but reflect a deliberate top-level orientation embedded by OpenAI. Nonetheless, the model’s persistence in reinforcing stereotypes during more nuanced, contextual conversations underscores unresolved challenges at the intersection of algorithmic architecture, data biases, and underlying ideological frameworks.</p>
<p>At the core of OpenAI’s mission is the equitable benefit of artificial general intelligence (AGI) for all humanity, enshrined through policies aimed at avoiding discriminatory or harmful outputs. This foundational commitment aligns with an explicit prioritization of gender equality in ChatGPT’s content generation. However, the study highlights a notable ideological tilt in the model’s responses, particularly its liberal feminist stance on contentious socioeconomic issues such as abortion rights, casual sex, and LGBTQ+ advocacy. These perspectives, grounded in principles of individual autonomy and pro-choice values, align strongly with progressive political ideologies, a bias that reflects broader patterns identified in previous research documenting ChatGPT’s political partialities. Specifically, the model’s output favors liberal figures and policy stances in the US, UK, and Brazil, mirroring the predominantly Global North-centric and US-originated training corpus.</p>
<p>A substantial portion of ChatGPT’s unconscious reproduction of gender bias stems directly from the vast training datasets underpinning its language capabilities. The Common Crawl dataset, a sprawling repository of multilingual internet text exceeding hundreds of billions of data points, encapsulates the ingrained societal prejudices and structural inequalities prevalent online. As language models learn statistical patterns from such data, they inadvertently internalize and replicate these historical biases. Compounding this issue is the “black box” complexity of multi-parameter neural architectures, which obscures the causal pathways leading to specific model outputs. Without transparent interpretability, developers face significant hurdles in detecting and mitigating nuanced biases, which can manifest as stereotypical associations of gender roles in both social and fictionalized narratives within ChatGPT’s responses.</p>
<p>This inability to grasp the subtleties of cultural and narrative contexts—what may be conceptualized as the “material fluidity of life”—accounts for much of the model’s problematic outputs. Unlike humans, whose social cognition is deeply rooted in lived experience and cultural immersion, ChatGPT operates purely on probabilistic language prediction devoid of material cultural understanding. As a consequence, its rendering of gender-related topics is stilted and often constrained by the binary and reductive frameworks embedded in its training data. The model’s repetitive linking of certain traits or behaviors to specific genders echoes human cognitive phenomena such as implicit bias, yet without the compensatory mechanisms that social learning affords humans to challenge and revise stereotypes.</p>
<p>The technical underpinnings of ChatGPT’s dialogue capabilities further explain its variable performance. Operating under an “instruction” or command-input framework rather than engaging in traditional Socratic questioning, the model generates responses through rapid probabilistic computations guided by input prompts. This “command input–rapid generation” paradigm defines the interactional style of generative AI and informs its characteristic unpredictability and occasional contradictions. In effect, this limits ChatGPT’s capacity for nuanced dialectics, contributing to inconsistencies observed in gender equality responses across different conversational depths and contexts.</p>
<p>Beyond technical intricacies, this research underscores the expanding ethical responsibilities of AI engineers amid the generative AI revolution. Professionals in this space are no longer mere algorithmic coders but must increasingly assume accountability for the social welfare implications of their creations. The complexity of bias, especially concerning gender equality, defies simplistic statistical fairness metrics that dominate current evaluation methods. Existing approaches, often centered on group fairness or parity, struggle to capture the context-dependent nature of bias and neglect dimensions such as individual fairness or culturally contingent norms. The study calls for a paradigm shift in bias testing frameworks to address the layered interplay between societal prejudices and algorithmic operation, fostering more holistic and equitable AI systems.</p>
<p>Integral to this ethical discourse is the recognition of parasocial relationships emerging from intimate human-AI interactions. Originally conceptualized to describe one-way emotional engagement between media consumers and public figures, parasocial theory now extends into AI user experiences. In contrast to traditional media’s unidirectionality, LLM-powered chatbots like ChatGPT facilitate dynamic, reciprocal dialogues that simulate social presence, nurturing pseudo-interactive spaces where users invest emotional and cognitive trust. This evolution raises the stakes of subtle bias transmission, as ingrained gender prejudices within AI outputs can subtly influence users’ perceptions in ostensibly personalized and meaningful ways, potentially reinforcing discriminatory norms undetected.</p>
<p>To confront these challenges, the study advocates for the infusion of specialized feminist expertise into the AI development pipeline. Generic training regimens and algorithmic debiasing efforts remain insufficient without deliberate incorporation of gender-aware perspectives and targeted datasets. Initiatives like AI Now’s emphasis on rectifying gender imbalances within AI research and fostering data diversity represent pivotal steps forward. Moreover, technical redesign must account for the intrinsic “language” mechanics of LLMs, acknowledging the stochastic processes driving model behaviors in varying dialogue scenarios—from factual prompts to imaginative storytelling. Balancing linguistic and cultural plurality in source data, especially in fictional narratives, stands as essential to overcoming the current blind spots in gender bias recognition.</p>
<p>Despite inherent limitations—including the influence of response randomness and the difficulty of exhaustively covering all dimensions of gender bias in experimental prompts—the research offers critical insights into the dialectical nature of ChatGPT’s gender-equality articulations. While version 4.0 of the model demonstrates promising capacity to affirm gender equity under clear-cut queries, it simultaneously reveals significant vulnerabilities under complex or contextually rich conditions. This ambivalence highlights the dual potential of generative AI: as both a powerful agent in promoting gender education on a global scale and as a conduit for perpetuating subtle prejudices.</p>
<p>Looking forward, the study calls for empirical investigations into the real-world impacts of engaging with AI systems like ChatGPT on users’ gender attitudes and beliefs. This direction is crucial for unpacking the dynamic feedback loops through which human cultural narratives are both reflected and reshaped by AI-mediated communication. Experimental panel designs and longitudinal user studies could elucidate how sustained interaction with gender-biased or gender-affirming AI outputs conditions social cognition, ultimately informing the responsible development of AI technologies that uphold and advance human values.</p>
<p>In sum, as generative AI takes center stage in mediating human knowledge and social discourse, rigorously addressing gender bias remains a thorny but essential endeavor. The intricate interplay between data-driven biases, architectural opacity, and human-machine parasocial dynamics demands concerted multidisciplinary efforts. Embedding feminist insights, advancing fairness metrics beyond superficial parity, and deepening interpretability of LLM behavior stand as critical paths toward realizing AI systems that genuinely promote inclusivity and equity. In their current iteration, generative models like ChatGPT straddle the line between reflecting societal progress and reproducing historical inequities, a tension that the scientific community must navigate to harness their transformative promise responsibly.</p>
<p>Only through transparent, ethically grounded, and culturally nuanced AI engineering can we move closer to technologies that not only mimic human conversational prowess but also embody the diversity and fairness that society aspires to achieve. The dialogues of tomorrow’s AI must transcend binary stereotypes, echoing the complexity of lived human experience rather than fossilizing outdated paradigms. This research provides a clarion call to the AI research world—innovation must be paired with vigilance and social accountability to cultivate trustworthy artificial interlocutors that elevate rather than undermine the cause of gender equality.</p>
<hr />
<p><strong>Subject of Research</strong>: Gender Bias and Equality in Large Language Models, specifically ChatGPT’s outputs in gender-related conversational contexts</p>
<p><strong>Article Title</strong>: Modes of asking as switches: prompt-driven inconsistency in ChatGPT’s gender equality perspective outputs</p>
<p><strong>Article References</strong>:<br />
Song, S., Liang, Z. &amp; Zhao, W. Modes of asking as switches: prompt-driven inconsistency in ChatGPT’s gender equality perspective outputs. <em>Humanit Soc Sci Commun</em> 13, 478 (2026). <a href="https://doi.org/10.1057/s41599-025-05577-2">https://doi.org/10.1057/s41599-025-05577-2</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: <a href="https://doi.org/10.1057/s41599-025-05577-2">https://doi.org/10.1057/s41599-025-05577-2</a></p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">150909</post-id>	</item>
		<item>
		<title>Unveiling the Hidden Biases, Emotions, Personalities, and Abstract Concepts Within Large Language Models</title>
		<link>https://scienmag.com/unveiling-the-hidden-biases-emotions-personalities-and-abstract-concepts-within-large-language-models/</link>
		
		<dc:creator><![CDATA[Reid Dalton]]></dc:creator>
		<pubDate>Thu, 19 Feb 2026 23:05:30 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[abstract concepts in AI]]></category>
		<category><![CDATA[AI emotional states modulation]]></category>
		<category><![CDATA[AI expert persona modeling]]></category>
		<category><![CDATA[AI transparency methods]]></category>
		<category><![CDATA[detecting hidden AI biases]]></category>
		<category><![CDATA[emotional representation in AI]]></category>
		<category><![CDATA[large language model biases]]></category>
		<category><![CDATA[manipulating AI model outputs]]></category>
		<category><![CDATA[MIT AI research breakthroughs]]></category>
		<category><![CDATA[personality traits in language models]]></category>
		<category><![CDATA[steering AI behavior]]></category>
		<category><![CDATA[UC San Diego AI collaboration]]></category>
		<guid isPermaLink="false">https://scienmag.com/unveiling-the-hidden-biases-emotions-personalities-and-abstract-concepts-within-large-language-models/</guid>

					<description><![CDATA[In the realm of artificial intelligence, large language models (LLMs) like OpenAI’s ChatGPT, Anthropic’s Claude, and Google’s Gemini have revolutionized how machines understand and generate human language. These models have transcended mere answer generation to embody complex, abstract concepts, representing nuanced tones, biases, personalities, and emotional states. Yet, despite their growing ubiquity and sophistication, the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the realm of artificial intelligence, large language models (LLMs) like OpenAI’s ChatGPT, Anthropic’s Claude, and Google’s Gemini have revolutionized how machines understand and generate human language. These models have transcended mere answer generation to embody complex, abstract concepts, representing nuanced tones, biases, personalities, and emotional states. Yet, despite their growing ubiquity and sophistication, the precise mechanisms through which these models encode and process such intangible attributes have remained largely enigmatic. Now, an innovative collaboration between researchers at MIT and the University of California San Diego has yielded a breakthrough methodology to both detect and manipulate hidden concepts embedded within LLMs, heralding a new era of transparency and control in AI behavior.</p>
<p>This pioneering technique goes beyond conventional prompting methods by incisively isolating the internal mathematical structures of LLMs tasked with encoding specific abstract notions. By harnessing these structures, the team can effectively “steer” the model’s outputs toward amplifying or attenuating targeted conceptual themes. Their experimental exploits encompassed more than 500 overarching concepts spread across personality traits, emotional dispositions, fears, locational preferences, and expert personas. For example, the researchers successfully identified and modulated LLM representations linked to personalities as disparate as “social influencer” and “conspiracy theorist” or stances ranging from “fear of marriage” to “enthusiasm for Boston.”</p>
<p>One particularly striking demonstration of the technique’s versatility involved augmenting the “conspiracy theorist” persona within a state-of-the-art vision-language model. When queried about the origins of the iconic “Blue Marble” photograph of Earth, the model, under the influence of the enhanced conspiracy theorist concept, produced an answer steeped in conspiracy-laden conjectures. Such vivid manipulations underscore both the power and potential pitfalls of this approach, emphasizing the critical necessity for responsible application.</p>
<p>Traditional methods to uncover latent abstractions in LLMs often rely on unsupervised learning algorithms that sift indiscriminately through vast arrays of unlabeled numerical representations, hoping to discern emergent patterns corresponding to concepts like “hallucination” or “deception.” These methods, while valuable, suffer from two main drawbacks: computational inefficiency and lack of specificity. Adityanarayanan “Adit” Radhakrishnan, assistant professor of mathematics at MIT and lead co-author on the study, analogizes conventional unsupervised tactics as casting wide, cumbersome nets in a vast ocean, hoping to catch a singular species, often overwhelmed by irrelevant captures.</p>
<p>To circumvent these issues, the research team employed a more surgical approach informed by recursive feature machines (RFMs), a predictive modeling framework designed to extract salient features from data by tapping into the implicit mathematical feature-learning mechanisms underlying neural networks. This approach, which Radhakrishnan and colleagues had previously developed, enables highly targeted identification of concept-specific numerical patterns within the dense vector spaces of LLMs, thereby sidestepping the noise and resource drain endemic to broader unsupervised methods.</p>
<p>Applying RFMs to LLMs, the researchers trained the algorithm on labeled sets of prompts — for instance, comparing 100 conspiracy-related queries against 100 neutral ones — to discern numerical fingerprints uniquely associated with the “conspiracy theorist” concept. Once trained to recognize these representations, the method can mathematically perturb the LLM’s internal activations complementing or suppressing the abstract concept’s influence. This granular modifiability allows precise steering of a model’s behavior, opening doors to tailor AI responses with unprecedented finesse.</p>
<p>Importantly, the team did not limit their exploration to a narrow class of concepts. They mapped representations for a diverse spectrum including psychological fears (such as fear of marriage or insects), expert identities (e.g., medievalist or social influencer), affective states (boastful or amused), geographic predilections (Boston or Kuala Lumpur), and historical or cultural personas (Ada Lovelace, Neil deGrasse Tyson). Through systematic application across several of today’s leading large language and multimodal vision-language models, the researchers established that these abstract concepts are intricately woven into the fabric of AI’s learned representations.</p>
<p>The technical heart of this breakthrough rests on an understanding of how LLMs process inputs. At their core, LLMs are sophisticated neural networks that ingest prompts by decomposing strings of natural language into tokens, each token encoded as a high-dimensional vector of numbers. These vectors are propagated through multiple computational layers, each performing linear algebraic transformations and nonlinear activations. Matrix representations evolve across layers as the model probabilistically infers summary representations poised to generate coherent, contextually appropriate outputs, ultimately decoded back into human-readable text. The RFM methodology effectively operates within this multi-layer numerical landscape to isolate and influence specific conceptual “coordinates.”</p>
<p>Beyond academic curiosity, the practical implications of this method are profound. The research showcased scenarios where typical model safeguards—such as refusal to engage with inappropriate queries—could be selectively deactivated by dialing up an “anti-refusal” representation, thereby highlighting potential vulnerabilities and risks. Conversely, positive modulation allows for the enhancement of beneficial attributes like brevity or rigorous reasoning in model outputs, promising pathways to customization that improve utility without sacrificing safety.</p>
<p>Radhakrishnan emphasizes that the revelation of these abstract conceptual embeddings within LLMs challenges conventional beliefs about the black-box nature of these models. With sufficient insight into how such representations manifest and interact, it is conceivable to engineer specialized LLMs finely tuned for particular tasks while simultaneously maintaining robust operational safety. The research team has prudently open-sourced the underlying code for their method, fostering transparency and encouraging wider community adoption for monitoring and refining AI models.</p>
<p>This breakthrough comes at a critical juncture as LLMs permeate countless applications, raising ethical and technical questions about underlying biases, hallucinations, and AI-generated misinformation. By advancing tools to untangle and modulate hidden conceptual layers, the study equips developers, policymakers, and researchers with a new lens to interrogate, understand, and ultimately govern AI behavior more effectively.</p>
<p>Furthermore, beyond immediate steering capabilities, this approach offers a scalable blueprint for universal monitoring and intervention protocols applicable to the burgeoning complexity of AI architectures. Such tools could form the backbone of next-generation AI safety frameworks, balancing flexibility with rigorous control.</p>
<p>As the authors note, while the potential benefits are substantial, caution remains imperative. Some extracted concepts, if manipulated irresponsibly, could exacerbate misinformation, prejudice, or unethical AI behaviors. Therefore, continued research and thoughtful governance are essential companions to technological advances.</p>
<p>In sum, this study represents a pivotal step towards demystifying the internal conceptual schema of AI language systems, transforming them from opaque behemoths into more interpretable, controllable entities. By enabling targeted activation and suppression of abstract notions, the research paves the way for AI that is not only smarter but safer, more ethical, and more aligned with human values.</p>
<p>Subject of Research: Understanding and steering abstract concept representations in large language models (LLMs).</p>
<p>Article Title: Toward universal steering and monitoring of AI models</p>
<p>News Publication Date: 19-Feb-2026</p>
<p>Web References: http://dx.doi.org/10.1126/science.aea6792</p>
<p>Keywords: Artificial intelligence, Large language models, Neural networks, Concept representations, Recursive feature machines, AI safety, Machine learning, Adaptive systems, Feature learning, Bias detection, AI steering, Computational linguistics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">138229</post-id>	</item>
	</channel>
</rss>
