<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>prompt design &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/prompt-design/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 07:34:53 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>prompt design &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Models Attach Gender to Workplace Roles, and Tone of Voice Matters More Than the Job</title>
		<link>https://scienmag.com/ai-models-attach-gender-to-workplace-roles-and-tone-of-voice-matters-more-than-the-job/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 07:34:53 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI fairness]]></category>
		<category><![CDATA[AI gender bias in workplace role descriptions]]></category>
		<category><![CDATA[algorithmic bias]]></category>
		<category><![CDATA[analysis of AI responses in professional contexts]]></category>
		<category><![CDATA[bias in AI language models across different configurations]]></category>
		<category><![CDATA[bootstrap inference]]></category>
		<category><![CDATA[communicative register]]></category>
		<category><![CDATA[effect of language phrasing on gender stereotypes in AI]]></category>
		<category><![CDATA[gender attribution]]></category>
		<category><![CDATA[impact of communicative style on AI-generated identities]]></category>
		<category><![CDATA[implications of AI]]></category>
		<category><![CDATA[influence of tone of voice on gender attribution]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models and demographic defaults]]></category>
		<category><![CDATA[large-scale testing of AI gender biases]]></category>
		<category><![CDATA[model heterogeneity]]></category>
		<category><![CDATA[occupational content and gender assignment in AI]]></category>
		<category><![CDATA[occupational stereotypes]]></category>
		<category><![CDATA[open-access audit on AI gender attribution]]></category>
		<category><![CDATA[prompt design]]></category>
		<category><![CDATA[response diversity]]></category>
		<category><![CDATA[role of tone and style in AI-generated demographic assumptions]]></category>
		<category><![CDATA[statistical audit]]></category>
		<category><![CDATA[workplace language]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252565</guid>

					<description><![CDATA[A live audit of nine large language models finds that directive workplace prompts shift gender attribution upward relative to deferential ones, but both remain majority female and model-to-model variation is extreme.]]></description>
										<content:encoded><![CDATA[<p>When a large language model is asked to imagine the writer of a workplace sentence, it almost always supplies a name, an age, and a gender, even though the prompt never mentions any of them. Those invented identities are not random. A new open-access audit published in the International Journal of Data Science and Analytics by Rajat Shukla of Nazareth University shows that the way a workplace sentence is phrased, not just the job it describes, dramatically changes which gender the model assigns to the imagined author. The finding offers one of the most statistically careful looks yet at how demographic defaults emerge from the intersection of occupational content and communicative style.</p>
<p>The audit was ambitious in scale. Shukla queried nine live model configurations, including GPT-5.5, Claude Opus 4.7, Claude Sonnet 4.6, three Gemini variants, two DeepSeek versions, and Gemma 4 31B, using twenty experimental phrases and five control phrases. Each model-phrase pair was tested with fifty stateless generations at temperature 1.0, producing 11,250 responses in total. Half of the experimental phrases bundled authority and resource-control content with direct, directive language, drawn from technology, business, medicine, and law. The other half bundled coordination and assistance content with reporting, deferential, or permission-seeking phrasing. Every response was parsed with deterministic regular expressions to extract the name, age, and gender the model had volunteered.</p>
<p>The headline numbers are striking. In the authority and directive condition, 63.58 percent of the imagined writers were labeled female, 35.89 percent male, 0.47 percent nonbinary, and 0.07 percent unparsed. In the support and deferential condition, the female share surged to 92.62 percent, while male attribution collapsed to 6.04 percent. In other words, both prompt bundles produced majority-female outputs, but the relative shift between them was enormous. The study is careful to frame this as a relative contrast within an overall female-skewed distribution rather than evidence of male dominance in authoritative roles, a nuance that distinguishes it from earlier, blunter headlines about occupational bias in language models.</p>
<p>Beneath the aggregate numbers lies an even more important story: the models disagreed wildly with one another. The difference in male-attribution risk between the two conditions ranged from minus 12.2 percentage points for Gemini 3.1 Pro, meaning it actually assigned more male writers to the supportive condition, to a staggering plus 93.6 points for Gemma 4 31B. Seven of the nine models showed positive differences and two showed negative ones. A Paule-Mandel random-effects synthesis across models produced an I-squared statistic of 98.6 percent on the log-odds scale, and a 95 percent prediction interval for the odds ratio stretched from 0.014 to 63,659. That interval spans the null effect and several orders of magnitude, which means no single pooled estimate can be trusted to describe any model outside the panel.</p>
<p>This heterogeneity is not a statistical nuisance; it is the dominant empirical result. A prediction interval that wide tells us that knowing the average behavior of these nine systems tells you almost nothing about the next model, or even the next version of the same model. The author argues, convincingly, that deployment audits should preserve model-specific estimates, exact provider routes, prompts, and collection dates rather than quoting a panel average as if it were a durable property of large language models in general. For organizations building on these systems, the practical implication is that demographic defaults must be measured on the exact endpoint they intend to use.</p>
<p>The study also confronts a subtle statistical trap: stateless API calls are not independent observations. Although each of the fifty generations per cell was a separate request with no conversational memory, the models often revisited the same narrow output modes. Within-cell name diversity, measured by Shannon entropy and effective name counts, ranged from roughly two to three effective names in several Gemma and Gemini 3 Flash strata to about twenty-seven to twenty-nine for DeepSeek v4 Pro. Some individual cells had a modal-name share as high as 98 percent, meaning nearly every generation in that cell produced the same name. Treating such responses as fifty independent draws would dramatically overstate the precision of any estimate.</p>
<p>To handle this dependence, the analysis employed four complementary techniques: a linear probability model with two-way cluster-robust covariance by phrase and model, a wild phrase-cluster bootstrap with Rademacher weights and 9,999 replications, a cross-classified block bootstrap that resampled models and phrases while keeping each fifty-response cell intact, and the random-effects synthesis. The cell-block bootstrap estimated an aggregate male-attribution gap of 29.9 percentage points with a 95 percent confidence interval of 6.0 to 55.6, while the two-way clustered model estimated 29.8 points. All approaches preserved a positive average contrast in this fixed panel, but with far wider uncertainty than naive response-level calculations would suggest.</p>
<p>The control conditions revealed another unsettling behavior. Five control phrases contained explicit demographic cues, and the models were asked to imagine a writer reproducing those cues. Control agreement ranged from 66.8 to 100 percent, and crucially, the mismatches were almost all explicit model outputs rather than parser failures. Some systems simply generated a gender inconsistent with the cue in the prompt, even when the task seemed to require preserving it. Excluding the two lowest-agreement models did not erase the aggregate gap; it actually widened from 29.9 to 33.5 points at an 80 percent threshold and to 36.5 points at 85 percent, showing that the contrast was not an artifact of poorly behaved endpoints.</p>
<p>Perhaps the most intellectually honest feature of the paper is its relabeling of its own estimand. An earlier draft described the comparison as an occupational-status effect, but inspection of the stimuli revealed that status-related role content covaried with speech act, mood, agency, and deference. Authority prompts use decisions and directives; support prompts use reports, questions, and requests for permission. The experiment therefore cannot distinguish whether models respond to the occupation, the communicative register, the degree of agency, or some combination. The published version renames the contrast a role-and-register effect and calls for a preregistered crossed design that independently varies role content and register, for example by expressing high-authority work deferentially and support work directively.</p>
<p>For anyone deploying language models in workplace contexts, the takeaways are concrete. Audit the exact register of your prompts, because directness and deference carry demographic associations of their own. Report full gender distributions rather than a single odds ratio, since a relative shift can obscure that both conditions share the same majority category. Measure response diversity, because repeated names signal that your apparent sample of outputs is far narrower than it looks. And treat cost and latency as separate axes from fairness: in this run, the inexpensive Gemma showed the largest positive risk difference while the pricier Gemini 3.1 Pro showed a negative one, a descriptive pattern with no causal meaning but a clear warning that fairness cannot be inferred from a price tag. As generated text increasingly drafts our examples, training materials, and simulations, these invisible demographic defaults quietly shape who gets imagined as authoritative and who gets imagined as helpful, one prompt at a time.</p>
<p><strong>Subject of Research:</strong> Gender attribution defaults in large language model outputs under workplace role-and-register prompts</p>
<p><strong>Article Title:</strong> Gender attribution under workplace role-and-register prompts in contemporary large language models</p>
<p><strong>Article References:</strong> Shukla, R. (2026). Gender attribution under workplace role-and-register prompts in contemporary large language models. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 333. <a href="https://doi.org/10.1007/s41060-026-01332-1" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01332-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01332-1" rel="noopener noreferrer">10.1007/s41060-026-01332-1</a></p>
<p><strong>Keywords:</strong> large language models, gender attribution, algorithmic bias, workplace language, model heterogeneity, prompt design, communicative register, response diversity, statistical audit, AI fairness, occupational stereotypes, bootstrap inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252565</post-id>	</item>
	</channel>
</rss>
