<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>large language model use taxonomy &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/large-language-model-use-taxonomy/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 06 Sep 2026 10:36:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>large language model use taxonomy &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Classifying AI uses: judgment and epistemic control, from delegation to abdication</title>
		<link>https://scienmag.com/classifying-ai-uses-judgment-and-epistemic-control-from-delegation-to-abdication/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 06 Sep 2026 10:36:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI decision-making thresholds]]></category>
		<category><![CDATA[AI deployment and moral responsibility]]></category>
		<category><![CDATA[AI ethics classification]]></category>
		<category><![CDATA[AI governance and regulation]]></category>
		<category><![CDATA[AI in education and healthcare decision-making]]></category>
		<category><![CDATA[AI in education and public administration]]></category>
		<category><![CDATA[AI moral abdication threshold]]></category>
		<category><![CDATA[AI transparency and contestability]]></category>
		<category><![CDATA[automated grading systems ethical analysis]]></category>
		<category><![CDATA[critique of sector-based AI classification]]></category>
		<category><![CDATA[epistemic control in AI systems]]></category>
		<category><![CDATA[epistemic control in artificial intelligence]]></category>
		<category><![CDATA[ethical implications of AI deployment]]></category>
		<category><![CDATA[ethical implications of AI in public sectors]]></category>
		<category><![CDATA[human-AI decision-making boundaries]]></category>
		<category><![CDATA[human-AI interaction and responsibility]]></category>
		<category><![CDATA[judgment delegation in AI]]></category>
		<category><![CDATA[judgment delegation in AI systems]]></category>
		<category><![CDATA[large language model application taxonomy]]></category>
		<category><![CDATA[large language model use taxonomy]]></category>
		<category><![CDATA[moral abdication by AI]]></category>
		<category><![CDATA[transparency and contestability in AI systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/classifying-ai-uses-judgment-and-epistemic-control-from-delegation-to-abdication/</guid>

					<description><![CDATA[The way we talk about artificial intelligence in public debate is dominated by broad labels: &#8220;AI in education,&#8221; &#8220;AI in healthcare,&#8221; &#8220;AI in public administration.&#8221; But according to a new study, these categories hide the most ethically important question of all—not where an AI system is used, but what, exactly, is being handed over to [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>The way we talk about artificial intelligence in public debate is dominated by broad labels: &#8220;AI in education,&#8221; &#8220;AI in healthcare,&#8221; &#8220;AI in public administration.&#8221; But according to a new study, these categories hide the most ethically important question of all—not where an AI system is used, but what, exactly, is being handed over to it. In a paper published in the journal AI &amp; Society, philosopher and AI researcher Rainer Mühlhoff of Osnabrück University and the Einstein Center Digital Future in Berlin introduces a taxonomy of large language model (LLM) use that classifies applications by two dimensions: how much judgment is delegated to the machine, and how much epistemic control—the ability to define, inspect, and contest the informational basis of a decision—remains with the human user. The framework&#8217;s most striking conclusion is that some current deployments, including an automated grading assistant used widely in German schools and a U.S. government system used to flag veterans&#8217; contracts for cancellation, do not merely assist human judgment. They cross what Mühlhoff calls the threshold from delegation into &#8220;moral abdication.&#8221;</p>
<p>The study begins from a critique of the classification schemes that currently dominate AI governance. Sector-based approaches, such as those used by UNESCO and the EU&#8217;s High-Level Expert Group on AI, distinguish domains like education, healthcare, and administration. Task-based frameworks, common in machine-learning research, sort uses by function: summarization, classification, question answering. Risk-based approaches, most prominently the EU AI Act, group applications by potential societal harm. Mühlhoff argues that each of these operates at a level of abstraction too coarse to register the ethically salient differences between concrete practices of use. Two practices within the same sector—asking an LLM to fix the grammar of a clinician&#8217;s notes versus asking it to interpret symptoms and recommend a diagnosis—fall under the same label, yet differ radically in how much interpretive authority is offloaded and what happens when the system errs. &#8220;Using AI,&#8221; he writes, is treated as a single, homogeneous practice when in fact the uses grouped under that label range from low-stakes linguistic assistance to high-stakes interpretive and normative tasks.</p>
<p>The alternative Mühlhoff proposes is a two-dimensional grid. The first axis runs from delegated execution to delegated judgment. Delegated execution means the system carries out a well-specified operation whose standards of correctness are fixed independently of it—a calculator multiplying numbers, a spell checker flagging orthographic errors. Delegated judgment means the system must interpret underdetermined criteria: deciding what counts as a relevant error, how competing considerations should be weighed, how qualities translate into a grade or a decision. When judgment is offloaded, the model does not merely execute a procedure but participates in sense-making itself. Importantly, the paper notes, judgment can be delegated at different stages—not only when the model performs a task but earlier, when it interprets what the task even is. A prompt like &#8220;use school marks to grade the factual quality of the following text&#8221; requires the system to determine what the task itself amounts to, since no criteria are provided.</p>
<p>The second axis concerns epistemic control, and it distinguishes closed-knowledge from open-knowledge tasks. In closed-knowledge settings, the material the model operates on is explicitly provided—an uploaded document, a fixed dataset—and users can in principle trace whether outputs are grounded in the given evidence. In open-knowledge settings, the model draws on latent training data, opaque retrieval mechanisms, or general background knowledge, and users have limited means of assessing what an output is grounded in or whether it rests on spurious associations. Errors in open contexts are harder to detect and contest, particularly when outputs are eloquent, confident, and presented to users whose own background knowledge is limited. Combining the two axes yields four quadrants, from minimal offloading (execution over bounded material) to maximal offloading (open-ended judgment in open-knowledge contexts)—the quadrant where, the paper argues, responsibility becomes hardest to sustain.</p>
<p>To make the taxonomy operational, Mühlhoff treats prompts themselves as what he calls &#8220;delegation artifacts&#8221;: textual configurations in which tasks, standards, and authority are distributed between humans and machines. Because a prompt specifies what the model is asked to do, with what degree of discretion, and on what basis, prompt-level analysis offers direct insight into how cognitive labor is actually divided in practice. This matters, he argues, because prompts are frequently hidden behind polished user interfaces, making the true extent of delegation invisible to end users. The approach has limits—it applies to prompt-mediated use rather than fully autonomous agentic systems, and proprietary systems often shield their templates entirely—but he argues that prompt opacity is itself an ethically relevant feature of current deployments, underscoring the case for prompt disclosure and auditability.</p>
<p>The first case study applies the framework to automated grading in education, examining the &#8220;grading assistant&#8221; of the German platform Fobizz, a national market leader in educational AI services whose tools are available free to a substantial share of German teachers through state-level licensing agreements. Technically, the grading assistant operates as a wrapper around OpenAI&#8217;s GPT-4 model: teachers enter an assignment description, a sample solution, a keyword list of grading criteria, and a student&#8217;s submission through a web form; the system programmatically inserts these into a fixed prompt template sent to the language model. Reconstructing that template, Mühlhoff finds that it instructs the model to perform &#8220;error analysis and preliminary correction&#8221; and to contribute to &#8220;a thorough and differentiated evaluation&#8221;—functions that require deciding what counts as a relevant error, how aspects of performance are weighed, and how judgments aggregate into a grade. That is delegated judgment, not execution.</p>
<p>The epistemic analysis of the grading tool is more subtle. At first glance it appears closed-knowledge: the prompt incorporates a bounded assignment, sample solution, and student text. But the criteria governing evaluation remain fundamentally open. The prompt does not specify how the expectation horizon is to be interpreted, what thresholds separate grades, or how conflicting qualities—originality versus conformity—should be balanced. The model must supply those standards from background knowledge and generalized pedagogical norms that are neither fixed nor contestable. The tool therefore occupies an epistemically hybrid position: bounded inputs, open norms, weak epistemic control. And because teachers interact only with a friendly web form, the interface conceals the judgment-intensive prompting underneath. Mühlhoff locates the tool high on the judgment axis under epistemically open conditions—precisely the quadrant where responsibility gaps, systematic bias, and eroded contestability concentrate. Prior work he cites found such tools systematically fail to detect false content in student work.</p>
<p>The second case study is starker. Drawing on a 2025 ProPublica investigation, Mühlhoff analyzes the system deployed under the U.S. Department of Government Efficiency (DOGE) to review thousands of Department of Veterans Affairs contracts and label them &#8220;MUNCHABLE&#8221; or &#8220;NOT MUNCHABLE&#8221;—candidates for termination. The disclosed prompts supplied no operational definitions for the categories: the model was told contracts involving &#8220;direct patient care&#8221; were not munchable, and those &#8220;multiple layers removed from veterans care&#8221; were munchable, without defining either. ProPublica noted that neither the developer nor the model had the knowledge required for such determinations. The model was shown only the first 10,000 characters of each document, even though government contracts routinely run far longer, so judgments rested on partial excerpts. In one documented case, the model labeled all extracted fields &#8220;not found&#8221; and still concluded the contract was munchable—converting evidential absence directly into a negative decision.</p>
<p>For Mühlhoff, this configuration sits at the extreme end of both axes: maximal delegation of judgment under minimal epistemic control, combining zero-shot reasoning with forced binary decisions in a domain requiring sophisticated understanding of medical care, institutional management, and human resources. Unlike the grading case, the extremity is not obscured by interface design—it is explicit in the prompts themselves. He describes the delegation here as approaching &#8220;abdication&#8221;: interpretive and normative judgment over high-stakes administrative outcomes handed wholesale to a system, in a form that is both risky and unaccountable.</p>
<p>The ethical framework built on these cases draws on the Aristotelian tradition, in which responsibility is tied to an agent&#8217;s knowledge of what they are doing and control over the action. LLMs, Mühlhoff emphasizes, are not moral agents and cannot bear responsibility; responsibility remains with human users and institutions. The question is whether it can be meaningfully exercised. Where judgment is heavily delegated under weak epistemic control, users lack the knowledge and control needed to answer for outcomes—a constellation the paper links to &#8220;epistemic deference&#8221; and to evidence that nominally keeping humans in the loop does not guarantee independent judgment, since users tend to conform to algorithmic outputs. The paper&#8217;s central normative claim is that ethically defensible LLM use must remain structured as tool use: users must retain epistemic access, independent judgment, and justificatory authority. Where interface design seals prompts behind benign frontends, formally responsible users are practically prevented from exercising responsibility. The taxonomy, Mühlhoff argues, is ultimately an instrument of reflexive ethics—letting users, designers, and institutions ask not whether AI should be used in a domain, but how it is being used, and whether the conditions of responsible action still hold.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> A use-based taxonomy for ethically classifying large language model applications according to delegated judgment and epistemic control, applied to automated grading in education and AI-driven contract termination in U.S. public administration.</p>
<p><strong>Article Title:</strong> From delegation to moral abdication: classifying large language model uses by judgment and epistemic control</p>
<p><strong>Article References:</strong> Mühlhoff, R. (2026). From delegation to moral abdication: classifying large language model uses by judgment and epistemic control. <em>AI &amp; SOCIETY</em>. <a href="https://doi.org/10.1007/s00146-026-03281-6" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s00146-026-03281-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00146-026-03281-6" target="_blank" rel="noopener noreferrer">10.1007/s00146-026-03281-6</a></p>
<p><strong>Keywords:</strong> large language models, ethics of AI, delegated judgment, epistemic control, cognitive offloading, automated grading, algorithmic decision-making, responsibility, prompt analysis, AI governance</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">188661</post-id>	</item>
	</channel>
</rss>
