<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-assisted sarcoma management &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-assisted-sarcoma-management/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 01:35:18 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-assisted sarcoma management &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos</title>
		<link>https://scienmag.com/ai-tumor-board-passes-reproducibility-test-structured-output-tames-sarcoma-llm-chaos/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 01:35:18 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI consistency in clinical recommendations]]></category>
		<category><![CDATA[AI decision code reproducibility]]></category>
		<category><![CDATA[AI framework for tumor boards]]></category>
		<category><![CDATA[AI in personalized cancer care]]></category>
		<category><![CDATA[AI tumor board reproducibility]]></category>
		<category><![CDATA[AI-assisted sarcoma management]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[challenges of LLM in rare cancers]]></category>
		<category><![CDATA[Claude Opus 4.5]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[clinical decision support with AI]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[improving AI reliability in oncology]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in cancer treatment]]></category>
		<category><![CDATA[multidisciplinary team]]></category>
		<category><![CDATA[oncology]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[sarcoma]]></category>
		<category><![CDATA[sarcoma multidisciplinary decision-making]]></category>
		<category><![CDATA[schema validation]]></category>
		<category><![CDATA[structured output]]></category>
		<category><![CDATA[structured output in medical AI]]></category>
		<category><![CDATA[tumor board]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260722</guid>

					<description><![CDATA[A schema-enforced large language model framework achieved 94 percent reproducibility across simulated sarcoma tumor board decisions, a leap over earlier free-text benchmarks, though the authors caution that consistent outputs are not necessarily correct ones.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has been edging closer to the rooms where cancer treatment decisions are made, but a stubborn problem has kept it at arm&#8217;s length: when you ask a large language model the same clinical question twice, you do not always get the same answer. Now a team at University Hospital Schleswig-Holstein in Lübeck, Germany, reports that a carefully engineered framework can make an AI&#8217;s tumor-board recommendations almost perfectly repeatable. In a study published in Scientific Reports, the researchers simulated multidisciplinary sarcoma tumor boards using the Claude Opus 4.5 model and found that 94 percent of its decision codes were identical across repeated runs of the same case — a dramatic improvement over the roughly 20 percent consistency documented in earlier, free-form benchmarks.</p>
<p>The stakes are higher than they might sound. Sarcomas, a family of more than 100 rare connective-tissue cancers, are among the most demanding cases any tumor board faces. Decisions about biopsy route, surgical margins, reconstruction, radiotherapy and chemotherapy must be woven together, and early missteps can cost patients a limb, a function, or a cure. Previous studies testing language models against human boards in sarcoma were sobering: in a benchmark against 21 European sarcoma centers, models including GPT-4 and Claude 3.5 Sonnet showed consistency in only one in five cases, and their best alignment with human consensus reached just 60 percent. Worse, the models occasionally invented guideline citations and fabricated survival statistics — the notorious hallucination problem — with explicit guideline derivation identifiable in fewer than half of responses.</p>
<p>The German team, led by Tekoshin Ammo and Moritz Englich of the Department of Plastic, Reconstructive and Aesthetic Surgery, attacked the problem architecturally rather than with better pleading. Their framework forces the model to return its answers through a tool-use interface as structured JSON data, validated at run time against a schema defined in the Python library Pydantic. The schema covers 21 individual decision codes across nine clinical domains — imaging, staging, surgical strategy, margin goals, reconstruction, functional reconstruction, radiotherapy strategy and dose, and systemic therapy. Categorical fields are locked to controlled vocabularies: the model cannot choose between &#8220;wide local excision&#8221; and &#8220;definitive surgical excision,&#8221; because only the enumerated value wide_excision passes validation. Inference was run at temperature 0.0, the setting that suppresses random sampling and pushes the model toward its most probable output.</p>
<p>On top of the schema, the system prompt layered seven coordinated prompting techniques, including role prompting, rule-based prompting, negative prompting, and conditional-logic rules that couple decisions to one another — for instance, instructing the model to record &#8220;none&#8221; for first-line chemotherapy when the overall systemic-therapy decision is negative. Two semantic clarifications resolved ambiguities that had plagued earlier iterations: whether a &#8220;True&#8221; for an imaging recommendation meant the scan still needed to be performed or merely belonged to the recommended workup, and whether a functional-reconstruction flag referred to a procedure definitively planned. These refinements emerged from three development phases run on a ten-case pilot subset before the framework was frozen and applied to the full dataset.</p>
<p>The evaluation was exhaustive by the standards of clinical AI studies. Fifty-one fully de-identified sarcoma cases, spanning 17 histological subtypes and collected between 2020 and 2025, were each processed three times — one case four times — yielding 154 runs under identical inputs and settings. Reproducibility was defined with deliberate harshness: a decision code counted as stable only if every run of a case returned exactly the same value, with no normalization, case-folding or semantic mapping. Of 1,071 decision-code instances, 1,007 were stable, for an overall reproducibility of 94.0 percent (95 percent confidence interval 92.0 to 95.8). In the 41 cases never touched during development, the figure was 94.2 percent, reassuringly close to the 93.3 percent seen in the development subset. Seventeen cases — a third of the dataset — achieved perfect agreement across all 21 codes.</p>
<p>The residual instability was not spread evenly. Imaging and systemic-therapy recommendations were the most reliable domains at 98.0 percent each, while radiotherapy lagged at 88.9 percent, driven largely by the model oscillating between different fractionation schemes such as 50 Gray in 25 fractions, 60 in 30, and 66 in 33. The single least stable code was the reconstruction decision, at 72.5 percent, where the model alternated between adjacent options such as no flap, local flap, and pedicled flap — choices that surgeons themselves often weigh against one another depending on intra-operative findings. Whether such alternation reflects legitimate clinical alternatives or model error, the authors stress, cannot be determined without expert adjudication, which this study deliberately did not perform.</p>
<p>Two comparison experiments sharpened the picture of what actually drives stability. In a post hoc ablation on 15 cases, removing the semantic and conditional-logic rules while keeping schema enforcement dropped reproducibility from 94.6 to 82.5 percent — a 12.1-point difference. But the improvement turned out to be concentrated almost entirely in three free-text fields, where exact wording agreement jumped by 60 points; across the 18 categorical and numeric codes, the difference was only 4.1 points, with a confidence interval that included zero. In other words, the rules mostly taught the model to phrase things identically, not to decide things identically. A separate free-text arm, in which the model answered in prose without any schema, was even more revealing: the model stated a decision for only 37.6 percent of code instances, and seven of the 21 codes — including all radiotherapy dose fields and the immunotherapy target — were never addressed at all. Schema enforcement, it appears, matters chiefly for completeness: it forces the model to answer every question, making its outputs answerable and comparable.</p>
<p>On the hallucination front, the results were cautiously encouraging. A single-reviewer screen of all 154 runs, using five categories defined before evaluation began, found no fabricated survival statistics, no invented references to ESMO, NCCN or other guidelines, and no anatomical or imaging measurements absent from the input. General clinical-knowledge claims and patient-specific inferences did occur, but none were judged factually incorrect under the study protocol. The authors are careful to note the limits of this finding: the screen was performed by one reviewer, covered only pre-specified categories, and was not independently adjudicated, so it should be read as the result of a restricted screen rather than proof that the framework is hallucination-proof.</p>
<p>The most important caveat in the entire study is also the simplest: reproducibility is not correctness. A model can be perfectly consistent and consistently wrong, and the authors state plainly that their results describe the consistency of outputs, not their clinical quality, and do not indicate readiness for clinical decision support. The estimates are also specific to a single model, Claude Opus 4.5, under a single schema; whether other models achieve comparable stability, and whether the schema itself rather than the model&#8217;s sampling behavior produced it, remains untested. Failed validation attempts were not logged, so the true schema-conformity rate across all attempts is unknown. The team is now preparing the study this pilot was designed to enable: a formal concordance analysis comparing the schema-enforced recommendations against documented historical tumor-board decisions in the same cases, alongside a multi-model comparison and expert plausibility assessment. Until then, the message is methodological rather than clinical — structured, schema-enforced output can turn a volatile AI into a measurable one, and measurement is the prerequisite for everything that follows.</p>
<p><strong>Subject of Research:</strong> Reproducibility of schema-enforced large language model decision support in simulated sarcoma multidisciplinary tumor boards</p>
<p><strong>Article Title:</strong> A schema-enforced large language model framework produces largely reproducible decision codes in simulated sarcoma tumor boards</p>
<p><strong>Article References:</strong> Ammo, T., Englich, M., Adrego, F. D. S., Jacobi, M., Wilckens, A., Kruppa, P., Bonaventura, B., Niemuth, S., Weiler, S. M., Wachenfeld-Teschner, V., Boos, A. M., &amp; Freund, G. (2026). A schema-enforced large language model framework produces largely reproducible decision codes in simulated sarcoma tumor boards. <em>Scientific Reports, 16</em>(1), Article 31714. <a href="https://doi.org/10.1038/s41598-026-74868-8" rel="noopener noreferrer">https://doi.org/10.1038/s41598-026-74868-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41598-026-74868-8" rel="noopener noreferrer">10.1038/s41598-026-74868-8</a></p>
<p><strong>Keywords:</strong> large language models, sarcoma, tumor board, reproducibility, schema validation, hallucination, clinical decision support, structured output, oncology, artificial intelligence, Claude Opus 4.5, multidisciplinary team</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260722</post-id>	</item>
	</channel>
</rss>
