Sunday, October 11, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos

October 11, 2026
in Technology and Engineering
Nathaniel Bowman
By Nathaniel Bowman Scienmag Editorial Profile - Precision Oncology
Reading Time: 5 mins read
0
AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos

AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence has been edging closer to the rooms where cancer treatment decisions are made, but a stubborn problem has kept it at arm’s length: when you ask a large language model the same clinical question twice, you do not always get the same answer. Now a team at University Hospital Schleswig-Holstein in Lübeck, Germany, reports that a carefully engineered framework can make an AI’s tumor-board recommendations almost perfectly repeatable. In a study published in Scientific Reports, the researchers simulated multidisciplinary sarcoma tumor boards using the Claude Opus 4.5 model and found that 94 percent of its decision codes were identical across repeated runs of the same case — a dramatic improvement over the roughly 20 percent consistency documented in earlier, free-form benchmarks.

The stakes are higher than they might sound. Sarcomas, a family of more than 100 rare connective-tissue cancers, are among the most demanding cases any tumor board faces. Decisions about biopsy route, surgical margins, reconstruction, radiotherapy and chemotherapy must be woven together, and early missteps can cost patients a limb, a function, or a cure. Previous studies testing language models against human boards in sarcoma were sobering: in a benchmark against 21 European sarcoma centers, models including GPT-4 and Claude 3.5 Sonnet showed consistency in only one in five cases, and their best alignment with human consensus reached just 60 percent. Worse, the models occasionally invented guideline citations and fabricated survival statistics — the notorious hallucination problem — with explicit guideline derivation identifiable in fewer than half of responses.

The German team, led by Tekoshin Ammo and Moritz Englich of the Department of Plastic, Reconstructive and Aesthetic Surgery, attacked the problem architecturally rather than with better pleading. Their framework forces the model to return its answers through a tool-use interface as structured JSON data, validated at run time against a schema defined in the Python library Pydantic. The schema covers 21 individual decision codes across nine clinical domains — imaging, staging, surgical strategy, margin goals, reconstruction, functional reconstruction, radiotherapy strategy and dose, and systemic therapy. Categorical fields are locked to controlled vocabularies: the model cannot choose between “wide local excision” and “definitive surgical excision,” because only the enumerated value wide_excision passes validation. Inference was run at temperature 0.0, the setting that suppresses random sampling and pushes the model toward its most probable output.

On top of the schema, the system prompt layered seven coordinated prompting techniques, including role prompting, rule-based prompting, negative prompting, and conditional-logic rules that couple decisions to one another — for instance, instructing the model to record “none” for first-line chemotherapy when the overall systemic-therapy decision is negative. Two semantic clarifications resolved ambiguities that had plagued earlier iterations: whether a “True” for an imaging recommendation meant the scan still needed to be performed or merely belonged to the recommended workup, and whether a functional-reconstruction flag referred to a procedure definitively planned. These refinements emerged from three development phases run on a ten-case pilot subset before the framework was frozen and applied to the full dataset.

The evaluation was exhaustive by the standards of clinical AI studies. Fifty-one fully de-identified sarcoma cases, spanning 17 histological subtypes and collected between 2020 and 2025, were each processed three times — one case four times — yielding 154 runs under identical inputs and settings. Reproducibility was defined with deliberate harshness: a decision code counted as stable only if every run of a case returned exactly the same value, with no normalization, case-folding or semantic mapping. Of 1,071 decision-code instances, 1,007 were stable, for an overall reproducibility of 94.0 percent (95 percent confidence interval 92.0 to 95.8). In the 41 cases never touched during development, the figure was 94.2 percent, reassuringly close to the 93.3 percent seen in the development subset. Seventeen cases — a third of the dataset — achieved perfect agreement across all 21 codes.

The residual instability was not spread evenly. Imaging and systemic-therapy recommendations were the most reliable domains at 98.0 percent each, while radiotherapy lagged at 88.9 percent, driven largely by the model oscillating between different fractionation schemes such as 50 Gray in 25 fractions, 60 in 30, and 66 in 33. The single least stable code was the reconstruction decision, at 72.5 percent, where the model alternated between adjacent options such as no flap, local flap, and pedicled flap — choices that surgeons themselves often weigh against one another depending on intra-operative findings. Whether such alternation reflects legitimate clinical alternatives or model error, the authors stress, cannot be determined without expert adjudication, which this study deliberately did not perform.

Two comparison experiments sharpened the picture of what actually drives stability. In a post hoc ablation on 15 cases, removing the semantic and conditional-logic rules while keeping schema enforcement dropped reproducibility from 94.6 to 82.5 percent — a 12.1-point difference. But the improvement turned out to be concentrated almost entirely in three free-text fields, where exact wording agreement jumped by 60 points; across the 18 categorical and numeric codes, the difference was only 4.1 points, with a confidence interval that included zero. In other words, the rules mostly taught the model to phrase things identically, not to decide things identically. A separate free-text arm, in which the model answered in prose without any schema, was even more revealing: the model stated a decision for only 37.6 percent of code instances, and seven of the 21 codes — including all radiotherapy dose fields and the immunotherapy target — were never addressed at all. Schema enforcement, it appears, matters chiefly for completeness: it forces the model to answer every question, making its outputs answerable and comparable.

On the hallucination front, the results were cautiously encouraging. A single-reviewer screen of all 154 runs, using five categories defined before evaluation began, found no fabricated survival statistics, no invented references to ESMO, NCCN or other guidelines, and no anatomical or imaging measurements absent from the input. General clinical-knowledge claims and patient-specific inferences did occur, but none were judged factually incorrect under the study protocol. The authors are careful to note the limits of this finding: the screen was performed by one reviewer, covered only pre-specified categories, and was not independently adjudicated, so it should be read as the result of a restricted screen rather than proof that the framework is hallucination-proof.

The most important caveat in the entire study is also the simplest: reproducibility is not correctness. A model can be perfectly consistent and consistently wrong, and the authors state plainly that their results describe the consistency of outputs, not their clinical quality, and do not indicate readiness for clinical decision support. The estimates are also specific to a single model, Claude Opus 4.5, under a single schema; whether other models achieve comparable stability, and whether the schema itself rather than the model’s sampling behavior produced it, remains untested. Failed validation attempts were not logged, so the true schema-conformity rate across all attempts is unknown. The team is now preparing the study this pilot was designed to enable: a formal concordance analysis comparing the schema-enforced recommendations against documented historical tumor-board decisions in the same cases, alongside a multi-model comparison and expert plausibility assessment. Until then, the message is methodological rather than clinical — structured, schema-enforced output can turn a volatile AI into a measurable one, and measurement is the prerequisite for everything that follows.

Subject of Research: Reproducibility of schema-enforced large language model decision support in simulated sarcoma multidisciplinary tumor boards

Article Title: A schema-enforced large language model framework produces largely reproducible decision codes in simulated sarcoma tumor boards

Article References: Ammo, T., Englich, M., Adrego, F. D. S., Jacobi, M., Wilckens, A., Kruppa, P., Bonaventura, B., Niemuth, S., Weiler, S. M., Wachenfeld-Teschner, V., Boos, A. M., & Freund, G. (2026). A schema-enforced large language model framework produces largely reproducible decision codes in simulated sarcoma tumor boards. Scientific Reports, 16(1), Article 31714. https://doi.org/10.1038/s41598-026-74868-8

Image Credits: AI Generated

DOI: 10.1038/s41598-026-74868-8

Keywords: large language models, sarcoma, tumor board, reproducibility, schema validation, hallucination, clinical decision support, structured output, oncology, artificial intelligence, Claude Opus 4.5, multidisciplinary team

Cite Scienmag News

Nathaniel Bowman. (October 11, 2026). AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos. Scienmag. https://scienmag.com/ai-tumor-board-passes-reproducibility-test-structured-output-tames-sarcoma-llm-chaos/

Nathaniel Bowman. "AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos." Scienmag, 11 October 2026, https://scienmag.com/ai-tumor-board-passes-reproducibility-test-structured-output-tames-sarcoma-llm-chaos/. Accessed 11 October 2026.

Nathaniel Bowman. "AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos." Scienmag. October 11, 2026. https://scienmag.com/ai-tumor-board-passes-reproducibility-test-structured-output-tames-sarcoma-llm-chaos/

Tags: AI consistency in clinical recommendationsAI decision code reproducibilityAI framework for tumor boardsAI in personalized cancer careAI tumor board reproducibilityAI-assisted sarcoma managementArtificial Intelligencechallenges of LLM in rare cancersClaude Opus 4.5clinical decision supportclinical decision support with AIhallucinationimproving AI reliability in oncologylarge language modelslarge language models in cancer treatmentmultidisciplinary teamoncologyreproducibilitysarcomasarcoma multidisciplinary decision-makingschema validationstructured outputstructured output in medical AItumor board
Share26Tweet16
Previous Post

Gut bacteria may decide who gets a fever after vaccination

Next Post

Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer

Related Posts

Prediction Set Size Doubles as a Free Early-Warning System for Corrupted Machine Learning Data
Technology and Engineering

Prediction Set Size Doubles as a Free Early-Warning System for Corrupted Machine Learning Data

October 11, 2026
Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer
Technology and Engineering

Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer

October 11, 2026
AI Learns Art History: Knowledge Graphs Help Machines Link Paintings to the World
Technology and Engineering

AI Learns Art History: Knowledge Graphs Help Machines Link Paintings to the World

October 11, 2026
Old-School Trees Beat Attention Models in Student Performance Prediction, Rigorous Benchmark Finds
Technology and Engineering

Old-School Trees Beat Attention Models in Student Performance Prediction, Rigorous Benchmark Finds

October 11, 2026
Deformation Rewrites the Grain Boundary Map of an Ordered Nickel Alloy
Technology and Engineering

Deformation Rewrites the Grain Boundary Map of an Ordered Nickel Alloy

October 11, 2026
Where the Power Comes From Could Decide the Carbon Cost of Your Concrete
Technology and Engineering

Where the Power Comes From Could Decide the Carbon Cost of Your Concrete

October 11, 2026
Next Post
Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer

Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Prediction Set Size Doubles as a Free Early-Warning System for Corrupted Machine Learning Data
  • Tumor Protein LRG1 Hijacks Neutrophil Mitochondria to Fuel Dangerous Vessel Growth in Bladder Cancer
  • AI Tumor Board Passes Reproducibility Test: Structured Output Tames Sarcoma LLM Chaos
  • Gut bacteria may decide who gets a fever after vaccination

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading