<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>combining NLP and deterministic mathematics in procurement &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/combining-nlp-and-deterministic-mathematics-in-procurement/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 07:32:54 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>combining NLP and deterministic mathematics in procurement &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Meets Human Judgment: A Governed Framework for Smarter Supplier Selection</title>
		<link>https://scienmag.com/ai-meets-human-judgment-a-governed-framework-for-smarter-supplier-selection/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 07:32:54 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI governance]]></category>
		<category><![CDATA[AI-driven supplier selection framework]]></category>
		<category><![CDATA[auditability]]></category>
		<category><![CDATA[challenges of trust and reliability in AI procurement tools]]></category>
		<category><![CDATA[combining NLP and deterministic mathematics in procurement]]></category>
		<category><![CDATA[criteria formalization in supplier selection]]></category>
		<category><![CDATA[evidence extraction]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[governed decision-making in procurement]]></category>
		<category><![CDATA[handling unstructured procurement documents with AI]]></category>
		<category><![CDATA[human oversight in AI-powered supplier ranking]]></category>
		<category><![CDATA[human-in-the-loop]]></category>
		<category><![CDATA[improving procurement accuracy with AI-human collaboration]]></category>
		<category><![CDATA[integration of large language models in supplier evaluation]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[multi-criteria decision making]]></category>
		<category><![CDATA[multi-criteria decision-making methods for sourcing]]></category>
		<category><![CDATA[procurement]]></category>
		<category><![CDATA[supplier selection]]></category>
		<category><![CDATA[Supply Chain Management]]></category>
		<category><![CDATA[sustainable and compliant supplier assessment]]></category>
		<category><![CDATA[sustainable sourcing]]></category>
		<category><![CDATA[TOPSIS]]></category>
		<category><![CDATA[transparency and traceability in supplier decisions]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226434</guid>

					<description><![CDATA[Researchers have built a governed framework that pairs large language models with deterministic multi-criteria analysis and human checkpoints to make supplier selection traceable, explainable, and accountable.]]></description>
										<content:encoded><![CDATA[<p>Choosing the right supplier has never been a simple matter of comparing prices. Modern procurement teams must weigh cost against quality, delivery reliability against sustainability credentials, and hard compliance rules against the messy, contradictory evidence that arrives in supplier documents, news reports, and certification records. A new study published in Machine Learning with Applications proposes a way to bring order to this complexity: a governed framework that combines the linguistic fluency of large language models with the unyielding precision of deterministic mathematics, all under the watchful eyes of human decision makers.</p>
<p>The research, conducted by Mohammadreza Rezaei, Sridhar Sreemulnath Iyer, and Omid Fatahi Valilai, tackles a gap that has long frustrated both researchers and practitioners. Multi-criteria decision-making methods such as TOPSIS can rank suppliers rigorously, but only after someone has formalised the criteria, weights, constraints, and evidence. Large language models, meanwhile, excel at reading natural-language procurement requests and extracting information from unstructured documents, yet they cannot be trusted to perform arithmetic or to rank suppliers autonomously. Existing approaches, the authors argue, address fragments of the problem while leaving the complete path from a spoken or written sourcing request to a traceable, defensible supplier decision largely uncharted.</p>
<p>The framework&#8217;s central design principle is a strict division of labour. Language model components handle the semantic work: interpreting what a procurement request actually means, formulating candidate evaluation scenarios, extracting claims from supplier documents, and explaining completed results in plain language. Deterministic services handle everything numerical: constructing weight vectors, applying constraints and evidence policies, calculating multi-criteria scores, resolving ties, and fixing the final supplier order. Crucially, the language models never touch the numbers. When a manager asks to place greater emphasis on geographic fit, the model proposes only a direction and a bounded strength category; a deterministic controller then converts that advice into a valid numerical weight vector through a constrained projection that enforces unit sums and policy-defined bounds.</p>
<p>Human authority is organised through four checkpoints. The data owner approves the structured supplier table and its quality record. The procurement analyst approves the scenario, its criteria, constraints, and the numerical weights. The evidence reviewer confirms whether extracted claims genuinely support their cited passages and what analytical role each claim may play. Finally, the decision owner reviews the complete package, including rankings, comparisons, explanations, and caveats, before any action is taken. Every consequential action generates an append-only audit event recording who acted, in what role, when, on what object, and why, creating a traceable chain from the original request to the final sourcing decision.</p>
<p>The empirical evaluation is unusually comprehensive for this field. Using more than 22,000 real procurement cases drawn from the European Union&#8217;s Tenders Electronic Daily database, the team first compared deterministic ranking methods. TOPSIS, which ranks suppliers by their distance from ideal and anti-ideal solutions after vector normalisation and weighting, outperformed frequency-based ranking, weighted-sum models, and VIKOR. But the more striking finding concerned the supplier representation itself. A generic four-feature baseline placed the historically awarded supplier in the top ten of only 5.259 percent of cases. When the researchers added procurement-specific feature groups capturing category fit, geographic relevance, and prior buyer relationships, that figure rose to 31.642 percent, and the median rank of the reference supplier improved from 914 to just 34.</p>
<p>The scenario experiments revealed how the framework responds when sourcing priorities change. Under a continuity profile that increases the weight on prior buyer relationships, the top-ten overlap with the balanced baseline remained high at 0.9369. Under a diversification profile that reduces reliance on existing relationships, overlap fell to 0.8304, yet suppliers entering the top ten had stronger geographic fit than those exiting in 96.984 percent of turnover cases. In other words, the rankings moved in traceable, explainable ways that followed the approved priorities, allowing procurement professionals to compare competing strategies and see exactly which suppliers enter or leave the shortlist under each.</p>
<p>The language model experiments were equally revealing. In intent formulation, the model correctly identified the scenario in 146 of 150 procurement requests and made the right proceed-or-clarify decision in all but one, with constraint and evidence-request extraction achieving F1 scores of 0.990 and 0.996 respectively. Evidence extraction proved robust: across 40 supplier evidence packets containing 160 expected claims, the module achieved a mean claim F1 of 0.928 and grounded 99.5 percent of its quotations in the actual source text, with unsupported quotations appearing in only 0.5 percent of cases. Yet the study is candid about weaknesses. Temporal-status accuracy fell to 0.834 and evidence-role accuracy to 0.748, meaning a claim can be quoted perfectly while still being assigned the wrong analytical meaning. Under conflicting evidence, performance degraded further, with hallucination rising to 4.3 percent.</p>
<p>Perhaps the most instructive experiment examined what happens when extracted qualitative evidence feeds into a deterministic ranking. In a controlled benchmark combining profit, lead time, and ESG evidence, all six material compliance violations were detected and eligibility accuracy was perfect. However, all twelve suppliers with mixed sustainability evidence, those with an emissions-reduction plan whose achieved reductions remained unverified, were mapped to an overly favourable certified state. The consequence was subtle: shortlists remained largely intact, with top-three overlap between 0.889 and 1.000, but the first-ranked supplier changed in many cases. The lesson is clear and important. High grounding and perfect violation detection can coexist with incorrect semantic qualification of evidence, which is precisely why the framework routes consequential evidence through human review before it can affect eligibility or scores.</p>
<p>Explanation fidelity completed the picture. In 24 controlled cases, the explanation module preserved the exact deterministic supplier order, used only valid supplier and evidence identifiers, and correctly escalated every case involving constraints, close score margins, negative evidence, or missing data. Its main weakness was over-inclusiveness: explanations sometimes mentioned all available feature groups rather than isolating those that genuinely drove the ranking. The authors treat this as a review item rather than a failure, presenting the narrative alongside the fixed score tables so the decision owner can inspect both.</p>
<p>The broader significance of this work lies in its refusal to choose between automation and augmentation. Rather than asking whether language models can replace procurement professionals, the framework specifies exactly which tasks each participant may perform and where approval is required. It also supports adaptive supplier reassessment: when market conditions, regulations, or supplier circumstances change, a revised scenario passes through the same validated, approved, deterministic pathway rather than someone quietly editing an existing ranking. The authors acknowledge that organisational deployment, including review effort and decision-owner behaviour, still requires field evaluation, and that benchmarking across multiple model families remains future work. But as a blueprint for bringing generative AI into high-stakes sourcing decisions without surrendering accountability, the study offers something rare: a complete, tested architecture in which the flexibility of language meets the rigour of mathematics, and where the final word always belongs to a human being.</p>
<p><strong>Subject of Research:</strong> A governed human-in-the-loop framework integrating large language models with deterministic multi-criteria supplier evaluation</p>
<p><strong>Article Title:</strong> A governed framework combining large language models and human oversight for supplier selection</p>
<p><strong>Article References:</strong> Rezaei, M., Iyer, S. S., &amp; Fatahi Valilai, O. (2026). A governed framework combining large language models and human oversight for supplier selection. <em>Machine Learning with Applications, 26</em>, Article 101028. <a href="https://doi.org/10.1016/j.mlwa.2026.101028" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101028</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101028" rel="noopener noreferrer">10.1016/j.mlwa.2026.101028</a></p>
<p><strong>Keywords:</strong> large language models, supplier selection, procurement, multi-criteria decision-making, TOPSIS, human-in-the-loop, supply chain management, explainable AI, evidence extraction, AI governance, sustainable sourcing, auditability</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226434</post-id>	</item>
	</channel>
</rss>
