<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>systematic AI trustworthiness evaluation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/systematic-ai-trustworthiness-evaluation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 13:08:35 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>systematic AI trustworthiness evaluation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Trustworthy AI Pipeline Slashes Bias and Dangerous Errors in Real-World Systems</title>
		<link>https://scienmag.com/new-trustworthy-ai-pipeline-slashes-bias-and-dangerous-errors-in-real-world-systems/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 13:08:35 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI fairness and security integration]]></category>
		<category><![CDATA[AI governance]]></category>
		<category><![CDATA[AI privacy protection methods]]></category>
		<category><![CDATA[bias reduction in AI systems]]></category>
		<category><![CDATA[ethical AI development practices]]></category>
		<category><![CDATA[Explainability]]></category>
		<category><![CDATA[explainability in artificial intelligence]]></category>
		<category><![CDATA[fairness]]></category>
		<category><![CDATA[healthcare AI]]></category>
		<category><![CDATA[holistic AI safety pipeline]]></category>
		<category><![CDATA[industrial AI]]></category>
		<category><![CDATA[interdisciplinary AI research collaboration]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[real-world AI error mitigation]]></category>
		<category><![CDATA[robust AI system design]]></category>
		<category><![CDATA[robustness]]></category>
		<category><![CDATA[safety]]></category>
		<category><![CDATA[safety and robustness in AI lifecycle]]></category>
		<category><![CDATA[SDG 9]]></category>
		<category><![CDATA[security]]></category>
		<category><![CDATA[systematic AI trustworthiness evaluation]]></category>
		<category><![CDATA[trustworthy AI]]></category>
		<category><![CDATA[trustworthy AI development]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222870</guid>

					<description><![CDATA[Researchers have unveiled a Trustworthy AI pipeline that embeds privacy, fairness, security, robustness, safety, and explainability across the entire development lifecycle, cutting demographic fairness gaps from 46 percent to under 5 percent while maintaining high accuracy on industrial and healthcare benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence systems now sit at the heart of decisions that shape lives, from hospital diagnostics to industrial automation, yet a persistent problem has haunted the field: the traits that make AI trustworthy are usually bolted on late, tested in isolation, and rarely guaranteed together. A team of researchers from Anglia Ruskin University in the United Kingdom, the Fraunhofer Institute for Biomedical Engineering IBMT in Germany, MAGGIOLI S.P.A. in Italy, and the University of Piraeus in Greece believes it has a way out of that trap. In a study published in the journal Complex and Intelligent Systems, they present a holistic Trustworthy AI pipeline that weaves six essential characteristics—privacy, fairness, security, robustness, safety, and explainability—directly into the development lifecycle rather than treating them as an afterthought.</p>
<p>The motivation behind the work is stark. AI failures do not merely produce inaccurate predictions; they can encode bias, deepen inequalities, and generate outcomes that no one can explain or audit. The authors argue that although the research community has identified the key trustworthiness characteristics, existing practices lack a unified pipeline for systematically aligning those traits across every phase of development. Performance metrics such as accuracy have dominated evaluation for years, but accuracy alone says nothing about whether a model leaks private data, discriminates against demographic groups, or collapses when confronted with adversarial inputs. The new pipeline is designed to close precisely that gap.</p>
<p>At the technical core of the framework sits a novel component the researchers call the Adaptive Trustworthiness Integration algorithm, or ATI. Rather than allowing developers to address trustworthiness concerns in any order they like, the algorithm enforces an irreversibility-based ordering of the six characteristics, ensuring that foundational properties such as privacy are settled before downstream properties such as explainability are layered on top. At each phase boundary of the development process, the algorithm applies conditional gating: if a trustworthiness violation remains unresolved, it is blocked from propagating into the next stage. This gating mechanism is the pipeline&#8217;s safeguard against the all-too-common scenario in which a flaw discovered late in development is quietly patched over instead of properly fixed.</p>
<p>One of the most intellectually interesting outputs of the approach is an empirically measured inter-characteristic trade-off matrix. Improving one trustworthiness property often degrades another—tightening privacy protections can reduce model accuracy, and aggressive fairness constraints can complicate explainability. By measuring these interactions directly, the ATI algorithm makes the costs and benefits of every intervention transparent and reproducible. Developers are no longer guessing whether a fairness fix will quietly undermine robustness; they can see the trade-offs quantified and decide deliberately which balance suits their application and its regulatory context.</p>
<p>The pipeline was put through its paces in two very different settings. The first was an industrial pilot at Fraunhofer IBMT, which the team extended with multi-seed validation across three random initialisations and a large-scale public microscopy benchmark known as PanNuke, comprising 7,904 images of tissue used in computational pathology. On that benchmark the system achieved a mean average precision of 99.5 percent at an intersection-over-union threshold of 0.5, a striking result given that trustworthiness constraints were being enforced throughout. The second experiment targeted healthcare, using the Synthea synthetic patient dataset, where the pipeline maintained a predictive accuracy of 89.5 percent while satisfying its trustworthiness requirements.</p>
<p>The fairness results may be the headline finding. Across the experiments, the pipeline reduced demographic fairness gaps from 46 percent to under 5 percent, meaning that the difference in performance between demographic groups shrank by roughly an order of magnitude. In a healthcare context, that kind of reduction is not an abstract statistical nicety; it translates directly into fewer misdiagnoses concentrated in underserved populations. The researchers also report a 0 percent rate of clinically dangerous hallucinations, achieved through a four-layer safety architecture that screens model outputs before they can reach a decision-maker.</p>
<p>Equally important is the claim of consistency. Because the team validated the pipeline across multiple random seeds and across two distinct datasets, the results suggest that the trustworthiness guarantees are not artifacts of a lucky training run. Reproducibility has long been a weak point in AI research, where reported gains often evaporate under different initialisations or data splits. By demonstrating that its guarantees hold across datasets and random initialisations, the study offers something rarer than a strong benchmark score: evidence of a process that behaves predictably.</p>
<p>The work also carries a broader framing. The authors explicitly connect the pipeline to United Nations Sustainable Development Goal 9, which concerns industry, innovation, and infrastructure. They argue that trustworthy AI contributes to that goal by ensuring fair and privacy-preserving data collection, preventing system failures through adversarial hardening, and enabling transparent decision-making. The framing signals a growing recognition that AI trustworthiness is not only a technical challenge but a societal one, tied to how emerging technologies are deployed in critical infrastructure and public services.</p>
<p>The research was supported by substantial European and British public funding. It received backing from the European Union&#8217;s Horizon Europe Programme under the CUSTODES project, which develops certification approaches for composite systems of ICT products and services, and from the Digital Europe Programme under the EuDoros project, which supports deployment of ready-to-offer cybersecurity services. Additional support came from UK Research and Innovation through the CHIST-ERA programme under the AI4MultiGIS project, which focuses on intelligent geospatial handling in multi-GIS applications. That mix of funding streams reflects the pipeline&#8217;s ambition to serve industrial, security, and geospatial domains alike.</p>
<p>For practitioners, the significance of the study lies less in any single benchmark number than in the architectural lesson it encodes: trustworthiness is a lifecycle property, not a final test. By ordering the six characteristics irreversibly, gating every phase boundary, and quantifying trade-offs empirically, the pipeline turns trustworthiness from a vague aspiration into an enforceable engineering discipline. As regulators worldwide move toward mandatory AI assurance, frameworks of this kind may become the template against which future AI-enabled applications—especially in high-stakes domains like medicine—are built, audited, and ultimately trusted.</p>
<p><strong>Subject of Research:</strong> A holistic trustworthy AI pipeline integrating six trustworthiness characteristics across the AI development lifecycle</p>
<p><strong>Article Title:</strong> A holistic trustworthy AI pipeline for building trusted AI-enabled applications</p>
<p><strong>Article References:</strong> Sardar, B., Islam, S., Amelin, D., &amp; Papastergiou, S. (2026). A holistic trustworthy AI pipeline for building trusted AI-enabled applications. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02508-9" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02508-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02508-9" rel="noopener noreferrer">10.1007/s40747-026-02508-9</a></p>
<p><strong>Keywords:</strong> trustworthy AI, fairness, privacy, security, robustness, safety, explainability, healthcare AI, industrial AI, machine learning, AI governance, SDG 9</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222870</post-id>	</item>
	</channel>
</rss>
