Retrieval-augmented generation, or RAG, has become the default prescription for making large language models smarter: when a model does not know something, fetch relevant documents and hand them to it before it answers. A new study in Results in Engineering delivers a striking counterexample. In metallurgical root cause analysis, the identification of why a component failed, adding retrieved literature made the models worse, not better, and for one model the penalty was statistically significant and substantial.
Researchers assembled a benchmark of 195 published failure analysis cases spanning nine metallurgical categories, from creep and fatigue to stress corrosion cracking, hydrogen embrittlement, and wear. For each case, they extracted a structured narrative describing the component, operating conditions, observations, and laboratory findings, while deliberately withholding the paper’s root cause conclusion. Three large language models were then asked to produce a 200 to 500 word narrative diagnosis, and an automated evaluator compared each diagnosis against the expert-determined ground truth from the original paper.
Without retrieval augmentation, the models performed surprisingly well. Xiaomi’s MiMo-v2.5 correctly identified the root cause in 76.4 percent of cases, DeepSeek-V4-Pro reached 80.0 percent, and DeepSeek-V4-Flash achieved 73.3 percent. When the dense RAG configuration was switched on, every model declined. MiMo-v2.5 fell to 64.1 percent under the filtered dense configuration, a drop of 12.3 percentage points that McNemar’s paired test confirmed as statistically significant (p = .003). The two DeepSeek models showed smaller, non-significant declines of 3.1 and 0.5 percentage points, revealing that the retrieval penalty is strongly model-dependent rather than universal.
The team did not stop at a single retriever. To test whether the degradation was an artifact of dense embedding search, they rebuilt the same 747-document corpus with a lexical BM25 index, the classic keyword-matching algorithm that underlies search engines. BM25 retrieval also failed to beat the no-retrieval baseline, scoring 70.3 percent against 76.4 percent, but it recovered more than half of the dense configuration’s loss. That pattern points the finger at the dense retriever itself: embedding similarity appears to select documents that are semantically similar on the surface yet mechanistically irrelevant to the failure at hand.
Digging into the retrieval logs, the authors identified several plausible culprits. In the default dense setup, duplicate file formats meant 74.4 percent of queries returned duplicate references, leaving only 3.9 unique documents among the requested five. More fundamentally, the dense retriever matched semantic similarity rather than mechanistic relevance, so a case of stress corrosion cracking could retrieve papers about superficially similar corrosion phenomena governed by entirely different physics. The authors frame these as hypotheses consistent with the data, not proven causes.
The error analysis revealed how retrieved context actively corrupts reasoning. In 42 cases, the model diagnosed correctly without retrieval but incorrectly with it, while only 18 cases moved the other way. The researchers describe three recurring patterns: reference anchoring, where the model aligns its conclusion with retrieved papers instead of the input evidence; mechanism contamination, where retrieved documents introduce failure mechanisms absent from the case; and confidence suppression. Common misdiagnoses included confusing fatigue with creep, mistaking stress corrosion cracking for corrosion fatigue, and blaming environmental factors when the true cause was a manufacturing defect.
Performance also varied by failure category, though the authors caution that small sample sizes make these patterns indicative rather than conclusive. Filtered dense RAG was most damaging for brittle fracture, where accuracy collapsed by 29.4 percentage points, followed by stress corrosion cracking and ductile overload, categories that often involve multi-factorial mechanisms sensitive to partially relevant context. Wear and fretting cases achieved the highest baseline accuracy at 91.7 percent. Notably, the retrieval corpus overlapped thematically with the ground-truth literature, which in principle could have inflated RAG’s apparent value, making the observed penalty all the more notable.
The evaluation methodology itself was carefully validated. Three independent failure analysis experts blind to the retrieval configuration reviewed stratified samples of both non-RAG and RAG outputs, achieving substantial agreement with one another. The automated evaluator agreed with expert consensus in 70.0 percent of non-RAG cases and 84.0 percent of RAG cases, and critically, it never rated a case correct when the experts rated it wrong. Using Gwet’s AC1, chosen over Cohen’s kappa for its stability with imbalanced categories, the team showed the reported accuracy differences are unlikely to be artifacts of evaluator unreliability.
The practical implications are pointed. Current LLMs appear viable as screening tools for text-based root cause analysis, with the best configuration correctly ruling out wrong mechanisms in 93.3 percent of cases, and the entire evaluation pipeline cost roughly US$0.43 in API fees. But the default assumption that bolting a literature database onto a model improves its analysis is not supported in this domain. The authors recommend deduplicating corpora before indexing, building mechanism-aware indexes rather than purely semantic ones, extracting mechanism-specific keywords for queries, and activating retrieval only when model confidence is low.
Limitations remain. Only one model family was tested across all four retrieval configurations, the pipeline processed text but not the micrographs and fractographs metallurgists routinely rely on, and some benchmark cases may have been encountered during pretraining, though memorization would affect both configurations equally and cannot explain the retrieval penalty. The conclusions apply to the specific retrieval setups evaluated, not to RAG in general. Still, the message lands with unusual force for the AI engineering community: in domains where models already carry deep internalized knowledge, indiscriminate retrieval does not add information, it adds noise, and the most expensive context can be the wrong context.
Subject of Research: Retrieval-augmented generation effects on large language model performance in metallurgical failure root cause analysis
Article Title: When RAG hurts: Retrieval-augmented generation degrades performance in metallurgical root cause analysis
Article References: Fauzi, A. R., Pratama, S. N., Purqon, A., Ramadhan, Y. Z., & Ardiansyah, M. (2026). When RAG hurts: Retrieval-augmented generation degrades performance in metallurgical root cause analysis. Results in Engineering, 32, Article 113136. https://doi.org/10.1016/j.rineng.2026.113136
Image Credits: AI Generated
DOI: 10.1016/j.rineng.2026.113136
Keywords: retrieval-augmented generation, large language models, metallurgical failure analysis, root cause analysis, RAG degradation, dense retrieval, BM25, materials science, artificial intelligence, failure mechanisms, DeepSeek, MiMo-v2.5
Cite Scienmag News
Denise Maddox. (October 2, 2026). When Retrieval Hurts: AI Gets Worse at Diagnosing Metal Failures With More References. Scienmag. https://scienmag.com/when-retrieval-hurts-ai-gets-worse-at-diagnosing-metal-failures-with-more-references/
Denise Maddox. "When Retrieval Hurts: AI Gets Worse at Diagnosing Metal Failures With More References." Scienmag, 2 October 2026, https://scienmag.com/when-retrieval-hurts-ai-gets-worse-at-diagnosing-metal-failures-with-more-references/. Accessed 2 October 2026.
Denise Maddox. "When Retrieval Hurts: AI Gets Worse at Diagnosing Metal Failures With More References." Scienmag. October 2, 2026. https://scienmag.com/when-retrieval-hurts-ai-gets-worse-at-diagnosing-metal-failures-with-more-references/

