<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>similarity retrieval &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/similarity-retrieval/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 08:21:23 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>similarity retrieval &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Assistants Forget Your Rules Because Similarity Search Is the Wrong Tool, Study Finds</title>
		<link>https://scienmag.com/ai-assistants-forget-your-rules-because-similarity-search-is-the-wrong-tool-study-finds/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 08:21:23 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI assistant behavior in programming tasks]]></category>
		<category><![CDATA[AI assistant memory limitations]]></category>
		<category><![CDATA[AI code assistant rule adherence]]></category>
		<category><![CDATA[AI hallucinations and memory failures]]></category>
		<category><![CDATA[BM25]]></category>
		<category><![CDATA[constraint adherence]]></category>
		<category><![CDATA[conversational memory]]></category>
		<category><![CDATA[data storage formats in AI systems]]></category>
		<category><![CDATA[dense bi-encoders]]></category>
		<category><![CDATA[episodic vs governing facts in AI]]></category>
		<category><![CDATA[evaluation benchmark]]></category>
		<category><![CDATA[governance fidelity]]></category>
		<category><![CDATA[improving AI long-term context retention]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[limitations of similarity-based search for AI recall]]></category>
		<category><![CDATA[long-term memory in conversational AI]]></category>
		<category><![CDATA[managing constraints and preferences in AI]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[similarity retrieval]]></category>
		<category><![CDATA[similarity search in AI]]></category>
		<category><![CDATA[structural flaws in AI information retrieval]]></category>
		<category><![CDATA[supersession]]></category>
		<category><![CDATA[typed graph memory]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=243777</guid>

					<description><![CDATA[A new study proves that similarity-based retrieval structurally cannot guarantee that an AI assistant's stored rules reach the model, and shows that type-conditional, supersession-aware retrieval fixes the failure.]]></description>
										<content:encoded><![CDATA[<p>Imagine telling your AI coding assistant on Monday, &#8220;never use pandas in this project; use polars instead.&#8221; The assistant acknowledges. All week the conversation drifts through logging, refactoring, and columnar storage formats. On Friday you ask for a quick utility to deduplicate some rows, and the assistant cheerfully writes code that imports pandas. The instruction was stored. It was indexed. It simply was not retrieved. A new study argues that this everyday failure is not a glitch or a hallucination but a structural flaw baked into the way nearly every conversational AI system handles long-term memory.</p>
<p>The research, published as an open-access paper in the International Journal of Data Science and Analytics by Oghenenefe Abeke, Rasheed Mohammad, and Haitham Mahmoud of Birmingham City University, draws a sharp line between two kinds of information an assistant accumulates. Episodic facts, such as a bug discussed last Tuesday or a code snippet shared in passing, are exactly what similarity-based retrieval was designed for: you want the items most relevant to the current question. But the authors identify a second class they call governing facts, including constraints, preferences, identity statements, commitments, and assigned roles. These items do not answer a query; they bind every query. Their force is constitutive rather than topical, and that distinction, the paper shows, is where the dominant architecture breaks down.</p>
<p>The core of the argument is a small but rigorous impossibility result. The authors prove that any retriever which ranks stored items by a type-blind similarity score and returns only a bounded top-k set can be defeated by an adversarial sequence of topically related noise. Because similarity scores cannot distinguish a binding rule from ordinary conversation, an attacker, or simply a chatty user, can always generate enough topically adjacent material to push the rule out of the retrieval window. The paper complements this worst-case theorem with an expected-case degradation law: under random topical noise, the probability that a governing fact survives in the top-k follows a binomial tail that falls below one-half at a predictable threshold roughly equal to the retrieval budget divided by the per-item outranking probability. In plain terms, the longer the conversation and the more the topic drifts, the more certain it becomes that your rules vanish.</p>
<p>The empirical evidence is striking. Across 5,040 retrieval measurements spanning three scenarios, five governing-fact types, eight noise levels, and seven retrieval substrates, every similarity-ranked retriever tested collapsed. Sparse methods such as TF-IDF and BM25, along with hybrid reciprocal-rank-fusion combinations, held perfect fidelity at zero noise but dropped below 0.10 governance fidelity by 100 adversarial noise pairs, with a phase transition between roughly 25 and 50 noise pairs. Dense bi-encoders fared even worse in the small-encoder panel: bge-small-en-v1.5 had already fallen to about 33 percent fidelity at just six noise pairs, and e5-small-v2 reached a floor of exactly zero from 50 pairs onward. The very property that makes dense retrievers powerful for open-domain question answering, recognising topical similarity beyond lexical overlap, is precisely what the adversarial noise exploits.</p>
<p>A natural objection is that bigger, better encoders might solve the problem. A follow-up sweep tested bge-large and e5-large, each roughly ten times larger than the small models, plus a cross-encoder reranker applied to an enlarged candidate shortlist. The verdict: scale and reranking postpone the collapse but do not prevent it. The large encoders still fell steeply, reaching 0.067 and 0.200 fidelity by 200 noise pairs, and the reranker, which lasted longest, still dropped to 0.140 at 500 pairs. Importantly, the same degradation appeared under random topical noise rather than adversarial construction, ruling out an artefact of the benchmark design. The collapse, the authors conclude, is intrinsic to similarity ranking as such, not to any particular encoder or its size.</p>
<p>The remedy the authors propose is disarmingly simple in principle: type-conditional retrieval. Instead of ranking governing facts against noise, retrieve them unconditionally by their type and scope. They instantiate this as Stratified Context Reconstruction (SCR) in Satchel, a typed graph memory where each node carries a type drawn from constraint, preference, identity, commitment, role, or episodic, along with polarity, a target predicate, and a scope expression. At read time, the governing stage applies a type-and-scope filter with no similarity scoring at all, while the episodic stage runs whatever retriever the practitioner prefers. Governing facts reach the language model along a path that never passes through a similarity ranker, so governance fidelity is 1.0 by construction whenever write-time typing is correct. In the benchmark, SCR held fidelity at 1.0 across every noise level, with a Cohen&#8217;s d of about 1.06 against the similarity retrievers and a Bonferroni-corrected p value near 2.6 times ten to the minus twenty-fifth.</p>
<p>But the story has a second act. Governing facts get revised. A user forbids pandas on Monday, permits it in a legacy module on Wednesday, and reverses the exception on Friday. Naive SCR, which returns every in-scope governing fact unconditionally, now hands the model a contradictory rule set containing all three statements. The authors prove a second impossibility: no conflict-blind retriever, one that ignores supersession relations and validity intervals, can guarantee governance consistency, and the error grows linearly in revision depth. Their revision-depth benchmark confirms the prediction. Naive SCR returned up to fifteen stale superseded facts per query, and its governance minimality, the precision of the governing block, decayed to 0.25. The extended remedy, SCR-T, consults validity intervals and supersession edges to return exactly the currently binding set, and it is the only strategy that holds fidelity, consistency, and minimality jointly at 1.0 across every revision depth. Even production-style baselines such as pinned rule blocks and metadata-filtered retrieval replicated the temporal failure, accumulating eight stale facts per query at the deepest revision level.</p>
<p>Does the retrieval-layer property matter downstream? The authors tested four small open-weight language models and found that SCR improved constraint adherence over a TF-IDF retrieval-augmented baseline by 17.59 percentage points, with a 95 percent confidence interval of 13.19 to 21.76 and p below 0.001. A further controlled four-arm experiment under the paper&#8217;s own adversarial noise, crossing retrieval strategy with prompt formatting, showed that the advantage is a high-noise phenomenon: at zero and mild noise the two arms were indistinguishable, but as noise grew the RAG baseline&#8217;s constraint recall fell monotonically while SCR&#8217;s held steady, widening the gap to about 0.15 at 500 noise pairs. Because the advantage survived when both strategies were rendered in identical formatting, the effect is attributable to what SCR retrieves rather than how the block looks.</p>
<p>The authors are candid about limitations. The guarantee depends on write-time classification: deterministic templates captured about 78 percent of governing facts in their deployment, with a small LLM-as-judge classifier handling the rest at roughly 86 percent accuracy, putting realistic deployed fidelity in the 0.95 to 0.99 range. SCR-T adds a second write-time dependency, inferring supersession edges, though the authors stress that both problems are moved to a place where they are checkable and auditable in a typed graph, unlike an implicit recency policy buried in a vector store. The scenarios are scripted and English-only, the taxonomy of five governing-fact types may not be exhaustive, and no full system-to-system comparison against frameworks such as Mem0, Zep, or A-MEM was run, though the authors note those systems share the similarity-ranked substrate whose failure the paper formalises.</p>
<p>The broader implication is a challenge to how the field evaluates conversational memory. The authors propose three governance invariants, fidelity, consistency, and minimality, reported jointly alongside standard recall, plus traceability as a structural property of typed memory. Standard recall captures episodic competence; the invariants capture whether a system structurally honours the rules users entrust to it, completely, exclusively, and economically. On cost, the approach is practical: the governing block averaged 612 tokens, a 3 to 30 percent overhead on typical episodic budgets, and unlike budget-enlargement fixes, whose required retrieval window scales linearly with conversation length, the governing stage does not grow as sessions lengthen. As AI assistants are asked to remember more, for longer, the message of this paper is blunt: a rule that is not retrieved cannot be followed, and no amount of embedding quality will change that.</p>
<p><strong>Subject of Research:</strong> Structural limitations of similarity-ranked retrieval for governing facts in long-term conversational memory of language model systems</p>
<p><strong>Article Title:</strong> Beyond similarity retrieval: governance fidelity, consistency, and minimality in long-term conversational memory</p>
<p><strong>Article References:</strong> Abeke, O., Mohammad, R., &amp; Mahmoud, H. (2026). Beyond similarity retrieval: governance fidelity, consistency, and minimality in long-term conversational memory. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 328. <a href="https://doi.org/10.1007/s41060-026-01290-8" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01290-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01290-8" rel="noopener noreferrer">10.1007/s41060-026-01290-8</a></p>
<p><strong>Keywords:</strong> conversational memory, retrieval-augmented generation, large language models, governance fidelity, similarity retrieval, typed graph memory, knowledge graphs, constraint adherence, supersession, evaluation benchmark, dense bi-encoders, BM25</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">243777</post-id>	</item>
	</channel>
</rss>
