<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>intelligent information systems &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/intelligent-information-systems/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 23:36:05 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>intelligent information systems &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Symptom Checkers Under Stress: Benchmark Reveals Which AI Models Truly Identify Disease</title>
		<link>https://scienmag.com/symptom-checkers-under-stress-benchmark-reveals-which-ai-models-truly-identify-disease/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:36:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI symptom checkers]]></category>
		<category><![CDATA[biomedical text mining]]></category>
		<category><![CDATA[biomedical transformers for disease diagnosis]]></category>
		<category><![CDATA[BlueBERT]]></category>
		<category><![CDATA[challenges in AI-based symptom analysis]]></category>
		<category><![CDATA[diagnostic accuracy of AI in medicine]]></category>
		<category><![CDATA[evaluation of symptom-to-disease mapping]]></category>
		<category><![CDATA[health condition identification]]></category>
		<category><![CDATA[health condition identification AI]]></category>
		<category><![CDATA[information retrieval]]></category>
		<category><![CDATA[intelligent information systems]]></category>
		<category><![CDATA[language technology in medical diagnosis]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in healthcare]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[medical language models evaluation]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in healthcare]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[robustness of AI health models]]></category>
		<category><![CDATA[symptom text classification]]></category>
		<category><![CDATA[symptom text dataset benchmarking]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[transformer models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203972</guid>

					<description><![CDATA[A harmonised benchmark of lexical models, biomedical transformers and retrieval-grounded large language models shows that internal test scores fail to predict how reliably AI systems identify health conditions when the data source changes.]]></description>
										<content:encoded><![CDATA[<p>Turning a short, casually written description of symptoms into a correct health condition is one of the most deceptively difficult tasks in applied artificial intelligence. A patient may write about a pounding headache, sensitivity to light and a stiff neck in a single sentence, and a useful system must map that fragmentary narrative onto the right diagnostic category. A new study published in the Journal of Intelligent Information Systems by Marina Bagić Babac of the University of Zagreb subjects the full spectrum of modern language technology to exactly this challenge, comparing classical lexical models, BERT-family biomedical transformers and retrieval-grounded large language models inside a single, carefully harmonised evaluation framework. The outcome is a sobering and highly practical message: the leaderboard that looks best on an internal benchmark tells only a fraction of the story, and the true measure of a medical language system is how it behaves when the data source shifts beneath it.</p>
<p>The study assembles a benchmark from multiple publicly available symptom-text datasets, each with its own writing style, vocabulary, label inventory and distributional quirks. Compiling these resources into one harmonised framework is itself a technical contribution, because short symptom narratives vary enormously in how formally they describe illness. Some datasets read like structured complaint forms, while others contain conversational phrasing with abbreviations, hedging and redundancy. Any model trained on one style may fail silently when confronted with another. Rather than rely on a single train-test split, the author therefore evaluates models under several stress conditions: performance on the in-distribution benchmark, transfer to external datasets that were never seen during training, cross-source transfer between distinct symptom-text sources, and leave-one-source-out validation, in which an entire data source is withheld to test generalisation to an unseen provenance.</p>
<p>Three families of models form the backbone of the comparison. The first consists of sparse lexical approaches, the direct descendants of classic information retrieval, which represent documents through term statistics such as weighted bag-of-words vectors. These methods are inexpensive, fast to train and easy to interpret, and they remain a sensible baseline for any text classification pipeline. The second family comprises transformer encoders in the BERT lineage, including domain-adapted checkpoints such as BlueBERT and BioBERT, which were pretrained on biomedical corpora and are widely regarded as state-of-the-art encoders for clinical and biomedical text. The third family is the newest and most computationally demanding: large language models, particularly instruction-tuned systems such as Meta-Llama-3-8B-Instruct and the medically oriented Llama3-Med42-8B, which can be prompted to produce condition labels in a generative manner rather than through a dedicated classification head.</p>
<p>On the internal benchmark, the hierarchy is clear. The biomedical transformers dominate, with BlueBERT achieving the strongest in-distribution performance, confirming that pretraining on biomedical text confers a substantial advantage when training and test data come from the same source. Domain-matched pretraining appears to align the model&#8217;s internal representations with clinical vocabulary, abbreviations and symptom phrasing, allowing relatively small encoder models to outperform both lexical baselines and much larger generative systems. This result alone would justify the enthusiasm for biomedical transformers that pervades the applied literature. But the study&#8217;s central insight emerges only when the rankings are re-examined under the harsher external evaluation protocols, where the picture changes in ways that a single internal score completely obscures.</p>
<p>When models are tested on external datasets and in leave-one-source-out validation, the previously dominant positions shift. Checkpoints that excelled in-distribution lose ground as the linguistic and distributional characteristics of the input change, revealing that part of their apparent superiority rested on source-specific patterns rather than a genuinely transferable understanding of symptom language. Cross-source transfer proves to be the acid test: models must cope with different label taxonomies, different levels of textual detail and different writing conventions simultaneously. The study demonstrates that a model&#8217;s rank on an internal benchmark is a poor predictor of its robustness under such source shift, and that any deployment decision based solely on in-distribution accuracy carries a real risk of overestimating a system&#8217;s reliability in production, where input provenance is never guaranteed to match the training corpus.</p>
<p>The large language models show a more textured profile, and this is where retrieval augmentation enters the story. Retrieval-grounded generation supplies the model with relevant retrieved examples or evidence before it answers, a technique popularised by retrieval-augmented generation methods for knowledge-intensive tasks. The study finds that grounding improves some instruction-tuned LLMs substantially, especially Meta-Llama-3-8B-Instruct and Llama3-Med42-8B, which benefit from having concrete reference cases and label candidates placed in their context. For other checkpoints, however, the benefit is only selective, suggesting that retrieval is not a universal upgrade but an interaction between the model&#8217;s parametric knowledge, its instruction-following behaviour and the quality of the retrieved shortlist. When the retrieved evidence is well matched to the query, generative models can approach or even rival the fine-tuned encoders; when the shortlist is noisy, the extra context can mislead the generation process.</p>
<p>This dependency on shortlist quality carries direct engineering consequences. A retrieval-grounded pipeline is a chain of components, and its accuracy is bounded by the weakest link. The study&#8217;s findings imply that practitioners must evaluate not only the language model but also the retrieval index, the embedding space used to find neighbours and the evidence format presented to the generator. A strong encoder classifier with a curated training set may outperform a sophisticated generative pipeline whose retrieval component surfaces irrelevant cases. Conversely, in low-resource or fast-changing label spaces where retraining a classifier is impractical, a well-grounded instruction-tuned LLM offers flexibility that no fixed classifier can match, because new conditions can be introduced by updating the evidence store rather than the weights.</p>
<p>Operational efficiency emerges as the third pillar of the decision framework. Sparse lexical models are nearly free at inference time and remain attractive for high-throughput triage at scale. Transformer encoders occupy a middle ground, delivering strong accuracy with moderate compute. Instruction-tuned large language models, particularly when combined with retrieval and multi-step generation, demand far more memory and latency, even when served with optimised inference stacks. The study&#8217;s conclusion is that model choice should be guided by source shift, evidence format, shortlist quality and inference cost rather than by one internal score alone. In other words, the right model is a function of the deployment environment: how stable the input distribution is, what kind of evidence can be retrieved, how reliable the shortlists are and what computational budget the application allows. For a healthcare chatbot serving thousands of users on commodity hardware, the calculus differs radically from a research system exploring rare conditions with expert supervision.</p>
<p>The broader significance of the work lies in its methodological discipline. By pooling multiple symptom-text sources and systematically withholding each one, the study offers a template for robust evaluation that the field of medical natural language processing urgently needs. Health condition identification sits at the intersection of information retrieval, machine learning and clinical practice, and erroneous automated predictions in this domain can misdirect patients in ways that have real consequences. Frameworks that expose fragility before deployment, rather than after, are therefore not merely academic exercises. The results also resonate with a wider lesson for the generative AI era: bigger and newer does not automatically mean better or more reliable, and the most impressive architecture on a single benchmark may be the least trustworthy when the ground shifts. For developers building the next generation of symptom checkers, diagnostic triage tools and patient-facing health assistants, the study provides both a caution and a roadmap, insisting that transfer robustness, retrieval quality and cost be weighed as carefully as headline accuracy before any model earns a place in a healthcare pipeline.</p>
<p><strong>Subject of Research:</strong> Robust benchmarking of language models for identifying health conditions from short symptom narratives</p>
<p><strong>Article Title:</strong> Robust evaluation of lexical, transformer, and retrieval-grounded language models for health condition identification</p>
<p><strong>Article References:</strong> Babac, M. B. (2026). Robust evaluation of lexical, transformer, and retrieval-grounded language models for health condition identification. <em>Journal of Intelligent Information Systems</em>. <a href="https://doi.org/10.1007/s10844-026-01092-1" rel="noopener noreferrer">https://doi.org/10.1007/s10844-026-01092-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10844-026-01092-1" rel="noopener noreferrer">10.1007/s10844-026-01092-1</a></p>
<p><strong>Keywords:</strong> health condition identification, symptom text classification, natural language processing, transformer models, large language models, retrieval-augmented generation, BlueBERT, transfer learning, machine learning, information retrieval, intelligent information systems, biomedical text mining</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203972</post-id>	</item>
		<item>
		<title>Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL</title>
		<link>https://scienmag.com/outcome-reward-models-improve-llm-based-text-to-sql-generation-with-gradesql/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 00:47:54 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[benchmarking and performance]]></category>
		<category><![CDATA[complex query translation]]></category>
		<category><![CDATA[database query verification]]></category>
		<category><![CDATA[democratized database access]]></category>
		<category><![CDATA[GradeSQL framework]]></category>
		<category><![CDATA[industry-standard benchmarks]]></category>
		<category><![CDATA[inference-time output validation]]></category>
		<category><![CDATA[intelligent information systems]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[machine-generated SQL accuracy]]></category>
		<category><![CDATA[machine-generated SQL correctness]]></category>
		<category><![CDATA[multi-table joins and nested queries]]></category>
		<category><![CDATA[natural language to SQL]]></category>
		<category><![CDATA[natural language to SQL translation]]></category>
		<category><![CDATA[Outcome Reward Models]]></category>
		<category><![CDATA[task-specific reward modeling]]></category>
		<category><![CDATA[Text-to-SQL generation]]></category>
		<guid isPermaLink="false">https://scienmag.com/outcome-reward-models-improve-llm-based-text-to-sql-generation-with-gradesql/</guid>

					<description><![CDATA[In an era when large language models are increasingly being asked to serve as intermediaries between everyday users and complex database systems, one of the most stubborn bottlenecks has been ensuring that the SQL queries these models generate are actually correct. Now, a team of researchers has unveiled GradeSQL, a framework that trains task-specific Outcome [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In an era when large language models are increasingly being asked to serve as intermediaries between everyday users and complex database systems, one of the most stubborn bottlenecks has been ensuring that the SQL queries these models generate are actually correct. Now, a team of researchers has unveiled GradeSQL, a framework that trains task-specific Outcome Reward Models (ORMs) to act as sophisticated judges of machine-generated database queries, consistently outperforming traditional verification methods on industry-standard benchmarks. The work, published in the Journal of Intelligent Information Systems, could reshape how intelligent information systems verify their own outputs at inference time.</p>
<p>The problem GradeSQL addresses is deceptively simple to state but notoriously difficult to solve. Text-to-SQL generation—the task of translating a natural language question like &#8220;Which hospitals in Boston treated more than 500 patients last year?&#8221; into an executable SQL query—has been a dream of computer scientists for more than five decades. The payoff is enormous: democratized database access for experts and non-experts alike, without requiring anyone to master query syntax. But while modern large language models have made impressive strides, they still stumble on queries involving multi-table joins, nested subqueries, and sophisticated aggregations. When a model generates a syntactically valid query that returns the wrong answer, the consequences in high-stakes database environments can range from merely confusing to genuinely harmful.</p>
<p>The dominant approach to filtering out bad queries has relied on what researchers call test-time inference strategies. Rather than retraining a model, these strategies generate multiple candidate outputs and then apply a heuristic to pick the best one. Two techniques have become standard: Best-of-N (BoN), which samples N candidate queries and selects the highest-scoring one according to some heuristic, and Majority Voting, which executes all candidates and chooses the output that appears most frequently. Prior empirical work on Text-to-SQL found that N=32 candidates represents an optimal trade-off between performance and computational cost. But both strategies share a fundamental weakness: they provide only coarse, discrete signals. Majority Voting depends on the frequency of execution results, while execution-based Best-of-N treats all executable queries that return non-empty result sets as equally plausible—discarding the rich semantic distinctions that separate a nearly correct query from a hopelessly wrong one.</p>
<p>GradeSQL&#8217;s central insight is borrowed from the reinforcement learning community, where Outcome Reward Models have proven their worth as verifiers of mathematical reasoning. Unlike Process Reward Models, which evaluate every intermediate reasoning step and require expensive step-level annotations, ORMs score only the final output. The concept was pioneered in multi-step mathematical reasoning, where a verifier assigns a probability of correctness to each candidate answer. But transplanting the idea into the Text-to-SQL domain demanded significant innovation, because a SQL verifier cannot simply check numeric equality. It must assess the semantic alignment between a natural language prompt, a structured database schema, and the resulting query execution—a far more nuanced judgment.</p>
<p>The GradeSQL framework unfolds in three carefully engineered stages. In the first stage, candidate generation, a generator LLM is prompted with a natural language question and its corresponding database schema to produce N candidate SQL queries using chain-of-thought reasoning, after which the reasoning trace is stripped away. In the second stage, data labeling, each candidate is executed against the target database and compared with the gold query&#8217;s result set. A candidate whose returned tuples exactly match the gold result set is labeled correct; one that executes successfully but returns a different result set is labeled a semantic mismatch; and one that crashes outright is discarded, since execution errors are trivially detectable. The crux of the framework lies in distinguishing semantically correct queries from those that are executable but wrong—a far subtler task.</p>
<p>The third stage is where the magic happens. The labeled candidates become training data for supervised fine-tuning, with the verification task cast as an autoregressive binary classification problem. The model receives an input sequence combining the database schema, the natural language question, and a candidate SQL query, and is trained to generate the token &#8220;Yes&#8221; if the query is semantically correct and &#8220;No&#8221; otherwise. Crucially, the researchers leverage the generator LLM&#8217;s latent ability to self-assess: they prompt the model to classify each candidate and extract the associated logits, which then serve as supervision signals during training. Fine-tuning is performed efficiently using Low-Rank Adaptation (LoRA) under a causal language modeling objective, minimizing the negative log-likelihood of the correct label. At inference time, the resulting ORM functions as a scoring function that maps any question-candidate pair to a probabilistic score between zero and one, inducing a ranked list from which the top-scoring candidate is selected.</p>
<p>To overcome the scarcity of reward-labeled data—a chronic obstacle that had stymied earlier attempts to apply ORMs to Text-to-SQL—the team built a scalable data synthesis pipeline that automates the generation and labeling of training candidates. This solved the cold-start problem that had left ORM-guided verification largely unexplored, and enabled the creation of task-specific reward models tailored to individual benchmark domains.</p>
<p>The evaluation was comprehensive, spanning multiple open-source LLM families and parameter sizes on the two most widely used Text-to-SQL benchmarks: BIRD and Spider. Spider, released in 2018, catalyzed the modern era of cross-domain schema generalization research, while BIRD poses real-world, Messy-data challenges that stress even state-of-the-art systems. The results were striking. ORM-guided Best-of-N achieved execution accuracy gains of up to 4.33 percentage points on BIRD and 2.10 points on Spider compared with execution-based Best-of-N, and gains of 2.91 points on BIRD and 0.93 points on Spider over Majority Voting. In a field where even minor gains represent substantial progress, these improvements are significant.</p>
<p>What makes ORMs particularly powerful is their ability to identify high-quality candidates that are underrepresented in the sample pool—a capability fundamentally absent from Majority Voting, which by design favors whatever output appears most often. ORMs can also recognize semantically equivalent queries that happen to differ syntactically, scoring two differently worded queries that return identical result sets as equally valid. This makes them a natural fit for re-ranking diverse candidate sets and for modular system architectures in which the generator and the verifier can be optimized independently. It also sidesteps the known failure mode of Best-of-N with coarse heuristics, which tends to over-prefer generic, high-probability outputs, and the vulnerability of Majority Voting when correct solutions are rare among the sampled candidates.</p>
<p>The research also situates itself within a broader trend of inference-time scaling, which allocates additional computational resources during inference rather than during training. The philosophy mirrors what has driven advances in mathematical reasoning: large language models can produce correct answers when given more inference time, rather than larger model sizes or additional training. Traditional decoding strategies, from greedy decoding and beam search to temperature scaling, top-k sampling, and nucleus sampling, operate only at the token level and cannot substantially improve the correctness of complete outputs, leaving systems vulnerable to hallucinations. Candidate-level verification with a trained reward model represents a fundamentally different lever—one that operates on the semantics of whole queries rather than the probabilities of individual tokens.</p>
<p>The team has released its codebase, synthesized datasets, and fine-tuned ORMs on GitHub and Hugging Face, an open approach intended to accelerate reproducibility and further research. The implications extend well beyond academic benchmarks. As enterprises increasingly deploy natural language interfaces over production databases—powering everything from customer support chatbots to internal analytics tools—the reliability of generated SQL becomes a first-order concern. A verification layer that provides continuous, probabilistic assessments of semantic quality, rather than binary pass-fail execution checks, offers a principled mechanism for deciding when to trust a machine-generated query and when to escalate to a human.</p>
<p>There remain open questions. The framework depends on the availability of gold queries for labeling during training, and its performance ceiling is presumably tied to the quality of the underlying generator&#8217;s self-assessment signals. The researchers also note that controlled ablations on prompt design, model scale, and training losses revealed nuances in how these factors interact, suggesting room for further optimization. Still, by demonstrating that task-specific Outcome Reward Models can consistently outperform both execution-based Best-of-N and Majority Voting, GradeSQL establishes a new baseline for test-time verification in Text-to-SQL—and points toward a future where language models not only write database queries, but grade their own work with genuine semantic understanding.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Outcome Reward Models for test-time verification in Text-to-SQL generation with large language models</p>
<p><strong>Article Title:</strong> GradeSQL: Outcome reward models for intelligent Text-to-SQL generation from LLMs</p>
<p><strong>Article References:</strong> Tritto, M., Farano, G., Di Palma, D., Rossiello, G., Subramanian, D., Narducci, F., &amp; Di Noia, T. (2026). GradeSQL: Outcome reward models for intelligent Text-to-SQL generation from LLMs. <em>Journal of Intelligent Information Systems</em>. <a href="https://doi.org/10.1007/s10844-026-01071-6" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s10844-026-01071-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10844-026-01071-6" target="_blank" rel="noopener noreferrer">10.1007/s10844-026-01071-6</a></p>
<p><strong>Keywords:</strong> Text-to-SQL, Large Language Models, Outcome Reward Models, test-time inference, Best-of-N, semantic verification, BIRD benchmark, Spider benchmark, LoRA fine-tuning, inference-time scaling</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">190494</post-id>	</item>
	</channel>
</rss>
