<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>missing data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/missing-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 04:14:44 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>missing data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Rising Sarcoma Surgery Rates in Germany May Be a Data Artifact, Researchers Warn</title>
		<link>https://scienmag.com/rising-sarcoma-surgery-rates-in-germany-may-be-a-data-artifact-researchers-warn/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:14:44 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[cancer registry]]></category>
		<category><![CDATA[challenges in rare cancer treatment]]></category>
		<category><![CDATA[clinical epidemiology]]></category>
		<category><![CDATA[evolution of surgical practice]]></category>
		<category><![CDATA[Germany]]></category>
		<category><![CDATA[Germany cancer treatment trends]]></category>
		<category><![CDATA[impact of clinical guidelines on surgery]]></category>
		<category><![CDATA[increase in multivisceral resection rates]]></category>
		<category><![CDATA[liposarcoma]]></category>
		<category><![CDATA[long-term registry data analysis]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[multivisceral resection]]></category>
		<category><![CDATA[national cancer registry analysis]]></category>
		<category><![CDATA[potential data artifacts in cancer registry]]></category>
		<category><![CDATA[procedure coding]]></category>
		<category><![CDATA[rare cancer]]></category>
		<category><![CDATA[rare cancers]]></category>
		<category><![CDATA[registry completeness]]></category>
		<category><![CDATA[retroperitoneal sarcoma]]></category>
		<category><![CDATA[surgical management of liposarcoma]]></category>
		<category><![CDATA[Surgical Oncology]]></category>
		<category><![CDATA[survival analysis]]></category>
		<category><![CDATA[tumor histological subtypes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225606</guid>

					<description><![CDATA[A new commentary argues that a reported surge in multivisceral resection for retroperitoneal sarcoma in Germany may partly reflect improving registry coding rather than a true shift in surgical practice.]]></description>
										<content:encoded><![CDATA[<p>A rare and formidable cancer has become the center of an unusually pointed scientific dispute. Retroperitoneal sarcoma, a tumor that grows in the deep space behind the abdominal organs, is so uncommon that most surgeons will encounter only a handful of cases in an entire career. Because no single hospital can accumulate meaningful experience with the disease, researchers increasingly turn to national cancer registries to answer questions that individual institutions cannot. A recent nationwide analysis from the German Cancer Registry Group, led by Beck and colleagues and published in the Journal of Cancer Research and Clinical Oncology, did exactly that: it examined treatment patterns for retroperitoneal sarcoma across Germany between 2000 and 2022, drawing on more than two decades of registry records to describe how surgical practice has evolved. The study reported a striking finding, namely that multivisceral resection, the removal of the sarcoma together with adjacent organs such as the kidney, spleen, pancreas, or colon, rose from roughly twelve percent of cases in 2001 to approximately fifty percent in recent years. That trajectory, the authors argued, reflects a genuine national shift toward guideline-concordant surgery for liposarcoma, the most common histological subtype of the disease.</p>
<p>Now a team of researchers from Meenakshi Medical College Hospital and Research Institute in Tamil Nadu, India, has published a formal commentary in the same journal challenging that interpretation. In a Matters Arising contribution, R. Saravanan, Paramanantham Madhavan, R. Sivayogana, and K. Anusha do not dispute the value of the underlying registry analysis, which they describe as an important framework for evaluating rare cancers using population-level data. Instead, they raise two methodological concerns that, if correct, could substantially alter how the German findings should be read. The first concerns whether the apparent rise in multivisceral resection reflects a real change in surgical practice or simply an improvement in how operations were documented and coded over time. The second concerns the handling of missing data on tumor grade and metastatic status, variables that were absent in a large fraction of cases yet still used to stratify survival analyses. Both concerns strike at the heart of a growing movement in oncology: the use of registry data to draw clinical conclusions about diseases too rare for randomized trials.</p>
<p>The argument about coding is technically subtle but conceptually compelling. The German analysis of organ resections relied on a subset of only 1,318 patients with codeable resection procedures, a considerably smaller group than the full surgical cohort, and it depended entirely on whether the Operationen- und Prozedurenschlüssel, the German procedure coding system, captured a code for every organ removed during each operation. The commentary points out that the original paper itself acknowledged a parallel phenomenon: the rise in annual case numbers over the same period was attributed to improving case ascertainment and documentation completeness following the 2014 legal framework for cancer registration, not to a true increase in disease incidence. If overall case documentation improved across two decades, the Indian authors argue, it is plausible and untested that procedure-level documentation improved as well. Under that scenario, earlier multivisceral resections may have been underdocumented, with only the single most prominent procedure code recorded, while later cases benefited from more complete coding that captured every organ removed in the same operation. Part or all of the observed temporal trend would then reflect better paperwork rather than bolder surgery.</p>
<p>The distinction is far more than methodological tidiness. If the trend is substantially a documentation artifact, the commentary notes, then the implicit message to clinicians and registry stewards, that German surgical practice has progressively converged on international multivisceral resection guidelines for liposarcoma, would be overstated. The more accurate takeaway would concern the maturation of registry coding rather than the maturation of surgical decision-making. The authors propose a concrete diagnostic test: reporting the proportion of surgical cases with at least one codeable resection procedure by diagnosis year, alongside the existing organ-count trend, would allow readers to judge whether the apparent rise in multivisceral resection tracks improving procedure documentation or represents an independent change in clinical practice. If the two curves move together, the documentation explanation gains weight; if they diverge, the practice-change interpretation becomes more credible. Such a check costs nothing beyond additional analysis of data already in hand, which is precisely why the commentary frames its absence as a missed opportunity.</p>
<p>The second major concern involves missing data in the survival analyses. According to the commentary, grading information was missing in twenty-four percent of cases, and distant metastasis status was missing in fifty-five percent, yet both variables were used to stratify the survival analyses and the primary treatment description without any sensitivity analysis addressing whether the missingness was informative. This matters because of a well-known statistical hazard: complete-case analysis, in which patients lacking a given variable are simply excluded, produces unbiased results only if the missing data occur randomly. If, as the original paper&#8217;s own discussion suggests, earlier-era cases suffered from less complete documentation, then those cases may be disproportionately represented among patients with missing grading or staging information. The Kaplan-Meier survival curves and treatment-pattern tables restricted to complete cases would then systematically overrepresent more recent, better-documented tumors, which may differ from the excluded cases in surgical approach, tumor biology ascertainment, or case mix. The resulting curves could paint a rosier or otherwise distorted picture of outcomes across the full registry cohort.</p>
<p>The commentary applies this concern specifically to the Kaplan-Meier estimates stratified by grade, which were calculated only among the roughly three-quarters of cases with recorded grading. No comparison of baseline characteristics between graded and ungraded cases was provided to establish that the two groups were otherwise similar. The authors suggest a remedy that mirrors their proposal on procedure coding: reporting whether missingness in grading and metastasis status varies systematically by diagnosis year, and providing a basic comparison of age, sex, and histology distribution between complete and incomplete cases. Such comparisons would allow readers to judge whether the survival and treatment-pattern findings generalize to the full registry cohort or are specific to the more thoroughly documented subset. The commentary also notes an internal inconsistency in the original paper&#8217;s logic: the study elsewhere treats improving documentation completeness as a plausible explanation for other temporal patterns in the dataset, yet does not extend that same reasoning to the grading and staging variables underlying the survival analysis.</p>
<p>What makes this exchange notable is its constructive tone and its broader relevance. The Indian authors explicitly state that neither concern displaces the substantial contribution of the German study, which established a reproducible, histology- and location-specific framework for extracting a rare tumor cohort from national registry data. They call that framework a genuinely useful methodological template for future rare-cancer registry research well beyond retroperitoneal sarcoma. The stakes they identify are correspondingly high: distinguishing genuine practice change from improved documentation completeness, and characterizing whether missing data correlate with diagnosis era, would allow clinicians and health-policy readers to rely on reported trends as evidence of evolving surgical practice with greater confidence, rather than treating documentation artifacts as clinical signals. In an era when health systems increasingly mine administrative and registry data to guide cancer care, the commentary is a reminder that the quality of the data pipeline can be as consequential as the quality of the medicine it records.</p>
<p>The technical issues at play are familiar to epidemiologists but deserve wider understanding. Procedure coding systems like the German OPS are designed for billing and administration, not research, and their granularity depends on local coding habits, staffing, and evolving regulations. Studies of complex surgical specimens in other fields have documented substantial variability in how thoroughly multi-organ procedures are captured in coded records, and the commentary cites work on procedural terminology optimization in genitourinary surgery as an illustration. Similarly, the problem of informative missingness in registry-based survival analysis has prompted a growing literature on imputation techniques for Kaplan-Meier estimation, which the commentary references. Neither problem is unique to the German study; both are endemic to registry research. The commentary&#8217;s contribution is to show how these general hazards map onto the specific claims of a high-profile national analysis, and to propose analyses that would resolve the ambiguity without requiring new data collection.</p>
<p>For patients with retroperitoneal sarcoma, the practical implications are indirect but real. Multivisceral resection is the single most important determinant of long-term survival in this disease, and international guidelines recommend aggressive compartmental surgery for liposarcoma whenever feasible. If German practice has genuinely converged on those guidelines over two decades, that is encouraging news for the roughly one in a hundred thousand people diagnosed each year. If instead the trend partly reflects coding maturation, the appropriate response is not complacency about surgery but investment in registry quality, so that future analyses can distinguish the two. The commentary, published open access in October 2026, ultimately asks for methodological consistency: the same skepticism about documentation quality that the original authors applied to case counts should be applied to procedure codes, tumor grades, and staging data. Whether the German group responds with the requested sensitivity analyses will determine how confidently the oncology community can cite this study as evidence that surgical practice, and not merely surgical record-keeping, has changed.</p>
<p><strong>Subject of Research:</strong> Methodological evaluation of German cancer registry data on retroperitoneal sarcoma treatment trends</p>
<p><strong>Article Title:</strong> Comment on “Treatment of retroperitoneal sarcoma in Germany between 2000 and 2022: a retrospective analysis from the German Cancer Registry Group”</p>
<p><strong>Article References:</strong> Comment on “Treatment of retroperitoneal sarcoma in Germany between 2000 and 2022: a retrospective analysis from the German Cancer Registry Group”. (n.d.). <a href="https://doi.org/10.1007/s00432-026-06591-w" rel="noopener noreferrer">https://doi.org/10.1007/s00432-026-06591-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00432-026-06591-w" rel="noopener noreferrer">10.1007/s00432-026-06591-w</a></p>
<p><strong>Keywords:</strong> retroperitoneal sarcoma, cancer registry, multivisceral resection, registry completeness, missing data, survival analysis, surgical oncology, liposarcoma, rare cancers, clinical epidemiology, procedure coding, Germany</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225606</post-id>	</item>
		<item>
		<title>New Mathematical Bridge Turns Incomplete Data into Reliable Concepts</title>
		<link>https://scienmag.com/new-mathematical-bridge-turns-incomplete-data-into-reliable-concepts/</link>
		
		<dc:creator><![CDATA[Reid Dalton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:00:18 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in knowledge extraction techniques]]></category>
		<category><![CDATA[application of formal concept analysis in real-world data]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[concept lattices]]></category>
		<category><![CDATA[concept-cognitive learning]]></category>
		<category><![CDATA[connections between mathematical concepts in incomplete data]]></category>
		<category><![CDATA[data incompleteness in formal concept analysis]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[data science and knowledge engineering]]></category>
		<category><![CDATA[derivation operators]]></category>
		<category><![CDATA[formal concept analysis]]></category>
		<category><![CDATA[formal concept lattice construction]]></category>
		<category><![CDATA[handling missing entries in datasets]]></category>
		<category><![CDATA[incomplete formal contexts]]></category>
		<category><![CDATA[knowledge acquisition]]></category>
		<category><![CDATA[knowledge extraction from missing information]]></category>
		<category><![CDATA[mathematical methods for incomplete data]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[partially-known formal concepts]]></category>
		<category><![CDATA[reliability in concept derivation]]></category>
		<category><![CDATA[systematic approach to incomplete data analysis]]></category>
		<category><![CDATA[three-way concept analysis]]></category>
		<category><![CDATA[uncertainty modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225510</guid>

					<description><![CDATA[A new Applied Intelligence study reveals how four types of mathematical concepts in incomplete data tables are structurally connected, enabling faster and more stable knowledge extraction from datasets with missing entries.]]></description>
										<content:encoded><![CDATA[<p>Real-world data is messy. Surveys go unanswered, sensors fail, medical records contain gaps, and databases accumulate missing entries the way old houses accumulate dust. For decades, mathematicians and computer scientists have wrestled with a deceptively simple question: when a table of objects and their attributes is incomplete, what can we still honestly say about the concepts hidden inside it? A new study published in Applied Intelligence by Xue Tian, Ruisi Ren, Ling Wei, and Yanhong She of Northwest University in Xi&#8217;an, China, offers one of the most systematic answers yet, showing how four different kinds of mathematical concepts arising from incomplete data are secretly connected to one another, and how those connections can be exploited to extract knowledge faster and more reliably than before.</p>
<p>The framework at the heart of this work is formal concept analysis, a mathematical theory introduced by Rudolf Wille in 1982 that has since become a cornerstone of knowledge engineering. In its classical form, formal concept analysis takes a formal context, essentially a table stating which objects possess which attributes, and derives from it a lattice of concepts. Each concept pairs a set of objects, called its extent, with the set of attributes they all share, called its intent. The resulting lattice is not just a list; it is a hierarchy, revealing how general and specific concepts nest inside one another. Biologists have used it to classify species, engineers to organize software libraries, and data miners to surface association rules in transaction data.</p>
<p>But the classical theory assumes the table is complete: every cell is a definite yes or no. The moment question marks appear, the elegant machinery stalls. A patient may or may not exhibit a symptom, a customer may or may not have bought a product, and the raw data simply does not say. Researchers have long responded by considering completions, hypothetical fully specified tables consistent with what is known. Among all possible completions, two stand out: the least completion, which fills every unknown with the most conservative answer possible, and the greatest completion, which fills them with the most generous one. Concepts in these two boundary tables bracket the truth, but computing them, and then computing their concept lattices, can be expensive, and the relationship between what they contain and what the incomplete table itself can reveal has remained murky.</p>
<p>That is precisely the gap the Xi&#8217;an team set out to close. Their starting point is a family of structures called partially-known formal concepts, introduced in earlier work to capture two intuitive modes of uncertainty: objects that jointly certainly possess a set of attributes, and objects that jointly possibly possess them. These ideas echo the possibility-theoretic reading of formal concept analysis developed by Didier Dubois, Florence Dupin de Saint-Cyr, and Henri Prade, and the interval-set perspective of Ying-Yu Yao, which treats an uncertain set as a pair of lower and upper bounds. Partially-known formal concepts come in three specific flavors, known as SE-ISI, ISE-SI, and ISE-ISI concepts, each defined through different combinations of upper and lower derivation operators that probe the incomplete context from different directions. Together with the classical formal concepts found in the least and greatest completions, they form four types of concepts, each illuminating a different facet of the same patchy data.</p>
<p>The first major contribution of the new paper is a thorough map of the connections among these four types. Because every one of them is ultimately defined by the same upper and lower derivation operators, the authors suspected, and then proved, that deep structural relationships must bind them. They show precisely how a partially-known concept relates to concepts in the least and greatest completions, establishing conditions under which one can be recovered from the other. This is more than an exercise in mathematical housekeeping. It means that knowledge extracted from a completed version of the data, which may be easier to compute in some settings, can be translated back into statements about the partially-known concepts of the raw incomplete table, and vice versa. The incomplete context, in other words, is not a impoverished cousin of a complete one; it carries enough structure to reconstruct much of what the completions would tell you, if you know how to read it.</p>
<p>The second contribution turns this theory into construction methods. The concept lattice of a partially-known structure, the hierarchical arrangement of all its concepts, can be built from the corresponding formal concept lattices of the completions, rather than from scratch. The authors also examine the relationships among the three kinds of partially-known concepts from two complementary viewpoints: the level of single concepts, and the level of entire collections of extents and intents. On both levels they demonstrate explicit methods for mutual conversion, showing how an SE-ISI concept can be transformed into an ISE-SI concept, and how the families of extents and intents of one type relate to those of another. These conversion procedures give practitioners a kind of Rosetta Stone: once any one of the four concept structures has been computed, the others become accessible through systematic translation rather than independent, redundant computation.</p>
<p>The practical payoff arrives in the third contribution: algorithms. Tian and colleagues designed algorithms that acquire partially-known formal concepts directly from already-computed formal concepts, leveraging the theoretical connections to avoid re-deriving everything from the incomplete table. They then stress-tested these algorithms experimentally, and crucially, they did so under different data missing mechanisms, the statistical regimes that statisticians distinguish as missing completely at random, missing at random, and more structured forms of absence. This matters because real datasets do not lose entries uniformly; missingness often correlates with the very attributes that matter most. An algorithm that is stable only under idealized random deletion would be of limited use in practice. The experiments, whose underlying datasets the authors have made available in a public repository, demonstrate that the proposed methods are both feasible and stable across these regimes, with a particularly striking result: when the incomplete contexts contain relatively few concepts, the algorithms exhibit superior performance.</p>
<p>That last finding may sound like a limitation, but it points to a sweet spot that is common in applied settings. Many real tables, especially those built from expert judgment or curated ontologies, are small in object and attribute counts even when riddled with unknowns. For such data, the new approach offers a way to squeeze every defensible drop of knowledge out of what is known, without pretending to know what is not. The work also connects to a broader research program on three-way concept analysis and concept-cognitive learning, in which the same Xi&#8217;an group and collaborators worldwide have been developing tools for reasoning under partial information, from attribute reduction in three-way concept lattices to dynamic updating of concepts as new data arrives. The new paper can be read as a unifying chapter in that program, tying together threads that had previously been studied in isolation.</p>
<p>For the wider field of artificial intelligence, the significance lies in the growing recognition that uncertainty is not noise to be cleaned away but structure to be modeled. Machine learning systems increasingly must operate on incomplete knowledge graphs, partially observed relational data, and sparse feature matrices, and formal concept analysis offers a symbolic, interpretable complement to statistical and neural approaches. By proving that the four concept types of an incomplete context form a coherent, inter-translatable system, and by delivering algorithms that exploit this coherence, the study gives knowledge engineers a principled toolkit for a problem that will only grow as data-hungry systems meet an imperfect world. The mathematics of missing entries, it turns out, has its own hidden lattice, and we are only beginning to climb it.</p>
<p><strong>Subject of Research:</strong> Connections among formal and partially-known formal concepts in incomplete formal contexts</p>
<p><strong>Article Title:</strong> A further understanding for four types of concepts in incomplete contexts based on their connections</p>
<p><strong>Article References:</strong> Tian, X., Ren, R., Wei, L., &amp; She, Y. (2026). A further understanding for four types of concepts in incomplete contexts based on their connections. <em>Applied Intelligence, 56</em>(15), Article 460. <a href="https://doi.org/10.1007/s10489-026-07311-0" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07311-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07311-0" rel="noopener noreferrer">10.1007/s10489-026-07311-0</a></p>
<p><strong>Keywords:</strong> formal concept analysis, incomplete formal contexts, partially-known formal concepts, concept lattices, three-way concept analysis, missing data, knowledge acquisition, derivation operators, concept-cognitive learning, data mining, applied intelligence, uncertainty modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225510</post-id>	</item>
		<item>
		<title>Borrowed Cases, Better Statistics: New Method Fights Missing Data by Pooling Outside Samples</title>
		<link>https://scienmag.com/borrowed-cases-better-statistics-new-method-fights-missing-data-by-pooling-outside-samples/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 19:21:03 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[auxiliary cases]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[combining independent research samples]]></category>
		<category><![CDATA[data integration]]></category>
		<category><![CDATA[data integration in behavioral research]]></category>
		<category><![CDATA[full information maximum likelihood]]></category>
		<category><![CDATA[handling missing questionnaire responses]]></category>
		<category><![CDATA[improving statistical power with external cases]]></category>
		<category><![CDATA[innovative approaches to missing data]]></category>
		<category><![CDATA[integrative data analysis]]></category>
		<category><![CDATA[longitudinal study data gaps]]></category>
		<category><![CDATA[measurement invariance]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[Missing data imputation]]></category>
		<category><![CDATA[multiple-group models]]></category>
		<category><![CDATA[novel techniques for missing data mitigation]]></category>
		<category><![CDATA[pooling external datasets]]></category>
		<category><![CDATA[sensor failure data recovery]]></category>
		<category><![CDATA[simulation study]]></category>
		<category><![CDATA[sleep research]]></category>
		<category><![CDATA[statistical methods for incomplete data]]></category>
		<category><![CDATA[statistical power]]></category>
		<category><![CDATA[structural equation modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218494</guid>

					<description><![CDATA[A new study in Behavior Research Methods shows that borrowing auxiliary cases from independent studies through multiple-group models can boost statistical power under missing data, but only if measurement invariance holds and between-dataset heterogeneity is properly modeled.]]></description>
										<content:encoded><![CDATA[<p>Missing data is the quiet saboteur of modern research. Whether a participant skips a questionnaire item, drops out of a longitudinal study, or a sensor fails mid-experiment, the gaps left behind can distort estimates, inflate uncertainty, and drain statistical power in ways that no amount of clever analysis can fully repair. Now, a new study published in Behavior Research Methods proposes an unexpected remedy: instead of treating data integration and missing data as separate problems, researchers can deliberately borrow cases from entirely independent studies to shore up their own incomplete datasets. The approach, developed by Po-Yi Chen of National Taiwan Normal University, reframes what it means to combine data across studies, turning integration itself into a weapon against missingness.</p>
<p>The core idea builds on a framework known as integrative data analysis, or IDA, which has traditionally been used to pool information from multiple studies in order to stabilize estimates and increase sample sizes. Past work on IDA has generally treated missing data as an obstacle to be overcome before datasets can be merged, a nuisance that complicates the harmonization of measures and models. Chen&#8217;s study flips that logic on its head. Rather than viewing integration as something that must survive missing data, the new research shows that integration can actively limit the damage that missing data inflicts on the analyses researchers actually care about.</p>
<p>The key concept is what Chen calls auxiliary cases. These are cases drawn from other independent samples or existing studies that are not part of the researcher&#8217;s target group of interest, but that carry sufficient data on the target variables to reduce the impact of missingness on the focal analyses. In other words, if a sleep researcher is studying the relationship between daytime sleepiness and cognitive performance in a clinical sample, but many participants are missing scores on some measures, cases from a completely different study that measured the same constructs can be brought in to help. These borrowed cases do not change the research question or contaminate the target population; they simply add statistical information that stabilizes the estimation process.</p>
<p>The mechanism for including auxiliary cases is a multiple-group model, a structural equation modeling technique in which cases from different sources are treated as belonging to different groups within a single analysis. The auxiliary cases form their own group, and the model imposes what are known as cross-group measurement invariance constraints on the measurement portion of the model. This means that the way latent constructs are measured, including factor loadings and related measurement parameters, is forced to be the same across the target group and the auxiliary group, ensuring that the borrowed cases are genuinely measuring the same things. Crucially, however, the focal structural parameters, such as the regression paths and structural relationships that constitute the actual research question, are allowed to be freely estimated for the target group. The auxiliary group thus contributes information about the measurement apparatus without dictating the substantive findings.</p>
<p>To test whether this strategy actually works, Chen conducted a series of simulation studies, the standard methodological tool for evaluating statistical techniques under known conditions. The simulations compared scenarios with complete data against scenarios with missing data, varying conditions that researchers commonly face. The results were striking in one direction: the benefits of including auxiliary cases were substantially larger under missing data conditions than under complete data conditions. When data were complete, adding auxiliary cases offered only modest gains, consistent with earlier findings that integrative data analysis can stabilize estimates. But when data were missing, the borrowed cases delivered meaningful improvements in both statistical power and the efficiency of the estimates, meaning narrower standard errors and a better chance of detecting true effects that would otherwise be drowned in uncertainty.</p>
<p>That asymmetry makes intuitive sense once the underlying statistics are considered. Modern missing data techniques, such as full information maximum likelihood estimation, extract as much information as possible from incomplete observations, but they cannot conjure information that simply is not there. Every missing value represents lost Fisher information, and the efficiency of parameter estimates degrades accordingly. Auxiliary cases inject fresh information about the joint distribution of the target variables, effectively replenishing what the missingness drained away. Under complete data, there is little to replenish, so the auxiliary cases have little to offer. Under missingness, they arrive precisely when the analysis is most starved of information.</p>
<p>Yet the study is equally clear about the limits of the approach, and its warnings are as important as its promises. If the cross-group invariance constraints imposed on the measurement model are improper, that is, if the measures do not actually function equivalently across the target and auxiliary samples, the inclusion of auxiliary cases can backfire, introducing additional bias into the estimates and even reducing statistical power. The problem is compounded when inter-dataset heterogeneities, genuine differences between the datasets in populations, procedures, or contexts, are not accounted for in the analyses. Under these conditions, the borrowed cases are not contributing neutral information about the measurement apparatus; they are smuggling in assumptions that are false, and the model&#8217;s estimates absorb the error.</p>
<p>Notably, these risks were found to be especially pronounced when full information maximum likelihood was used as the estimation method. FIML is one of the most widely recommended approaches for handling missing data in structural equation models, prized for its efficiency and its ability to use all available information without deleting cases. But the same property that makes FIML powerful, its reliance on the model&#8217;s assumptions to fill in the informational gaps, also makes it sensitive to violations. When invariance constraints are wrong or heterogeneity is ignored, FIML propagates those errors through the entire model, and the auxiliary cases become a liability rather than an asset. The lesson for practitioners is that the auxiliary-case strategy is not a plug-and-play fix; it demands careful verification that the measures are truly comparable across datasets and that between-study differences are explicitly modeled.</p>
<p>To demonstrate the approach on real data rather than simulated numbers, Chen provided an empirical example based on two independent experimental datasets drawn from the National Sleep Research Resource, a publicly accessible repository supported by the National Heart, Lung, and Blood Institute. The two datasets, the Apnea Positive Pressure Long-term Efficacy Study, known as APPLES, and the Best Apnea Interventions in Research study, or BestAIR, both included the Epworth Sleepiness Scale, a widely used measure of daytime sleepiness whose psychometric properties have been extensively validated. Treating one dataset as the target sample and the other as a source of auxiliary cases, the example illustrated how the multiple-group framework operates in practice, complete with the invariance testing steps needed to justify the cross-group constraints. The analysis code for both the simulations and the empirical example has been made available through the Open Science Framework, and the datasets themselves are freely available upon request from the repository, giving other researchers a concrete template to follow.</p>
<p>The broader significance of this work lies in how it repositions two familiar tools on the methodological landscape. Data integration has long been framed as a way to answer bigger questions by combining evidence, and missing data handling has been framed as a way to salvage what a flawed dataset can still tell us. Chen&#8217;s study shows that these framings are not merely parallel but intertwined: the act of integrating data can itself be a missing data intervention, and the choice of which outside cases to borrow becomes a design decision with direct consequences for power, bias, and efficiency. As open data repositories grow and secondary analysis becomes ever more central to the research enterprise, the pool of potential auxiliary cases will only deepen. The approach will not suit every study, since it requires overlapping measures, defensible invariance, and honest attention to heterogeneity across sources. But for researchers staring down a dataset riddled with gaps, the message is genuinely encouraging: the remedy for missing data may already exist in someone else&#8217;s study, waiting to be borrowed, constrained carefully, and put to work.</p>
<p><strong>Subject of Research:</strong> Using auxiliary cases from independent samples in multiple-group models to mitigate the impact of missing data on structural equation analyses</p>
<p><strong>Article Title:</strong> Including auxiliary cases to address missing data issues through multiple-group models</p>
<p><strong>Article References:</strong> Chen, P.-Y. (2026). Including auxiliary cases to address missing data issues through multiple-group models. <em>Behavior Research Methods, 58</em>(11), Article 304. <a href="https://doi.org/10.3758/s13428-026-03180-0" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03180-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03180-0" rel="noopener noreferrer">10.3758/s13428-026-03180-0</a></p>
<p><strong>Keywords:</strong> missing data, auxiliary cases, multiple-group models, integrative data analysis, measurement invariance, full information maximum likelihood, structural equation modeling, statistical power, data integration, simulation study, Behavior Research Methods, sleep research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218494</post-id>	</item>
		<item>
		<title>New Tensor Method Cleans Up Messy Multi-View Data in One Efficient Step</title>
		<link>https://scienmag.com/new-tensor-method-cleans-up-messy-multi-view-data-in-one-efficient-step/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:08:16 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced data mining methods]]></category>
		<category><![CDATA[alternating optimization]]></category>
		<category><![CDATA[cross-view consistency]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[efficient data clustering techniques]]></category>
		<category><![CDATA[handling sensor failures and privacy restrictions]]></category>
		<category><![CDATA[incomplete multi-view clustering]]></category>
		<category><![CDATA[incomplete multi-view data]]></category>
		<category><![CDATA[low-frequency nuclear norm]]></category>
		<category><![CDATA[low-rank approximation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[missing data handling in machine learning]]></category>
		<category><![CDATA[multi-channel dataset integration]]></category>
		<category><![CDATA[multi-view clustering]]></category>
		<category><![CDATA[multi-view data fusion]]></category>
		<category><![CDATA[multi-view data imputation]]></category>
		<category><![CDATA[one-step optimization]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[tensor learning]]></category>
		<category><![CDATA[tensor low-frequency learning]]></category>
		<category><![CDATA[tensor-based data analysis]]></category>
		<category><![CDATA[TLF-IMVC method]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217314</guid>

					<description><![CDATA[A new one-step tensor learning method uses a novel low-frequency nuclear norm to jointly filter noise and cluster incomplete multi-view data efficiently.]]></description>
										<content:encoded><![CDATA[<p>Modern datasets rarely arrive through a single channel. A video clip may be described by its pixels, its audio track, its subtitles, and the metadata surrounding it; a patient may be characterized by imaging scans, laboratory tests, and clinical notes. Each of these parallel descriptions is called a view, and combining them so that a machine learning algorithm can group similar examples together is the task of multi-view clustering. In the real world, however, this task is complicated by a stubborn inconvenience: some views are simply missing. Sensor failures, incomplete records, privacy restrictions, and expensive measurement procedures all conspire to leave gaps in the data, and the field that grapples with this problem is known as incomplete multi-view clustering. A newly published study in the journal Data Mining and Knowledge Discovery proposes a method called one-step Tensor Low-Frequency learning for Incomplete Multi-View Clustering, or TLF-IMVC, that addresses several long-standing weaknesses of existing approaches at once.</p>
<p>The work, authored by Lisha Zhao and Hongwei Ge of Jiangnan University and Shuzhi Su of Anhui University of Science and Technology, targets three problems that the authors identify as persistent obstacles in the literature. The first is computational cost: many tensor-based methods are powerful but slow. The second is a limited ability to characterize higher-order consistency, meaning the structured agreement that should exist across all views simultaneously rather than merely pairwise. The third, and perhaps most subtle, is the degradation that arises from two-step optimization frameworks, in which an algorithm first repairs or completes the missing data and then performs clustering on the result as if the two operations were independent. TLF-IMVC is designed to dissolve that separation, fusing the steps into a single joint optimization.</p>
<p>To understand why the two-step design is problematic, it helps to picture what happens when an algorithm fills in missing views before clustering. The imputation stage makes decisions based on incomplete information, and whatever errors it introduces are then frozen into the completed dataset. The clustering stage has no way to push back and correct those errors, because it only ever sees the finished product. Errors compound rather than cancel. By contrast, a one-step framework lets the clustering objective directly shape how missing information is recovered, so that the two processes inform each other continuously during optimization. The authors of the new study adopt this philosophy, joining the estimation of cluster structure and the exploitation of cross-view relationships inside a unified mathematical formulation.</p>
<p>The mathematical heart of the method is the tensor, a higher-dimensional generalization of the matrix. Where a matrix is a grid of numbers with rows and columns, a tensor can stack many such grids into a three-dimensional or higher-order object. In multi-view clustering, a natural construction is to take the spectral embeddings, the low-dimensional coordinate representations produced by spectral clustering, for each view and stack them across views to form a tensor. This stacked object encodes cross-view correlations as higher-order structure. Tensor-based methods have attracted considerable attention precisely because they can model these higher-order relationships, which pairwise matrix techniques inevitably flatten away. TLF-IMVC builds its core representation on exactly such stacked spectral embedding tensors.</p>
<p>What distinguishes the new approach is the particular structure it imposes on that tensor. The method jointly models two properties: low-frequency structure and low-rank structure. The low-rank assumption, familiar from decades of work on robust principal component analysis and related techniques, holds that the essential information in the data lives in a small number of dominant patterns, so that the tensor can be well approximated by one with far fewer degrees of freedom. The low-frequency assumption is newer in this context and more evocative. It treats the tensor as a signal that can be decomposed, in effect, into components of varying frequency, with genuine cluster-consistent information concentrated in the smooth, slowly varying low-frequency components while noise and artifacts accumulate in the high-frequency ones.</p>
<p>To make this idea operational, the authors introduce what they call a Tensor Low-Frequency nuclear norm. A nuclear norm is a standard device in low-rank optimization: it is a convex surrogate for the rank of a matrix or tensor, and minimizing it encourages solutions to be low-rank without requiring one to know the rank in advance. The novel norm proposed in this study adds a discriminative frequency-domain constraint, meaning it does not treat all components of the tensor equally. Instead, it selectively preserves the informative low-frequency components, which carry the consistent cross-view signal, and suppresses the noisy high-frequency components, which tend to encode corruption and missing-view artifacts. The result is a regularizer that simultaneously encourages low-rank structure and frequency-domain cleanliness, filtering the data representation as part of the optimization itself rather than as a separate preprocessing step.</p>
<p>Solving the resulting optimization problem is nontrivial, since it involves coupled variables, tensor operations, and non-smooth regularizers. The authors develop an efficient alternating optimization algorithm, a strategy in which the variables are updated in turn, each optimized while the others are held fixed, cycling until convergence. Alternating schemes of this kind are the workhorses of multi-view clustering research, and their theoretical behavior under nonconvex settings has been studied extensively in the optimization literature the paper builds upon. The practical payoff claimed for the new algorithm is efficiency: by operating on the stacked spectral embeddings and exploiting the joint one-step formulation, TLF-IMVC avoids the heavy cost of completing raw missing views while still enforcing consistency across all of them.</p>
<p>The empirical evaluation spans eight benchmark datasets, on which the authors report extensive experiments comparing TLF-IMVC against state-of-the-art incomplete multi-view clustering methods. According to the study, the results demonstrate the effectiveness and competitiveness of the proposed approach, supporting the central claims that joint one-step optimization, tensor-based higher-order modeling, and frequency-domain regularization combine into a method that outperforms alternatives that address these issues in isolation. The datasets used in the experiments are publicly available from standard repositories, and the analysis code and supplementary materials are available from the corresponding author upon reasonable request, a transparency measure that should make it straightforward for other researchers to verify and extend the results.</p>
<p>The significance of the work lies less in any single technical gadget than in the direction it points. Incomplete multi-view clustering has become a crowded subfield, with a steady stream of methods based on matrix completion, graph learning, contrastive prediction, anchor graphs, and deep generative networks, and tensor frameworks have featured prominently among them. What TLF-IMVC contributes is a specific answer to the question of what structure the fused representation should obey. By identifying low-frequency content as the signature of genuine cross-view consistency and encoding that intuition directly into a tensor nuclear norm, the method offers a principled filter that is learned jointly with the clustering rather than bolted on beforehand. If the frequency-domain perspective proves portable, it could inform the design of future methods well beyond the specific algorithm introduced here.</p>
<p>For practitioners, the practical appeal is the combination of accuracy and efficiency in a setting that is all too common. Data pipelines in medicine, multimedia analysis, and sensor networks routinely produce multi-view collections with missing entries, and methods that are either too slow or too brittle to handle the gaps end up shelved. A one-step framework that jointly handles recovery and clustering, grounded in a theoretically motivated norm, addresses both concerns. The study, published in volume 40 of Data Mining and Knowledge Discovery as article 108, arrived after peer review beginning in April 2026 and appearing in September 2026, and was supported by funding from the National Natural Science Foundation of China and provincial science foundations. As datasets grow richer and messier in parallel, techniques of this kind, which extract clean, consistent structure from noisy, incomplete observations, are likely to become an increasingly standard part of the machine learning toolkit.</p>
<p><strong>Subject of Research:</strong> One-step tensor low-frequency learning for clustering data with incomplete multiple views</p>
<p><strong>Article Title:</strong> One-step tensor low-frequency learning for incomplete multi-view clustering</p>
<p><strong>Article References:</strong> Zhao, L., Ge, H., &amp; Su, S. (2026). One-step tensor low-frequency learning for incomplete multi-view clustering. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 108. <a href="https://doi.org/10.1007/s10618-026-01263-2" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01263-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01263-2" rel="noopener noreferrer">10.1007/s10618-026-01263-2</a></p>
<p><strong>Keywords:</strong> incomplete multi-view clustering, tensor learning, low-frequency nuclear norm, spectral clustering, low-rank approximation, one-step optimization, cross-view consistency, data mining, machine learning, alternating optimization, missing data, unsupervised learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217314</post-id>	</item>
		<item>
		<title>Dueling Imputations: Deterministic Framework Sharpens AI Time-Series Forecasting</title>
		<link>https://scienmag.com/dueling-imputations-deterministic-framework-sharpens-ai-time-series-forecasting/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 02:45:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI forecasting]]></category>
		<category><![CDATA[AI robustness in incomplete data]]></category>
		<category><![CDATA[data imputation for predictive modeling]]></category>
		<category><![CDATA[data preprocessing]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deterministic data repair methods]]></category>
		<category><![CDATA[deterministic imputation]]></category>
		<category><![CDATA[evaluating imputation techniques]]></category>
		<category><![CDATA[forecasting performance-driven data filling]]></category>
		<category><![CDATA[forecasting reliability]]></category>
		<category><![CDATA[GRU]]></category>
		<category><![CDATA[innovative imputation frameworks]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[missing data handling in AI]]></category>
		<category><![CDATA[PM2.5]]></category>
		<category><![CDATA[RNN]]></category>
		<category><![CDATA[sensor data imputation]]></category>
		<category><![CDATA[sensor data reliability]]></category>
		<category><![CDATA[sensor network data gaps]]></category>
		<category><![CDATA[time-series continuity restoration]]></category>
		<category><![CDATA[time-series forecasting accuracy]]></category>
		<category><![CDATA[time-series imputation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200972</guid>

					<description><![CDATA[Researchers have developed a deterministic duel-based imputation framework that selects missing-data replacements by directly testing candidates against forecasting performance, cutting forecasting error by up to 70 percent.]]></description>
										<content:encoded><![CDATA[<p>Missing data is one of the oldest and most stubborn enemies of artificial intelligence. In sensor networks that monitor air quality, energy grids, hospitals, and industrial plants, readings vanish for all sorts of mundane reasons: a sensor fails, a transmission drops, a report arrives late, or extreme weather interrupts collection. For time-series forecasting systems, those gaps are more than an inconvenience. They break the temporal continuity on which predictive models depend, and the way engineers fill them can quietly determine whether a forecast is trustworthy or wildly off the mark. A new study published in Discover Informatics argues that the fix lies not in smarter models but in smarter, and strikingly disciplined, data repair.</p>
<p>Researchers led by Agung Bella Putra Utama, Aji Prasetya Wibawa, and Anik Nur Handayani of Universitas Negeri Malang, together with Andrew Nafalski of Adelaide University, have introduced a deterministic duel-based imputation framework that puts forecasting performance itself in charge of deciding how missing values should be filled. The central insight is deceptively simple: most imputation methods are judged by how statistically similar their reconstructions are to the original data, not by whether they actually help a forecasting model make better predictions. A filled-in series can look numerically plausible on paper and still wreck a downstream forecast by smoothing away the very peaks and rhythms the model needs to learn.</p>
<p>The framework works like a structured tournament. For every missing observation, the system generates a set of candidate values using familiar baseline techniques, including mean, median, and mode substitution, K-Nearest Neighbors, Multiple Imputation by Chained Equations, and Last-Observation-Carried-Forward. Each candidate is then scored with a composite loss function that combines three complementary forecasting metrics: Mean Absolute Percentage Error, which captures proportional error; Root Mean Square Error, which penalizes large absolute deviations; and the coefficient of determination, or R-squared, which measures how much of the variance in the real series the reconstruction explains. The weights are set at 0.4, 0.4, and 0.2 respectively, so error-based criteria dominate while explanatory power still plays a supporting role.</p>
<p>What separates this approach from conventional candidate ranking is the way winners are chosen. Instead of collapsing every candidate into a single aggregated score and picking the top one, the framework stages deterministic pairwise duels. Candidates face each other one-on-one, and a sensitivity threshold of 0.001 prevents negligible score differences from flipping outcomes. After all comparisons, candidates are ranked by cumulative wins, and the top three advance to a final aggregation stage. Four aggregation modes are available: averaging the survivors for low-variance stability, taking the median for outlier robustness, taking the maximum to emphasize peak responsiveness, or a winner-take-all selection of the candidate with the lowest forecasting loss. Because every step follows fixed rules, the same input always produces the same output, a property the authors argue is essential for auditing, validation, and accountability in operational systems.</p>
<p>The team evaluated the framework on the Beijing PM2.5 dataset, a widely used benchmark of hourly meteorological and air-pollution measurements collected between 2010 and 2014. After cleaning and temporal alignment, 41,776 valid records remained out of 43,800, with roughly 2,068 missing values concentrated in the PM2.5 variable itself. Crucially, the missingness is not random: gaps cluster in cold seasons and during periods of rapid atmospheric change, exactly the conditions where forecasting matters most and where naive imputation does the most damage. The dataset includes dew point, temperature, pressure, cumulative wind speed, snowfall, rainfall, and temporal indicators as explanatory variables, with PM2.5 concentration serving as the prediction target.</p>
<p>The imputed datasets were then fed into four recurrent forecasting architectures trained under identical conditions: a vanilla Recurrent Neural Network, a Long Short-Term Memory network, a Bidirectional LSTM, and a Gated Recurrent Unit. Hyperparameters for all models were tuned with Particle Swarm Optimization, though the researchers are careful to note that this stochastic tuning happens entirely outside the imputation stage and does not compromise its determinism. The optimized configurations converged on surprisingly lightweight designs: three hidden layers of 24 neurons each, sigmoid activations, the Adam optimizer with mean squared error loss, a batch size of 32, 46 training epochs, and a dropout rate of 0.2.</p>
<p>The results are striking. Compared with conventional imputation methods, the duel-based framework reduced MAPE by up to 70 percent and RMSE by 12 to 15 percent, while maintaining R-squared values above 0.95 across all four architectures. The Mean aggregation variant delivered the strongest overall performance, preserving both proportional variation and the amplitude structure of the signal. Dropping missing observations entirely, by contrast, produced the worst forecasts, confirming that simply discarding incomplete rows destroys the temporal coherence recurrent models rely on. Even BRITS, a sophisticated deep-learning imputation method, showed less consistent gains, suggesting that reconstruction quality alone does not guarantee forecasting quality.</p>
<p>Statistical testing reinforced the case. Paired t-tests and Wilcoxon signed-rank tests comparing the proposed Mean variant against KNN, the strongest conventional baseline, produced p-values below 0.01 across all architectures, with each experiment repeated ten times under fixed random seeds. Visual analysis of predicted versus observed PM2.5 trajectories showed close alignment across smooth and volatile segments alike, with the Bi-LSTM achieving the highest R-squared values and the GRU delivering the most efficient runtime while matching LSTM accuracy. The framework&#8217;s computational overhead scales quadratically with the number of candidates but linearly with data size, and because the candidate set is small and fixed, runtimes remain predictable even as data volume grows.</p>
<p>To test generalizability, the researchers extended the evaluation beyond air quality to three additional domains: the Heart Disease dataset, which has weak temporal structure; the KEDS e-journal dataset, which shows irregular behavioral patterns and moderate sparsity; and the Sunspot dataset, which exhibits strong periodic behavior. The framework achieved lower MAPE and RMSE than baselines including drop-missing, mean imputation, KNN, and BRITS across all of them, with the largest gains appearing in datasets with strong sequential patterns such as Sunspot. The authors acknowledge limitations, including the quadratic cost of pairwise comparison for large candidate pools and the fixed forecasting horizons tested, and they point to cluster-based candidate reduction and online learning as future directions. But the broader message stands: imputation should not be a passive preprocessing chore. Treated as a forecasting-driven decision process, it becomes a strategic lever for building AI systems that are not only accurate but consistent, transparent, and worthy of trust.</p>
<p><strong>Subject of Research:</strong> A deterministic duel-based imputation framework that integrates forecasting performance metrics into missing-data selection for reliable AI-driven time-series forecasting.</p>
<p><strong>Article Title:</strong> A deterministic duel-based imputation framework for reliable AI-driven time-series forecasting</p>
<p><strong>Article References:</strong> Utama, A. B. P., Wibawa, A. P., Handayani, A. N., &amp; Nafalski, A. (2026). A deterministic duel-based imputation framework for reliable AI-driven time-series forecasting. <em>Discover Informatics, 1</em>(1), Article 3. <a href="https://doi.org/10.1007/s44564-026-00001-6" rel="noopener noreferrer">https://doi.org/10.1007/s44564-026-00001-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44564-026-00001-6" rel="noopener noreferrer">10.1007/s44564-026-00001-6</a></p>
<p><strong>Keywords:</strong> time-series imputation, missing data, AI forecasting, deterministic imputation, deep learning, LSTM, GRU, RNN, PM2.5, forecasting reliability, machine learning, data preprocessing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200972</post-id>	</item>
	</channel>
</rss>
