<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>COVID-19 spread modeling &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/covid-19-spread-modeling/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 00:04:57 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>COVID-19 spread modeling &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Machine Learning Reconstructs Missing Infection Links in COVID-19 Networks</title>
		<link>https://scienmag.com/machine-learning-reconstructs-missing-infection-links-in-covid-19-networks/</link>
		
		<dc:creator><![CDATA[Kristina Jarvis]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 00:04:57 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[contact tracing]]></category>
		<category><![CDATA[contact tracing in infectious disease outbreaks]]></category>
		<category><![CDATA[COVID-19]]></category>
		<category><![CDATA[COVID-19 infection networks]]></category>
		<category><![CDATA[COVID-19 spread modeling]]></category>
		<category><![CDATA[Cyprus]]></category>
		<category><![CDATA[Cyprus COVID-19 case studies]]></category>
		<category><![CDATA[data-driven pandemic surveillance]]></category>
		<category><![CDATA[epidemiological attribute analysis]]></category>
		<category><![CDATA[epidemiology]]></category>
		<category><![CDATA[gradient boosting]]></category>
		<category><![CDATA[graph representation learning]]></category>
		<category><![CDATA[GraphSAGE]]></category>
		<category><![CDATA[infection source identification]]></category>
		<category><![CDATA[link prediction]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning applications in public health]]></category>
		<category><![CDATA[machine learning for epidemiological data]]></category>
		<category><![CDATA[missing links in COVID-19 transmission]]></category>
		<category><![CDATA[network reconstruction]]></category>
		<category><![CDATA[network reconstruction in epidemiology]]></category>
		<category><![CDATA[Public health]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[reconstructing transmission chains using AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250661</guid>

					<description><![CDATA[Researchers in Cyprus used classical machine learning and GraphSAGE-based graph representation learning to predict missing transmission links in COVID-19 infection networks, reconstructing more connected and epidemiologically plausible spread structures across the first four pandemic waves.]]></description>
										<content:encoded><![CDATA[<p>When a new infectious disease sweeps through a population, epidemiologists face a problem that is as much about information as it is about biology: the map of who infected whom is never complete. Contact tracing teams work under intense time pressure, interviewing cases, reconstructing movements, and piecing together chains of transmission, yet a substantial fraction of infections inevitably ends up recorded as isolated cases with no known source. During the COVID-19 pandemic this gap was particularly visible, with thousands of positive tests that could not be tied to any documented epidemiological link. A new study published in PLOS Complex Systems by Pavlos Alexandros Dimitriou, Valentinos Silvestros, Elisavet Constantinou, Costas Pitris, and Panayiotis Kolios tackles this reconstruction problem directly, asking whether machine learning methods can infer the missing connections in a real national infection network and thereby restore a more faithful picture of how the virus actually spread.</p>
<p>The researchers focused on Cyprus, a country that maintained detailed surveillance data across the first four waves of the COVID-19 pandemic. Each confirmed case in their dataset carried a set of epidemiological attributes, including age, gender, date of infection, residency, and a classification of the case according to NACE, the European statistical classification of economic activities that encodes occupational context. Cases that had been linked by contact tracing formed the observed edges of an infection network, while the unlinked cases represented the missing structure the team hoped to recover. Framing the problem this way turns source attribution into a link prediction task: given two cases, can a model estimate how likely it is that one infected the other, even when no tracer ever documented that connection?</p>
<p>The first of the two approaches evaluated in the study relied on classical machine learning classifiers trained on engineered edge-level features. For every candidate pair of cases, the researchers combined the node-level attributes of the two individuals using a variety of operators, producing a feature vector that described the pair rather than either case alone. This design allowed the models to learn patterns such as whether infections were more likely between people of similar ages, overlapping infection dates, the same region of residence, or compatible occupational categories. Because the raw feature space constructed from all these combinations was large, the team applied Minimum Redundancy Maximum Relevance, or mRMR, to select features that were strongly associated with true transmission links while avoiding redundant information, and then refined the selection with greedy forward feature selection. Permutation importance was subsequently used to quantify how much each retained feature contributed to the classifiers&#8217; decisions.</p>
<p>Among the classifiers tested, Random Forest and Gradient Boosting consistently delivered the strongest performance. During the first two pandemic waves, these models achieved F1-scores ranging from approximately 0.82 to 0.90, indicating a strong balance between precision and recall in distinguishing plausible transmission pairs from implausible ones. Mean Reciprocal Rank values, which measure how highly the true infector appears in a model&#8217;s ranked list of candidate sources, ranged from 0.55 to 0.64 for those early waves. In practical terms, this means that for a substantial share of unlinked cases, the correct or nearly correct infector sat at or near the top of the model&#8217;s shortlist, exactly the kind of ranking an epidemiologist would need to prioritize follow-up investigations.</p>
<p>Performance declined in the later waves, and the decline is itself informative. For the third and fourth waves, F1-scores ranged from 0.68 to 0.75 and MRR values dropped to between 0.23 and 0.33. The authors&#8217; results suggest that the structure of transmission changed as the pandemic evolved: later waves in Cyprus were shaped by widespread community transmission, variants of concern, and vaccination, all of which blur the simple epidemiological signatures, such as tight date proximity or shared workplace, that made early-wave links easier to predict. When many candidate infectors share similar attributes with a case, the ranking task becomes fundamentally harder, and the numbers reflect that added ambiguity rather than any weakness in the method itself.</p>
<p>The second approach moved beyond hand-engineered features and turned to graph representation learning. The team employed a GraphSAGE model, a neural architecture that generates node embeddings by aggregating information from a node&#8217;s attributes and from the topology of the observed network around it. Each case is thus represented as a dense vector that encodes both what the case looks like epidemiologically and where it sits in the web of known connections. To score a candidate link between two cases, the researchers evaluated several strategies for combining the two embeddings, including concatenation, absolute difference, squared difference, the Hadamard product, and the dot product, each of which captures a different notion of similarity or compatibility between the pair.</p>
<p>Link prediction with graph representation learning achieved F1-scores ranging from roughly 0.70 to 0.79 and MRR values from 0.23 to 0.42 across all four pandemic waves. Compared with the classical classifiers, the GraphSAGE approach was somewhat less accurate in the early waves, where the engineered features captured the epidemiological signal very effectively, but it offered a more uniform level of performance across the entire pandemic period. This consistency matters for public health practice, because a tool that degrades sharply between epidemic phases is harder to integrate into routine surveillance workflows than one whose reliability is more stable. The embedding-based method also has the advantage of learning representations directly from the network, potentially capturing relational patterns that a fixed feature-engineering pipeline would miss.</p>
<p>The most consequential step of the study came after validation. Having confirmed that the best-performing classifier could reliably recover known links held out from the training data, the researchers applied it to the genuinely unlinked cases in the Cypriot dataset, asking the model to infer the most likely infector for each isolated node. The resulting reconstructed networks showed markedly fewer isolated components and larger connected structures than the original traced networks. Importantly, the reconstruction preserved epidemiologically relevant characteristics of the network, most notably the outdegree distribution, which reflects how many secondary infections are attributed to each case. Maintaining this property indicates that the inferred links did not distort the underlying transmission dynamics, such as the presence of superspreading individuals, but instead extended the observed network in a plausible direction.</p>
<p>The implications of this work extend well beyond one country or one pathogen. Incomplete infection networks are a chronic feature of outbreak response, from foodborne disease investigations to emerging zoonotic threats, and the quality of public health decisions depends heavily on how accurately those networks reflect reality. A reconstructed network with fewer isolated cases gives epidemiologists a more complete picture of transmission chains, enabling more targeted interventions: contact investigations can be prioritized toward the candidate sources ranked highest by the model, and containment resources can be directed at the connection points that matter most for interrupting spread. The authors note that such tools are particularly valuable during small epidemic waves, when case numbers are limited and every unexplained infection carries a proportionally larger informational cost.</p>
<p>The study also illustrates a broader lesson in applied machine learning for public health: classical feature-based models and modern graph neural networks are complementary rather than competing tools. Engineered epidemiological features, when chosen carefully with methods like mRMR and validated with permutation importance, can deliver excellent performance when the transmission signal is strong and interpretable. Graph representation learning, by contrast, offers robustness across changing epidemic conditions and the flexibility to incorporate network structure directly. As surveillance systems grow richer and more digitized, hybrid pipelines that combine both strategies, validated rigorously on known links before being trusted with unknown ones, are likely to become a standard component of the computational epidemiology toolkit, turning the frustrating gaps in contact tracing data from dead ends into testable hypotheses about the hidden architecture of disease spread.</p>
<p><strong>Subject of Research:</strong> Link prediction in COVID-19 infection networks using machine learning and graph representation learning</p>
<p><strong>Article Title:</strong> Predicting missing links in COVID-19 infection networks</p>
<p><strong>Article References:</strong> Dimitriou, P. A., Silvestros, V., Constantinou, E., Pitris, C., &amp; Kolios, P. (2026). Predicting missing links in COVID-19 infection networks. <em>PLOS Complex Systems, 3</em>(8), e0000124. <a href="https://doi.org/10.1371/journal.pcsy.0000124" rel="noopener noreferrer">https://doi.org/10.1371/journal.pcsy.0000124</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pcsy.0000124" rel="noopener noreferrer">10.1371/journal.pcsy.0000124</a></p>
<p><strong>Keywords:</strong> COVID-19, link prediction, machine learning, graph representation learning, GraphSAGE, contact tracing, epidemiology, Cyprus, Random Forest, Gradient Boosting, network reconstruction, public health</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250661</post-id>	</item>
	</channel>
</rss>
