<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data sparsity &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/data-sparsity/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 25 Sep 2026 22:30:58 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>data sparsity &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Graph AI Meets Learning Automata to Crush Recommender System Cold Starts</title>
		<link>https://scienmag.com/graph-ai-meets-learning-automata-to-crush-recommender-system-cold-starts/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 22:30:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive decision-making in recommendation models]]></category>
		<category><![CDATA[addressing new user onboarding challenges]]></category>
		<category><![CDATA[Amazon dataset]]></category>
		<category><![CDATA[autoencoders]]></category>
		<category><![CDATA[cold-start problem]]></category>
		<category><![CDATA[collaborative filtering]]></category>
		<category><![CDATA[combating sparse rating data in recommender systems]]></category>
		<category><![CDATA[data sparsity]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning applications in recommender systems]]></category>
		<category><![CDATA[denoising autoencoders for personalized recommendations]]></category>
		<category><![CDATA[graph convolutional networks]]></category>
		<category><![CDATA[graph deep learning for recommendations]]></category>
		<category><![CDATA[hybrid recommendation]]></category>
		<category><![CDATA[hybrid recommender system architectures]]></category>
		<category><![CDATA[innovative approaches to improve user engagement in streaming and shopping platforms]]></category>
		<category><![CDATA[intelligent graph-based recommendation algorithms]]></category>
		<category><![CDATA[learning automata]]></category>
		<category><![CDATA[learning automata in machine learning]]></category>
		<category><![CDATA[machine learning techniques for cold start problem]]></category>
		<category><![CDATA[MovieLens]]></category>
		<category><![CDATA[Netflix Prize]]></category>
		<category><![CDATA[recommender system cold start problem]]></category>
		<category><![CDATA[recommender systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=214984</guid>

					<description><![CDATA[Researchers have built a hybrid recommender system that combines a deep denoising graph convolutional autoencoder, demographic side information, and learning automata to outperform state-of-the-art methods on four benchmark datasets while resisting cold-start and sparsity problems.]]></description>
										<content:encoded><![CDATA[<p>Every time a new user signs up for a streaming platform or opens a shopping app for the first time, the algorithms behind the scenes face one of the most stubborn problems in machine learning: they know almost nothing about this person. With no rating history to learn from, conventional recommenders stumble, often serving generic suggestions that frustrate users and cost businesses engagement. The same fragility appears when ratings are sparse, which is nearly always the case, since even the most active users touch only a sliver of a platform&#8217;s catalog. A new study published in the International Journal of Data Science and Analytics tackles these twin weaknesses head-on with a hybrid architecture that weaves together graph deep learning, denoising autoencoders, and an adaptive decision-making mechanism known as learning automata.</p>
<p>The system, called IGDHRS for intelligent graph-based deep hybrid recommender system, was developed by Milad Payandeh, Seyed Mahdi Jameii, and Mostafa Haghi Kashani of the Department of Computer Engineering at Islamic Azad University in Iran. Their starting point is a familiar taxonomy: recommender systems generally fall into collaborative filtering, which learns from patterns in user behavior; content-based filtering, which matches item attributes to user preferences; and hybrid models that blend the two. Each family carries its own liabilities. Collaborative approaches collapse when interaction data is thin or missing, while content-based methods struggle to capture the subtle, evolving tastes that ratings reveal. The Iranian team&#8217;s answer is a hybrid that treats the user population itself as a graph and then learns rich representations from that structure.</p>
<p>The first architectural step is the construction of a user–user similarity graph, in which nodes represent individual users and edges encode how alike their rating behaviors are. Building such a graph requires deciding, for every pair of users, whether their measured similarity is strong enough to justify a connection, and that decision hinges on a similarity threshold. Set the threshold too low and the graph becomes a dense tangle of weakly related users, diluting the signal; set it too high and the graph fragments, cutting off genuinely helpful neighborhood information. Rather than fixing this threshold by hand, the researchers let it adapt dynamically using learning automata, a class of reinforcement-driven stochastic decision units that adjust their actions based on feedback from the environment. In effect, the system tunes its own notion of who counts as a similar user as training proceeds.</p>
<p>Learning automata deserve a closer look because they are a departure from the gradient-based optimization that dominates modern deep learning. An automaton maintains a probability distribution over a set of possible actions, selects one, observes a reward or penalty, and updates its probabilities accordingly. Over many iterations it converges toward actions that consistently earn rewards. In IGDHRS, that feedback loop nudges the similarity thresholds toward values that ultimately improve recommendation quality, a form of automatic hyperparameter adaptation that relieves engineers of a delicate tuning burden and lets the graph topology evolve to match the data at hand.</p>
<p>To give the graph more expressive power, the authors enrich each user node with auxiliary demographic information, including age, gender, and occupation. This is a deliberate countermeasure to the cold-start problem: even a brand-new user with zero ratings carries demographic attributes that can anchor them in the similarity graph, linking them to established users with comparable profiles. Previous work has shown that demographic profile expansion can buffer sparse rating matrices, but integrating such side information directly into a graph neural architecture is what makes this system distinctive. The demographics act as a bridge across the rating desert, allowing information to propagate from well-modeled users to newcomers along graph edges that would not otherwise exist.</p>
<p>At the heart of the architecture sits the study&#8217;s central technical contribution: a deep denoising graph convolutional autoencoder, abbreviated DDGCAE. An autoencoder is a neural network trained to compress its input into a low-dimensional latent code and then reconstruct the original signal from that code, forcing it to learn the essential structure of the data. A denoising autoencoder raises the stakes by deliberately corrupting the input, for example by masking or perturbing entries, and requiring the network to recover the clean version, which cultivates robustness to the missing and noisy values that pervade real rating matrices. A graph convolutional autoencoder extends this idea to graph-structured data: graph convolution layers aggregate information from each node&#8217;s neighbors, so the learned embeddings encode not just a user&#8217;s own behavior but the behavior of the surrounding network neighborhood.</p>
<p>Combining all three ingredients means that DDGCAE operates on an enriched, adaptively thresholded user similarity graph, learns compressed representations by reconstructing denoised graph signals, and produces latent user profiles that capture both interaction patterns and demographic context. Those latent representations then drive rating prediction. The design echoes and extends a lineage of prior systems, from classic autoencoder-based collaborative filtering models such as AutoRec and collaborative denoising autoencoders to graph-based methods like Neural Graph Collaborative Filtering and LightGCN, but the authors argue that the joint interplay of denoising, graph convolution, and automata-driven graph construction is what sets their approach apart.</p>
<p>The team implemented IGDHRS in Python and evaluated it on four widely used benchmark datasets spanning different scales and domains: MovieLens 100K, MovieLens 1M, a stratified random sample of the Netflix Prize data, and the Amazon Movies and TV dataset. Because the Netflix and Amazon corpora are enormous, the researchers extracted stratified random samples, and they have made the sampling scripts and generated sample indices publicly available in a GitHub repository, alongside the full system implementation, a level of openness that supports reproducibility. Performance was measured with standard regression and ranking metrics: root mean squared error and mean absolute error to quantify how far predictions deviate from true ratings, and precision and recall to gauge the quality of the top recommendations actually surfaced to users.</p>
<p>The reported results are striking. Across all four datasets, the proposed system significantly outperformed several state-of-the-art comparison methods, and the advantages held in the conditions that matter most: domains with higher data sparsity and datasets where demographic information was partially missing. The authors attribute this robustness directly to the three-way combination of auxiliary user data, learning automata, and the DDGCAE architecture, and they specifically highlight improved resilience to cold-start situations, the scenario in which traditional collaborative filtering degrades most severely. Statistical rigor was addressed as well, with the study employing cross-validation practices and nonparametric significance testing in the tradition of the Wilcoxon ranking method to substantiate that observed gains were not artifacts of a lucky split.</p>
<p>For the broader field, the study suggests that the path past the cold-start and sparsity bottleneck may lie not in any single clever component but in architectures that let multiple adaptive mechanisms reinforce one another. Graph structures supply the relational scaffolding, denoising objectives harden the learned embeddings against missing data, demographic side channels keep new users connected from day one, and learning automata quietly optimize the structural choices that humans would otherwise guess. The code and data are public, the benchmarks are the community&#8217;s standards, and the message is clear: recommender systems that can rebuild their own wiring while learning from corrupted signals are a promising blueprint for the next generation of personalization engines, from streaming catalogs to e-commerce and beyond.</p>
<p><strong>Subject of Research:</strong> A deep graph-based hybrid recommender system addressing cold-start and data sparsity using learning automata</p>
<p><strong>Article Title:</strong> An intelligent recommender system based on deep denoising graph convolutional autoencoder and learning automata</p>
<p><strong>Article References:</strong> Payandeh, M., Jameii, S. M., &amp; Kashani, M. H. (2026). An intelligent recommender system based on deep denoising graph convolutional autoencoder and learning automata. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 312. <a href="https://doi.org/10.1007/s41060-026-01293-5" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01293-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01293-5" rel="noopener noreferrer">10.1007/s41060-026-01293-5</a></p>
<p><strong>Keywords:</strong> recommender systems, graph convolutional networks, autoencoders, learning automata, cold-start problem, data sparsity, collaborative filtering, deep learning, MovieLens, Netflix Prize, Amazon dataset, hybrid recommendation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">214984</post-id>	</item>
		<item>
		<title>Machine Learning Model Tames Massive Heterogeneous Music-Streaming Data While Slashing Energy Use</title>
		<link>https://scienmag.com/machine-learning-model-tames-massive-heterogeneous-music-streaming-data-while-slashing-energy-use/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 04:43:52 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[big data analytics for digital music services]]></category>
		<category><![CDATA[data sparsity]]></category>
		<category><![CDATA[energy-aware scheduling]]></category>
		<category><![CDATA[energy-efficient computing]]></category>
		<category><![CDATA[energy-efficient data processing]]></category>
		<category><![CDATA[heterogeneous music-traffic data]]></category>
		<category><![CDATA[heterogeneous music-traffic data modeling]]></category>
		<category><![CDATA[large-scale digital music platforms]]></category>
		<category><![CDATA[long short-term preferences]]></category>
		<category><![CDATA[long-term listener preference modeling]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for music recommendation]]></category>
		<category><![CDATA[music recommendation]]></category>
		<category><![CDATA[Music streaming data analysis]]></category>
		<category><![CDATA[non-negative matrix factorization]]></category>
		<category><![CDATA[optimizing data integration in music streaming]]></category>
		<category><![CDATA[real-world music-traffic data challenges]]></category>
		<category><![CDATA[recommender systems]]></category>
		<category><![CDATA[reducing energy consumption in AI models]]></category>
		<category><![CDATA[scalable recommender systems]]></category>
		<category><![CDATA[sparse and imbalanced listening data]]></category>
		<category><![CDATA[user preference modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193770</guid>

					<description><![CDATA[Researchers have developed a machine learning framework that models long and short-term music preferences, decomposes sparse heterogeneous traffic data, and schedules computation for energy-efficient recommendation at scale.]]></description>
										<content:encoded><![CDATA[<p>Music streaming platforms sit atop one of the largest and most chaotic data streams in the modern digital economy. Every skip, replay, search, playlist addition and listening session generates signals about what a listener wants, yet these signals arrive in wildly different formats, at different speeds, and with wildly different levels of reliability. A new study published in the Journal of Big Data tackles this problem head-on, presenting a machine learning framework that models listener preferences over both short and long time horizons while confronting two challenges that are often treated as afterthoughts in recommendation research: the sparse, imbalanced nature of real-world music-traffic data, and the mounting energy cost of processing it at scale.</p>
<p>The research, led by Ke Zhang of Henan Normal University together with Achyut Shankar of the University of Warwick, Sang-Bing Tsai of the International Engineering and Technology Institute in Hong Kong, and Wattana Viriyasitavat of Chulalongkorn University in Bangkok, addresses what the authors identify as the central obstacle facing contemporary recommender systems: not a shortage of data, but the difficulty of integrating large-scale heterogeneous music-traffic data in a way that maximizes its value. As information overload intensifies and data volumes grow, the sheer computational burden of extracting useful signals from the noise has become as important as the accuracy of the recommendations themselves.</p>
<p>At the heart of the new work is a music recommendation model built on long and short-term preference modeling. The underlying insight is intuitive: a listener&#8217;s taste is layered. Some patterns are stable for years, such as a durable affinity for jazz or a favorite era of rock, while others flicker in and out over days or weeks, driven by mood, season, or a single catchy song discovered on a commute. A system that treats all history equally risks drowning durable preferences in transient noise, while one that focuses only on recent behavior loses the deep context that makes long-term recommendations feel personal. The proposed model constructs a user preference model from historical music behaviors, explicitly separating these temporal layers so that recommendations can draw on both the enduring core and the volatile surface of a listener&#8217;s habits.</p>
<p>Technically, the modeling leans on machine learning techniques suited to sequential behavioral data, in line with the long short-term modeling tradition that underpins modern sequence-aware recommenders. By learning representations of user interactions that preserve temporal structure, the system can weigh the recency and the persistence of different signals. The authors report that ablation trials, in which components of the model are systematically removed to test their individual contributions, validate the effectiveness of this design, and that the resulting recommendation model outperforms benchmark methods in their experiments.</p>
<p>But accuracy alone is not the paper&#8217;s real ambition. Large-scale music-traffic data is plagued by imbalance and sparsity: a small fraction of extremely popular tracks attracts the overwhelming majority of interactions, while the long tail of the catalog is listened to so rarely that the user-item matrix is mostly empty. Traditional collaborative filtering approaches struggle in this regime, producing unreliable estimates for the sparse regions where novelty-seeking listeners actually live. To cope, the study proposes a two-stage decomposition method within non-negative matrix factorization, a technique that factorizes the large user-item interaction matrix into lower-dimensional non-negative components whose additive structure makes them interpretable as latent preferences and latent item attributes.</p>
<p>The two-stage strategy effectively breaks the hard problem into more tractable pieces. Instead of forcing a single factorization to explain both the dense, popularity-dominated region of the matrix and the sparse long tail simultaneously, the method decomposes the problem in stages, which the authors show mitigates the distortion that imbalance otherwise introduces. This matters commercially as much as scientifically: recommender systems that only amplify hits trap users in feedback loops, while systems that can model sparse interactions credibly can surface catalog depth, benefiting artists and listeners alike. The paper frames this as the key to alleviating the data sparsity problems that have long limited recommendation quality on massive heterogeneous platforms.</p>
<p>Perhaps the most distinctive contribution, however, is aimed at a problem that rarely appears in recommendation papers: energy consumption. As the scale of music-traffic data grows, so does the power draw of the server fleets that crunch it. Training and serving recommendation models over billions of interactions is an energy-intensive operation, and the authors argue that achieving low-power processing of these algorithms has become an urgent requirement for energy-efficient data analysis. In response, they design an energy-efficient scheduling strategy specifically for heterogeneous music-traffic workloads, orchestrating computational tasks so that the analytical pipeline consumes less power without sacrificing the quality of the resulting model.</p>
<p>The experimental results reported in the study support both halves of this dual objective. The proposed recommendation model performs better than the benchmarks against which it was tested, and the experiments also verify the effectiveness of the proposed algorithm on energy efficiency, suggesting that accuracy and sustainability need not be traded off against each other. For an industry in which streaming platforms operate some of the largest machine learning deployments in existence, the demonstration that scheduling-aware, energy-conscious design can coexist with improved recommendations is a notable datapoint in a broader conversation about the carbon footprint of artificial intelligence.</p>
<p>The work also reflects a wider shift in how big data research frames its problems. Rather than treating a recommender as an isolated algorithm, the authors treat it as a system embedded in a data pipeline with physical costs: heterogeneous inputs must be integrated, sparse signals must be strengthened, and every matrix operation has an electricity bill attached. Their framework, spanning preference modeling, two-stage matrix decomposition, and energy-aware scheduling, reads as an attempt to close that loop from raw traffic data all the way to sustainable serving. The article was received in October 2023, accepted in September 2026, and published as an open-access paper that is citable under a permanent DOI, with the authors declaring no competing interests and no specific funding support.</p>
<p>For listeners, the practical upshot is subtle but real: better long and short-term preference modeling means the next recommended track is more likely to feel like a genuine reflection of taste rather than an echo of the last three songs played. For operators, the message is louder. As catalogs and user bases expand, the bottleneck is shifting from model accuracy to data integration and energy economics, and methods like those proposed here, which attack sparsity, heterogeneity and power consumption in a single design, offer a template for building recommendation systems that can scale responsibly into an era of ever-bigger music-traffic data.</p>
<p>Non-negative matrix factorization has a long history in recommendation research precisely because of its interpretability. Unlike factorization methods that allow negative values, the non-negativity constraint means latent factors can only be added together, not subtracted, which encourages parts-based representations: a user&#8217;s profile becomes a weighted combination of coherent taste components rather than an abstract vector that resists human inspection. The two-stage decomposition proposed in this study builds on that foundation, and the reported ablation trials, a methodology in which individual components are removed one at a time to measure their contribution, offer a level of component-level accountability that single end-to-end accuracy comparisons often lack.</p>
<p>The emphasis on temporal preference modeling also connects to a broader lineage of sequence-aware recommendation. Recurrent architectures in the long short-term memory tradition were designed to preserve information over long input sequences while selectively forgetting irrelevant detail, a property that maps naturally onto listening behavior, where a single skipped track carries different weight than a track played to completion dozens of times. Treating short-term and long-term preferences as distinct modeling targets, rather than collapsing all history into one aggregate profile, reflects a growing consensus that recency and persistence encode different kinds of user intent.</p>
<p>The energy dimension of the work sits within a wider research conversation about the computational cost of machine learning at scale. Large recommendation deployments run continuously rather than in discrete training bursts, meaning that inference and data processing, not just model training, dominate lifetime energy use. Scheduling strategies that route heterogeneous workloads intelligently across computing resources can therefore yield savings that compound over millions of daily recommendation requests, which is why the authors frame low-power processing as an urgent requirement rather than an optimization afterthought.</p>
<p>It is also worth noting the publication trajectory of the paper itself. The manuscript was received in late 2023 and accepted nearly three years later, a timeline that reflects the extended peer review cycles common for work spanning multiple technical domains. It appears as an open-access article under a Creative Commons license that permits non-commercial sharing with attribution, and it is published as a citable, DOI-bearing version ahead of final editorial formatting, an increasingly common practice intended to accelerate access to accepted research.</p>
<p>The collaborative composition of the author team, spanning institutions in China, the United Kingdom, Hong Kong, and Thailand, mirrors the global character of the problem being studied. Music-traffic data crosses borders effortlessly, and the engineering challenges of integrating heterogeneous streams, correcting for sparsity, and constraining energy use are shared by platforms regardless of where their users live. Work that treats these as a single coupled design problem, rather than as separable concerns handed to different teams, offers a useful reference point for how large-scale data systems research may continue to evolve.</p>
<p><strong>Subject of Research:</strong> Machine learning-based analysis of large-scale heterogeneous music-traffic data for energy-efficient personalized music recommendation</p>
<p><strong>Article Title:</strong> ML-driven large-scale heterogeneous music-traffic data analysis</p>
<p><strong>Article References:</strong> Zhang, K., Shankar, A., Tsai, S.-B., &amp; Viriyasitavat, W. (2026). ML-driven large-scale heterogeneous music-traffic data analysis. <em>Journal of Big Data</em>. <a href="https://doi.org/10.1186/s40537-026-01557-8" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01557-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01557-8" rel="noopener noreferrer">10.1186/s40537-026-01557-8</a></p>
<p><strong>Keywords:</strong> music recommendation, machine learning, heterogeneous music-traffic data, non-negative matrix factorization, LSTM, user preference modeling, energy-efficient computing, data sparsity, big data, recommender systems, long short-term preferences, energy-aware scheduling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193770</post-id>	</item>
	</channel>
</rss>
