<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Q-learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/q-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 21:15:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Q-learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Q-Learning Meets RPL: New Routing Protocol Keeps Mobile IoT Networks Fast and Efficient</title>
		<link>https://scienmag.com/q-learning-meets-rpl-new-routing-protocol-keeps-mobile-iot-networks-fast-and-efficient/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 21:15:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive routing for mobile sensors]]></category>
		<category><![CDATA[Contiki OS]]></category>
		<category><![CDATA[Cooja simulator]]></category>
		<category><![CDATA[distributed routing algorithms]]></category>
		<category><![CDATA[dynamic IoT network management]]></category>
		<category><![CDATA[energy efficiency]]></category>
		<category><![CDATA[Internet of Mobile Things]]></category>
		<category><![CDATA[IoT]]></category>
		<category><![CDATA[IoT data collection challenges]]></category>
		<category><![CDATA[IoT routing protocols]]></category>
		<category><![CDATA[low-power IoT device communication]]></category>
		<category><![CDATA[low-power networks]]></category>
		<category><![CDATA[mobile IoT networks]]></category>
		<category><![CDATA[mobility support]]></category>
		<category><![CDATA[MSQ-RPL protocol]]></category>
		<category><![CDATA[multi-sink networks]]></category>
		<category><![CDATA[multipath forwarding]]></category>
		<category><![CDATA[Q-learning]]></category>
		<category><![CDATA[reinforcement learning in networking]]></category>
		<category><![CDATA[routing protocol]]></category>
		<category><![CDATA[RPL]]></category>
		<category><![CDATA[RPL protocol limitations]]></category>
		<category><![CDATA[scalable routing schemes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229075</guid>

					<description><![CDATA[Researchers have developed MSQ-RPL, a distributed routing protocol that uses Q-learning-inspired reward propagation to keep mobile, multi-sink Internet of Things networks reliable and energy efficient.]]></description>
										<content:encoded><![CDATA[<p>The Internet of Mobile Things has quietly become one of the most demanding environments in modern networking. Billions of small, battery-powered sensors now move through hospitals, factories, smart cities, and agricultural fields, streaming data toward collection points known as sinks. When those devices move and when there are many sinks at once, the routing protocols that hold these networks together begin to strain. A new study published in Cluster Computing by Mahmoud Alilou, Amin Babazadeh Sangar, Kambiz Majidzadeh, and Mohammad Masdari of Islamic Azad University in Urmia, Iran, tackles this problem head-on with a protocol called MSQ-RPL, a distributed and scalable routing scheme that borrows its decision-making logic from reinforcement learning.</p>
<p>The starting point for the research is a well-known weakness in the standard routing stack for low-power networks. The IPv6 Routing Protocol for Low-Power and Lossy Networks, or RPL, was standardized by the Internet Engineering Task Force in 2012 and remains the dominant routing protocol for constrained IoT devices. RPL builds a logical tree, called a Destination Oriented Directed Acyclic Graph, rooted at a sink, and every node forwards packets upward toward that root. The design works beautifully in static deployments with a single collection point. But when devices move, the tree constantly breaks. Nodes drift out of range of their preferred parents, routes fail, and the protocol responds by flooding the network with control messages to rebuild its topology. The result is a cascade of packet losses, rising latency, and wasted energy, precisely the resources that battery-operated sensors cannot afford to squander.</p>
<p>Multi-sink deployments compound the difficulty. In principle, distributing several sinks across a network should relieve congestion, shorten transmission paths, and balance the energy load, since each node can hand its traffic to the nearest collection point. In practice, classical RPL was never designed to exploit that freedom. Its centralized structure and relatively static topology assumptions mean that a mobile node with multiple sinks available still tends to lock onto one parent and cling to it until the link degrades. The Iranian team&#8217;s answer is to let every node learn, in a distributed fashion, which forwarding choice currently earns the best reward, and to keep those learned scores fresh as the network churns around it.</p>
<p>The core mechanism of MSQ-RPL is a reward-propagation scheme inspired by Q-learning, the classic reinforcement learning algorithm introduced by Richard Sutton and Andrew Barto&#8217;s framework of trial-and-error value estimation. In a Q-learning system, an agent learns the long-term value of taking a particular action in a particular state by updating a score, or Q-value, each time it observes a reward. MSQ-RPL transplants this idea into the routing layer: each node maintains adaptive scores for candidate next hops, and those scores are refreshed using feedback that reflects real-time link quality and the residual energy of neighboring devices. A neighbor that reliably delivers packets while conserving its battery accumulates a high score; a neighbor with a flaky link or a depleted power reserve sees its score decay. Forwarding decisions then follow the reward landscape rather than a rigid tree structure, which means the protocol can reroute around failures before they become catastrophic.</p>
<p>Crucially, the researchers paired this learning mechanism with lightweight control dissemination so that the score updates themselves do not choke the network. The protocol employs scope-aware control distribution, meaning control information is propagated only as far as it is useful, rather than broadcast network-wide. This is combined with decentralized path optimization, so each node improves its own local view of the network without waiting for a central authority to issue routing tables. The design is also clustering-friendly: because the scoring system is built on real-time link and energy metrics, nodes can naturally organize into clusters around strong, energy-rich forwarding nodes, a behavior that meshes well with the clustered architectures common in large-scale IoT deployments.</p>
<p>Two further features target the realities of mobility and scale. MSQ-RPL supports sink-agnostic multipath forwarding, which allows a node to maintain several viable routes toward any sink rather than committing to a single parent. If one path collapses as the node or an intermediate device moves, traffic can shift to an alternative without triggering an expensive rediscovery process. Alongside this, the protocol performs dynamic score diffusion, continuously spreading updated reward information through the neighborhood so that forwarding choices track the current topology rather than a snapshot taken minutes ago. Together, these mechanisms are intended to sustain reliable performance under the kind of topological variation that would cripple a conventional RPL instance.</p>
<p>To find out whether the theory holds up, the team implemented MSQ-RPL in Contiki OS, the widely used open-source operating system for constrained IoT devices, and evaluated it with the Cooja network simulator. The benchmarks pitted the new protocol against standard RPL and against several recent multi-sink and mobility-aware routing schemes, including mRPL, multi-REBTAM, multi-EKF-MRPL, and QCM2R. The evaluation focused on the four metrics that matter most in low-power networks: packet delivery ratio, end-to-end delay, energy consumption, and control overhead, the last of which measures how much of the network&#8217;s scarce bandwidth is consumed by routing bookkeeping rather than useful data.</p>
<p>The reported results show that in most cases MSQ-RPL delivers comparable or improved performance across all four metrics relative to its competitors. The reinforcement-learning-driven scoring appears to pay its largest dividends under mobility, where the ability to switch paths quickly and locally prevents the route failures that plague tree-based protocols. The scope-aware dissemination of control messages keeps overhead in check even as the network grows, supporting the authors&#8217; claim that the protocol scales to realistic distributed IoMT infrastructures with multiple mobile sinks. The researchers also note that the protocol&#8217;s stability across varying conditions confirms its practical applicability, an important point because laboratory-friendly protocols often falter when confronted with the messy, unpredictable link quality of real wireless environments.</p>
<p>The work sits within a rapidly growing research movement that applies machine learning to the routing layer of the Internet of Things. Recent years have produced Q-learning variants of RPL such as RI-RPL and QSec-RPL, reinforcement-learning-based mobility support schemes like RIMS-RPL, and deep reinforcement learning approaches to clustering and load balancing. What distinguishes MSQ-RPL is the combination of that learning capability with an explicitly multi-sink, mobility-first design and a deliberate emphasis on keeping the protocol lightweight enough for the microcontrollers and radios found in real IoMT hardware. The authors&#8217; earlier work on QFS-RPL, a mobility and energy aware multipath protocol, laid groundwork that this new scheme extends with reward-driven scoring and scope-limited control traffic.</p>
<p>The implications reach well beyond academic benchmarking. Healthcare monitoring, industrial automation, logistics tracking, and environmental sensing all depend on fleets of mobile, energy-constrained devices that must deliver data reliably despite constant motion. A routing layer that learns from its own successes and failures, distributes its intelligence across every node, and treats multiple sinks as opportunities rather than complications could extend network lifetimes and improve data delivery in exactly those settings. The simulation datasets supporting the study are available from the corresponding author upon reasonable request, and the full results appear in the manuscript published in Cluster Computing, volume 29, article 799, dated 27 September 2026. As the Internet of Mobile Things continues to expand, protocols like MSQ-RPL suggest that the path forward may lie not in heavier central coordination, but in teaching every small device to make smarter decisions for itself.</p>
<p><strong>Subject of Research:</strong> A Q-learning-based distributed routing protocol for mobile multi-sink low-power IoT networks</p>
<p><strong>Article Title:</strong> MSQ-RPL: a distributed and scalable reward-driven routing protocol for multi-sink IoMT environments</p>
<p><strong>Article References:</strong> Alilou, M., Babazadeh Sangar, A., Majidzadeh, K., &amp; Masdari, M. (2026). MSQ-RPL: a distributed and scalable reward-driven routing protocol for multi-sink IoMT environments. <em>Cluster Computing, 29</em>(14), Article 799. <a href="https://doi.org/10.1007/s10586-026-06547-2" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06547-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06547-2" rel="noopener noreferrer">10.1007/s10586-026-06547-2</a></p>
<p><strong>Keywords:</strong> Internet of Mobile Things, RPL, Q-learning, routing protocol, multi-sink networks, low-power networks, IoT, Contiki OS, Cooja simulator, energy efficiency, mobility support, multipath forwarding</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229075</post-id>	</item>
	</channel>
</rss>
