<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>employing deep reinforcement learning to optimize video quality &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/employing-deep-reinforcement-learning-to-optimize-video-quality/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 04:56:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>employing deep reinforcement learning to optimize video quality &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Stream Video That Saves Your Battery Without Ruining Picture Quality</title>
		<link>https://scienmag.com/ai-learns-to-stream-video-that-saves-your-battery-without-ruining-picture-quality/</link>
		
		<dc:creator><![CDATA[Faith Mcneil]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 04:56:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[ABR improves user experience. The new ESM system enhances this by also considering battery life]]></category>
		<category><![CDATA[adaptive bitrate]]></category>
		<category><![CDATA[and energy efficiency simultaneously.]]></category>
		<category><![CDATA[Cluster Computing]]></category>
		<category><![CDATA[deep reinforcement learning]]></category>
		<category><![CDATA[employing deep reinforcement learning to optimize video quality]]></category>
		<category><![CDATA[energy efficiency]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[mobile devices]]></category>
		<category><![CDATA[network throughput]]></category>
		<category><![CDATA[PPO]]></category>
		<category><![CDATA[QoE]]></category>
		<category><![CDATA[Quality of Experience]]></category>
		<category><![CDATA[smooth playback]]></category>
		<category><![CDATA[video stream based on network conditions]]></category>
		<category><![CDATA[video streaming]]></category>
		<category><![CDATA[VMAF]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=257478</guid>

					<description><![CDATA[Researchers at Huzhou University have developed ESM, a deep reinforcement learning system that adapts video streaming bitrates to improve perceived quality by nearly 10 percent while cutting modeled energy consumption by over 14 percent.]]></description>
										<content:encoded><![CDATA[<p>Every time you watch a video on your phone, an invisible algorithm is making a split-second decision about how much data to pull down from the network. Choose a higher bitrate and the picture looks crisp; choose a lower one and you save bandwidth but risk a blurry image. Now a team of researchers at Huzhou University in China has added a new dimension to that decision: how much energy the whole process burns through your battery. Their system, called ESM, uses deep reinforcement learning to balance picture quality, playback smoothness and power consumption, and in trace-driven simulations it delivered an average Quality of Experience improvement of 9.97 percent alongside an average modeled energy reduction of 14.44 percent compared with the average performance of the baseline algorithms they evaluated.</p>
<p>The work, published in the journal Cluster Computing by Wenlong Ke, Ying Wang and Yong Zhang, addresses a blind spot in what is known as Adaptive Bitrate, or ABR, streaming. Traditional ABR algorithms exist to solve one central problem: network throughput fluctuates constantly, especially on cellular connections, so a fixed video quality would either stutter when the connection dips or waste capacity when it surges. By dynamically adjusting the bitrate of each successive video chunk, ABR systems such as the buffer-based BOLA approach or control-theoretic methods like the one behind DASH keep playback smooth. What those classic designs largely ignore, the researchers argue, is that downloading and playing video is one of the most power-hungry activities on a mobile device, and that the energy cost of a streaming session depends on far more than just the bitrate the client selects.</p>
<p>ESM reframes the adaptation problem in two important ways. First, instead of driving its decisions with raw bitrate figures, it evaluates candidate chunks using perceptual quality metrics, specifically VMAF, the Video Multimethod Assessment Fusion score that estimates how good a video actually looks to a human viewer. This matters because the relationship between bitrate and perceived quality is nonlinear and content-dependent: a fast-moving sports scene may need far more bits than a static talking-head shot to reach the same visual quality. By feeding VMAF values for the next chunk at each available bitrate level directly into the decision process, the agent can target perceptual quality rather than a proxy number.</p>
<p>Second, the system explicitly models energy. In the researchers&#8217; formulation, the total energy consumed for each video chunk is decomposed into two components: the energy spent acquiring the data over the network and the energy spent playing the video back locally. The energy cost of data acquisition depends on the size of the chunk and the actual throughput the network delivers, because a slow, strained radio link keeps the transmitter active longer per byte. Local playback energy, in turn, scales with the video bitrate being decoded and rendered. Crucially, the energy model incorporates user experience preferences through dynamic scaling factors, allowing the trade-off between quality and battery life to be tuned rather than hard-coded.</p>
<p>To actually learn a policy under this richer objective, the team built ESM on deep reinforcement learning, and in particular on an improved version of Proximal Policy Optimization, or PPO. PPO is an actor-critic algorithm: a policy network, the actor, learns to select bitrate decisions, while a value network, the critic, estimates how good each state is so that the actor&#8217;s updates are guided by an advantage signal rather than raw rewards. PPO&#8217;s signature trick is clipped surrogate optimization, which constrains how far the policy can move in a single update and thereby stabilizes training. The ESM authors go further with a dual clipping parameter, extending the bounds on policy updates in both directions, which they combine with an entropy term weighted toward a target entropy to keep the agent exploring rather than collapsing prematurely onto a mediocre strategy.</p>
<p>The state that the agent observes at each decision point is deliberately rich. It includes throughput measurements and download times from past observations, the sizes and VMAF values of the next video chunk at every candidate bitrate level, the current playback buffer occupancy, the VMAF score of the previously downloaded chunk, and the number of chunks remaining in the video. The action space is the choice of bitrate for the next chunk. Rewards combine the perceptual quality achieved with penalties for rebuffering time and, distinctively, the energy consumed, with penalty coefficients weighting the components. Because the QoE weights themselves are dynamic scaling factors drawn from a tunable range, the same trained framework can express different user priorities, from quality-maximal viewing to battery-sipping endurance mode.</p>
<p>Evaluating such a system honestly requires more than a toy setup. The researchers ran trace-driven simulations using network traces and video content configured to match diverse real-world conditions, drawing on the kind of bandwidth recordings gathered from mobile and cellular environments in prior measurement studies. Against the average performance of all evaluated baseline algorithms, which include both rule-based ABR schemes and learning-based methods developed in recent years, ESM&#8217;s gains of 9.97 percent in QoE and 14.44 percent in model-calculated relative energy consumption represent a meaningful margin. The authors are careful to note a caveat that is often glossed over in machine learning papers: these figures are aggregated statistically from the specific network traces and video samples used in the study, and the optimization gains will fluctuate somewhat across different combinations of content and network conditions.</p>
<p>The broader significance of the work lies in where streaming is heading. Mobile video already dominates global internet traffic, and short-video platforms have made continuous, hours-long consumption the norm, multiplying the energy bill paid in batteries and, upstream, in data centers and radio networks. Researchers studying quality of experience have begun quantifying the trade-off between viewer satisfaction and carbon emissions across current and future network generations, and dedicated measurement datasets for content-consumption energy are now emerging. Earlier energy-aware streaming efforts tended to bolt energy considerations onto existing ABR logic or to sacrifice quality aggressively in exchange for power savings. ESM&#8217;s contribution is to fold energy directly into the learning objective alongside perceptual quality, so the agent discovers the Pareto frontier itself rather than following a hand-tuned heuristic.</p>
<p>The choice of PPO as the learning backbone also reflects lessons the field has absorbed since deep reinforcement learning first entered networking with systems like Pensieve in 2017. Value-based methods such as deep Q-networks can be brittle when the action space and state complexity grow, and streaming presents a partially observable, nonstationary environment in which the network behaves almost adversarially. Policy-gradient methods with trust-region-style constraints, entropy regularization and advantage estimation have proven more robust in exactly these settings, and variants of soft actor-critic and PPO now appear across the learning-based streaming literature. ESM&#8217;s dual-clipped, entropy-targeted refinement is a natural evolution of that trend, aimed at making training converge reliably across the wildly different bandwidth regimes a phone encounters over a single commute.</p>
<p>The authors have released their code and data in a public GitHub repository, which matters for a subfield where reproducibility has been a persistent complaint; randomized in-situ experiments have shown that simulation results can diverge from live deployments in surprising ways. The next frontier for systems like ESM is validation on physical devices, where Wi-Fi and 5G power behavior, thermal throttling and vendor-specific decoding pipelines complicate any analytical energy model. Still, the direction is clear. As the researchers demonstrate with their trace-driven results, an agent that understands both what viewers see and what their batteries pay can deliver better streaming on both fronts at once, and that multi-objective balance under hardware constraints is precisely what the next billion hours of mobile video will demand.</p>
<p><strong>Subject of Research:</strong> Energy-efficient adaptive bitrate video streaming using deep reinforcement learning</p>
<p><strong>Article Title:</strong> ESM: energy-efficient bitrate adaptation for video streaming with deep reinforcement learning</p>
<p><strong>Article References:</strong> Ke, W., Wang, Y., &amp; Zhang, Y. (2026). ESM: energy-efficient bitrate adaptation for video streaming with deep reinforcement learning. <em>Cluster Computing, 29</em>(13), Article 746. <a href="https://doi.org/10.1007/s10586-026-06576-x" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06576-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06576-x" rel="noopener noreferrer">10.1007/s10586-026-06576-x</a></p>
<p><strong>Keywords:</strong> adaptive bitrate, video streaming, deep reinforcement learning, energy efficiency, quality of experience, PPO, VMAF, mobile devices, QoE, network throughput, Cluster Computing, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">257478</post-id>	</item>
	</channel>
</rss>
