HVAC systems are the quiet giants of commercial building energy consumption, and few environments expose their inefficiencies more clearly than the open-plan office. Because unpartitioned zones exchange heat through convection, the thermal behavior of one area of the floor is inseparable from its neighbors, and conventional controllers that treat each zone in isolation routinely waste energy chasing setpoints they cannot actually hold. A new study published in Neural Computing and Applications by M. A. Attia, M. A. Abdelaal, and E. A. Sallam of Tanta University in Egypt argues that the solution does not require the deep neural networks that dominate current headlines. Instead, the researchers resurrected and re-engineered SARSA, a tabular reinforcement learning algorithm first introduced in the 1990s, and showed that a carefully designed decentralized version can cut annual HVAC energy consumption in a simulated six-zone office by 39.39 percent compared with a schedule-based baseline, while keeping thermal comfort violations to just 2.85 percent of occupied hours.
The appeal of the approach lies in what it leaves out. Deep reinforcement learning methods such as Deep Q-Networks have demonstrated impressive results in building control, learning nonlinear strategies directly from sensor data without an explicit model of the building’s thermodynamics. Yet their adoption in real building energy management systems has been hampered by practical friction: training instability, high computational cost, and hardware demands that exceed what typical controllers installed in commercial buildings can provide. A tabular learner, by contrast, stores its knowledge as an explicit table mapping situations to action values, which makes it transparent, compact, and fast. The Tanta team’s best-performing controller made decisions in about 45 microseconds per zone and required a Q-table of only 1.6 megabytes, figures that place the method squarely within the reach of resource-constrained building automation hardware.
SARSA, which stands for State-Action-Reward-State-Action, belongs to the family of on-policy temporal-difference learning algorithms. Unlike its off-policy cousin Q-learning, which learns the value of the greedy action regardless of what the agent actually does, SARSA evaluates the policy it is currently following, updating its action values based on the action it actually takes next. This distinction matters in a building context, where exploration must be gentle and where an overly aggressive update can translate into uncomfortable temperature swings for occupants. The on-policy character of SARSA tends to produce more conservative, stable learning, which the authors identified as a key reason for choosing it over deep alternatives for a safety-sensitive application like climate control.
The central technical challenge was inter-zone coupling. In an open-plan office, warm air drifting from a sunlit zone near the facade alters the cooling load of the interior zone beside it, so a controller that observes only its own local temperature is fundamentally flying blind. The researchers addressed this with a coupling-aware observation design: each zone agent receives, in addition to its local conditions, an aggregate summary of the temperatures in neighboring zones. This lightweight information sharing allows each tabular learner to compensate for the heat exchanged across the invisible boundaries of the open floor, without requiring a centralized controller that would scale poorly as the number of zones grows. The result is a decentralized architecture in which simple agents, each with a modest memory footprint, collectively manage a thermally entangled space.
Training took place in a simulated six-zone office environment driven by TMY3 typical meteorological year weather data from the EnergyPlus weather library, ensuring that the agents experienced a realistic annual profile of outdoor temperatures and solar conditions. Occupancy was modeled with stochastic hybrid-work schedules, reflecting the irregular attendance patterns that have become common in post-pandemic offices and that routinely defeat fixed timetables. The reward function was deliberately multi-objective, balancing thermal comfort against energy consumption and penalizing excessive equipment switching, since frequent ON/OFF cycling of HVAC components accelerates wear and shortens equipment life. The best-trained policy settled at an average of 4.1 switching actions per zone per day, a level the authors suggest is compatible with practical equipment longevity.
The headline result, a 39.39 percent reduction in annual HVAC energy use relative to a clearly defined schedule-based baseline, comes with an important caveat of scientific honesty: the comparison target is a conventional scheduled controller, not a state-of-the-art model predictive control system. Even so, the robustness of the finding is strengthened by replication. Across ten independent training runs, the final energy savings converged to 38.0 plus or minus 0.9 percent, indicating that the learning process is reliable rather than the product of a lucky random seed. The authors also conducted a sensitivity analysis over the reward weights, mapping out the trade-off curve between comfort and energy and showing how the balance can be tuned to a building operator’s priorities.
Thermal comfort was assessed against established standards for human occupancy, and the trained controller held violations to 2.85 percent of occupied hours, a figure that suggests the energy savings did not come at the expense of shivering or sweltering occupants. This comfort-energy balance is the perennial tension of building automation: the cheapest way to save energy is simply to switch everything off, but the value of a learning controller lies in discovering subtler strategies, such as pre-conditioning zones ahead of occupancy, exploiting thermal inertia, and modulating output in response to predicted heat flows between zones. The coupling-aware observations appear to be what enables these strategies, since agents that understand their neighbors’ states can coordinate implicitly through the physics of the shared air volume.
The study situates itself within a rapidly expanding literature on reinforcement learning for building energy management, a field that has produced deep learning controllers demonstrated in real buildings and multi-agent frameworks for allocating energy across zones. What distinguishes this contribution is its deliberate minimalism. Where much of the field has pursued ever-larger neural function approximators, the Tanta researchers pursued the opposite direction, asking how much intelligence can be extracted from the simplest possible learner if the state representation is designed thoughtfully. Their answer is that a tabular method, often dismissed as incapable of handling high-dimensional control problems, becomes competitive when the problem is decomposed sensibly and the agents are given just enough information about coupling to act intelligently.
The practical implications are significant for the building industry, where retrofitting advanced control into existing stock is often blocked by cost and complexity. A controller whose entire learned knowledge fits in 1.6 megabytes and whose decision latency is measured in microseconds can run on the embedded processors already present in many building energy management systems, without cloud connectivity or GPU acceleration. That said, the authors are explicit that their results are simulation-based, and embedded field validation remains future work. The gap between simulated and real buildings, with unmodeled sensor noise, actuator degradation, and occupant behavior, is where many promising control algorithms have stumbled, and the transition from the EnergyPlus simulation to a physical open-plan office will be the true test.
Still, the study offers a compelling counterpoint to the assumption that progress in building intelligence must ride on ever-deeper networks. As buildings account for a substantial share of global energy demand and HVAC represents one of the largest end uses in commercial spaces, even fractional improvements compound into meaningful climate impact. If a decades-old, on-policy learning rule, given a clever observation design and a well-balanced reward, can deliver nearly 40 percent savings with hardware that costs almost nothing extra, the barrier to adoption drops dramatically. The work suggests that for many practical control problems, the smartest move may not be a bigger model, but a better-posed question, and it hands building engineers a lightweight, stable, and interpretable tool that could bring learning-based control out of the laboratory and into the ordinary office floor.
Subject of Research: On-policy reinforcement learning with SARSA for energy-efficient multi-zone HVAC control in open-plan office buildings
Article Title: Enhancing energy management in multi-zone buildings using the on-policy reinforcement learning algorithm SARSA
Article References: Attia, M. A., Abdelaal, M. A., & Sallam, E. A. (2026). Enhancing energy management in multi-zone buildings using the on-policy reinforcement learning algorithm SARSA. Neural Computing and Applications, 38(17), Article 717. https://doi.org/10.1007/s00521-026-12445-w
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12445-w
Keywords: SARSA, reinforcement learning, HVAC control, building energy management, open-plan offices, multi-zone buildings, thermal comfort, energy optimization, tabular learning, smart buildings, EnergyPlus simulation, decentralized control
Cite Scienmag News
Denise Maddox. (October 7, 2026). A 40-Year-Old Learning Algorithm Slashes Office HVAC Energy Use by Nearly 40 Percent. Scienmag. https://scienmag.com/a-40-year-old-learning-algorithm-slashes-office-hvac-energy-use-by-nearly-40-percent/
Denise Maddox. "A 40-Year-Old Learning Algorithm Slashes Office HVAC Energy Use by Nearly 40 Percent." Scienmag, 7 October 2026, https://scienmag.com/a-40-year-old-learning-algorithm-slashes-office-hvac-energy-use-by-nearly-40-percent/. Accessed 7 October 2026.
Denise Maddox. "A 40-Year-Old Learning Algorithm Slashes Office HVAC Energy Use by Nearly 40 Percent." Scienmag. October 7, 2026. https://scienmag.com/a-40-year-old-learning-algorithm-slashes-office-hvac-energy-use-by-nearly-40-percent/

