<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>EnergyPlus &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/energyplus/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 00:17:12 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>EnergyPlus &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Agents With Human Memories Take the Guesswork Out of Building Energy Simulation</title>
		<link>https://scienmag.com/ai-agents-with-human-memories-take-the-guesswork-out-of-building-energy-simulation/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 00:17:12 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI agents with long-term memory]]></category>
		<category><![CDATA[AI-powered occupant modeling]]></category>
		<category><![CDATA[American Time Use Survey]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[building energy consumption accuracy]]></category>
		<category><![CDATA[building energy modeling standards]]></category>
		<category><![CDATA[building energy simulation]]></category>
		<category><![CDATA[BuildOcc]]></category>
		<category><![CDATA[demand response]]></category>
		<category><![CDATA[divergence between modeled and actual energy use]]></category>
		<category><![CDATA[energy modeling]]></category>
		<category><![CDATA[EnergyPlus]]></category>
		<category><![CDATA[generative agents]]></category>
		<category><![CDATA[human behavior variability in energy models]]></category>
		<category><![CDATA[HVAC]]></category>
		<category><![CDATA[improving energy efficiency through AI]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in energy research]]></category>
		<category><![CDATA[occupant behavior]]></category>
		<category><![CDATA[open-source building energy platforms]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[persistent memory in AI agents]]></category>
		<category><![CDATA[real-world occupant behavior simulation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220226</guid>

					<description><![CDATA[An open-source platform called BuildOcc uses survey-grounded AI agents with persistent memories to simulate realistic human decision-making in building energy models.]]></description>
										<content:encoded><![CDATA[<p>For decades, the people inside building energy models have been ghosts. They follow the same schedule every day: asleep at eleven, at work by nine, home at six, lights off at midnight. Real humans, of course, are nothing like that. They work from home on Tuesdays, linger over dinner, crank the air conditioning when a heatwave hits, and ignore utility companies&#8217; pleas to save power. That gap between tidy assumption and messy reality is one of the main reasons simulated energy consumption so often diverges from what the meter actually records. Now a researcher at the University of Arizona has built an open-source platform that replaces those ghostly occupants with something far stranger and far more promising: artificial intelligence agents powered by large language models, grounded in national survey data, and equipped with memories that persist across a simulated day.</p>
<p>The platform, called BuildOcc and described in the journal SoftwareX, tackles a stubborn problem at the heart of building energy research. Traditional models represent occupants through fixed reference schedules, such as those published under ASHRAE Standard 90.1, which assign identical activity profiles to every day and every person. Later generations of stochastic models improved on this by sampling behavior from sensor data or national time-use diaries, using techniques like first-order Markov chains to generate realistic day-to-day variation. But these remain statistical descriptions: they can reproduce the shape of human activity without capturing any of the reasoning behind it. An occupant who turns off the lights before leaving is not following a probability distribution; they are making a decision based on context, habit, and intent.</p>
<p>BuildOcc&#8217;s central insight is that large language models, with their capacity for contextual reasoning and memory retention, are unusually well suited to simulating that decision-making. Each agent in the platform is built from six interacting components: a structured demographic persona, an activity scheduler, a memory stream, a reasoning engine, a simulation environment, and a persistent store that saves the agent&#8217;s cognitive state between operations. At every fifteen-minute timestep, the reasoning engine assembles the persona, the most relevant retrieved memories, and the current environment into a prompt, and the language model returns a structured action—adjust the thermostat, toggle a device, move to another room, or do nothing—along with a one-sentence explanation of why.</p>
<p>What separates BuildOcc from earlier demonstrations of language-model occupants is its grounding in real population data. Personas and activity schedules are drawn from the American Time Use Survey, a Bureau of Labor Statistics dataset covering more than sixteen thousand respondents, of whom 6,611 fall within the platform&#8217;s four demographic strata: employed single adults, retired couples, employed parents, and not-employed adults. Activity probabilities are aggregated by hour of day and day type, weighted so that each estimate represents the U.S. population. Appliance ownership comes from the Residential Energy Consumption Survey, and comfort bands widen for lower income brackets, reflecting the documented tendency of lower-income households to tolerate wider temperature swings to reduce heating and cooling costs. Work-from-home behavior is handled separately, with each agent carrying a probability of working remotely on any given day, drawn from where survey respondents actually reported working.</p>
<p>The memory architecture borrows directly from the influential generative agents work of Park and colleagues at Stanford. Every observation and action the agent takes is written to a memory stream with an importance score assigned by the model itself. Retrieval balances recency, which decays exponentially with a twenty-four-hour half-life, against that importance score. When the accumulated importance of new memories crosses a threshold, the agent pauses to reflect, distilling its recent experiences into higher-order insights that are fed back into the stream. The result is cross-timestep coherence: an agent that raised its thermostat an hour ago will recall that decision when a new request arrives, and judge the request against it. In one demonstration walkthrough, an agent raised its setpoint during peak electricity pricing while still away from home, then declined an educational demand-response signal after arriving, reasoning that thirty-five cents in savings was not worth the discomfort given the comfort band it had already stretched.</p>
<p>That demand-response interface is one of the platform&#8217;s most experimentally useful features. BuildOcc supports three mutually exclusive signal types drawn from established behavioral intervention typologies: direct commands, educational price information, and social norm comparisons. In validation experiments, the results were strikingly patterned. Educational price signals were the only type accepted in every demographic stratum, accepted by between one and three agents in five depending on the group. Direct commands were declined in all but one case, with agents reasoning that a thirty-minute air conditioning shutdown on a hot afternoon would push the zone past its comfort limits. Social norm messages—telling agents that most similar households had reduced their use—failed to persuade a single agent in any stratum. A fixed-schedule model, which has no mechanism to respond to a signal at all, could never produce these differences.</p>
<p>The platform&#8217;s validation proceeded in two tiers. The first tested the activity scheduler alone, comparing simulated activity distributions against the survey reference data using Kullback-Leibler divergence, a measure of how far one probability distribution departs from another. Every survey-grounded condition fell within or below the sampling noise floor expected from a correct sampler, while a deterministic rule-based baseline diverged by two orders of magnitude—worst for retired and not-employed strata, whose real schedules depart furthest from the standard working-adult assumption baked into conventional models. The second tier showed that demographic priors propagate into distinguishable behavior: retired-couple agents, home all day, were the most active, toggling devices and moving between rooms far more often than employed agents under identical environmental conditions.</p>
<p>Practical considerations have not been ignored. Because the agent queries a language model at nearly every timestep, cost scales linearly with simulated agent-days; measured runs came to roughly twenty cents per agent-day using Anthropic&#8217;s models, with wall-clock time dominated by provider latency. Researchers can eliminate per-call costs entirely by running a local open-weight model through Ollama, and the default temperature of zero means a fixed model version and prompt reproduce the same action on every run—though that determinism does not survive across providers or model updates, so long-term studies are advised to pin a local model. Integration is handled through three layers: a Python library for batch simulations, a stateless REST API suited to EnergyPlus co-simulation, and a Model Context Protocol server that lets tools like Claude Desktop, LangChain, or Home Assistant drive simulations without custom integration code. A plugin registry allows other groups to add new demographic strata, substitute alternative time-use surveys such as the Harmonised European Time Use Surveys, or swap in different memory architectures without touching the core library.</p>
<p>The limitations are stated with unusual candor. Each timestep is drawn independently from the hourly activity distribution, so simulated activity episodes are far shorter than the diaries record—simulated sleep averages under seventy minutes against more than five hours in the survey data—and no duration model exists yet. Agents take one action per timestep, the strata cover only U.S. demographics, and the memory importance scores are self-assigned by the agent with no external calibration. The validation establishes internal consistency, not behavioral realism; comparing simulated agents against measured occupant data, including field data from smart thermostats, is left for future work, along with multi-agent household simulation. Still, the conceptual shift is considerable. By treating occupant behavior as a structured, demographically parameterizable experimental variable rather than an uncontrolled source of uncertainty, BuildOcc turns the messiest part of building energy modeling into something researchers can finally hold constant, vary deliberately, and reproduce across laboratories—a small step toward buildings designed not for ghosts, but for us.</p>
<p><strong>Subject of Research:</strong> Large language model-based occupant agents for building energy simulation</p>
<p><strong>Article Title:</strong> BuildOcc: A large language model occupant agent platform for building energy research</p>
<p><strong>Article References:</strong> BuildOcc: A large language model occupant agent platform for building energy research. (n.d.). <a href="https://doi.org/10.1016/j.softx.2026.103068" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103068</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103068" rel="noopener noreferrer">10.1016/j.softx.2026.103068</a></p>
<p><strong>Keywords:</strong> BuildOcc, large language models, building energy simulation, occupant behavior, demand response, American Time Use Survey, generative agents, EnergyPlus, open-source software, HVAC, artificial intelligence, energy modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220226</post-id>	</item>
		<item>
		<title>AI Agents Can Now Build and Test Building Energy Models From Plain Language</title>
		<link>https://scienmag.com/ai-agents-can-now-build-and-test-building-energy-models-from-plain-language/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:13:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent benchmark]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI agents in construction planning]]></category>
		<category><![CDATA[AI-driven building design optimization]]></category>
		<category><![CDATA[AI-powered building energy modeling]]></category>
		<category><![CDATA[ASHRAE 90.1]]></category>
		<category><![CDATA[automated energy model creation and modification]]></category>
		<category><![CDATA[building energy modeling]]></category>
		<category><![CDATA[building envelope and equipment simulation]]></category>
		<category><![CDATA[Department of Energy building research]]></category>
		<category><![CDATA[EnergyPlus]]></category>
		<category><![CDATA[HVAC synthesis]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Measure authoring]]></category>
		<category><![CDATA[Model Context Protocol]]></category>
		<category><![CDATA[natural language interface for energy simulation]]></category>
		<category><![CDATA[open-source EnergyPlus engine integration]]></category>
		<category><![CDATA[OpenStudio SDK]]></category>
		<category><![CDATA[OpenStudio-MCP]]></category>
		<category><![CDATA[OpenStudio-MCP software for energy modeling]]></category>
		<category><![CDATA[physics-based building performance diagnostics]]></category>
		<category><![CDATA[plain language building performance analysis]]></category>
		<category><![CDATA[reducing human coding in energy modeling]]></category>
		<category><![CDATA[sandboxing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202051</guid>

					<description><![CDATA[Researchers have unveiled OpenStudio-MCP, an open-source server that lets AI agents create, simulate, and diagnose building energy models from plain-language requests.]]></description>
										<content:encoded><![CDATA[<p>A building&#8217;s lifetime energy bill is largely decided before anyone pours a foundation. Choices about form, envelope, equipment, controls, and operation lock in performance years in advance, and building energy modeling is the discipline that prices those choices before construction. Now researchers at the National Laboratory of the Rockies, working for the U.S. Department of Energy, have unveiled a system that hands that entire modeling workflow to artificial intelligence agents. The software, called OpenStudio-MCP, is described in the journal SoftwareX and lets large language models create, modify, simulate, and diagnose physics-based building energy models from nothing more than a plain-language request, with no human writing code at any point.</p>
<p>The open-source EnergyPlus engine, developed by the Department of Energy, performs the underlying calculations that predict how a design will perform. Most tools operate on its text input files, but practitioners typically work one layer up, in the OpenStudio software development kit, which provides a typed model structure and a library of reusable transformation scripts known as Measures. The new server deliberately builds on that SDK layer rather than raw input files, because the typed object structure is more composable and less error-prone to manipulate, and it connects AI agents to the broader ecosystems of OpenStudio-standards, ComStock, and the Building Component Library. Earlier protocol-based servers, such as EnergyPlus-MCP, work only at the input-file layer and defer geometry creation and HVAC loop construction to future work; OpenStudio-MCP tackles the full lifecycle from creation through simulation and evaluation.</p>
<p>The technical heart of the system is the Model Context Protocol, introduced by Anthropic in late 2024, which allows a language model to drive external software by calling validated, schema-typed tools instead of emitting free-form code. OpenStudio-MCP exposes 197 such tools, organized into self-contained skill modules that are discovered automatically at startup. They span model creation, geometry and thermal zoning, construction and schedule assignment, HVAC synthesis, simulation control, results extraction, quality assurance, and even large-scale parametric analysis on OpenStudio-server. A single call to add_baseline_system, for example, wires a complete air- and plant-loop topology corresponding to one of the ten baseline system types defined in ASHRAE Standard 90.1 Appendix G, a task that has historically required careful manual scripting.</p>
<p>Because language models can call operations out of order, select unsuitable systems, or simply invent SDK methods that do not exist, the server wraps the SDK at two levels. Low-level tools expose explicit OpenStudio operations as typed calls, while higher-level tools encode common workflows such as whole-building creation, baseline HVAC assignment, and result extraction. A separate knowledge layer serves curated workflow guides covering object dependencies, ASHRAE system-selection rules, and Measure authoring, which agents retrieve on demand. The server also guards against the finite context window of any language model: instead of dumping a 5,000-space model or a full 8,760-hour annual result into the conversation, it returns compact structured summaries, previews models without loading them, and distills entire simulations into a handful of summary numbers while the model files themselves remain server-side.</p>
<p>Perhaps the most striking capability is automated Measure authoring. Measures are programs, and extending analysis beyond the off-the-shelf library has traditionally required engineers to double as software developers, putting custom analysis out of reach for many architects and engineers who best understand the building. With OpenStudio-MCP, an agent scaffolds a new Measure from a plain-language request, writes its logic, runs its tests, and applies it inside a sandboxed environment that runs child processes as unprivileged users under a fail-closed Landlock filesystem policy, a seccomp filter denying outbound network access, and strict resource limits. The authors stress these are implemented controls rather than security guarantees, since the software has not undergone a formal external penetration test, but targeted adversarial probes using canary listeners and decoy secrets exercised cross-tenant file reads, environment-secret capture, filesystem escapes, and forged transfer requests.</p>
<p>To demonstrate the system end to end, the researchers gave an AI agent a natural-language request specifying a medium office building, its location, a baseline HVAC system, a comfort criterion, and a four-pipe active-chilled-beam retrofit, naming no tools. Working from an empty session with Claude Opus 4.8, the agent created and simulated a 27-zone, three-story, 53,600-square-foot office using the ASHRAE 90.1-2019 template with a variable-air-volume reheat system. The initial model narrowly failed the comfort criterion with 301.7 occupied unmet hours. Digging into sizing reports, the agent found that heating capacity was adequate but a 50 percent reheat-mode airflow cap was constraining morning warm-up after nighttime setback, raised the cap on all 27 terminals, and brought the model to 78.0 unmet hours at a site energy use intensity of 42.4 kBtu per square foot, squarely within the middle half of the observed U.S. office stock from the 2018 Commercial Buildings Energy Consumption Survey.</p>
<p>The agent then authored a Measure that replaced each terminal with a four-pipe beam connected to the existing chilled- and hot-water plants, verified SDK class names, passed its own tests, and validated 27 beam terminals with no errors, all without a human writing or reviewing code. The retrofit maintained comfort but increased site energy by 8.7 percent, driven by a 179 percent surge in fan energy, because the constant-volume beams forfeited the variable-volume air handler&#8217;s part-load fan savings and economizer hours. Crucially, the agent reported this adverse result, explained both mechanisms, and proposed remedies including a right-sized dedicated outdoor air system with energy recovery. The session used 98 calls to 36 tools, three annual simulations, and roughly 21 minutes, with no human intervention after the initial request.</p>
<p>Feasibility is not reliability, so the team built a reproducible benchmark of 16 graded tasks across six families, tested with Claude Opus 4.8, Opus 4.6, Sonnet 4.6, and Haiku 4.5 alongside GPT-5.4 and GPT-5.4-mini. Each trial was graded by two deterministic checks without any AI judging: whether the agent called an acceptable tool, and whether the saved model passed physical checks such as assembly R-values, HVAC loop membership, and pinned EnergyPlus outputs. The distinction proved essential, because in 18 of 23 outcome failures an agent replaced a roof assembly with one up to 1.86 square meters kelvin per watt worse while reporting success, an error visible only in the saved artifact, not in the agent&#8217;s confident report. Under a common configuration with all tool schemas loaded, GPT-5.4 and Opus 4.8 achieved 100 percent outcome rates, with Opus 4.6 at 95.8 percent, GPT-5.4-mini at 93.8, Sonnet at 91.7, and Haiku at 85.4.</p>
<p>The ablation results carry a practical lesson for anyone deploying agentic AI. Loading every tool schema up front raised the weakest model&#8217;s success rate by 12.5 points but inflated costs by 36 to 68 percent for stronger Claude tiers while changing outcomes by at most two tasks, meaning deferred schema discovery saves money for capable models at no accuracy loss. The curated knowledge layer, surprisingly, changed no model&#8217;s completion rate by more than 6.3 points. The completion budget also mattered: at a 120-second limit, 17 of Opus 4.8&#8217;s 18 failures were timeouts, yet it passed every trial in four of five configurations when given 600 seconds, showing that slow but productive work was being misclassified as failure. An unscaffolded baseline without the server showed agents can handle basic OpenStudio operations through direct scripting, but both tested models failed a task requiring exact counting of warnings in an EnergyPlus error file.</p>
<p>The researchers frame the work as broadening access rather than replacing rigor. Engineers and architects could request, test, and run bespoke retrofit analyses without writing Ruby, while organizations could offer shared modeling capacity to design firms, classrooms, or utility programs by issuing authentication tokens instead of provisioning workstations, turning energy modeling from a per-seat desktop activity into shared infrastructure. Validation and quality-assurance tools, audit records of every tool call, and artifact-based grading make the checking explicit, but the authors are careful to note that engineering judgment is not automated: generated artifacts and conclusions require review by a qualified practitioner before use in real engineering decisions. The code, benchmark harness, and archived trial records are openly available under a BSD-3-Clause-style license, inviting the building science community to put AI-driven modeling to the test.</p>
<p><strong>Subject of Research:</strong> An open-source Model Context Protocol server enabling AI agents to perform full-lifecycle building energy modeling with the OpenStudio SDK.</p>
<p><strong>Article Title:</strong> OpenStudio-MCP: a model context protocol (MCP) server for AI agent-driven building energy modeling with the OpenStudio SDK</p>
<p><strong>Article References:</strong> Ball, B. L., Long, N., Fleming, K., &amp; Goldwasser, D. (2026). OpenStudio-MCP: a model context protocol (MCP) server for AI agent-driven building energy modeling with the OpenStudio SDK. <em>SoftwareX, 36</em>, Article 103020. <a href="https://doi.org/10.1016/j.softx.2026.103020" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103020</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103020" rel="noopener noreferrer">10.1016/j.softx.2026.103020</a></p>
<p><strong>Keywords:</strong> OpenStudio-MCP, building energy modeling, Model Context Protocol, large language models, AI agents, EnergyPlus, OpenStudio SDK, HVAC synthesis, Measure authoring, ASHRAE 90.1, agent benchmark, sandboxing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202051</post-id>	</item>
	</channel>
</rss>
