For decades, the people inside building energy models have been ghosts. They follow the same schedule every day: asleep at eleven, at work by nine, home at six, lights off at midnight. Real humans, of course, are nothing like that. They work from home on Tuesdays, linger over dinner, crank the air conditioning when a heatwave hits, and ignore utility companies’ pleas to save power. That gap between tidy assumption and messy reality is one of the main reasons simulated energy consumption so often diverges from what the meter actually records. Now a researcher at the University of Arizona has built an open-source platform that replaces those ghostly occupants with something far stranger and far more promising: artificial intelligence agents powered by large language models, grounded in national survey data, and equipped with memories that persist across a simulated day.
The platform, called BuildOcc and described in the journal SoftwareX, tackles a stubborn problem at the heart of building energy research. Traditional models represent occupants through fixed reference schedules, such as those published under ASHRAE Standard 90.1, which assign identical activity profiles to every day and every person. Later generations of stochastic models improved on this by sampling behavior from sensor data or national time-use diaries, using techniques like first-order Markov chains to generate realistic day-to-day variation. But these remain statistical descriptions: they can reproduce the shape of human activity without capturing any of the reasoning behind it. An occupant who turns off the lights before leaving is not following a probability distribution; they are making a decision based on context, habit, and intent.
BuildOcc’s central insight is that large language models, with their capacity for contextual reasoning and memory retention, are unusually well suited to simulating that decision-making. Each agent in the platform is built from six interacting components: a structured demographic persona, an activity scheduler, a memory stream, a reasoning engine, a simulation environment, and a persistent store that saves the agent’s cognitive state between operations. At every fifteen-minute timestep, the reasoning engine assembles the persona, the most relevant retrieved memories, and the current environment into a prompt, and the language model returns a structured action—adjust the thermostat, toggle a device, move to another room, or do nothing—along with a one-sentence explanation of why.
What separates BuildOcc from earlier demonstrations of language-model occupants is its grounding in real population data. Personas and activity schedules are drawn from the American Time Use Survey, a Bureau of Labor Statistics dataset covering more than sixteen thousand respondents, of whom 6,611 fall within the platform’s four demographic strata: employed single adults, retired couples, employed parents, and not-employed adults. Activity probabilities are aggregated by hour of day and day type, weighted so that each estimate represents the U.S. population. Appliance ownership comes from the Residential Energy Consumption Survey, and comfort bands widen for lower income brackets, reflecting the documented tendency of lower-income households to tolerate wider temperature swings to reduce heating and cooling costs. Work-from-home behavior is handled separately, with each agent carrying a probability of working remotely on any given day, drawn from where survey respondents actually reported working.
The memory architecture borrows directly from the influential generative agents work of Park and colleagues at Stanford. Every observation and action the agent takes is written to a memory stream with an importance score assigned by the model itself. Retrieval balances recency, which decays exponentially with a twenty-four-hour half-life, against that importance score. When the accumulated importance of new memories crosses a threshold, the agent pauses to reflect, distilling its recent experiences into higher-order insights that are fed back into the stream. The result is cross-timestep coherence: an agent that raised its thermostat an hour ago will recall that decision when a new request arrives, and judge the request against it. In one demonstration walkthrough, an agent raised its setpoint during peak electricity pricing while still away from home, then declined an educational demand-response signal after arriving, reasoning that thirty-five cents in savings was not worth the discomfort given the comfort band it had already stretched.
That demand-response interface is one of the platform’s most experimentally useful features. BuildOcc supports three mutually exclusive signal types drawn from established behavioral intervention typologies: direct commands, educational price information, and social norm comparisons. In validation experiments, the results were strikingly patterned. Educational price signals were the only type accepted in every demographic stratum, accepted by between one and three agents in five depending on the group. Direct commands were declined in all but one case, with agents reasoning that a thirty-minute air conditioning shutdown on a hot afternoon would push the zone past its comfort limits. Social norm messages—telling agents that most similar households had reduced their use—failed to persuade a single agent in any stratum. A fixed-schedule model, which has no mechanism to respond to a signal at all, could never produce these differences.
The platform’s validation proceeded in two tiers. The first tested the activity scheduler alone, comparing simulated activity distributions against the survey reference data using Kullback-Leibler divergence, a measure of how far one probability distribution departs from another. Every survey-grounded condition fell within or below the sampling noise floor expected from a correct sampler, while a deterministic rule-based baseline diverged by two orders of magnitude—worst for retired and not-employed strata, whose real schedules depart furthest from the standard working-adult assumption baked into conventional models. The second tier showed that demographic priors propagate into distinguishable behavior: retired-couple agents, home all day, were the most active, toggling devices and moving between rooms far more often than employed agents under identical environmental conditions.
Practical considerations have not been ignored. Because the agent queries a language model at nearly every timestep, cost scales linearly with simulated agent-days; measured runs came to roughly twenty cents per agent-day using Anthropic’s models, with wall-clock time dominated by provider latency. Researchers can eliminate per-call costs entirely by running a local open-weight model through Ollama, and the default temperature of zero means a fixed model version and prompt reproduce the same action on every run—though that determinism does not survive across providers or model updates, so long-term studies are advised to pin a local model. Integration is handled through three layers: a Python library for batch simulations, a stateless REST API suited to EnergyPlus co-simulation, and a Model Context Protocol server that lets tools like Claude Desktop, LangChain, or Home Assistant drive simulations without custom integration code. A plugin registry allows other groups to add new demographic strata, substitute alternative time-use surveys such as the Harmonised European Time Use Surveys, or swap in different memory architectures without touching the core library.
The limitations are stated with unusual candor. Each timestep is drawn independently from the hourly activity distribution, so simulated activity episodes are far shorter than the diaries record—simulated sleep averages under seventy minutes against more than five hours in the survey data—and no duration model exists yet. Agents take one action per timestep, the strata cover only U.S. demographics, and the memory importance scores are self-assigned by the agent with no external calibration. The validation establishes internal consistency, not behavioral realism; comparing simulated agents against measured occupant data, including field data from smart thermostats, is left for future work, along with multi-agent household simulation. Still, the conceptual shift is considerable. By treating occupant behavior as a structured, demographically parameterizable experimental variable rather than an uncontrolled source of uncertainty, BuildOcc turns the messiest part of building energy modeling into something researchers can finally hold constant, vary deliberately, and reproduce across laboratories—a small step toward buildings designed not for ghosts, but for us.
Subject of Research: Large language model-based occupant agents for building energy simulation
Article Title: BuildOcc: A large language model occupant agent platform for building energy research
Article References: BuildOcc: A large language model occupant agent platform for building energy research. (n.d.). https://doi.org/10.1016/j.softx.2026.103068
Image Credits: AI Generated
DOI: 10.1016/j.softx.2026.103068
Keywords: BuildOcc, large language models, building energy simulation, occupant behavior, demand response, American Time Use Survey, generative agents, EnergyPlus, open-source software, HVAC, artificial intelligence, energy modeling
Cite Scienmag News
Denise Maddox. (October 1, 2026). AI Agents With Human Memories Take the Guesswork Out of Building Energy Simulation. Scienmag. https://scienmag.com/ai-agents-with-human-memories-take-the-guesswork-out-of-building-energy-simulation/
Denise Maddox. "AI Agents With Human Memories Take the Guesswork Out of Building Energy Simulation." Scienmag, 1 October 2026, https://scienmag.com/ai-agents-with-human-memories-take-the-guesswork-out-of-building-energy-simulation/. Accessed 1 October 2026.
Denise Maddox. "AI Agents With Human Memories Take the Guesswork Out of Building Energy Simulation." Scienmag. October 1, 2026. https://scienmag.com/ai-agents-with-human-memories-take-the-guesswork-out-of-building-energy-simulation/

