Smart homes promise to watch over elderly relatives and patients with chronic diseases, alerting caregivers the moment something goes wrong. But most of today’s activity recognition systems share a fundamental blind spot: they can label what a person is doing right now, yet they struggle to predict what that person will do next. A new study published in Neural Computing and Applications by Samyan Qayyum Wahla and Muhammad Usman Ghani Khan of the University of Engineering and Technology Lahore argues that this predictive gap matters most precisely for the people who need assistive technology the most, because their daily routines rarely follow the tidy, generalizable scripts that conventional forecasting models are trained on.
The researchers’ central insight is that human behavior is deeply individual. A generalized forecasting model, trained on the pooled habits of many people, may learn that most people make breakfast after waking up, but it will fail when a particular patient’s morning ritual diverges from the statistical average. For dependent individuals, such as older adults living alone or people managing chronic illness, those divergences are not noise to be smoothed away. They are the earliest signals that something has changed, potentially something alarming. The team therefore proposes an N-of-1 framework, a term borrowed from personalized medicine meaning a model built for a single subject rather than a population, that learns the behavioral fingerprint of one specific person from the life-logs generated by activity recognition modules.
Technically, the framework is organized as a hierarchy that captures human behavior at multiple levels of granularity simultaneously. At the lowest level, the system recognizes atomic actions, the short, indivisible units of movement such as reaching for a cup or opening a cabinet. At a higher level, it recognizes composite actions, which are sequences of atomic actions that together constitute meaningful activities such as preparing a meal or tidying a room. Existing systems, the authors note, typically categorize actions at only one of these levels, either unit actions, composite actions, or broad behaviors. By contrast, the new framework brings all levels together in a single pipeline, so that the model can reason about a person’s day both in fine-grained motion terms and in higher-level semantic terms.
The temporal backbone of the framework is the transformer, the attention-based architecture that has reshaped modern machine learning since its introduction in 2017. In this system, the transformer serves as an encoder of the temporal timeline: it takes the stream of recognized actions produced by the activity recognition modules and learns the patterns embedded in that sequence, capturing which actions tend to follow which, at what times, and in what contexts. Once those temporal patterns have been encoded, a decoder processes the resulting feature representations to forecast the next action in the sequence. This encode-then-decode design mirrors the way transformers handle language, but here the vocabulary consists of human actions rather than words, and the sentences are the unfolding routines of daily life.
What elevates the framework beyond simple next-action prediction is its use of the learned patterns for safety monitoring. Because the model builds a personalized picture of how a specific individual normally behaves, deviations from that picture become detectable. The learned patterns are used to recognize anomalies and to forecast alarming situations earlier than would otherwise be possible. In an assistive living context, this matters enormously: a system that only reacts after a fall or a medical emergency has occurred offers limited protection, whereas one that notices a person’s routine quietly drifting away from its usual shape can raise an alert while intervention is still easy and effective.
To validate the approach, the researchers conducted a case study focused on vision modalities, testing the framework on the Toyota SmartHome Untrimmed dataset, known as TSU. This dataset is an important benchmark because, unlike curated clips that show one clean action at a time, it consists of real-world, untrimmed video recordings of daily activities in smart home environments. Untrimmed footage is far closer to what an actual assistive system would face: long stretches of continuous living in which relevant actions are embedded in a stream of ordinary, unlabeled moments. The experimental results showed that the proposed framework achieves better accuracy compared to existing activity forecasting models on this challenging benchmark.
The study situates itself within a rich body of prior work on human activity recognition and prediction. Earlier research has explored activity recognition using wearable accelerometers, Wi-Fi channel state information, smartwatches paired with Bluetooth beacons, and depth cameras combined with inertial sensors. Other lines of work have applied long short-term memory networks to predict activities in smart homes, used change-point detection to segment activities more effectively, and built ontology-based systems to infer high-level context from low-level sensor data. Anomaly detection in the daily routines of older persons in single-resident smart homes has also been studied as a proof of concept. The new framework distinguishes itself by combining personalization, multi-level action categorization, and forecasting within one architecture, rather than treating these as separate problems.
The choice of the TSU dataset also reflects a deliberate methodological stance. The authors report that the data supporting their findings are available on reasonable request, and the research was approved by the research committee of the University of Engineering and Technology Lahore. Because the study relies on a publicly available dataset, the consent protocols followed for the subjects participating in the underlying data collection apply. The authors declare that no funds, grants, or other support were received during the preparation of the manuscript, and they report no relevant financial or non-financial conflicts of interest. All authors contributed to the study’s conception and design, with data collection, analysis, and experiments performed by Wahla, who also wrote the first stable draft.
The implications for assistive living are considerable. As populations age across much of the world, the demand for technologies that allow people to live independently and safely at home is growing faster than the supply of human caregivers. Systems like the one described in this study point toward a model of care in which the home itself becomes a quiet, continuous observer, one that knows its resident well enough to distinguish a harmless quirk from a warning sign. Personalization is the key to that distinction: a model that has learned one person’s rhythms can be sensitive to that person’s changes in a way no population-level model can.
There remain, of course, practical hurdles between a validated research framework and a deployed assistive product. The current validation rests on vision modalities, and real deployments would need to handle variable lighting, occlusion, privacy concerns around cameras, and the integration of other sensor types such as wearables and ambient sensors. The N-of-1 approach also implies that each new user requires a period during which the system learns their individual patterns before it can forecast and detect anomalies reliably. Nevertheless, the study demonstrates that treating behavior prediction as a personalized, multi-level, sequence-modeling problem, rather than a generalized classification problem, produces measurably better forecasting accuracy on realistic smart home data. For a field whose ultimate goal is to protect vulnerable people in their own homes, that shift in perspective, from the average person to the individual, may prove to be the most consequential contribution of all.
Subject of Research: Personalized transformer-based human activity forecasting for assistive living using smart home sensor data
Article Title: N-of-1 multi-modal human activity forecasting framework for assistive living
Article References: Wahla, S. Q., & Khan, M. U. G. (2026). N-of-1 multi-modal human activity forecasting framework for assistive living. Neural Computing and Applications, 38(19), Article 762. https://doi.org/10.1007/s00521-026-12079-y
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12079-y
Keywords: human activity forecasting, assistive living, smart homes, transformer, N-of-1 personalization, activity recognition, anomaly detection, Toyota SmartHome Untrimmed dataset, elderly care, machine learning, computer vision, behavioral modeling
Cite Scienmag News
Blake Davidson. (October 1, 2026). AI Learns One Person at a Time to Predict Daily Actions and Flag Danger Early. Scienmag. https://scienmag.com/ai-learns-one-person-at-a-time-to-predict-daily-actions-and-flag-danger-early/
Blake Davidson. "AI Learns One Person at a Time to Predict Daily Actions and Flag Danger Early." Scienmag, 1 October 2026, https://scienmag.com/ai-learns-one-person-at-a-time-to-predict-daily-actions-and-flag-danger-early/. Accessed 1 October 2026.
Blake Davidson. "AI Learns One Person at a Time to Predict Daily Actions and Flag Danger Early." Scienmag. October 1, 2026. https://scienmag.com/ai-learns-one-person-at-a-time-to-predict-daily-actions-and-flag-danger-early/

