Large language models have already rewritten the rules of how we search, write, and code. Now, according to a sweeping systematic review published in Artificial Intelligence Review, they are being retooled for something far more intimate: understanding your health, your habits, and your unique circumstances well enough to nudge you toward a better life. A team of researchers at Dalhousie University, led by Japheth Mumo Kimeu and including Gerry Chan, Rita Orji, and Oladapo Oyebode, analyzed 149 peer-reviewed studies to answer a deceptively simple question: how well can these models actually adapt to individual people, rather than dispensing the same generic advice to everyone?
The answer, the review finds, is both promising and sobering. The researchers identified three core dimensions along which LLM-driven health systems personalize their behavior, eight distinct technical architectures that developers use to adapt these models to individual users, and four categories of implementation challenges that consistently undermine effectiveness. That structure matters because personalization in health is not a single feature you can bolt onto a chatbot. It is a design problem that spans what the system knows about you, how it reasons about that knowledge, and how it changes its outputs over time as your needs evolve.
To understand why this review lands at such a pivotal moment, consider what makes large language models different from the health apps that came before them. Earlier generations of digital health tools relied on rigid rule engines and decision trees: if a user logged three workouts in a week, the app sent a congratulatory notification; if a diabetic patient’s glucose reading crossed a threshold, an alert fired. These systems were transparent but brittle. They could not hold a conversation, could not interpret the messy, ambiguous way people actually describe their lives, and could not adjust their tone or content to a user’s literacy level, culture, or emotional state. LLMs, by contrast, are trained on vast corpora of text and can generate fluent, contextually sensitive responses to essentially any prompt, which makes them natural candidates for conversational health coaching, mental wellness support, medication guidance, and lifestyle intervention.
But fluency is not the same as personalization. A model that answers every question beautifully but identically for a teenager in Nairobi, a retiree in Halifax, and a shift nurse in Mumbai is not adaptive; it is merely articulate. The Dalhousie team’s synthesis of the literature shows that researchers have attacked this problem from multiple technical directions. Among the eight adaptation architectures catalogued in the review are approaches such as fine-tuning, where a base model is retrained on domain-specific health data to sharpen its medical reasoning; retrieval-augmented generation, where the model consults external knowledge sources or a user’s personal records before responding; prompt engineering and in-context learning, where user profiles, preferences, and history are packed into the prompt itself so the model conditions its answers on them; and agent-based designs, where multiple specialized model components collaborate, with one handling medical accuracy, another managing conversational tone, and a third tracking long-term user goals.
Each architecture carries distinct trade-offs that the review helps clarify for the field. Fine-tuning can produce deep domain competence but requires curated datasets, computational resources, and careful validation to avoid degrading a model’s general capabilities or introducing subtle biases. Retrieval-augmented approaches keep personal data outside the model weights, which eases privacy concerns and allows information to be updated in real time, but they depend on the quality and currency of the underlying data stores. Prompt-based personalization is cheap and flexible, yet it consumes context window space and can be fragile when a user’s profile grows complex. Agentic pipelines offer modularity and auditability but add engineering complexity and latency. The review’s contribution is to map these options against the health domains where they have been deployed, giving designers a conceptual framework for choosing the right tool for the right intervention.
The health domains covered in the synthesized literature are strikingly broad. LLM-driven systems have been explored for mental health support, where conversational agents can offer around-the-clock availability and nonjudgmental interaction that some users find easier than talking to a human; for chronic disease management, where models can interpret self-reported symptoms and medication logs; for physical activity and nutrition coaching, where adaptive feedback loops respond to a user’s progress; and for health information access, where models translate clinical jargon into plain language tailored to a reader’s background. Across these domains, the review highlights evidence that LLM-driven applications can be effective, while emphasizing that their success hinges on how well they meet individual needs and preferences rather than on raw model capability alone.
Here is where the review delivers its most uncomfortable finding: inclusivity and contextual sensitivity remain underexplored across the entire field. Most systems studied to date are built and evaluated with narrow populations in mind, often English-speaking, digitally literate, and from high-income settings. A personalization engine that models a user’s preferences but ignores their cultural context, language, disability status, or socioeconomic reality risks producing interventions that work beautifully in a pilot study and fail, or even cause harm, in the real world. The authors frame this as a central gap for the human-computer interaction community, arguing that future systems must be designed to be adaptive, personalized, inclusive, effective, and responsive to evolving individual characteristics and needs, a five-part standard that raises the bar well beyond current practice.
The four categories of implementation challenges identified in the review compound that concern. Technical hurdles include the tendency of LLMs to generate plausible but incorrect health information, a failure mode that is dangerous when the subject is medication dosing or symptom triage. Data challenges revolve around the scarcity of high-quality, representative datasets for training and evaluation, particularly for underrepresented groups. Ethical and privacy challenges loom largest for deployment: health conversations are among the most sensitive data a person can share, and questions about consent, data retention, and secondary use remain unsettled in most jurisdictions. Finally, evaluation challenges persist because there is no consensus on how to measure personalization quality itself; a system can score well on generic benchmarks while failing catastrophically at adapting to any single real user.
The ethical dimension receives particular attention in the review, which situates technical choices inside questions of fairness, autonomy, and global impact. If LLM health assistants become a mainstream interface to care, the populations excluded from their training data could face systematically worse automated advice, deepening existing health disparities rather than closing them. Conversely, the authors point to genuine opportunities for global reach: conversational agents that speak low-resource languages, that operate on inexpensive hardware, and that deliver evidence-informed guidance to communities with limited access to clinicians could extend the frontier of wellness support to billions of people. Whether that promise materializes depends on the design decisions being made in laboratories right now, which is precisely why a structured map of the field arrives at such a useful moment.
The Dalhousie team’s conceptual framework, built from the three personalization dimensions, eight adaptation architectures, and four challenge categories, is offered as a shared vocabulary for the next generation of research. Rather than each lab reinventing its own taxonomy of what personalization means, future studies can locate their contributions within a common structure, compare results across domains, and identify which combinations of architecture and dimension remain untested. The review’s dataset, including all coding records from the 149 articles, is available in supplementary material, and the article is open access, lowering the barrier for researchers in resource-constrained settings to build on it. The work was published on 30 September 2026 in Artificial Intelligence Review, with no conflicts of interest declared by the authors.
What emerges from this synthesis is a field in transition. The first wave of LLM health applications proved that the technology could hold convincing, helpful conversations about well-being. The second wave, which this review both documents and aims to accelerate, must prove something harder: that these systems can know who they are talking to, respect who that person is, and change what they say accordingly, safely, and fairly. The 149 studies analyzed here suggest the ingredients are in place, from mature adaptation architectures to growing awareness of ethical pitfalls. What remains is the disciplined, inclusive engineering work of turning a chatty generalist into a genuinely personal health companion, one carefully validated system at a time.
Subject of Research: Systematic review of large language model-driven personalized and adaptive systems for health and wellness
Article Title: LLM-driven personalized and adaptive systems for health and wellness: a systematic review
Article References: Kimeu, J. M., Chan, G., Orji, R., & Oyebode, O. (2026). LLM-driven personalized and adaptive systems for health and wellness: a systematic review. Artificial Intelligence Review. https://doi.org/10.1007/s10462-026-11703-6
Image Credits: AI Generated
DOI: 10.1007/s10462-026-11703-6
Keywords: large language models, personalization, adaptive systems, digital health, wellness, human-computer interaction, systematic review, health informatics, artificial intelligence, machine learning, e-health, conceptual framework
Cite Scienmag News
Blake Davidson. (September 30, 2026). AI Health Coaches Get Personal: Massive Review Maps How Chatbots Could Tailor Care to You. Scienmag. https://scienmag.com/ai-health-coaches-get-personal-massive-review-maps-how-chatbots-could-tailor-care-to-you/
Blake Davidson. "AI Health Coaches Get Personal: Massive Review Maps How Chatbots Could Tailor Care to You." Scienmag, 30 September 2026, https://scienmag.com/ai-health-coaches-get-personal-massive-review-maps-how-chatbots-could-tailor-care-to-you/. Accessed 30 September 2026.
Blake Davidson. "AI Health Coaches Get Personal: Massive Review Maps How Chatbots Could Tailor Care to You." Scienmag. September 30, 2026. https://scienmag.com/ai-health-coaches-get-personal-massive-review-maps-how-chatbots-could-tailor-care-to-you/

