Fever in a returning traveller is one of the most diagnostically intimidating presentations in medicine. A patient who has just stepped off a long-haul flight may be harbouring anything from a self-limiting viral illness to malaria, dengue, typhoid, or a host of other tropical and non-tropical infections, and the initial hours of assessment can determine whether the outcome is a routine discharge or an admission to intensive care. Clinicians must weigh travel itineraries, vaccination histories, exposure timelines, and subtle patterns in vital signs against a backdrop of geographic variation in disease epidemiology. It is precisely this kind of complex, high-stakes decision-making that has drawn the attention of artificial intelligence researchers, who see an opportunity for machine learning systems to support, rather than replace, the clinical judgement of emergency physicians and travel medicine specialists.
A new scoping review published in PLOS Digital Health has now mapped the landscape of AI research applied to this specific clinical problem. The review, conducted by Bhavya Gandhi, Leo Morjaria, Imeth Illamperuma, Meha Bhatt, and Andrew Kapoor, followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews guidelines and was registered on the Open Science Framework, giving the work a transparent and reproducible methodology. The central question was deceptively simple: what does the published literature actually tell us about the use of artificial intelligence in assessing fever in the returning traveller? The answer, it turns out, is that the evidence base is remarkably thin, consisting of just three eligible studies that each took a distinct approach to the problem.
The scarcity of studies is itself a striking finding. Fever in returning travellers sits at the intersection of emergency medicine, infectious diseases, and global health, and it generates substantial clinical anxiety because missing a diagnosis of malaria or another rapidly progressive infection can be fatal. Yet despite decades of enthusiasm for clinical decision support systems, only three studies met the review’s inclusion criteria. This narrow evidence base means that any enthusiasm about AI in this domain must be tempered by the recognition that the field is essentially at its starting line, with far more questions than answers about how algorithms would perform in the messy reality of a busy emergency department.
The three studies identified were notably heterogeneous, each evaluating a different AI strategy and each measuring success against its own study-specific endpoints. That heterogeneity makes direct comparison difficult, a familiar challenge in the broader literature on machine learning in healthcare, where models are trained on different datasets, validated on different outcomes, and reported with different metrics. Some approaches focused on predicting which patients were at risk of serious disease, while others explored ways of structuring the diagnostic workup itself. What united them was that all reported promising performance within their own controlled settings, a phrase that carries important caveats, because controlled settings are precisely where AI systems tend to shine before encountering the unpredictable conditions of real clinical practice.
The review’s authors were careful to spell out the limitations that temper the promising headline results. None of the three studies had undergone external validation, meaning the models were evaluated on data similar to those on which they were developed rather than on independent patient populations from different institutions or countries. External validation is widely regarded as a critical gatekeeper in medical AI, because models frequently degrade when moved to new environments where patient demographics, local disease prevalence, and documentation practices differ from the original training data. A malaria prediction model trained in one referral centre, for example, may encode the baseline prevalence of that centre’s catchment population and misjudge risk when deployed elsewhere.
The populations studied were also narrow, another recurring weakness in the clinical AI literature. If the patients included in a study do not reflect the full spectrum of returning travellers, including children, elderly patients, immunocompromised individuals, and travellers returning from a wide range of destinations, then the model’s reported performance cannot be generalised to the diverse patients who actually present with fever after international travel. The review also found an absence of real-world implementation studies and, crucially, no user acceptance testing, meaning that nobody has yet examined whether clinicians would trust these tools, how they would integrate them into their workflow, or whether the systems would change decisions in ways that improve patient outcomes.
These gaps matter because the history of clinical decision support is littered with tools that performed well in retrospective evaluations but failed to influence care when deployed. An algorithm that achieves impressive discrimination between high-risk and low-risk febrile travellers on historical data may still falter when confronted with incomplete travel histories, atypical presentations, or the time pressures of a night shift. Implementation science has repeatedly shown that the success of a clinical tool depends as much on workflow design, clinician trust, and alert fatigue as on statistical performance. Without studies measuring real-world effectiveness and safety, the clinical utility of AI in this setting remains, as the review puts it, unestablished.
One of the most topical observations in the review concerns large language models, the class of AI systems behind modern conversational agents. These models represent an emerging area of interest for clinical assessment because of their ability to process free-text clinical narratives, synthesize scattered information, and reason across broad medical knowledge. Yet the review found that large language models were not evaluated in any of the studies included, which means that the fastest-moving corner of AI has not yet been formally tested against the specific challenge of the febrile returning traveller. Given how quickly these systems are being adopted in other areas of medicine, their absence from this literature highlights how far behind this particular clinical niche remains.
The technical challenges specific to this domain help explain why progress has been slow. Fever in returning travellers is a low-prevalence, high-consequence problem, and rare but dangerous diagnoses such as severe malaria or viral haemorrhagic fever are uncommon even in specialist centres, making it difficult to assemble datasets large enough to train robust models. Travel itineraries and exposure histories are recorded inconsistently across health systems, and the ground truth against which an algorithm would be judged, the definitive diagnosis, is often uncertain or delayed. Any AI system would need to handle this uncertainty gracefully and communicate risk in a way that supports rather than undermines clinical reasoning, requirements that go well beyond achieving a good area under the receiver operating characteristic curve.
The overall message from the review is one of cautious, conditional optimism. The limited available evidence suggests that AI may have future potential to support the assessment of fever in returning travellers, and the three existing studies demonstrate that researchers are beginning to engage with the problem. But the authors are explicit that further research is needed before any claims of clinical benefit can be made, and they call for work that explores these technologies in real-world clinical settings, with external validation, broader populations, and evaluation of user acceptance and safety. For now, the febrile returning traveller will continue to be assessed by human clinicians drawing on travel medicine expertise, with artificial intelligence waiting in the wings, promising but not yet proven.
Subject of Research: Artificial intelligence for the clinical assessment of fever in returning international travellers
Article Title: The use of artificial intelligence in assessing fever in the returning traveller: A scoping review
Article References: Gandhi, B., Morjaria, L., Illamperuma, I., Bhatt, M., & Kapoor, A. (2026). The use of artificial intelligence in assessing fever in the returning traveller: A scoping review. PLOS Digital Health, 5(10), e0001742. https://doi.org/10.1371/journal.pdig.0001742
Image Credits: AI Generated
DOI: 10.1371/journal.pdig.0001742
Keywords: artificial intelligence, fever, returning traveller, travel medicine, scoping review, machine learning, clinical decision support, infectious disease, external validation, large language models, emergency medicine, PLOS Digital Health
Cite Scienmag News
Ophelia Keating. (October 11, 2026). AI Shows Promise but Remains Unproven for Assessing Fever in Returning Travellers. Scienmag. https://scienmag.com/ai-shows-promise-but-remains-unproven-for-assessing-fever-in-returning-travellers/
Ophelia Keating. "AI Shows Promise but Remains Unproven for Assessing Fever in Returning Travellers." Scienmag, 11 October 2026, https://scienmag.com/ai-shows-promise-but-remains-unproven-for-assessing-fever-in-returning-travellers/. Accessed 11 October 2026.
Ophelia Keating. "AI Shows Promise but Remains Unproven for Assessing Fever in Returning Travellers." Scienmag. October 11, 2026. https://scienmag.com/ai-shows-promise-but-remains-unproven-for-assessing-fever-in-returning-travellers/








