The Nipah virus, one of the deadliest pathogens on the World Health Organization’s priority list, has long haunted public health officials across South and Southeast Asia. With case fatality rates that can exceed 70 percent, no licensed antiviral therapy, and outbreaks that erupt unpredictably in Bangladesh, India and Malaysia, the virus represents one of the most daunting challenges in modern virology. Now, researchers at the ICMR-National Institute of Virology in Pune, India, have turned to artificial intelligence to close the therapeutic gap, deploying machine learning models to sift through thousands of existing drugs and identify compounds that could be repurposed against the virus. The study, published in the journal Molecular Diversity, describes a computational pipeline that trained algorithms on known Nipah inhibitors, screened a library of more than 9,000 compounds, and validated the most promising candidates through molecular docking and molecular dynamics simulations.
Led by Shivangi Sharma, with Pragya D. Yadav and Sarah Cherian as senior authors, the research team assembled a training dataset of 211 compounds drawn from three publicly curated sources: the Anti-Nipah database, the Nipah Virus Inhibitor Knowledgebase (NVIK), and PubChem, supplemented by a systematic review of the literature. This dataset contained both known inhibitors and inactive compounds, providing the labeled examples needed for supervised learning. Each molecule was converted into a numerical fingerprint using molecular descriptors calculated with the Mordred software, capturing features such as molecular weight, topological indices, electronic properties and atom-type electrotopological states. These descriptors serve as the language through which algorithms perceive chemical structure, allowing a model to learn which patterns of atoms, bonds and charge distributions correlate with antiviral activity against Nipah virus.
The investigators benchmarked seven different supervised machine learning algorithms: Support Vector Machines, Random Forest, Logistic Regression, Decision Tree, k-Nearest Neighbors, Artificial Neural Networks, and Ridge Classifier. Each was trained and evaluated using standard performance metrics, including accuracy, receiver operating characteristic analysis, and the Matthews correlation coefficient, a measure favored in cheminformatics because it remains robust even when classes are imbalanced. Among all seven, the Random Forest model, an ensemble method that aggregates the votes of hundreds of decision trees each trained on random subsets of the data, emerged as the clear winner. It achieved 95 percent accuracy on the training data and, critically, 86 percent on the held-out test set, indicating that the model had learned generalizable chemical patterns rather than simply memorizing its training examples. This balance between training and testing performance is the crucial test of any predictive model, and the Random Forest classifier passed it convincingly.
With a validated model in hand, the team unleashed it on a massive virtual haystack. The screening library comprised 9,021 compounds spanning FDA-approved drugs, molecules in preclinical development, investigational agents in clinical trials, and a dedicated collection of known antivirals. This breadth is the essence of drug repurposing: rather than spending a decade and billions of dollars synthesizing and testing novel chemicals, researchers can ask whether a molecule already optimized for safety, pharmacokinetics and manufacturability might also inhibit a new target. For a virus like Nipah, whose outbreaks are sporadic and unpredictable, the speed advantage is decisive. The machine learning classifier assigned each compound a probability of anti-Nipah activity, and the highest-scoring molecules advanced to the next stage of the pipeline.
That next stage was structural. The researchers focused on two of the virus’s most critical proteins: the attachment glycoprotein G, which sits on the viral surface and latches onto ephrin-B2 and ephrin-B3 receptors on human cells, initiating the entry process; and the RNA-dependent RNA polymerase (RdRp), the L-P protein complex that copies the viral genome and is essential for replication. Blocking either target can cripple the virus, and recent cryo-electron microscopy structures of the Nipah polymerase complex have finally given computational scientists an accurate map to work from. Using the Glide docking engine and the OPLS4 force field, the team computed binding poses and affinity scores for the shortlisted candidates within the binding pockets of both proteins, after careful preparation of protonation states and protein geometry.
Docking scores alone can be misleading, so the researchers added a second layer of physical rigor: molecular dynamics simulations. These simulations, run on the NAKSHATRA high-performance computing facility developed under India’s PM-Ayushman Bharat Health Infrastructure Mission, allow the protein-ligand complexes to flex and move in a simulated aqueous environment over time. A compound that binds well in a static docked pose but falls out of the binding pocket during dynamic simulation is unlikely to be a real inhibitor. By monitoring structural stability, root-mean-square deviations and persistent intermolecular contacts, the team separated genuine binders from computational artifacts.
The final verdict yielded eight candidate molecules, distributed across the two targets. Against the glycoprotein, three compounds stood out: 2,3,4,5,6-pentagalloylglucose (PGG), echinacoside, and parishin A. Against the RNA-dependent RNA polymerase, five compounds showed stable, high-affinity binding: neohesperidin dihydrochalcone, naringin dihydrochalcone, diosmin, orientin, and amikacin. Several of these names will be familiar to natural products chemists. PGG, a heavily galloylated tannin found in oak bark and various medicinal plants, has previously been shown to block influenza A virus and to inhibit the interaction between the SARS-CoV-2 spike protein and the human ACE2 receptor. Echinacoside, derived from Echinacea species, has documented antiviral activity against respiratory viruses and was recently flagged in computational studies against the Zika virus polymerase. Parishin A, a bioactive constituent of the orchid Gastrodia elata, has similarly been implicated in blocking viral entry in prior structural studies.
The polymerase-directed candidates are equally intriguing. Neohesperidin dihydrochalcone is, remarkably, a widely used artificial sweetener approved as a food additive in many countries, and it carries an extensive safety dossier along with documented anti-inflammatory and antioxidant properties; it has also been docked against SARS-CoV-2 proteins in earlier work. Diosmin, a flavonoid used as a vascular tonic in human medicine, has recently demonstrated antiviral potential against influenza A. Orientin, a luteolin glycoside found in passionflower and other botanicals, has been studied experimentally against the SARS-CoV-2 spike protein. Amikacin is the most surprising entry on the list: an aminoglycoside antibiotic in clinical use for decades, it belongs to a drug class that has been shown in independent research to enhance host resistance to viral infections through microbiota-independent mechanisms. Its strong binding to the Nipah polymerase adds a completely new dimension to its potential therapeutic profile.
The choice of targets reflects a deep understanding of Nipah virus biology. The G glycoprotein is the virus’s key to the cell, and its receptor-binding domain has been mapped at atomic resolution in multiple crystal structures, revealing precisely which residues engage ephrin-B2. Antibodies and small molecules that occlude this interface prevent attachment and fusion. The polymerase, meanwhile, is the engine of viral replication, and it has proven to be an Achilles’ heel for related paramyxoviruses: allosteric polymerase inhibitors developed against measles virus and respiratory syncytial virus suppress all RNA synthesis activity in those pathogens. A recent non-nucleotide allosteric inhibitor with pan-coronavirus activity has further demonstrated that polymerase targets can yield broad-spectrum antivirals, making the Nipah L-P complex a highly strategic point of attack. The availability of the cryo-EM structure of the Nipah L-P complex, published in late 2024, transformed this target from an aspirational goal into a practical docking substrate.
The current standard of care for Nipah infection remains rudimentary. Ribavirin, used empirically during the 1998-1999 Malaysian outbreak, showed only ambiguous benefit in observational studies of acute encephalitis patients. The monoclonal antibody m102.4, which neutralizes the G glycoprotein, has been administered on a compassionate basis but is not licensed. Remdesivir and favipiravir both protect animals in challenge experiments, and remdesivir has been tested in compassionate-use settings, yet neither is an approved Nipah therapy. Vaccine development is accelerating: a recombinant vesicular stomatitis virus vector vaccine has advanced toward human trials in outbreak regions, and the University of Oxford launched the world’s first Phase II Nipah vaccine trial with CEPI support. But vaccines alone cannot treat established infections, and the unpredictable geography of spillovers from fruit bats to humans through contaminated date palm sap or direct animal contact means that a stockpiled, orally available antiviral would be an invaluable component of outbreak preparedness.
The authors emphasize that their integrative approach, combining machine learning prediction with structure-based validation, is designed precisely to compress the timeline between an outbreak’s emergence and the availability of candidate therapeutics. Because every compound that survived the pipeline already exists in the pharmacopeia, at least in some form, downstream development could in principle bypass many early-stage hurdles of traditional drug discovery. The team also notes that all data used in the study are publicly available, and the models were built entirely from open resources, a transparency that should allow other groups to reproduce, refine and extend the approach.
Caveats remain, as they do for all purely computational studies. Docking and molecular dynamics predictions must ultimately be confirmed in live-virus assays, which require high-containment BSL-4 laboratories, and subsequently in animal models and human trials. Physicochemical liabilities, metabolic stability and oral bioavailability of the larger natural-product-derived molecules will need careful evaluation using tools such as SwissADME and ADMETlab, both of which the authors employed in their computational assessment. Nevertheless, the study stands as a template for how artificial intelligence can transform the response to neglected, high-consequence pathogens. By converting scattered inhibitor data into a predictive engine and then filtering thousands of real-world drugs through docking and dynamics, the Pune team has handed the Nipah research community a short, chemically diverse and mechanistically grounded list of candidates worth testing the moment the next outbreak appears. In a field where every month of delay can cost lives, that kind of head start could prove invaluable.
Cite Scienmag News
Kristina Jarvis. (September 11, 2026). Machine learning accelerates drug repurposing against Nipah virus with computational validation. Scienmag. https://scienmag.com/machine-learning-accelerates-drug-repurposing-against-nipah-virus-with-computational-validation/
Kristina Jarvis. "Machine learning accelerates drug repurposing against Nipah virus with computational validation." Scienmag, 11 September 2026, https://scienmag.com/machine-learning-accelerates-drug-repurposing-against-nipah-virus-with-computational-validation/. Accessed 11 September 2026.
Kristina Jarvis. "Machine learning accelerates drug repurposing against Nipah virus with computational validation." Scienmag. September 11, 2026. https://scienmag.com/machine-learning-accelerates-drug-repurposing-against-nipah-virus-with-computational-validation/

