Saturday, August 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Chemistry

Insilico Medicine Launches Industry’s First Drug Discovery Benchmark-as-a-Service for Frontier AI Models

August 1, 2026
in Chemistry
Reading Time: 3 mins read
0
Insilico Medicine Launches Industry’s First Drug Discovery Benchmark-as-a-Service for Frontier AI Models

Insilico Medicine Launches Industry’s First Drug Discovery Benchmark-as-a-Service for Frontier AI Models

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Insilico Medicine has launched what it describes as the first Drug Discovery and Development Benchmark as a Service, a testing framework designed to determine whether artificial intelligence systems can make meaningful scientific decisions—or merely excel at recalling information from their training data. The initiative targets a growing weakness in AI evaluation: many existing benchmarks use publicly available questions and datasets that may have been absorbed during model training. As a result, a system can achieve impressive scores without demonstrating that it can solve the uncertain, sequential problems faced by drug researchers.

The new DDD Benchmark is built around the idea that drug discovery should be evaluated in conditions resembling real research rather than conventional examinations. In pharmaceutical development, scientists must combine disease biology, chemistry, molecular modeling, experimental interpretation and clinical reasoning over a period of years. A model that correctly predicts an isolated molecular property may still fail to identify a viable drug candidate because drug development depends on a chain of interdependent decisions involving potency, selectivity, safety, pharmacokinetics, manufacturability and clinical relevance.

Insilico says its benchmark uses carefully decontaminated public datasets alongside proprietary, out-of-distribution test sets. Decontamination is intended to reduce the possibility that an AI model has encountered the exact questions or answers during training. Out-of-distribution testing goes further by assessing performance on examples that differ from the data used to develop a model, offering a more realistic indication of whether it has learned transferable scientific principles rather than memorized familiar patterns.

The evaluation framework is divided into two complementary suites. The first, called Drug Discovery Foundations, contains more than 300 assessments covering core capabilities across the drug development process. These include understanding disease mechanisms, predicting and optimizing molecular properties, designing chemical structures, planning synthetic routes through retrosynthesis, applying structure-based drug design and reasoning about aspects of clinical development. Together, the tests are intended to reveal where a model is scientifically reliable and where its apparent expertise breaks down.

The second suite, Drug Candidate Essentials, examines whether an AI system can navigate an entire discovery program from hit identification to the nomination of a preclinical candidate. This is a substantially more demanding task than answering individual technical questions. A candidate molecule must survive a series of decisions in which each result changes the next step. For example, improving potency may damage solubility, increasing exposure may raise toxicity concerns, and a chemically elegant compound may prove difficult to manufacture. The benchmark therefore focuses on the quality and consistency of sequential decisions.

Insilico plans to evaluate models through standard chat-completions application programming interfaces, allowing organizations to submit systems without rebuilding them for a specialized testing environment. The company will compare model outputs with expert reference baselines and provide a standardized scorecard. Participants can request private assessments for internal or commercial use, while organizations seeking public recognition may publish results on a leaderboard designed to offer a like-for-like comparison between competing models.

The company says the benchmark is also intended for a new generation of AI agents that do more than generate text. These systems may plan experiments, interpret laboratory results, reason across scientific literature and call external tools through protocols such as the Model Context Protocol. Measuring such systems requires evaluating not only whether an answer sounds plausible, but whether the proposed action is scientifically justified, appropriately cautious and useful within a real discovery workflow. An agent that confidently recommends an invalid experiment could be more dangerous than a system that simply admits uncertainty.

Alex Zhavoronkov, founder and chief executive of Insilico Medicine, said the benchmark was developed from the company’s experience building AI systems for the drug discovery value chain. Insilico reports that it has nominated 31 preclinical candidates, received more than 10 investigational new drug clearances and reduced the time required to nominate a preclinical candidate to approximately 12 to 18 months. The company’s lead program, rentosertib, also known as ISM001-055, is described as an AI-discovered and AI-designed TNIK inhibitor currently in Phase III development for idiopathic pulmonary fibrosis.

The DDD Benchmark builds on Insilico’s Pharma.AI platform and MMAI Gym, a post-training environment developed for scientific AI systems. Its broader purpose is to address a credibility problem emerging as pharmaceutical companies increasingly adopt generative models: impressive demonstrations do not necessarily prove that an AI can produce medicines. By testing performance against confidential or newly constructed problems and anchoring at least part of the evaluation to validated discovery programs, Insilico is attempting to shift attention from general benchmark scores to measurable performance in high-stakes scientific work.

The service is available to organizations developing AI for drug discovery or using foundation models in research. Insilico says interested groups can request an evaluation, a private report or placement on the public leaderboard through dddbench.insilico.com. If widely adopted, the framework could create a more transparent method for comparing systems that currently make very different claims about scientific capability—and could help distinguish models that merely sound like scientists from those capable of contributing to the difficult, uncertain process of developing new medicines.

Subject of Research: Artificial intelligence evaluation for drug discovery and development

Article Title: Insilico Medicine Launches Benchmark to Test Whether AI Can Really Discover Drugs

Web References: http://dddbench.insilico.com

Image Credits: Insilico Medicine

Keywords

Generative AI, artificial intelligence, drug discovery, pharmaceutical research, foundation models, medicinal chemistry, retrosynthesis, clinical development, AI benchmarking, Insilico Medicine

Tags: AI decision-making in medicineAI drug discovery evaluationAI model validation for safety and efficacybenchmark for pharmaceutical researchdeep learning for drug developmentimproving drug discovery AI performanceinterdisciplinary drug development challengesmolecular property prediction accuracyout-of-distribution testing in AI modelspharmaceutical research AI benchmarksproprietary datasets for AI testingreal-world drug discovery assessment
Share26Tweet16
Previous Post

Imaging Reveals How Brain–Body Interactions Shape Systemic Disease

Next Post

Pan-cancer pro-angiogenic atlas reveals tumor-educated pericyte-driven anti-angiogenic resistance

Related Posts

New wearable sensors improve uric acid monitoring, review finds
Chemistry

New wearable sensors improve uric acid monitoring, review finds

August 1, 2026
Spectroscopy System Detects Aerosols Using Common Surfaces
Chemistry

Spectroscopy System Detects Aerosols Using Common Surfaces

August 1, 2026
Netherlands Joins CTAO ERIC as Observer
Chemistry

Netherlands Joins CTAO ERIC as Observer

August 1, 2026
Smart sensor identifies current molecules using memories of past measurements
Chemistry

Smart sensor identifies current molecules using memories of past measurements

August 1, 2026
Targeted cooling boosts insulation efficiency in liquid-hydrogen tanks
Chemistry

Targeted cooling boosts insulation efficiency in liquid-hydrogen tanks

August 1, 2026
Oxford chemists unlock Appel fluorination using potassium fluoride
Chemistry

Oxford chemists unlock Appel fluorination using potassium fluoride

August 1, 2026
Next Post
Pan-cancer pro-angiogenic atlas reveals tumor-educated pericyte-driven anti-angiogenic resistance

Pan-cancer pro-angiogenic atlas reveals tumor-educated pericyte-driven anti-angiogenic resistance

  • Mothers who receive childcare support from maternal grandparents show more

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Scalable In-Process Inspection for Direct-Ink-Writing Additive Manufacturing
  • New wearable sensors improve uric acid monitoring, review finds
  • Spectroscopy System Detects Aerosols Using Common Surfaces
  • Divergent Porcine Astrovirus 4 Forms Distinct Respiratory and Gastrointestinal Genetic Clades

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,147 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading