Insilico Medicine has launched what it describes as the first Drug Discovery and Development Benchmark as a Service, a testing framework designed to determine whether artificial intelligence systems can make meaningful scientific decisions—or merely excel at recalling information from their training data. The initiative targets a growing weakness in AI evaluation: many existing benchmarks use publicly available questions and datasets that may have been absorbed during model training. As a result, a system can achieve impressive scores without demonstrating that it can solve the uncertain, sequential problems faced by drug researchers.
The new DDD Benchmark is built around the idea that drug discovery should be evaluated in conditions resembling real research rather than conventional examinations. In pharmaceutical development, scientists must combine disease biology, chemistry, molecular modeling, experimental interpretation and clinical reasoning over a period of years. A model that correctly predicts an isolated molecular property may still fail to identify a viable drug candidate because drug development depends on a chain of interdependent decisions involving potency, selectivity, safety, pharmacokinetics, manufacturability and clinical relevance.
Insilico says its benchmark uses carefully decontaminated public datasets alongside proprietary, out-of-distribution test sets. Decontamination is intended to reduce the possibility that an AI model has encountered the exact questions or answers during training. Out-of-distribution testing goes further by assessing performance on examples that differ from the data used to develop a model, offering a more realistic indication of whether it has learned transferable scientific principles rather than memorized familiar patterns.
The evaluation framework is divided into two complementary suites. The first, called Drug Discovery Foundations, contains more than 300 assessments covering core capabilities across the drug development process. These include understanding disease mechanisms, predicting and optimizing molecular properties, designing chemical structures, planning synthetic routes through retrosynthesis, applying structure-based drug design and reasoning about aspects of clinical development. Together, the tests are intended to reveal where a model is scientifically reliable and where its apparent expertise breaks down.
The second suite, Drug Candidate Essentials, examines whether an AI system can navigate an entire discovery program from hit identification to the nomination of a preclinical candidate. This is a substantially more demanding task than answering individual technical questions. A candidate molecule must survive a series of decisions in which each result changes the next step. For example, improving potency may damage solubility, increasing exposure may raise toxicity concerns, and a chemically elegant compound may prove difficult to manufacture. The benchmark therefore focuses on the quality and consistency of sequential decisions.
Insilico plans to evaluate models through standard chat-completions application programming interfaces, allowing organizations to submit systems without rebuilding them for a specialized testing environment. The company will compare model outputs with expert reference baselines and provide a standardized scorecard. Participants can request private assessments for internal or commercial use, while organizations seeking public recognition may publish results on a leaderboard designed to offer a like-for-like comparison between competing models.
The company says the benchmark is also intended for a new generation of AI agents that do more than generate text. These systems may plan experiments, interpret laboratory results, reason across scientific literature and call external tools through protocols such as the Model Context Protocol. Measuring such systems requires evaluating not only whether an answer sounds plausible, but whether the proposed action is scientifically justified, appropriately cautious and useful within a real discovery workflow. An agent that confidently recommends an invalid experiment could be more dangerous than a system that simply admits uncertainty.
Alex Zhavoronkov, founder and chief executive of Insilico Medicine, said the benchmark was developed from the company’s experience building AI systems for the drug discovery value chain. Insilico reports that it has nominated 31 preclinical candidates, received more than 10 investigational new drug clearances and reduced the time required to nominate a preclinical candidate to approximately 12 to 18 months. The company’s lead program, rentosertib, also known as ISM001-055, is described as an AI-discovered and AI-designed TNIK inhibitor currently in Phase III development for idiopathic pulmonary fibrosis.
The DDD Benchmark builds on Insilico’s Pharma.AI platform and MMAI Gym, a post-training environment developed for scientific AI systems. Its broader purpose is to address a credibility problem emerging as pharmaceutical companies increasingly adopt generative models: impressive demonstrations do not necessarily prove that an AI can produce medicines. By testing performance against confidential or newly constructed problems and anchoring at least part of the evaluation to validated discovery programs, Insilico is attempting to shift attention from general benchmark scores to measurable performance in high-stakes scientific work.
The service is available to organizations developing AI for drug discovery or using foundation models in research. Insilico says interested groups can request an evaluation, a private report or placement on the public leaderboard through dddbench.insilico.com. If widely adopted, the framework could create a more transparent method for comparing systems that currently make very different claims about scientific capability—and could help distinguish models that merely sound like scientists from those capable of contributing to the difficult, uncertain process of developing new medicines.
Subject of Research: Artificial intelligence evaluation for drug discovery and development
Article Title: Insilico Medicine Launches Benchmark to Test Whether AI Can Really Discover Drugs
Web References: http://dddbench.insilico.com
Image Credits: Insilico Medicine
Keywords
Generative AI, artificial intelligence, drug discovery, pharmaceutical research, foundation models, medicinal chemistry, retrosynthesis, clinical development, AI benchmarking, Insilico Medicine

