Old version — revision 1
This is a fixed snapshot of AI drug discovery, saved by Import as part of the initial corpus import. It is not edited and it is not updated; the article may have changed since.
Edit summary: Initial import of content/ai-drug-discovery.md — the filesystem corpus, unchanged. Not an edit.
The use of machine learning to select drug targets, generate candidate molecules and predict failure, so far demonstrably faster at making compounds than at making medicines.
AI drug discovery is the application of machine learning to the early stages of pharmaceutical research: choosing which protein to drug, generating molecules that might bind it, predicting which of those will be soluble, selective and non-toxic, and planning routes to synthesize them. The chemistry claims are the better supported ones: software now proposes synthesizable, potent-looking compounds against a defined target at a speed no medicinal-chemistry team matches by hand, though how much of that speed the models themselves explain is disputed. The clinical claims are not settled. As of mid-2026 several AI-derived candidates have entered phase 1 and phase 2 trials, and none has completed phase 3 or been approved.
The label bundles at least five tasks that succeed and fail independently.
Target selection ranks proteins whose modulation might change a disease, using expression data, genetic association studies, knowledge graphs assembled from the literature, and increasingly language models reading that literature directly. The most consequential public case is BenevolentAI's identification of baricitinib, an approved rheumatoid arthritis drug, as a plausible COVID-19 treatment in early 2020.1 Baricitinib went on to be authorized for that use. It is a genuine result and a narrow one: the system reranked known pharmacology rather than producing a molecule.
Molecule generation trains models over SMILES strings, molecular graphs or three-dimensional poses, then steers them with reinforcement learning or conditions them on a binding pocket. Insilico Medicine's 2019 report on the kinase DDR1 was the first widely noticed demonstration, claiming designed inhibitors in weeks rather than months.2
Property prediction estimates absorption, metabolism, cardiac ion-channel liability and toxicity before synthesis. This is the least glamorous application and probably the most useful, because it kills bad series cheaply.
Screening triage trains a classifier on a modest set of measured compounds and applies it to libraries too large to test. A graph neural network trained on a few thousand molecules tested against E. coli nominated halicin, a compound originally investigated for diabetes, which killed several drug-resistant bacterial species in culture and in mice.3
Structure and synthesis. Protein structure prediction, covered under AI protein design, opened pocket-based design to proteins with no experimental structure, though benchmark studies of docking have found predicted structures to perform worse than experimental ones. Retrosynthesis models propose routes, sometimes coupled to automated synthesis platforms so that design, make and test run as a closed loop.
The most advanced case is Insilico's ISM001-055, later named rentosertib, an inhibitor of the kinase TNIK for idiopathic pulmonary fibrosis. Both the target and the molecule came out of the company's own platforms, which makes it the cleanest available test of the full pipeline. It cleared phase 1, and results from a twelve-week randomized placebo-controlled phase 2a trial in China were published in 2025: one dose arm showed a mean improvement in forced vital capacity where the placebo arm declined.4 That trial was not sized to establish clinical benefit, and the compound had not entered phase 3 as of mid-2026.
The rest of the first wave has behaved like ordinary pharmacology. DSP-1181 was announced in 2020 as the first AI-designed molecule to enter a clinical trial and did not progress beyond phase 1; a later Exscientia oncology candidate was also dropped after early data. In 2024 Exscientia merged with Recursion, whose approach centred on high-throughput cellular imaging rather than generative chemistry. Neither company had an approved drug at the time.
The neighbouring field of computational protein design has produced no approved medicine of its own either. A designed nanoparticle scaffold is a component of a COVID-19 vaccine approved in South Korea, but as of 2026 no de novo designed protein has been approved anywhere as a therapeutic drug.
What "AI-discovered" meansNo regulator recognizes the category, and companies apply the label to everything from a molecule drawn by a generative model to a conventional campaign in which a classifier filtered one plate. The US Food and Drug Administration's 2025 draft guidance addresses AI used to produce evidence submitted in support of a decision, not the provenance of a molecule.5 There is no registry, no definition, and therefore no denominator against which to count successes.
The mainstream objection is structural rather than technical. Roughly one compound in ten that enters phase 1 is eventually approved, and the largest single cause of failure is lack of efficacy in phase 2: the drug engages its target and the patients do not improve. AstraZeneca's published post-mortem on its own pipeline reached that conclusion and rebuilt the company's project criteria around target and patient selection rather than chemistry.6 That is a statement about biology: the target was wrong, or the animal model misled, or the disease is heterogeneous in ways the trial did not stratify. Generating better molecules faster addresses a step that was not the binding constraint.
The strongest evidence on what does move the number points the same way. Drug targets supported by human genetic evidence are substantially more likely to reach approval than those without it, which is an argument for better target biology rather than better chemistry.7 Target-selection models are trained largely on the published literature that produced the current failure rate, so it is unclear whether they can do more than re-rank the field's existing beliefs efficiently.
The disagreementProponents argue the objection misses compounding effects: cheaper cycles mean more targets tested, cleaner selectivity means fewer toxicity failures, and structure prediction opens proteins that had no tractable starting point. Critics answer that none of this has yet changed a phase 2 outcome, that the industry has absorbed several such tool revolutions without moving approval rates, and that the case rests on an argument rather than on a result. Both sides accept that the question is empirically open until a candidate designed this way succeeds or fails in phase 3.
Public bioactivity data is a record of what people chose to measure and publish. Negative results are under-reported, assays are not comparable across laboratories, and chemical space is sampled around scaffolds that were already interesting. Models trained on it inherit that shape.
Retrospective benchmarks compound the problem. Widely used virtual-screening benchmark sets contain biases that allow a model to score well by learning artefacts of how decoy molecules were assembled rather than anything about binding.8 Because the counterfactual campaign is never run, a company cannot demonstrate that its molecule arrived faster than it would have otherwise, and almost every published speed claim is a comparison against an industry average rather than a control.
Aging biology is where the approach is most enthusiastically applied and least easily judged. Insilico began as an aging-focused company and uses deep aging clocks internally; Calico Life Sciences, Altos Labs and others run computational groups of their own. Screening models have produced real hits: a classifier trained on known senolytic compounds identified three natural products that killed senescent human cells in culture, at a small fraction of the cost of the physical screen.9 Those hits then join every other senolytic candidate in the queue for human evidence that does not yet exist.
The constraint here is sharper than in oncology. A trial of a geroprotector has no accepted endpoint, no validated surrogate measure, and a natural readout measured in decades. Compressing discovery from four years to one changes little when the limiting step is a trial nobody knows how to design. Better preclinical models, whether Organoids or simulation approaches such as Human digital twins, plausibly matter more to healthspan research than faster chemistry does — though the simulation side is least developed in exactly the systems aging research needs, metabolism and immunity, and no model of a whole person exists.
The clearest demonstrated risk is inversion. Researchers who rewarded a toxicity-prediction model for lethality rather than penalizing it generated tens of thousands of candidate toxic molecules, including analogues of known nerve agents, in under a day of computing.10 The molecules were never synthesized, and the barrier to weaponization remains chemistry and delivery rather than design, but the episode is the standard reference in Dual-use research of concern debates about model release and sits inside the broader catastrophic-risk literature on biological misuse.
Less dramatic risks are more likely to bite. Mining the same public data pushes many groups toward the same targets, which narrows rather than broadens the pipeline. Discovery costs are a small fraction of total development spending, so savings there do little for drug pricing and access. And a claim of computational provenance is a marketing asset, which creates pressure to describe conventional programmes in those terms.
The reversibility of anything produced this way is a property of the molecule, not the method: a small molecule with a short half-life clears, a covalent inhibitor does not, and the same tools are being applied to gene therapies whose effects are permanent. The design route carries no distinct risk profile of its own, which is one reason regulators treat AI-derived candidates like any others.
Judgment will take years. The decisive test is whether AI-nominated targets fail in phase 2 at the historical rate or better, which requires dozens of shots and roughly a decade of readouts, not the handful of trials available in 2026. Meanwhile the most defensible gains are in places that attract no headlines: assay triage, guide design that reduces off-target editing, formulation work on delivery vehicles and targeting ligands, and routine property prediction. Nothing about the field's readiness supports the compressed-timeline claims made for it, or the assumption that progress toward more general AI transfers cleanly into biology. The drug class most often named as the past decade's largest pharmacological advance, the GLP-1 receptor agonists, came from decades of peptide endocrinology with no machine learning involved.
paperRichardson, P. et al. "Baricitinib as potential treatment for 2019-nCoV acute respiratory disease." The Lancet, 2020.↩Baricitinib was already approved for rheumatoid arthritis, so the system reranked known pharmacology rather than proposing a new molecule.
paperZhavoronkov, A. et al. "Deep learning enables rapid identification of potent DDR1 kinase inhibitors." Nature Biotechnology, 2019.↩DDR1 was a well-precedented target and the compounds resembled known inhibitors; chemists disputed how much of the speed the model explained.
paperStokes, J.M. et al. "A Deep Learning Approach to Antibiotic Discovery." Cell, 2020.↩Halicin's activity was confirmed in mouse infection models; it has not entered human trials.
paperXu, Z. et al. "A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial." Nature Medicine, 2025.↩Seventy-one patients across three dose arms and placebo, treated for twelve weeks; a trial that size and length cannot establish clinical benefit in fibrosis.
regulatorUS Food and Drug Administration. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. Draft guidance, 2025. ↩
paperCook, D. et al. "Lessons learned from the fate of AstraZeneca's drug pipeline: a five-dimensional framework." Nature Reviews Drug Discovery, 2014. ↩
paperNelson, M.R. et al. "The support of human genetic evidence for approved drug indications." Nature Genetics, 2015.↩The association is between genetic support and eventual approval; it does not show that adding genetics to a programme causes success.
paperChen, L. et al. "Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening." PLoS ONE, 2019.↩The bias sits in the benchmark rather than in any one model, which makes published virtual-screening gains hard to compare.
paperSmer-Barreto, V. et al. "Discovery of senolytics using machine learning." Nature Communications, 2023. ↩
paperUrbina, F. et al. "Dual use of artificial-intelligence-powered drug discovery." Nature Machine Intelligence, 2022.↩The molecules were generated in silico and never synthesized; the paper reports a capability, not a released agent.