Hogyan segít az AI a rákgyógyszerek keresésében: Bepillantás a BIOPTIC molekuláris felfedezőmotorjába
Szerző: Wendy Frey
Finding a new drug is often described as looking for a needle in a haystack. In reality, the problem is much bigger: the “haystack” can contain billions, or even trillions, of possible molecules.
This is where artificial intelligence enters the laboratory.
BIOPTIC, an AI-first biotech company associated with entrepreneur and former Google executive Andrey Doronichev, is developing technology that uses machine learning to search enormous chemical libraries for molecules that could eventually become new medicines. One of the company's most interesting programs focuses on cancer, where its AI system has been used to identify selective inhibitors of VEGFR-3, a biological target involved in tumor angiogenesis.
The important word here is eventually. AI does not press a button and produce a finished cancer drug. Instead, it helps scientists decide which molecules are worth synthesizing and testing in the real world.
What Is BIOPTIC and What Does Its AI Actually Do?
BIOPTIC operates at the intersection of artificial intelligence, computational chemistry and drug discovery. Its core idea is to treat chemical discovery partly as a search problem.
A conventional drug discovery program may begin with a biological target, a protein involved in a disease, and then search for a molecule capable of interacting with that target in the desired way. The challenge is that the number of theoretically possible drug-like molecules is astronomical.
BIOPTIC's B1 system attempts to make this search manageable.
The system represents molecules mathematically, creating compact numerical embeddings that allow chemical structures to be compared at high speed. According to the company's peer-reviewed research, B1 can search libraries containing tens of billions of compounds and retrieve molecules that may have similar biological properties even when their chemical structures differ significantly.
| Stage | What happens |
|---|---|
| 1. Target selection | Scientists choose a disease-related protein |
| 2. AI search | The model scans huge virtual molecular libraries |
| 3. Ranking | Promising candidates are prioritized |
| 4. Synthesis | Selected molecules are physically produced |
| 5. Lab testing | Researchers measure real biological activity |
| 6. Optimization | Results guide the next search cycle |
The AI is therefore not replacing the laboratory. It is acting as an extremely fast filter before the laboratory work begins.
How Does BIOPTIC B1 Search Billions of Molecules?
The technical foundation of B1 is a ligand-based virtual screening system.
A ligand is a molecule that can bind to a biological target, such as a protein. If scientists already know several molecules with useful activity, those molecules can serve as starting points for a new search.
B1 converts molecular information into a compact vector representation. In the published system, each molecule is represented as a 60-dimensional embedding stored using float16 precision. The model can then compare a query molecule with pre-indexed molecular representations and calculate similarity at enormous scale.
The system was tested on Enamine REAL Space, a virtual chemical library containing approximately 40 billion enumerated molecules.
That scale creates an engineering problem as much as a scientific one. BIOPTIC's researchers separated the process into two phases:
- Scanning: neural networks process molecules and generate embeddings.
- Searching: the precomputed embeddings are compared to retrieve the most relevant candidates.
In the published benchmarks, a search across the 40-billion-compound library could be performed in about two minutes and fifteen seconds using the SSD-based search setup described in the paper. A RAM-based configuration was considerably faster.
This does not mean the AI understands biology in the same way a scientist does. It means the system can rapidly navigate a mathematical map of chemical space and identify candidates that deserve closer attention.
The Cancer Project: Searching for Selective VEGFR-3 Inhibitors
BIOPTIC's cancer program provides a concrete example of how this approach can work.
The target is VEGFR-3, also known as FLT4, a receptor involved in signaling pathways connected to angiogenesis, the formation of new blood vessels. Tumors can exploit angiogenesis to support their growth, making this pathway an important area of cancer research.
The challenge is selectivity.
Drugs that broadly affect multiple VEGF receptors can also interact with other biological systems, potentially increasing toxicity. BIOPTIC's project aimed to identify small molecules with stronger selectivity for VEGFR-3.
The company used a two-model pipeline:
| AI component | Function |
|---|---|
| SMILES-based language model | Searches and ranks molecules using chemical representations |
| Graph Neural Network | Re-scores top candidates for kinase selectivity |
| Oncobox analysis | Uses RNA-seq data to explore potentially responsive cancer types |
SMILES is a text-based notation used to describe molecular structures. This makes chemistry surprisingly compatible with language-model techniques: a molecule can be represented as a sequence of symbols, allowing AI systems to learn patterns in chemical data.
After the first model screened an ultra-large library of 40 billion compounds, a Graph Neural Network was used as a second filter. Unlike a text sequence, a graph represents atoms as nodes and chemical bonds as connections, making GNNs useful for modelling molecular structure and relationships.
From a Digital Prediction to a Real Molecule
This is the point where the AI story becomes real science.
The computer can predict that a molecule looks promising, but the molecule must still be synthesized and tested experimentally.
In the VEGFR-3 project, BIOPTIC reported testing 110 compounds using a kinase profiling assay. Several molecules showed activity against VEGFR-3, and three were selective for a single VEGFR target. One candidate demonstrated more than 45-fold selectivity over VEGFR-1 and VEGFR-2 in the reported experiments. The findings were presented in an abstract for the 2025 ASCO Annual Meeting and published as a meeting abstract in the Journal of Clinical Oncology.
The researchers then combined molecular discovery with RNA sequencing analysis. Using the Oncobox algorithm, they analyzed thousands of tumor expression profiles to estimate which cancer types might be most responsive to selective VEGFR-3 inhibition. The reported candidates for further investigation included papillary thyroid cancer, clear-cell renal cancer, pancreatic cancer, ovarian cancer and sarcomas.
These are research findings, not evidence that a finished cancer medicine is ready for patients.
The Scientific Evidence Behind BIOPTIC B1
BIOPTIC's technology has also been described in peer-reviewed scientific literature.
In 2025, researchers published a study in the Journal of Chemical Information and Modeling describing BIOPTIC B1 and its use in a prospective drug discovery experiment targeting LRRK2, a protein associated with Parkinson's disease.
The system searched the same 40-billion-molecule chemical space and identified novel ligands for both normal and mutant forms of LRRK2. The best experimentally validated compounds achieved dissociation constants as low as 110 nanomolar.
The next stage is particularly important for understanding AI-assisted drug discovery. Initial hits were not treated as final answers. Scientists used the best compounds as starting points for an analog search, retrieved structurally related molecules, filtered them for practical properties and selected candidates that could realistically be synthesized.
Of 48 chosen analogs, 47 were successfully synthesized. The resulting compounds were then tested again, creating a feedback loop between computation, chemistry and biology.
Why AI Cannot Replace Drug Scientists
The phrase “AI discovers drugs” is catchy, but incomplete.
AI can rank molecules based on patterns learned from data. It can accelerate virtual screening and identify chemical relationships that would be difficult to explore manually. But a prediction remains a prediction until experimental work confirms it.
Scientists still need to answer fundamental questions:
- Does the molecule actually bind to the target?
- Does it affect the intended biological pathway?
- Is it selective enough?
- Is it toxic?
- Can it reach the right tissue?
- Is it metabolized safely?
- Does it work in animals?
- Does it help human patients?
A promising molecule can fail at any point.
That is why the real value of systems such as BIOPTIC B1 may be speed rather than magic. If AI can reduce billions of theoretical possibilities to a few dozen molecules worth testing, it can make the earliest stages of drug discovery more efficient.
The Future of AI-Driven Medicine
BIOPTIC's work illustrates a broader transformation in biotechnology. AI is increasingly becoming part of the research infrastructure itself, a system for searching chemical space, prioritizing experiments and connecting molecular data with biological information.
The company's cancer research is especially interesting because it combines several layers of computation: molecular screening, selectivity modelling and RNA-seq analysis of tumor biology.
But the final product of this process is not an AI-generated prescription. It is a shortlist of hypotheses for scientists to test.
That distinction matters.
AI may help researchers search billions of molecules faster than traditional methods can. It may uncover unfamiliar chemical scaffolds and point laboratories toward candidates that would otherwise remain hidden. Yet the molecule still has to survive chemistry, biology, toxicology, preclinical research and clinical trials.
In other words, AI is not replacing the scientist who searches for a cure for cancer.
It is giving that scientist a vastly more powerful search engine.




