The method, in full
What the project used, how its four models were trained, and how to read each kind of result. Everything here is computational: no new patient or laboratory experiments were performed, and every candidate is a lead for laboratory testing.
- 35,232labelled laboratory records to learn from
- 4models, one per pathogen species
- 7,044predictions for 1,761 approved medicines
- 1,019repurposing candidates, leads for the laboratory
How the models learned
Four models, one for each pathogen, learned from published laboratory measurements. This is what they saw, how they learned, and what they are for.
Each laboratory measurement becomes a label
A record says how much of a molecule it took to stop the pathogen growing. Less is more potent.
Records in the middle are not forced into a yes or no. Leaving them out keeps borderline chemistry from teaching the model either way.
One model per pathogen, each with its own records
The bar is every labelled record for that species. The ochre line beneath is the share measured on a named resistant strain.
- MRSA7,403 active · 1,974 inactive12.9% resistant strain
- E. coli7,265 active · 2,464 inactive0.0% resistant strain
- K. pneumoniae4,801 active · 3,671 inactive13.9% resistant strain
- M. tuberculosis5,026 active · 2,628 inactive1.9% resistant strain
So the models describe each species. They were rarely shown the resistant strain itself, and for some species not at all.
It learns a pattern, then estimates for medicines it has not seen
Labelled structures go in as fingerprints. The model learns which structural features go with activity, then scores each approved medicine.
The output is an AI-predicted activity from 0 to 100%. Medicines at ≥40% or more are kept for a closer look. That is a discovery filter, not a clinical cutoff.
What the models are for
- Sorting the approved-medicine library into a short list worth testing in the laboratory.
- Comparing medicines against one of the four species the models cover.
- Pointing at structures that resemble molecules already measured as active.
What they cannot tell you
- Whether a medicine treats an infection in a patient.
- Whether it works against the resistant strain specifically.
- What dose would be needed, or whether that dose is safe.
- Anything about pathogens other than these four.
What we used
- ChEMBLLaboratory bioactivity records and molecular structures35,232labelled laboratory records used for training
- Molecular structuresThe broader molecular dataset the pipeline works from20,258valid structures, each stored as a fingerprint
- FDA Orange BookWhich medicines are approved1,761approved medicines in the screening library
- ClinicalTrials.govRegistered clinical-study history1,761 / 1,761medicines checked · 39,548 registered studies
- Protein Data Bank (PDB)3D protein structures for docking4targets: Dihydrofolate reductase (3FRE); Dihydrofolate reductase (1RX2); KPC-2 carbapenemase (2OV5); Enoyl-ACP reductase (InhA) (4TZK)
- RDKit and AutoDock VinaSoftware: RDKit processes and standardises structures; AutoDock Vina runs the docking898library medicines docked so far (a subset)
The 20,258 molecular structures and the 1,761 approved medicines are different populations. They are never added together.
How the training worked, step by step
Prepare the data
From published measurements to labelled examples
- 01
Gather published measurements for each species
47,804 records gathered
Laboratory records from ChEMBL for MRSA, E. coli, K. pneumoniae and M. tuberculosis: how much of a molecule it took to stop the pathogen growing.
- 02
Put every measurement on one scale
Records come in different units. Each is converted to a molar concentration, so a value in micrograms per millilitre and a value in micromolar can be compared. A record whose structure cannot be read is dropped rather than guessed.
- 03
Label each record, or leave it out
35,232 kept, 12,572 left out
Active if growth stopped at 10 µM or less (24,495 records). Inactive if it took 100 µM or more (10,737). In between is too close to call: 12,572 records left out. A “greater than” measurement can only ever count as inactive, and a “less than” one only as active.
- 04
Describe each molecule as a fingerprint
20,258 fingerprints
Each structure is turned into a molecular fingerprint: a fixed-length pattern of structural features that lets a computer compare molecules.
Learn
One model per species, checked on molecules it has not seen
- 05
Train one model per species
4 models
A machine-learning model learns which fingerprint features go with activity for that species. It learns from earlier laboratory observations; it does not simulate a patient.
- 06
Check it on molecules it has not seen
Part of the data is held back during training, and the model is judged on it. Molecules that share a core structure are kept on the same side of that split, so the check is not flattered by near-copies.
Use
Scoring the approved medicines, then setting aside what is already known
- 07
Score every approved medicine
7,044 predictions
Each of the 1,761 medicines gets an AI-predicted activity for each of the four species (7,044 predictions). 1,276 reach ≥40% for at least one species, a discovery filter rather than a clinical cutoff.
- 08
Set aside what is already an antimicrobial
1,019 candidates
Existing antimicrobials (156) are identified from WHO ATC codes and FDA pharmacologic classes, never from a name. 101 that neither source classifies are kept for review, not counted. 1,019 repurposing candidates remain.
A known limit
Few training records name a resistant strain (MRSA 12.9%, E. coli 0.0%, K. pneumoniae 13.9%, M. tuberculosis 1.9%). The models therefore predict activity against the species, not against the resistant phenotype.
How to read a result
AI-predicted activity, e.g. 81%
Means: How strongly the model expects the molecule to be active against that species in a laboratory test, given the patterns it learned.
Does not mean: A chance of curing anyone, a measurement, or a statement about the resistant strain.
Compare medicines against the same species. ≥40% is where the site starts listing a medicine as worth a closer look.
Docking score, e.g. −9.7 kcal/mol
Means: A computer estimate of how well the molecule fits one selected pathogen protein. More negative is a better fit.
Does not mean: Proof that the molecule binds, or that it has an antimicrobial effect.
This project screens at -7.0 kcal/mol, its own target rather than a universal cutoff. Only a subset of medicines has been docked.
Registered study
Means: A study listed on ClinicalTrials.gov names the medicine.
Does not mean: That the study succeeded, or that the medicine is approved for that use.
Read the listed condition: most registered studies are for the medicine's existing use.
Laboratory record
Means: A measured result from a published laboratory test, recorded in ChEMBL.
Does not mean: A clinical result. Laboratory activity does not show a medicine works in people.
These are the measurements the models learned from, shown separately from any prediction.
- No evidence found
- The source was searched and nothing matched. It is not the same as “does not work”.
- Not yet checked
- The source has not been searched for this medicine yet.
- ≠No effect
- A different statement altogether, which needs actual negative evidence. This site never makes it.
How newly approved medicines are scored
A self-hosted update worker, run separately from this website, checks the FDA Orange Book and ChEMBL for newly approved medicines, scores each new one with the same four trained models, and publishes the results here. It runs on demand as a batch job.
- 1Finds new approved medicines and their structures
- 2Scores each with the four trained models; it never retrains on new data
- 3Checks itself model files are checksum-verified and a known-answer test runs on every start
- 4Publishes the new predictions to the database this site reads
The worker does not yet classify new medicines as antimicrobial or not. Until that step runs, a new medicine reads “not yet checked” and is not counted as a repurposing candidate.
The code behind it
The pipeline, the models’ training, the checks and this website are in one public repository on GitHub.