Screening services

Library triage run for you, rather than software you run yourself. Send a compound list and a target; get back a ranked, filtered, annotated shortlist and a written read on what to trust in it.

Prefer to run it yourself? The app and API do the same scoring on a self-serve plan from $99/month. This page is for teams who would rather hand over the list and get an answer back.

The pilot

Send up to 250 compounds that are already plausible binders for a target you have assay data on — hits from a screen you have run, analogues around a known active, a series you have measured. Make the potency range as wide as you can: the scorer separates orders of magnitude, not twofold differences. Public, blinded, or internal. Scored free, returned within 48 hours, so you can grade it against results you already trust.

What not to send. An unfiltered vendor or screening library. The scorer cannot tell binders from non-binders and ranks such a set no better than chance, which is measured and published rather than hedged. Nor a tight analogue series inside one log unit, which is the opposite failure.

If the ordering doesn't help on your chemistry, that is a genuinely useful answer and it costs you a CSV. I would rather find that out on 250 compounds than after an invoice.

Three worked examples, including one failure

Run on PDB structures released after the model's training cutoff, so these are targets it had never seen. “Lift” is the mean true pKd of the compounds the model ranked top, minus the mean of a random selection of the same size.

TargetSeries spreadSpearmanLift over randomVerdict
CA2 (n=12)7.0 log+0.72+2.51 pKdWorked — captured 88% of the achievable gain
Q96MU7 (n=21)3.3 log+0.21+0.36 pKdMarginal — 43% of achievable
NOS1 (n=18)2.3 log−0.06+0.09 pKdFailed — no better than picking at random

The pattern is the spread of the series, not the protein. These are small sets (n = 12–21) and should be read as illustrations rather than a benchmark — the validation page has the 516-complex version. But they show the shape of it: wide range, useful; narrow range, not.

What a project includes

  1. Ranked CSV. Predicted pKd, QED, LogP, TPSA, HBD/HBA, rotatable bonds, Lipinski compliance, PAINS/BRENK/NIH structural-alert flags, and your own series tags and IDs preserved.
  2. ChEMBL analog mapping. Nearest known compounds by Tanimoto similarity, with clinical-phase flags where an analog reached the clinic.
  3. Selectivity heatmap for multi-target runs, showing predicted margins across a family.
  4. Pocket viewer files with residue listings for the targets in scope.
  5. A short written read — what I would prioritise, and, more usefully, which parts of the output I would not rely on.

What the scoring can and cannot do

Measured on 516 protein–ligand complexes across 179 targets, every structure released after the model's training cutoff. Full workings, including a correction that moved these numbers, are on the validation page.

MeasureResultWhat it means
Top-quintile enrichment1.70×The predicted top fifth of a diverse library holds ~1.7× as many potent compounds as a random fifth.
Top-decile enrichment2.23×51% hit rate against a 23% base rate.
Rank correlationρ = 0.390.63 on clean Kd labels; weaker on assay-dependent IC50.
Within one potency bandρ = −0.01No ordering ability. It will not rank an SAR series, and I will not sell it as though it will.

This is not a hit-finding tool. Run against four real primary screens from PubChem BioAssay — 120,000 compounds at hit rates between 0.21% and 1.10% — it ranks no better than chance (mean ROC AUC 0.499). It is trained on solved protein–ligand complexes, so every example it has seen is a compound that binds; it has no representation of a non-binder, and a screening library is almost entirely non-binders. The measurement is published in full.

What it does is order compounds that are already plausible binders and that span a wide potency range: a hit list from a screen you have already run, a set of analogues bought around a known active, a series where potencies differ by orders of magnitude rather than by twofold. On that job it matches AutoDock Vina at roughly 1/125th the compute, and it runs on targets where docking cannot.

It is not a replacement for FEP or an assay, and the absolute pKd should not be read as a predicted Ki.

How it compares to docking

Because "ρ = 0.39" means little on its own, the same 513 complexes were scored with AutoDock Vina, a reference anyone can download and run. The two tie on rank correlation — a paired bootstrap puts the difference at +0.011 in Vina's favour, 95% CI [−0.076, +0.094]. I am not going to tell you this scoring beats docking, because it does not.

What it does is reach the same ranking quality about 125× faster (0.46 s versus 58 s per compound), pick a better top decile (2.25× versus 1.65× enrichment), and run on targets where docking cannot — docking needs a deposited structure, and 43% of carved pockets do not enclose a site it can use. For a 5,000-compound library that difference is roughly forty minutes against three days.

The comparison was checked two ways before it went up: redocking 507 known ligands into their own receptors reproduced the crystal pose to a median 0.99 Å, and re-running at four times the search effort moved the correlation by 0.028. Neither method can rank within a potency band — docking does not solve that either. Full numbers and failure modes.

Pricing

EngagementScopePrice
Pilot250 compounds, one target, returned in 48 hours.Free
ScreenUp to 5,000 compounds across one to three targets, with everything in the deliverables list above.$1,500
Generation campaignGoal-directed molecule generation with REINVENT4 against your target and property profile, then re-scoring and selectivity profiling of the output.Scope-dependent
Ongoing triageRecurring screening as new target campaigns open. New libraries, not iteration within an existing series.Let's talk

On generation specifically: the reinforcement-learning loop optimises against the affinity model described above, so anything it produces inherits that model's limits. Treat generated structures as hypotheses to filter and inspect, not as a ranked shortlist. That is why it is priced per project rather than per compound.

Your compounds

Structures sent for scoring are held in memory, scored, returned, and dropped. They are not written to disk, not retained, and not shared. The specifics are published rather than summarised. If you need that in a signed agreement first, or would rather send a blinded list with identifiers stripped, both are fine.

Start with the pilot

Email a CSV of up to 250 compounds that are already plausible binders for the target you want them scored against, ideally spanning a wide potency range, and say what you already know about the actives. Reply within 48 hours. If your set is an unfiltered library or a tight analogue series, say so — those are the two cases the scorer handles worst, and I will tell you up front rather than after.

[email protected]

Zach Chiappini · VectaBind · validation · methods