New OpenBind-0 model and dataset benchmark co-folding tools in drug discovery
Key Takeaways
- OpenBind-0 (OB0) is a fully open-source co-folding model built on OpenFold3, trained on Protein Data Bank (PDB) data, and optimized for predicting protein–small-molecule complexes.
- Alongside the model, a discovery-relevant benchmark dataset of 717 new ligand-bound structures capturing fragment-to-hit progression across FatA and RdRp targets was released.
- Evaluation across targets shows wide performance variability, with chemical steering during diffusion sampling improving joint accuracy and physical validity.
OpenBind-0 (OB0) is an open-source molecular structure prediction model built on OpenFold3 and trained on PDB data through June 2025. Accompanied by a public dataset of 717 ligand-bound structures, the release provides an open baseline for evaluating protein–small-molecule co-folding tools in drug discovery contexts. Read the full blog post by OpenBind.
OpenBind is the UK’s open science initiative to create the world’s largest dataset of drug-protein interactions led by collaboration among a robust team of researchers, including Frank DiMaio, Mohammed AlQuraishi, Frank von Delft, John Chodera, and Karmen Čondić-Jurkić.
Existing co-folding tools face generalizability challenges on novel targets
Predicting how small-molecule ligands interact with protein targets remains a critical challenge in structure-based drug design. While co-folding architectures have expanded structural prediction capabilities, downstream applications—such as free energy calculations and medicinal chemistry decisions—require both precise placement and physically plausible ligand geometries. Furthermore, models often encounter performance drops on flexible protein systems or targets that are highly dissimilar to their pre-training distributions.
Inference-time chemical steering enhances structural validity
OB0 uses an architecture based on OpenFold3 and was selected to optimize protein-ligand prediction while maintaining performance on other molecular modalities. To address unphysical ligand geometries, the developers incorporated inference-time chemical steering during diffusion sampling, following approaches used in Boltz and Protenix. The model is released under an Apache 2.0 license, making code, weights, data, and training recipes publicly accessible.
Model accuracy varies substantially across benchmark systems
On post-cutoff benchmark complexes, OB0 shows performance comparable to models like AlphaFold3 and Protenix-v1-20250630. Implementing chemical steering increased the joint success rate—defined by pose accuracy and PoseBusters physical validity—from 48% to 61%. When tested on newly released fragment-to-lead targets, performance varied widely:
- EV-A71 2A protease: OB0 achieved a top-25 success rate of 92.2% and top-1 success rate of 73.8%.
- FatA thioesterase: While OB0 doubled OpenFold3-preview2 on FatA (28.2% vs. 14.5%), Protenix-v1-20250630 achieved the highest top-25 success rate on this specific target at 39.4%.
- RdRp (DENV-2 and ZIKV): Co-folding models achieved top-25 success rates below 10% (7.7% for DENV-2 and 6.7% for ZIKV).
Target-specific evaluations confirm performance dependencies
Model predictions were validated across 547 fragment binding events and 170 hit binding events across FatA and RdRp campaigns. Failure mode analysis revealed that for RdRp, models frequently failed to locate the correct binding pocket or generated incorrect pocket conformations due to high target flexibility and low training set similarity. Experiments fine-tuning models on target-specific fragment structures showed variable gains, demonstrating that fine-tuning alone does not resolve predictive gaps for all complex targets.
Open model establishes benchmarks for future model development
As new experimental data is generated, OB0 provides a reference baseline to evaluate whether incorporating additional structural data improves generalizability and active learning workflows. The authors hope the community will find value in the model and continue to find ways to improve model performance.
