A computational method for virtual screening reported stronger performance near the top of ranked compound lists while changing only a small part of a pretrained model. The method, called PETA, updated approximately 0.03% of the full model parameters for each target pocket and then screened the molecular library.
The findings appear in an arXiv version-1 preprint dated 20 August 2026. The paper asks whether a pretrained virtual-screening model can be specialized to an unseen target pocket at test time without retraining the entire model or relying on binding annotations.
A pocket-specific adjustment
PETA first retrieves a structurally matched reference ligand, constructs pocket-aware negative examples and adapts the ligand encoder before screening the test library. For each target pocket, it updates only LayerNorm parameters in that encoder rather than the full model.
For all evaluated targets, the reference ligand was retrieved from PDBbind and excluded from the screening library. The method generated invalid candidates and divided them into easy negatives, with relatively low pretrained scores, and hard negatives, with scores in the upper half of the generated candidates.
The screening tests used AUROC, BEDROC and enrichment factor. The FEP+ test used pairwise accuracy and Kendall rank correlation K for its fine-grained ligand-ranking evaluation.
Strong early-ranking scores, with a trade-off
On the DUD-E virtual-screening benchmark, PETA reported an AUROC of 82.62, a BEDROC of 53.86, and enrichment factors of 42.34 at the 0.5% cutoff, 34.65 at 1%, and 11.44 at 5%. The paper states that these were the best values among the evaluated methods.
Relative to pretrained DrugCLIP, the reported enrichment factor at the 0.5% cutoff was 4.44 units higher on DUD-E and 3.38 units higher on LIT-PCBA. The authors also report stronger early-enrichment performance than fully trained DrugHash and BindCLIP.
The LIT-PCBA results were more mixed. PETA reported an AUROC of 57.56%, below GNINA at 60.93% and BindCLIP at 59.15%, but it reported the strongest BEDROC and enrichment-factor results at all evaluated cutoffs. Its BEDROC was 8.84, compared with 7.88 for BindCLIP and 6.41 for DrugCLIP.
On the four-target FEP+ ligand-ranking benchmark, PETA reported pairwise accuracy of 70.4%. That was 4.7 percentage points above BindCLIP and 14.0 points above DrugCLIP. Its Kendall rank correlation K was 0.41, compared with 0.31 and 0.14 for the two comparators.
What the results show, and what they do not
An ablation analysis found that the complete PETA configuration achieved the highest BEDROC on both DUD-E and LIT-PCBA. The study found no partial configuration that was consistently best across the two datasets.
The evaluation covered DUD-E and LIT-PCBA for virtual screening and four FEP+ targets for fine-grained ligand ranking. Its reported endpoints were computational ranking metrics, not direct experimental binding or hit-validation outcomes.
The supplied analysis reports no confidence intervals, error bars or significance tests for the benchmark comparisons. The results therefore describe the reported performance on these tests without an inferential estimate of how much the scores might vary.
The method's setup depends on finding a structurally matched reference ligand in PDBbind. The evaluation does not establish how PETA performs when such a reference is unavailable, or whether improved ranking metrics would lead to higher experimental hit rates.
The study does not establish universal superiority beyond the evaluated computational benchmarks. Testing additional target pockets, compound libraries and ligand series, together with prospective experimental validation, would be needed to determine how well the approach transfers beyond these results.
Paper data and sources
Original title: PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening
Authors: Jia-Qi Lin, Yinghua Yao, Chang-Dong Wang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text