Drug Repurposing & Reformulation Leaderboard

targettrajectorycandidateslaunchfield

Public-100 results · three-seed mean ± sample SD

DRR overall ranking

No-env and with-Apodex Env entries are ranked together by normalized continuous score. The table also reports categorical accuracy and raw continuous score for every submitted model and setting.

Normalized continuous score anchors mean quality at random = 0 and perfect = 100; categorical accuracy measures exact-tier matches (or at-least-attained matches for lower bounds); raw continuous score is the unnormalized mean quality under the same loss.

No-env

with-Apodex Env

RankModelSettingNormalized continuous scoreCategorical accuracyRaw continuous score
1Claude Opus 5with-Apodex Env66.6287± 1.17080.8867± 0.00580.9564± 0.0015
2Claude Opus 5No-env64.1813± 1.08790.8767± 0.00580.9532± 0.0014
3GPT-5.6-solwith-Apodex Env61.8104± 3.91500.8500± 0.01730.9501± 0.0051
4GPT-5.5with-Apodex Env56.3038± 1.56240.8367± 0.01150.9429± 0.0020
5Kimi K3with-Apodex Env54.2388± 3.09950.8200± 0.01730.9402± 0.0041
6GPT-5.6-solNo-env54.2133± 2.19720.8067± 0.01150.9401± 0.0029
7GPT-5.5No-env53.7799± 3.22340.8100± 0.01730.9396± 0.0042
8Qwen3.8 Maxwith-Apodex Env52.1483± 2.98310.8000± 0.02000.9374± 0.0039
9Apodex 1.1with-Apodex Env46.4632± 1.80830.7933± 0.00580.9300± 0.0024
10Kimi K3No-env44.7042± 4.62950.7500± 0.02650.9277± 0.0061
11DeepSeek V4 ProNo-env44.6787± 2.56220.7667± 0.01530.9277± 0.0034
12Apodex 1.1No-env44.4747± 1.67740.7700± 0.01000.9274± 0.0022
13DeepSeek V4 Prowith-Apodex Env39.2230± 0.76990.7600± 0.03000.9205± 0.0010
14GLM-5.2No-env37.7444± 4.73750.7467± 0.01150.9186± 0.0062
15GLM-5.2with-Apodex Env36.5088± 4.86330.7425± 0.03030.9170± 0.0064
16Qwen3.8 MaxNo-env24.0543± 3.49560.6400± 0.02650.9007± 0.0046

GLM-5.2 with-Apodex Env: 299 valid submissions; the failed row is omitted. All other entries: 300/300. Reasoning effort default to medium.

Interpretation boundary. These are descriptive Public-100 development-set results from three seeds, not definitive estimates of hidden-test performance.