Drug Repurposing & Reformulation Leaderboard
Public-100 results · three-seed mean ± sample SD
DRR overall ranking
No-env and with-Apodex Env entries are ranked together by normalized continuous score. The table also reports categorical accuracy and raw continuous score for every submitted model and setting.
Normalized continuous score anchors mean quality at random = 0 and perfect = 100; categorical accuracy measures exact-tier matches (or at-least-attained matches for lower bounds); raw continuous score is the unnormalized mean quality under the same loss.
No-env
with-Apodex Env
| Rank | Model | Setting | Normalized continuous score | Categorical accuracy | Raw continuous score |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 | with-Apodex Env | 66.6287± 1.1708 | 0.8867± 0.0058 | 0.9564± 0.0015 |
| 2 | Claude Opus 5 | No-env | 64.1813± 1.0879 | 0.8767± 0.0058 | 0.9532± 0.0014 |
| 3 | GPT-5.6-sol | with-Apodex Env | 61.8104± 3.9150 | 0.8500± 0.0173 | 0.9501± 0.0051 |
| 4 | GPT-5.5 | with-Apodex Env | 56.3038± 1.5624 | 0.8367± 0.0115 | 0.9429± 0.0020 |
| 5 | Kimi K3 | with-Apodex Env | 54.2388± 3.0995 | 0.8200± 0.0173 | 0.9402± 0.0041 |
| 6 | GPT-5.6-sol | No-env | 54.2133± 2.1972 | 0.8067± 0.0115 | 0.9401± 0.0029 |
| 7 | GPT-5.5 | No-env | 53.7799± 3.2234 | 0.8100± 0.0173 | 0.9396± 0.0042 |
| 8 | Qwen3.8 Max | with-Apodex Env | 52.1483± 2.9831 | 0.8000± 0.0200 | 0.9374± 0.0039 |
| 9 | Apodex 1.1 | with-Apodex Env | 46.4632± 1.8083 | 0.7933± 0.0058 | 0.9300± 0.0024 |
| 10 | Kimi K3 | No-env | 44.7042± 4.6295 | 0.7500± 0.0265 | 0.9277± 0.0061 |
| 11 | DeepSeek V4 Pro | No-env | 44.6787± 2.5622 | 0.7667± 0.0153 | 0.9277± 0.0034 |
| 12 | Apodex 1.1 | No-env | 44.4747± 1.6774 | 0.7700± 0.0100 | 0.9274± 0.0022 |
| 13 | DeepSeek V4 Pro | with-Apodex Env | 39.2230± 0.7699 | 0.7600± 0.0300 | 0.9205± 0.0010 |
| 14 | GLM-5.2 | No-env | 37.7444± 4.7375 | 0.7467± 0.0115 | 0.9186± 0.0062 |
| 15 | GLM-5.2 | with-Apodex Env | 36.5088± 4.8633 | 0.7425± 0.0303 | 0.9170± 0.0064 |
| 16 | Qwen3.8 Max | No-env | 24.0543± 3.4956 | 0.6400± 0.0265 | 0.9007± 0.0046 |
GLM-5.2 with-Apodex Env: 299 valid submissions; the failed row is omitted. All other entries: 300/300. Reasoning effort default to medium.
Interpretation boundary. These are descriptive Public-100 development-set results from three seeds, not definitive estimates of hidden-test performance.