AAV Capsid Leaderboard
Model rankings
Standings across the four AAV tasks. Each cell reports a result score (the task’s headline metric) alongside its TRACES / HDS6 process score (0–4). Every score is deterministic, auto-computed and leakage-audited. Submit your model to be added.
Operator disclosure. TRACES is built and operated by Apodex, which also submits its own solvers. Apodex entries are marked.
How these results were produced. All results on this page were produced independently by Apodex. No system shown here was submitted by its developer, and no developer participated in, configured, or reviewed these evaluations.
Conflict of interest. Apodex designed, operates, and scores this benchmark, and also develops the apodex model family. apodex-1.1 and apodex-1.0 are shown for reference only and are not ranked. They were run under the same harness, budgets, and hidden verifiers as every other system; they received no access to ground truth, verifiers, or task instances beyond what any evaluated system receives; and they were not trained or fine-tuned on any TRACES environment data, task instance, or answer key.
v1.1 · 11 systems evaluated · 9 ranked
Overview
Ranked by the mean of the first three tasks (Viability, Tropism, Structure). Each cell shows the result score with its TRACES/HDS6 process score (0–4) beside it. Design is reported but excluded from the ranking, since several systems were not run on it (—).
| # | Solver | Viability | Tropism | Structure | Design | Mean · T1–T3 |
|---|---|---|---|---|---|---|
| 1 | kimi-k3 | 0.9363.49 | 0.6493.61 | 0.6793.68 | 0.1943.59 | 0.755 |
| 2 | claude-opus-5 | 0.9363.46 | 0.6533.55 | 0.6683.61 | — | 0.752 |
| 3 | glm-5.2 | 0.9323.27 | 0.6443.76 | 0.6653.39 | 0.1913.32 | 0.747 |
| 4 | claude-opus-4-8 | 0.9313.65 | 0.6353.30 | 0.6573.36 | — | 0.741 |
| 5 | Claude Code (opus 4.8) | 0.921— | 0.633— | 0.595— | — | 0.716 |
| 6 | deepseek-v4-pro | 0.9183.37 | 0.6132.97 | 0.6102.19 | 0.1292.64 | 0.714 |
| 7 | claude-sonnet-4-6 | 0.9223.45 | 0.6282.77 | 0.5752.73 | 0.0942.97 | 0.709 |
| 8 | gpt-5.5 | 0.9092.99 | 0.6003.57 | 0.5441.22 | 0.1422.24 | 0.684 |
| 9 | qwen3.5-397b | 0.8963.27 | 0.6052.93 | 0.5431.97 | 0.0852.21 | 0.682 |
| Reference — benchmark operator's own models (not ranked) | ||||||
| apodex-1.1 | 0.9043.36 | 0.6353.37 | 0.6493.52 | 0.1803.23 | 0.729 | |
| apodex-1.0 | 0.8843.23 | 0.5182.90 | 0.5261.84 | 0.0722.34 | 0.643 | |
| Reference — published SOTA envelope | ||||||
| published SOTA | 0.878 | 0.622 | 0.605 | 0.116 | 0.702 | |
published SOTA is, per task, the single strongest reference method on that task, scored by its mean over instances. The specific method and citation for each task are given on that task’s tab.
Viability
Out-of-distribution viability discrimination (mean per-distance AUROC; chance 0.5). Multi-mutation = count extrapolation (Bryant 2021); Single-mutation = substitution→indel extrapolation (Ogden 2019).
| # | Solver | Multi-mutation | Single-mutation | Mean |
|---|---|---|---|---|
| 1 | kimi-k3 | 0.9623.48 | 0.9103.51 | 0.936 |
| 2 | claude-opus-5 | 0.9593.43 | 0.9133.49 | 0.936 |
| 3 | glm-5.2 | 0.9603.37 | 0.9053.18 | 0.932 |
| 4 | claude-opus-4-8 | 0.9513.62 | 0.9123.68 | 0.931 |
| 5 | claude-sonnet-4-6 | 0.9523.36 | 0.8933.55 | 0.922 |
| 6 | Claude Code (opus 4.8) | 0.950— | 0.891— | 0.921 |
| 7 | deepseek-v4-pro | 0.9413.33 | 0.8943.41 | 0.918 |
| 8 | gpt-5.5 | 0.9272.85 | 0.8913.14 | 0.909 |
| 9 | qwen3.5-397b | 0.8903.24 | 0.9023.31 | 0.896 |
| Reference — benchmark operator's own models (not ranked) | ||||
| apodex-1.1 | 0.9313.27 | 0.8773.44 | 0.904 | |
| apodex-1.0 | 0.9053.17 | 0.8633.30 | 0.884 | |
| Reference — baselines | ||||
| CAP-PLM (published SOTA) | 0.897 | 0.859 | 0.878 | |
| ALICE-GB | 0.885 | 0.696 | 0.791 | |
| augmented ridge | 0.874 | 0.848 | 0.861 | |
| logistic additive | 0.888 | 0.852 | 0.870 | |
Sources. CAP-PLM (published SOTA), a protein-language-model fitness predictor — Wu et al., “Prediction of adeno-associated virus fitness with a protein language-based machine learning model,” Human Gene Therapy 36:823–829, 2024. ALICE-GB — Guo et al., “Mapping AAV capsid sequences to functions through function-guided in silico evolution,” Cell Press Blue 1(3), 2026. augmented ridge — Hsu et al., “Learning protein fitness models from evolutionary and assay-labeled data,” Nature Biotechnology 40:1114–1122, 2022. logistic additive, a per-site additive logistic model — Bryant et al., “Deep diversification of an AAV capsid protein by machine learning,” Nature Biotechnology 39:691–696, 2021. Task data: Bryant et al. 2021 above (multi-mutation); Ogden et al., “Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design,” Science 366:1139–1143, 2019 (single-mutation).
Tropism
7-mer enrichment prediction (Spearman / Pearson). A mouse multi-organ tropism · B mouse→macaque transfer · C human in-vitro cell assays. F4F is the Fit4Function reproduction.
| # | Solver | A · Mouse tropism | B · Cross-species | C · In-vitro cell | Mean |
|---|---|---|---|---|---|
| 1 | claude-opus-5 | 0.5623.65 | 0.6383.10 | 0.7603.90 | 0.653 |
| 2 | kimi-k3 | 0.5593.70 | 0.6373.63 | 0.7513.51 | 0.649 |
| 3 | glm-5.2 | 0.5533.73 | 0.6263.84 | 0.7523.72 | 0.644 |
| 4 | claude-opus-4-8 | 0.5483.70 | 0.6193.40 | 0.7402.80 | 0.635 |
| 5 | Claude Code (opus 4.8) | 0.545— | 0.617— | 0.738— | 0.633 |
| 6 | claude-sonnet-4-6 | 0.5493.42 | 0.6101.91 | 0.7252.97 | 0.628 |
| 7 | deepseek-v4-pro | 0.5083.23 | 0.6022.81 | 0.7282.88 | 0.613 |
| 8 | qwen3.5-397b | 0.5002.98 | 0.6002.84 | 0.7162.96 | 0.605 |
| 9 | gpt-5.5 | 0.5343.56 | 0.5753.75 | 0.6913.39 | 0.600 |
| Reference — benchmark operator's own models (not ranked) | |||||
| apodex-1.1 | 0.5393.16 | 0.6213.57 | 0.7463.37 | 0.635 | |
| apodex-1.0 | 0.5252.43 | 0.5402.64 | 0.4893.62 | 0.518 | |
| Reference — baseline | |||||
| F4F (SOTA) | 0.532 | 0.614 | 0.719 | 0.622 | |
Sources. F4F (published SOTA), our reproduction of the Fit4Function production model — Eid et al., “Systematic multi-trait AAV capsid engineering for efficient gene delivery,” Nature Communications 15:6602, 2024. Task data: the same screens.
Structure
Sequence→3D structure. Loop lDDT on surface loops · Assembly 60-mer geometry × stability · Complex DockQ + interface Fnat. All in [0,1], higher is better. The published SOTA is AlphaFold 3 with template-based symmetry expansion, the strongest single method on the task.
| # | Solver | Loop | Assembly | Complex | Mean |
|---|---|---|---|---|---|
| 1 | kimi-k3 | 0.9243.56 | 0.7543.60 | 0.3593.89 | 0.679 |
| 2 | claude-opus-5 | 0.9133.43 | 0.7683.79 | 0.3223.60 | 0.668 |
| 3 | glm-5.2 | 0.9223.31 | 0.7363.53 | 0.3363.33 | 0.665 |
| 4 | claude-opus-4-8 | 0.9273.28 | 0.7693.82 | 0.2752.99 | 0.657 |
| 5 | deepseek-v4-pro | 0.9272.50 | 0.7492.94 | 0.1541.13 | 0.610 |
| 6 | Claude Code (opus 4.8) | 0.889— | 0.660— | 0.236— | 0.595 |
| 7 | claude-sonnet-4-6 | 0.9152.69 | 0.6623.42 | 0.1482.09 | 0.575 |
| 8 | gpt-5.5 ‡ | 0.9271.38 | 0.4900.94 | 0.2161.35 | 0.544 |
| 9 | qwen3.5-397b | 0.9271.29 | 0.5553.18 | 0.1481.45 | 0.543 |
| Reference — benchmark operator's own models (not ranked) | |||||
| apodex-1.1 | 0.9283.47 | 0.7423.60 | 0.2773.49 | 0.649 | |
| apodex-1.0 | 0.9271.53 | 0.5043.06 | 0.1480.93 | 0.526 | |
| Reference — baselines (partial coverage) | |||||
| AlphaFold 3 + symmetry expansion § (published SOTA) | 0.927 | 0.740 | 0.148 | 0.605 | |
| Boltz-1 + symmetry expansion § | 0.867 | 0.695 | 0.074 | 0.545 | |
| AlphaFold 2 + symmetry expansion § | 0.901 | 0.732 | 0.073 | 0.569 | |
| CapBuild | — | 0.593 | — | — | |
| Template docking | — | — | 0.252 | — | |
§ Folding engines cannot predict the 60-mer shell directly. For the Assembly level we fold the single subunit with the named engine, then expand it onto the icosahedral operators of a fixed AAV capsid template.
‡ gpt-5.5 has low HDS6 scores across all three instances, it consistently skips the self-validation step the task requires and submits the cheapest answer instead. Loop: folds every target with one default engine without checking or comparison. Assembly: copies the highest sequence-identity template’s coordinates verbatim for every target. Complex: hardcodes outcome with severe collapses.
Sources. AlphaFold 3 (published SOTA) — Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3,” Nature 630:493–500, 2024. Boltz-1 — Wohlwend et al., “Boltz-1: democratizing biomolecular interaction modeling,” bioRxiv, 2025. AlphaFold 2 — Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature 596:583–589, 2021. CapBuild — Loeb et al., “Complete neutralizing antibody evasion by serodivergent non-mammalian AAVs enables gene therapy redosing,” Cell Reports Medicine 6, 2025. template docking, a superposition-based docking baseline built by us on PDB templates (this work). Interface quality uses DockQ — Basu & Wallner, “DockQ: a quality measure for protein–protein docking models,” PLoS ONE 11, 2016.
Design
Generative design: 1000 novel AAV9 7-mers per run under a metered oracle budget and a hard viability gate. Track A (organs) = selective pass rate averaged over the top-1%/2%/5% thresholds · Track B (receptors) = specific pass rate. The published SOTA is AAVDiff, the strongest of the three reference generators on the task. All three are our own reproductions under the identical episode budget (1000 submitted designs, 100 scoring queries), differing only in the generator.
| # | Solver | Track A · tissue tropism | Track B · receptor | Mean | brain | spinalcord | heart | LY6A | LY6C1 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | kimi-k3 | 0.1953.28 | 0.3023.70 | 0.2233.75 | 0.1323.53 | 0.1173.71 | 0.194 | ||
| 2 | glm-5.2 | 0.2493.73 | 0.2653.57 | 0.2423.42 | 0.0822.62 | 0.1173.28 | 0.191 | ||
| 3 | gpt-5.5 | 0.1171.94 | 0.2652.33 | 0.1922.55 | 0.0732.53 | 0.0611.85 | 0.142 | ||
| 4 | deepseek-v4-pro | 0.1122.91 | 0.2573.03 | 0.2043.06 | 0.0381.61 | 0.0322.58 | 0.129 | ||
| 5 | claude-sonnet-4-6 | 0.0782.95 | 0.2013.02 | 0.1083.00 | 0.0413.35 | 0.0432.55 | 0.094 | ||
| 6 | qwen3.5-397b | 0.0662.58 | 0.1751.67 | 0.0992.28 | 0.0542.25 | 0.0302.25 | 0.085 | ||
| 7 | Claude Code (sonnet 4-6) | 0.077— | 0.187— | 0.100— | 0.035— | 0.002— | 0.080 | ||
| Reference — benchmark operator's own models (not ranked) | |||||||||
| apodex-1.1 | 0.2593.26 | 0.2653.50 | 0.2203.81 | 0.0353.21 | 0.1222.36 | 0.180 | |||
| apodex-1.0 | 0.0252.37 | 0.2282.94 | 0.0622.78 | 0.0391.67 | 0.0081.92 | 0.072 | |||
| Reference — baselines | |||||||||
| AAVGen | 0.104 | 0.247 | 0.160 | 0.026 | 0.009 | 0.109 | |||
| AAVDiff (published SOTA) | 0.113 | 0.267 | 0.177 | 0.021 | 0.003 | 0.116 | |||
| ALICE | 0.093 | 0.247 | 0.168 | 0.024 | 0.019 | 0.110 | |||
Sources. AAVGen — Ghaffarzadeh-Esfahani & Gheisari, “AAVGen: precision engineering of adeno-associated viral capsids for renal selective targeting,” arXiv 2602.18915, 2026. AAVDiff — Liu et al., “AAVDiff: experimental validation of enhanced viability and diversity in recombinant adeno-associated virus (AAV) capsids through diffusion generation,” arXiv 2404.10573, 2024. ALICE — Guo et al., “Mapping AAV capsid sequences to functions through function-guided in silico evolution,” Cell Press Blue 1(3), 2026. Target data: Eid et al., “Systematic multi-trait AAV capsid engineering for efficient gene delivery,” Nature Communications 15:6602, 2024 (tissue selectivity); Huang et al., “Targeting AAV vectors to the central nervous system by engineering capsid–receptor interactions that enable crossing of the blood–brain barrier,” PLOS Biology 21, 2023 (LY6A / LY6C1).
TRACES · process
Process capability (TRACES / HDS6), 0–4, from the blind trajectory judge — scored independently of task outcome. Process mean is over available tasks; Claude Code omitted (no judged trajectory).
| # | Solver | Viability | Tropism | Structure | Design | Process mean |
|---|---|---|---|---|---|---|
| 1 | kimi-k3 | 3.49 | 3.61 | 3.68 | 3.59 | 3.59 |
| 2 | claude-opus-5 | 3.46 | 3.55 | 3.61 | — | 3.54 |
| 3 | glm-5.2 | 3.27 | 3.76 | 3.39 | 3.32 | 3.44 |
| 4 | claude-opus-4-8 | 3.65 | 3.30 | 3.36 | — | 3.44 |
| 5 | claude-sonnet-4-6 | 3.45 | 2.77 | 2.73 | 2.97 | 2.98 |
| 6 | deepseek-v4-pro | 3.37 | 2.97 | 2.19 | 2.64 | 2.79 |
| 7 | qwen3.5-397b | 3.27 | 2.93 | 1.97 | 2.21 | 2.60 |
| 8 | gpt-5.5 | 2.99 | 3.57 | 1.22 | 2.24 | 2.51 |
| Reference — benchmark operator's own models (not ranked) | ||||||
| apodex-1.1 | 3.36 | 3.37 | 3.52 | 3.23 | 3.37 | |
| apodex-1.0 | 3.23 | 2.90 | 1.84 | 2.34 | 2.58 | |
apodex models shown for reference · not ranked
score = headline metric · small
best per column shown in accent
Third-party system names and trademarks are the property of their respective owners. Inclusion does not indicate endorsement of, or any affiliation with, Apodex.
FULL RESULT MATRIX
Full result matrix across models and harnesses
The complete solver × track matrix for the four AAV environments, copied from the harness × model result appendix with no rewriting. Rows are solvers (harness · model); columns are that environment’s own tracks or settings. Every environment keeps its own metric and its own table — nothing is merged and nothing is averaged across environments.
The grey italic rows at the foot of each table are published SOTA, method-baseline and pass-threshold references; they do not compete for “best in column”. Claude Code only runs Opus-4.8 and Codex CLI only runs GPT-5.6-sol — by design, not a missing run.
LLM solver (harness × model)
Single-model harness (Claude Code / Codex CLI)
ApodexHarness (in-house)
published SOTA / method baseline (ref)
best in column
OUTCOME MARKERS
Refused
safety or content-policy refusal
Fault
harness bug, API, compatibility or quota fault
Gate not met
a hard gate was missed (loss, speed, throughput, no-regression) or training failed
Not submitted
nothing submitted, or not submitted to protocol
Timed out
budget exhausted; nothing submitted inside the limit
Note
a remark on the score, not a failure
Grounding
stopped to ask the user instead of acting
Out of sweep
not in the sweep this metric was computed from
N/A
route not applicable
A number with a superscript letter (for example 0.000a) means the run happened and either scored a genuine zero or needs a note. — means the configuration was never run. A marker in place of a number means it ran and produced nothing scorable. Heat bars scale to the maximum of their own table and are comparable inside it only.
Every number above comes from the harness × model result-matrix appendix, copied cell for cell with no rewriting. — means the configuration was never run; a marker in place of a number means it ran and produced nothing scorable; a number with a superscript letter is a real score that needs a note. Scores from different environments are not comparable.