AAV Capsid Leaderboard

obstacletargetdetourblocked

Model rankings

Standings across the four AAV tasks. Each cell reports a result score (the tasks headline metric) alongside its TRACES / HDS6 process score (04). Every score is deterministic, auto-computed and leakage-audited. Submit your model to be added.

Operator disclosure. TRACES is built and operated by Apodex, which also submits its own solvers. Apodex entries are marked.

How these results were produced. All results on this page were produced independently by Apodex. No system shown here was submitted by its developer, and no developer participated in, configured, or reviewed these evaluations.

Conflict of interest. Apodex designed, operates, and scores this benchmark, and also develops the apodex model family. apodex-1.1 and apodex-1.0 are shown for reference only and are not ranked. They were run under the same harness, budgets, and hidden verifiers as every other system; they received no access to ground truth, verifiers, or task instances beyond what any evaluated system receives; and they were not trained or fine-tuned on any TRACES environment data, task instance, or answer key.

v1.1 · 11 systems evaluated · 9 ranked

Overview

Ranked by the mean of the first three tasks (Viability, Tropism, Structure). Each cell shows the result score with its TRACES/HDS6 process score (0–4) beside it. Design is reported but excluded from the ranking, since several systems were not run on it (—).

#SolverViabilityTropismStructureDesignMean · T1–T3
1kimi-k30.9363.490.6493.610.6793.680.1943.590.755
2claude-opus-50.9363.460.6533.550.6683.610.752
3glm-5.20.9323.270.6443.760.6653.390.1913.320.747
4claude-opus-4-80.9313.650.6353.300.6573.360.741
5Claude Code (opus 4.8)0.9210.6330.5950.716
6deepseek-v4-pro0.9183.370.6132.970.6102.190.1292.640.714
7claude-sonnet-4-60.9223.450.6282.770.5752.730.0942.970.709
8gpt-5.50.9092.990.6003.570.5441.220.1422.240.684
9qwen3.5-397b0.8963.270.6052.930.5431.970.0852.210.682
Reference — benchmark operator's own models (not ranked)
apodex-1.10.9043.360.6353.370.6493.520.1803.230.729
apodex-1.00.8843.230.5182.900.5261.840.0722.340.643
Reference — published SOTA envelope
published SOTA0.8780.6220.6050.1160.702

published SOTA is, per task, the single strongest reference method on that task, scored by its mean over instances. The specific method and citation for each task are given on that task’s tab.

Viability

Out-of-distribution viability discrimination (mean per-distance AUROC; chance 0.5). Multi-mutation = count extrapolation (Bryant 2021); Single-mutation = substitution→indel extrapolation (Ogden 2019).

#SolverMulti-mutationSingle-mutationMean
1kimi-k30.9623.480.9103.510.936
2claude-opus-50.9593.430.9133.490.936
3glm-5.20.9603.370.9053.180.932
4claude-opus-4-80.9513.620.9123.680.931
5claude-sonnet-4-60.9523.360.8933.550.922
6Claude Code (opus 4.8)0.9500.8910.921
7deepseek-v4-pro0.9413.330.8943.410.918
8gpt-5.50.9272.850.8913.140.909
9qwen3.5-397b0.8903.240.9023.310.896
Reference — benchmark operator's own models (not ranked)
apodex-1.10.9313.270.8773.440.904
apodex-1.00.9053.170.8633.300.884
Reference — baselines
CAP-PLM (published SOTA)0.8970.8590.878
ALICE-GB0.8850.6960.791
augmented ridge0.8740.8480.861
logistic additive0.8880.8520.870

Sources. CAP-PLM (published SOTA), a protein-language-model fitness predictor — Wu et al., “Prediction of adeno-associated virus fitness with a protein language-based machine learning model,” Human Gene Therapy 36:823–829, 2024. ALICE-GB — Guo et al., “Mapping AAV capsid sequences to functions through function-guided in silico evolution,” Cell Press Blue 1(3), 2026. augmented ridge — Hsu et al., “Learning protein fitness models from evolutionary and assay-labeled data,” Nature Biotechnology 40:1114–1122, 2022. logistic additive, a per-site additive logistic model — Bryant et al., “Deep diversification of an AAV capsid protein by machine learning,” Nature Biotechnology 39:691–696, 2021. Task data: Bryant et al. 2021 above (multi-mutation); Ogden et al., “Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design,” Science 366:1139–1143, 2019 (single-mutation).

Tropism

7-mer enrichment prediction (Spearman / Pearson). A mouse multi-organ tropism · B mouse→macaque transfer · C human in-vitro cell assays. F4F is the Fit4Function reproduction.

#SolverA · Mouse tropismB · Cross-speciesC · In-vitro cellMean
1claude-opus-50.5623.650.6383.100.7603.900.653
2kimi-k30.5593.700.6373.630.7513.510.649
3glm-5.20.5533.730.6263.840.7523.720.644
4claude-opus-4-80.5483.700.6193.400.7402.800.635
5Claude Code (opus 4.8)0.5450.6170.7380.633
6claude-sonnet-4-60.5493.420.6101.910.7252.970.628
7deepseek-v4-pro0.5083.230.6022.810.7282.880.613
8qwen3.5-397b0.5002.980.6002.840.7162.960.605
9gpt-5.50.5343.560.5753.750.6913.390.600
Reference — benchmark operator's own models (not ranked)
apodex-1.10.5393.160.6213.570.7463.370.635
apodex-1.00.5252.430.5402.640.4893.620.518
Reference — baseline
F4F (SOTA)0.5320.6140.7190.622

Sources. F4F (published SOTA), our reproduction of the Fit4Function production model — Eid et al., “Systematic multi-trait AAV capsid engineering for efficient gene delivery,” Nature Communications 15:6602, 2024. Task data: the same screens.

Structure

Sequence→3D structure. Loop lDDT on surface loops · Assembly 60-mer geometry × stability · Complex DockQ + interface Fnat. All in [0,1], higher is better. The published SOTA is AlphaFold 3 with template-based symmetry expansion, the strongest single method on the task.

#SolverLoopAssemblyComplexMean
1kimi-k30.9243.560.7543.600.3593.890.679
2claude-opus-50.9133.430.7683.790.3223.600.668
3glm-5.20.9223.310.7363.530.3363.330.665
4claude-opus-4-80.9273.280.7693.820.2752.990.657
5deepseek-v4-pro0.9272.500.7492.940.1541.130.610
6Claude Code (opus 4.8)0.8890.6600.2360.595
7claude-sonnet-4-60.9152.690.6623.420.1482.090.575
8gpt-5.5 ‡0.9271.380.4900.940.2161.350.544
9qwen3.5-397b0.9271.290.5553.180.1481.450.543
Reference — benchmark operator's own models (not ranked)
apodex-1.10.9283.470.7423.600.2773.490.649
apodex-1.00.9271.530.5043.060.1480.930.526
Reference — baselines (partial coverage)
AlphaFold 3 + symmetry expansion § (published SOTA)0.9270.7400.1480.605
Boltz-1 + symmetry expansion §0.8670.6950.0740.545
AlphaFold 2 + symmetry expansion §0.9010.7320.0730.569
CapBuild0.593
Template docking0.252

§ Folding engines cannot predict the 60-mer shell directly. For the Assembly level we fold the single subunit with the named engine, then expand it onto the icosahedral operators of a fixed AAV capsid template.

gpt-5.5 has low HDS6 scores across all three instances, it consistently skips the self-validation step the task requires and submits the cheapest answer instead. Loop: folds every target with one default engine without checking or comparison. Assembly: copies the highest sequence-identity template’s coordinates verbatim for every target. Complex: hardcodes outcome with severe collapses.

Sources. AlphaFold 3 (published SOTA) — Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3,” Nature 630:493–500, 2024. Boltz-1 — Wohlwend et al., “Boltz-1: democratizing biomolecular interaction modeling,” bioRxiv, 2025. AlphaFold 2 — Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature 596:583–589, 2021. CapBuild — Loeb et al., “Complete neutralizing antibody evasion by serodivergent non-mammalian AAVs enables gene therapy redosing,” Cell Reports Medicine 6, 2025. template docking, a superposition-based docking baseline built by us on PDB templates (this work). Interface quality uses DockQ — Basu & Wallner, “DockQ: a quality measure for protein–protein docking models,” PLoS ONE 11, 2016.

Design

Generative design: 1000 novel AAV9 7-mers per run under a metered oracle budget and a hard viability gate. Track A (organs) = selective pass rate averaged over the top-1%/2%/5% thresholds · Track B (receptors) = specific pass rate. The published SOTA is AAVDiff, the strongest of the three reference generators on the task. All three are our own reproductions under the identical episode budget (1000 submitted designs, 100 scoring queries), differing only in the generator.

#SolverTrack A · tissue tropismTrack B · receptorMeanbrainspinalcordheartLY6ALY6C1
1kimi-k30.1953.280.3023.700.2233.750.1323.530.1173.710.194
2glm-5.20.2493.730.2653.570.2423.420.0822.620.1173.280.191
3gpt-5.50.1171.940.2652.330.1922.550.0732.530.0611.850.142
4deepseek-v4-pro0.1122.910.2573.030.2043.060.0381.610.0322.580.129
5claude-sonnet-4-60.0782.950.2013.020.1083.000.0413.350.0432.550.094
6qwen3.5-397b0.0662.580.1751.670.0992.280.0542.250.0302.250.085
7Claude Code (sonnet 4-6)0.0770.1870.1000.0350.0020.080
Reference — benchmark operator's own models (not ranked)
apodex-1.10.2593.260.2653.500.2203.810.0353.210.1222.360.180
apodex-1.00.0252.370.2282.940.0622.780.0391.670.0081.920.072
Reference — baselines
AAVGen0.1040.2470.1600.0260.0090.109
AAVDiff (published SOTA)0.1130.2670.1770.0210.0030.116
ALICE0.0930.2470.1680.0240.0190.110

Sources. AAVGen — Ghaffarzadeh-Esfahani & Gheisari, “AAVGen: precision engineering of adeno-associated viral capsids for renal selective targeting,” arXiv 2602.18915, 2026. AAVDiff — Liu et al., “AAVDiff: experimental validation of enhanced viability and diversity in recombinant adeno-associated virus (AAV) capsids through diffusion generation,” arXiv 2404.10573, 2024. ALICE — Guo et al., “Mapping AAV capsid sequences to functions through function-guided in silico evolution,” Cell Press Blue 1(3), 2026. Target data: Eid et al., “Systematic multi-trait AAV capsid engineering for efficient gene delivery,” Nature Communications 15:6602, 2024 (tissue selectivity); Huang et al., “Targeting AAV vectors to the central nervous system by engineering capsid–receptor interactions that enable crossing of the blood–brain barrier,” PLOS Biology 21, 2023 (LY6A / LY6C1).

TRACES · process

Process capability (TRACES / HDS6), 0–4, from the blind trajectory judge — scored independently of task outcome. Process mean is over available tasks; Claude Code omitted (no judged trajectory).

#SolverViabilityTropismStructureDesignProcess mean
1kimi-k33.493.613.683.593.59
2claude-opus-53.463.553.613.54
3glm-5.23.273.763.393.323.44
4claude-opus-4-83.653.303.363.44
5claude-sonnet-4-63.452.772.732.972.98
6deepseek-v4-pro3.372.972.192.642.79
7qwen3.5-397b3.272.931.972.212.60
8gpt-5.52.993.571.222.242.51
Reference — benchmark operator's own models (not ranked)
apodex-1.13.363.373.523.233.37
apodex-1.03.232.901.842.342.58

apodex models shown for reference · not ranked

score = headline metric · small

best per column shown in accent

Third-party system names and trademarks are the property of their respective owners. Inclusion does not indicate endorsement of, or any affiliation with, Apodex.

FULL RESULT MATRIX

Full result matrix across models and harnesses

The complete solver × track matrix for the four AAV environments, copied from the harness × model result appendix with no rewriting. Rows are solvers (harness · model); columns are that environments own tracks or settings. Every environment keeps its own metric and its own table nothing is merged and nothing is averaged across environments.

The grey italic rows at the foot of each table are published SOTA, method-baseline and pass-threshold references; they do not compete for “best in column”. Claude Code only runs Opus-4.8 and Codex CLI only runs GPT-5.6-sol — by design, not a missing run.

LLM solver (harness × model)

Single-model harness (Claude Code / Codex CLI)

ApodexHarness (in-house)

published SOTA / method baseline (ref)

best in column

OUTCOME MARKERS

Refused

safety or content-policy refusal

Fault

harness bug, API, compatibility or quota fault

Gate not met

a hard gate was missed (loss, speed, throughput, no-regression) or training failed

Not submitted

nothing submitted, or not submitted to protocol

Timed out

budget exhausted; nothing submitted inside the limit

Note

a remark on the score, not a failure

Grounding

stopped to ask the user instead of acting

Out of sweep

not in the sweep this metric was computed from

N/A

route not applicable

A number with a superscript letter (for example 0.000a) means the run happened and either scored a genuine zero or needs a note. means the configuration was never run. A marker in place of a number means it ran and produced nothing scorable. Heat bars scale to the maximum of their own table and are comparable inside it only.

Every number above comes from the harness × model result-matrix appendix, copied cell for cell with no rewriting. means the configuration was never run; a marker in place of a number means it ran and produced nothing scorable; a number with a superscript letter is a real score that needs a note. Scores from different environments are not comparable.

14AAV · enrichment / tropismaav-capsid-enrichmenttrack_total (per-assay Pearson)
harness · modelmouse · multi-organmacaque · cross-speciesin vitro · cell lines
Claude Code · Opus-4.8single-model0.6800.5770.756
Codex CLI · GPT-5.6-solsingle-model0.5830.5230.754
DeerFlow · Opus-4.80.5400.4850.731
DeerFlow · GPT-5.6-sol0.6020.5800.742
DeerFlow · DeepSeek-v4-pro0.5490.5010.000a
DeerFlow · Kimi-K30.6980.5610.741
DeerFlow · GLM-5.20.6330.3200.727
OpenHands · Opus-4.80.5750.6170.748
OpenHands · GPT-5.6-sol0.4090.6310.762
OpenHands · DeepSeek-v4-pro0.5020.5850.707
OpenHands · Kimi-K30.5150.4940.750
OpenHands · GLM-5.20.5170.5890.747
A-Evolve · Opus-4.8N/A bN/A bN/A b
A-Evolve · GPT-5.6-sol0.6280.4720.749
A-Evolve · DeepSeek-v4-pro0.3710.5940.649
A-Evolve · Kimi-K30.4680.5680.751
A-Evolve · GLM-5.20.5160.6330.612
F4F (Fit4Function) · paper reproductionpublished SOTA0.5320.6140.719
k-mer lookup floorfloor0.1820.0520.000
Reading: macaque is the discriminating axis: all 16 scored cells clear the k-mer lookup floor, 9/16 beat published SOTA on mouse and 12/16 in vitro (F4F 0.532 / 0.719), but on cross-species macaque only 3/16 (A-Evolve · GLM 0.633, OpenHands · GPT 0.631, OpenHands · Opus 0.617) beat F4F’s 0.614 — the specialised method still leads most agents on the hardest track. GPT’s first-pass timeouts on this host-bash world were recovered by a low-concurrency re-run, so that was submission discipline rather than a capability limit; the only unrecovered cell is DeerFlow · DeepSeek’s in vitro.
Grounding aGrounding failure: the agent stopped to ask the user instead of acting, and stronger prompting did not recover it — the only unrecovered cell in the sweep.
N/A bA-Evolve’s provider layer rejects non-OpenAI model names, so all three cells are unavailable by construction rather than failures.
15AAV · structure predictionaav-capsid-structureStructure accuracy 0–1 (best of ≤3)
harness · modelloop · VR monomerassembly · 60-mercomplex · interface
Claude Code · Opus-4.8single-model0.9290.6920.136
Codex CLI · GPT-5.6-solsingle-model0.9270.7740.379
DeerFlow · Opus-4.80.9270.000a0.163
DeerFlow · GPT-5.6-sol0.9270.6640.163
DeerFlow · DeepSeek-v4-pro0.9270.000b0.000c
DeerFlow · Kimi-K30.9270.000d0.000c
DeerFlow · GLM-5.20.928e0.000d0.163
OpenHands · Opus-4.80.9270.6680.120
OpenHands · GPT-5.6-sol0.9300.8190.342
OpenHands · DeepSeek-v4-pro0.9270.000b0.163
OpenHands · Kimi-K30.9250.000d0.017
OpenHands · GLM-5.20.9270.7090.163
A-Evolve · Opus-4.8N/A fN/A fN/A f
A-Evolve · GPT-5.6-sol0.9300.5910.128
A-Evolve · DeepSeek-v4-pro0.9270.7040.136
A-Evolve · Kimi-K30.9270.000d0.365
A-Evolve · GLM-5.20.9270.6340.163
AlphaFold 3 + symmetry expansionpublished SOTA0.9270.7400.148
Boltz-1 + symmetry expansionref0.8670.6950.074
AlphaFold 2 + symmetry expansionref0.9010.7320.073
CapBuildref0.593
Template dockingref0.252
nearest-template baseline (loop pass threshold)threshold0.709
Reading: loop is saturated: 11 of the 16 scored cells sit at exactly 0.927 — the same number as published SOTA / AlphaFold 3 — with 4 more at 0.928–0.930 and one at 0.925, and the submitted structures are byte-identical across harnesses (the default fold path is deterministic and cached), so this track measures nothing (DeerFlow · GLM’s 0.928 even fails the correctness gate: only 5 of 12 targets folded, coverage 0.417). Six complex cells land on exactly 0.163, the floor for submitting the default fold unchanged. assembly is the only strongly discriminating axis this round: 7 of 16 scored cells are 0 (4 built none of the 60-mer, 2 placed all 60 chains but with whole-capsid RMSD 100–155 Å, 1 had no submission to grade), and of the 9 that built something only 2 beat published SOTA (AlphaFold 3 with symmetry expansion, 0.740) — OpenHands · GPT 0.819 and Codex 0.774 — though 8 of the 9 clear CapBuild’s 0.593. On complex only 3 cells (Codex 0.379, A-Evolve · Kimi 0.365, OpenHands · GPT 0.342) beat published SOTA (template docking 0.252).
Not submitted aThe scorer found no submission to grade.
Gate not met bAll 60 chains were placed, but whole-capsid RMSD is 100–155 Å, so the RMSD score is zero.
Note cAll 15 complexes were graded but DockQ and interface Fnat are both zero — verified to be a real failure, not a silently swallowed export error.
Gate not met dNone of the 60-mer was built.
Gate not met eThe score itself is fine, but only 5 of 12 targets were folded (coverage 0.417), so the correctness gate fails.
N/A fA-Evolve’s provider layer rejects non-OpenAI model names.
16AAV · generative designaav-sequence-designPass rate (mean of three bars + viability gate)
harness · modelbrainheartspinalcordly6a · receptorly6c1 · receptor
Claude Code · Opus-4.8single-modelRefused dRefused dRefused dRefused dRefused d
Codex CLI · GPT-5.6-solsingle-model0.1190.2900.4150.0700.135
DeerFlow · Opus-4.8Refused dRefused dRefused dRefused dRefused d
DeerFlow · GPT-5.6-sol0.059a0.076a0.2210.0620.000b
DeerFlow · DeepSeek-v4-pro0.0250.0510.1380.0100.013
DeerFlow · Kimi-K30.0400.1020.2410.0210.015
DeerFlow · GLM-5.20.090a0.120a0.2200.0880.056
OpenHands · Opus-4.8Out of sweep cOut of sweep cOut of sweep cOut of sweep cOut of sweep c
OpenHands · GPT-5.6-sol0.1850.1530.3920.1320.175
OpenHands · DeepSeek-v4-pro0.0750.1500.2690.0800.021
OpenHands · Kimi-K30.1000.0950.2140.0760.038
OpenHands · GLM-5.20.1090.1930.3570.0660.001
A-Evolve · Opus-4.8N/A eN/A eN/A eN/A eN/A e
A-Evolve · GPT-5.6-sol0.1220.1740.3490.1410.088
A-Evolve · DeepSeek-v4-pro0.0600.0910.1740.0250.073
A-Evolve · Kimi-K30.0470.0750.1480.0660.051
A-Evolve · GLM-5.20.1360.1510.0910.0870.059
AAVDiffpublished SOTA0.1130.1770.2670.0210.003
AAVGenref0.1040.1600.2470.0260.009
ALICEref0.0930.1680.2470.0240.019
Reading: The metric moved from a single top-2% bar to the mean of three bars (5% / 2% / 1%), because the strictest bar is counting-noise dominated (out of a 1000-design budget the median is 24 qualifying designs at the 2% bar and 2 at the 1% bar). Under the new rule only 4 of 13 scored configurations (Codex · GPT, OpenHands · GPT, OpenHands · GLM, A-Evolve · GPT) beat published SOTA on the three-organ mean (= AAVDiff’s 0.113 / 0.177 / 0.267, mean 0.186) — agents do not hold the edge on tropism. Receptors are the opposite: 10/13 on ly6a and 9/13 on ly6c1 beat published SOTA (AAVGen 0.026 on ly6a, ALICE 0.019 on ly6c1), and the leaders (OpenHands · GPT ly6c1 0.175, A-Evolve · GPT ly6a 0.141) are more than 5× it. spinalcord is consistently the easiest organ (top: Codex · GPT 0.415). Refusals: Opus refuses this generative-design task at the product level on Claude Code and DeerFlow; OpenHands · Opus was not part of the three-bar sweep.
Not submitted aThis organ’s mean includes one split that scored 0.000 because nothing was submitted.
Note bIsolated incident; not reproduced on other targets under the same configuration.
Out of sweep cOpus was not part of the sweep these three-bar scores were computed from; the earlier Opus figures came from the superseded single top-2% bar and are not comparable.
Refused dClaude product-level safety refusal: Opus refuses this generative-design task on all targets under Claude Code and DeerFlow.
N/A eA-Evolve’s provider layer rejects non-OpenAI model names.
17AAV · viabilityaav-viabilitymean per-distance AUROC
harness · modelBryant · labelledBryant · oracleOgden · labelledOgden · oracle
Claude Code · Opus-4.8single-model0.9470.968Refused aRefused a
Codex CLI · GPT-5.6-solsingle-model0.9800.9550.9280.903
DeerFlow · Opus-4.80.9600.9520.9080.862
DeerFlow · GPT-5.6-sol0.9760.9480.9360.854
OpenHands · Opus-4.80.9820.9570.8670.851
OpenHands · GPT-5.6-sol0.9820.9430.9400.913
A-Evolve · Opus-4.8N/A cN/A cN/A cN/A c
A-Evolve · GPT-5.6-sol0.9860.9700.930Refused b
CAP-PLM protein language modelpublished SOTA0.9490.8970.8830.859
logistic additiveref0.8880.8880.8520.852
ALICE-GBref0.9400.8850.8150.696
augmented ridgeref0.9080.8740.8660.848
pass threshold (app-17 additive baseline)threshold0.8850.8850.8510.851
Reading: All 25 scored cells clear the environment’s pass threshold (Bryant 0.885 / Ogden 0.851; 33 cells by the appendix’s count). published SOTA is set per split — CAP-PLM 0.949 labelled / 0.897 oracle on Bryant and 0.884 labelled / 0.859 oracle on Ogden — and against it the margin is far thinner than the pass threshold suggests: on Bryant 13/14 cells beat CAP-PLM, the exception being Claude Code · Opus 0.947 on the labelled split; on Ogden · labelled 5/6 do, with OpenHands · Opus 0.867 below 0.884; and on the hardest setting, Ogden · oracle, only 3/5 (OpenHands · GPT 0.913, Codex 0.903, DeerFlow · Opus 0.862 — DeerFlow · GPT 0.854 and OpenHands · Opus 0.851 fall short of 0.859). CAP-PLM is the binding reference in every column: all 14 Bryant and all 11 Ogden cells still clear ALICE-GB (0.940 / 0.885 Bryant, 0.815 / 0.696 Ogden). Hidden oracle × single-mutation is the one setting that is genuinely not saturated. The difficulty gradient Ogden > Bryant, oracle > labelled reads straight off the scores (Ogden · oracle 0.851–0.913 vs Bryant · labelled up to 0.986). The 11 first-pass zeros were all concurrency artefacts and recovered on a low-concurrency re-run — zero capability failures; effort is uncorrelated with score (10 to 1970 actions). Both refusals are deterministic provider-side refusals of the Ogden capsid-engineering framing (Anthropic usage policy / OpenAI invalid_prompt), harness-specific and not capability results.
Refused aAnthropic usage-policy refusal (one action, never submitted): a deterministic provider-side refusal of the Ogden capsid-engineering framing, harness-specific and not a capability result.
Refused bOpenAI invalid_prompt, reproduced at three concurrency levels.
N/A cA-Evolve’s provider layer rejects non-OpenAI model names.