The TRACES Leaderboard
Six capability scores, 0-4, higher is better. One table per capability, ranked.
Scoring
How to read
Each capability is scored 0–4 by a process scorer; the higher the better. The band a score falls in says how the system under test (SuT) did.
0–1
1–2
2–3
3–4
Model Arena
The Model Arena
A live leaderboard for foundation models, under the shared pi-harness to compare their performances on our TRACES capability dimensions.
Bars are drawn on the native 0 to 4 scale and the gridlines mark the four anchors.
Harness Arena
The Harness Arena
A leaderboard for agent harnesses, under the shared model apodex-1.1 (Apodex’s self-owned foundation model) to compare harness performances on our TRACES capability dimensions.
Capabilities
The six TRACES capabilities
T
Tools
Selecting, calling and correctly interpreting external tools
R
Repair
Locating and correcting its own errors once feedback arrives
A
Alternatives
Laying out competing hypotheses and keeping or discarding them as evidence accumulates
C
Coherence
Holding state, constraints and logic intact across a long chain of work
E
Evidence
Grounding every conclusion in observation, data, experiment or citation
S
Scope
Stating the conditions under which a conclusion holds, and where it does not apply
Detailed scores
Detailed scores
All scores range from 0 to 4.
| # | Model | Harness | Tools | Repair | Alternatives | Coherence | Evidence | Scope | % policy refusal | % incomplete | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-opus-5 | pi | 2.80 | 2.69 | 2.72 | 3.08 | 2.48 | 2.65 | 12.8 | 10.6 | 2.73 |
| 2 | glm-5.2 | pi | 2.75 | 2.71 | 2.76 | 3.11 | 2.49 | 2.53 | 0.0 | 12.6 | 2.70 |
| 3 | kimi-k3 | pi | 2.81 | 2.58 | 2.59 | 2.90 | 2.50 | 2.43 | 0.0 | 17.2 | 2.64 |
| 4 | apodex-1.1 | openhands | 2.56 | 2.64 | 2.53 | 2.93 | 2.54 | 2.48 | 0.0 | 15.7 | 2.59 |
| 5 | apodex-1.1 | pi | 2.62 | 2.51 | 2.55 | 2.95 | 2.44 | 2.44 | 0.0 | 16.8 | 2.59 |
| 6 | apodex-1.1 | freephdlabor | 2.59 | 2.71 | 2.52 | 2.77 | 2.52 | 2.34 | 0.0 | 12.1 | 2.54 |
| 7 | deepseek-v4-flash | pi | 2.56 | 2.38 | 2.45 | 2.70 | 2.34 | 2.25 | 0.0 | 21.4 | 2.45 |
| 8 | apodex-1.1 | aevolve | 2.35 | 2.04 | 2.13 | 2.67 | 2.40 | 2.18 | 0.0 | 23.1 | 2.32 |
| 9 | gpt-5.5 | pi | 2.74 | 2.32 | 1.92 | 2.57 | 2.29 | 1.98 | 0.0 | 14.3 | 2.31 |
| 10 | gpt-5.6-sol | pi | 2.44 | 1.79 | 2.05 | 2.68 | 2.22 | 2.13 | 0.8 | 16.1 | 2.28 |
| 11 | apodex-1.1 | deerflow | 2.20 | 2.36 | 2.23 | 2.57 | 2.15 | 2.19 | 0.0 | 30.0 | 2.25 |
Methods and caveats
A capability is scored inside each environment first and the environments are then combined as a weighted average. An environment with many instances would be weighted slightly higher, while an environment with relatively few instances would be weighted lower to avoid too much variation. The weights are normalised so that the biomedical and the non-biomedical environments carry equal total weight, keeping the two domains balanced.
A model that declines a task on policy grounds is scored based on the average score of other scorable SuTs on that particular task. Other than policy refusals, incomplete trajectories are penalised separately. A trajectory is incomplete if it did not finish within the prescribed time limit, never submitted a final solution, or failed the anti-hacking checks; each SuT’s score is multiplied by the square root of its share of complete trajectories, while policy refused or declined tasks count as complete since they are already scored above.