Live

The TRACES Leaderboard

Six capability scores, 0-4, higher is better. One table per capability, ranked.

obstacletargetdetourblocked

Scoring

How to read

Each capability is scored 0–4 by a process scorer; the higher the better. The band a score falls in says how the system under test (SuT) did.

0–1
Flawed
1–2
Insufficient
2–3
Good
3–4
Excellent

Model Arena

The Model Arena

A live leaderboard for foundation models, under the shared pi-harness to compare their performances on our TRACES capability dimensions.

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.81
2.80
2.75
2.74
2.56
2.44
kimi-k3
opus-5
glm-5.2
gpt-5.5
deepseek-v4-flash
gpt-5.6-sol

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.69
2.58
2.38
2.32
1.79
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.76
2.72
2.59
2.45
2.05
1.92
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
3.11
3.08
2.90
2.70
2.68
2.57
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.50
2.49
2.48
2.34
2.29
2.22
kimi-k3
glm-5.2
opus-5
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.65
2.53
2.43
2.25
2.13
1.98
opus-5
glm-5.2
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.81
2.80
2.75
2.74
2.56
2.44
kimi-k3
opus-5
glm-5.2
gpt-5.5
deepseek-v4-flash
gpt-5.6-sol

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.69
2.58
2.38
2.32
1.79
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.76
2.72
2.59
2.45
2.05
1.92
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
3.11
3.08
2.90
2.70
2.68
2.57
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.50
2.49
2.48
2.34
2.29
2.22
kimi-k3
glm-5.2
opus-5
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.65
2.53
2.43
2.25
2.13
1.98
opus-5
glm-5.2
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.81
2.80
2.75
2.74
2.56
2.44
kimi-k3
opus-5
glm-5.2
gpt-5.5
deepseek-v4-flash
gpt-5.6-sol

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.69
2.58
2.38
2.32
1.79
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.76
2.72
2.59
2.45
2.05
1.92
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
3.11
3.08
2.90
2.70
2.68
2.57
glm-5.2
opus-5
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.50
2.49
2.48
2.34
2.29
2.22
kimi-k3
glm-5.2
opus-5
deepseek-v4-flash
gpt-5.5
gpt-5.6-sol

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.65
2.53
2.43
2.25
2.13
1.98
opus-5
glm-5.2
kimi-k3
deepseek-v4-flash
gpt-5.6-sol
gpt-5.5

Bars are drawn on the native 0 to 4 scale and the gridlines mark the four anchors.

Harness Arena

The Harness Arena

A leaderboard for agent harnesses, under the shared model apodex-1.1 (Apodexs self-owned foundation model) to compare harness performances on our TRACES capability dimensions.

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.62
2.59
2.56
2.35
2.20
P
pi
F
freephdlabor
O
openhands
A
aevolve
D
deerflow

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.64
2.51
2.36
2.04
F
freephdlabor
O
openhands
P
pi
D
deerflow
A
aevolve

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.55
2.53
2.52
2.23
2.13
P
pi
O
openhands
F
freephdlabor
D
deerflow
A
aevolve

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.95
2.93
2.77
2.67
2.57
P
pi
O
openhands
F
freephdlabor
A
aevolve
D
deerflow

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.54
2.52
2.44
2.40
2.15
O
openhands
F
freephdlabor
P
pi
A
aevolve
D
deerflow

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.48
2.44
2.34
2.19
2.18
O
openhands
P
pi
F
freephdlabor
D
deerflow
A
aevolve

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.62
2.59
2.56
2.35
2.20
P
pi
F
freephdlabor
O
openhands
A
aevolve
D
deerflow

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.64
2.51
2.36
2.04
F
freephdlabor
O
openhands
P
pi
D
deerflow
A
aevolve

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.55
2.53
2.52
2.23
2.13
P
pi
O
openhands
F
freephdlabor
D
deerflow
A
aevolve

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.95
2.93
2.77
2.67
2.57
P
pi
O
openhands
F
freephdlabor
A
aevolve
D
deerflow

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.54
2.52
2.44
2.40
2.15
O
openhands
F
freephdlabor
P
pi
A
aevolve
D
deerflow

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.48
2.44
2.34
2.19
2.18
O
openhands
P
pi
F
freephdlabor
D
deerflow
A
aevolve

Tools

Selecting, calling and correctly interpreting external tools

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.62
2.59
2.56
2.35
2.20
P
pi
F
freephdlabor
O
openhands
A
aevolve
D
deerflow

Repair

Locating and correcting its own errors once feedback arrives

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.71
2.64
2.51
2.36
2.04
F
freephdlabor
O
openhands
P
pi
D
deerflow
A
aevolve

Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.55
2.53
2.52
2.23
2.13
P
pi
O
openhands
F
freephdlabor
D
deerflow
A
aevolve

Coherence

Holding state, constraints and logic intact across a long chain of work

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.95
2.93
2.77
2.67
2.57
P
pi
O
openhands
F
freephdlabor
A
aevolve
D
deerflow

Evidence

Grounding every conclusion in observation, data, experiment or citation

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.54
2.52
2.44
2.40
2.15
O
openhands
F
freephdlabor
P
pi
A
aevolve
D
deerflow

Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Flawed
0-1 means the system under test (SuT) is flawed
Insufficient
1-2 means the SuT is insufficient
Good
2-3 means the SuT is good
Excellent
3-4 means the SuT is excellent
2.48
2.44
2.34
2.19
2.18
O
openhands
P
pi
F
freephdlabor
D
deerflow
A
aevolve

Capabilities

The six TRACES capabilities

T
Tools

Selecting, calling and correctly interpreting external tools

R
Repair

Locating and correcting its own errors once feedback arrives

A
Alternatives

Laying out competing hypotheses and keeping or discarding them as evidence accumulates

C
Coherence

Holding state, constraints and logic intact across a long chain of work

E
Evidence

Grounding every conclusion in observation, data, experiment or citation

S
Scope

Stating the conditions under which a conclusion holds, and where it does not apply

Detailed scores

Detailed scores

All scores range from 0 to 4.

#ModelHarnessToolsRepairAlternativesCoherenceEvidenceScope% policy refusal% incompleteOverall
1claude-opus-5pi2.802.692.723.082.482.6512.810.62.73
2glm-5.2pi2.752.712.763.112.492.530.012.62.70
3kimi-k3pi2.812.582.592.902.502.430.017.22.64
4apodex-1.1openhands2.562.642.532.932.542.480.015.72.59
5apodex-1.1pi2.622.512.552.952.442.440.016.82.59
6apodex-1.1freephdlabor2.592.712.522.772.522.340.012.12.54
7deepseek-v4-flashpi2.562.382.452.702.342.250.021.42.45
8apodex-1.1aevolve2.352.042.132.672.402.180.023.12.32
9gpt-5.5pi2.742.321.922.572.291.980.014.32.31
10gpt-5.6-solpi2.441.792.052.682.222.130.816.12.28
11apodex-1.1deerflow2.202.362.232.572.152.190.030.02.25

Methods and caveats

A capability is scored inside each environment first and the environments are then combined as a weighted average. An environment with many instances would be weighted slightly higher, while an environment with relatively few instances would be weighted lower to avoid too much variation. The weights are normalised so that the biomedical and the non-biomedical environments carry equal total weight, keeping the two domains balanced.

A model that declines a task on policy grounds is scored based on the average score of other scorable SuTs on that particular task. Other than policy refusals, incomplete trajectories are penalised separately. A trajectory is incomplete if it did not finish within the prescribed time limit, never submitted a final solution, or failed the anti-hacking checks; each SuT’s score is multiplied by the square root of its share of complete trajectories, while policy refused or declined tasks count as complete since they are already scored above.